Architecture at a glance

LLM (Traditional)
272K tokens entire input LLM single forward pass attention over ALL tokens ⚠ degrades Output quality on long inputs
RLM (Recursive)
10M+ tokens stored as env variable decompose LLM chunk 1 LLM chunk 2 LLM chunk N recurse aggregate REPL Environment merge partial results Output quality on long inputs

Head-to-head comparison

Dimension LLM RLM
Max effective input 4K–272K tokens (degrades at scale) 10M+ tokens (tested)
Processing model Single forward pass — all tokens attend to all tokens Recursive self-calls — decompose, process chunks, aggregate
How input is accessed Loaded entirely into context window Stored as environment variable, examined programmatically
Scaling behavior O(n²) attention — quality drops as input grows Recursive decomposition — quality maintained at any scale
Selective attention No — must attend to everything Yes — model decides what to examine
Code execution Not part of inference Central — model writes code in REPL to slice and process
Cost at scale Linear or worse — paying for all tokens Often cheaper — only processes relevant chunks
Failure mode "Context rot" — gradually loses information Can miss connections across chunks (mitigated by overlap strategies)
Best suited for Short-to-medium inputs, conversational tasks Book-length docs, codebases, legal corpora, deep research
Training required Standard pretraining + fine-tuning None to start (wraps any LLM); optional post-training, e.g. ~1,000 samples for RLM-Qwen3-8B
Core insight

Same model, different paradigm

An RLM is not a different model architecture. It's the same transformer — the same weights, the same attention mechanism — wrapped in a recursive execution framework.

Think of it this way: an LLM is a person trying to read an entire library by cramming all the books into their field of vision at once. An RLM is the same person, but now they have a desk, a notepad, and a system. They pick up one book at a time, take notes, cross-reference, and build understanding incrementally.

The key components that make this work:

1. REPL Environment — The input becomes a variable in a code sandbox. The model doesn't "see" the full text. It writes Python to examine it.

2. Recursive Self-Calls — The model can call itself on sub-problems. Process chunk 1, get a partial answer, process chunk 2 with that context, repeat.

3. Programmatic Decomposition — The model decides how to split the input. It's not fixed chunking — it's task-aware. A summarization task splits differently than a search task.

The numbers

Performance where it matters

The MIT OASYS lab compared vanilla GPT-5 against RLM(GPT-5), an RLM with GPT-5 as the root model and GPT-5-mini sub-calls at recursion depth 1. Scores below are from Table 1 of the May 2026 revision of the paper (arXiv:2512.24601v3):

BenchmarkGPT-5RLM(GPT-5)Delta
S-NIAH (retrieval)HighComparable≈0%
OOLONG (131K tokens)44.056.0+12.0 pts
OOLONG-Pairs (quadratic, 32K tokens)0.1 F158.0 F1+580x
BrowseComp-Plus (1K docs, 6M-11M tokens)0.0 (exceeds context)91.3+91.3 pts
CodeQA (23K-4.2M tokens)24.062.0+38.0 pts

At small scale, the lab also post-trained RLM-Qwen3-8B on 1,000 samples. It beats the base Qwen3-8B by 28.3% on average and approaches vanilla GPT-5 on three of the four tasks.

The OOLONG-Pairs result is the most striking. This benchmark requires comparing information across the entire input — the kind of task where attention mechanisms fundamentally struggle at scale. GPT-5 essentially fails. The RLM version handles it because it doesn't try to attend to everything at once.

On cost: at the median, RLM calls on GPT-5 are cheaper than vanilla GPT-5, because the model selectively examines context rather than paying for attention over all tokens.

Practical guidance

When to use which

Use a standard LLM when:

  • Input fits comfortably in context (<50K tokens)
  • Task is conversational or generative (not analytical)
  • Latency matters more than thoroughness
  • You need real-time streaming responses

Use an RLM when:

  • Input exceeds the model's effective context window
  • Task requires dense reasoning over the entire input
  • You need to cross-reference information across documents
  • Accuracy matters more than speed
  • Processing codebases, legal documents, research corpora, or book-length content

Keep the recursion shallow. An independent reproduction published in March 2026 rebuilt the RLM framework on DeepSeek v3.2 and Kimi K2 and measured results by recursion depth. Depth-1 recursion improved complex reasoning tasks, but depth-2 recursion made results worse — clearly so on simple retrieval tasks — while runtime and token use climbed steeply, from 3.6 seconds to 344.5 seconds on one measured task. The author's summary: deeper recursion causes models to overthink (Wang, arXiv:2603.02615, 2026). Treat depth as a cost you justify per task, not a dial you turn up.

The authors' own May 2026 revision shows the depth effect depends on the model. With GPT-5, OOLONG-Pairs rose from 58.0 at depth 1 to 76.0 at depth 3. With Qwen3-Coder-480B, CodeQA fell from 56.0 at depth 1 to 44.0 at depth 3 (arXiv:2512.24601v3, Table 1). Test deeper recursion on your own model and task before you rely on it.

The two approaches aren't mutually exclusive. An RLM uses standard LLM calls internally — it just orchestrates them recursively. You can think of it as a meta-layer on top of any LLM.

Adoption

The ecosystem is moving

DSPy v3.1.2+ ships with built-in RLM support. If you're already using DSPy for prompt programming, adding recursive processing is a configuration change.

Google's Agent Development Kit (ADK) has an enterprise-ready implementation with lazy file loading and parallel sub-calls — optimized for production workloads.

Prime Intellect's Prime Agent (open-sourced August 2026) is a coding harness built directly on the RLM abstraction. Context is held as variables in a persistent IPython kernel, and sub-agents are invoked as function calls from inside the REPL. Prime Intellect reports that, running Claude Opus 5, it scored 95.5% on ARC-AGI-3 (best-of-1), above the 95.4% human-expert baseline (Prime Intellect, August 2026).

The original MIT paper (arXiv:2512.24601) includes the full algorithm, training data, and post-training recipe. The authors now list it as a NeurIPS 2026 paper (repo README, September 2026).The model weights for RLM-Qwen3-8B are available on HuggingFace (mit-oasys/rlm-qwen3-8b-v0.1). The pip install rlms package (alexzhang13/rlm, Python 3.11+) provides the reference implementation; its import name is rlm.

In May 2026, the reference library (v0.1.2) added an RL training harness built on Prime Intellect's prime-rl. The lab used it to release RLM-Qwen3-30B-A3B v0.1, a LoRA adapter for Qwen3-30B-A3B-Instruct-2507. On the lab's own evals, it scores 45.0 on OOLONG-Pairs at 32K tokens, against 42.9 for the untrained base model. The same release changed how runs end: the root model now writes its result to answer["content"] in the REPL and sets answer["ready"] = True, in place of FINAL() or FINAL_VAR() (PR #162).

This isn't a research curiosity anymore. It's being deployed in production systems for document processing, code analysis, and deep research applications.


Ready to go deeper? Start with how RLMs work, or see the benchmark results.