Both use the same underlying neural networks. The difference is how they interact with long inputs. One brute-forces it. The other thinks programmatically.
| Dimension | LLM | RLM |
|---|---|---|
| Max effective input | 4K–272K tokens (degrades at scale) | 10M+ tokens (tested) |
| Processing model | Single forward pass — all tokens attend to all tokens | Recursive self-calls — decompose, process chunks, aggregate |
| How input is accessed | Loaded entirely into context window | Stored as environment variable, examined programmatically |
| Scaling behavior | O(n²) attention — quality drops as input grows | Recursive decomposition — quality maintained at any scale |
| Selective attention | No — must attend to everything | Yes — model decides what to examine |
| Code execution | Not part of inference | Central — model writes code in REPL to slice and process |
| Cost at scale | Linear or worse — paying for all tokens | Often cheaper — only processes relevant chunks |
| Failure mode | "Context rot" — gradually loses information | Can miss connections across chunks (mitigated by overlap strategies) |
| Best suited for | Short-to-medium inputs, conversational tasks | Book-length docs, codebases, legal corpora, deep research |
| Training required | Standard pretraining + fine-tuning | None to start (wraps any LLM); optional post-training, e.g. ~1,000 samples for RLM-Qwen3-8B |
An RLM is not a different model architecture. It's the same transformer — the same weights, the same attention mechanism — wrapped in a recursive execution framework.
Think of it this way: an LLM is a person trying to read an entire library by cramming all the books into their field of vision at once. An RLM is the same person, but now they have a desk, a notepad, and a system. They pick up one book at a time, take notes, cross-reference, and build understanding incrementally.
The key components that make this work:
1. REPL Environment — The input becomes a variable in a code sandbox. The model doesn't "see" the full text. It writes Python to examine it.
2. Recursive Self-Calls — The model can call itself on sub-problems. Process chunk 1, get a partial answer, process chunk 2 with that context, repeat.
3. Programmatic Decomposition — The model decides how to split the input. It's not fixed chunking — it's task-aware. A summarization task splits differently than a search task.
The MIT OASYS lab compared vanilla GPT-5 against RLM(GPT-5), an RLM with GPT-5 as the root model and GPT-5-mini sub-calls at recursion depth 1. Scores below are from Table 1 of the May 2026 revision of the paper (arXiv:2512.24601v3):
| Benchmark | GPT-5 | RLM(GPT-5) | Delta |
|---|---|---|---|
| S-NIAH (retrieval) | High | Comparable | ≈0% |
| OOLONG (131K tokens) | 44.0 | 56.0 | +12.0 pts |
| OOLONG-Pairs (quadratic, 32K tokens) | 0.1 F1 | 58.0 F1 | +580x |
| BrowseComp-Plus (1K docs, 6M-11M tokens) | 0.0 (exceeds context) | 91.3 | +91.3 pts |
| CodeQA (23K-4.2M tokens) | 24.0 | 62.0 | +38.0 pts |
At small scale, the lab also post-trained RLM-Qwen3-8B on 1,000 samples. It beats the base Qwen3-8B by 28.3% on average and approaches vanilla GPT-5 on three of the four tasks.
The OOLONG-Pairs result is the most striking. This benchmark requires comparing information across the entire input — the kind of task where attention mechanisms fundamentally struggle at scale. GPT-5 essentially fails. The RLM version handles it because it doesn't try to attend to everything at once.
On cost: at the median, RLM calls on GPT-5 are cheaper than vanilla GPT-5, because the model selectively examines context rather than paying for attention over all tokens.
Keep the recursion shallow. An independent reproduction published in March 2026 rebuilt the RLM framework on DeepSeek v3.2 and Kimi K2 and measured results by recursion depth. Depth-1 recursion improved complex reasoning tasks, but depth-2 recursion made results worse — clearly so on simple retrieval tasks — while runtime and token use climbed steeply, from 3.6 seconds to 344.5 seconds on one measured task. The author's summary: deeper recursion causes models to overthink (Wang, arXiv:2603.02615, 2026). Treat depth as a cost you justify per task, not a dial you turn up.
The authors' own May 2026 revision shows the depth effect depends on the model. With GPT-5, OOLONG-Pairs rose from 58.0 at depth 1 to 76.0 at depth 3. With Qwen3-Coder-480B, CodeQA fell from 56.0 at depth 1 to 44.0 at depth 3 (arXiv:2512.24601v3, Table 1). Test deeper recursion on your own model and task before you rely on it.
The two approaches aren't mutually exclusive. An RLM uses standard LLM calls internally — it just orchestrates them recursively. You can think of it as a meta-layer on top of any LLM.
DSPy v3.1.2+ ships with built-in RLM support. If you're already using DSPy for prompt programming, adding recursive processing is a configuration change.
Google's Agent Development Kit (ADK) has an enterprise-ready implementation with lazy file loading and parallel sub-calls — optimized for production workloads.
Prime Intellect's Prime Agent (open-sourced August 2026) is a coding harness built directly on the RLM abstraction. Context is held as variables in a persistent IPython kernel, and sub-agents are invoked as function calls from inside the REPL. Prime Intellect reports that, running Claude Opus 5, it scored 95.5% on ARC-AGI-3 (best-of-1), above the 95.4% human-expert baseline (Prime Intellect, August 2026).
The original MIT paper (arXiv:2512.24601) includes the full algorithm, training data, and post-training recipe. The authors now list it as a NeurIPS 2026 paper (repo README, September 2026).The model weights for RLM-Qwen3-8B are available on HuggingFace (mit-oasys/rlm-qwen3-8b-v0.1). The pip install rlms package (alexzhang13/rlm, Python 3.11+) provides the reference implementation; its import name is rlm.
In May 2026, the reference library (v0.1.2) added an RL training harness built on Prime Intellect's prime-rl. The lab used it to release RLM-Qwen3-30B-A3B v0.1, a LoRA adapter for Qwen3-30B-A3B-Instruct-2507. On the lab's own evals, it scores 45.0 on OOLONG-Pairs at 32K tokens, against 42.9 for the untrained base model. The same release changed how runs end: the root model now writes its result to answer["content"] in the REPL and sets answer["ready"] = True, in place of FINAL() or FINAL_VAR() (PR #162).
This isn't a research curiosity anymore. It's being deployed in production systems for document processing, code analysis, and deep research applications.
Ready to go deeper? Start with how RLMs work, or see the benchmark results.