Back to Blog
September 8, 2026

The short version: an RLM does not read the corpus. It loads the corpus into a Python REPL as a variable. The root model then writes code to search, filter, and slice that variable. Each promising slice goes to a sub-model call, and each hop of the question becomes one more round of code, sub-calls, and stored variables. On BrowseComp-Plus with 1,000 documents (6M to 11M tokens per query), RLM(GPT-5) scored 91.3% at recursion depth 1. Base GPT-5 scored 0.0%, because the input did not fit its 272K-token window.

Where does the corpus live during the run?

In the environment, not in the model. Algorithm 1 of the RLM paper starts by initializing a REPL with the prompt stored as a variable. The root model never sees the full string. It sees only constant-size metadata: the length, a short prefix, and instructions on how to access parts of it. Printed output is truncated the same way. The authors call this the key design choice, because it forces the model to rely on variables and sub-calls instead of polluting its own window.

The root model gets two handles. A context variable holds the prompt. An llm_query function sends a string to a sub-model that can handle around 500K characters. The system prompt tells the root model to batch about 200K characters into each call rather than making thousands of small ones. The run ends when the model emits FINAL(answer) or FINAL_VAR(variable_name), which returns a variable it built in the REPL. In the main runs, GPT-5 is the root model and GPT-5-mini answers the sub-calls.

For BrowseComp-Plus, the variable holds 1,000 documents. In the worked example the authors publish (Query 74), that is 8.3M tokens. The model has a symbolic handle to all of it. This is the same setup our decompose-recurse-aggregate post describes, applied to a corpus instead of a single long document.

How does the root model find candidate documents?

With code, not with embeddings. The paper reports that RLMs frequently filter and interact with their context through regex. On Query 74, GPT-5 first probes the 1,000-document list with regex queries. It uses two sources for its patterns: its own priors about the event in the question, and literal keywords from the prompt such as "beauty pageant" and "festival". The paper's summary: on BrowseComp-Plus, the model "uses model priors to programmatically narrow the search space of sub-calls."

This step is a loop, not a lookup. A filter runs across the whole corpus, the model inspects the count and a truncated sample, and then it adjusts the pattern. Nothing is dropped by a top-k cutoff, because every pass scans the full variable. A scan costs CPU time in the REPL, not model tokens, so the root model can afford several passes before it spends a sub-call.

How does one hop chain into the next?

Through variables. In Query 74, the regex pass surfaces an interesting snippet in the chunk at index 6. The root model launches a recursive LM call over that snippet and asks it for information relevant to the original query. The sub-call stores its result in a variable named answer6 and prints it for the root model to read. The sub-call reports that the answer is likely "Maria Dalmacio". The root model then decides it has enough, but it spends two more recursive calls to confirm the answer before returning it with FINAL. The answer is correct. The trajectory cost $0.079.

The general pattern is the same at every hop. Extract a fact from a candidate through a sub-call. Store the fact in a variable. Use the fact to write the next filter or the next sub-call prompt. The REPL is the working memory. The root context holds only code and truncated outputs, so a five-hop chain does not cost five documents of context.

At depth 2 and above, the root model also gets rlm_query(context, query). That function spawns a full RLM loop with its own REPL over a sub-context. The prompt reserves it for sub-tasks that need their own chunking, and it falls back to llm_query at the depth limit. Figure 10 of the paper shows GPT-5 uses significantly more sub-calls on BrowseComp-Plus than on any other task.

What do the BrowseComp-Plus numbers say?

The benchmark comes from Chen et al.. The full corpus has 100,195 documents and 830 queries. Each query has 6.1 evidence documents and 2.9 gold documents on average, and each document averages 5,179.2 words. The RLM paper samples 150 queries and gives each one 1,000 randomly chosen documents that include the gold and evidence documents.

The GPT-5 rows of Table 1 (accuracy, then average API cost per query):

  • Base GPT-5: 0.0%. The input exceeded the context window.
  • CodeAct + BM25 retriever: 51.0%, $0.71.
  • CodeAct + sub-calls (context loaded into the model): 0.0%.
  • Compaction agent: 70.5%, $0.57.
  • Claude Code with context offloading (Claude Opus 4.1): 84.0%, $2.03.
  • OpenCode with context offloading: 94.0%, cost not reported.
  • RLM depth 0 (REPL, no sub-calls): 88.0%, $0.44.
  • RLM depth 1: 91.3%, $0.99.
  • RLM depth 2: 92.0%, $0.55.
  • RLM depth 3: 92.0%, $0.51.

Two things stand out. First, the depth-0 ablation already scores 88.0%. Most of the gain on this task comes from the REPL handle on the corpus, and recursion adds a few points on top. The paper says this directly in Observation 2. Second, OpenCode with the corpus offloaded to a file beat the RLM by about two points. A coding agent that treats the corpus as a filesystem uses the same core idea.

The open model does worse. With Qwen3-Coder-480B-A35B as the root, RLM depth 1 scored 44.7%, compaction 38.0%, and CodeAct + BM25 12.7%. The paper attributes the gap to syntax errors in Qwen3-Coder's trajectories. Worse search code finds fewer candidates.

Appendix D.2 adds a scaling run on 20 queries as the document count grows toward 1,000. RLM(GPT-5) is the only method that keeps perfect performance at 1,000 documents on that subset, and only the iterative methods, RLM and ReAct, hold up past 100 documents. A linear extrapolation of GPT-5-mini reading 6M to 11M tokens directly comes to $1.50 to $2.75 per query, against the RLM's $0.99.

How does this compare with retrieval and long-context baselines?

Briefly, because our RLM vs RAG post covers the architecture argument. Inside the RLM paper, the retrieval agent (CodeAct + BM25) scored 51.0% against the RLM's 91.3%. In the BrowseComp-Plus paper itself, over the full 100,195-document corpus with a search tool that returns the top 5 hits truncated to 512 tokens each, GPT-5 scored 55.9% with BM25 and 70.1% with the Qwen3-Embedding-8B retriever. That is a different setup (full corpus, 830 queries), so it is not a head-to-head. It does show that retriever quality caps the agent: swapping the retriever moved GPT-5 by 14 points. The RLM avoids that cap by running its own filters over full documents.

There is no long-context baseline at this scale. GPT-5's window is 272K tokens and the inputs are 6M to 11M. The compaction agent is the stand-in, and it lands 21 points below the RLM.

What does the reproduction add?

Wang's reproduction did not run BrowseComp-Plus. It ran S-NIAH and OOLONG with DeepSeek v3.2 and Kimi K2, 20 samples per condition, one run each. The findings still matter for hop chains, because they show what happens when sub-calls get their own REPL.

On OOLONG, DeepSeek v3.2 went from 0.0% as a base model to 42.1% at depth 1, then fell to 33.7% at depth 2. Kimi K2 scored 86.6% natively, 60.0% at depth 1, and 55.0% at depth 2. On S-NIAH, both base models scored 100.0%; DeepSeek fell to 85.0% at depth 1 and 70.0% at depth 2. Wall-clock time for DeepSeek on S-NIAH went from 3.6 seconds to 89.3 seconds to 344.5 seconds.

The failure modes are the useful part. At depth 2, DeepSeek sometimes abandoned the corpus and answered from its weights: asked for fictional "magic numbers" in the text, it returned the nuclear shell numbers 2, 8, 20, 28, 50, 82, 126. For a multi-hop chain, that is the dangerous failure. A sub-call that answers from priors breaks the chain silently, and the root model stores the bad fact and builds the next hop on it. Keep sub-calls at depth 1 unless a hop needs its own chunking. Our recursion depth post covers the trade-off.

The bottom line

An RLM answers a multi-hop question over thousands of documents by keeping the corpus in a REPL variable, filtering it with regex and keyword code, sending only the surviving slices to sub-calls, and storing each hop's fact in a variable that shapes the next filter. On BrowseComp-Plus at 1,000 documents, that workflow took GPT-5 from 0.0% to 91.3% at $0.99 per query, and the REPL handle alone accounted for 88.0 of those points. Depth 1 is enough. Deeper recursion adds under a point here and, in the reproduction, invites sub-calls that answer from priors. The remaining gap is latency, and that is an implementation problem.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025 (v3 May 2026). Algorithm 1, the REPL system prompt, Table 1 BrowseComp-Plus scores and costs, Appendix D.2 document scaling, and the Query 74 trajectory in Appendix E.1. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. DeepSeek v3.2 and Kimi K2 accuracy, wall-clock time, and failure modes on S-NIAH and OOLONG at depths 0 to 2. arxiv.org/abs/2603.02615
  3. Chen, Z., Ma, X., Zhuang, S., et al. "BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent." arXiv:2508.06600, August 2025. Corpus size, evidence and gold document counts, retriever setup, and GPT-5 accuracy with BM25 and Qwen3-Embedding-8B. arxiv.org/abs/2508.06600
  4. Zhang, A. L. "Recursive Language Models." Project blog post, 2025. Peek, grep, and partition-plus-map patterns, and the 1,000-document BrowseComp-Plus claim. alexzhang13.github.io/blog/2025/rlm
FAQ

Frequently asked questions

Does the RLM read all 1,000 documents?

No. The root model never sees the full corpus. It sees the length, a short prefix, and instructions for accessing the variable. It runs code over the whole corpus to filter it, then sends only the surviving slices to sub-model calls. Scanning is CPU time in the REPL, so repeated passes cost no model tokens.

Which model answered the sub-calls in the BrowseComp-Plus runs?

GPT-5-mini. The RLM paper uses GPT-5 as the root model and GPT-5-mini for recursive calls in its main GPT-5 experiments. The authors chose this split to balance capability against the cost of many sub-calls. The average cost at depth 1 was $0.99 per query.

Why did base GPT-5 score 0.0% on this benchmark?

The inputs were 6M to 11M tokens and GPT-5's context window is 272K tokens. The paper marks these runs as hitting input context limits. CodeAct with sub-calls and plain Claude Code scored 0.0% for the same reason, because they load the corpus into the model's context.

Is recursion depth 1 enough for multi-hop questions?

On BrowseComp-Plus, yes. Depth 1 scored 91.3% and depths 2 and 3 scored 92.0%, so deeper recursion added under one point. The reproduction found that depth 2 made DeepSeek v3.2 and Kimi K2 worse on OOLONG and S-NIAH and multiplied wall-clock time. Use depth 2 only when a single hop needs its own chunking strategy.

Did the reproduction test BrowseComp-Plus?

No. Wang's reproduction ran S-NIAH and OOLONG with DeepSeek v3.2 and Kimi K2, using 20 samples per condition in a single run. Its value for multi-hop work is the failure analysis, which shows sub-calls at depth 2 answering from model priors instead of the corpus.