Back to Blog
August 2026

The short version: an RLM root model needs three things. It must write correct Python in a REPL, because code is the interface to the long input. It must follow a long orchestration prompt without drifting, because the harness is instructions all the way down. And it must have output-token budget left over after any thinking, because a model that reasons itself out of tokens cannot emit the program that does the work. Context window size is the least important line on the spec sheet, and raw parameter count matters less than post-training: an 8B model tuned on 1,000 trajectories beat its own base by a median of 28.3%.

Each of those claims comes from the original evaluation by Zhang, Kraska, and Khattab, or from the reproduction work that followed it. Here is the evidence, requirement by requirement.

Why is coding ability the first requirement?

Because the root model never reads the long input directly. The input sits in a Python REPL as a variable, and the root model interacts with it the only way a REPL allows: by writing code to peek into it, slice it, search it, and dispatch recursive sub-calls over pieces of it. The root model's output is a program. If the program has a syntax error, that turn produced nothing.

The paper states the consequence flatly: models without sufficient coding capabilities struggle as RLMs. And its evaluation shows what "struggle" costs. RLM trajectories from Qwen3-Coder-480B-A35B contained significantly more syntax errors than GPT-5 trajectories -- even on trajectories that reached a correct answer. On OOLONG-Pairs at depth 1, RLM(GPT-5) scored 58.0% F1 while RLM(Qwen3-Coder) scored 23.1%. Same harness, same task, same recursion budget. The difference is almost entirely how reliably each model operates the REPL.

Note what kind of coding this is. It is not competitive programming. The root model writes short, mundane snippets: index into a string, run a regex, split on delimiters, loop over chunks, call a function. The bar is reliability on boring code, not brilliance on hard code. A model that produces one malformed cell in ten wastes a tenth of its turns, and when sub-calls inherit the same weakness, the errors compound instead of averaging out. That compounding is the first failure mode in our post on RLM limitations.

Does the root model need a big context window?

No, and this is the part most people get backwards. The whole point of the RLM construction is that the full input never enters the root model's context. The root context holds the system prompt, the code the model has written, and the REPL's observations -- previews, line counts, sub-call results. Everything else stays in the environment.

That is how the method processes inputs up to two orders of magnitude beyond model context limits, at the 10M+ token scale, while the root models themselves have ordinary windows -- GPT-5's is 272K tokens, and the post-trained Qwen3-8B root operates with a far smaller one. What the window has to fit is the orchestration loop, not the data. A long trajectory with many turns of code and observations needs room, so a few-thousand-token window would choke; but the difference between a 128K window and a 1M window is close to irrelevant for the root's job.

If anything, the causality runs the other way: the smaller the base model's window, the more the recursive strategy is worth. Our piece on why context windows are the wrong abstraction makes the general version of that argument.

Do thinking models make good RLM roots?

Only with a caveat that the paper is unusually specific about: thinking models without sufficient output tokens struggle as RLMs. The authors attribute some smaller-than-expected gaps directly to trajectories that ran out of output tokens.

The mechanism is simple. A reasoning model spends part of its per-turn output allowance on deliberation before it emits anything actionable. In a chat setting that is fine; the deliberation is the product. In an RLM turn, the deliverable is a code cell, and the code comes after the thinking. Cap the output at a budget the model's reasoning habit can eat, and the turn ends mid-thought with no program emitted -- a wasted turn that looks, from the harness's perspective, identical to a syntax error.

So a thinking model can be an excellent root -- the strongest results in the paper come from one -- but the output ceiling has to be provisioned for thinking plus code, not for code alone. If you benchmark several roots under a uniform output cap, you have quietly handicapped the ones that reason longest.

Can an open-source model be the root?

Yes, and the record is specific about which and how.

Qwen3-Coder-480B-A35B works at depth 1, with real gains over running the same model as a plain long-context call: 36.0% to 48.0% on OOLONG, 20.0% to 56.0% on CodeQA. But it trails GPT-5 badly on hard aggregation (the 23.1% versus 58.0% gap above), its trajectories carry more syntax errors, and deeper recursion amplifies both problems rather than fixing them -- see our post on how deep an RLM should recurse.

DeepSeek v3.2 and Kimi K2, tested in a follow-up reproduction study, held up as depth-1 roots on complex reasoning tasks. At depth 2 they overthought: accuracy degraded on simple retrieval and execution time inflated from 3.6 seconds to 344.5 seconds. Capable open-source agentic models clear the bar, but the depth knob is less forgiving for them than for a frontier coder.

Qwen3-8B is the interesting case, because out of the box it does not clear the bar -- and post-training gets it there cheaply. RLM-Qwen3-8B was trained on just 1,000 filtered trajectories generated by Qwen3-Coder-480B-A35B, and outperforms its own base model by a median of 28.3% across the evaluation tasks, approaching vanilla GPT-5 on three of the long-context benchmarks. The lesson: for small models, RLM competence is a trainable skill, not an emergent property of scale. Our post on small models and big contexts works through the economics of that result.

Root model Works as RLM root? Evidence
GPT-5 Yes, best documented 58.0% F1 on OOLONG-Pairs at depth 1; gains grow with depth
Qwen3-Coder-480B-A35B Depth 1 only Gains on OOLONG and CodeQA; more syntax errors; 23.1% F1 on OOLONG-Pairs
DeepSeek v3.2 / Kimi K2 Depth 1 only Gains on complex reasoning; overthinking at depth 2 (3.6s to 344.5s)
Qwen3-8B (vanilla) No Insufficient without tuning
RLM-Qwen3-8B (post-trained) Yes Median +28.3% over base; approaches GPT-5 on three tasks

Does the same setup transfer between base models?

No, and the authors say so themselves: using the exact same RLM system prompt across all models can be problematic. Their concrete example is telling -- they had to add a sentence to the system prompt for Qwen3-Coder specifically to stop it from firing off too many recursive sub-calls. One model needs a leash; another needs a nudge.

This is an instruction-following requirement hiding inside a portability warning. The RLM harness is a long prompt describing an unusual job: you are in a REPL, the input is a variable, here is how to call yourself, here is when to stop. Models differ in how faithfully they inhabit that role, and the differences do not show up until you run trajectories. Budget for per-model prompt tuning; a harness tuned for one root is, in practice, a harness for that root.

Do the root and the sub-calls need the same model?

No. The original evaluation ran GPT-5 as the root and dispatched recursive sub-calls to GPT-5-mini, explicitly for cost reasons. The asymmetry is principled, not just cheap: the root does the hard job -- decomposition, orchestration, deciding when the answer is settled -- while a typical sub-call reads a bounded slice and answers a narrow question. The capability bar for the leaf is genuinely lower than for the root.

This is also the shape of the post-training story: RLM-Qwen3-8B learned its root behavior from trajectories written by a much larger teacher. Strong root, cheap leaves, distill downward when you need the root itself to be small.

The checklist

Choosing a base model for an RLM, in order of importance: reliable code emission in a REPL (measure the syntax-error rate per cell, not a coding benchmark score); faithful instruction following on a long orchestration prompt, tuned per model; output-token headroom after thinking; and a context window big enough for the trajectory, which almost any modern window is. Parameter count buys you depth tolerance -- frontier coders profit from depth 2 and 3 where everyone else should stay at depth 1 -- and post-training on a thousand good trajectories buys a small model a seat at the table.

The surprise, given the technique's name, is that nothing on that list says "long context." The base model you need is a careful junior programmer with good working habits. The recursion supplies the rest.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025 (revised May 2026). arxiv.org/abs/2512.24601
  2. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615. Depth ablations with DeepSeek v3.2 and Kimi K2 on S-NIAH and OOLONG. arxiv.org/abs/2603.02615
  3. Qwen Team. "Qwen3 Technical Report." arXiv:2505.09388. Base models for Qwen3-8B and Qwen3-Coder-480B-A35B. arxiv.org/abs/2505.09388
  4. OOLONG and OOLONG-Pairs: semantic labeling and pairwise aggregation benchmarks used in the RLM evaluation.
  5. CodeQA: repository-scale code understanding, 23K to 4.2M tokens.
FAQ

Frequently asked questions

What is the most important capability for an RLM base model?

Coding ability. The root model of an RLM operates a Python REPL where the long input is stored as a variable, so its output is a program, not prose. The original paper states plainly that models without sufficient coding capabilities struggle as RLMs. In its evaluation, Qwen3-Coder trajectories contained significantly more syntax errors than GPT-5 trajectories, even on trajectories that ultimately reached a correct answer, and that gap tracked the accuracy gap: 58.0% versus 23.1% F1 at depth 1 on OOLONG-Pairs.

Does the RLM root model need a large context window?

No. The window only needs to hold the orchestration loop: the system prompt, the code the model writes, and REPL observations. The full input never enters the root context; it lives in the environment as a variable. That is how the method processes inputs two orders of magnitude beyond model context limits, at the 10M+ token scale, and how an 8B model with a modest window can serve as a root after post-training.

Do reasoning (thinking) models make good RLM root models?

Only if they have output-token budget to spare. The paper reports that thinking models without sufficient output tokens struggle as RLMs, and attributes some smaller-than-expected gaps to trajectories that ran out of output tokens. A model that spends its allowance deliberating has nothing left to emit the code that does the actual work.

Can an open-source model run as an RLM root?

Yes, with caveats. Qwen3-Coder-480B-A35B gained from the harness at depth 1 (36.0% to 48.0% on OOLONG) but trailed GPT-5 badly on aggregation and needed its own prompt adjustments. A reproduction study found DeepSeek v3.2 and Kimi K2 worked at depth 1 on complex reasoning but overthought at depth 2. The strongest path for small models is post-training: RLM-Qwen3-8B, trained on 1,000 filtered trajectories, beat its base model by a median of 28.3%.

Do the root model and the sub-call model have to be the same?

No, and in practice they usually should not be. The original evaluation ran GPT-5 as the root and dispatched recursive sub-calls to GPT-5-mini for cost reasons. The root does orchestration, which is the hard job; sub-calls mostly read a bounded slice and answer a narrow question, which a cheaper model handles.