Back to Blog
September 6, 2026

The short version: an RLM and a coding agent with a REPL run the same outer loop: the model writes code, an interpreter runs it, and the output feeds the next step. Three things differ. In an RLM the prompt is a variable inside the REPL, the code can call the same model on slices of it and collect the results, and the run ends when a REPL variable is marked final. A coding agent treats the filesystem as its environment and fills its window with what it reads; an RLM treats the input as the environment and keeps its window nearly empty.

The RLM authors say so in arXiv:2512.24601: their Algorithm 2, a "deceptively similar" scaffold, also "support[s] some notion of sub-calls, external objects, and code execution." The difference is "where the prompt and intermediate values live and where recursion occurs."

What do an RLM and a coding agent actually share?

The tool loop. Anthropic describes Claude Code as cycling through "gather context, take action, and verify results," with tools that "read your code, edit files, run commands" and feed each result into the next decision. OpenAI's Codex CLI docs say the agent can "inspect code, make changes, run commands, and automate repeatable work without leaving your terminal." SWE-agent's agent-computer interface lets the model "create and edit code files, navigate entire repositories, and execute tests and other programs." OpenHands agents act "by writing code, interacting with a command line, and browsing the web."

An RLM does the same thing with a narrower toolset. The paper "equip[s] an LLM with a Python REPL, where all tools, including sub-LM or sub-RLM calls, are available as modules." Each iteration "executes code in the REPL, updates REPL state (intermediate variables), and collects in stdout any printed text." That is a CodeAct loop. CodeAct (arXiv:2402.01030) proposed "executable Python code to consolidate LLM agents' actions into a unified action space."

Both also grep before they read. The RLM blog lists this as an emergent strategy: the model "look[s] for keywords or regex patterns to narrow down lines of interest."

Where does the context live?

In a coding agent, the environment is the filesystem and the shell. The task prompt, every file read, and every command output land in the context window. The paper: "prior coding agents and retrieval agents treat some designated external data source (e.g., a filesystem or a corpus of search documents) as an environment for fetching snippets. However, they can only fill up the underlying LLM's context window with snippets before facing compaction." Claude Code's docs confirm it: "As you work, context fills up. Claude compacts automatically."

In an RLM, the prompt is the environment. "Given a prompt P, the RLM initializes a Read-Eval-Print Loop (REPL) programming environment in which P is set as the value of a variable." The root model starts with "only (constant-size) metadata about the user prompt, like its length, a short prefix, and how to access parts of it." The paper calls this "a symbolic handle to the user prompt P, so the model can manipulate it without copying text into the root context window."

So a coding agent reasons about what it has loaded. An RLM reasons about a 10M-token string it never loaded, because slicing a variable costs no context. Still, the paper's Observation 2 credits the shell: "RLM(depth=0) and coding agents like Claude Code and OpenCode are able to scale beyond the context limit of the model and outperform other task-agnostic baselines on most long context settings." A file on disk plus grep gets you most of the way on retrieval-shaped work. We covered the compaction side in RLM vs context compaction.

What does recursion add that a shell does not?

A model call as a function inside the code. The RLM system prompt gives the root model llm_query, "a function that allows you to query an LLM (that can handle around 500K chars) inside your REPL environment." For depth above 1 it adds rlm_query(context, query) for "complex sub-tasks that benefit from iterative" processing. The model then writes a loop: chunk the variable, call llm_query on each chunk, store the returns, and call llm_query once more to aggregate.

A shell cannot do this. grep matches a pattern. It cannot classify a line or judge whether two paragraphs contradict each other. OOLONG-Pairs needs that work over every pair of chunks, and the authors saw the RLM "perform the necessary semantic transformation line-by-line through recursive sub-calls, while the ablation without sub-calls is forced to use keyword heuristics." This is the pattern in decompose, recurse, aggregate.

The recursion is also programmatic, not conversational. The paper contrasts RLMs with "self-delegation approaches" that let LLMs "invoke themselves as sub-agents" but are "handicapped by the underlying LLM's limited output lengths because they are designed to verbalize sub-calls autoregressively rather than producing them programmatically." A model that types each delegation as a message can launch a handful. A model that writes for chunk in chunks: can launch a thousand.

Is a coding agent with subagents already an RLM?

Closer, but no. Claude Code subagents run in their "own context window," return "only the summary," and by default nest "up to three layers below the main conversation." Two things are still missing. The parent still holds the task and its reads in its window; the input is not a variable. And the parent delegates by writing a task message, one at a time, which is the autoregressive verbalization the paper flags. No loop dispatches a sub-call per slice and collects the returns into a variable.

The blog names the split: RLMs take "a context-centric view rather than a problem-centric view." An RLM decomposes the context, on "the principle that fundamentally, LMs should decide how to break down a problem."

How does each one decide it is done?

A coding agent is done when the world is right: tests pass, the diff is written, the PR is open. Its output is a side effect on the filesystem, and its final message is a report about that side effect. Claude Code's docs describe the loop as repeating "until task complete."

An RLM is done when a variable is set. "Once the RLM sets the variable Final inside the REPL, iteration stops and the value in Final is returned as the response." The prompted version wraps the answer in FINAL(answer) or points at a variable with FINAL_VAR(variable_name). The output is a string, because "an RLM exposes the same external interface as an LLM or a reasoning model: it accepts a string prompt of arbitrary structure and produces a string response." It replaces a model call, not an agent.

The authors admit this boundary is fragile. "Distinguishing between a final answer and a thought is brittle for RLMs." The model sometimes "outputs its plan as a final answer." A coding agent has no equivalent failure, because its stopping condition is external.

How did the two compare when measured?

The paper ran Claude Code as a baseline, "Claude Opus 4.1 with Claude Code v2.0.0," in two variants: context as the initial prompt, or "offloaded to a file." The RLM used GPT-5 as root and GPT-5-mini for sub-calls, so the rows use different models and the comparison is not clean. With that caveat, Table 1 reports:

  • CodeQA (23K to 4.2M tokens): Claude Code with offloading 62.0, RLM(GPT-5, depth=1) 62.0. A shell over a file was enough.
  • BrowseComp-Plus (6M to 11M tokens): Claude Code with offloading 84.0, RLM(GPT-5, depth=1) 91.3.
  • OOLONG (131K tokens): Claude Code with offloading 48.0, RLM(GPT-5, depth=1) 56.0.
  • OOLONG-Pairs (32K tokens): Claude Code with offloading 6.5 F1, RLM(GPT-5, depth=1) 58.0, depth=3 76.0.

Without offloading, Claude Code hit input limits on CodeQA and BrowseComp-Plus and scored 0.1 on OOLONG-Pairs. The paper's headline is a median gain of "13% against Claude Code" and "130% against CodeAct with sub-calls."

Recursion is not free. The reproduction in arXiv:2603.02615 found that depth 2, or RLM use "on simple retrieval tasks," "paradoxically degrades performance" and inflated runtime "from 3.6s to 344.5s." The original implementation made every sub-call "blocking / sequential." Our post on recursion depth covers when to stop.

The bottom line

A coding agent with a REPL and an RLM share the code-execution loop and the habit of grepping before reading. They differ on three points. The RLM keeps the input as a variable outside the window, so the root never pays context for it. The RLM's code can call the model on slices and aggregate the returns, which a shell cannot. And the RLM returns a string when a variable is set, where a coding agent stops when the world is right. For "find and change something in this repo," a coding agent with a file on disk is already most of an RLM(depth=0), and the paper's numbers say so. For a model's judgment applied to every slice of a very long input, the recursion is the part a shell cannot fake.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025 (v3 May 2026). RLM definition, Algorithm 1 vs Algorithm 2, symbolic handle, critique of coding agents and self-delegation, Claude Code and CodeAct baselines, Table 1 scores, FINAL/FINAL_VAR brittleness. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Depth-2 degradation, retrieval-task regression, 3.6s to 344.5s runtime. arxiv.org/abs/2603.02615
  3. Zhang, A. L. "Recursive Language Models." Project blog, 2025. Context-centric vs problem-centric framing, peek and grep strategies, FINAL and FINAL_VAR, blocking sub-calls. alexzhang13.github.io/blog/2025/rlm
  4. Zhang, A. L. "rlm" reference implementation, GitHub. llm_query and rlm_query functions, supported execution environments, batched sub-calls. github.com/alexzhang13/rlm
  5. Anthropic. "How Claude Code works." Claude Code docs. Agentic loop phases, built-in tool categories, automatic compaction behavior. code.claude.com/docs/en/how-claude-code-works
  6. Anthropic. "Subagents." Claude Code docs. Separate context windows, summary-only return, three-layer nesting default. code.claude.com/docs/en/sub-agents
  7. OpenAI. "Codex CLI." OpenAI developer docs. Terminal agent that inspects code, makes changes, and runs commands locally with permission settings. learn.chatgpt.com/docs/codex/cli
  8. Wang, X. et al. "Executable Code Actions Elicit Better LLM Agents." ICML 2024 (arXiv:2402.01030). CodeAct: executable Python as a unified action space. arxiv.org/abs/2402.01030
  9. Yang, J. et al. "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering." arXiv:2405.15793, 2024. ACI for editing files, navigating repositories, running tests. arxiv.org/abs/2405.15793
  10. Wang, X. et al. "OpenHands: An Open Platform for AI Software Developers as Generalist Agents." ICLR 2025 (arXiv:2407.16741). Agents that write code, use a command line, and browse the web in sandboxed runtimes. arxiv.org/abs/2407.16741
FAQ

Frequently asked questions

Can I turn Claude Code or Codex CLI into an RLM?

Partly. Write the long input to a file and the agent behaves like an RLM(depth=0), which the paper found competitive on retrieval-shaped tasks. To get depth 1 you need a function the agent's code can call that runs a model on a slice and returns the string. Subagents are a weaker substitute because each one is dispatched as a typed message, not in a loop.

Is CodeAct the same thing as an RLM?

No, and the paper tests both. CodeAct runs Python inside a ReAct loop but puts the user prompt directly in the model's context. The RLM puts the prompt in the REPL as a variable. The paper reports a 130% median gain over CodeAct with sub-calls, which isolates the effect of where the prompt lives.

Does an RLM need a sandbox the way a coding agent does?

Yes, for the same reason: the model writes and runs arbitrary code. The reference implementation supports local, IPython, Docker, Modal, and other execution environments. The paper lists sandboxed REPLs as future work alongside asynchronous sub-calls.

Which model runs the recursive sub-calls?

It does not have to be the root model. In the paper's GPT-5 experiments the root was GPT-5 and every sub-call went to GPT-5-mini, chosen for cost. The Qwen3-Coder runs used the same model at both levels. The fine-tuned RLM-Qwen3-8B was trained on trajectories with Qwen3-8B sub-calls.

Why did the RLM paper benchmark Claude Code at all?

Because a coding agent with a file on disk is the closest existing system to an RLM. The authors wanted to show which gains come from offloading context to an environment, which any shell provides, and which come from programmatic recursion, which only the RLM provides. Claude Code with offloading matched the RLM on CodeQA and lost badly on OOLONG-Pairs.