In wall-clock time, yes -- often dramatically. But the measured slowdown is mostly an implementation choice, and the cost story runs the opposite direction from the speed story.
The short version: yes, current RLMs are slower than a single long-context call, usually by one to two orders of magnitude. A reproduction study measured DeepSeek v3.2 at 3.6 seconds for a base S-NIAH query, 89.3 seconds through an RLM at depth 1, and 344.5 seconds at depth 2. The cause is not the paradigm; it is that today's implementations run every sub-call sequentially and blocking, with no prefix caching. And speed is not cost: the median RLM(GPT-5) run was cheaper than the median GPT-5 run in the original evaluation.
So the honest answer has three parts: how slow, why, and whether slow also means expensive. Each part has measured numbers behind it.
The clearest measurements come from a reproduction study that tracked execution time explicitly (arXiv:2603.02615). On S-NIAH, a simple needle-in-a-haystack retrieval task, DeepSeek v3.2 answered a base query in 3.6 seconds. Wrapping the same model in an RLM harness at depth 1 inflated that to 89.3 seconds -- roughly 25x. Pushing recursion to depth 2 took it to 344.5 seconds, nearly 100x the base call. Kimi K2 showed the same pattern and peaked at 545.5 seconds per query at depth 2. One OOLONG run at depth 1 clocked 1,223.6 seconds -- over twenty minutes for a single query.
The original authors do not dispute the direction of this result. They report that, depending on how the root model partitions the input, queries "range from a few seconds to several minutes," and they concede they have no strong guarantees on the total runtime of any given call. Latency is the least flattering column in the RLM ledger, which is why our guide to evaluating an RLM insists on reporting the latency tail, not just accuracy.
| Setup (S-NIAH) | Wall-clock time | Accuracy |
|---|---|---|
| DeepSeek v3.2, base call | 3.6s | 100.0% |
| DeepSeek v3.2, RLM depth 1 | 89.3s | 85.0% |
| DeepSeek v3.2, RLM depth 2 | 344.5s | 70.0% |
| Kimi K2, RLM depth 2 | 545.5s (peak) | degraded vs base |
Note what that table also says about accuracy. On a task the base model already solves -- both models scored a perfect 100.0% on S-NIAH natively -- the RLM was slower and worse. The latency question only becomes interesting on tasks where the single call fails.
Because of how the reference implementation is built, not because recursion is inherently slow. The original paper is unusually candid here: every LM call in the implementation is blocking and sequential, and no call takes advantage of prefix caching. Each recursive sub-call waits for the previous one to finish. Total latency is therefore the sum of every sub-call, every REPL execution, and every root-model turn, laid end to end.
A single long-context call has none of that structure. The provider streams one response from one forward pass over one (very large) prompt. It may be expensive in prefill compute, but it is one round trip. An RLM trajectory on a hard task can involve dozens of round trips, each carrying API overhead, and in the current implementation none of them overlap. The authors state the consequence directly: "RLMs without asynchronous LM calls are slow," especially compared to just the base model.
Depth multiplies the problem. Each level of recursion nests another sequential loop inside a sequential loop, which is why the reproduction study saw time grow superlinearly with depth. Curiously, tokens did not always follow: DeepSeek v3.2's token usage on S-NIAH actually dropped from 25.2k at depth 1 to 20.1k at depth 2 while execution time nearly quadrupled. Time and tokens are not the same axis. Deep trajectories stall in many small isolated loops rather than burning tokens in bulk -- one more reason depth should stay shallow unless the task demands it.
Mostly artifact. When the root model splits a long input into chunks and queries each chunk, those sub-calls are independent of one another; nothing in the paradigm requires chunk 2 to wait for chunk 1. The original authors say runtime "can be significantly improved through asynchrony of LM calls," and they flag async execution as the obvious engineering fix they deliberately skipped. Prefix caching is the second free win: sub-calls that share a system prompt currently re-pay for it on every call.
What asynchrony cannot remove is the serial floor. The root model's turns are genuinely sequential -- turn N depends on what the REPL returned at turn N-1. A trajectory with 15 orchestration turns keeps 15 round trips of latency no matter how well the leaves parallelize. So a well-engineered RLM should land much closer to a single call than today's numbers suggest, but it will not match one on tasks that need many adaptive turns. That residual is the price of the harness's flexibility, and it belongs on the same list as the other documented RLM limitations.
No, and this is where the comparison flips. In the original evaluation, the median RLM(GPT-5) run was cheaper than the median GPT-5 run. On OOLONG at 132K tokens, RLM(GPT-5-mini) beat GPT-5 by 114% at roughly the same total API cost per query; at 263K tokens it was cheaper per query on average. On BrowseComp-Plus at 1,000 documents -- inputs of 6-11M tokens -- RLM(GPT-5) averaged $0.99 per query and beat both the summarization and retrieval baselines by over 29%, while the summarization-agent baseline cost up to 3x more.
The mechanism is simple. A single long-context call pays prefill on every token whether or not it matters. An RLM reads selectively: peek, grep, slice, and only send the relevant pieces through sub-calls. Slow does not mean wasteful; the wall-clock cost is round trips, not tokens.
The caveat is variance. RLM cost distributions are long-tailed: a run that fails to converge keeps making sub-calls, and the paper's cost curves show sharp increases at the 95th percentile. Median cheap, tail expensive, with no hard guarantee on where any single query lands. Budget for the tail, not the median.
When the input does not fit. A single long-context call is only an option inside the model's window; RLMs process inputs up to two orders of magnitude beyond it, at the 10M+ token scale. In that regime there is no single call to be slower than -- the alternatives are retrieval pipelines, summarization agents, or nothing. "Slower than a single call" is a question about the overlap zone, where both approaches are feasible, and that zone is exactly where the accuracy gap is smallest.
Treat it as a three-way decision. If the task is simple retrieval inside the window, use the single call: it was 25x faster and more accurate in the measured case. If the task is hard aggregation or reasoning over a long input -- the regime where base models collapse and RLMs more than double their scores -- accept the minutes, or engineer them down with async sub-calls, and pocket the median cost savings. If the input exceeds the window entirely, the question answers itself.
The reproduction study's verdict for interactive, industrial deployment is blunt: massive latency penalties currently outweigh the benefits when a frontier model can natively handle the context "at a fraction of the time and cost." That is a fair reading of today's unoptimized harnesses on tasks that fit in a window. It is not a verdict on the paradigm. The slowness lives in the plumbing, and plumbing gets fixed.
Roughly one to two orders of magnitude in current implementations. In a reproduction study, DeepSeek v3.2 answered a base S-NIAH query in 3.6 seconds. The same query through an RLM at depth 1 took 89.3 seconds, and depth 2 took 344.5 seconds. Kimi K2 peaked at 545.5 seconds per query at depth 2. The original authors report that queries range from a few seconds to several minutes depending on how the model partitions the input.
Because the reference implementation runs every LM call as blocking and sequential, with no prefix caching. Each recursive sub-call waits for the previous one to finish, so total latency is the sum of every sub-call plus every REPL turn. The original paper says this plainly: RLMs without asynchronous LM calls are slow, and the authors did not optimize their implementation for speed.
No. Sub-calls over disjoint chunks are independent of each other, so they can run concurrently. The original authors state that runtime can be significantly improved through asynchrony of LM calls. What is fundamental is the serial floor: the root model's own turns depend on prior REPL observations, so orchestration turns cannot be parallelized away.
No. Latency and cost separate cleanly. The original evaluation found the median RLM(GPT-5) run cheaper than the median GPT-5 run, with RLM(GPT-5-mini) matching GPT-5's per-query API cost on OOLONG at 132K tokens while scoring 114% higher. On BrowseComp-Plus at 1,000 documents, RLM(GPT-5) averaged $0.99 per query over 6-11M token inputs. The caveat is variance: the cost distribution has a heavy tail from long trajectories.