Back to Blog
September 3, 2026

The short version: compaction handles long sessions well enough when the work is linear and the facts that matter are recent. An RLM handles them better when the session holds specific facts that must be recalled verbatim, when the answer depends on many scattered parts of the history, or when a summary would have to guess what matters later. The measured evidence: in the original RLM evaluation, RLM(GPT-5) beat a compaction agent by a median of 26% across four long-context benchmarks. Compaction is cheaper and faster per turn. Recursion keeps the original text and pays for it in round trips.

Both approaches attack the same problem. A session that runs for hours fills the window, and quality drops long before the hard limit. The RLM authors describe it in their blog: "as the conversation goes on, the model gets...dumber?" Chroma's context rot study measured it across 18 models, and Liu et al. showed the U-shaped position curve two years earlier. The question is what to do with the parts you shrink.

What does compaction actually keep and drop?

Compaction replaces old turns with a model-written summary. Anthropic's server-side compaction triggers when input tokens reach a threshold (150,000 by default, with a floor of 50,000). The API generates a summary, writes it into a compaction block, and "automatically drops all content blocks prior to the compaction block." OpenAI's Responses API does the same job through a compact_threshold setting or a standalone /responses/compact endpoint. Its compaction item is "opaque and not intended to be human-interpretable," and it "carries forward key prior state and reasoning into the next run using fewer tokens."

Context editing is the lighter cousin. Anthropic's clear_tool_uses_20250919 strategy removes old tool results once the prompt passes 100,000 input tokens, keeps the three most recent by default, and swaps each cleared result for placeholder text.

Both mechanisms are lossy by design, and the docs say so. Anthropic's context engineering guide describes what Claude Code's compaction keeps ("architectural decisions, unresolved bugs, and implementation details") and what it discards ("redundant tool outputs or messages"). The summarizer decides what is important before it knows what the next question will be. On Claude Fable 5.1 and Claude Mythos 5.1 the compaction docs add a second loss: thinking blocks from before a compaction block "aren't carried forward, so the summary is all the model has of that earlier work."

What does an RLM keep that a summary cannot?

The original text. An RLM treats the long prompt as "part of an external environment" and gives the root model "a symbolic handle to the user prompt" that it can manipulate "without copying text into the root context window." A session history stored this way is never summarized away. The root model peeks at it, greps it, slices it, and hands slices to fresh sub-calls. Any fact from turn 3 is still retrievable at turn 300, byte for byte.

That is the core difference. A compaction summary is a single, early, irreversible decision about relevance. An RLM makes that decision late, per query, and can make it again. The RLM paper's compaction baseline is "an iterative agent that compacts the context as it is filled," one that will "iteratively accumulate the documents and summarize when full." On OOLONG-Pairs, the task that requires aggregating pairs of chunks across the whole input, that baseline scored 0.1% F1 with GPT-5. RLM(GPT-5) scored 76%. On plain OOLONG the gap was 46% against 56.5%. No summary can hold the pairwise structure the task needs, because the summary was written before the pairs were asked for.

This is the same argument made in our post on why context windows are the wrong abstraction. The question is not how much the model can see at once. It is whether the model can go back and look.

When is compaction enough?

Most of the time, for coding agents. A coding session has a strong recency bias. The current file, the current error, and the current plan matter. Compaction's failure mode is losing a detail the summary did not flag, and on linear tasks that detail is usually recoverable by re-reading the file or re-running the command.

Compaction is enough when:

  • The task is one long thread, and old state is superseded rather than needed again.
  • The facts that matter live on disk or in a database, so the agent can re-fetch them.
  • You can steer the summary. Claude Code accepts /compact Focus on code samples and API usage, or a compact-instructions block in CLAUDE.md.

The Chroma result cuts both ways. On LongMemEval, a focused prompt of about 300 tokens holding only the relevant context beat a full prompt of about 113k tokens across the models tested. The difference is whether the discarded tokens are gone or merely out of the window.

When does recursion win?

Whenever the answer depends on a specific earlier fact that the summary cannot be trusted to have kept. Four cases show up repeatedly:

  1. Needle retrieval across the session. "What port did we set in the config at the start?" A summary might have kept it. An RLM greps for it. LongMemEval found "a 30% accuracy drop" for commercial chat assistants and long-context LLMs "on memorizing information across sustained interactions."
  2. Verifiable facts versus paraphrase. A summary of a stack trace is not a stack trace. If the output must quote or match the source, recursion returns the source.
  3. Aggregation over the whole history. Counting, pairing, or cross-referencing many turns is the OOLONG-Pairs shape. Compaction scored 0.1% there.
  4. Knowledge updates. A fact stated at turn 10 and corrected at turn 200. A summary written at turn 150 carries the stale version and cannot know it is stale. An RLM reading the raw history sees both and can order them.

Recursion loses on the opposite cases. The reproduction study found that "using RLMs on simple retrieval tasks paradoxically degrades performance," and that depth 2 degraded results where depth 1 had helped. If the base model can answer from the compacted context, do not recurse.

What do the two approaches cost in latency and tokens?

Compaction costs one summary call per trigger. Anthropic's docs note that it "requires an additional sampling step, which contributes to rate limits and billing," and the Claude Code cost guide warns that "compacting a large context is itself a large request." After that, every turn runs on a small context. Context editing is cheaper still, but it invalidates the prompt cache at the clearing point, which is why its clear_at_least parameter exists.

Recursion costs round trips. The reference RLM implementation runs every sub-call blocking and sequential, and the reproduction study measured a base call at 3.6 seconds against 344.5 seconds at depth 2. The RLM paper reports that the median RLM run is cheaper than the median base run, "but more expensive on average due to outlier trajectories." The full accounting is in our latency post. Compaction is cheap and fast on every turn. Recursion is cheap at the median, slow every time, and expensive in the tail.

Can you combine them?

Yes, and the best current systems already do. The pattern is: compact the working context, but never delete the original. Keep the raw history in a store the model can query, and let a recursive call reach back into it when the summary is not enough. Voltropy's LCM does this with an immutable message store under a hierarchical summary; we covered it in our LCM write-up. Anthropic's own guidance points the same way. It describes sub-agents that explore "using tens of thousands of tokens or more" but return "only a condensed, distilled summary" of "often 1,000-2,000 tokens." That is a depth-1 RLM with a compaction step at the boundary.

A practical hybrid for a long agent session:

  • Use context editing for tool results. They are the bulk of a coding session and the easiest to re-fetch.
  • Use compaction for the conversational thread, with an instruction to keep decisions, open bugs, and named values.
  • Persist the full transcript as it happens.
  • Give the agent a tool that runs an RLM-style query over the transcript: grep first, sub-call over the matching slices, return the verbatim span.
  • Route to that tool only when the question references the past. Recursion on easy questions makes things worse.

The bottom line

Compaction is a bet that the summarizer can predict what you will need. It is a good bet for linear work with recent state, and it is cheap. Recursion refuses the bet. It keeps the source and pays in latency to consult it. For needle retrieval across a long session, for anything that must be quoted rather than paraphrased, and for aggregation over the whole history, recursion wins, and the 26% median gap in the RLM paper is the measured size of that win. For everything else, compact, keep the transcript, and add a recursive lookup for the moments when the summary runs out.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025 (revised May 2026). Compaction baseline definition, 26% median gap, OOLONG and OOLONG-Pairs scores, cost variance. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Depth-1 gains, depth-2 degradation, retrieval-task regression, 3.6s to 344.5s latency. arxiv.org/abs/2603.02615
  3. Zhang, A. L. "Recursive Language Models." Project blog, 2025. Context rot motivation, peek/grep/partition strategies, blocking sub-call caveat. alexzhang13.github.io/blog/2025/rlm
  4. Anthropic. "Compaction." Claude Developer Platform docs. Trigger thresholds, compaction block semantics, thinking-block loss, billing note. platform.claude.com/docs/en/build-with-claude/compaction
  5. Anthropic. "Context editing." Claude Developer Platform docs. clear_tool_uses_20250919 defaults, placeholder replacement, clear_at_least and cache invalidation. platform.claude.com/docs/en/build-with-claude/context-editing
  6. OpenAI. "Compaction." OpenAI API docs. compact_threshold, /responses/compact endpoint, opaque compaction item. developers.openai.com/api/docs/guides/compaction
  7. Anthropic. "Effective context engineering for AI agents." Anthropic Engineering, 2025. What Claude Code compaction keeps and discards; sub-agent summary sizes. www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  8. Anthropic. "Manage costs effectively." Claude Code docs. /compact instructions, CLAUDE.md compact block, compaction as a large request. code.claude.com/docs/en/costs
  9. Hong, K., Troynikov, A., Huber, J. "Context Rot: How Increasing Input Tokens Impacts LLM Performance." Chroma Research, 2025. 18-model evaluation, focused vs full prompt on LongMemEval. www.trychroma.com/research/context-rot
  10. Liu, N. F. et al. "Lost in the Middle: How Language Models Use Long Contexts." TACL, 2023 (arXiv:2307.03172). U-shaped position curve. arxiv.org/abs/2307.03172
  11. Wu, D. et al. "LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory." arXiv:2410.10813, 2024. 30% accuracy drop across sustained interactions. arxiv.org/abs/2410.10813
FAQ

Frequently asked questions

Does compaction lose information permanently?

On the server side, yes. Anthropic's API drops all content blocks before the compaction block on later requests, and OpenAI's compaction item is opaque. Your client can still keep the full history locally, but the model no longer sees it unless you build a way to feed pieces back in.

Is an RLM the same as giving the agent a memory database?

No. A memory database returns pre-indexed chunks by similarity. An RLM holds the raw history as a variable and lets the model write code to peek, grep, slice, and sub-call over it. The retrieval strategy is chosen per query by the model, not fixed at indexing time.

Which is cheaper per session, compaction or an RLM?

Compaction, on most sessions. It adds one summary call per trigger and then runs every turn on a small context. RLM runs are cheaper than a base call at the median in the original paper but more expensive on average because of outlier trajectories, and they are slower on every query.

Should a coding agent switch from compaction to an RLM?

Not wholesale. Keep compaction and context editing for the main thread, persist the full transcript, and add a recursive lookup tool the agent calls only when a question references the past. The reproduction study found that recursion on simple retrieval degrades accuracy, so routing matters.

Does context editing count as compaction?

It is a narrower operation. Context editing clears old tool results or thinking blocks and replaces them with placeholders; nothing is summarized. Compaction rewrites the whole prior conversation into a summary. Both shrink the working context, and both are lossy.