Compaction summarizes the past and moves on. An RLM keeps the past and reads it on demand. Which one you want depends on whether the summary can be trusted to guess your next question.
The short version: compaction handles long sessions well enough when the work is linear and the facts that matter are recent. An RLM handles them better when the session holds specific facts that must be recalled verbatim, when the answer depends on many scattered parts of the history, or when a summary would have to guess what matters later. The measured evidence: in the original RLM evaluation, RLM(GPT-5) beat a compaction agent by a median of 26% across four long-context benchmarks. Compaction is cheaper and faster per turn. Recursion keeps the original text and pays for it in round trips.
Both approaches attack the same problem. A session that runs for hours fills the window, and quality drops long before the hard limit. The RLM authors describe it in their blog: "as the conversation goes on, the model gets...dumber?" Chroma's context rot study measured it across 18 models, and Liu et al. showed the U-shaped position curve two years earlier. The question is what to do with the parts you shrink.
Compaction replaces old turns with a model-written summary. Anthropic's server-side compaction triggers when input tokens reach a threshold (150,000 by default, with a floor of 50,000). The API generates a summary, writes it into a compaction block, and "automatically drops all content blocks prior to the compaction block." OpenAI's Responses API does the same job through a compact_threshold setting or a standalone /responses/compact endpoint. Its compaction item is "opaque and not intended to be human-interpretable," and it "carries forward key prior state and reasoning into the next run using fewer tokens."
Context editing is the lighter cousin. Anthropic's clear_tool_uses_20250919 strategy removes old tool results once the prompt passes 100,000 input tokens, keeps the three most recent by default, and swaps each cleared result for placeholder text.
Both mechanisms are lossy by design, and the docs say so. Anthropic's context engineering guide describes what Claude Code's compaction keeps ("architectural decisions, unresolved bugs, and implementation details") and what it discards ("redundant tool outputs or messages"). The summarizer decides what is important before it knows what the next question will be. On Claude Fable 5.1 and Claude Mythos 5.1 the compaction docs add a second loss: thinking blocks from before a compaction block "aren't carried forward, so the summary is all the model has of that earlier work."
The original text. An RLM treats the long prompt as "part of an external environment" and gives the root model "a symbolic handle to the user prompt" that it can manipulate "without copying text into the root context window." A session history stored this way is never summarized away. The root model peeks at it, greps it, slices it, and hands slices to fresh sub-calls. Any fact from turn 3 is still retrievable at turn 300, byte for byte.
That is the core difference. A compaction summary is a single, early, irreversible decision about relevance. An RLM makes that decision late, per query, and can make it again. The RLM paper's compaction baseline is "an iterative agent that compacts the context as it is filled," one that will "iteratively accumulate the documents and summarize when full." On OOLONG-Pairs, the task that requires aggregating pairs of chunks across the whole input, that baseline scored 0.1% F1 with GPT-5. RLM(GPT-5) scored 76%. On plain OOLONG the gap was 46% against 56.5%. No summary can hold the pairwise structure the task needs, because the summary was written before the pairs were asked for.
This is the same argument made in our post on why context windows are the wrong abstraction. The question is not how much the model can see at once. It is whether the model can go back and look.
Most of the time, for coding agents. A coding session has a strong recency bias. The current file, the current error, and the current plan matter. Compaction's failure mode is losing a detail the summary did not flag, and on linear tasks that detail is usually recoverable by re-reading the file or re-running the command.
Compaction is enough when:
/compact Focus on code samples and API usage, or a compact-instructions block in CLAUDE.md.The Chroma result cuts both ways. On LongMemEval, a focused prompt of about 300 tokens holding only the relevant context beat a full prompt of about 113k tokens across the models tested. The difference is whether the discarded tokens are gone or merely out of the window.
Whenever the answer depends on a specific earlier fact that the summary cannot be trusted to have kept. Four cases show up repeatedly:
Recursion loses on the opposite cases. The reproduction study found that "using RLMs on simple retrieval tasks paradoxically degrades performance," and that depth 2 degraded results where depth 1 had helped. If the base model can answer from the compacted context, do not recurse.
Compaction costs one summary call per trigger. Anthropic's docs note that it "requires an additional sampling step, which contributes to rate limits and billing," and the Claude Code cost guide warns that "compacting a large context is itself a large request." After that, every turn runs on a small context. Context editing is cheaper still, but it invalidates the prompt cache at the clearing point, which is why its clear_at_least parameter exists.
Recursion costs round trips. The reference RLM implementation runs every sub-call blocking and sequential, and the reproduction study measured a base call at 3.6 seconds against 344.5 seconds at depth 2. The RLM paper reports that the median RLM run is cheaper than the median base run, "but more expensive on average due to outlier trajectories." The full accounting is in our latency post. Compaction is cheap and fast on every turn. Recursion is cheap at the median, slow every time, and expensive in the tail.
Yes, and the best current systems already do. The pattern is: compact the working context, but never delete the original. Keep the raw history in a store the model can query, and let a recursive call reach back into it when the summary is not enough. Voltropy's LCM does this with an immutable message store under a hierarchical summary; we covered it in our LCM write-up. Anthropic's own guidance points the same way. It describes sub-agents that explore "using tens of thousands of tokens or more" but return "only a condensed, distilled summary" of "often 1,000-2,000 tokens." That is a depth-1 RLM with a compaction step at the boundary.
A practical hybrid for a long agent session:
Compaction is a bet that the summarizer can predict what you will need. It is a good bet for linear work with recent state, and it is cheap. Recursion refuses the bet. It keeps the source and pays in latency to consult it. For needle retrieval across a long session, for anything that must be quoted rather than paraphrased, and for aggregation over the whole history, recursion wins, and the 26% median gap in the RLM paper is the measured size of that win. For everything else, compact, keep the transcript, and add a recursive lookup for the moments when the summary runs out.
On the server side, yes. Anthropic's API drops all content blocks before the compaction block on later requests, and OpenAI's compaction item is opaque. Your client can still keep the full history locally, but the model no longer sees it unless you build a way to feed pieces back in.
No. A memory database returns pre-indexed chunks by similarity. An RLM holds the raw history as a variable and lets the model write code to peek, grep, slice, and sub-call over it. The retrieval strategy is chosen per query by the model, not fixed at indexing time.
Compaction, on most sessions. It adds one summary call per trigger and then runs every turn on a small context. RLM runs are cheaper than a base call at the median in the original paper but more expensive on average because of outlier trajectories, and they are slower on every query.
Not wholesale. Keep compaction and context editing for the main thread, persist the full transcript, and add a recursive lookup tool the agent calls only when a question references the past. The reproduction study found that recursion on simple retrieval degrades accuracy, so routing matters.
It is a narrower operation. Context editing clears old tool results or thinking blocks and replaces them with placeholders; nothing is summarized. Compaction rewrites the whole prior conversation into a summary. Both shrink the working context, and both are lossy.