Back to Blog
September 13, 2026

The short version: partly. Provider prompt caching bills repeated prefix tokens at about a tenth of the normal input rate, but only when the start of the prompt matches a recent request exactly. The root loop of an RLM fits that shape well, because each turn resends the whole earlier trajectory. Sub-calls mostly do not, because each one carries a different chunk, so caching saves money on them only when you put a long shared prefix first on purpose.

How does provider prompt caching work?

All three major APIs cache a prefix. The provider stores the processed state of the first part of a prompt. A later request that starts with the same tokens reuses it and pays a reduced rate for those tokens. A single changed token early in the prompt breaks the match for everything after it.

  • Anthropic. The prompt caching docs give the prefix order as tools, then system, then messages. You mark up to 4 breakpoints with cache_control, or set one top-level cache_control for automatic caching, which moves the breakpoint forward as a conversation grows. Cache writes cost 1.25x the base input price for the default 5-minute TTL and 2x for the 1-hour TTL. Cache reads cost 0.1x. The minimum cacheable length ranges from 512 to 4,096 tokens depending on the model, and a shorter prompt is simply not cached, with no error.
  • OpenAI. The OpenAI guide says caching "is enabled by default for supported OpenAI models" and that "cache reuse requires the entire rendered prefix to match." For GPT-5.6 and later, the minimum is 1,024 tokens, cached tokens cost 0.1x the uncached rate, and cache writes cost 1.25x. In-memory retention lasts "around 5 to 10 minutes of inactivity, up to one hour," and extended retention can keep entries for up to 24 hours.
  • Google. The Gemini caching docs say implicit caching "is enabled by default for all Gemini 2.5 and newer models," with minimums of 2,048 tokens for Gemini 2.5 Flash and Pro and 4,096 for newer models. Explicit caches default to a 1-hour TTL and bill storage by the hour. The pricing page lists Gemini 2.5 Pro input at $1.25 per million tokens and cached input at $0.125, plus $4.50 per million tokens per hour of storage.

The common rules are simple. Stable content goes first. The prefix must pass a minimum length. The entry expires after minutes of idle time unless you pay for longer retention.

What does the RLM paper say about cost and caching?

It reports cost, not caching. The RLM paper finds that "the median RLM run is cheaper than the median base model run, but more expensive on average due to outlier trajectories." On BrowseComp-Plus with 1,000 documents, RLM(GPT-5) averaged $0.99 per query, against a linearly extrapolated $1.50 to $2.75 for GPT-5-mini reading 6 to 11 million input tokens. The authors state that "all LM calls are blocking / sequential" and list "exploding sub-call costs" as future work.

We searched the full text of the paper for any mention of prompt, prefix, or KV caching and found none. The reproduction study reports that API costs grow by orders of magnitude once the RLM harness is used, and it does not mention caching either. So published RLM cost figures are almost certainly uncached figures. Our post on RLM latency and cost covers those numbers. This post covers what caching could change.

Which parts of an RLM call are a stable prefix?

The root loop is the best case. In the reference rlm library, _setup_prompt builds a system message with the REPL instructions, followed by a user message that states the context type and total length. Each turn then appends to the same list. A code comment calls it a "fully prefixed trajectory," a continuous [system, metadata, user_0, assistant_0, repl_0, user_1, ...] chain. Turn 12 therefore resends turns 1 through 11 unchanged, and the prefix grows every turn. That is the multi-turn chat shape that every provider cache is built for.

Two things break it. The metadata message includes the character count of the context, so the prefix differs between tasks, although it stays stable within one run. Compaction also rewrites history: the library keeps the first two messages and replaces the rest with a summary, so the next turn misses the cache after that point. Our compaction comparison covers why RLMs compact at all.

Sub-calls are the hard case. In the library, llm_query sends one prompt string, and the Anthropic client wraps a plain string as a single user message with no system prompt. When the root model writes something like llm_query(f"Find mentions of X in: {chunk}"), the only shared prefix is the short instruction. It is almost always below the 512 to 4,096 token minimums, so no entry is written, and each chunk pays full price. The chunks themselves never repeat, because splitting the input into distinct pieces is the whole point of the decompose, recurse, aggregate pattern.

How should you structure sub-call prompts to hit the cache?

  1. Put the shared part first and make it worth caching. If every sub-call needs the same long rubric, schema, glossary, or set of worked examples, put it at the top of the prompt, above the provider minimum, and put the chunk last. The first call writes the entry, and the rest read it.
  2. Invert the order when one chunk gets many questions. If the root asks several questions about the same large section, put that section first and the question last.
  3. Keep the prefix byte-identical. No timestamps, turn counters, or chunk indices in the shared part. OpenAI says to place dynamic content "at the end rather than the beginning," and Google gives the same advice.
  4. Write the entry before you fan out. A request can only read an entry that an earlier request wrote. Send one sub-call first, then send the batch through llm_query_batched while the TTL is still live.
  5. Mark breakpoints on Anthropic. The rlm library client code we read sets no cache_control, so nothing is cached on Anthropic until you add it. OpenAI and Gemini implicit caching need no code change.
  6. Check the usage fields. Read cache_read_input_tokens on Anthropic or cached_tokens on OpenAI. A zero there means the prefix is too short or it changed.

The write premium sets the break-even point. With Anthropic multipliers, a prefix written at 1.25x and read once at 0.1x costs 1.35x, against 2x uncached, so two uses already save money. A 1-hour write at 2x needs a third use to beat the uncached price. A prefix used once is a small loss.

What about local inference?

Self-hosted engines do the same thing without a price list. vLLM automatic prefix caching "caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix." Its design notes say it hashes each KV block "by the tokens in the block and the tokens in the prefix before the block," caches only full blocks, and evicts the least recently used block first. It is set with enable_prefix_caching. The docs add that it "only reduces the time of processing the queries (the prefilling phase) and does not reduce the time of generating new tokens."

SGLang goes further. Its paper describes RadixAttention, which keeps prompts and generation results "in a radix tree, enabling efficient prefix search, reuse, insertion, and eviction," with a cache-aware scheduler that orders requests by matched prefix length. The authors report up to 6.4x higher throughput than other inference systems of the time. For an RLM on local hardware, the gain shows up as GPU time instead of a smaller invoice, and the same rule applies: shared tokens first.

The bottom line

Prompt caching does cut RLM cost, but mostly in the root loop, where each turn resends a growing and unchanged trajectory. Sub-calls save money only when they share a long prefix that passes the provider minimum, which the typical short instruction plus chunk does not. If you want cached sub-calls, put a real shared prefix first and the chunk last, keep that prefix identical, warm it before the batch, and read the usage fields. Published RLM cost figures, including the paper's, do not use caching, so treat them as a ceiling, not a floor.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025. Median vs average cost, the $0.99 BrowseComp-Plus figure, blocking sequential calls, and no mention of caching. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Cost growth under the RLM harness, with no caching discussed. arxiv.org/abs/2603.02615
  3. Anthropic. "Prompt caching." Claude Developer Platform documentation, accessed September 2026. Prefix order, breakpoints, automatic caching, write and read multipliers, TTLs, minimum lengths, and usage fields. platform.claude.com/docs/en/build-with-claude/prompt-caching
  4. OpenAI. "Prompt caching." API documentation, accessed September 2026. Default-on caching, exact prefix match, 1,024-token minimum, 0.1x cached rate, 1.25x writes, and retention windows. developers.openai.com/api/docs/guides/prompt-caching
  5. Google. "Context caching." Gemini API documentation, accessed September 2026. Implicit caching on Gemini 2.5 and newer, minimum token counts, and prompt ordering advice. ai.google.dev/gemini-api/docs/caching
  6. Google. "Context caching" (generateContent). Gemini API documentation, accessed September 2026. The 1-hour default TTL and storage billing for explicit caches. ai.google.dev/gemini-api/docs/generate-content/caching
  7. Google. "Gemini Developer API pricing." Accessed September 2026. Gemini 2.5 Pro input, cached input, and storage prices. ai.google.dev/gemini-api/docs/pricing
  8. vLLM. "Automatic Prefix Caching." Documentation, accessed September 2026. KV reuse for shared prefixes and the prefill-only limitation. docs.vllm.ai/en/stable/features/automatic_prefix_caching
  9. vLLM. "Automatic Prefix Caching" design document. Accessed September 2026. Block hashing, full-block caching, and LRU eviction. docs.vllm.ai/en/stable/design/prefix_caching
  10. Zheng, L., et al. "SGLang: Efficient Execution of Structured Language Model Programs." arXiv:2312.07104, December 2023. RadixAttention, LRU leaf eviction, cache-aware scheduling, and up to 6.4x throughput. arxiv.org/abs/2312.07104
  11. Zhang, A. L., et al. "rlm." GitHub repository, accessed September 2026. The append-only root message history, the metadata message, compaction, and the llm_query client path. github.com/alexzhang13/rlm
FAQ

Frequently asked questions

Does the RLM paper use prompt caching?

No mention of it appears in the paper. It reports that all LM calls were blocking and sequential, and it gives uncached API costs. The reproduction study does not mention caching either.

Why do most RLM sub-calls miss the cache?

Each sub-call usually sends a short instruction followed by a unique chunk. The shared part is shorter than the provider minimum, which ranges from 512 to 4,096 tokens, so nothing is written. The chunk differs every time, so it cannot match.

Is caching worth it if a prefix is only used once?

No. On Anthropic a 5-minute cache write costs 1.25x the base input price, so a single use costs more than no caching. Two uses at 1.25x plus 0.1x already cost less than two uncached calls.

Does the rlm library turn on prompt caching?

The client code we read sets no cache_control, so Anthropic requests are not cached unless you add it. OpenAI and Gemini 2.5 and newer cache matching prefixes automatically. Check the usage fields to confirm hits.

Does vLLM prefix caching make RLM sub-calls generate faster?

It speeds up the prefill phase only. The vLLM docs state that it does not reduce the time spent generating new tokens. Sub-calls with long outputs see smaller gains.