Prompt caching bills repeated prefixes at a fraction of the input price. Here is which parts of an RLM run have a stable prefix, and how to shape sub-calls so they hit the cache.
The short version: partly. Provider prompt caching bills repeated prefix tokens at about a tenth of the normal input rate, but only when the start of the prompt matches a recent request exactly. The root loop of an RLM fits that shape well, because each turn resends the whole earlier trajectory. Sub-calls mostly do not, because each one carries a different chunk, so caching saves money on them only when you put a long shared prefix first on purpose.
All three major APIs cache a prefix. The provider stores the processed state of the first part of a prompt. A later request that starts with the same tokens reuses it and pays a reduced rate for those tokens. A single changed token early in the prompt breaks the match for everything after it.
tools, then system, then messages. You mark up to 4 breakpoints with cache_control, or set one top-level cache_control for automatic caching, which moves the breakpoint forward as a conversation grows. Cache writes cost 1.25x the base input price for the default 5-minute TTL and 2x for the 1-hour TTL. Cache reads cost 0.1x. The minimum cacheable length ranges from 512 to 4,096 tokens depending on the model, and a shorter prompt is simply not cached, with no error.The common rules are simple. Stable content goes first. The prefix must pass a minimum length. The entry expires after minutes of idle time unless you pay for longer retention.
It reports cost, not caching. The RLM paper finds that "the median RLM run is cheaper than the median base model run, but more expensive on average due to outlier trajectories." On BrowseComp-Plus with 1,000 documents, RLM(GPT-5) averaged $0.99 per query, against a linearly extrapolated $1.50 to $2.75 for GPT-5-mini reading 6 to 11 million input tokens. The authors state that "all LM calls are blocking / sequential" and list "exploding sub-call costs" as future work.
We searched the full text of the paper for any mention of prompt, prefix, or KV caching and found none. The reproduction study reports that API costs grow by orders of magnitude once the RLM harness is used, and it does not mention caching either. So published RLM cost figures are almost certainly uncached figures. Our post on RLM latency and cost covers those numbers. This post covers what caching could change.
The root loop is the best case. In the reference rlm library, _setup_prompt builds a system message with the REPL instructions, followed by a user message that states the context type and total length. Each turn then appends to the same list. A code comment calls it a "fully prefixed trajectory," a continuous [system, metadata, user_0, assistant_0, repl_0, user_1, ...] chain. Turn 12 therefore resends turns 1 through 11 unchanged, and the prefix grows every turn. That is the multi-turn chat shape that every provider cache is built for.
Two things break it. The metadata message includes the character count of the context, so the prefix differs between tasks, although it stays stable within one run. Compaction also rewrites history: the library keeps the first two messages and replaces the rest with a summary, so the next turn misses the cache after that point. Our compaction comparison covers why RLMs compact at all.
Sub-calls are the hard case. In the library, llm_query sends one prompt string, and the Anthropic client wraps a plain string as a single user message with no system prompt. When the root model writes something like llm_query(f"Find mentions of X in: {chunk}"), the only shared prefix is the short instruction. It is almost always below the 512 to 4,096 token minimums, so no entry is written, and each chunk pays full price. The chunks themselves never repeat, because splitting the input into distinct pieces is the whole point of the decompose, recurse, aggregate pattern.
llm_query_batched while the TTL is still live.cache_control, so nothing is cached on Anthropic until you add it. OpenAI and Gemini implicit caching need no code change.cache_read_input_tokens on Anthropic or cached_tokens on OpenAI. A zero there means the prefix is too short or it changed.The write premium sets the break-even point. With Anthropic multipliers, a prefix written at 1.25x and read once at 0.1x costs 1.35x, against 2x uncached, so two uses already save money. A 1-hour write at 2x needs a third use to beat the uncached price. A prefix used once is a small loss.
Self-hosted engines do the same thing without a price list. vLLM automatic prefix caching "caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix." Its design notes say it hashes each KV block "by the tokens in the block and the tokens in the prefix before the block," caches only full blocks, and evicts the least recently used block first. It is set with enable_prefix_caching. The docs add that it "only reduces the time of processing the queries (the prefilling phase) and does not reduce the time of generating new tokens."
SGLang goes further. Its paper describes RadixAttention, which keeps prompts and generation results "in a radix tree, enabling efficient prefix search, reuse, insertion, and eviction," with a cache-aware scheduler that orders requests by matched prefix length. The authors report up to 6.4x higher throughput than other inference systems of the time. For an RLM on local hardware, the gain shows up as GPU time instead of a smaller invoice, and the same rule applies: shared tokens first.
Prompt caching does cut RLM cost, but mostly in the root loop, where each turn resends a growing and unchanged trajectory. Sub-calls save money only when they share a long prefix that passes the provider minimum, which the typical short instruction plus chunk does not. If you want cached sub-calls, put a real shared prefix first and the chunk last, keep that prefix identical, warm it before the batch, and read the usage fields. Published RLM cost figures, including the paper's, do not use caching, so treat them as a ceiling, not a floor.
No mention of it appears in the paper. It reports that all LM calls were blocking and sequential, and it gives uncached API costs. The reproduction study does not mention caching either.
Each sub-call usually sends a short instruction followed by a unique chunk. The shared part is shorter than the provider minimum, which ranges from 512 to 4,096 tokens, so nothing is written. The chunk differs every time, so it cannot match.
No. On Anthropic a 5-minute cache write costs 1.25x the base input price, so a single use costs more than no caching. Two uses at 1.25x plus 0.1x already cost less than two uncached calls.
The client code we read sets no cache_control, so Anthropic requests are not cached unless you add it. OpenAI and Gemini 2.5 and newer cache matching prefixes automatically. Check the usage fields to confirm hits.
It speeds up the prefill phase only. The vLLM docs state that it does not reduce the time spent generating new tokens. Sub-calls with long outputs see smaller gains.