Back to Blog
September 20, 2026

The short version: usually yes, but not for the reason most people assume. The reference RLM library already exposes the two settings separately, as sampling_args for the root and sub_sampling_args for depth-1 sub-calls, so the split is a supported knob and not a hack. No published RLM result ablates temperature, so any specific pair of numbers is untested folklore. On most frontier roots the question is moot anyway, because reasoning models reject a non-default temperature outright.

Does the reference implementation let you set them separately?

Yes. The RLM constructor in the reference library takes both sampling_args and sub_sampling_args. The code comment states the division plainly: sampling_args "applies to the root model (depth=0)" and sub_sampling_args to "depth=1 sub-LLM calls." If you pass sub_sampling_args without naming a second backend, the library mirrors the root backend so that depth 1 routes through its own client with its own sampling args.

Both default to None. Out of the box, the root and the sub-calls therefore inherit the same provider defaults, and there is no split at all. We searched the repository and found no temperature value set anywhere, including the training configs, which set max_completion_tokens and disable thinking but never touch temperature.

One trap matters more than the choice of number. Only the OpenAI-compatible client actually forwards these args: it unpacks them into the chat.completions.create call and renames max_tokens to max_completion_tokens on the way. The Anthropic client builds its request from model, max_tokens, messages, and system only, and never reads sampling_args. The Gemini client builds a config that carries a system instruction and nothing else. On those two backends, a temperature you set is dropped in silence. Confirm the value reaches the wire before you interpret any result.

Do the RLM papers report a temperature ablation?

No. We read the latest version of the RLM paper and found no temperature, top-p, greedy decoding, or seed setting anywhere in the text. The methods section says only that GPT-5 ran "with medium reasoning and default sampling parameters," and that Qwen3-Coder-480B-A35B ran "using the sampling parameters described in" the Qwen release. Those published Qwen settings are temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05. That is a vendor default applied to every call the model makes, not a root-versus-leaf decision.

The reproduction study by Daren Wang varies recursion depth, not sampling. It reports nothing about temperature either. So the honest position is that the root-versus-sub-call temperature question has no published answer. Anyone quoting you a tuned pair is reasoning from general practice, as we are about to.

Can you even set a temperature on the models people use as roots?

Often not, and this collapses most of the question. Microsoft's reasoning-model documentation for the Azure OpenAI models states that "reasoning models other than GPT-6 Astra don't support the following parameters: temperature, top_p, presence_penalty, frequency_penalty, logprobs, top_logprobs, logit_bias, max_tokens." The whole GPT-5 family sits in that group, and the paper's own root is GPT-5 at medium reasoning.

Anthropic's thinking documentation draws a similar line and is explicit about where it applies. On its newer models, "non-default temperature, top_p, or top_k values return a 400 error on every request, regardless of whether thinking is used." On older models the restriction applies only while thinking is on, where "temperature and top_k are incompatible with thinking, and top_p is allowed at values between 0.95 and 1."

So the split is genuinely available only when at least one side is an open-weights or non-reasoning model. If your root is a frontier reasoning model, the practical question is not "which two temperatures" but "the provider default at the root, and what at the leaf." That is also why the root and the leaf are usually different models to begin with, as our post on choosing an RLM base model covers.

Would a low root temperature make the REPL loop reproducible?

No, and this is the most common bad reason for turning the root dial down. Hosted endpoints do not offer determinism at any temperature. Thinking Machines Lab traced the cause: "the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies," and standard kernels are not batch invariant, so "when the batch size changes, each element in the batch can get different results." With batch-invariant kernels their 1,000 runs produced identical completions; the default implementation produced 80 unique ones. You cannot buy that property through an API.

The second reason is that a root model's run-to-run variance comes mostly from which decomposition it chooses, not from token-level jitter inside a chosen plan. Renze and Guven swept temperature from 0.0 to 1.6 across nine models and five prompting techniques and found that "changes in temperature from 0.0 to 1.0 do not have a statistically significant impact on LLM performance." That was multiple-choice problem solving, not REPL code, so do not over-read it. It does argue against expecting a large accuracy win from the root dial alone.

There is one real argument for keeping the root low. The root's output is a program, and a malformed cell wastes an entire turn. Work on code decoding splits code tokens into "challenging tokens that are difficult to predict and confident tokens that can be easily inferred," and lowers the temperature on the confident ones to avoid "the influence of tail randomness noises." Most of what an RLM root writes is the confident kind: slice a string, run a regex, loop over chunks, call llm_query. If your provider lets you set it, there is no upside to sampling that code hot. If you want stable evaluation numbers, get them from a fixed harness and repeated runs, which our guide to evaluating an RLM sets out.

When does a higher sub-call temperature actually help?

Only when you sample the same sub-call more than once and aggregate the answers. That is the self-consistency setup: Wang and colleagues "first sample a diverse set of reasoning paths instead of only taking the greedy one, and then select the most consistent answer by marginalizing out the sampled reasoning paths," reporting gains such as 17.9 points on GSM8K. Diversity is the whole mechanism, and it requires repeated samples of one question.

The default RLM fan-out is not that shape. Sub-calls go out over distinct chunks, so each one asks a different question and each answer is used exactly once. Diversity between those calls buys nothing, because nothing compares them. Raising temperature there only adds noise to extraction, and bad extractions are precisely what the root aggregates into a confidently wrong final answer. Our post on RLM hallucination of REPL output describes how that failure looks from the trajectory.

The rule that follows is simple: temperature should follow the aggregation, not the recursion depth. A sub-call whose result is used once belongs at the low end. A sub-call you issue several times over the same slice and then majority-vote belongs at a sampling temperature. Budget for it, because k samples cost k times as much, and sub-call cost is already the part of an RLM bill that blows out on the tail.

What should you actually set?

  1. Check the value reaches the API. On the reference library, only the OpenAI-compatible client forwards sampling args. Log the outgoing request once and confirm.
  2. Leave the root at the provider default. With a reasoning model you have no choice. Without one, you have no evidence that moving it helps.
  3. Put single-shot extraction sub-calls at the low end. These calls read a bounded slice and answer a narrow question. Creativity is not the job.
  4. Raise temperature only where you vote. Repeated samples over the same slice, aggregated by majority, is the one pattern where diversity is paying for something.
  5. Do not shop for determinism. Report medians and spread over repeated runs instead.
  6. Log whatever you set. The papers do not, which is exactly why this question has no published answer yet.

The bottom line

The reference implementation supports different temperatures at the root and in sub-calls, and for most workloads a split is reasonable: provider default at the root, low for single-use sub-calls. What no source supports is a tuned pair of numbers. The RLM paper and its reproduction never report a temperature at all, and the models most often used as roots reject the parameter, so the decision is really about the leaves. Set the leaf temperature by what you do with the answer. One use means low. A vote means diversity. Everything else is guesswork with a decimal point on it.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025 (v3, May 2026). No temperature or sampling ablation; GPT-5 run with medium reasoning and default sampling parameters. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Depth ablations with DeepSeek v3.2 and Kimi K2, and no temperature ablation. arxiv.org/abs/2603.02615
  3. Zhang, A. L., et al. "rlm." GitHub repository, accessed September 2026. The sampling_args and sub_sampling_args constructor arguments, their depth-0 and depth-1 scoping, and which clients forward them. github.com/alexzhang13/rlm
  4. Microsoft. "Azure OpenAI reasoning models." Microsoft Learn, updated September 2026. The list of parameters reasoning models do not support, including temperature and top_p. learn.microsoft.com/en-us/azure/foundry/openai/how-to/reasoning
  5. Anthropic. "Thinking." Claude Developer Platform documentation, accessed September 2026. Non-default temperature, top_p, and top_k returning 400 on newer models, and the thinking-only restriction on older ones. platform.claude.com/docs/en/build-with-claude/thinking
  6. Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., Zhou, D. "Self-Consistency Improves Chain of Thought Reasoning in Language Models." arXiv:2203.11171, March 2022. Sampling diverse reasoning paths and marginalizing to the most consistent answer. arxiv.org/abs/2203.11171
  7. Renze, M., Guven, E. "The Effect of Sampling Temperature on Problem Solving in Large Language Models." arXiv:2402.05201, February 2024. No statistically significant effect of temperature 0.0 to 1.0 on problem-solving accuracy. arxiv.org/abs/2402.05201
  8. Zhu, Y., Li, J., Li, G., Zhao, Y., Li, J., Jin, Z., Mei, H. "Hot or Cold? Adaptive Temperature Sampling for Code Generation with Large Language Models." arXiv:2309.02772, September 2023. Challenging versus confident code tokens and tail randomness noise. arxiv.org/abs/2309.02772
  9. He, H., Thinking Machines Lab. "Defeating Nondeterminism in LLM Inference." September 2025. Batch-size variation as the cause of nondeterminism, and the 1,000-run batch-invariance experiment. thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference
  10. Qwen Team. "Qwen3-Coder-480B-A35B-Instruct" model card. Hugging Face, accessed September 2026. The recommended temperature 0.7, top_p 0.8, top_k 20, repetition penalty 1.05 settings referenced by the RLM paper. huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct
FAQ

Frequently asked questions

What are sampling_args and sub_sampling_args in the rlm library?

They are two constructor arguments on the RLM class that carry sampling settings such as temperature, top_p, max_tokens, and seed. The first applies to the root model at depth 0, the second to depth-1 sub-LLM calls. Both default to None, so an unconfigured run uses provider defaults on both sides.

Why does my temperature setting seem to do nothing on Claude or Gemini?

The reference library only forwards sampling args through its OpenAI-compatible client. The Anthropic client builds its request from model, max_tokens, messages, and system, and the Gemini client only passes a system instruction. Anything else you set on those backends is dropped without an error.

Does temperature 0 give the same output every time?

No. Thinking Machines Lab showed that hosted endpoints vary because server load changes the batch size, and standard kernels are not batch invariant. Their default setup produced 80 unique completions across 1,000 runs at the same settings. Batch-invariant kernels fixed it, but that is not something an API exposes.

Is self-consistency voting worth adding to an RLM?

Only for sub-calls you sample several times over the same slice and then aggregate. Self-consistency works by comparing diverse reasoning paths for one question. The normal RLM fan-out asks a different question of each chunk, so there is nothing to compare and the extra samples only multiply cost.

What temperature did the original RLM evaluation use?

The paper never states one. It says GPT-5 ran with medium reasoning and default sampling parameters, and that Qwen3-Coder ran with the sampling parameters from its model card, which are temperature 0.7, top_p 0.8, top_k 20, and repetition penalty 1.05. Neither is a root-versus-sub-call choice.