The reference implementation lets you set root and sub-call sampling separately. No published RLM result tells you what to put in them, so here is what the code and the provider docs do support.
The short version: usually yes, but not for the reason most people assume. The reference RLM library already exposes the two settings separately, as sampling_args for the root and sub_sampling_args for depth-1 sub-calls, so the split is a supported knob and not a hack. No published RLM result ablates temperature, so any specific pair of numbers is untested folklore. On most frontier roots the question is moot anyway, because reasoning models reject a non-default temperature outright.
Yes. The RLM constructor in the reference library takes both sampling_args and sub_sampling_args. The code comment states the division plainly: sampling_args "applies to the root model (depth=0)" and sub_sampling_args to "depth=1 sub-LLM calls." If you pass sub_sampling_args without naming a second backend, the library mirrors the root backend so that depth 1 routes through its own client with its own sampling args.
Both default to None. Out of the box, the root and the sub-calls therefore inherit the same provider defaults, and there is no split at all. We searched the repository and found no temperature value set anywhere, including the training configs, which set max_completion_tokens and disable thinking but never touch temperature.
One trap matters more than the choice of number. Only the OpenAI-compatible client actually forwards these args: it unpacks them into the chat.completions.create call and renames max_tokens to max_completion_tokens on the way. The Anthropic client builds its request from model, max_tokens, messages, and system only, and never reads sampling_args. The Gemini client builds a config that carries a system instruction and nothing else. On those two backends, a temperature you set is dropped in silence. Confirm the value reaches the wire before you interpret any result.
No. We read the latest version of the RLM paper and found no temperature, top-p, greedy decoding, or seed setting anywhere in the text. The methods section says only that GPT-5 ran "with medium reasoning and default sampling parameters," and that Qwen3-Coder-480B-A35B ran "using the sampling parameters described in" the Qwen release. Those published Qwen settings are temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05. That is a vendor default applied to every call the model makes, not a root-versus-leaf decision.
The reproduction study by Daren Wang varies recursion depth, not sampling. It reports nothing about temperature either. So the honest position is that the root-versus-sub-call temperature question has no published answer. Anyone quoting you a tuned pair is reasoning from general practice, as we are about to.
Often not, and this collapses most of the question. Microsoft's reasoning-model documentation for the Azure OpenAI models states that "reasoning models other than GPT-6 Astra don't support the following parameters: temperature, top_p, presence_penalty, frequency_penalty, logprobs, top_logprobs, logit_bias, max_tokens." The whole GPT-5 family sits in that group, and the paper's own root is GPT-5 at medium reasoning.
Anthropic's thinking documentation draws a similar line and is explicit about where it applies. On its newer models, "non-default temperature, top_p, or top_k values return a 400 error on every request, regardless of whether thinking is used." On older models the restriction applies only while thinking is on, where "temperature and top_k are incompatible with thinking, and top_p is allowed at values between 0.95 and 1."
So the split is genuinely available only when at least one side is an open-weights or non-reasoning model. If your root is a frontier reasoning model, the practical question is not "which two temperatures" but "the provider default at the root, and what at the leaf." That is also why the root and the leaf are usually different models to begin with, as our post on choosing an RLM base model covers.
No, and this is the most common bad reason for turning the root dial down. Hosted endpoints do not offer determinism at any temperature. Thinking Machines Lab traced the cause: "the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies," and standard kernels are not batch invariant, so "when the batch size changes, each element in the batch can get different results." With batch-invariant kernels their 1,000 runs produced identical completions; the default implementation produced 80 unique ones. You cannot buy that property through an API.
The second reason is that a root model's run-to-run variance comes mostly from which decomposition it chooses, not from token-level jitter inside a chosen plan. Renze and Guven swept temperature from 0.0 to 1.6 across nine models and five prompting techniques and found that "changes in temperature from 0.0 to 1.0 do not have a statistically significant impact on LLM performance." That was multiple-choice problem solving, not REPL code, so do not over-read it. It does argue against expecting a large accuracy win from the root dial alone.
There is one real argument for keeping the root low. The root's output is a program, and a malformed cell wastes an entire turn. Work on code decoding splits code tokens into "challenging tokens that are difficult to predict and confident tokens that can be easily inferred," and lowers the temperature on the confident ones to avoid "the influence of tail randomness noises." Most of what an RLM root writes is the confident kind: slice a string, run a regex, loop over chunks, call llm_query. If your provider lets you set it, there is no upside to sampling that code hot. If you want stable evaluation numbers, get them from a fixed harness and repeated runs, which our guide to evaluating an RLM sets out.
Only when you sample the same sub-call more than once and aggregate the answers. That is the self-consistency setup: Wang and colleagues "first sample a diverse set of reasoning paths instead of only taking the greedy one, and then select the most consistent answer by marginalizing out the sampled reasoning paths," reporting gains such as 17.9 points on GSM8K. Diversity is the whole mechanism, and it requires repeated samples of one question.
The default RLM fan-out is not that shape. Sub-calls go out over distinct chunks, so each one asks a different question and each answer is used exactly once. Diversity between those calls buys nothing, because nothing compares them. Raising temperature there only adds noise to extraction, and bad extractions are precisely what the root aggregates into a confidently wrong final answer. Our post on RLM hallucination of REPL output describes how that failure looks from the trajectory.
The rule that follows is simple: temperature should follow the aggregation, not the recursion depth. A sub-call whose result is used once belongs at the low end. A sub-call you issue several times over the same slice and then majority-vote belongs at a sampling temperature. Budget for it, because k samples cost k times as much, and sub-call cost is already the part of an RLM bill that blows out on the tail.
The reference implementation supports different temperatures at the root and in sub-calls, and for most workloads a split is reasonable: provider default at the root, low for single-use sub-calls. What no source supports is a tuned pair of numbers. The RLM paper and its reproduction never report a temperature at all, and the models most often used as roots reject the parameter, so the decision is really about the leaves. Set the leaf temperature by what you do with the answer. One use means low. A vote means diversity. Everything else is guesswork with a decimal point on it.
They are two constructor arguments on the RLM class that carry sampling settings such as temperature, top_p, max_tokens, and seed. The first applies to the root model at depth 0, the second to depth-1 sub-LLM calls. Both default to None, so an unconfigured run uses provider defaults on both sides.
The reference library only forwards sampling args through its OpenAI-compatible client. The Anthropic client builds its request from model, max_tokens, messages, and system, and the Gemini client only passes a system instruction. Anything else you set on those backends is dropped without an error.
No. Thinking Machines Lab showed that hosted endpoints vary because server load changes the batch size, and standard kernels are not batch invariant. Their default setup produced 80 unique completions across 1,000 runs at the same settings. Batch-invariant kernels fixed it, but that is not something an API exposes.
Only for sub-calls you sample several times over the same slice and then aggregate. Self-consistency works by comparing diverse reasoning paths for one question. The normal RLM fan-out asks a different question of each chunk, so there is nothing to compare and the extra samples only multiply cost.
The paper never states one. It says GPT-5 ran with medium reasoning and default sampling parameters, and that Qwen3-Coder ran with the sampling parameters from its model card, which are temperature 0.7, top_p 0.8, top_k 20, and repetition penalty 1.05. Neither is a root-versus-sub-call choice.