Back to Blog
October 10, 2026

The short version: start with the library default, then measure. The reference RLM library sets max_iterations=30. DSPy names the setting max_iters and sets it to 20. The RLM paper tests no values for this cap, so no number has research behind it. Log the turns your tasks use, set the cap a little above what successful runs need, and use a budget or timeout limit to control cost.

What does max_iterations control?

It caps the turns of the root loop. In one turn, the root model writes code, the REPL runs the code, and the output goes back to the model. The loop ends when the model submits a final answer or when the turn count reaches the cap.

This is a different setting from recursion depth. max_depth controls how many levels of child RLMs can exist. max_iterations controls how many turns each level can take. Our post on how deep an RLM should recurse covers depth. This post covers only the turn cap.

The RLM paper writes its core loop as while True, with no cap in the algorithm. The text then adds that "we typically want to limit the iterations at any level of recursion irrespective." The paper also notes that one root iteration "can launch arbitrarily many sub-calls." So the turn cap limits the length of the loop. It does not limit the work inside one turn.

What are the defaults in the two libraries?

The parameter name is not the same in both libraries, and the defaults differ.

  • Reference library. The constructor in rlm/core/rlm.py has max_iterations: int = 30. The docstring says "The maximum number of iterations of the RLM." The loop is for i in range(self.max_iterations).
  • DSPy. The constructor in dspy/predict/rlm.py has max_iters: int = 20. The docstring says "Maximum REPL interaction iterations." The same constructor has max_llm_calls: int = 50, which is a separate cap on sub-LLM calls.

Both libraries show the cap to the model. The reference library starts each turn with a user message built from the template Turn {iter_1}/{max_iter}: in rlm/utils/prompts.py. DSPy passes an iteration field that it describes as the "Current iteration number (1-indexed) out of max_iters." So the cap is also a signal. The model can see how many turns it has left.

The README of the reference library does not document max_iterations. You find the default only in the source. Both repositories change often, so read the source for the release you install. Our comparison of dspy.RLM and the reference library lists the other constructor differences.

What happens when the loop hits the cap?

Neither library raises an error. Each one forces an answer, and that is the risk.

In the reference library, the code after the loop calls _default_answer. The method appends one message to the history: "Please provide a final answer to the user's question based on the information provided." It then makes one more model call and returns that text as the response. The root model writes this answer from its message history. No code runs, so the answer does not come from a REPL variable.

In DSPy, the loop falls through to an extract predictor. The DSPy guide says the step "reads the variable metadata and the full REPL history and produces the signature’s output fields directly." The source logs a warning, "RLM reached max iterations, using extract to get final output," and sets final_reasoning to "Extract forced final output."

The paper shows how a forced answer can fail. In its Example E.2, RLM with Qwen3-Coder built the answer in the REPL in the first iteration. It then repeated the same work for steps 6 to 11, "before finally returning an answer after being prompted to provide a final answer." The paper says the returned answer was "the root LM generating an answer," and that it was wrong. The correct answer was in a variable, and the model never returned it.

A run that hits the cap is therefore a failed run with an answer attached. Flag it. Do not count it as a normal result.

Does the research recommend a value?

No. The paper sweeps recursion depth from 0 to 3, but it reports no sweep of the turn cap. In the text we read, the paper does not state the cap for its main benchmark runs. It gives one number, in the training appendix: for the MRCRv2 experiment the authors "set the max number of RLM iterations to 20." That is a training configuration. It is not a recommendation for inference.

The paper does describe the shape of the problem. It finds "that the median RLM run is cheaper than the median base model run, but more expensive on average due to outlier trajectories where the RLM struggles to find an answer." It says the longest runs happen infrequently and "can be early-stopped with timeout logic."

The reproduction study does not report an iteration cap. It names a failure mode, "Performative Reasoning and Endless Verification." In one OOLONG run at depth 2, DeepSeek v3.2 spent 741.5 seconds to generate 11,715 tokens. The author asks for "better stopping mechanisms within the REPL environment to prevent redundant loops."

The rest of this article is engineering reasoning. It is not a research result.

How do you choose a value for your tasks?

Measure the turns your tasks use, then set the cap from that data.

  1. Start at the default. Use 30 in the reference library or 20 in DSPy.
  2. Log every run. In the reference library, pass logger=RLMLogger(). The README says completion.metadata then holds "the full trajectory (run config + all iterations and sub-calls)." In DSPy, each Prediction has a trajectory list.
  3. Count turns for each task. Use a sample that matches your real workload. Keep correct runs and wrong runs in separate groups.
  4. Find the cap hits. In DSPy, look for final_reasoning equal to "Extract forced final output." In the reference library, _default_answer logs one extra iteration with an empty code_blocks list.
  5. Set the cap above the turn count of your correct runs. Leave a small margin. If almost all correct runs finish in 8 turns, a cap of 30 mostly pays for failed runs.
  6. Raise the cap only with evidence. If correct runs often finish on the last turns, the cap is too low. If cap hits come from repeated work, a higher cap will not help. Fix the prompt or change the model.

Do the count again when you change the base model. The paper found that Qwen3-Coder made "hundreds to thousands of recursive sub-calls for a single simple task, while GPT-5 makes on the order of ten." Turn counts can differ between models in the same way. Our guide on how to evaluate an RLM covers what else to log.

Is the iteration cap enough to control cost?

No. A turn has no fixed price, because one turn can start many sub-calls. Use the cap to stop loops and a second limit to stop spend.

  • Reference library. The constructor has max_budget in USD, max_timeout in seconds, max_tokens and max_errors. All four default to None, so they are off until you set them. Each one raises an exception when exceeded. The docstring says max_budget "requires cost-tracking backend (e.g., OpenRouter)."
  • DSPy. max_llm_calls counts each prompt in a batch as one call. Past the limit, the call fails with "LLM call limit exceeded," and the message says "Use Python code for aggregation instead of making more LLM calls."

Depth multiplies the turn cap in the reference library. When a parent starts a child RLM, the code passes max_iterations=self.max_iterations. Each child gets the full turn count again. At max_depth=2, each child that the root starts can take up to 30 turns of its own. The children receive the remaining budget and the remaining timeout, so those two limits do apply to the full tree.

History length is the other cost. The reference library sends the full message history to the root model on each turn, so later turns cost more input tokens than early turns. The library has an optional compaction flag that summarizes the root history near the context limit. It is off by default.

The bottom line

No published result gives a correct value for the turn cap. The reference library uses 30. DSPy uses 20 and calls the setting max_iters. The paper used 20 in one training experiment and tested no other values. Keep the default, log the turns your tasks use, and set the cap a little above what correct runs need. Treat each cap hit as a failure to inspect, because both libraries return a forced answer with no error. Control spend with max_budget, max_timeout or max_llm_calls.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, HTML version. The while-True loop in Algorithm 1, the note on limiting iterations, the MRCRv2 training cap of 20, Example E.2 and the outlier trajectory findings. arxiv.org/html/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, HTML version. The endless verification failure mode, the 741.5 second run and the call for better stopping mechanisms. arxiv.org/html/2603.02615
  3. Zhang, A. L., et al. "rlm/core/rlm.py" source. max_iterations default of 30, the root loop, _default_answer, the max_budget, max_timeout, max_tokens and max_errors limits, and the child RLM constructor call. raw.githubusercontent.com/alexzhang13/rlm/main/rlm/core/rlm.py
  4. Zhang, A. L., et al. "rlm/utils/prompts.py" source. The per-turn user prompt template that shows the turn number and the cap to the model. raw.githubusercontent.com/alexzhang13/rlm/main/rlm/utils/prompts.py
  5. Zhang, A. L., et al. "rlm" GitHub repository, README. RLMLogger and the trajectory in completion.metadata. Has no entry for max_iterations. raw.githubusercontent.com/alexzhang13/rlm/main/README.md
  6. Stanford NLP. "dspy/predict/rlm.py" source. max_iters default of 20, max_llm_calls default of 50, the call limit error and the extract fallback. raw.githubusercontent.com/stanfordnlp/dspy/main/dspy/predict/rlm.py
  7. Stanford NLP. "RLM: exploring large contexts with code." DSPy documentation source. Description of max_iters, max_llm_calls and the extract fallback. raw.githubusercontent.com/stanfordnlp/dspy/main/docs/docs/diving-deeper/rlm.md
FAQ

Frequently asked questions

What is the default max_iterations in the reference RLM library?

The default is 30. The constructor in rlm/core/rlm.py has max_iterations: int = 30, and the root loop runs for at most that many turns. The README does not document the setting, so read the source for the release you install.

Does dspy.RLM have a max_iterations parameter?

No. DSPy calls the same idea max_iters, and its default is 20. DSPy also has max_llm_calls, default 50, which caps sub-LLM calls across one run. The two limits are independent.

Does an RLM raise an error when it runs out of iterations?

No, not in these two libraries. The reference library asks the root model for a final answer in one more call and returns that text. DSPy runs an extract predictor over the REPL history and fills the output fields. Both return a normal result, so you must detect the cap hit yourself.

Is max_iterations the same as max_depth?

No. max_depth sets how many levels of child RLMs can exist. max_iterations sets how many REPL turns each level can take. In the reference library a child RLM gets the same max_iterations value as its parent.

Did the RLM paper test different iteration limits?

No. The paper sweeps recursion depth, but it reports no sweep of the iteration cap. The one stated value is in the training appendix, where the MRCRv2 experiment set the maximum number of RLM iterations to 20.