Back to Blog
September 24, 2026

The short version: a small sub-call model repeats itself for the same reason any language model does: once a sentence is in the context, the model rates it as more likely, and each copy raises the odds of the next. Greedy or low-temperature decoding, a missing stop signal, and no output cap turn that bias into a loop that runs until the token limit. The RLM papers do not study this inside llm_query, but a Qwen3-4B report on the reference repo shows it, and the fixes are sampling settings, an output cap, and a check in the REPL.

What does the repetition look like inside an RLM?

The clearest public case is issue #146 on the reference RLM repository, opened in April 2026. The reporter used Qwen3-4B as the sub-call model and called llm_query with a short extraction prompt over a finance table. The model first invented context that was not in the table, including a date of "35 December 2019." It then found the right values and, in the reporter's words, "enters an infinite repetition loop" of Answer: £1.2 million (2018), £0.8 million (2019) "repeated hundreds of times until max tokens is exhausted."

Note the order. The correct answer appears first, and the loop starts after it. The model knew the answer. It did not know how to stop.

That is a different failure from the ones our post on RLM limitations and failure modes covers. The RLM paper does describe repetition, but at the root, not inside a sub-call. In a worked trajectory in its appendix, the RLM paper says Qwen3-Coder "will continue to repeatedly verify its answers" and re-ran its whole sub-call process five times before it returned a wrong answer. The reproduction study by Daren Wang reports the same pattern at depth 2: models "continuously launching new API calls to re-verify already extracted answers." Both are loops of whole steps by large root models. Neither paper measures token-level loops inside a single sub-call response.

Why do language models fall into loops at all?

The research on this predates RLMs by years, and it points at two causes: the decoding rule and the model itself.

On decoding, Holtzman et al. found that "using likelihood as a decoding objective leads to text that is bland and strangely repetitive." Their key figure shows that "the probability of a repeated phrase increases with each repetition, creating a positive feedback loop," and they found this held "for the vast majority of phrases we tested." In their GPT-2 experiments, "sampling with temperatures lower than 0.9 severely increase repetition."

Xu et al. put a name on the mechanism: a self-reinforcement effect. "The more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence," and sentences that start with a high probability reinforce faster. Their analysis explains why the confident sentence is the one that loops: greedy decoding picks high-likelihood sentences, and high-likelihood sentences are the ones most likely to repeat. That fits issue #146 well. Answer: £1.2 million (2018), £0.8 million (2019) is exactly the kind of short, confident line the model rates highly.

On the model side, Welleck et al. argue that the training objective is also at fault: likelihood training gives "too much probability to sequences containing repeats and frequent words." So a better sampler reduces the loop, but it does not remove the bias that causes it.

Is this specific to small models?

Not as far as the evidence goes. We found no published measure of repetition rate against model size for RLM sub-calls, and the research above covers models of many sizes. The model Xu et al. trained has 750M parameters, and Holtzman et al. studied GPT-2.

What is specific to small models in practice is who controls the sampler. Frontier sub-call models such as GPT-5-mini, which the RLM paper used for sub-calls, run behind a provider API with the provider's defaults. Small open-weights models such as Qwen3-4B or Qwen3-8B usually run on your own server, and then the decoding settings are your job. The model vendor knows this failure well. The Qwen3-4B model card says of thinking mode: "DO NOT use greedy decoding, as it can lead to performance degradation and endless repetitions." It also says: "If you encounter significant endless repetitions," set presence_penalty to 1.5.

Small models also have less headroom. The RLM authors report that "smaller models like Qwen3-8B" struggled as RLMs without enough coding ability, which is why they post-trained RLM-Qwen3-8B. Our guide to choosing a base model for an RLM covers that side. Sub-calls, the paper says, need to behave roughly like "a general purpose reasoning model," and a 4B model with poor sampling settings falls short of that.

Is the chat template the cause?

Probably not in the current library, although the reporter suggested it. Issue #146 proposed that llm_query "passes the prompt string directly to the model as a raw completion," so the model gets no end-of-turn signal. We read the reference library source (v0.1.3). Its OpenAI-compatible client wraps a string prompt as [{"role": "user", "content": prompt}] and sends it to chat.completions.create. A chat server such as vLLM applies the model's own chat template to that request; the vLLM server docs say the chat API needs a model with a chat template. So the turn markers are there.

The template theory is still worth a check if you use a custom backend that calls a raw completions endpoint. As of this writing, the issue is open and has no maintainer reply.

Two other library facts matter more. First, sub-call sampling is off by default: sub_sampling_args defaults to None, so the library sends no output cap and no penalty. Second, vLLM "applies generation_config.json from the Hugging Face model repository if it exists." For Qwen3-4B, the card says that file holds the thinking-mode settings. Your sub-call settings may therefore come from the model repo, not your code.

Which fixes actually work?

  1. Cap the output. A sub-call that extracts two numbers does not need thousands of tokens. Pass max_tokens in sub_sampling_args; the OpenAI client renames it to max_completion_tokens. This does not stop the loop, but it bounds the cost and the time.
  2. Use the vendor sampling settings, not greedy. For Qwen3 in non-thinking mode the card suggests Temperature=0.7, TopP=0.8, TopK=20. This is a tension with our post on root and sub-call temperatures, which puts extraction sub-calls at the low end. The two fit together: go low, but not to zero, and follow the model card for small open models.
  3. Prefer presence or frequency penalty to repetition penalty for extraction. In vLLM, presence_penalty and frequency_penalty act on "the generated text so far." repetition_penalty acts on "the prompt and the generated text so far." An extraction sub-call has to copy values from its prompt, so a penalty on prompt tokens pushes against the exact tokens you want. The original CTRL repetition penalty was defined on generated tokens only, with θ ≈ 1.2, and its authors note it "succeeds only if the model has learned a sufficiently reliable distribution." The Qwen card also warns that a high presence_penalty "may occasionally result in language mixing."
  4. Ask for a short, closed format. A prompt that asks for one JSON object or one line gives the model a clear end point. You can then add a stop string at the end of that format. vLLM exposes stop strings, and generation ends when one appears.
  5. Clean the result in the REPL. The root model gets the sub-call output as a Python string. Parse the first valid answer, or cut repeated lines, before you store it. A loop that the root never sees cannot pollute the aggregate.

Track the rate. Log output length and a simple repeated-line count for each sub-call, and treat any response that hits the cap as a failure in your eval. Run the same slices more than once, since a loop may appear on one sample and not the next.

The bottom line

Repetition inside llm_query is ordinary neural text degeneration, not a new RLM failure. The research explains it: each repeat makes the next one more likely, and maximization-style decoding lets that feedback run. Small open models show it most in RLM setups because you, not a provider, own the sampler, and the defaults are often wrong for the job. Cap sub-call output, follow the model card instead of greedy decoding, avoid penalties on prompt tokens for extraction, and parse the answer in code before the root uses it.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, v3 May 2026. Root-level repeated verification by Qwen3-Coder (Appendix E), GPT-5-mini sub-calls, Qwen3-8B struggles, sub-call role. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Redundant sub-call loops and re-verification at depth 2. arxiv.org/abs/2603.02615
  3. YWenxi. "llm_query produces repetitive output with Qwen3-4B." alexzhang13/rlm issue #146, GitHub, April 2026. Public report of a sub-call repetition loop with Qwen3-4B. github.com/alexzhang13/rlm/issues/146
  4. Zhang, A. L., et al. "rlm: inference library for Recursive Language Models." GitHub, v0.1.3. OpenAI client message wrapping, sub_sampling_args default, max_tokens rename. github.com/alexzhang13/rlm
  5. Holtzman, A., Buys, J., Du, L., Forbes, M., Choi, Y. "The Curious Case of Neural Text Degeneration." ICLR 2020, arXiv:1904.09751. Repetition feedback loop and effect of low temperature. arxiv.org/abs/1904.09751
  6. Xu, J., Liu, X., Yan, J., Cai, D., Li, H., Li, J. "Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation." NeurIPS 2022, arXiv:2206.02369. Self-reinforcement effect of sentence-level repetition. arxiv.org/abs/2206.02369
  7. Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., Weston, J. "Neural Text Generation with Unlikelihood Training." arXiv:1908.04319, August 2019. Likelihood training over-weights repeats. arxiv.org/abs/1908.04319
  8. Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., Socher, R. "CTRL: A Conditional Transformer Language Model for Controllable Generation." arXiv:1909.05858, September 2019. Original penalized sampling (repetition penalty) definition. arxiv.org/abs/1909.05858
  9. Qwen Team. "Qwen3-4B" model card. Hugging Face, accessed September 2026. Greedy decoding warning, presence_penalty guidance, recommended sampling settings. huggingface.co/Qwen/Qwen3-4B
  10. vLLM Project. "SamplingParams" source and "OpenAI-Compatible Server" docs. GitHub, accessed September 2026. Penalty definitions, stop strings, generation_config defaults, chat template requirement. github.com/vllm-project/vllm/blob/main/vllm/sampling_params.py
FAQ

Frequently asked questions

Does the RLM paper report repetition loops inside sub-calls?

No. The paper reports repetition at the root, where Qwen3-Coder re-verified its answer and re-ran its sub-call process five times. It does not measure token-level loops inside a single llm_query response. The only public report of that we found is issue #146 on the reference repository.

Should I set repetition_penalty on an extraction sub-call?

Be careful. In vLLM, repetition_penalty counts tokens in the prompt as well as the output, so it can push against copying values out of the context. presence_penalty and frequency_penalty only count generated tokens, which makes them a safer choice for extraction.

Is temperature 0 safe for small sub-call models?

Not always. The Qwen3 model card warns against greedy decoding in thinking mode because it can cause endless repetitions, and Holtzman et al. found that low temperatures raise repetition in open-ended text. A low but nonzero temperature with the vendor top_p and top_k is the safer default.

Does the reference RLM library cap sub-call output by default?

No. sub_sampling_args defaults to None, so the OpenAI-compatible client sends no output cap and the server default applies. Pass max_tokens in sub_sampling_args and the client forwards it as max_completion_tokens.

Will a bigger sub-call model stop the loops?

It may help, but the research does not show that repetition goes away with scale. The self-reinforcement effect was found across model sizes. A bigger hosted model mostly helps because a provider tunes its sampling defaults for you.