A small sub-call model finds the right answer, then prints it again and again until the token limit. That loop has a well-studied cause, and most of the fixes live in the sampler.
The short version: a small sub-call model repeats itself for the same reason any language model does: once a sentence is in the context, the model rates it as more likely, and each copy raises the odds of the next. Greedy or low-temperature decoding, a missing stop signal, and no output cap turn that bias into a loop that runs until the token limit. The RLM papers do not study this inside llm_query, but a Qwen3-4B report on the reference repo shows it, and the fixes are sampling settings, an output cap, and a check in the REPL.
The clearest public case is issue #146 on the reference RLM repository, opened in April 2026. The reporter used Qwen3-4B as the sub-call model and called llm_query with a short extraction prompt over a finance table. The model first invented context that was not in the table, including a date of "35 December 2019." It then found the right values and, in the reporter's words, "enters an infinite repetition loop" of Answer: £1.2 million (2018), £0.8 million (2019) "repeated hundreds of times until max tokens is exhausted."
Note the order. The correct answer appears first, and the loop starts after it. The model knew the answer. It did not know how to stop.
That is a different failure from the ones our post on RLM limitations and failure modes covers. The RLM paper does describe repetition, but at the root, not inside a sub-call. In a worked trajectory in its appendix, the RLM paper says Qwen3-Coder "will continue to repeatedly verify its answers" and re-ran its whole sub-call process five times before it returned a wrong answer. The reproduction study by Daren Wang reports the same pattern at depth 2: models "continuously launching new API calls to re-verify already extracted answers." Both are loops of whole steps by large root models. Neither paper measures token-level loops inside a single sub-call response.
The research on this predates RLMs by years, and it points at two causes: the decoding rule and the model itself.
On decoding, Holtzman et al. found that "using likelihood as a decoding objective leads to text that is bland and strangely repetitive." Their key figure shows that "the probability of a repeated phrase increases with each repetition, creating a positive feedback loop," and they found this held "for the vast majority of phrases we tested." In their GPT-2 experiments, "sampling with temperatures lower than 0.9 severely increase repetition."
Xu et al. put a name on the mechanism: a self-reinforcement effect. "The more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence," and sentences that start with a high probability reinforce faster. Their analysis explains why the confident sentence is the one that loops: greedy decoding picks high-likelihood sentences, and high-likelihood sentences are the ones most likely to repeat. That fits issue #146 well. Answer: £1.2 million (2018), £0.8 million (2019) is exactly the kind of short, confident line the model rates highly.
On the model side, Welleck et al. argue that the training objective is also at fault: likelihood training gives "too much probability to sequences containing repeats and frequent words." So a better sampler reduces the loop, but it does not remove the bias that causes it.
Not as far as the evidence goes. We found no published measure of repetition rate against model size for RLM sub-calls, and the research above covers models of many sizes. The model Xu et al. trained has 750M parameters, and Holtzman et al. studied GPT-2.
What is specific to small models in practice is who controls the sampler. Frontier sub-call models such as GPT-5-mini, which the RLM paper used for sub-calls, run behind a provider API with the provider's defaults. Small open-weights models such as Qwen3-4B or Qwen3-8B usually run on your own server, and then the decoding settings are your job. The model vendor knows this failure well. The Qwen3-4B model card says of thinking mode: "DO NOT use greedy decoding, as it can lead to performance degradation and endless repetitions." It also says: "If you encounter significant endless repetitions," set presence_penalty to 1.5.
Small models also have less headroom. The RLM authors report that "smaller models like Qwen3-8B" struggled as RLMs without enough coding ability, which is why they post-trained RLM-Qwen3-8B. Our guide to choosing a base model for an RLM covers that side. Sub-calls, the paper says, need to behave roughly like "a general purpose reasoning model," and a 4B model with poor sampling settings falls short of that.
Probably not in the current library, although the reporter suggested it. Issue #146 proposed that llm_query "passes the prompt string directly to the model as a raw completion," so the model gets no end-of-turn signal. We read the reference library source (v0.1.3). Its OpenAI-compatible client wraps a string prompt as [{"role": "user", "content": prompt}] and sends it to chat.completions.create. A chat server such as vLLM applies the model's own chat template to that request; the vLLM server docs say the chat API needs a model with a chat template. So the turn markers are there.
The template theory is still worth a check if you use a custom backend that calls a raw completions endpoint. As of this writing, the issue is open and has no maintainer reply.
Two other library facts matter more. First, sub-call sampling is off by default: sub_sampling_args defaults to None, so the library sends no output cap and no penalty. Second, vLLM "applies generation_config.json from the Hugging Face model repository if it exists." For Qwen3-4B, the card says that file holds the thinking-mode settings. Your sub-call settings may therefore come from the model repo, not your code.
max_tokens in sub_sampling_args; the OpenAI client renames it to max_completion_tokens. This does not stop the loop, but it bounds the cost and the time.Temperature=0.7, TopP=0.8, TopK=20. This is a tension with our post on root and sub-call temperatures, which puts extraction sub-calls at the low end. The two fit together: go low, but not to zero, and follow the model card for small open models.presence_penalty and frequency_penalty act on "the generated text so far." repetition_penalty acts on "the prompt and the generated text so far." An extraction sub-call has to copy values from its prompt, so a penalty on prompt tokens pushes against the exact tokens you want. The original CTRL repetition penalty was defined on generated tokens only, with θ ≈ 1.2, and its authors note it "succeeds only if the model has learned a sufficiently reliable distribution." The Qwen card also warns that a high presence_penalty "may occasionally result in language mixing."stop strings, and generation ends when one appears.Track the rate. Log output length and a simple repeated-line count for each sub-call, and treat any response that hits the cap as a failure in your eval. Run the same slices more than once, since a loop may appear on one sample and not the next.
Repetition inside llm_query is ordinary neural text degeneration, not a new RLM failure. The research explains it: each repeat makes the next one more likely, and maximization-style decoding lets that feedback run. Small open models show it most in RLM setups because you, not a provider, own the sampler, and the defaults are often wrong for the job. Cap sub-call output, follow the model card instead of greedy decoding, avoid penalties on prompt tokens for extraction, and parse the answer in code before the root uses it.
No. The paper reports repetition at the root, where Qwen3-Coder re-verified its answer and re-ran its sub-call process five times. It does not measure token-level loops inside a single llm_query response. The only public report of that we found is issue #146 on the reference repository.
Be careful. In vLLM, repetition_penalty counts tokens in the prompt as well as the output, so it can push against copying values out of the context. presence_penalty and frequency_penalty only count generated tokens, which makes them a safer choice for extraction.
Not always. The Qwen3 model card warns against greedy decoding in thinking mode because it can cause endless repetitions, and Holtzman et al. found that low temperatures raise repetition in open-ended text. A low but nonzero temperature with the vendor top_p and top_k is the safer default.
No. sub_sampling_args defaults to None, so the OpenAI-compatible client sends no output cap and the server default applies. Pass max_tokens in sub_sampling_args and the client forwards it as max_completion_tokens.
It may help, but the research does not show that repetition goes away with scale. The self-reinforcement effect was found across model sizes. A bigger hosted model mostly helps because a provider tunes its sampling defaults for you.