Back to Blog
September 17, 2026

The short version: the model does not stop at the end of its code block. Nothing in a plain chat turn forces it to, so it keeps predicting tokens, writes what it expects the output to be, and sometimes commits to an answer in the same turn. The official RLM repo hit exactly this with Claude 4.6 models in PR #127. The fixes are mechanical: end the answer only from REPL state, stop or truncate generation at the closing fence, or run code through native tool calling.

What does the failure look like?

The clearest public report is PR #127 in the alexzhang13/rlm repository, opened in February 2026. It says Claude 4.6 Sonnet and Opus, "instead of generating a code block and stopping to wait for execution feedback," continued past the block, invented the execution output, reasoned over it, and wrote a FINAL() answer, all in one turn.

The example in the PR is a customer spending question. The model wrote a correct print(context) block. It then wrote made-up rows such as "Alice: $1200, Bob: $450, Kevin: $5200," none of which were in the data, and returned a ranked top 10 built on them. The harness did run the real code. But the parser found FINAL() in the same response and exited the loop, so the real output never reached the model.

A second report, issue #157, describes what its author calls the gpt-oss-120b analog. On a 25-file corpus of 126,349 input tokens, the run returned a summary of searches as its final result, and a downstream reviewer model turned that into a false AES-CBC finding on a file that uses AES-GCM. That issue also involves context pruning, so it is a mixed case, not a clean copy of #127.

Does the RLM paper or its reproduction discuss this?

Not directly. We searched the full text of the RLM paper (v3, May 2026) and found no discussion of fabricated REPL output. The closest passage is in Appendix B, the negative results. It says "distinguishing between a final answer and a thought is brittle for RLMs," that the answer was marked with FINAL() or FINAL_VAR() tags, and that a model sometimes "outputs its plan as a final answer." The authors say they "added minor safeguards" and do not describe them.

Appendix A adds a related number. In the Qwen3-Coder trajectories used to train RLM-Qwen3-8B, "16% of turns incorrectly used FINAL answers, and 13% of turns incorrectly called a variable from the REPL (i.e. FINAL_VAR) as a final answer," so the authors added a programmatic step to patch common template mistakes. That is answer-format misuse, not invented stdout.

The reproduction study by Daren Wang reports hallucination of a different kind. At depth 2, DeepSeek v3.2 answered a needle task with real nuclear magic numbers (2, 8, 20, 28, 50, 82, 126) instead of the planted value, which the author calls parametric hallucination. It also reports Kimi K2 runs where every RLM sample failed to parse because <thinking> tags wrapped the output. Neither is the root model writing fake execution results. Our post on RLM limitations and failure modes covers those broader problems.

Why does the model keep writing after the code fence?

Because a chat model ends its turn when it decides it is done, not when a code block closes. The Anthropic Messages API reference says models "normally stop when they have naturally completed their turn," with a stop_reason of end_turn. In an RLM loop, a closed ```repl fence is just more text. If the model predicts that output usually follows code, it writes the output.

The reference harness did not add a hard stop. Its OpenAI client sends the model, messages, and any sampling arguments you pass, and sets no stop sequence of its own. The parser in parsing.py pulls every block that matches ```repl\s*\n(.*?)\n``` out of the whole response and runs them all. Before May 2026 it also scanned the whole response text for FINAL(...). Put together, a model that writes code, fake output, and FINAL() in one turn gets its fake answer accepted.

The maintainer did not treat it as a single-model bug. Replying on PR #127, Alex Zhang called it "a weird problem that's quite model dependent" and said he wanted "less hacky solutions." That matches the paper's point that one prompt does not transfer cleanly across models, which our post on choosing a base model covers.

What did the official repo change?

PR #127 proposed ignoring text FINAL() whenever a response also held code blocks. It was not merged. Instead, PR #162, merged on 13 May 2026, removed text-based FINAL parsing. The REPL now holds an answer dict with "content" and "ready" keys, and the run ends only when code sets answer["ready"] = True during a real execution. The PR says it follows DSPy.RLM and the verifiers RLMEnv.

A contributor then wrote regression tests for a repl block followed by a hallucinated trailing FINAL(...), reported that they pass on main, and the PR author closed #127 as resolved by the dict change. The current system prompt also says "execute one ```repl``` block every turn," and the first user turn warns: "Look at the context first; do not provide a final answer yet."

Two gaps remain, from our reading of the code rather than a filed bug. The assistant message, fake output included, is still appended to history before the real REPL output. And because every block in a response runs, a model could write invented numbers into answer["content"] in a second block of the same turn.

How do you stop it in your own harness?

  1. End runs from REPL state, never from text. Copy the answer["ready"] pattern. Prose after a code block then cannot end the run.
  2. Set a stop sequence at the closing fence. A commenter on #127 suggested this. Anthropic returns stop_reason: "stop_sequence" and the matched string in stop_sequence, per the stop reasons guide. OpenAI Chat Completions accepts up to 4 stop sequences, notes that "the returned text will not contain the stop sequence," and marks stop as not supported on o3 and o4-mini. If the fence is stripped, add it back, or the repo regex above will not match.
  3. Truncate on the client. Where a model rejects stop sequences, cut the response after the first closing fence before you run code or save history. The model then never sees its own fake output.
  4. Use native tool calling for execution. Expose code execution as a tool. Anthropic ends the turn with stop_reason tool_use, and the guide says to "run the tool and return the result." OpenAI reports finish_reason tool_calls. The provider enforces the turn boundary, not your regex.
  5. Allow one block per turn. Run the first block and drop the rest, so a second block cannot submit an answer built on invented output.
  6. Log trajectories and look for text after the fence. Anything that reads like output right after ``` in an assistant turn is a fabrication signal. Treat the REPL like any exec surface, as our sandboxing post describes.

The bottom line

Fabricated REPL output is a turn-boundary bug, not a reasoning bug. The model predicts the output because nothing stops it. The RLM paper does not document this case, and the reproduction reports other hallucinations. The official repo saw it with Claude 4.6 models and fixed the worst part by ending runs only from REPL state. Add a stop sequence or client truncation at the closing fence, or move execution into tool calls, and the model has to wait for real output.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025 (v3 May 2026). Appendix B brittle final-answer passage, Appendix A FINAL misuse rates (16% and 13%), Algorithm 1 termination on the Final variable; no discussion of fabricated REPL output. arxiv.org/abs/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Depth-2 parametric hallucination, formatting collapse, and Kimi K2 thinking-tag parse failures. arxiv.org/abs/2603.02615
  3. alexzhang13/rlm. "Fix hallucinated FINAL() answers with Claude 4.6 models." GitHub PR #127, February 2026. Failure description, customer-data example, stop-token suggestion, maintainer reply, and closure as superseded. github.com/alexzhang13/rlm/pull/127
  4. alexzhang13/rlm. "New dict-based Final Answer format." GitHub PR #162, merged May 13, 2026. The answer dict with content and ready keys that replaced text FINAL parsing. github.com/alexzhang13/rlm/pull/162
  5. alexzhang13/rlm. "compress_corpus produces fabricated 'search summary' output instead of compressed code at >100K input tokens (gpt-oss-120b via Cerebras)." GitHub issue #157, May 2026. Second fabrication report and the downstream false AES-CBC finding. github.com/alexzhang13/rlm/issues/157
  6. alexzhang13/rlm. "parsing.py." GitHub source, accessed September 2026. The repl code-block regex and format_iteration history handling. github.com/alexzhang13/rlm/blob/main/rlm/utils/parsing.py
  7. alexzhang13/rlm. "prompts.py." GitHub source, accessed September 2026. One block per turn instruction and first-turn safeguard text. github.com/alexzhang13/rlm/blob/main/rlm/utils/prompts.py
  8. alexzhang13/rlm. "clients/openai.py." GitHub source, accessed September 2026. Chat Completions call with no default stop sequence. github.com/alexzhang13/rlm/blob/main/rlm/clients/openai.py
  9. Anthropic. "Messages API reference." Claude Platform docs, accessed September 2026. stop_sequences parameter and end_turn behavior. platform.claude.com/docs/en/api/messages
  10. Anthropic. "Handling stop reasons." Claude Platform docs, accessed September 2026. stop_sequence and tool_use stop reasons. platform.claude.com/docs/en/build-with-claude/handling-stop-reasons
  11. OpenAI. "Create chat completion." API reference, accessed September 2026. stop parameter limits, stripped stop text, o3 and o4-mini exclusion, tool_calls finish reason. developers.openai.com/api/reference/resources/chat/subresources/completions/methods/create
FAQ

Frequently asked questions

Is this the same as the model hallucinating facts from memory?

No. Parametric hallucination is when a model answers from its training data instead of the input, which the reproduction study saw at depth 2. Fabricated REPL output is a turn-boundary failure: the model writes what it expects the code to print. The fixes differ, because the second one is solved by stopping generation at the right place.

Which models are known to do this?

The public report in the official rlm repo names Claude 4.6 Sonnet and Opus. A separate issue describes a similar pattern with gpt-oss-120b on a large corpus, mixed with context pruning. The maintainer called the problem model dependent, so test any root model you plan to use.

Will a stronger system prompt fix it?

Not reliably. The old reference prompt already told models not to use FINAL tags until the task was complete, and the Claude 4.6 behavior still appeared. The repo fixed it by changing how a run ends, not by adding more instructions.

Does upgrading the rlm library solve the problem?

It removes the worst case. Since PR #162, text FINAL answers are ignored and only answer ready set in real code ends a run. The fabricated text can still enter the message history, so a stop sequence or client-side truncation is still worth adding.

Can I use stop sequences with reasoning models?

It depends on the provider and model. The OpenAI Chat Completions reference says the stop parameter is not supported on o3 and o4-mini. In that case, truncate after the closing fence on the client or use native tool calling.