The root model writes a code block, then keeps going: it invents the output and commits to an answer before the REPL runs. Here is why that happens and how to stop it.
The short version: the model does not stop at the end of its code block. Nothing in a plain chat turn forces it to, so it keeps predicting tokens, writes what it expects the output to be, and sometimes commits to an answer in the same turn. The official RLM repo hit exactly this with Claude 4.6 models in PR #127. The fixes are mechanical: end the answer only from REPL state, stop or truncate generation at the closing fence, or run code through native tool calling.
The clearest public report is PR #127 in the alexzhang13/rlm repository, opened in February 2026. It says Claude 4.6 Sonnet and Opus, "instead of generating a code block and stopping to wait for execution feedback," continued past the block, invented the execution output, reasoned over it, and wrote a FINAL() answer, all in one turn.
The example in the PR is a customer spending question. The model wrote a correct print(context) block. It then wrote made-up rows such as "Alice: $1200, Bob: $450, Kevin: $5200," none of which were in the data, and returned a ranked top 10 built on them. The harness did run the real code. But the parser found FINAL() in the same response and exited the loop, so the real output never reached the model.
A second report, issue #157, describes what its author calls the gpt-oss-120b analog. On a 25-file corpus of 126,349 input tokens, the run returned a summary of searches as its final result, and a downstream reviewer model turned that into a false AES-CBC finding on a file that uses AES-GCM. That issue also involves context pruning, so it is a mixed case, not a clean copy of #127.
Not directly. We searched the full text of the RLM paper (v3, May 2026) and found no discussion of fabricated REPL output. The closest passage is in Appendix B, the negative results. It says "distinguishing between a final answer and a thought is brittle for RLMs," that the answer was marked with FINAL() or FINAL_VAR() tags, and that a model sometimes "outputs its plan as a final answer." The authors say they "added minor safeguards" and do not describe them.
Appendix A adds a related number. In the Qwen3-Coder trajectories used to train RLM-Qwen3-8B, "16% of turns incorrectly used FINAL answers, and 13% of turns incorrectly called a variable from the REPL (i.e. FINAL_VAR) as a final answer," so the authors added a programmatic step to patch common template mistakes. That is answer-format misuse, not invented stdout.
The reproduction study by Daren Wang reports hallucination of a different kind. At depth 2, DeepSeek v3.2 answered a needle task with real nuclear magic numbers (2, 8, 20, 28, 50, 82, 126) instead of the planted value, which the author calls parametric hallucination. It also reports Kimi K2 runs where every RLM sample failed to parse because <thinking> tags wrapped the output. Neither is the root model writing fake execution results. Our post on RLM limitations and failure modes covers those broader problems.
Because a chat model ends its turn when it decides it is done, not when a code block closes. The Anthropic Messages API reference says models "normally stop when they have naturally completed their turn," with a stop_reason of end_turn. In an RLM loop, a closed ```repl fence is just more text. If the model predicts that output usually follows code, it writes the output.
The reference harness did not add a hard stop. Its OpenAI client sends the model, messages, and any sampling arguments you pass, and sets no stop sequence of its own. The parser in parsing.py pulls every block that matches ```repl\s*\n(.*?)\n``` out of the whole response and runs them all. Before May 2026 it also scanned the whole response text for FINAL(...). Put together, a model that writes code, fake output, and FINAL() in one turn gets its fake answer accepted.
The maintainer did not treat it as a single-model bug. Replying on PR #127, Alex Zhang called it "a weird problem that's quite model dependent" and said he wanted "less hacky solutions." That matches the paper's point that one prompt does not transfer cleanly across models, which our post on choosing a base model covers.
PR #127 proposed ignoring text FINAL() whenever a response also held code blocks. It was not merged. Instead, PR #162, merged on 13 May 2026, removed text-based FINAL parsing. The REPL now holds an answer dict with "content" and "ready" keys, and the run ends only when code sets answer["ready"] = True during a real execution. The PR says it follows DSPy.RLM and the verifiers RLMEnv.
A contributor then wrote regression tests for a repl block followed by a hallucinated trailing FINAL(...), reported that they pass on main, and the PR author closed #127 as resolved by the dict change. The current system prompt also says "execute one ```repl``` block every turn," and the first user turn warns: "Look at the context first; do not provide a final answer yet."
Two gaps remain, from our reading of the code rather than a filed bug. The assistant message, fake output included, is still appended to history before the real REPL output. And because every block in a response runs, a model could write invented numbers into answer["content"] in a second block of the same turn.
answer["ready"] pattern. Prose after a code block then cannot end the run.stop_reason: "stop_sequence" and the matched string in stop_sequence, per the stop reasons guide. OpenAI Chat Completions accepts up to 4 stop sequences, notes that "the returned text will not contain the stop sequence," and marks stop as not supported on o3 and o4-mini. If the fence is stripped, add it back, or the repo regex above will not match.stop_reason tool_use, and the guide says to "run the tool and return the result." OpenAI reports finish_reason tool_calls. The provider enforces the turn boundary, not your regex.``` in an assistant turn is a fabrication signal. Treat the REPL like any exec surface, as our sandboxing post describes.Fabricated REPL output is a turn-boundary bug, not a reasoning bug. The model predicts the output because nothing stops it. The RLM paper does not document this case, and the reproduction reports other hallucinations. The official repo saw it with Claude 4.6 models and fixed the worst part by ending runs only from REPL state. Add a stop sequence or client truncation at the closing fence, or move execution into tool calls, and the model has to wait for real output.
No. Parametric hallucination is when a model answers from its training data instead of the input, which the reproduction study saw at depth 2. Fabricated REPL output is a turn-boundary failure: the model writes what it expects the code to print. The fixes differ, because the second one is solved by stopping generation at the right place.
The public report in the official rlm repo names Claude 4.6 Sonnet and Opus. A separate issue describes a similar pattern with gpt-oss-120b on a large corpus, mixed with context pruning. The maintainer called the problem model dependent, so test any root model you plan to use.
Not reliably. The old reference prompt already told models not to use FINAL tags until the task was complete, and the Claude 4.6 behavior still appeared. The repo fixed it by changing how a run ends, not by adding more instructions.
It removes the worst case. Since PR #162, text FINAL answers are ignored and only answer ready set in real code ends a run. The fabricated text can still enter the message history, so a stop sequence or client-side truncation is still worth adding.
It depends on the provider and model. The OpenAI Chat Completions reference says the stop parameter is not supported on o3 and o4-mini. In that case, truncate after the closing fence on the client or use native tool calling.