Back to Blog
October 5, 2026

The short version: yes. Several 2025 studies measured accuracy against reasoning length and found a curve that rises, peaks, then falls. Past the peak, extra thinking tokens add distraction, variance, and second-guessing. The same pattern shows up in RLMs as recursion depth: one level of sub-calls helps, and a second level can cut accuracy while it multiplies latency.

What does the evidence say about longer reasoning?

The clearest statement comes from Wu et al., who studied chain-of-thought length directly. Their abstract reports that "task accuracy typically follows an inverted U-shaped curve with CoT length, where performance initially improves but eventually decreases as the number of CoT steps increases." The peak is not fixed. The optimal length "increases with task difficulty but decreases with model capability." A stronger model needs fewer steps on the same problem.

Hassid et al. looked at the same question inside a single prompt. They sampled several thinking chains for each question and compared them by length. Shorter chains were "up to 34.5% more accurate than the longest chain sampled for the same question." That is a within-question comparison, so task difficulty does not explain it. When the same model thinks longer on the same problem, the long attempt is more often the wrong one.

Ghosal et al. forced the issue. They extended thinking traces with prompts like "Wait" and "Let me rethink" and reported "a consistent pattern of initial performance improvements from additional thinking followed by a decline, due to 'overthinking'." Their explanation is that additional thinking increases output variance. The early gains look like better reasoning, but part of that gain is an artifact of how uncertainty interacts with the metric.

Which tasks get worse with a bigger thinking budget?

An Anthropic-led team built tasks to find out. In "Inverse Scaling in Test-Time Compute", Gema et al. constructed evaluations "where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance." The tasks span four categories:

  • simple counting tasks with distractors
  • regression tasks with spurious features
  • deduction tasks with constraint tracking
  • advanced AI risks

The failure modes differ by model family. The abstract says Claude models "become increasingly distracted by irrelevant information," while OpenAI o-series models "resist distractors but overfit to problem framings." On the regression tasks, models "shift from reasonable priors to spurious correlations." On the deduction tasks, "all models show difficulties in maintaining focus."

Read those four lines together and a rule appears. Longer reasoning hurts most when the prompt contains something tempting and irrelevant. More tokens give the model more chances to pick it up. The authors do not reject test-time compute. They call it "promising" and ask that models be evaluated "across diverse reasoning lengths."

Apple's "The Illusion of Thinking" adds a second regime. Shojaee et al. compared reasoning models with their standard counterparts under the same inference compute. On low-complexity tasks, "standard models outperform LRMs." On medium-complexity tasks, the reasoning models win. On high-complexity tasks, both collapse. They also saw that reasoning effort "increases with problem complexity up to a point, then declines despite having remaining token budget." A large budget does not mean the model uses it well.

Does overthinking also hurt agents?

Yes, and the cost is larger because an agent can act on a bad plan. Cuadron et al. analyzed 4018 trajectories on SWE Bench Verified. They define overthinking as a model that favors "extended internal reasoning chains over environmental interaction," and they name three patterns: Analysis Paralysis, Rogue Actions, and Premature Disengagement.

Their headline result: "higher overthinking scores correlate with decreased performance," and reasoning models overthink more than non-reasoning models. A simple fix worked. Selecting the solution with the lower overthinking score improved performance "by almost 30% while reducing computational costs by 43%."

This matters for RLMs because an RLM root is an agent in a REPL. A root that reasons at length about what the context probably contains, without running code to check, shows the same pattern. The fix is the same: run the code, then reason about the result.

What do the vendor docs say about thinking budgets?

Vendor guidance is more careful than "more is better," but it stops short of the papers. Anthropic's extended thinking documentation says "larger budgets can improve response quality by enabling more thorough analysis for complex problems." It then tells you to start low: "For simple tasks, start near the 1,024-token minimum and increase incrementally to find the optimal range for your use case." For complex tasks it suggests 16,000 tokens or more, "with diminishing returns that depend on the task."

Two details on that page are easy to miss. The budget "is a target rather than a strict cap," so a high number does not force a long trace. And the fixed budget is on its way out. The page states that budget_tokens is deprecated on the Claude 4.6 models and rejected by Claude 4.7 and later, which use adaptive thinking with an effort setting. The model then decides how much to think on each request.

The docs describe diminishing returns. The papers describe negative returns on specific task types. Both can be true. The practical reading is the same: treat thinking as a setting you tune per task, and measure it.

How does this show up in RLM recursion depth?

Recursion depth is the RLM version of a thinking budget. The RLM paper defines it this way: "Max recursion depth 0 is an RLM without sub-calling capabilities. Max recursion depth 1 allows sub-calling LLMs, while max depth >1 allows sub-calling RLMs." Each level lets the system spend more compute on the same question.

The reproduction by Wang tested what the extra level buys on open models (DeepSeek v3.2 and Kimi K2). The title gives the answer: "Think, But Don't Overthink." Depth 1 improved accuracy on complex reasoning tasks. But "applying deeper recursion (depth=2) or using RLMs on simple retrieval tasks paradoxically degrades performance and exponentially inflates execution time (e.g., from 3.6s to 344.5s) and token costs."

That is the inverted U again, with depth on the horizontal axis. It also matches the Apple result: on a simple task, the plain model beats the system that thinks more. Our post on how deep an RLM should recurse walks through the depth numbers task by task.

How should you set thinking and sub-call budgets in an RLM?

The research points to five rules.

  1. Start low and sweep. Run your evaluation set at several thinking levels and several depths. Keep the lowest setting that reaches peak accuracy. Our guide on how to evaluate an RLM covers the setup.
  2. Give sub-calls less thinking than the root. A sub-call usually does a narrow job on one snippet: extract, classify, or count. Those are the low-complexity tasks where long reasoning helps least. This is an inference from the studies above, so test it on your workload.
  3. Remove distractors before the model sees them. Gema et al. found that distraction grows with reasoning length. An RLM can filter with code first, so each sub-call gets a short, relevant snippet.
  4. Spend extra budget in parallel. Ghosal et al. generated independent reasoning paths within the same budget and took a majority vote, "achieving up to 20% higher accuracy compared to extended thinking." Hassid et al. report that their short-1@k method matched or beat majority voting in low-compute settings while "using up to 40% fewer thinking tokens." An RLM fan-out already has this shape.
  5. Cap depth at 1 unless a test says otherwise. Depth 1 is the default in the RLM paper. Raise it only when your own numbers show a gain.

The bottom line

More thinking tokens can make a reasoning model less accurate. The evidence is consistent across chain-of-thought length, forced extended thinking, agent trajectories, and RLM recursion depth. Accuracy rises with reasoning up to a peak that depends on the task and the model, then it falls. Simple tasks and prompts with distractors reach that peak early.

For RLM builders, the lesson is to treat thinking effort and recursion depth as tuned settings. Sweep them, keep the smallest value that holds accuracy, and spend spare compute on parallel short calls. For the wider case for test-time compute, see inference-time compute scaling.

References & Further Reading

  1. Gema AP, Hagele A, Chen R, et al. "Inverse Scaling in Test-Time Compute." arXiv:2507.14417, July 2025; published in TMLR, December 2025. Tasks where longer reasoning lowers accuracy, and five failure modes by model family. arxiv.org/abs/2507.14417
  2. Wu Y, Wang Y, Ye Z, Du T, Jegelka S, Wang Y. "When More is Less: Understanding Chain-of-Thought Length in LLMs." arXiv:2502.07266, February 2025. Inverted U-shaped accuracy curve against chain-of-thought length; optimal length scales with task difficulty and model capability. arxiv.org/abs/2502.07266
  3. Hassid M, Synnaeve G, Adi Y, Schwartz R. "Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning." arXiv:2505.17813, May 2025. Shorter chains up to 34.5% more accurate than the longest chain for the same question; short-m@k method. arxiv.org/abs/2505.17813
  4. Ghosal SS, Chakraborty S, Reddy A, et al. "Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models." arXiv:2506.04210, June 2025; NeurIPS 2025. Gains then decline from extended thinking; parallel thinking with majority vote. arxiv.org/abs/2506.04210
  5. Shojaee P, Mirzadeh I, Alizadeh K, Horton M, Bengio S, Farajtabar M. "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity." arXiv:2506.06941, June 2025; NeurIPS 2025. Three complexity regimes; reasoning effort declines despite remaining token budget. arxiv.org/abs/2506.06941
  6. Cuadron A, Li D, Ma W, et al. "The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks." arXiv:2502.08235, February 2025. 4018 SWE Bench Verified trajectories; overthinking scores correlate with lower performance. arxiv.org/abs/2502.08235
  7. Anthropic. "Extended thinking." Claude API documentation, accessed October 5, 2026. Budget minimum, tuning guidance, diminishing returns, deprecation of budget_tokens in favor of adaptive thinking. platform.claude.com/docs/en/build-with-claude/extended-thinking
  8. Zhang AL, Kraska T, Khattab O. "Recursive Language Models." arXiv:2512.24601. Definition of max recursion depth 0, 1, and greater than 1; depth 1 as the default. arxiv.org/abs/2512.24601
  9. Wang D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, March 2026. Depth 2 degrades performance on open models and inflates execution time from 3.6s to 344.5s. arxiv.org/abs/2603.02615
FAQ

Frequently asked questions

What is overthinking in a reasoning model?

Overthinking is reasoning that continues past the point where it helps. Ghosal et al. observed initial gains from additional thinking followed by a decline. Cuadron et al. use the term for agents that favor long internal reasoning over interaction with the environment.

Is there one optimal number of thinking tokens?

No. Wu et al. found that the optimal chain-of-thought length increases with task difficulty and decreases with model capability. You have to find the peak for each task type and each model with an evaluation sweep.

Is it better to run several short reasoning attempts than one long attempt?

Two 2025 papers say it often is. Ghosal et al. report up to 20% higher accuracy from parallel thinking with a majority vote, compared with extended thinking at the same budget. Hassid et al. report that preferring the shortest chains matched or beat standard majority voting with up to 40% fewer thinking tokens in low-compute settings.

Does a high thinking budget force the model to think for that long?

No. Anthropic's documentation says the budget is a target, not a strict cap, and that the model may stop reasoning well before the budget is exhausted. Newer Claude models replace the fixed budget with adaptive thinking and an effort setting.

Does deeper recursion in an RLM count as more thinking?

In practice, yes. Each added depth level lets the system spend more compute on the same question. The Wang reproduction found that depth 2 degraded performance on open models and raised execution time from 3.6 seconds to 344.5 seconds in one example.