Fable 5 Median Thinking Decline in August – Evidence, Causes, and Community Reactions

Takeaway

Lon Lundgren’s six‑week measurement shows that Fable 5’s median "thinking" token count fell dramatically in August, and the drop persisted across multiple workloads. The evidence points to a change in the inference regime (e.g., compute allocation or sampling settings) rather than a deliberate model "nerf".


What the Data Shows

  • Median reasoning tokens fell: Across a diverse set of production projects, the number of tokens the model spent on internal reasoning (the "thinking" tokens) dropped from the levels seen in early August to near‑zero for many invocations.
  • Fluctuations align with releases: Peaks and troughs in the token counts corresponded to specific product announcements and version releases, suggesting that the service’s configuration changes over time.
  • Consistent across effort levels: Even when users selected the highest effort settings ("xhigh" or "max"), the majority of calls received little or no thinking tokens.
  • Longer runs still under‑perform: When the model did engage in extended reasoning, the token usage never reached the benchmark levels reported in the model’s published papers.

"I was consistently using an xhigh or max effort level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all. And when longer thinking runs did happen, they almost never reached published benchmark levels." – Lon Lundgren (Twitter thread)


Why Inference Regime Matters More Than Model Architecture

  • Inference regime = compute budget, sampling temperature, token limits, and internal routing. Changing any of these can reduce the amount of internal reasoning without altering the underlying weights.
  • User‑visible performance drops: Users report that the model feels "dumber" after a few weeks, even when they keep the same settings. This aligns with Lundgren’s finding that the service, not the model, changed.
  • Hidden A/B testing: Several commenters suspect that the provider runs covert A/B experiments, adjusting compute allocation to balance cost against perceived performance.

"It is clear by now to me that Anthropic is constantly trying to find a kind of ‘auto’ degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3‑4 weeks." – @theplumber


Community Observations Supporting the Trend

  • Anecdotal reports: Multiple users independently noted a rapid decline in reasoning ability within weeks of a new release (e.g., gpt‑5.6‑luna, Claude Code, Opus).
  • Consistent pattern across models: The phenomenon is not unique to Fable 5; similar drops have been reported for Anthropic’s Opus and Claude Code, suggesting a broader industry practice.
  • Quantitative trackers: Some community‑maintained trackers (e.g., marginlab.ai) show a trend toward fewer tokens needed for the same tasks, though the quality impact is debated.

"I have found the same. I spend a lot of time with these frontier models, brainstorming, etc., and the drop in performance from, say, week 1 to week 8 is often massive." – @mlmonkey


Methodological Critiques

  • Workload variability: Critics note that the token count metric conflates model effort with task difficulty; as projects mature, they may require fewer reasoning steps.
  • Benchmark choice: Lundgren compares thinking tokens to ARC‑AGI‑2, a benchmark designed for extremely hard problems, which may overstate the expected token usage for everyday coding tasks.
  • Data collection transparency: The analysis relied on a man‑in‑the‑middle proxy capturing wire logs, a method that reveals inference‑side changes but does not expose the exact server‑side configuration.

"The corpus analyzed comes exclusively from Fable 5, at xhigh and max effort levels, during sustained production work across a diverse set of projects and workloads. Data was aggregated from transcripts and live wire logs. The analysis gets worse from there…" – @Aurornis (commentary on methodology)


Possible Explanations

  1. Cost‑driven compute throttling – Providers may reduce per‑request compute after an initial launch window to manage operating expenses.
  2. Intentional A/B experiments – Rotating inference settings across user cohorts can hide performance regressions while still delivering headline benchmark scores.
  3. Model‑agnostic degradation – As the underlying model stays static, only the service layer changes, meaning the same model can appear weaker without any weight updates.
  4. User workload drift – Over time, developers may rely on the model for more routine tasks, naturally requiring fewer reasoning tokens.

Legal and Transparency Implications

  • Potential liability: If a provider deliberately reduces service quality without clear disclosure, users could argue breach of contract or deceptive practices.
  • Calls for proof‑of‑compute: Community members suggest requiring providers to return a checksum‑like proof of the quantization level or compute budget used for each request.

"All llm api providers should be compelled to return a checksum‑like proof of quantization level of the model that served the request. Basic transparency should be the bare minimum." – @reilly3000


What Users Can Do Now

  • Log inference details: Capture token usage, model version, and effort level for each request to detect regressions.
  • Switch to self‑hosted models: Tools like Ollama let users run comparable models locally, avoiding opaque service changes.
  • Demand public benchmarks: Encourage providers to publish regular, unfiltered benchmark runs (e.g., every 2–3 days) to reduce the incentive for hidden throttling.

Conclusion

Lon Lundgren’s six‑week deep dive provides strong evidence that Fable 5’s reasoning capability has been systematically reduced through changes in the inference regime, not by altering the model itself. The pattern mirrors broader community reports of similar degradations across frontier LLM providers, raising concerns about cost‑driven throttling, transparency, and user trust. Until providers disclose inference configurations or adopt verifiable proof‑of‑compute mechanisms, users will need to rely on independent monitoring and self‑hosted alternatives to safeguard performance.

Sources

Related