Claude Code Effort Level A/B Testing and User Feedback
Anthropic A/B Tests Effort Level Mappings in Claude Code
Anthropic is conducting server-side A/B tests on how numerical effort values are mapped to the "effort" setting in Claude Code (versions 2.1.236+). In these tests, the numerical value associated with a "high" effort setting may be mapped to a lower number (e.g., 10 out of 100) than in previous versions, leading some users to perceive a reduction in model performance or "intelligence."
Technical Implementation and Official Response
Thariq from the Claude Code team has clarified that these changes are server-side API serving configurations. The numerical values displayed or logged—such as a "10" appearing for a high effort setting—are internal mappings and do not necessarily represent a a 0-100 scale of actual effort.
According to the Claude Code team, the effort level selected by the user remains the same in terms of actual model performance. The team maintains that in-depth evaluations have confirmed that these mapping changes do not affect model performance. Users experiencing clear regressions are encouraged to use the /feedback command to report issues with their session IDs.
User Perceptions of Model Performance
Despite official assurances, several users report a noticeable decline in the quality of outputs from newer model versions, specifically referencing "Fable" and "Opus 5."
- Inefficiency in Task Execution: One user reported that a simple config file update that previously took under two minutes on version 4.6 took 43 minutes on Opus 5, involving unnecessary sandboxes and testing suites that exceeded the scope of the request.
- Tangents and Over-Engineering: Some users have observed that Opus 5 and Sonnet 5 tend to go on unasked-for tangents, particularly when set to "high" effort, compared to older models.
- Model Degradation: Some users have downgraded their subscriptions or switched to alternative models (such as Codex or GLM-5.3) due to perceived quality drops in the Fable model.
Community Discussion on AI Incentives and Billing
The incident has sparked a broader community discussion regarding the transparency of LLM providers. Users expressed concerns over "enshitification"—the process of where products become less useful to the user to maximize profit—and the same opaque nature of token-based billing.
"Why are we allowing billing to take place in tokens that are nebulous and fully controlled by the operators who have no aligned incentives?"
Critics argue that the billing model based on tokens, rather than raw compute or resource usage, allows providers to change model behavior or routing in the backend without user visibility, potentially leading to a higher cost for lower-quality responses.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch