livenerf: A Deterministic Benchmark for Detecting Post‑Launch Capability Drift in Claude Opus 5.5
TL;DR – What livenerf does and why it matters
livenerf continuously runs a frozen panel of 78 sometimes‑right questions against Claude Opus 5.5 (via Claude Code Max) and reports any statistically significant change in accuracy or output‑token count relative to the first‑10‑day launch baseline. The project provides the first public, reproducible time‑series that can confirm or refute claims that Anthropic “nerfs” its models after release.
Core contribution: a deterministic, append‑only drift detector
livenerf isolates every source of nondeterminism except the model’s own stochasticity. It pins the Claude Code CLI version, freezes the system prompt, uses exact‑match graders, and logs every request‑level artifact (prompt, response, usage, latency, harness SHA, etc.). By running the same panel daily for 30 days, the benchmark measures paired per‑item score differences with clustered standard errors, following the statistical framework of Anthropic’s Adding Error Bars to Evals (Miller 2024). This design eliminates drift from prompts, tools, or environment changes, making any observed shift attributable to the model or serving stack.
How the benchmark is built
1. Calibration isolates informative items
Only questions that the model gets right sometimes (≈50 %‑70 % pass rate) are kept. Fully‑right or fully‑wrong items carry no information about degradation. Calibration sampled 2,336 GPQA‑Diamond, MMLU‑Pro, competition‑math, and AIME 2025‑26 questions with four samples each, yielding 78 panel items (plus two excluded later). These items have a measured fresh‑sample pass rate of 62 % after correcting for selection bias.
2. Paired daily runs eliminate item difficulty
Each day the same panel is evaluated and the result is compared to the same item’s baseline score from days 1‑10. Because the comparison is per‑item, difficulty cancels out, and the primary metric becomes the mean paired difference across items.
3. Two‑arm design controls for platform effects
Primary arm: Opus 5.5 via Claude Code Max. Control arm: Claude Opus 5 (the previous generation) runs the same GPQA subset. If both arms move together, the change is attributed to the harness or infrastructure rather than the new model.
4. Pre‑registered validation proves sensitivity
Before collecting baseline data, the rig runs a known degradation (effort medium vs. high) and an A/A check. The validation shows that a 26 % token‑usage reduction corresponds to a –4.2 ± 3.9‑point accuracy drop, while a 62 % token reduction yields –8.3 ± 4.5 points. This establishes that the instrument can detect real effort‑related regressions.
Running the experiment: timeline and current status (as of 2026‑09‑29)
| Phase | Days | Samples | Status |
|---|---|---|---|
| Baseline | 1‑10 (launch 2026‑09‑24) | 90 samples/day | 6 days collected, no misses |
| First 10‑day window | 11‑20 | – | Pending |
| Second 10‑day window | 21‑30 | – | Pending |
All runs use the pinned CLI version 2.1.280 (hash 461391b6fce64167). Day 5 required a one‑time budget‑guard override, which is logged in the deviations file.
Detectable effect size and limits
- A single daily run of the full panel can detect an accuracy change of ~7.5 points over a 10‑day window (≈3.6 % of the weekly plan’s statistical power).
- The instrument cannot reliably distinguish Opus 5 from Opus 5.5: the observed difference was –3.8 ± 6.3 points, not statistically significant at 99 % confidence.
- Token‑usage is a more sensitive early indicator; a modest effort reduction shows up as a large token drop before accuracy moves.
How to reproduce the benchmark
- Environment – Python 3.11+,
uv, and a logged‑in Claude Code Max subscription. - Pin the CLI – Disable auto‑updates, record
claude --versiontoCLAUDE_CLI_VERSION, and place a copy of the binary in~/.local/share/livenerf. - Install –
git clone https://github.com/ninjahawk/livenerf && cd livenerf && uv sync. - Calibrate & design – Run
python -m livenerf.benchmarks.calibrate …thenpython -m livenerf.design …to generatedata/standard_panel.jsonand lock the panel. - Validate – Execute
python -m livenerf.validate run …to confirm the rig detects the known effort degradation. - Pre‑flight –
python -m livenerf.preflight --probemust printREADYbefore scheduling. - Schedule – Add the provided
scripts/daily.shto cron (Linux/macOS) or Task Scheduler (Windows) to run once per day, respecting the weekly usage guard (skip if >75 % weekly meter or >60 % 5‑hour meter). - Analyze – After each 10‑day window, run
python -m livenerf.analysisto obtain paired deltas, clustered SEs, and the automated decision (Δ > 3 pts, 99 % CI excludes zero, control arm stable).
All raw logs are stored as Inspect .eval files in logs/ (append‑only, never edited).
Community reaction on Hacker News
@jug – “BridgeBench’s Nerf Bench detected a degradation of Opus 4.6 that Anthropic later blogged about.”
@johnfn – “Most ‘nerf’ reports are perception bias; a graphic explains the honeymoon‑effect where new models feel infinitely capable until users hit a complexity ceiling.”
@Cider9986 – “I saw a nerf last night in a Claude Code chat; a quick Nitter search turned up the same claim.”
@zeroonetwothree – “If the benchmark can’t tell Opus 5 from Opus 5.5, it isn’t useful.”
@nomilk – “Great idea; it holds providers accountable and is a little embarrassing that it’s needed.”
The comments illustrate a split between anecdotal “nerf” experiences and skepticism that systematic measurement is possible. livenerf directly addresses the latter by providing a pre‑registered, statistically rigorous protocol.
Limitations and open questions
- Model‑specificity – The benchmark only measures Claude Opus 5.5 as served through Claude Code Max, not the raw API model. Results may differ on Azure or Bedrock deployments.
- Panel privacy – Individual question texts are kept private (only hashes are published) to avoid leaking proprietary evaluation data.
- Detection ceiling – Same‑family swaps (Opus 5 → 5.5) are below the current statistical power; larger sample sizes or longer windows would be needed.
- External factors – Load‑dependent serving changes, safety‑classifier rejections, or quota throttling can affect token usage and are logged but not fully disentangled.
Future directions
- Extending the protocol to other frontier models (e.g., GPT‑6 Astra) and to API‑only deployments.
- Publishing a public synthetic panel for community replication while preserving the private calibrated items.
- Automating cross‑provider comparisons (Azure, Bedrock) to test whether “nerfing” is provider‑specific.
- Investigating finer‑grained token‑usage signals (e.g., per‑step thinking time) as early‑warning indicators.
How to contribute
Contributions should only add new pure‑function graders, synthetic generators, or analysis tools. Changing the frozen panel or prompts creates a new version and must be accompanied by a new pre‑registration commit. All PRs must disclose any LLM‑generated content, as the repository itself was partially authored with Claude.
Citation
@misc{livenerf,
author = {ninjahawk},
title = {livenerf: tracking post‑launch capability drift in frontier models},
year = {2026},
publisher = {GitHub},
url = {https://github.com/ninjahawk/livenerf}
}
Sources
Related
- Dispatch
- Dispatch
- Project
- Dispatch
- Dispatch