ninjahawk/livenerf

Benchmark for tracking model capability after release.

What it solves

livenerf is designed to detect "nerfing"—the phenomenon where a frontier AI model's performance quietly degrades after its initial launch. It replaces anecdotal reports of quality drops with a deterministic, long-running benchmark that tracks a model's performance over time against a launch-day baseline.

How it works

The project uses a rigorous statistical approach to measure drift in a model's output:

  • Informative Panel Selection: It filters large datasets (GPQA Diamond, MMLU-Pro, competition-math, AIME) to keep only questions the model gets right sometimes (roughly 30-70% pass rate), as these are the most sensitive to change.
  • Deterministic Harness: To eliminate noise, it pins the CLI version, freezes system prompts, and uses exact-match graders instead of LLM judges.
  • Statistical Tracking: It runs the same panel daily and compares results against a launch-week baseline using paired per-item score differences and clustered standard errors.
  • Control Arms: It runs a control model (e.g., Claude Opus 5) alongside the target model to distinguish between model degradation and platform-wide infrastructure changes.
  • Token Monitoring: It tracks output token counts as a secondary signal, as reductions in "thinking" tokens often precede accuracy drops.

Who it’s for

Researchers and developers who want objective, data-driven evidence of model drift in frontier LLMs, specifically those served via subscription interfaces like Claude Code.

Highlights

  • Pre-registered Protocol: Uses a public git timestamp to commit to a measurement plan and decision rule before data collection begins.
  • Inspect AI Integration: Built on the UK AI Security Institute's open-source evaluation framework.
  • Rigorous Validation: Includes a "positive control" check to ensure the rig can detect known degradations (e.g., changing effort levels) before trusting null results.
  • Hermetic Execution: Ensures every call is isolated with no external tools, memory, or settings to prevent leakage.

Related

  • Project
  • Project
  • Project
  • Project