Artificial Analysis Intelligence Index v4.2 release analysis

New evaluations raise the bar for realistic, private testing

AA‑Briefcase and Surge’s GDP.pdf are the headline additions in Index v4.2. AA‑Briefcase evaluates multi‑week, agentic knowledge‑work projects built by industry experts, using a private held‑out test set, rubric‑based scoring, and pairwise grading to measure task success, analytical quality, and presentation quality. GDP.pdf, created by Surge AI, requires models to synthesize evidence across 4,592 PDF pages (text, tables, charts, footnotes) and meet 1,275 atomic criteria; the headline metric is an All‑pass Rate that only counts a task as correct when every criterion is satisfied.

Private held‑out weighting doubles to curb gaming

The proportion of private, held‑out test data in the overall Index weighting has risen from 20 % in v4.1 to 40 % in v4.2. Held‑out sets now include AA‑Briefcase, AA‑Omniscience, and solutions for CritPt. This change is intended to reduce the ability of model developers to over‑fit public benchmarks and to make the Index more reflective of real‑world performance. The team signals that the private weighting will increase further in the upcoming v5.

Grading infrastructure upgrades improve score stability

Version v4.2 introduces a new grading‑system prompt in AA‑LCR v1.1, fixes ambiguities in answer keys, and re‑anchors the Elo scale for GDPval‑AA v2 and AA‑Briefcase. SciCode’s sandbox has been hardened so that slow‑but‑correct code no longer registers as a failure. These upgrades aim to make scores more accurate and less volatile as new models are added.

Leaderboard highlights: Claude Fable 5.1 and GPT‑6 Astra dominate

  • Anthropic’s Claude Fable 5.1 holds the top spot on the overall Index.
  • OpenAI’s GPT‑6 Astra ranks second overall and leads the output‑token efficiency frontier, achieving a substantially lower token cost per unit of intelligence compared with most competitors.
  • Meta, SpaceXAI, Moonshot/Kimi, Z.AI, and Google round out the top six.

Cost‑per‑Task Pareto frontier

Four labs—Anthropic, OpenAI, Meta, and Z.AI—occupy the updated cost‑per‑task Pareto frontier, indicating they deliver the highest intelligence scores for the lowest compute cost.

Token‑efficiency frontier

GPT‑6 Astra is the most token‑efficient model among those scoring above 25 on the Index, with Claude Fable 5.1, Grok 4.5, and Gemini 3.5 Flash‑Lite forming the extremes of the curve.

Detailed benchmark performance

  • AA‑Briefcase: Claude Fable 5.1 and Opus 5 lead, followed by GPT‑6 Astra and Muse Spark 1.3. GPT‑6 Astra shows an ~85 Elo‑point gain over GPT‑5.6 Sol.
  • GDP.pdf: OpenAI’s GPT‑6 Astra achieves a 33.2 % All‑pass Rate, ahead of GPT‑5.6 Sol (28.2 %) and Claude Fable 5.1 (26.2 %).

Community reactions on Hacker News

  • Positive view of private held‑out tests: @jascha_eng notes that the Omniscience index, which measures knowledge reliability and penalizes hallucinations, correlates strongly with real‑world usefulness. They highlight that Fable outperforms Opus 5 on this metric, aligning with perceived model strength.
  • Skepticism about post‑hoc adjustments: @redox99 argues that the index was tweaked to align Astra’s score with expectations, calling the change “unscientific.”
  • Criticism of benchmark relevance: @throwaway13337 dismisses the new benchmark set as “not useful,” questioning the inclusion of models like Gemini 3.8.
  • Praise for token‑efficiency claims: @jl emphasizes Astra’s dominance on the output‑token frontier, noting it has the second‑highest overall score and third‑lowest token usage among displayed models.
  • Calls for broader benchmark coverage: @stared and @6thbit ask why ARC‑AGI‑3 is absent, suggesting its inclusion would shift rankings.
  • Concerns about weighting changes: @sanxiyn and @dgacmu criticize the increased weight of SciCode, labeling it a “broken benchmark” and advocating for newer versions (e.g., Terminal‑Bench 4.0) to better differentiate models.

What’s next for the Artificial Analysis Index?

The team confirms that v4.2 is an interim release accelerating parts of the upcoming v5 roadmap. They plan additional incremental releases before the full v5 launch, with further increases in private test weighting and more sophisticated real‑world tasks.


Read the full methodology and detailed per‑model breakdowns at the Artificial Analysis website.

Sources

Related