EvalEval and UK AISI Release Open Evaluation Cards for Frontier LLM Benchmarks
TL;DR
EvalEval and the UK AI Security Institute (AISI) have released publicly verified Evaluation Cards for five high‑profile benchmarks (HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, Terminal‑Bench 2.0) covering six frontier LLMs, providing the full configuration, compute, and protocol details needed for reproducible evaluation.
Why reproducible evaluation reporting matters
Reproducibility is essential because evaluation results increasingly serve as the primary evidence for model performance and safety. Current reporting is fragmented across formats and platforms, often omitting critical details such as inference compute, evaluation protocol, and data splits, which makes re‑running experiments prohibitively expensive. EvalEval addresses this gap with two core artifacts:
- Every Eval Ever (EEE) schema – a shared, machine‑readable specification for benchmark metadata, model metadata, and run‑time parameters.
- Evaluation Cards – an open‑access portal that stores EEE‑compliant records, linking scores to the exact experimental setup.
Together with AISI’s research on efficient and statistically rigorous evaluation (e.g., OptStop, HiBayES), the partnership aims to diagnose reporting gaps and provide a standardized infrastructure for transparent evaluation.
What AISI is sharing
AISI has uploaded Evaluation Cards for the five benchmarks listed in its recent paper How Inference Compute Shapes Frontier LLM Evaluation (arXiv:2606.17930). Each card contains:
- Verified scores for each benchmark.
- Full context including model version (Claude Opus 4, 4.5, 4.6; GPT‑5, 5.2, 5.4), inference‑time compute budget, token limits, and evaluation protocol.
- Configuration details such as prompt templates, sampling settings, and any oracle feedback mechanisms.
The release also includes two auxiliary cyber‑security evaluations (Cyber CTFs and The Last Ones) that use a partially overlapping model set. By exposing these details, researchers can:
- Examine how changes in compute or protocol affect performance (e.g., the token‑efficiency curves for Humanity’s Last Exam).
- Compare results across studies that appear numerically similar but were obtained under different conditions.
- Use the cards as verified reference points for meta‑research and policy analysis.
Performance on Humanity’s Last Exam changes with evaluation protocol and inference compute. Each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task.
How the community can contribute
The EvalEval Coalition invites three groups to help expand the ecosystem:
- Model developers – submit verified Evaluation Cards for their models via the Get Verified workflow.
- Benchmark developers – encode new benchmarks and run data using the open‑source EEE schema (GitHub link provided).
- Researchers in evaluation, governance, and policy – explore the public card repository to assess reporting quality, conduct meta‑analyses, or inform regulatory frameworks.
About the EvalEval Coalition
The EvalEval Coalition is a research community focused on building robust, scientifically grounded infrastructure for AI evaluation. Its flagship projects are:
- Every Eval Ever – a universal schema and public repository for evaluation results.
- Evaluation Cards – an interface that aggregates benchmark metadata, run‑time data, and model metadata into interpretable, comparable records.
These tools make it possible to detect when identical scores arise from materially different experimental conditions, thereby improving the reliability of evaluation‑driven decisions.
About the UK AI Security Institute (AISI)
AISI is a UK government research organization within the Department for Science, Innovation and Technology. Its mission is to provide governments with scientific insight into the risks of advanced AI. AISI’s work includes:
- OptStop – methods for reducing evaluation cost by early‑stopping unpromising runs.
- HiBayES – hierarchical Bayesian modeling to produce statistically rigorous LLM evaluation.
- Standardizing transcript analysis and capability elicitation across benchmarks.
Further reading
- How Inference Compute Shapes Frontier LLM Evaluation (arXiv:2606.17930)
- HiBayES: Improving LLM evaluation with hierarchical Bayesian modelling (AISI blog & arXiv:2505.05602)
- OptStop paper (arXiv:2608.14425)
- Every Eval Ever project page
- Evaluation Cards portal