centerforaisafety/hle
Humanity's Last Exam
What it solves
It addresses the need for a high-difficulty, closed-ended academic benchmark to test the limits of AI models' knowledge across a wide range of subjects. It is designed to be the final benchmark of its kind, pushing the frontier of human knowledge to prevent models from simply memorizing existing test sets.
How it works
The project provides a dataset of 2,500 multi-modal questions across dozens of subjects, including mathematics, humanities, and natural sciences. These questions are multiple-choice or short-answer, allowing for automated grading. It includes a simple evaluation pipeline using the openai-python interface to generate predictions and judge results.
Who it’s for
AI researchers and model builders who need to evaluate the same level of advanced academic knowledge that human subject-matter experts possess.
Highlights
- Multi-modal benchmark covering dozens of subjects.
- 2,500 expert-developed questions.
- Includes a canary string to help model builders filter the dataset from training data to prevent data contamination.
- Automated grading for multiple-choice and short-answer responses.
Related
- Project
- Project
- Dispatch
- Project
- Project