harbor-framework/terminal-bench-science
Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains
What it solves
Terminal-Bench-Science provides a standardized way to measure how well AI agents can handle complex, real-world scientific research workflows. It addresses the gap between general AI capabilities and the specialized, expert-level tasks required in scientific domains like the life, physical, earth, mathematical, and engineering sciences.
How it works
The project is a continuous benchmark consisting of expert-curated tasks that are objectively verified in a terminal environment. It uses the Harbor framework to run agents against these tasks. The benchmark evolves through a community-driven pipeline where domain experts propose, build, and review tasks to ensure they remain challenging and relevant to frontier AI models.
Who it’s for
This benchmark is designed for AI researchers developing agents and for scientific domain experts who want to contribute tasks that reflect the actual work they need AI systems to support.
Highlights
- Expert-Curated: Tasks are authored and reviewed by domain experts across five broad scientific domains.
- Objectively Verifiable: All tasks produce outcomes that can be verified within a terminal environment.
- Community-Driven: Features a transparent pipeline for task proposals, implementation, and review.
- Contamination Protection: Uses canary strings to help training-data pipelines detect and exclude benchmark data from training corpora.
Related
- Project
- Project
- Project
- Project
- Project