harbor-framework/terminal-bench
Measuring and evolving with the frontier of agent work
What it solves
Terminal-Bench is a benchmark designed to measure the capabilities of frontier AI agents. It provides a diverse and difficult set of high-quality tasks that evolve over time, allowing agent builders to track progress and compare the performance of different agents and models.
How it works
Built on the Harbor framework, the benchmark consists of a dataset of tasks that can be run in a sandboxed environment (such as Modal). Users can test their agents by specifying the agent and model to be used, and run the benchmark using the Harbor CLI tool. The project maintains a continuous benchmark with tagged releases and a public leaderboard.
Who it’s for
This tool is for frontier AI agent builders who want to evaluate their agents' ability to handle complex, real-world terminal-based tasks.
Highlights
- Continuous benchmark with evolving tasks to prevent stagnation.
- Integration with the Harbor framework for standardized execution.
- Support for sandboxed environments to ensure safe and agent execution.
- Public leaderboard for comparing agent and model performance.
Related
- Project
- Project
- Project
- Project