harbor-framework/harbor

Framework for evaluating and improving agents

What it solves

Harbor provides a unified framework for evaluating and optimizing AI agents and language models. It simplifies the process of running benchmarks across diverse environments, allowing developers to avoid the manual setup of complex testbeds for different agents (like Claude Code or OpenHands).

How it works

Harbor acts as an execution harness that connects agents, models, and datasets. It allows users to run evaluations in parallel across thousands of environments using cloud providers like Daytona, Modal, and LangSmith. It can bet used to run existing benchmarks (such as Terminal-Bench-2.0 or SWE-Bench) or to build and share custom environments and benchmarks.

Who it’s for

AI researchers and developers building agents or LLMs who need a scalable way to test their systems against standardized datasets and generate rollouts for reinforcement learning (RL) optimization.

Highlights

  • Scalable Execution: Run experiments in parallel across thousands of environments via various cloud providers.
  • Agent Agnostic: Supports evaluating arbitrary agents including Claude Code, OpenHands, and Codex CLI.
  • Benchmark Integration: Official harness for Terminal-Bench-2.0 and supports other third-party benchmarks like SWE-Bench.
  • RL Support: Capable of generating rollouts for RL optimization.

관련

  • 프로젝트
  • 프로젝트
  • 프로젝트
  • 프로젝트
  • 프로젝트