harveyai/harvey-labs

A benchmark built to evaluate and improve agent capabilities for supporting legal work.

What it solves

It provides a standardized way to evaluate how well LLM-based agents can handle real-world legal work. Because legal tasks are complex and require high precision, this benchmark helps developers measure agent performance across various legal practice areas.

How it works

The project consists of two main components: a dataset of tasks (which include specific instructions, relevant documents, and scoring rubrics) and an execution harness. This harness allows developers to run their agents against the tasks and automatically evaluate their performance based on thedefined rubrics.

Who it’s for

It is designed for AI researchers and developers building LLM agents specifically for the legal domain.

Highlights

  • Comprehensive coverage with over 1,600 tasks across 24+ legal practice areas.
  • Includes a dedicated execution harness for running and scoring agents.
  • Uses an all-pass rubric scoring system and LLM judges for evaluation.
  • Open-source framework for comparing agent performance via dashboards.

Related

  • Project
  • Project
  • Project
  • Project
  • Project