aisa-group/PostTrainBench

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

What it solves

PostTrainBench addresses the challenge of automating the research and development (R&D) process for post-training large language models (LLMs). It provides a standardized benchmark to test whether AI agents can autonomously improve a base LLM's performance on specific tasks without human intervention.

How it works

The system sets up a controlled environment where a CLI-based agent (such as Claude Code or Gemini CLI) is given a base LLM, an evaluation script, and a fixed amount of compute resources (10 hours on an H100 GPU). The agent must then autonomously research, experiment, and post-train the model to maximize its score on a target benchmark.

Evaluation is conducted across seven diverse benchmarks covering:

  • Reasoning & Math: AIME 2025, GSM8K
  • Knowledge: GPQA, HealthBench Easy
  • Tool Use: BFCL
  • Coding: HumanEval
  • Writing: Arena Hard Writing

Who it’s for

This tool is designed for AI researchers and developers interested in agentic AI and automated machine learning (AutoML), specifically those looking to evaluate how well LLM agents can handle complex, multi-step software engineering and R&D tasks.

Highlights

  • Autonomous R&D: Tests agents on their ability to select data sources and training methods independently.
  • Capped Resources: Enforces a strict time limit (10 hours) and hardware constraint (H100 GPU) to simulate real-world constraints.
  • Diverse Task Set: Spans multiple domains from medical knowledge to competitive mathematics.
  • Multi-Agent Support: Compatible with various CLI scaffolds including Claude Code, Codex CLI, Gemini CLI, and OpenCode.

関連

  • プロジェクト
  • プロジェクト
  • プロジェクト
  • プロジェクト
  • プロジェクト