Terminal-Bench

The terminal-task benchmark coding agents now cite next to SWE-bench: hard, containerized terminal tasks scored end to end. Terminal-Bench 2.0 runs on the harbor evaluation framework; the 1.0 tasks live on in the org's terminal-bench-1 repo.

evalsclipython
Stars
738
Adoption surface
slightly complex
Autonomy
headless
Recovery
none
License
✅ open-source
Category
Evaluation and benchmarking harnesses

Repository ↗ Example: Terminal-Bench leaderboard ↗

Related in Evaluation and benchmarking harnesses