Terminal-Bench
The terminal-task benchmark coding agents now cite next to SWE-bench: hard, containerized terminal tasks scored end to end. Terminal-Bench 2.0 runs on the harbor evaluation framework; the 1.0 tasks live on in the org's terminal-bench-1 repo.
evalsclipython
- Stars
- 738
- Adoption surface
- slightly complex
- Autonomy
- headless
- Recovery
- none
- License
- ✅ open-source
Repository ↗ Example: Terminal-Bench leaderboard ↗