ClawBench

Open web-agent evaluation **harness**: runs selectable agents in isolated Docker containers across 153 live-site tasks (plus 130 in V2), intercepts irreversible requests, and records video, screenshots, HTTP traffic, actions, and agent messages for replayable scoring.

evalsvisionsandboxpython
Stars
801
Adoption surface
complex
Autonomy
headless
Recovery
none
License
✅ open-source
Category
Evaluation and benchmarking harnesses

Repository ↗ Example: ClawBench live leaderboard ↗

Related in Evaluation and benchmarking harnesses