ClawBench
Open web-agent evaluation **harness**: runs selectable agents in isolated Docker containers across 153 live-site tasks (plus 130 in V2), intercepts irreversible requests, and records video, screenshots, HTTP traffic, actions, and agent messages for replayable scoring.
evalsvisionsandboxpython
- Stars
- 801
- Adoption surface
- complex
- Autonomy
- headless
- Recovery
- none
- License
- ✅ open-source
Repository ↗ Example: ClawBench live leaderboard ↗