Best of Agent Harnesses
Hand-curated, ranked list of 167 AI agent harnesses — the runtimes that close the loop between a stateless model and the outside world.
167 harnesses · 12 categories · MCP-ready · weekly-rescored
claude mcp add agent-harnesses -- uvx agent-harnesses-mcp
Browse all 167
| Project | Category | Stars | Tier | OSS |
|---|---|---|---|---|
| OpenClaw typescriptmulti-agent | Personal agent runtimes | 390k | complex | ✅ |
| superpowers memorycliide | Coding harness configs and SDKs | 289k | complex | ✅ |
| Hermes memorypythonprovider-agnostic | Personal agent runtimes | 247k | slightly complex | ✅ |
| opencode mcpprovider-agnosticclituitypescript | Coding agent products (IDEs, CLIs, full suites) | 209k | slightly complex | ✅ |
| n8n workflowlocaltypescript | Frameworks | 205k | complex | ⚠️ Fair-code |
| AutoGPT memoryevalspython | Frameworks | 187k | complex | ⚠️ Polyform-SU |
| Anthropic Skills | Coding harness configs and SDKs | 177k | mostly simple | ⚠️ Anthropic terms |
| Dify low-coderagpython | Frameworks | 157k | complex | ⚠️ Fair-code |
| langflow low-codepython | Frameworks | 155k | complex | ✅ |
| langchain python | Frameworks | 147k | complex | ✅ |
| GStack typescript | Coding harness configs and SDKs | 134k | slightly complex | ✅ |
| Codex sandboxprovider-agnosticcli | Coding agent products (IDEs, CLIs, full suites) | 125k | slightly complex | ✅ |
| browser-use browserpython | Frameworks | 116k | slightly complex | ✅ |
| pi provider-agnostictuirust | Coding agent products (IDEs, CLIs, full suites) | 108k | slightly complex | ❓ |
| Gemini CLI mcpclitypescript | Coding agent products (IDEs, CLIs, full suites) | 107k | slightly complex | ✅ |
| addyosmani/agent-skills workflowide | Coding harness configs and SDKs | 97.5k | mostly simple | ✅ |
| claude-mem memory | Memory and state | 94.3k | slightly complex | ✅ |
| MCP Servers mcpmemorytypescript | Plugins, MCPs, CLI tools | 90.5k | mostly simple | ✅ |
| OpenHands memorybrowsersandboxpython | Coding agent products (IDEs, CLIs, full suites) | 88.6k | complex | ⚠️ (multi-license) |
| DeerFlow memorymulti-agentsandboxpython | Research and task-specific harnesses | 82.8k | complex | ✅ |
| Headroom mcprag | Progressive disclosure harnesses | 73.2k | mostly simple | ✅ |
| Daytona sandbox | Libraries and SDKs | 71.7k | slightly complex | ✅ |
| MetaGPT multi-agentpython | Multi-agent and orchestration | 70.5k | complex | ✅ |
| Cline idetypescript | Coding agent products (IDEs, CLIs, full suites) | 68.9k | slightly complex | ✅ |
| Open Interpreter clipython | Coding agent products (IDEs, CLIs, full suites) | 68.4k | mostly simple | ✅ |
| AnythingLLM ragtypescript | Personal agent runtimes | 66.3k | complex | ✅ |
| Mem0 memorypython | Memory and state | 65.7k | slightly complex | ✅ |
| Context7 mcptrainingtypescript | Plugins, MCPs, CLI tools | 62.2k | super simple | ✅ |
| autogen multi-agentpython | Multi-agent and orchestration | 61.1k | complex | ✅ CC-BY |
| LiteLLM provider-agnosticpython | Libraries and SDKs | 59.2k | mostly simple | ✅ |
| crewAI python | Multi-agent and orchestration | 58.8k | complex | ✅ |
| OpenManus multi-agentpython | Multi-agent and orchestration | 58.4k | complex | ✅ |
| Flowise low-codetypescript | Frameworks | 55.5k | complex | ⚠️ Apache+CLA |
| goose mcprust | Coding agent products (IDEs, CLIs, full suites) | 54.5k | slightly complex | ✅ |
| awesome-claude-code | Coding harness configs and SDKs | 54.3k | super simple | ❓ |
| chrome-devtools-mcp mcpbrowsertypescript | Plugins, MCPs, CLI tools | 52.4k | mostly simple | ✅ |
| llama-index ragpython | Frameworks | 52.2k | complex | ✅ |
| aider mcpclipython | Plugins, MCPs, CLI tools | 49.1k | slightly complex | ✅ |
| nanobot mcpmemorylocalpython | Personal agent runtimes | 48.4k | mostly simple | ❓ |
| CowAgent memorypython | Personal agent runtimes | 47.1k | slightly complex | ❓ |
| agno memoryevalspython | Frameworks | 42.3k | complex | ✅ |
| langgraph workflowpython | Frameworks | 42k | slightly complex | ✅ |
| awesome-cursorrules ide | Progressive disclosure harnesses | 40.8k | super simple | ✅ |
| wshobson/agents multi-agentcliide | Coding harness configs and SDKs | 39.8k | super simple | ✅ |
| Khoj python | Personal agent runtimes | 37.4k | complex | ✅ |
| Playwright MCP mcpvisionbrowsertypescript | Plugins, MCPs, CLI tools | 37.4k | mostly simple | ✅ |
| continue idetypescript | Plugins, MCPs, CLI tools | 36k | complex | ✅ |
| DeepSeek-Reasonix memoryclituitypescript | Coding agent products (IDEs, CLIs, full suites) | 35.6k | slightly complex | ❓ |
| Langfuse evalstypescript | Observability and eval-ops | 34.9k | slightly complex | ✅ |
| ChatDev python | Multi-agent and orchestration | 34.3k | slightly complex | ✅ |
| github-mcp-server mcp | Plugins, MCPs, CLI tools | 33.1k | slightly complex | ✅ |
| oh-my-pi browserprovider-agnosticcliiderust | Coding agent products (IDEs, CLIs, full suites) | 32.1k | slightly complex | ✅ |
| Graphiti (Zep) memoryragworkflowpython | Memory and state | 31k | slightly complex | ✅ |
| cognee memoryragworkflowpython | Memory and state | 30.9k | slightly complex | ✅ |
| Composio sandboxtool-discoverypythontypescript | Libraries and SDKs | 30.3k | complex | ✅ |
| deepagents multi-agentsandboxpythontypescript | Libraries and SDKs | 29.6k | slightly complex | ✅ |
| openai-agents-python python | Multi-agent and orchestration | 29.6k | mostly simple | ✅ |
| gpt-researcher multi-agentpython | Research and task-specific harnesses | 29.5k | complex | ✅ |
| smolagents sandboxpython | Libraries and SDKs | 29.4k | mostly simple | ✅ |
| semantic-kernel python | Frameworks | 28.6k | complex | ✅ |
| mastra typedtypescript | Frameworks | 28.2k | slightly complex | ⚠️ Elastic-2.0 |
| crush memoryclitui | Coding agent products (IDEs, CLIs, full suites) | 28.2k | slightly complex | ⚠️ FSL-1.1-MIT |
| vibe-kanban | Coding agent products (IDEs, CLIs, full suites) | 28.1k | slightly complex | ❓ |
| MLflow evalspython | Observability and eval-ops | 28.1k | complex | ✅ |
| qwen-code sandboxclitypescript | Coding agent products (IDEs, CLIs, full suites) | 28k | slightly complex | ❓ |
| Kilo Code mcpcliidetypescript | Coding agent products (IDEs, CLIs, full suites) | 27.4k | slightly complex | ❓ |
| beads memory | Memory and state | 27.3k | mostly simple | ❓ |
| Symphony sandbox | Coding agent products (IDEs, CLIs, full suites) | 27.3k | complex | ❓ |
| planning-with-files memory | Coding harness configs and SDKs | 27k | mostly simple | ❓ |
| vercel/ai provider-agnostictypescript | Libraries and SDKs | 26.9k | slightly complex | ✅ |
| Haystack memoryragpython | Frameworks | 26.6k | complex | ✅ |
| letta memorypython | Frameworks | 24.8k | mostly simple | ✅ |
| Stagehand browsertypescript | Frameworks | 24.6k | slightly complex | ✅ |
| agents.md idetypescript | Progressive disclosure harnesses | 24.5k | super simple | ✅ |
| MCP Python SDK mcppython | Plugins, MCPs, CLI tools | 24.4k | mostly simple | ✅ |
| Roo Code mcpworkflowidetypescript | Coding agent products (IDEs, CLIs, full suites) | 24.3k | slightly complex | ✅ |
| context-mode mcpmemorysandbox | Progressive disclosure harnesses | 23.8k | mostly simple | ⚠️ Elastic-2.0 |
| Opik evalspython | Observability and eval-ops | 22.2k | slightly complex | ✅ |
| Google ADK evalssandboxpython | Frameworks | 21.6k | complex | ✅ |
| rasa voicepython | Frameworks | 21.3k | complex | ✅ |
| Prime Agent memorymulti-agentclitypescript | Coding agent products (IDEs, CLIs, full suites) | 21.1k | slightly complex | ✅ |
| SWE-agent memoryevalspython | Coding harness configs and SDKs | 20.4k | slightly complex | ✅ |
| pydantic-ai mcptypedprovider-agnosticpython | Libraries and SDKs | 20.1k | slightly complex | ✅ |
| jcode mcpmemoryprovider-agnosticclirust | Coding agent products (IDEs, CLIs, full suites) | 19.9k | slightly complex | ❓ |
| Eliza memorymulti-agenttypescript | Personal agent runtimes | 19.4k | complex | ✅ |
| Agent Zero memorymulti-agentbrowsersandboxpython | Personal agent runtimes | 19.2k | slightly complex | ❓ |
| Agent Lightning evalstrainingpython | Evaluation and benchmarking harnesses | 18.4k | complex | ✅ |
| OpenHarness (HKUDS) memorymulti-agent | Personal agent runtimes | 15.8k | complex | ✅ |
| eigent multi-agentlocal | Coding agent products (IDEs, CLIs, full suites) | 15.3k | complex | ❓ |
| QM memorysandboxtypescript | Personal agent runtimes | 15.2k | complex | ✅ |
| botpress low-codetypescript | Frameworks | 14.9k | complex | ✅ |
| cc-haha memorymulti-agenttypescript | Coding agent products (IDEs, CLIs, full suites) | 14.7k | complex | ❓ |
| AutoResearchClaw multi-agent | Research and task-specific harnesses | 14.5k | complex | ❓ |
| E2B sandboxpython | Libraries and SDKs | 13.9k | slightly complex | ✅ |
| Microsoft Agent Framework multi-agentworkflowpython | Multi-agent and orchestration | 13.6k | slightly complex | ✅ |
| MCP TypeScript SDK mcptypescript | Plugins, MCPs, CLI tools | 13.4k | mostly simple | ✅ |
| Arize Phoenix evalspython | Observability and eval-ops | 11.5k | slightly complex | ⚠️ Elastic-2.0 |
| hive multi-agentpython | Multi-agent and orchestration | 11.1k | complex | ❓ |
| MCP Inspector mcptypescript | Plugins, MCPs, CLI tools | 10.9k | super simple | ✅ |
| omnigent sandboxidepython | Multi-agent and orchestration | 10.1k | complex | ❓ |
| OpenJarvis memorylocalpython | Personal agent runtimes | 10k | slightly complex | ✅ |
| get-shit-done clipython | Coding harness configs and SDKs | 9.7k | mostly simple | ✅ |
| PraisonAI multi-agentpython | Multi-agent and orchestration | 9.1k | mostly simple | ✅ |
| MiroThinker evals | Research and task-specific harnesses | 8.4k | slightly complex | ❓ |
| Claude Agent SDK mcpmemorypythontypescript | Coding harness configs and SDKs | 8.1k | complex | ✅ |
| R2R visionragworkflowpython | Frameworks | 8k | complex | ✅ |
| agent-squad multi-agent | Frameworks | 7.8k | slightly complex | ✅ |
| Steel memorybrowserlocal | Libraries and SDKs | 7.7k | slightly complex | ✅ |
| strands-agents mcpmulti-agenttypedpython | Libraries and SDKs | 7.4k | mostly simple | ✅ |
| MCP Registry mcp | Plugins, MCPs, CLI tools | 7.3k | slightly complex | ✅ |
| Agent Governance Toolkit sandboxpython | Plugins, MCPs, CLI tools | 6.3k | slightly complex | ✅ |
| agents-cli evalscli | Coding harness configs and SDKs | 6k | mostly simple | ❓ |
| SWE-bench evalssandboxpython | Evaluation and benchmarking harnesses | 5.9k | slightly complex | ✅ |
| Cloudflare Agents memorytypescript | Libraries and SDKs | 5.6k | slightly complex | ✅ |
| AgentVerse multi-agentpython | Frameworks | 5.1k | complex | ✅ |
| skillhub localcliide | Coding harness configs and SDKs | 5.1k | mostly simple | ❓ |
| AG2 multi-agentpython | Multi-agent and orchestration | 4.9k | complex | ❓ |
| youtu-agent | Frameworks | 4.6k | mostly simple | ❓ |
| mcp-context-forge mcppython | Plugins, MCPs, CLI tools | 4.5k | complex | ❓ |
| Agent Sandbox memorysandboxlocal | Libraries and SDKs | 4k | slightly complex | ✅ |
| openai-agents-js multi-agentvoicetypescript | Libraries and SDKs | 3.8k | slightly complex | ✅ |
| AgentBench evalssandboxragworkflowpython | Evaluation and benchmarking harnesses | 3.7k | complex | ✅ |
| Bee Agent Framework mcpmulti-agentpythontypescript | Frameworks | 3.4k | complex | ✅ |
| inspect_ai evalssandboxpython | Evaluation and benchmarking harnesses | 2.8k | complex | ✅ |
| cocoindex-code mcpcli | Plugins, MCPs, CLI tools | 2.7k | mostly simple | ❓ |
| agent-vault | Plugins, MCPs, CLI tools | 2.2k | mostly simple | ❓ |
| AgentStack | Frameworks | 2.2k | slightly complex | ✅ |
| WebArena python | Evaluation and benchmarking harnesses | 1.6k | complex | ✅ |
| Meta-Harness | Coding harness configs and SDKs | 1.6k | slightly complex | ❓ |
| Docker MCP Gateway mcpsandboxcli | Plugins, MCPs, CLI tools | 1.6k | slightly complex | ✅ |
| AIlice sandboxpython | Personal agent runtimes | 1.4k | slightly complex | ✅ |
| WebVoyager evalsvision | Evaluation and benchmarking harnesses | 1.1k | slightly complex | ✅ |
| agent-qa mcpmemorysandboxclitypescript | Evaluation and benchmarking harnesses | 887 | slightly complex | ⚠️ FSL-1.1-ALv2 |
| ClawBench evalsvisionsandboxpython | Evaluation and benchmarking harnesses | 801 | complex | ✅ |
| swe-smith trainingpython | Evaluation and benchmarking harnesses | 775 | slightly complex | ✅ |
| SWE-Gym evalstrainingpython | Evaluation and benchmarking harnesses | 742 | slightly complex | ✅ |
| ARC-AGI-2 | Evaluation and benchmarking harnesses | 738 | super simple | ✅ |
| Terminal-Bench evalsclipython | Evaluation and benchmarking harnesses | 738 | slightly complex | ✅ |
| inspect_evals evalssandbox | Evaluation and benchmarking harnesses | 675 | slightly complex | ✅ |
| open-harness mcpmulti-agenttypescript | Libraries and SDKs | 612 | slightly complex | ✅ |
| langgraph-bigtool tool-discoverypython | Progressive disclosure harnesses | 558 | slightly complex | ✅ |
| RepoMaster workflowpython | Coding harness configs and SDKs | 553 | slightly complex | ❓ |
| claw-code-agent mcprustpythontypescript | Coding agent products (IDEs, CLIs, full suites) | 545 | slightly complex | ❓ |
| MCP-Zero tool-discovery | Progressive disclosure harnesses | 512 | complex | ✅ |
| Proliferate multi-agentsandboxidetypescript | Coding agent products (IDEs, CLIs, full suites) | 505 | complex | ✅ |
| AgentBox sandboxtypescript | Coding agent products (IDEs, CLIs, full suites) | 467 | slightly complex | ✅ |
| AgentSilex python | Frameworks | 456 | super simple | ✅ |
| openagents | Research and task-specific harnesses | 450 | complex | ✅ |
| AutoHarness memorymulti-agentprovider-agnosticpython | Coding harness configs and SDKs | 378 | super simple | ✅ |
| arc-agi-benchmarking evalsprovider-agnosticpython | Evaluation and benchmarking harnesses | 362 | mostly simple | ✅ |
| AgentRL trainingpython | Multi-agent and orchestration | 354 | complex | ✅ |
| SuperAgentX multi-agentpython | Frameworks | 203 | mostly simple | ✅ |
| ToolGen tool-discoverypython | Progressive disclosure harnesses | 184 | complex | ❓ |
| VitaBench | Evaluation and benchmarking harnesses | 177 | complex | ✅ |
| LoopTroop typescript | Coding harness configs and SDKs | 150 | mostly simple | ✅ |
| AgencyBench evalssandboxpython | Evaluation and benchmarking harnesses | 100 | complex | ✅ |
| Talon mcpmemoryclitypescript | Personal agent runtimes | 83 | slightly complex | ✅ |
| letta-evals memorypython | Evaluation and benchmarking harnesses | 82 | mostly simple | ✅ |
| YYLO multi-agenttypescript | Coding agent products (IDEs, CLIs, full suites) | 60 | slightly complex | ✅ |
| SUPER sandboxpython | Evaluation and benchmarking harnesses | 58 | slightly complex | ✅ |
| ToolRAG mcptool-discovery | Progressive disclosure harnesses | 34 | mostly simple | ✅ |
| puppeteer-real-browser-mcp mcpbrowsertypescript | Plugins, MCPs, CLI tools | 26 | mostly simple | ❓ |
| TRAIL | Evaluation and benchmarking harnesses | 24 | mostly simple | ✅ |
| Community-curated agent lists | Libraries and SDKs | 15 | super simple | ❓ |
| Better-OpenCodeMCP mcptypescript | Plugins, MCPs, CLI tools | 9 | mostly simple | ✅ |
| pmstack evals | Coding harness configs and SDKs | 8 | super simple | ✅ |
| agentlog memoryclipython | Plugins, MCPs, CLI tools | 1 | super simple | ✅ |
Categories
Progressive disclosure harnesses8 projectsCoding agent products (IDEs, CLIs, full suites)24 projectsCoding harness configs and SDKs17 projectsPersonal agent runtimes13 projectsFrameworks26 projectsMulti-agent and orchestration12 projectsPlugins, MCPs, CLI tools19 projectsMemory and state5 projectsEvaluation and benchmarking harnesses19 projectsObservability and eval-ops4 projectsResearch and task-specific harnesses5 projectsLibraries and SDKs15 projects
Pick by use case
- I want a turnkey coding agent today — opencode, Cline, Codex, Gemini CLI, OpenHands
- I want an always-on personal agent that lives in my chat apps — OpenClaw, Hermes, Khoj, Agent Zero, OpenHarness (HKUDS)
- I want to extend Claude Code, Codex, or OpenCode with skills and slash commands — Anthropic Skills, wshobson/agents, superpowers, GStack, pmstack
- I want to build my own coding harness from scratch — Claude Agent SDK, Google ADK, AutoHarness, SWE-agent, RepoMaster
- I want a drop-in memory layer for agents — Mem0, Graphiti (Zep), claude-mem, agentlog, letta
- I want to plug hundreds to thousands of tools without context bloat — MCP-Zero, ToolGen, ToolRAG, langgraph-bigtool
- I want multi-agent orchestration — openai-agents-python, crewAI, autogen, Microsoft Agent Framework, PraisonAI
- I want a general LLM app framework — langgraph, langchain, llama-index, pydantic-ai, agno
- I want low-code / visual workflows — langflow, Flowise, Dify, n8n
- I want browser-using agents — browser-use, Stagehand, WebVoyager, puppeteer-real-browser-mcp
- I want sandboxed code execution for agent-generated code — E2B, Agent Sandbox, Daytona, smolagents, OpenHands
- I want to evaluate or benchmark agents — SWE-bench, Terminal-Bench, AgencyBench, inspect_ai, WebArena
- I want a deep research / autonomous research agent — deepagents, gpt-researcher, openagents
- I want a provider-agnostic LLM pipe (not a framework) — LiteLLM, vercel/ai
Decision guides
- Agent evals: SWE-bench vs inspect_ai vs Terminal-Bench — You changed your agent: new model, new prompt, new tools. Did it get better or worse? An eval is how you answer that with a number instead o
- The best AI agent harnesses in 2026, ranked by category — These are the top three agent harnesses in each of the 12 categories of best-of-Agent-Harnesses, a hand-curated list of 167 harnesses re-ran
- Browser agents: browser-use vs Stagehand vs Playwright MCP vs chrome-devtools-mcp — "Browser agent" covers three different kinds of product, and most bad picks here come from comparing across the lanes instead of within one.
- Browser infrastructure for agents: Browserbase vs Steel vs Hyperbrowser — Agent libraries like browser-use and Stagehand decide what to click; something still has to run the browsers they click in. At small scale t
- Claude Code skill packs: superpowers vs GStack vs get-shit-done vs Anthropic Skills — A skill is a folder of instructions (a SKILL.md file, plus any scripts it needs) that a coding agent loads only when the task matches, inste
- Eval and observability platforms: Langfuse vs LangSmith vs Braintrust vs Phoenix — Benchmarks tell you how a model ranks; your production agent still fails in ways no public exam covers. An eval and observability platform i
- How to pick a harness — This is the decision guide for [best-of-Agent-Harnesses](../README.md), a curated, ranked list of the runtimes that turn an AI model into a
- How to test-drive a harness — Spec sheets cannot answer "which harness should I use," because an agent's performance is a property of the *pairing* between harness and mo
- Managed vs self-hosted always-on agents — An always-on agent keeps working between your messages, and in 2026 you can rent one (Grok Bot, Claude Managed Agents), run one for your who
- Agent memory layers: Mem0 vs Zep vs Letta vs claude-mem — Agents forget. A model keeps nothing between sessions, so anything your agent should still know tomorrow (who the user is, what was decided,
- Multi-agent orchestration: OpenAI Agents SDK vs CrewAI vs AutoGen vs Agent Framework vs LangGraph — Orchestration is the layer that coordinates several AI agents working on one job: who acts next, what they share, and what happens when a st
- OpenClaw vs Hermes: the always-on personal-agent debate — An always-on personal agent is a program that runs all day on your own machine, talks to you through the chat apps you already use (WhatsApp
- Context files for agents: AGENTS.md vs CLAUDE.md vs skills vs MCP tool search — A model has a context window: a fixed amount of text it can consider at once. Everything competes for that space: your instructions, the def
- Agent sandboxing: what it is and how to pick — An AI agent does not just suggest code. It runs code, opens web pages, and edits files on a real computer. Agent sandboxing means making tha
- Terminal coding agents: opencode vs Codex vs Gemini CLI vs crush vs goose — The most-asked pick in this list: *"I want a turnkey coding agent in my terminal today."* A terminal coding agent is a program you run in yo
- Why the harness matters more than the model — The same model weights score about 30% on ARC-AGI-3 as a bare model and 95.5% inside a good harness, and that gap is the number agent builde
For agents
This list is published machine-readable so coding and research agents can recommend harnesses directly:
- harnesses.json — every project with tier, tags, axes, license, example, and use-case index.
- llms.txt — the whole list in one agent-readable file.
- llms-full.txt: llms.txt plus the full text of every decision guide, one file.
- harnesses.jsonld — schema.org Dataset + ItemList.
- feed.json — JSON Feed of refreshes.
- MCP server:
uvx agent-harnesses-mcp— pick_harness, search_harnesses, get_harness, comparisons.