Best of Agent Harnesses
Hand-curated, ranked list of 154 AI agent harnesses — the runtimes that close the loop between a stateless model and the outside world.
154 harnesses · 12 categories · MCP-ready · weekly-rescored
claude mcp add agent-harnesses -- uvx agent-harnesses-mcp
Browse all 154
| Project | Category | Stars | Tier | OSS |
|---|---|---|---|---|
| OpenClaw typescriptmulti-agent | Personal agent runtimes | 385k | complex | ✅ |
| superpowers memoryide | Coding harness configs and SDKs | 265k | complex | ✅ |
| Hermes memorypythonprovider-agnostic | Personal agent runtimes | 224k | slightly complex | ✅ |
| n8n workflowlocaltypescript | Frameworks | 199k | complex | ⚠️ Fair-code |
| opencode mcpprovider-agnosticclituitypescript | Coding agent products (IDEs, CLIs, full suites) | 192k | slightly complex | ✅ |
| AutoGPT memoryevalspython | Frameworks | 186k | complex | ⚠️ Polyform-SU |
| Anthropic Skills | Coding harness configs and SDKs | 166k | mostly simple | ✅ |
| langflow low-codepython | Frameworks | 153k | complex | ✅ |
| Dify low-coderagpython | Frameworks | 151k | complex | ⚠️ Fair-code |
| langchain python | Frameworks | 143k | complex | ✅ |
| GStack typescript | Coding harness configs and SDKs | 126k | slightly complex | ✅ |
| browser-use mcpbrowserpython | Frameworks | 108k | slightly complex | ✅ |
| Gemini CLI mcpclitypescript | Coding agent products (IDEs, CLIs, full suites) | 106k | slightly complex | ✅ |
| Codex sandboxprovider-agnosticcli | Coding agent products (IDEs, CLIs, full suites) | 103k | slightly complex | ✅ |
| claude-mem memory | Memory and state | 89.3k | slightly complex | ✅ |
| MCP Servers mcpmemorytypescript | Plugins, MCPs, CLI tools | 89.1k | mostly simple | ✅ |
| OpenHands memorybrowsersandboxpython | Coding agent products (IDEs, CLIs, full suites) | 82.9k | complex | ⚠️ (multi-license) |
| pi provider-agnostictuirust | Coding agent products (IDEs, CLIs, full suites) | 82.2k | slightly complex | ❓ |
| addyosmani/agent-skills workflowide | Coding harness configs and SDKs | 81.3k | mostly simple | ✅ |
| DeerFlow memorymulti-agentsandboxpython | Research and task-specific harnesses | 78.9k | complex | ✅ |
| Daytona sandbox | Libraries and SDKs | 72.1k | slightly complex | ✅ |
| MetaGPT multi-agentpython | Multi-agent and orchestration | 69.6k | complex | ✅ |
| Open Interpreter clipython | Coding agent products (IDEs, CLIs, full suites) | 67.5k | mostly simple | ✅ |
| Cline idetypescript | Coding agent products (IDEs, CLIs, full suites) | 65.5k | slightly complex | ✅ |
| Headroom mcprag | Progressive disclosure harnesses | 64k | mostly simple | ✅ |
| Mem0 memorypython | Memory and state | 62.3k | slightly complex | ✅ |
| autogen multi-agentpython | Multi-agent and orchestration | 60.2k | complex | ✅ CC-BY |
| Context7 mcptrainingtypescript | Plugins, MCPs, CLI tools | 60.2k | super simple | ✅ |
| OpenManus multi-agentpython | Multi-agent and orchestration | 57.8k | complex | ✅ |
| crewAI python | Multi-agent and orchestration | 56.5k | complex | ✅ |
| LiteLLM provider-agnosticpython | Libraries and SDKs | 55.4k | mostly simple | ✅ |
| Flowise low-codetypescript | Frameworks | 55.1k | complex | ⚠️ Apache+CLA |
| goose mcprust | Coding agent products (IDEs, CLIs, full suites) | 52.1k | slightly complex | ✅ |
| awesome-claude-code | Coding harness configs and SDKs | 51.5k | super simple | ❓ |
| llama-index ragpython | Frameworks | 51.3k | complex | ✅ |
| chrome-devtools-mcp mcpbrowsertypescript | Plugins, MCPs, CLI tools | 48.4k | mostly simple | ❓ |
| aider mcpclipython | Plugins, MCPs, CLI tools | 47.9k | slightly complex | ✅ |
| nanobot mcpmemorylocalpython | Personal agent runtimes | 46.5k | mostly simple | ❓ |
| CowAgent memorypython | Personal agent runtimes | 46.3k | slightly complex | ❓ |
| agno memoryevalspython | Frameworks | 41.5k | complex | ✅ |
| awesome-cursorrules ide | Progressive disclosure harnesses | 40.5k | super simple | ✅ |
| langgraph workflowpython | Frameworks | 38.7k | slightly complex | ✅ |
| wshobson/agents multi-agentcliide | Coding harness configs and SDKs | 38.4k | super simple | ✅ |
| Khoj python | Personal agent runtimes | 36.2k | complex | ✅ |
| Playwright MCP mcpvisionbrowsertypescript | Plugins, MCPs, CLI tools | 35.7k | mostly simple | ✅ |
| continue idetypescript | Plugins, MCPs, CLI tools | 35.3k | complex | ✅ |
| ChatDev python | Multi-agent and orchestration | 33.9k | slightly complex | ✅ |
| Langfuse evalstypescript | Observability and eval-ops | 32.3k | slightly complex | ✅ |
| github-mcp-server mcp | Plugins, MCPs, CLI tools | 31.9k | slightly complex | ✅ |
| cognee memoryragworkflowpython | Memory and state | 29.7k | slightly complex | ✅ |
| Composio sandboxtool-discoverypythontypescript | Libraries and SDKs | 29.5k | complex | ✅ |
| DeepSeek-Reasonix memoryclituitypescript | Coding agent products (IDEs, CLIs, full suites) | 28.8k | slightly complex | ❓ |
| gpt-researcher multi-agentpython | Research and task-specific harnesses | 28.8k | complex | ✅ |
| smolagents sandboxpython | Libraries and SDKs | 28.6k | mostly simple | ✅ |
| semantic-kernel python | Frameworks | 28.4k | complex | ✅ |
| openai-agents-python python | Multi-agent and orchestration | 28.3k | mostly simple | ✅ |
| vibe-kanban | Coding agent products (IDEs, CLIs, full suites) | 27.6k | slightly complex | ❓ |
| MLflow evalspython | Observability and eval-ops | 27.3k | complex | ✅ |
| deepagents multi-agentsandboxpythontypescript | Libraries and SDKs | 27.2k | slightly complex | ✅ |
| crush memoryclitui | Coding agent products (IDEs, CLIs, full suites) | 27k | slightly complex | ⚠️ FSL-1.1-MIT |
| mastra typedtypescript | Frameworks | 26.8k | slightly complex | ⚠️ Elastic-2.0 |
| Kilo Code mcpcliidetypescript | Coding agent products (IDEs, CLIs, full suites) | 26.7k | slightly complex | ❓ |
| qwen-code sandboxclitypescript | Coding agent products (IDEs, CLIs, full suites) | 26.5k | slightly complex | ❓ |
| Symphony sandbox | Coding agent products (IDEs, CLIs, full suites) | 26.4k | complex | ❓ |
| Haystack memoryragpython | Frameworks | 26.1k | complex | ✅ |
| vercel/ai provider-agnostictypescript | Libraries and SDKs | 26k | slightly complex | ✅ |
| planning-with-files memory | Coding harness configs and SDKs | 25.9k | mostly simple | ❓ |
| beads memory | Memory and state | 25.8k | mostly simple | ❓ |
| Roo Code mcpworkflowidetypescript | Coding agent products (IDEs, CLIs, full suites) | 24.4k | slightly complex | ✅ |
| letta memorypython | Frameworks | 24.1k | mostly simple | ✅ |
| MCP Python SDK mcppython | Plugins, MCPs, CLI tools | 23.8k | mostly simple | ✅ |
| agents.md typescript | Progressive disclosure harnesses | 23.4k | super simple | ✅ |
| rasa voicepython | Frameworks | 21.3k | complex | ✅ |
| oh-my-pi browserprovider-agnosticcliiderust | Coding agent products (IDEs, CLIs, full suites) | 21.2k | slightly complex | ✅ |
| Google ADK evalssandboxpython | Frameworks | 21k | complex | ✅ |
| SWE-agent memoryevalspython | Coding harness configs and SDKs | 20k | slightly complex | ✅ |
| context-mode mcpmemorysandbox | Progressive disclosure harnesses | 19.6k | mostly simple | ❓ |
| pydantic-ai mcptypedprovider-agnosticpython | Libraries and SDKs | 19k | slightly complex | ✅ |
| Eliza memorymulti-agenttypescript | Personal agent runtimes | 18.9k | complex | ✅ |
| Agent Zero memorymulti-agentbrowsersandboxpython | Personal agent runtimes | 18.7k | slightly complex | ❓ |
| Agent Lightning evalstrainingpython | Evaluation and benchmarking harnesses | 17.4k | complex | ✅ |
| OpenHarness (HKUDS) memorymulti-agent | Personal agent runtimes | 15.2k | complex | ✅ |
| jcode mcpmemoryprovider-agnosticclirust | Coding agent products (IDEs, CLIs, full suites) | 15.2k | slightly complex | ❓ |
| botpress low-codetypescript | Frameworks | 14.8k | complex | ✅ |
| eigent multi-agentlocal | Coding agent products (IDEs, CLIs, full suites) | 14.7k | complex | ❓ |
| AutoResearchClaw multi-agent | Research and task-specific harnesses | 13.9k | complex | ❓ |
| cc-haha memorymulti-agenttypescript | Coding agent products (IDEs, CLIs, full suites) | 13.9k | complex | ❓ |
| E2B sandboxpython | Libraries and SDKs | 13.2k | slightly complex | ✅ |
| MCP TypeScript SDK mcptypescript | Plugins, MCPs, CLI tools | 13k | mostly simple | ✅ |
| Microsoft Agent Framework multi-agentworkflowpython | Multi-agent and orchestration | 12.5k | slightly complex | ✅ |
| hive multi-agentpython | Multi-agent and orchestration | 10.8k | complex | ❓ |
| MCP Inspector mcptypescript | Plugins, MCPs, CLI tools | 10.6k | super simple | ✅ |
| PraisonAI multi-agentpython | Multi-agent and orchestration | 8.5k | mostly simple | ✅ |
| MiroThinker evals | Research and task-specific harnesses | 8.4k | slightly complex | ❓ |
| omnigent sandboxidepython | Multi-agent and orchestration | 8k | complex | ❓ |
| R2R visionragworkflowpython | Frameworks | 8k | complex | ✅ |
| Claude Agent SDK mcpmemorypythontypescript | Coding harness configs and SDKs | 7.8k | complex | ✅ |
| agent-squad multi-agent | Frameworks | 7.7k | slightly complex | ✅ |
| get-shit-done clipython | Coding harness configs and SDKs | 7.6k | mostly simple | ✅ |
| MCP Registry mcp | Plugins, MCPs, CLI tools | 7.1k | slightly complex | ✅ |
| strands-agents mcpmulti-agenttypedpython | Libraries and SDKs | 6.8k | mostly simple | ✅ |
| Agent Governance Toolkit sandboxpython | Plugins, MCPs, CLI tools | 5.6k | slightly complex | ✅ |
| SWE-bench evalssandboxpython | Evaluation and benchmarking harnesses | 5.5k | slightly complex | ✅ |
| agents-cli evalscli | Coding harness configs and SDKs | 5.5k | mostly simple | ❓ |
| Cloudflare Agents memorytypescript | Libraries and SDKs | 5.3k | slightly complex | ✅ |
| AgentVerse multi-agentpython | Frameworks | 5.1k | complex | ✅ |
| skillhub localcliide | Coding harness configs and SDKs | 4.8k | mostly simple | ❓ |
| AG2 multi-agentpython | Multi-agent and orchestration | 4.8k | complex | ❓ |
| youtu-agent | Frameworks | 4.6k | mostly simple | ❓ |
| mcp-context-forge mcppython | Plugins, MCPs, CLI tools | 4.2k | complex | ❓ |
| AgentBench evalssandboxragworkflowpython | Evaluation and benchmarking harnesses | 3.6k | complex | ✅ |
| openai-agents-js multi-agentvoicetypescript | Libraries and SDKs | 3.5k | slightly complex | ✅ |
| Bee Agent Framework mcpmulti-agentpythontypescript | Frameworks | 3.3k | complex | ✅ |
| cocoindex-code mcpcli | Plugins, MCPs, CLI tools | 2.6k | mostly simple | ❓ |
| inspect_ai evalssandboxpython | Evaluation and benchmarking harnesses | 2.4k | complex | ✅ |
| AgentStack | Frameworks | 2.2k | slightly complex | ✅ |
| agent-vault | Plugins, MCPs, CLI tools | 2k | mostly simple | ❓ |
| WebArena python | Evaluation and benchmarking harnesses | 1.6k | complex | ✅ |
| Docker MCP Gateway mcpsandboxcli | Plugins, MCPs, CLI tools | 1.5k | slightly complex | ✅ |
| AIlice sandboxpython | Personal agent runtimes | 1.4k | slightly complex | ✅ |
| Meta-Harness | Coding harness configs and SDKs | 1.4k | slightly complex | ❓ |
| WebVoyager evalsvision | Evaluation and benchmarking harnesses | 1.1k | slightly complex | ✅ |
| ARC-AGI-2 | Evaluation and benchmarking harnesses | 731 | super simple | ✅ |
| swe-smith trainingpython | Evaluation and benchmarking harnesses | 722 | slightly complex | ✅ |
| SWE-Gym evalstrainingpython | Evaluation and benchmarking harnesses | 714 | slightly complex | ✅ |
| inspect_evals evalssandbox | Evaluation and benchmarking harnesses | 611 | slightly complex | ✅ |
| open-harness mcpmulti-agenttypescript | Libraries and SDKs | 592 | slightly complex | ✅ |
| langgraph-bigtool tool-discoverypython | Progressive disclosure harnesses | 552 | slightly complex | ✅ |
| RepoMaster workflowpython | Coding harness configs and SDKs | 541 | slightly complex | ❓ |
| claw-code-agent mcprustpythontypescript | Coding agent products (IDEs, CLIs, full suites) | 537 | slightly complex | ❓ |
| MCP-Zero tool-discovery | Progressive disclosure harnesses | 503 | complex | ✅ |
| AgentSilex python | Frameworks | 453 | super simple | ✅ |
| openagents | Research and task-specific harnesses | 444 | complex | ✅ |
| AutoHarness memorymulti-agentprovider-agnosticpython | Coding harness configs and SDKs | 363 | super simple | ✅ |
| arc-agi-benchmarking evalsprovider-agnosticpython | Evaluation and benchmarking harnesses | 356 | mostly simple | ✅ |
| AgentBox sandboxtypescript | Coding agent products (IDEs, CLIs, full suites) | 332 | slightly complex | ✅ |
| AgentRL trainingpython | Multi-agent and orchestration | 330 | complex | ✅ |
| SuperAgentX multi-agentpython | Frameworks | 202 | mostly simple | ✅ |
| ToolGen tool-discoverypython | Progressive disclosure harnesses | 183 | complex | ❓ |
| agent-qa mcpmemorysandboxclitypescript | Evaluation and benchmarking harnesses | 170 | slightly complex | ⚠️ FSL-1.1-ALv2 |
| VitaBench | Evaluation and benchmarking harnesses | 162 | complex | ✅ |
| Proliferate multi-agentsandboxidetypescript | Coding agent products (IDEs, CLIs, full suites) | 156 | complex | ✅ |
| LoopTroop typescript | Coding harness configs and SDKs | 106 | mostly simple | ✅ |
| AgencyBench evalssandboxpython | Evaluation and benchmarking harnesses | 90 | complex | ✅ |
| letta-evals memorypython | Evaluation and benchmarking harnesses | 80 | mostly simple | ✅ |
| Talon mcpmemoryclitypescript | Personal agent runtimes | 71 | slightly complex | ✅ |
| SUPER sandboxpython | Evaluation and benchmarking harnesses | 57 | slightly complex | ✅ |
| ToolRAG mcptool-discovery | Progressive disclosure harnesses | 29 | mostly simple | ✅ |
| puppeteer-real-browser-mcp mcpbrowsertypescript | Plugins, MCPs, CLI tools | 25 | mostly simple | ❓ |
| TRAIL | Evaluation and benchmarking harnesses | 22 | mostly simple | ✅ |
| Community-curated agent lists | Libraries and SDKs | 14 | super simple | ❓ |
| Better-OpenCodeMCP mcptypescript | Plugins, MCPs, CLI tools | 9 | mostly simple | ✅ |
| pmstack evals | Coding harness configs and SDKs | 8 | super simple | ✅ |
| agentlog memoryclipython | Plugins, MCPs, CLI tools | 1 | super simple | ✅ |
Categories
Progressive disclosure harnesses8 projectsCoding agent products (IDEs, CLIs, full suites)22 projectsCoding harness configs and SDKs17 projectsPersonal agent runtimes10 projectsFrameworks25 projectsMulti-agent and orchestration12 projectsPlugins, MCPs, CLI tools19 projectsMemory and state4 projectsEvaluation and benchmarking harnesses17 projectsObservability and eval-ops2 projectsResearch and task-specific harnesses5 projectsLibraries and SDKs13 projects
Pick by use case
- I want a turnkey coding agent today — opencode, Cline, Codex, Gemini CLI, OpenHands
- I want an always-on personal agent that lives in my chat apps — OpenClaw, Hermes, Khoj, Agent Zero, OpenHarness (HKUDS)
- I want to extend Claude Code, Codex, or OpenCode with skills and slash commands — Anthropic Skills, wshobson/agents, superpowers, GStack, pmstack
- I want to build my own coding harness from scratch — Claude Agent SDK, Google ADK, AutoHarness, SWE-agent, RepoMaster
- I want a drop-in memory layer for agents — Mem0, claude-mem, agentlog, agno, letta
- I want to plug hundreds to thousands of tools without context bloat — MCP-Zero, ToolGen, ToolRAG, langgraph-bigtool
- I want multi-agent orchestration — openai-agents-python, crewAI, autogen, Microsoft Agent Framework, PraisonAI
- I want a general LLM app framework — langgraph, langchain, llama-index, pydantic-ai, agno
- I want low-code / visual workflows — langflow, Flowise, Dify, n8n
- I want browser-using agents — browser-use, WebVoyager, puppeteer-real-browser-mcp
- I want sandboxed code execution for agent-generated code — E2B, Daytona, smolagents, OpenHands
- I want to evaluate or benchmark agents — SWE-bench, AgencyBench, inspect_ai, WebArena, ARC-AGI-2
- I want a deep research / autonomous research agent — deepagents, gpt-researcher, openagents
- I want a provider-agnostic LLM pipe (not a framework) — LiteLLM, vercel/ai
Decision guides
- How to pick a harness — Six questions, in order. Each one eliminates most of the list; by the end you should be choosing between two or three projects, not 103. The
- Agent memory layers: Mem0 vs Letta vs claude-mem — "Add memory to my agent" hides three different products: a **memory API** you call from any agent (Mem0), an **agent runtime** where memory
- Multi-agent orchestration: OpenAI Agents SDK vs CrewAI vs AutoGen vs LangGraph — Four very different answers to "how should multiple agents coordinate?" — and the differences are architectural, not cosmetic. Picking wrong
- OpenClaw vs Hermes: the always-on personal-agent debate — The loudest harness argument of 2026. Both are open-source (MIT), self-hosted, always-on personal agents you talk to from chat apps — and th
- Terminal coding agents: opencode vs Codex vs Gemini CLI vs crush vs goose — The most-asked pick in this list: *"I want a turnkey coding agent in my terminal today."* These five are the open(ish)-source field. The clo
For agents
This list is published machine-readable so coding and research agents can recommend harnesses directly:
- harnesses.json — every project with tier, tags, axes, license, example, and use-case index.
- llms.txt — the whole list in one agent-readable file.
- harnesses.jsonld — schema.org Dataset + ItemList.
- feed.json — JSON Feed of refreshes.
- MCP server:
uvx agent-harnesses-mcp— pick_harness, search_harnesses, get_harness, comparisons.