Best agent orchestration frameworks in 2026: 10 open-source options
Agent frameworks have converged on the same feature list, but they still differ in how they model control flow, where state lives and how much they assume about your model provider. Ten code-first, open-source frameworks, compared against published criteria with GitHub activity checked on 5 October 2026.
TL;DR
- An agent orchestration framework is a library that decides what runs next in an agent system, keeps state between steps, and lets a run pause for a person or recover after a crash. This comparison covers ten open-source, code-first frameworks, not hosted platforms.
- LangGraph and Microsoft Agent Framework rank highest because they combine an explicit graph model, persistent checkpoints, human-in-the-loop pauses, MCP tools and OpenTelemetry or equivalent tracing in a permissive licence.
- OpenAI Agents SDK, Google ADK and Pydantic AI follow. Each is provider-neutral in its documentation, and each reaches durable execution either natively or through engines such as Temporal and DBOS.
- AutoGen and Semantic Kernel are now superseded by Microsoft Agent Framework. The Claude Agent SDK is not ranked because its use is governed by Anthropic’s Commercial Terms of Service.
- No framework replaces the layer underneath it. Provider failover, spend limits and tool permissions are easier to enforce in one gateway than in ten framework configurations.
Most agent frameworks in 2026 advertise the same list: tools, memory, multi-agent patterns, human approval, tracing. The differences that matter in production are less visible. They are the shape of the control flow (a graph, a handoff, an event, a crew of roles), where a paused run’s state is stored, whether a crash means starting again, and how much the framework assumes about which model you call. This guide compares ten open-source agent orchestration frameworks on those points. Capabilities come from each project’s own documentation and repository, and release and GitHub figures were retrieved through the GitHub API on 5 October 2026. None of the frameworks was benchmarked for this article.
What an orchestration framework does
An agent is a loop: a model reads context, proposes an action, a tool runs, and the result goes back to the model. Orchestration is everything around that loop. It covers which agent or step runs next, what state each one sees, what happens when a step fails, and when a person must approve something.
Four terms recur below and are worth defining once.
- Orchestration model. The abstraction a framework uses to express control flow. Graphs use nodes and edges. Handoffs let one agent transfer the conversation to another. Event-driven systems trigger steps when an event of a given type is emitted. Crews assign roles and tasks to a group of agents.
- Durable execution. The ability to persist progress so that a run interrupted by a crash, a deploy or a long human wait resumes from its last completed step instead of from the beginning. Some frameworks implement this with checkpoints in a database; others delegate it to a workflow engine such as Temporal.
- Human-in-the-loop (HITL). A pause in the run that waits for a person to approve, reject or edit an action before execution continues.
- MCP and A2A. The Model Context Protocol connects an agent to tools. The Agent2Agent protocol, maintained by the Linux Foundation with a technical steering committee that includes AWS, Google, IBM Research, Microsoft and others, connects one agent to another. The A2A documentation summarises the split as “MCP is for agent-to-tool communication” and “A2A is for agent-to-agent communication.”
This piece covers libraries a developer imports and runs inside their own process. Hosted agent platforms with their own control planes are the subject of a companion comparison of managed agent orchestration platforms.
Evaluation criteria
The criteria are published before the ranking so a reader can apply different weights. Each was checked against the framework’s documentation or repository.
| Criterion | What was checked | Why it matters |
|---|---|---|
| Orchestration model | Graph, handoff, event-driven, crew or role-based, or a mix | Decides how readable and testable control flow is |
| State and durable execution | Checkpoint stores, resumable runs, integrations with durable engines | A crash or deploy should not discard an hour of agent work |
| Human-in-the-loop | Documented pause, approve or reject, and resume | Tools that write, pay or delete need a person in the path |
| MCP | Consuming MCP tools; exposing agents as MCP servers | MCP is the common tool interface across vendors |
| A2A | Exposing and calling agents over A2A | Lets agents built in different frameworks cooperate |
| Provider neutrality | Documented support for several model providers | Avoids rewriting the agent when the model changes |
| Observability | OpenTelemetry or tracing hooks; default trace destination | Agent failures are diagnosed from traces, not logs |
| Language and licence | Supported languages; licence of the core | Determines who can adopt it and on what terms |
| Maturity | Latest release date, release cadence, GitHub stars on 5 October 2026 | Signals whether the project is maintained |
Stars are a weak signal of quality and a reasonable signal of community size. Release dates are the more useful maturity check: every framework ranked below shipped a release in the 30 days before publication.
Frameworks compared at a glance
The ranking reflects how completely each framework covers the criteria above for a general production team. A lower rank does not mean a worse library; several entries are the right choice for a specific language or use case.
| Rank and framework | Orchestration model | Durable state | MCP / A2A | Languages | Licence | Latest release (date) | GitHub stars |
|---|---|---|---|---|---|---|---|
| 1. LangGraph | Graph (state machine) | Checkpointers: in-memory, SQLite, Postgres | MCP via adapter; A2A via Agent Server | Python, JS | MIT | 1.2.12 (21 Sep) | 42.7k |
| 2. Microsoft Agent Framework | Graph workflows plus sequential, concurrent, handoff, group chat | Checkpointing and hydration | Both, built in | .NET, Python; Go in preview | MIT | python-1.20.0 (2 Oct) | 13.9k |
| 3. OpenAI Agents SDK | Handoffs and agents as tools | Sessions; Temporal, Restate, DBOS, Dapr integrations | MCP (four transports) | Python, TypeScript | MIT | v0.23.1 (2 Oct) | 29.8k (Python) |
| 4. Google ADK | Graph workflow runtime, task delegation, template workflows | Database and Vertex AI session services | Both | Python, TypeScript, Go, Java, Kotlin | Apache 2.0 | v2.11.0 (2 Oct) | 21.7k (Python) |
| 5. Pydantic AI | Typed agent loop; Pydantic Graph | Eight durable engines, including Temporal, DBOS, Prefect, Restate | MCP; A2A via FastA2A | Python | MIT | v2.54.0 (3 Oct) | 20.4k |
| 6. CrewAI | Crews (roles) plus event-driven Flows | @persist on Flows | Both | Python | MIT | 1.15.23 (28 Sep) | 59.4k |
| 7. Mastra | Agents plus graph workflows | Storage-backed suspend and resume | Both | TypeScript | Apache 2.0 core; ee/ source-available | @mastra/core 1.74.0 (5 Oct) | 28.6k |
| 8. Strands Agents | Model-driven loop; graph, swarm, workflow | Session managers, including S3 | Both | Python, TypeScript | Apache 2.0 | python/v1.57.2 (1 Oct) | 8.7k |
| 9. LlamaIndex Workflows | Event-driven steps | Pluggable checkpoint and resume | MCP (tools and workflow-as-server) | Python | MIT | llama-index-workflows 2.25.0 (25 Sep) | 52.4k (llama_index) |
| 10. Agno | Agents, teams (four modes), workflows | Sessions and memory in your database | Both | Python | Apache 2.0 | v3.1.1 (2 Oct) | 42.6k |
Graph-first frameworks
The two highest-ranked frameworks make control flow explicit. A developer draws the graph, and the framework persists state at each step.
1. LangGraph
LangGraph describes itself as a “low-level orchestration framework for building, managing, and deploying long-running, stateful agents.” It is MIT-licensed, has a JavaScript counterpart in LangGraph.js, and can be used without LangChain. The README credits Pregel and Apache Beam as inspiration, which explains the model: a graph of nodes that read and write a shared state object.
Strengths. Persistence is the core feature. The persistence documentation separates checkpointers, which store the state of one thread, from stores, which hold long-term memory across threads. Checkpointers ship for memory, SQLite and Postgres; the docs warn that the in-memory saver loses everything on restart and recommend a persistent checkpointer for production. Human-in-the-loop builds on that: calling interrupt() inside a node saves the graph state and suspends execution “at the exact point where interrupt is called,” and the run resumes when the graph is invoked again with a Command carrying the person’s answer. MCP tools load through an MCP adapter built on FastMCP. Maturity is high: separate packages for the core library, checkpointers and CLI are released on their own schedules.
Limits. A2A is documented as a feature of Agent Server, at /a2a/{assistant_id}, rather than of the open-source library, and the graph state must include a messages key to accept A2A parts. Tracing documentation is centred on LangSmith, LangChain’s commercial product. The low-level API also means more code than a role-based framework for simple cases; the README points newcomers to a higher-level package, Deep Agents, for that reason.
Best fit. Teams that want every branch, retry and approval point visible in code, and a mature checkpoint model they can run on their own Postgres.
2. Microsoft Agent Framework
Microsoft Agent Framework (MAF) is the successor to both Semantic Kernel and AutoGen. Its overview says it “combines AutoGen’s simple abstractions for single- and multi-agent patterns with Semantic Kernel’s enterprise-grade features” and adds workflows “that give developers explicit control over multi-agent execution paths.” Microsoft announced version 1.0 for .NET and Python on 3 April 2026, with “stable APIs, and a commitment to long-term support.” The licence is MIT.
Strengths. MAF covers more criteria out of the box than any other entry. The 1.0 release lists connectors for Microsoft Foundry, Azure OpenAI, OpenAI, Anthropic Claude, Amazon Bedrock, Google Gemini and Ollama; a graph-based workflow engine; sequential, concurrent, handoff, group chat and Magentic-One orchestration patterns; checkpointing and hydration so long-running processes “survive interruptions”; human-in-the-loop approvals with pause and resume; and both MCP and A2A. The README lists built-in OpenTelemetry integration and YAML-defined declarative agents. It is the only ranked framework with first-class .NET support.
Limits. Several pieces are not yet stable. The 1.0 announcement marks DevUI, Foundry hosted agents, the Agent Harness and Skills packages as preview, and the Go SDK is in public preview without declarative agents or functional workflows. The documentation’s quick starts lean on Azure credentials, and its overview carries a notice that using third-party models and servers is “at your own risk,” with data-flow responsibilities on the developer.
Best fit. .NET shops, teams migrating from Semantic Kernel or AutoGen, and anyone who wants graph workflows, MCP and A2A in one supported package.
Provider SDKs with multi-agent support
The next two frameworks come from model vendors but are documented as provider-neutral. Their orchestration style starts from agents rather than graphs.
3. OpenAI Agents SDK
The OpenAI Agents SDK is a “lightweight yet powerful framework for building multi-agent workflows,” MIT-licensed, in Python with a TypeScript version. It describes itself as provider-agnostic, supporting OpenAI’s Responses and Chat Completions APIs “as well as 100+ other LLMs.”
Strengths. The orchestration model is small and easy to reason about. Handoffs “are represented as tools to the LLM,” so a handoff to a refund agent appears to the model as a tool named transfer_to_refund_agent; agents can also be called as tools. Human approval is per tool through needs_approval, and pending approvals surface as interruptions. The run converts to a RunState that can be serialised with to_json(), stored in a database or queue, and resumed later, which makes multi-day approvals practical. For durability, the running agents guide documents integrations with Temporal, Restate, DBOS and Dapr, and session backends include SQLAlchemy, SQLite, Redis and MongoDB. MCP support spans hosted MCP tools, Streamable HTTP, SSE and stdio, each with a require_approval policy.
Limits. Defaults point at OpenAI. Tracing uploads traces to OpenAI’s servers unless disabled or redirected to a custom processor, and it is “unavailable for organizations that use OpenAI’s APIs under a Zero Data Retention (ZDR) policy.” Non-OpenAI models go through OpenAI-compatible endpoints or the LiteLLM and any-llm adapters, which the models page labels beta. The version number is still below 1.0, and A2A was not among the features documented on the pages reviewed.
Best fit. Teams that want a thin layer over the agent loop, with handoffs as the main pattern and a choice of durable engine underneath.
4. Google Agent Development Kit
Google ADK is “an open-source, code-first Python framework for building, evaluating, and deploying” agents, under Apache 2.0, with versions in TypeScript, Go, Java and Kotlin. The README says that although it is “optimized for Gemini, ADK is model-agnostic, deployment-agnostic.”
Strengths. ADK 2.0 added a graph-based workflow runtime with routing, fan-out and fan-in, loops, retries, state management, dynamic nodes, human-in-the-loop and nested workflows, alongside a Task API for structured agent-to-agent delegation. The workflows guide also keeps the older template workflows (sequential, loop and parallel agents). A2A is unusually well covered: the A2A docs include quick starts for exposing agents in Python, Go and Java and for consuming them in Python, Go, Java and Kotlin. Non-Gemini models connect through Anthropic, LiteLLM, Ollama and vLLM connectors. Session services include a DatabaseSessionService for PostgreSQL, MySQL or SQLite and a Vertex AI service, and tool calls can be gated by a confirmation flow.
Limits. The deepest integrations, such as Vertex AI Agent Engine and Vertex session storage, assume Google Cloud. The graph runtime is newer than LangGraph’s and arrived with a major version change. MCP toolsets have a deployment constraint: agents deployed to production must define McpToolset synchronously.
Best fit. Polyglot organisations, teams that need A2A across languages, and Google Cloud users who still want the option to run other models.
Typed and role-based Python frameworks
These two Python frameworks start from opposite premises. Pydantic AI treats an agent as a typed function; CrewAI treats it as a team member with a role.
5. Pydantic AI
Pydantic AI, from the team behind the Pydantic validation library, describes itself as “a typed, extensible agent loop with every model a string swap away.” It is MIT-licensed. Outputs are declared as Pydantic models, tool arguments are validated before the tool runs, and the run returns a validated object.
Strengths. Durable execution is the broadest in this comparison. The durable execution docs list eight supported engines plus a builder for others; Temporal, DBOS, Prefect and Restate are among those co-maintained with the engine vendors. Each engine “wraps every model request and tool call as its own durable unit.” The framework is OpenTelemetry-native, so traces go to any OTel backend rather than only to Pydantic’s own Logfire service. Pydantic Graph adds typed graph workflows when a plain loop is not enough, and the optional Harness package bundles capabilities such as memory, guardrails and sub-agents.
Limits. It is Python-only. A2A support moved out of the core: the FastA2A server now lives in a separate repository maintained by Datalayer, and Pydantic AI 2 removed Agent.to_a2a() in favour of a bridge in that package. Only one durable engine can be attached per agent. Multi-agent coordination is expressed through sub-agents and tool calls rather than a dedicated crew or handoff abstraction.
Best fit. Python teams that value type safety and want to choose their own durable execution engine.
6. CrewAI
CrewAI is the most-starred framework in this list, at about 59,400 stars. It is MIT-licensed and, per its FAQ, “a standalone Python framework” with its own primitives for agents, tasks, crews and flows. A commercial control plane, CrewAI AMP, is sold separately and not assessed here.
Strengths. CrewAI offers two models in one library. Crews assign roles, goals and tasks to agents and suit open-ended collaboration. Flows are event-driven: methods marked @start() begin a flow, @listen() runs when another method completes, and @router() branches on output. The @persist decorator “enables automatic state persistence,” a run can be resumed by passing its state ID, and @human_feedback pauses a flow for a person. MCP tools attach through an mcps field or an adapter over stdio, SSE and Streamable HTTP, and A2A works in both directions through client and server configurations on an agent.
Limits. Two models mean two mental models; teams need to decide where crews end and flows begin. Agents default to the OpenAI API unless configured otherwise. The README states that anonymous usage telemetry is collected by default and documents what is and is not sent. Persistence is a Flow feature, not a crew feature.
Best fit. Teams whose problem decomposes naturally into roles, and who want event-driven flows around those crews for the deterministic parts.
TypeScript and multi-language frameworks
Most agent tooling began in Python. These two frameworks treat TypeScript as a first-class target.
7. Mastra
Mastra is “a framework for building AI-powered applications and agents with a modern TypeScript stack.” Its licence is split: the core and most of the codebase are Apache 2.0, while any directory named ee/ is source-available under the Mastra Enterprise License and needs a licence for production use.
Strengths. Mastra pairs autonomous agents with a graph-based workflow engine using .then(), .branch() and .parallel(). Agents and workflows can suspend for human input and resume later, with state held in storage “so you can pause indefinitely.” The README describes model routing to more than 40 providers through one interface. MCP is covered in both directions: MCPClient consumes servers, and MCPServer exposes Mastra agents, tools and workflows to other clients. A2A support covers protocol version 0.3.0 and v1.0 on the same endpoints, for exposing agents and for wrapping remote agents as sub-agents. Built-in evals and observability round out the feature list.
Limits. Readers should check which features sit under ee/ before depending on them; the licence file names enterprise authentication and agent-builder code among them. Managed durable execution relies on external workflow runners such as Inngest. It is TypeScript-only.
Best fit. TypeScript and Next.js teams that want agents, workflows and MCP in the same codebase as the application.
8. Strands Agents
Strands Agents, from AWS, is “an open-source SDK for building and running AI agents in Python and TypeScript” under Apache 2.0. Its GitHub repository has moved to a monorepo, strands-agents/harness-sdk, which also contains a pre-assembled “harness” agent.
Strengths. Strands is model-driven: the agent loop lets the model plan, and the README says the loop “traces every decision by default” while hooks can intercept any step. For explicit orchestration, the multi-agent patterns page offers agents as tools, A2A across processes, a Swarm of agents that hand off autonomously, a Graph with conditional edges and loops, and a fixed-DAG Workflow. Interrupts pause single agents and multi-agent graphs for human approval. Session managers persist conversations to files or S3, and tracing uses OpenTelemetry. Bedrock, Anthropic, OpenAI and Gemini are first-class providers.
Limits. Amazon Bedrock is the default model provider, so non-AWS teams must configure a provider explicitly. The project has fewer stars than its peers, at about 8,700. Some session and interrupt features are still marked in development for the bidirectional streaming agent.
Best fit. AWS teams and anyone who prefers to start from a model-driven loop and add graph or swarm structure only where needed.
Event-driven and platform-style frameworks
The last two entries sit at opposite ends: a minimal event-driven library and a framework that ships its own runtime.
9. LlamaIndex Workflows
LlamaIndex’s Agent Workflows, packaged as llama-index-workflows under MIT, is “an event-driven orchestration library where steps are async Python functions that emit and consume events.” The code now lives in the run-llama/llama-agents repository, which has far fewer stars than the main LlamaIndex repository.
Strengths. The event model replaces explicit edges with ordinary Python: a step receives an event, does work and returns another event, and branching and loops use plain if statements. Durability is pluggable, from saving and resuming a run in a file to connecting a database or coordination backend, and a workflow server adds streaming, persistence and human-in-the-loop over REST. Multi-agent patterns include AgentWorkflow handoffs, an orchestrator agent that calls specialists as tools, and a custom planner. MCP tools load through llama-index-tools-mcp, and workflow_as_mcp turns a workflow into an MCP server.
Limits. The project positions itself around document-centric agents, so its examples and deployment path lean toward LlamaParse. A2A was not documented on the pages reviewed. Observability beyond event streaming depends on integrations.
Best fit. Document-processing pipelines with OCR, extraction and human review, and Python developers who prefer events to graph DSLs.
10. Agno
Agno is “a framework and runtime for agent platforms,” under Apache 2.0, with about 42,600 stars. It pairs an SDK with AgentOS, a runtime that serves agents as an API.
Strengths. Agno’s teams have four documented modes: coordinate (decompose and delegate), route (send to one specialist), broadcast (send to all and synthesise) and tasks (loop over a task list). Agents connect to MCP servers through MCPTools over stdio or Streamable HTTP. AgentOS exposes agents over A2A and AG-UI as well as chat channels, stores sessions, memory and traces in the team’s own database, and adds OpenTelemetry tracing, JWT-based role-based access control and human approval that can persist pending approvals for an administrator.
Limits. Much of the operational value sits in AgentOS rather than in the embeddable library, which makes Agno closer to a platform than the other entries. A2A requires no authentication by default unless authorisation is enabled. The README states that a telemetry event is sent per agent run unless AGNO_TELEMETRY=false is set.
Best fit. Teams that want a self-hosted runtime and API around their agents without adopting a hosted platform.
Frameworks not ranked
Several well-known projects were considered and left out, each for a stated reason.
- AutoGen. The AutoGen README says the project “is now in maintenance mode” and “will not receive new features,” and directs new users to Microsoft Agent Framework. Its last release on GitHub was in September 2025.
- Semantic Kernel. Still actively released, but its README calls Microsoft Agent Framework “the enterprise-ready successor to Semantic Kernel.” New projects are better served by MAF.
- Claude Agent SDK. The Python repository is MIT-licensed, but the SDK documentation states that its use is governed by Anthropic’s Commercial Terms of Service, and the TypeScript package carries an all-rights-reserved notice. It also runs the Claude Code binary and is built around Claude models. It offers subagents with isolated context, hooks, permissions, MCP and resumable sessions, and is a strong option for teams committed to Claude, but it does not meet this list’s open-source bar.
- smolagents. Hugging Face’s Apache 2.0 library for agents that write their actions as code remains a useful minimal option, but its last tagged release was 1.26.0 in May 2026, the least recent of any candidate reviewed apart from AutoGen.
- Temporal. A durable execution engine rather than an agent framework. Its Python SDK ships contrib integrations for OpenAI Agents, Google ADK, LangGraph and Strands, which is how several ranked frameworks reach durable execution.
What sits under the framework in production
A framework decides what runs next. It does not, by itself, govern the traffic that results. Every ranked framework eventually makes two kinds of outbound call: a request to a model provider, and a call to a tool, increasingly over MCP. In a company with several agent teams, those calls come from several frameworks, each with its own retry settings, keys and tracing defaults.
One gateway for model calls
Provider neutrality in a framework means it can format a request for several APIs. It does not mean the request survives a provider outage, stays within a team’s budget, or uses a key that can be revoked centrally. Those controls sit more naturally in a gateway that every agent calls through an OpenAI- or Anthropic-compatible endpoint. The trade-offs of retries, fallbacks and weighted routing are covered in the Frontier Wire explainer on LLM failover and load balancing, and the cost side in the piece on cost control with an AI gateway.
Bifrost, an open-source Go gateway under Apache 2.0 from Maxim AI, is one example of this layer, and its documentation includes drop-in setups for LangChain and Pydantic AI clients alongside OpenAI, Anthropic, Bedrock and Google GenAI SDKs. The Bifrost LLM gateway page describes automatic fallback to the next provider when the primary fails, virtual keys that scope providers and models per consumer, and budgets and rate limits at key, team and customer level. Per the fallback documentation, each fallback provider gets its own retry budget and is treated as a fresh request, so caching, governance and logging run again. Bifrost’s own benchmark reports 11 µs of added overhead per request at 5,000 requests per second on a t3.xlarge instance against a mocked upstream; that is a vendor figure, not an independent test. Clustering, adaptive load balancing, guardrails and SSO are in its enterprise tier.
One gateway for tools
Tool access needs the same treatment. Most frameworks above let a developer attach any MCP server to any agent, which is convenient in development and hard to audit later. An MCP gateway puts the tool catalogue behind one endpoint and decides which caller may use which tool. The Bifrost MCP gateway applies the same virtual keys to tools, lets administrators bundle permitted tools into a curated virtual server, and by default returns a model’s tool calls as suggestions that require explicit approval; an agent mode auto-executes only an allow-listed subset. That default lines up with the human-in-the-loop controls the frameworks provide, rather than bypassing them. A broader field of options is compared in the best MCP gateways.
Traces that cross frameworks
Framework tracing tells you what one agent did. A gateway sees every model and tool call from every framework at a single point, which is useful when an incident spans two teams’ agents. Gateway-level observability in Bifrost records tokens, cost and latency per request, emits OpenTelemetry traces with GenAI semantic conventions, and exposes Prometheus metrics. Framework traces and gateway traces complement each other; neither replaces evaluation of whether the agent’s answers were right, which is covered in the guide on how to compare LLM evaluation frameworks.
Choosing by orchestration model
The fastest way to narrow the list is to start from the shape of the problem and the language the team already uses.
| Starting point | First framework to evaluate | Reason |
|---|---|---|
| Explicit, auditable control flow in Python or JS | LangGraph | Mature checkpointers and interrupt-based approval |
| .NET, or migrating from Semantic Kernel or AutoGen | Microsoft Agent Framework | Stable 1.0 APIs, graph workflows, MCP and A2A |
| Thin agent loop with handoffs | OpenAI Agents SDK | Small API; serialisable approval state; Temporal, Restate, DBOS integrations |
| Agents in several languages that must interoperate | Google ADK | A2A quick starts in Python, Go, Java and Kotlin |
| Typed outputs and a durable engine of your choice | Pydantic AI | Eight durable engines; OpenTelemetry-native |
| Work that maps onto roles and tasks | CrewAI | Crews for collaboration, Flows for deterministic steps |
| TypeScript application code | Mastra | Agents, workflows, MCP and A2A in one TypeScript stack |
| AWS-first, model-driven agents | Strands Agents | Bedrock default; swarm and graph patterns |
| Document pipelines with human review | LlamaIndex Workflows | Event-driven steps; workflow-as-MCP-server |
| Self-hosted agent runtime with an API | Agno | AgentOS with storage, RBAC, A2A and AG-UI |
The orchestration model matters more than the feature list. Graphs suit processes whose steps are known in advance and must be audited. Handoffs suit conversational routing between specialists. Event-driven designs suit pipelines where steps fan out and rejoin. Crews suit open-ended work where the decomposition itself is part of the task. Most production systems mix two of these, which is why several frameworks now ship both an agent abstraction and a workflow engine.
Limits of this comparison
This is a documentation and repository review, not a benchmark. No framework was installed, load-tested or run against an agent task for this article, and no vendor performance claim was reproduced. Feature descriptions reflect documentation read on 5 October 2026; agent frameworks release weekly, and several features cited here are labelled preview or beta by their own maintainers.
GitHub stars and release dates were retrieved from the GitHub API on the same day. Stars measure attention, not quality, and repositories that split code across several packages (LlamaIndex Workflows, the OpenAI Agents SDK, Google ADK) are not directly comparable to single-repository projects.
Some areas were out of scope. Managed agent platforms, visual and low-code builders, and commercial add-ons such as LangSmith, CrewAI AMP and AgentOS’s hosted control plane were not assessed. Security properties, such as sandboxing of code-executing agents and resistance to prompt injection through tool results, were not tested. Pricing of the models each framework calls was not compared. A2A support was recorded only where documentation described it; a framework marked without A2A may still interoperate through a third-party adapter.
Sources
- LangGraph repository (GitHub, MIT)
- LangGraph docs: persistence and durable execution
- LangGraph docs: interrupts
- LangChain docs: MCP adapter
- LangSmith docs: A2A endpoint in Agent Server
- LangGraph docs: observability
- Microsoft Agent Framework repository (GitHub, MIT)
- Microsoft Agent Framework overview (Microsoft Learn)
- Microsoft Agent Framework version 1.0 announcement (3 April 2026)
- AutoGen repository and maintenance-mode notice
- Semantic Kernel repository and successor notice
- OpenAI Agents SDK for Python repository (GitHub, MIT)
- OpenAI Agents SDK: handoffs
- OpenAI Agents SDK: human in the loop
- OpenAI Agents SDK: running agents, sessions and durable execution integrations
- OpenAI Agents SDK: MCP
- OpenAI Agents SDK: models and non-OpenAI providers
- OpenAI Agents SDK: tracing
- Google Agent Development Kit for Python repository (GitHub, Apache 2.0)
- ADK docs: workflows
- ADK docs: A2A
- ADK docs: models
- ADK docs: session services
- Pydantic AI repository (GitHub, MIT)
- Pydantic AI docs: durable execution
- FastA2A repository (A2A server for Pydantic AI)
- CrewAI repository (GitHub, MIT)
- CrewAI docs: Flows
- CrewAI docs: A2A agent delegation
- Mastra repository and licence mapping (GitHub, Apache 2.0 with ee/ directories)
- Mastra docs: A2A
- Strands Agents monorepo (GitHub, Apache 2.0)
- Strands docs: multi-agent patterns
- LlamaAgents and Agent Workflows repository (GitHub, MIT)
- Agno repository (GitHub, Apache 2.0)
- Claude Agent SDK overview, including licence and terms
- Temporal Python SDK: contrib integrations for agent frameworks
- A2A protocol documentation (Linux Foundation)
- Bifrost docs: fallbacks
- Bifrost docs: benchmarking
Questions readers ask
What is an agent orchestration framework?
It is a library you embed in your own application to coordinate model calls, tool calls and, often, several agents. It decides what runs next (a graph edge, a handoff, an event or a role assignment), keeps state between steps, and lets the run pause for a human or resume after a failure. Managed platforms that host agents for you are a separate category.
Which agent framework is best for production in 2026?
For explicit control flow with checkpoints and human approval, LangGraph and Microsoft Agent Framework are the most complete on the criteria used here. Teams on .NET should start with Microsoft Agent Framework, TypeScript teams with Mastra, and Python teams that want typed outputs and pluggable durable execution with Pydantic AI.
What is the difference between MCP and A2A in agent frameworks?
The Model Context Protocol (MCP) connects an agent to tools and data. The Agent2Agent protocol (A2A), now governed by the Linux Foundation, connects an agent to another agent that keeps its own tools, memory and prompts private. Most frameworks reviewed here consume MCP tools; fewer can both expose and call agents over A2A.
Is AutoGen still maintained?
AutoGen is in maintenance mode. Its README says it will not receive new features and points new users to Microsoft Agent Framework, which reached version 1.0 for .NET and Python on 3 April 2026. Semantic Kernel's README also names Microsoft Agent Framework as its successor.
Do I need an AI gateway if my agent framework already supports many model providers?
Provider support in a framework means it can format requests for several APIs. It does not give you failover across providers, organisation-wide budgets, a single audit trail for tool calls, or one place to rotate keys. Those controls sit in a gateway that every framework instance calls, so they stay consistent when teams use different frameworks.
