Skip to content

Arena

Agent Frameworks & SDKs arenaAgent Frameworks & SDKs

Frameworks and SDKs for building LLM agents — tool definitions, multi-agent orchestration, durable state, human-in-the-loop control, and the observability and deployment story around them — judged on how quickly a team ships a reliable agent and how little the framework locks them in.

51 user stories · 459 judged cells · updated 2026-09-16 · Evidence as of 2026-09-16

Buyer checklist →Procurement report →

Leaderboard — every product ranked by evidenceLeaderboard

Rank by
1Claude Agent SDK logoClaude Agent SDK
usage-based
vs OpenAI Agents SDK
67/100
2OpenAI Agents SDK logoOpenAI Agents SDK
usage-based
vs Claude Agent SDK
55/100
3Pydantic AI logoPydantic AI
free-tier
vs Claude Agent SDK
62/100
4Google ADK logoGoogle ADK
vs Claude Agent SDK
46/100
5smolagents logosmolagents
vs Claude Agent SDK
52/100
6CrewAI logoCrewAI
free-tier
vs Claude Agent SDK
47/100
7Mastra logoMastra
free-tier
vs Claude Agent SDK
51/100
8LangGraph logoLangGraph🔥
free-tier
vs Claude Agent SDK
43/100
9AutoGen logoAutoGen
vs Claude Agent SDK
31/100

Best by user type — persona-weighted winnersBest by user type

Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.

Best for developer

Mastra logo

Mastra

64/100

Runner-up: OpenAI Agents SDK logo OpenAI Agents SDK (60/100)

13 developer stories scored

Best for engineering-lead

Mastra logo

Mastra

61/100

Runner-up: Google ADK logo Google ADK (60/100)

6 engineering-lead stories scored

Best for ai-native

Pydantic AI logo

Pydantic AI

36/100

Runner-up: Claude Agent SDK logo Claude Agent SDK (35/100)

32 ai-native stories scored

Story matrix — every product × every judged storyStory matrix

51/51 stories shown · legend

Agenticness — how well agents can access and operate the productAgenticness

Agent access

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docsai-native
fullT
9/10
fullT
8/10
fullT
9/10
fullT
9/10
fullT
8/10
fullT
9/10
disputedD
3/10
none
0/10
partialT
4/10
Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automationai-native
partialC
6/10
fullC
7/10
fullX
8/10
fullT
8/10
fullC
8/10
partialC
6/10
fullC
8/10
partialC
6/10
fullC
7/10
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their toolsai-native
fullC
6/10
fullC
9/10
fullC
8/10
fullC
8/10
fullC
8/10
fullC
8/10
fullC
8/10
partialC
6/10
partialC
6/10
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP serverai-native
none
0/10
n/a
n/a
none
0/10
n/a
fullT
8/10
partialC
6/10
n/a
n/a
Agenticness — how well agents can access and operate the productUse an official CLIai-native
fullT
9/10
none
0/10
fullT
7/10
fullT
8/10
fullT
8/10
partialC
6/10
fullC
8/10
none
0/10
fullC
8/10
Agenticness — how well agents can access and operate the productDrive the product through a documented public APIai-native
partialT
7/10
fullT
8/10
fullX
9/10
partialT
6/10
fullT
8/10
partialT
6/10
partialT
6/10
fullT
7/10
fullT
7/10
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agentai-native
none
0/10
none
0/10
partialC
4/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productBuild against official SDKsai-native
fullT
8/10
fullX
9/10
fullX
8/10
partialT
6/10
fullX
8/10
fullT
8/10
fullC
9/10
fullT
8/10
fullC
8/10
Agenticness — how well agents can access and operate the productSubscribe to events via webhooksai-native
none
0/10
none
0/10
none
0/10
partialC
5/10
none
0/10
none
0/10
none
0/10
none
0/10
n/a

Agentic features

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the productai-native
n/a
n/a
partialC
6/10
partialX
5/10
n/a
n/a
n/a
partialC
4/10
partialX
5/10
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the backgroundai-native
partialX
7/10
none
0/10
partialX
6/10
fullX
7/10
partialX
5/10
fullC
8/10
partialC
6/10
partialX
6/10
none
0/10
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the productai-native
n/a
partialC
6/10
fullC
8/10
fullC
8/10
partialT
6/10
n/a
none
0/10
partialC
6/10
n/a
Agenticness — how well agents can access and operate the productOperate the product with natural-language commandsai-native
none
0/10
fullC
7/10
fullC
8/10
none
0/10
fullC
7/10
partialT
4/10
partialC
5/10
partialC
6/10
fullC
7/10

Api quality

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examplesai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)ai-native
none
0/10
none
0/10
none
0/10
none
0/10
partialT
5/10
none
0/10
none
0/10
none
0/10
n/a
Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production dataai-native
partialC
4/10
fullC
7/10
partialC
3/10
none
0/10
partialX
5/10
partialC
4/10
partialC
5/10
partialX
4/10
partialC
6/10
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policyai-native
none
0/10
none
0/10
none
0/10
partialT
3/10
none
0/10
none
0/10
none
0/10
partialC
3/10
none
0/10

Agents tools — stories about agents tools in this arenaAgents tools

Agent authoring

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Agents tools — stories about agents tools in this arenaDefine an agent with typed custom tools in a few lines of codedeveloper
partialX
4/10
fullX
9/10
fullC
7/10
partialC
6/10
fullX
9/10
fullC
9/10
partialC
6/10
partialC
6/10
fullX
8/10

Ai buildability

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Agents tools — stories about agents tools in this arenaHave a coding agent scaffold a new agent project from an official CLI or template in one commandai-native
partialT
4/10
none
0/10
none
0/10
fullT
8/10
none
0/10
partialC
6/10
fullC
8/10
none
0/10
partialC
4/10
Agents tools — stories about agents tools in this arenaRun the framework's example agents headlessly from a terminal so an agent can verify what it just builtai-native
partialT
4/10
partialC
5/10
partialX
5/10
partialT
5/10
partialC
5/10
none
0/10
fullC
7/10
partialX
4/10
partialC
6/10
Agents tools — stories about agents tools in this arenaRely on strict typing and schema validation so a coding agent catches its own mistakes at build timeai-native
partialC
3/10
partialC
6/10
partialC
4/10
none
0/10
disputedD
5/10
fullC
6/10
none
0/10
none
0/10
none
0/10

Automation depth — how much of the product can run unattendedAutomation depth

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Automation depth — how much of the product can run unattendedPerform bulk operations across many items at onceai-native
none
0/10
none
0/10
partialX
5/10
partialX
5/10
partialC
4/10
partialC
4/10
partialC
3/10
none
0/10
partialX
4/10
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on eventsai-native
partialC
4/10
partialC
4/10
fullC
7/10
partialX
6/10
none
0/10
partialC
6/10
fullC
7/10
partialC
4/10
n/a
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflowsai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
fullC
7/10
none
0/10
none
0/10
none
0/10
Automation depth — how much of the product can run unattendedVersion, review, and roll back my automationsai-native
partialC
5/10
none
0/10
none
0/10
partialX
2/10
none
0/10
none
0/10
none
0/10
none
0/10
partialC
4/10

Deployment portability — stories about deployment portability in this arenaDeployment portability

Deployment

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Deployment portability — stories about deployment portability in this arenaDeploy an agent to a managed runtime and call it as an API endpointengineering-lead
partialX
6/10
none
0/10
partialC
6/10
partialC
6/10
none
0/10
partialT
6/10
fullC
8/10
partialC
4/10
none
0/10
Deployment portability — stories about deployment portability in this arenaRun my agents entirely on my own infrastructure with no dependence on the vendor's platformengineering-lead
fullX
7/10
partialC
6/10
disputedD
4/10
fullX
7/10
fullC
7/10
fullX
8/10
fullC
8/10
fullX
8/10
fullC
8/10

Portability

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Deployment portability — stories about deployment portability in this arenaSwap the underlying LLM provider or model without rewriting my agentdeveloper
none
0/10
fullC
8/10
none
0/10
fullC
7/10
disputedD
6/10
fullC
8/10
fullC
8/10
partialX
6/10
fullC
9/10

Evals observability — stories about evals observability in this arenaEvals observability

Evals

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Evals observability — stories about evals observability in this arenaScore agent quality with built-in evals and run them as part of CIengineering-lead
none
0/10
none
0/10
none
0/10
partialC
4/10
partialC
6/10
partialC
5/10
fullC
8/10
none
0/10
none
0/10

Testing

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Evals observability — stories about evals observability in this arenaUnit-test agents with mocked models and toolsdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
fullX
8/10
none
0/10
partialC
5/10
none
0/10
none
0/10

Tracing

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Evals observability — stories about evals observability in this arenaTrace every LLM call and tool invocation of an agent run in an observability UIdeveloper
fullX
8/10
fullC
9/10
disputedD
4/10
partialC
7/10
fullX
9/10
fullC
8/10
partialC
5/10
partialC
4/10
fullC
8/10

Guardrails safety — stories about guardrails safety in this arenaGuardrails safety

Guardrails

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Guardrails safety — stories about guardrails safety in this arenaAttach input/output guardrails that validate, transform, or block unsafe contentdeveloper
none
0/10
fullC
8/10
partialC
6/10
none
0/10
partialC
3/10
fullC
7/10
partialC
5/10
none
0/10
partialC
5/10
Guardrails safety — stories about guardrails safety in this arenaRestrict what an agent may do with fine-grained tool permissions and sandboxed executionengineering-lead
none
0/10
partialC
7/10
fullC
8/10
partialX
3/10
partialC
6/10
fullC
7/10
partialC
4/10
partialX
5/10
partialX
6/10

Human in the loop — stories about human in the loop in this arenaHuman in the loop

Approval flows

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Human in the loop — stories about human in the loop in this arenaPause an agent mid-run for human input or approval and resume with the human's decisiondeveloper
fullC
9/10
fullC
8/10
fullC
8/10
partialX
5/10
fullC
8/10
fullC
8/10
fullC
8/10
partialC
6/10
partialC
4/10
Human in the loop — stories about human in the loop in this arenaRequire human approval before specific sensitive tool calls executeengineering-lead
fullX
8/10
fullC
8/10
fullC
8/10
disputedD
4/10
fullC
8/10
fullC
8/10
fullC
8/10
partialC
5/10
none
0/10

Memory context — stories about memory context in this arenaMemory context

Memory

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Memory context — stories about memory context in this arenaTrim, summarize, or filter conversation history to keep an agent inside its context windowdeveloper
none
0/10
none
0/10
partialC
4/10
none
0/10
partialC
3/10
none
0/10
fullC
7/10
none
0/10
partialC
4/10
Memory context — stories about memory context in this arenaGive agents long-term memory that persists across sessions and threadsdeveloper
fullX
8/10
partialC
5/10
partialC
6/10
fullC
8/10
partialX
5/10
fullC
8/10
partialC
3/10
partialC
5/10
none
0/10

Openness — open source, data portability, and self-hosting storiesOpenness

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UIai-native
partialT
4/10
n/a
partialX
7/10
partialT
4/10
n/a
partialT
4/10
partialT
5/10
partialC
6/10
n/a
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leaveai-native
partialC
4/10
n/a
partialC
5/10
none
0/10
partialX
5/10
none
0/10
n/a
partialC
4/10
partialC
4/10
Openness — open source, data portability, and self-hosting storiesRead the product's source under an open licenseai-native
partialC
5/10
partialX
5/10
none
0/10
partialC
6/10
none
0/10
disputedD
4/10
fullC
7/10
partialC
6/10
partialC
5/10
Openness — open source, data portability, and self-hosting storiesSelf-host the core productai-native
fullC
8/10
fullC
8/10
none
0/10
fullC
8/10
fullX
7/10
disputedD
5/10
fullC
8/10
fullX
8/10
fullC
8/10

Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestration multi agent

Multi agent

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestrate multiple agents — handoffs, subagents, or crews — inside one workflowdeveloper
fullX
8/10
fullX
9/10
fullC
8/10
fullX
9/10
fullX
7/10
fullX
7/10
fullC
9/10
fullX
9/10
fullC
8/10

Workflow control

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Orchestration multi agent — stories about orchestration multi agent in this arenaCompose agents into an explicit graph or workflow with branching, loops, and parallel stepsdeveloper
fullX
9/10
partialX
5/10
partialC
4/10
partialC
6/10
partialC
6/10
partialX
7/10
fullC
9/10
partialC
6/10
partialC
4/10

Privacy posture — data-handling and privacy storiesPrivacy posture

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency)ai-native
none
0/10
none
0/10
none
0/10
none
0/10
n/a
partialC
4/10
none
0/10
n/a
n/a
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI modelsai-native
none
0/10
none
0/10
none
0/10
n/a
n/a
n/a
n/a
n/a
n/a
Privacy posture — data-handling and privacy storiesControl data retention and deletionai-native
none
0/10
none
0/10
none
0/10
none
0/10
n/a
none
0/10
none
0/10
none
0/10
n/a
Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage trackingai-native
none
0/10
none
0/10
none
0/10
none
0/10
partialC
4/10
none
0/10
none
0/10
none
0/10
none
0/10

State durability — stories about state durability in this arenaState durability

Durable state

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
State durability — stories about state durability in this arenaCheckpoint agent state so a run can resume exactly where it left off after a crash or restartdeveloper
fullC
9/10
partialX
5/10
fullC
7/10
disputedD
4/10
fullC
7/10
fullC
8/10
none
0/10
partialC
3/10
partialC
4/10
State durability — stories about state durability in this arenaRun long-lived agents durably across process restarts and deploys, natively or via durable-execution integrationsengineering-lead
fullC
8/10
partialX
4/10
partialC
7/10
disputedD
3/10
partialX
6/10
fullC
7/10
partialC
3/10
none
0/10
none
0/10

Streaming output — stories about streaming output in this arenaStreaming output

Streaming

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Streaming output — stories about streaming output in this arenaStream tokens and intermediate agent events (tool calls, steps) to my UI in real timedeveloper
partialX
7/10
fullC
9/10
fullC
8/10
partialC
4/10
partialX
6/10
fullC
8/10
partialC
4/10
none
0/10
partialC
4/10

Structured output

StoryPersona
LangGraph logoLangGraph
OpenAI Agents SDK logoOpenAI Agents SDK
Claude Agent SDK logoClaude Agent SDK
CrewAI logoCrewAI
Pydantic AI logoPydantic AI
Mastra logoMastra
Google ADK logoGoogle ADK
AutoGen logoAutoGen
smolagents logosmolagents
Streaming output — stories about streaming output in this arenaGet schema-validated structured output from an agent, with automatic retries when validation failsdeveloper
none
0/10
partialC
5/10
partialC
6/10
none
0/10
disputedD
5/10
partialC
5/10
none
0/10
none
0/10
partialC
3/10
Verdict✓ fullclear evidence~ partialwith caveats! disputedevidence conflicts— noneno evidence foundn/aquestion doesn't apply to this kind of product
ProofT probedtested by usX communityusers back itC claimedvendor claim onlyD contradictedevidence disagrees⚿ auth-gatedprobe hit a live sign-in wall — verified reachable, untestable keylessly
quality 0–10 · PA Score /100 · A–D = evidence confidence · full guide

Adjacent arenas — categories often shopped togetherAdjacent arenas

Shopping this category often means shopping these too.