Rank #2 of 9 in Agent Frameworks & SDKs
Access
Showcase


Products
OpenAI, product by product →OpenAI ships more than one product — each judged line competes in its own arena on the same stories as everyone else.
| Line | Arena | Rank | PA Score | Agent-ready |
|---|---|---|---|---|
| ChatGPT | AI Assistants | #1/9 | 38/100 | 49/100 |
| Codex | AI Coding Agents | #6/13 | 33/100 | 46/100 |
| Agents SDKthis page | Agent Frameworks & SDKs | #2/9 | 37/100 | 55/100 |
| Plugins | Agent Skills & Extensions | #4/5 | 19/100 | 33/100 |
Not yet judged (6 — no arena where they compete): ChatGPT Work · Image generation (GPT-Image-2.5) · Codex Security · ChatKit · Agents API · API Platform
Verified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Agents tools — stories about agents tools in this arenaAgents toolsevidence →
Stories about agents tools in this arena
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment portability — stories about deployment portability in this arenaDeployment portabilityevidence →
Stories about deployment portability in this arena
Evals observability — stories about evals observability in this arenaEvals observabilityevidence →
Stories about evals observability in this arena
Guardrails safety — stories about guardrails safety in this arenaGuardrails safetyevidence →
Stories about guardrails safety in this arena
Human in the loop — stories about human in the loop in this arenaHuman in the loopevidence →
Stories about human in the loop in this arena
Memory context — stories about memory context in this arenaMemory contextevidence →
Stories about memory context in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestration multi agentevidence →
Stories about orchestration multi agent in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
State durability — stories about state durability in this arenaState durabilityevidence →
Stories about state durability in this arena
Streaming output — stories about streaming output in this arenaStreaming outputevidence →
Stories about streaming output in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 2 free · 0 paid · 0 enterprise · 26 not stated in evidence
Follow the green: where the map greys out is where OpenAI Agents SDK stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Webhooks · Scoped API keys · Machine-readable spec · Versioning policy · Official CLI · Have a coding agent scaffold a new agent project from an official CLI or template in one command
Subscribe to events via webhooks
—–
Build against official SDKs
✓9/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
n/an/a
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
✓7/10
Explore an interactive API reference with runnable examples
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
~6/10
Operate the product with natural-language commands
✓7/10
unlocks → Autonomous automations
Plug MCP servers into this product so it can use their tools
✓9/10
Get AI-generated insights and suggestions from my data inside the product
n/an/a
Set up automations that run autonomously in the background
—0/10
Agents tools — stories about agents tools in this arenaAgents tools
Stories about agents tools in this arena
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment portability — stories about deployment portability in this arenaDeployment portability
Stories about deployment portability in this arena
Evals observability — stories about evals observability in this arenaEvals observability
Stories about evals observability in this arena
Guardrails safety — stories about guardrails safety in this arenaGuardrails safety
Stories about guardrails safety in this arena
Human in the loop — stories about human in the loop in this arenaHuman in the loop
Stories about human in the loop in this arena
Memory context — stories about memory context in this arenaMemory context
Stories about memory context in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestration multi agent
Stories about orchestration multi agent in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
State durability — stories about state durability in this arenaState durability
Stories about state durability in this arena
Streaming output — stories about streaming output in this arenaStreaming output
Stories about streaming output in this arena
Sorted by importance (agentic first) (high → low) · 51/51 stories · click a row’s chevron for the rationale and evidence
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 6/10 | Cclaimed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Xcommunity | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | full | 7/10 | Cclaimed | |
Define an agent with typed custom tools in a few lines of code C Agent authoring | developer | Agents tools — stories about agents tools in this arenaAgents tools | 3 | full | 9/10 | Xcommunity | |
Orchestrate multiple agents — handoffs, subagents, or crews — inside one workflow C Multi agent | developer | Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestration multi agent | 3 | full | 9/10 | Xcommunity | |
Stream tokens and intermediate agent events (tool calls, steps) to my UI in real time C Streaming | developer | Streaming output — stories about streaming output in this arenaStreaming output | 3 | full | 9/10 | Cclaimed | |
Trace every LLM call and tool invocation of an agent run in an observability UI C Tracing | developer | Evals observability — stories about evals observability in this arenaEvals observability | 3 | full | 9/10 | Cclaimed | |
Attach input/output guardrails that validate, transform, or block unsafe content C Guardrails | developer | Guardrails safety — stories about guardrails safety in this arenaGuardrails safety | 3 | full | 8/10 | Cclaimed | |
Pause an agent mid-run for human input or approval and resume with the human's decision C Approval flows | developer | Human in the loop — stories about human in the loop in this arenaHuman in the loop | 3 | full | 8/10 | Cclaimed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullfree | 8/10 | Cclaimed | |
Swap the underlying LLM provider or model without rewriting my agent C Portability | developer | Deployment portability — stories about deployment portability in this arenaDeployment portability | 3 | full | 8/10 | Cclaimed | |
Checkpoint agent state so a run can resume exactly where it left off after a crash or restart C Durable state | developer | State durability — stories about state durability in this arenaState durability | 3 | partial | 5/10 | Xcommunity | |
Get schema-validated structured output from an agent, with automatic retries when validation fails C Structured output | developer | Streaming output — stories about streaming output in this arenaStreaming output | 3 | partial | 5/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | n/a | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Require human approval before specific sensitive tool calls execute C Approval flows | engineering-lead | Human in the loop — stories about human in the loop in this arenaHuman in the loop | 2 | full | 8/10 | Cclaimed | |
Restrict what an agent may do with fine-grained tool permissions and sandboxed execution C Guardrails | engineering-lead | Guardrails safety — stories about guardrails safety in this arenaGuardrails safety | 2 | partial | 7/10 | Cclaimed | |
Rely on strict typing and schema validation so a coding agent catches its own mistakes at build time C Ai buildability | ai-native user | Agents tools — stories about agents tools in this arenaAgents tools | 2 | partial | 6/10 | Cclaimed | |
Run my agents entirely on my own infrastructure with no dependence on the vendor's platform C Deployment | engineering-lead | Deployment portability — stories about deployment portability in this arenaDeployment portability | 2 | partialfree | 6/10 | Cclaimed | |
Compose agents into an explicit graph or workflow with branching, loops, and parallel steps C Workflow control | developer | Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestration multi agent | 2 | partial | 5/10 | Xcommunity | |
Give agents long-term memory that persists across sessions and threads C Memory | developer | Memory context — stories about memory context in this arenaMemory context | 2 | partial | 5/10 | Cclaimed | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Xcommunity | |
Run the framework's example agents headlessly from a terminal so an agent can verify what it just built C Ai buildability | ai-native user | Agents tools — stories about agents tools in this arenaAgents tools | 2 | partial | 5/10 | Cclaimed | |
Run long-lived agents durably across process restarts and deploys, natively or via durable-execution integrations C Durable state | engineering-lead | State durability — stories about state durability in this arenaState durability | 2 | partial | 4/10 | Xcommunity | |
Deploy an agent to a managed runtime and call it as an API endpoint C Deployment | engineering-lead | Deployment portability — stories about deployment portability in this arenaDeployment portability | 2 | none | 0/10 | ||
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Trim, summarize, or filter conversation history to keep an agent inside its context window C Memory | developer | Memory context — stories about memory context in this arenaMemory context | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | n/a | untested | none yet | |
Have a coding agent scaffold a new agent project from an official CLI or template in one command C Ai buildability | ai-native user | Agents tools — stories about agents tools in this arenaAgents tools | 2 | none | untested | none yet | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
Score agent quality with built-in evals and run them as part of CI C Evals | engineering-lead | Evals observability — stories about evals observability in this arenaEvals observability | 2 | none | untested | none yet | |
Unit-test agents with mocked models and tools C Testing | developer | Evals observability — stories about evals observability in this arenaEvals observability | 2 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 31 stories with headroom
What would move OpenAI Agents SDK’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
The evidence pack covers SDK features (tools, handoffs, tracing, sessions, MCP, voice, sandboxing) but contains no mention of data-training opt-out, privacy controls, or API data usage policies relevant to model training.
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background
nonemoves Built-in AIimpact 30
Missing: scheduling/cron or event-trigger support, a deployment/hosting mechanism for unattended background execution, and any documentation of persistent autonomous operation outside a developer-invoked run.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
The evidence pack is entirely about the Agents SDK library (Python/JS) — agent primitives, tools, tracing, sessions, MCP, streaming — with no mention of an official CLI tool for scaffolding, running, or managing agents; OpenAPI/CLI probes returned 404s.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
Missing: any mention of scoped/least-privilege credential issuance, API key scoping, or permission boundaries for agents.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
The evidence shows in-process streaming/tracing for observing agent run events, but nothing about registering webhook URLs or subscribing to events via HTTP callbacks.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
The evidence pack shows only static documentation pages describing SDK concepts (agents, tools, handoffs, tracing) with no interactive API reference, no runnable code sandbox, and the OpenAPI/swagger probe explicitly returned 404 on all candidate paths, indicating no machine-readable interactive reference exists.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
The probe explicitly checked common OpenAPI spec locations and all returned 404, and no other evidence shows a machine-readable API spec being published for the Agents SDK.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
The evidence pack documents SDK features (agents, tools, tracing, handoffs) but contains no mention of API versioning scheme or a documented deprecation policy; the OpenAPI probe even returned 404s, and no versioning/deprecation docs are cited.
Showing the top 8 of 31 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 28 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Openai agents python docs25 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Plug MCP servers into this product so it can use their tools
- Drive the product through a documented public API
- Build against official SDKs
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Define an agent with typed custom tools in a few lines of code
- Run the framework's example agents headlessly from a terminal so an agent can verify what it just built
- Rely on strict typing and schema validation so a coding agent catches its own mistakes at build time
- Define rules that trigger actions automatically on events
- Run my agents entirely on my own infrastructure with no dependence on the vendor's platform
- Swap the underlying LLM provider or model without rewriting my agent
- Trace every LLM call and tool invocation of an agent run in an observability UI
- Attach input/output guardrails that validate, transform, or block unsafe content
- Restrict what an agent may do with fine-grained tool permissions and sandboxed execution
- Give agents long-term memory that persists across sessions and threads
- Self-host the core product
- Orchestrate multiple agents — handoffs, subagents, or crews — inside one workflow
- Compose agents into an explicit graph or workflow with branching, loops, and parallel steps
- Checkpoint agent state so a run can resume exactly where it left off after a crash or restart
- Run long-lived agents durably across process restarts and deploys, natively or via durable-execution integrations
- Stream tokens and intermediate agent events (tool calls, steps) to my UI in real time
- Get schema-validated structured output from an agent, with automatic retries when validation fails
Hacker News8 stories
- Drive the product through a documented public API
- Build against official SDKs
- Define an agent with typed custom tools in a few lines of code
- Read the product's source under an open license
- Orchestrate multiple agents — handoffs, subagents, or crews — inside one workflow
- Compose agents into an explicit graph or workflow with branching, loops, and parallel steps
- Checkpoint agent state so a run can resume exactly where it left off after a crash or restart
- Run long-lived agents durably across process restarts and deploys, natively or via durable-execution integrations
GitHub README6 stories
- Drive the product through a documented public API
- Build against official SDKs
- Run my agents entirely on my own infrastructure with no dependence on the vendor's platform
- Swap the underlying LLM provider or model without rewriting my agent
- Read the product's source under an open license
- Self-host the core product
Openai agents js docs6 stories
- Define rules that trigger actions automatically on events
- Restrict what an agent may do with fine-grained tool permissions and sandboxed execution
- Pause an agent mid-run for human input or approval and resume with the human's decision
- Require human approval before specific sensitive tool calls execute
- Checkpoint agent state so a run can resume exactly where it left off after a crash or restart
- Run long-lived agents durably across process restarts and deploys, natively or via durable-execution integrations
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
4 of 15 testable claims verified · 0 contradicted → integrity 27/100
23 distinct capability claims found in OpenAI Agents SDK’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
4
Verified
11
Unverified
0
Contradicted
13
Undersold
Verified (6)
“An agent is an LLM configured with instructions, tools, and optional handoffs/guardrails/structured outputs”
Define an agent with typed custom tools in a few lines of codefullproof ↗
“Any Python function can be wrapped as a tool (FunctionTool)”
Define an agent with typed custom tools in a few lines of codefullproof ↗
“Agents can hand off tasks to other specialized agents”
Orchestrate multiple agents — handoffs, subagents, or crews — inside one workflowfullproof ↗
“Agents can be run via Runner.run(), Runner.run_sync(), or Runner.run_streamed()”
Drive the product through a documented public APIfullproof ↗
“Sandbox agents run inside isolated workspaces with manifest-defined files, sandbox client selection, and resumable sessions”
Checkpoint agent state so a run can resume exactly where it left off after a crash or restartpartialproof ↗
“An agent can be exposed as a callable tool for another agent without a full handoff”
Orchestrate multiple agents — handoffs, subagents, or crews — inside one workflowfullproof ↗
Unverified (14)
“Guardrails let you validate/check user input and agent output”
Attach input/output guardrails that validate, transform, or block unsafe contentfullproof ↗
“Streaming lets you subscribe to live updates of an agent run, including partial responses”
Stream tokens and intermediate agent events (tool calls, steps) to my UI in real timefullproof ↗
“Built-in session memory automatically maintains conversation history across multiple agent runs”
Give agents long-term memory that persists across sessions and threadspartialproof ↗
“Built-in tracing captures a full record of LLM generations, tool calls, handoffs, guardrails, and custom events during a run”
Trace every LLM call and tool invocation of an agent run in an observability UIfullproof ↗
“SDK supports multiple MCP transports, letting agents use filesystem, HTTP, or connector-backed MCP tool servers”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“When a tool call needs approval, the SDK pauses the run and lets you resume later from the same RunState”
Pause an agent mid-run for human input or approval and resume with the human's decisionfullproof ↗
“When a tool call needs approval, the SDK pauses the run and lets you resume later from the same RunState”
Require human approval before specific sensitive tool calls executefullproof ↗
“Sandbox agents run inside isolated workspaces with manifest-defined files, sandbox client selection, and resumable sessions”
Restrict what an agent may do with fine-grained tool permissions and sandboxed executionpartialproof ↗
“Traces dashboard lets you debug, visualize, and monitor workflows in dev and production”
Trace every LLM call and tool invocation of an agent run in an observability UIfullproof ↗
“Function tools get automatic schema generation and Pydantic-powered validation”
Rely on strict typing and schema validation so a coding agent catches its own mistakes at build timepartialproof ↗
“Function tools get automatic schema generation and Pydantic-powered validation”
Get schema-validated structured output from an agent, with automatic retries when validation failspartialproof ↗
“Can use OpenAI-managed tools: web search, file search, code interpreter, hosted MCP, image generation”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Can set OPENAI_DEFAULT_MODEL env var to consistently use a specific model for agents without a custom model set”
Swap the underlying LLM provider or model without rewriting my agentfullproof ↗
“Provider-agnostic: supports OpenAI Responses/Chat Completions APIs plus 100+ other LLMs”
Swap the underlying LLM provider or model without rewriting my agentfullproof ↗
Undersold (13)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Delegate tasks to a built-in AI assistant inside the productpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Test against a sandbox environment without touching production datafullproof ↗
Run the framework's example agents headlessly from a terminal so an agent can verify what it just builtpartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Run my agents entirely on my own infrastructure with no dependence on the vendor's platformpartialproof ↗
Read the product's source under an open licensepartialproof ↗
Compose agents into an explicit graph or workflow with branching, loops, and parallel stepspartialproof ↗
Run long-lived agents durably across process restarts and deploys, natively or via durable-execution integrationspartialproof ↗
Claims outside our story set (6)
Real capability claims found in OpenAI Agents SDK’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Built-in WebSearchTool lets an agent search the web”
source ↗“Out-of-the-box support for OpenAI models in two flavors”
source ↗“Voice agents combine speech-to-text, an agent workflow, and text-to-speech into voice pipelines”
source ↗“Supports building text, sandbox, and voice agents with a small set of primitives”
source ↗“Built-in tools include WebSearchTool, FileSearchTool (OpenAI Vector Stores), and CodeInterpreterTool (sandboxed code execution)”
source ↗“RealtimeAgent/RealtimeSession automatically connects microphone and audio output in the browser via WebRTC”
source ↗
Business model
The Agents SDK is MIT-licensed and free; running agents is billed per token through the OpenAI API (or any provider you plug in), with tracing free on the OpenAI platform.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
