Access
Install
npm install -g @openai/codexShowcase


Products
OpenAI, product by product →OpenAI ships more than one product — each judged line competes in its own arena on the same stories as everyone else.
| Line | Arena | Rank | PA Score | Agent-ready |
|---|---|---|---|---|
| ChatGPT | AI Assistants | #1/9 | 38/100 | 49/100 |
| Codexthis page | AI Coding Agents | #6/13 | 33/100 | 46/100 |
| Agents SDK | Agent Frameworks & SDKs | #2/9 | 37/100 | 55/100 |
| Plugins | Agent Skills & Extensions | #4/5 | 19/100 | 33/100 |
Not yet judged (6 — no arena where they compete): ChatGPT Work · Image generation (GPT-Image-2.5) · Codex Security · ChatKit · Agents API · API Platform
Try itExperimental
See what an agent can do with Codex before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$codex --versionrecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Autonomy agents — stories about autonomy agents in this arenaAutonomy agentsevidence →
Stories about autonomy agents in this arena
Code generation — quality of generated code — correctness, style, fit to the codebaseCode generationevidence →
Quality of generated code — correctness, style, fit to the codebase
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understandingevidence →
How deeply the tool maps your repo — cross-file context, architecture awareness, history
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystemevidence →
Integrations, plugins, and third-party ecosystem stories
Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integrationevidence →
Meeting you in the IDE and terminal — extensions, inline flows, context
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safetyevidence →
Keeping generated changes safe — diffs, approvals, guardrails
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 2 paid · 0 enterprise · 53 not stated in evidence
Follow the green: where the map greys out is where Codex stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
~5/10
unlocks → Webhooks · Full data export · View interactive diffs and share selected code as context from within my JetBrains IDE
Subscribe to events via webhooks
—–
Build against official SDKs
~6/10
Issue scoped/least-privilege API credentials for an agent
~5/10
Connect an agent via an official MCP server
~5/10
Download a machine-readable API spec (OpenAPI or equivalent)
~4/10
Rely on versioned APIs with a documented deprecation policy
~3/10
Test against a sandbox environment without touching production data
~6/10
Explore an interactive API reference with runnable examples
~5/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
~4/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓8/10
Operate the product with natural-language commands
✓8/10
Plug MCP servers into this product so it can use their tools
✓9/10
Get AI-generated insights and suggestions from my data inside the product
~6/10
Set up automations that run autonomously in the background
✓7/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Autonomy agents — stories about autonomy agents in this arenaAutonomy agents
Stories about autonomy agents in this arena
Background execution
Parallel agents
Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
~5/10
Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation
Quality of generated code — correctness, style, fit to the codebase
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding
How deeply the tool maps your repo — cross-file context, architecture awareness, history
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem
Integrations, plugins, and third-party ecosystem stories
Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration
Meeting you in the IDE and terminal — extensions, inline flows, context
Start a task on one device and continue it later from another device or browser
✓8/10
Ide integration
Session management
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety
Keeping generated changes safe — diffs, approvals, guardrails
Sorted by importance (agentic first) (high → low) · 74/74 stories · click a row’s chevron for the rationale and evidence
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Cclaimed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Cclaimed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 5/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 5/10 | Xcommunity | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Xcommunity | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Xcommunity | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 3/10 | Cclaimed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 6/10 | Xcommunity | |
Run a coding agent locally from my terminal C Terminal workflow | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 3 | full | 9/10 | Xcommunity | |
Chat with the coding assistant directly inside my IDE for contextual help C Ide integration | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 3 | full | 8/10 | Xcommunity | |
Describe a feature or bug in plain language and have the agent implement or fix it across multiple files C Feature implementation | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | full | 8/10 | Xcommunity | |
Inspect diffs and run checks to catch problems before merging C Pr review | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 3 | full | 8/10 | Cclaimed | |
Turn a tracked issue into a complete pull request end-to-end C Feature implementation | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | full | 8/10 | Cclaimed | |
Delegate longer-running coding tasks to run in the background in an isolated cloud environment C Background execution | developer | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 3 | full | 7/10 | Cclaimed | |
Reproduce issues, narrow down root causes, and verify fixes C Issue diagnosis | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | full | 7/10 | Cclaimed | |
Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context C Tool integration | developer | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 3 | partial | 6/10 | Cclaimed | |
Have the agent stage changes, write commit messages, create branches, and open pull requests C Pr review | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 3 | partial | 6/10 | Cclaimed | |
Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me C Maintenance automation | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | partial | 6/10 | Xcommunity | |
Understand how a codebase fits together to find where to start making changes C Codebase mapping | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | partial | 6/10 | Cclaimed | |
Get automatic code review with contextual feedback on every pull request C Pr review | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 3 | partial | 5/10 | Cclaimed | |
Have the agent map and explain an entire unfamiliar codebase without manually selecting context files C Codebase mapping | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | partial | 5/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Add a project instructions file to set coding standards and conventions the agent follows C Context management | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | none | untested | none yet | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Receive inline code completions and next-edit suggestions as I type C Code completion | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | n/a | untested | none yet | |
Sign in with my existing product subscription plan to use the coding agent C Authentication | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | fullpaid | 9/10 | Cclaimed | |
Have a cloud agent build, test, and demo a feature end-to-end for my review C Background execution | ai-native user | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | full | 8/10 | Xcommunity | |
Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack C Tool integration | developer | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 2 | full | 8/10 | Cclaimed | |
Run the agent non-interactively in scripts for workflow automation C Terminal workflow | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | full | 8/10 | Cclaimed | |
Start a task on one device and continue it later from another device or browser C Cross device continuity | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | full | 8/10 | Cclaimed | |
Debug issues and troubleshoot using natural-language queries C Debugging | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 2 | partial | 7/10 | Xcommunity | |
Have the agent operate inside a sandbox when interacting with code, tools, and network resources C Safe execution | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 2 | full | 7/10 | Xcommunity | |
Manage multiple agent-driven coding sessions from one unified workspace C Session management | engineering-lead | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | full | 7/10 | Xcommunity | |
Configure a reproducible cloud environment with the dependencies and setup steps my repository needs C Background execution | developer | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | partial | 6/10 | Cclaimed | |
Control which external tools and integrations the agent is allowed to access C Safe execution | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 2 | partial | 6/10 | Cclaimed | |
Launch fleets of autonomous agents that work in parallel on different tasks for hours or days C Parallel agents | ai-native user | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | partial | 6/10 | Xcommunity | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Authenticate through an enterprise identity or cloud platform for compliance and scalability G Authentication | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | partial | 5/10 | Cclaimed | |
Authenticate with an API key instead of an account login G Authentication | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | partial | 5/10 | Xcommunity | |
Review diffs visually and run multiple sessions side by side in a desktop app C Session management | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | partial | 5/10 | Xcommunity | |
Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously C Scheduled automation | ai-native user | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | partial | 5/10 | Cclaimed | |
Generate a working app from a sketch, image, or PDF design C Multimodal generation | ai-native user | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 2 | partial | 4/10 | Cclaimed | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 4/10 | Xcommunity | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 3/10 | Xcommunity | |
Get contextual explanations and automatic fixes for security vulnerabilities C Security checks | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 2 | partial | 3/10 | Cclaimed | |
Have the agent build and recall memory automatically across sessions C Context management | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 2 | partial | 3/10 | Cclaimed | |
Choose which underlying AI model powers my session from multiple providers C Model choice | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Include multiple project directories in a single session for broader context C Context management | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Sign in with a personal account to get free-tier access without managing API keys G Authentication | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | fullpaid | 8/10 | Cclaimed | |
Equip the agent with custom skills to perform specialized tasks C Marketplace | developer | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 1 | full | 7/10 | Cclaimed | |
Integrate third-party partner-built agent apps into my workflows C Marketplace | engineering-lead | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 1 | partial | 6/10 | Cclaimed | |
Run several task attempts in parallel and compare results before choosing one C Parallel agents | developer | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 1 | partial | 6/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 4/10 | Cclaimed | |
Create a shared workspace from my docs and repos as a common source of truth for the team C Team knowledge | engineering-lead | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 1 | none | 0/10 | ||
Let the tool automatically pick the best model for each task C Model choice | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | none | 0/10 | ||
View interactive diffs and share selected code as context from within my JetBrains IDE C Ide integration | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 1 | none | 0/10 | ||
Debug a live running web application directly from my coding assistant C Debugging | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 1 | n/a | untested | none yet | |
Opt out of having my code and prompts used for AI model training C Data governance | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 1 | none | untested | none yet | |
See license and public-code matching references for AI-suggested code C Security checks | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 50 stories with headroom
What would move Codex’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events
nonemoves PA Scoreimpact 30
Missing: evidence of a rules/trigger definition interface, conditional logic configuration, or event-to-action mapping system that users can author themselves.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
Missing: any documented data export tool, open-format export (JSON/Markdown dump), or data portability statement.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Codex CLI runs locally but requires signing into a ChatGPT account or OpenAI API key, and the core inference/model and cloud environments are OpenAI-hosted only; there is no self-hosted backend option.
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyAdd a project instructions file to set coding standards and conventions the agent follows
nonemoves PA Scoreimpact 30
The evidence pack covers Codex's CLI, cloud, MCP, and review features but contains no mention of a project-level instructions/config file (e.g., AGENTS.md or similar) for setting coding standards or conventions the agent should follow.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
The evidence pack contains no mention of data-training opt-out controls, enterprise data usage policies, or privacy settings for excluding user data from model training; it covers CLI features, MCP, RBAC, and community sentiment but nothing about training-data exclusion.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
No evidence in the pack mentions webhooks or event subscription capabilities for Codex; the product exposes MCP servers, CLI, and cloud task integrations but nothing about outbound webhook events for AI-native consumers.
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
partialq5/10moves agent-readyimpact 22.5
Missing: independent corroboration that the current 'Codex app server' MCP mode works reliably in production, and clearer first-party documentation of its interface now that the original is deprecated.
Agenticness — how well agents can access and operate the productDrive the product through a documented public API
partialq5/10moves agent-readyimpact 22.5
Missing: a single stable, non-deprecated documented public API surface, confirmation that the current model is API-accessible, and independent corroboration that third parties successfully drive Codex via this API.
Showing the top 8 of 50 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 55 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs41 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Plug MCP servers into this product so it can use their tools
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Rely on versioned APIs with a documented deprecation policy
- Perform bulk operations across many items at once
- Version, review, and roll back my automations
- Have a cloud agent build, test, and demo a feature end-to-end for my review
- Delegate longer-running coding tasks to run in the background in an isolated cloud environment
- Debug issues and troubleshoot using natural-language queries
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
- Generate a working app from a sketch, image, or PDF design
- Understand how a codebase fits together to find where to start making changes
- Have the agent map and explain an entire unfamiliar codebase without manually selecting context files
- Have the agent build and recall memory automatically across sessions
- Reproduce issues, narrow down root causes, and verify fixes
- Equip the agent with custom skills to perform specialized tasks
- Integrate third-party partner-built agent apps into my workflows
- Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
- Start a task on one device and continue it later from another device or browser
- Chat with the coding assistant directly inside my IDE for contextual help
- Review diffs visually and run multiple sessions side by side in a desktop app
- Manage multiple agent-driven coding sessions from one unified workspace
- Run a coding agent locally from my terminal
- Run the agent non-interactively in scripts for workflow automation
- Do everything through the API that I can do in the UI
- Have the agent stage changes, write commit messages, create branches, and open pull requests
- Get automatic code review with contextual feedback on every pull request
- Inspect diffs and run checks to catch problems before merging
- Control which external tools and integrations the agent is allowed to access
- Have the agent operate inside a sandbox when interacting with code, tools, and network resources
- Get contextual explanations and automatic fixes for security vulnerabilities
docs25 stories
- Point an agent at llms.txt or agent-oriented docs
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Have a cloud agent build, test, and demo a feature end-to-end for my review
- Delegate longer-running coding tasks to run in the background in an isolated cloud environment
- Configure a reproducible cloud environment with the dependencies and setup steps my repository needs
- Launch fleets of autonomous agents that work in parallel on different tasks for hours or days
- Run several task attempts in parallel and compare results before choosing one
- Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
- Turn a tracked issue into a complete pull request end-to-end
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Integrate third-party partner-built agent apps into my workflows
- Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
- Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack
- Start a task on one device and continue it later from another device or browser
- Review diffs visually and run multiple sessions side by side in a desktop app
- Manage multiple agent-driven coding sessions from one unified workspace
- Have the agent stage changes, write commit messages, create branches, and open pull requests
- Get automatic code review with contextual feedback on every pull request
- Inspect diffs and run checks to catch problems before merging
- Have the agent operate inside a sandbox when interacting with code, tools, and network resources
GitHub README24 stories
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Rely on versioned APIs with a documented deprecation policy
- Delegate longer-running coding tasks to run in the background in an isolated cloud environment
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Have the agent map and explain an entire unfamiliar codebase without manually selecting context files
- Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
- Chat with the coding assistant directly inside my IDE for contextual help
- Review diffs visually and run multiple sessions side by side in a desktop app
- Run a coding agent locally from my terminal
- Run the agent non-interactively in scripts for workflow automation
- Do everything through the API that I can do in the UI
- Read the product's source under an open license
- Authenticate with an API key instead of an account login
- Authenticate through an enterprise identity or cloud platform for compliance and scalability
- Sign in with my existing product subscription plan to use the coding agent
- Sign in with a personal account to get free-tier access without managing API keys
Hacker News17 stories
- Drive the product through a documented public API
- Build against official SDKs
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Have a cloud agent build, test, and demo a feature end-to-end for my review
- Launch fleets of autonomous agents that work in parallel on different tasks for hours or days
- Debug issues and troubleshoot using natural-language queries
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
- Chat with the coding assistant directly inside my IDE for contextual help
- Review diffs visually and run multiple sessions side by side in a desktop app
- Manage multiple agent-driven coding sessions from one unified workspace
- Run a coding agent locally from my terminal
- Do everything through the API that I can do in the UI
- Read the product's source under an open license
- Authenticate with an API key instead of an account login
- Have the agent operate inside a sandbox when interacting with code, tools, and network resources
API reference8 stories
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Rely on versioned APIs with a documented deprecation policy
- Do everything through the API that I can do in the UI
- Authenticate through an enterprise identity or cloud platform for compliance and scalability
- Control which external tools and integrations the agent is allowed to access
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$codex --versionreproduced$ codex --version codex-cli 0.142.0
$codex exec --helpreproduced$ codex exec --help
Run Codex non-interactively
Usage: codex exec [OPTIONS] [PROMPT]
codex exec [OPTIONS] <COMMAND> [ARGS]
Commands:
resume Resume a previous session by id or pick the most recent with --last
review Run a code review against the current repository
help Print this message or the help of the given subcommand(s)
Arguments:
[PROMPT]
Initial instructions for the agent. If not provided as an argument (or if `-` is used),
instructions are read from stdin. If stdin is piped and a prompt is also provided, stdin
is appended as a `<stdin>` block
Options:
-c, --config <[redacted]=value>
Override a configuration value that would otherwise be loaded from `~/.codex/config.toml`.
Use a dotted path (`foo.bar.baz`) to override nested values. The `value` portion is parsed
as TOML. If it fails to parse as TOML, the raw string is used as a literal.
Examples: - `-c model="o3"` - `-c 'sandbox_permissions=["di[redacted]full-read-access"]'` - `-c
shell_environment_policy.inherit=all`
--enable <FEATURE>
Enable a feature (repeatable). Equivalent to `-c features.<name>=true`
--disable <FEATURE>
Disable a feature (repeatable). Equivalent to `-c features.<name>=false`
--strict-config
Error out when config.toml contains fields that are not recognized by this version of
Codex
-i, --image <FILE>...
Optional image(s) to attach to the initial prompt
-m, --model <MODEL>
Model the agent should use
--oss
Use open-source provider
--local-provider <OSS_PROVIDER>
Specify which local provider to use (lmstudio or ollama). If not specified with --oss,
will use config default or show selection
-p, --profile <CONFIG_PROFILE_V2>
Layer $CODEX_HOME/<name>.config.toml on top of the base user config
-s, --sandbox <SANDBOX_MODE>
Select the sandbox policy to use when executing model-generated shell commands
[possible values: read-only, workspace-write, danger-full-access]
--dangerously-bypass-approvals-and-sandbox
Skip all confirmation prompts and execute commands without sandboxing. EXTREMELY
DANGEROUS. Intended solely for running in environments that are externally sandboxed
--dangerously-bypass-hook-trust
Run enabled hooks without requiring persisted hook trust for this invocation. DANGEROUS.
Intended only for automation that already vets hook sources
-C, --cd <DIR>
Tell the agent to use the specified directory as its working root
--add-dir <DIR>
Additional directories that should be writable alongside the primary workspace
--skip-git-repo-check
Allow running Codex outside a Git repository
--ephemeral
Run without persisting session files to disk
--ignore-user-config
Do not load `$CODEX_HOME/config.toml`; auth still uses `CODEX_HOME`
--ignore-rules
Do not load user or project execpolicy `.rules` files
--output-schema <FILE>
Path to a JSON Schema file describing the model's final response shape
--color <COLOR>
Specifies color settings for use in the output
[default: auto]
[possible values: always, never, auto]
--json
Print events to stdout as JSONL
-o, --output-last-message <FILE>
Specifies file where the last message from the agent should be written
-h, --help
Print help (see a summary with '-h')
-V, --version
Print version
$echo '<jsonrpc initialize>' | codex mcp-serverreproduced$ echo '<jsonrpc initialize>' | codex mcp-server
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"codex-mcp-server","title":"Codex","version":"0.142.0","user_agent":"codex_cli_rs/0.142.0 (Mac OS 26.5.2; arm64) xterm-256color (productarena-probe; 1.0)"}}}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
6 of 22 testable claims verified · 0 contradicted → integrity 27/100
28 distinct capability claims found in Codex’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
6
Verified
16
Unverified
0
Contradicted
33
Undersold
Verified (7)
“Run coding tasks in isolated cloud environments in parallel without using local machine”
Launch fleets of autonomous agents that work in parallel on different tasks for hours or dayspartialproof ↗
“Codex CLI is a coding agent that runs locally on your computer”
“Install Codex directly inside code editors like VS Code, Cursor, and Windsurf”
Chat with the coding assistant directly inside my IDE for contextual helpfullproof ↗
“Use Codex with an API key as an alternative authentication method”
Authenticate with an API key instead of an account loginpartialproof ↗
“Inspect code, make changes, run commands, and automate repeatable work from the terminal”
“Configure when Codex may edit files or run commands without approval, and inspect the sandbox/writable roots”
Have the agent operate inside a sandbox when interacting with code, tools, and network resourcesfullproof ↗
“Install Codex CLI via npm package”
Unverified (23)
“Run coding tasks in isolated cloud environments in parallel without using local machine”
Delegate longer-running coding tasks to run in the background in an isolated cloud environmentfullproof ↗
“Start tasks from web, GitHub, GitLab, Linear, or Slack”
Kick off agent tasks directly from GitHub, GitLab, Linear, or Slackfullproof ↗
“Configure dependencies, tools, variables, and setup steps needed per repository for cloud environment”
Configure a reproducible cloud environment with the dependencies and setup steps my repository needspartialproof ↗
“Inspect a summary and diff, request follow-up work, or open a pull request once a task completes”
Inspect diffs and run checks to catch problems before mergingfullproof ↗
“Delegate a longer task to run and check back later when it's done”
Delegate longer-running coding tasks to run in the background in an isolated cloud environmentfullproof ↗
“Start and review agent work from either the web or the CLI”
Start a task on one device and continue it later from another device or browserfullproof ↗
“Sign in with existing ChatGPT Plus/Pro/Business/Edu/Enterprise plan to use Codex”
Sign in with my existing product subscription plan to use the coding agentfullproof ↗
“Inspect code, make changes, run commands, and automate repeatable work from the terminal”
Run the agent non-interactively in scripts for workflow automationfullproof ↗
“Explore unfamiliar code, plan changes, edit files, and run local dev tools from within a repo session”
Understand how a codebase fits together to find where to start making changespartialproof ↗
“Run a dedicated code review against uncommitted changes, a commit, or a base branch with prioritized findings”
Get automatic code review with contextual feedback on every pull requestpartialproof ↗
“Reopen or search past local chat sessions from a repository to return to earlier work”
Have the agent build and recall memory automatically across sessionspartialproof ↗
“Attach an image like an error screenshot or design reference to a prompt”
Generate a working app from a sketch, image, or PDF designpartialproof ↗
“Delegate focused subtasks to specialized sub-agents and merge their findings back into the main session”
Equip the agent with custom skills to perform specialized tasksfullproof ↗
“Browse chats, submit work to a cloud environment, and apply results back to local repo from the terminal”
Start a task on one device and continue it later from another device or browserfullproof ↗
“Add local or remote MCP servers and inspect available tools before use”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Configure when Codex may edit files or run commands without approval, and inspect the sandbox/writable roots”
Control which external tools and integrations the agent is allowed to accesspartialproof ↗
“Run Codex non-interactively for repeatable scripted workflows”
Run the agent non-interactively in scripts for workflow automationfullproof ↗
“Package repeatable instructions as skills and add plugins to connect Codex to team tools/data”
Equip the agent with custom skills to perform specialized tasksfullproof ↗
“Package repeatable instructions as skills and add plugins to connect Codex to team tools/data”
Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its contextpartialproof ↗
“Shared MCP server configuration works across ChatGPT desktop app, CLI, and IDE extension without redoing setup”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Add MCP servers via CLI command with environment variables”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Run Codex as an MCP server exposing a JSON-RPC API over MCP transport so other MCP clients can control it”
“Legacy MCP server command still usable for existing integrations, though deprecated in favor of app server”
Undersold (33)
Point an agent at llms.txt or agent-oriented docspartialproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Drive the product through a documented public APIpartialproof ↗
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Set up automations that run autonomously in the backgroundfullproof ↗
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
Operate the product with natural-language commandsfullproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)partialproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Rely on versioned APIs with a documented deprecation policypartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Have a cloud agent build, test, and demo a feature end-to-end for my reviewfullproof ↗
Run several task attempts in parallel and compare results before choosing onepartialproof ↗
Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomouslypartialproof ↗
Debug issues and troubleshoot using natural-language queriespartialproof ↗
Turn a tracked issue into a complete pull request end-to-endfullproof ↗
Describe a feature or bug in plain language and have the agent implement or fix it across multiple filesfullproof ↗
Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for mepartialproof ↗
Have the agent map and explain an entire unfamiliar codebase without manually selecting context filespartialproof ↗
Reproduce issues, narrow down root causes, and verify fixesfullproof ↗
Integrate third-party partner-built agent apps into my workflowspartialproof ↗
Review diffs visually and run multiple sessions side by side in a desktop apppartialproof ↗
Manage multiple agent-driven coding sessions from one unified workspacefullproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Read the product's source under an open licensepartialproof ↗
Authenticate through an enterprise identity or cloud platform for compliance and scalabilitypartialproof ↗
Sign in with a personal account to get free-tier access without managing API keysfullproof ↗
Have the agent stage changes, write commit messages, create branches, and open pull requestspartialproof ↗
Get contextual explanations and automatic fixes for security vulnerabilitiespartialproof ↗
Claims outside our story set (2)
Real capability claims found in Codex’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Switch a run to live web search when task needs current external information”
source ↗“Generate shell completions, choose a syntax theme, and open long prompts in your configured editor”
source ↗
Business model
Free tier, plus ChatGPT Plus/Pro/Business/Enterprise subscriptions include Codex; also API pay-per-token usage and purchasable credits.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
