Skip to content

Rank #6 of 13 in AI Coding Agents

123.1k87.2k/yr +2.1k

Showcase

Codex homepage screenshot
homepage · captured Sep 2026 · view live ↗
Codex docs screenshot
docs · captured Sep 2026 · view live ↗

OpenAI ships more than one product — each judged line competes in its own arena on the same stories as everyone else.

LineArenaRankPA Score
ChatGPTAI Assistants#1/938/100
Codexthis pageAI Coding Agents#6/1333/100
Agents SDKAgent Frameworks & SDKs#2/937/100
PluginsAgent Skills & Extensions#4/519/100

Not yet judged (6 — no arena where they compete): ChatGPT Work · Image generation (GPT-Image-2.5) · Codex Security · ChatKit · Agents API · API Platform

Try itExperimental

See what an agent can do with Codex before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$codex --versionrecorded session — replayed, not live
recorded 2026-09-03 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

47.6/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

12.0/100

Autonomy agents — stories about autonomy agents in this arenaAutonomy agentsevidence →

Stories about autonomy agents in this arena

50.8/100

Code generation — quality of generated code — correctness, style, fit to the codebaseCode generationevidence →

Quality of generated code — correctness, style, fit to the codebase

55.4/100

Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understandingevidence →

How deeply the tool maps your repo — cross-file context, architecture awareness, history

27.7/100

Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystemevidence →

Integrations, plugins, and third-party ecosystem stories

46.8/100

Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integrationevidence →

Meeting you in the IDE and terminal — extensions, inline flows, context

68.7/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

8.4/100

Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →

Free-tier ceilings, usage caps, and rate limits before you have to pay

38.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safetyevidence →

Keeping generated changes safe — diffs, approvals, guardrails

40.4/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 0 free · 2 paid · 0 enterprise · 53 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 74/74 stories · click a row’s chevron for the rationale and evidence

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10C

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10C

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3partial5/10C

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3partial5/10X

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10X

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10X

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10C

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10C

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10C

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10C

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10T

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial3/10C

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial6/10X

Run a coding agent locally from my terminal C

Terminal workflow

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration3full9/10X

Chat with the coding assistant directly inside my IDE for contextual help C

Ide integration

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration3full8/10X

Describe a feature or bug in plain language and have the agent implement or fix it across multiple files C

Feature implementation

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3full8/10X

Inspect diffs and run checks to catch problems before merging C

Pr review

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety3full8/10C

Turn a tracked issue into a complete pull request end-to-end C

Feature implementation

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3full8/10C

Delegate longer-running coding tasks to run in the background in an isolated cloud environment C

Background execution

developerAutonomy agents — stories about autonomy agents in this arenaAutonomy agents3full7/10C

Reproduce issues, narrow down root causes, and verify fixes C

Issue diagnosis

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3full7/10C

Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context C

Tool integration

developerEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem3partial6/10C

Have the agent stage changes, write commit messages, create branches, and open pull requests C

Pr review

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety3partial6/10C

Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me C

Maintenance automation

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3partial6/10X

Understand how a codebase fits together to find where to start making changes C

Codebase mapping

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3partial6/10C

Get automatic code review with contextual feedback on every pull request C

Pr review

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety3partial5/10C

Have the agent map and explain an entire unfamiliar codebase without manually selecting context files C

Codebase mapping

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3partial5/10C

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3none0/10

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Add a project instructions file to set coding standards and conventions the agent follows C

Context management

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3noneuntestednone yet

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3noneuntestednone yet

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

Receive inline code completions and next-edit suggestions as I type C

Code completion

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3n/auntestednone yet

Sign in with my existing product subscription plan to use the coding agent C

Authentication

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2fullpaid9/10C

Have a cloud agent build, test, and demo a feature end-to-end for my review C

Background execution

ai-native userAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2full8/10X

Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack C

Tool integration

developerEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem2full8/10C

Run the agent non-interactively in scripts for workflow automation C

Terminal workflow

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2full8/10C

Start a task on one device and continue it later from another device or browser C

Cross device continuity

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2full8/10C

Debug issues and troubleshoot using natural-language queries C

Debugging

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation2partial7/10X

Have the agent operate inside a sandbox when interacting with code, tools, and network resources C

Safe execution

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety2full7/10X

Manage multiple agent-driven coding sessions from one unified workspace C

Session management

engineering-leadIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2full7/10X

Configure a reproducible cloud environment with the dependencies and setup steps my repository needs C

Background execution

developerAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2partial6/10C

Control which external tools and integrations the agent is allowed to access C

Safe execution

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety2partial6/10C

Launch fleets of autonomous agents that work in parallel on different tasks for hours or days C

Parallel agents

ai-native userAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2partial6/10X

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial6/10C

Authenticate through an enterprise identity or cloud platform for compliance and scalability G

Authentication

engineering-leadPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2partial5/10C

Authenticate with an API key instead of an account login G

Authentication

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2partial5/10X

Review diffs visually and run multiple sessions side by side in a desktop app C

Session management

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2partial5/10X

Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously C

Scheduled automation

ai-native userAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2partial5/10C

Generate a working app from a sketch, image, or PDF design C

Multimodal generation

ai-native userCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation2partial4/10C

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial4/10X

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial3/10X

Get contextual explanations and automatic fixes for security vulnerabilities C

Security checks

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety2partial3/10C

Have the agent build and recall memory automatically across sessions C

Context management

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding2partial3/10C

Choose which underlying AI model powers my session from multiple providers C

Model choice

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2none0/10

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Include multiple project directories in a single session for broader context C

Context management

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Sign in with a personal account to get free-tier access without managing API keys G

Authentication

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits1fullpaid8/10C

Equip the agent with custom skills to perform specialized tasks C

Marketplace

developerEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem1full7/10C

Integrate third-party partner-built agent apps into my workflows C

Marketplace

engineering-leadEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem1partial6/10C

Run several task attempts in parallel and compare results before choosing one C

Parallel agents

developerAutonomy agents — stories about autonomy agents in this arenaAutonomy agents1partial6/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial4/10C

Create a shared workspace from my docs and repos as a common source of truth for the team C

Team knowledge

engineering-leadEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem1none0/10

Let the tool automatically pick the best model for each task C

Model choice

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits1none0/10

View interactive diffs and share selected code as context from within my JetBrains IDE C

Ide integration

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration1none0/10

Debug a live running web application directly from my coding assistant C

Debugging

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation1n/auntestednone yet

Opt out of having my code and prompts used for AI model training C

Data governance

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety1noneuntestednone yet

See license and public-code matching references for AI-suggested code C

Security checks

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety1noneuntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 50 stories with headroom

What would move Codex’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events

    nonemoves PA Scoreimpact 30

    Missing: evidence of a rules/trigger definition interface, conditional logic configuration, or event-to-action mapping system that users can author themselves.

  2. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    Missing: any documented data export tool, open-format export (JSON/Markdown dump), or data portability statement.

  3. Openness — open source, data portability, and self-hosting storiesSelf-host the core product

    nonemoves PA Scoreimpact 30

    Codex CLI runs locally but requires signing into a ChatGPT account or OpenAI API key, and the core inference/model and cloud environments are OpenAI-hosted only; there is no self-hosted backend option.

  4. Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyAdd a project instructions file to set coding standards and conventions the agent follows

    nonemoves PA Scoreimpact 30

    The evidence pack covers Codex's CLI, cloud, MCP, and review features but contains no mention of a project-level instructions/config file (e.g., AGENTS.md or similar) for setting coding standards or conventions the agent should follow.

  5. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    The evidence pack contains no mention of data-training opt-out controls, enterprise data usage policies, or privacy settings for excluding user data from model training; it covers CLI features, MCP, RBAC, and community sentiment but nothing about training-data exclusion.

  6. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    No evidence in the pack mentions webhooks or event subscription capabilities for Codex; the product exposes MCP servers, CLI, and cloud task integrations but nothing about outbound webhook events for AI-native consumers.

  7. Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server

    partialq5/10moves agent-readyimpact 22.5

    Missing: independent corroboration that the current 'Codex app server' MCP mode works reliably in production, and clearer first-party documentation of its interface now that the original is deprecated.

  8. Agenticness — how well agents can access and operate the productDrive the product through a documented public API

    partialq5/10moves agent-readyimpact 22.5

    Missing: a single stable, non-deprecated documented public API surface, confirmation that the current model is API-accessible, and independent corroboration that third parties successfully drive Codex via this API.

Showing the top 8 of 50 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 55 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

docs41 stories

docs25 stories

GitHub README24 stories

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$codex --versionreproduced
$ codex --version
codex-cli 0.142.0
proves: Use an official CLIrecorded 2026-09-03
$codex exec --helpreproduced
$ codex exec --help
Run Codex non-interactively

Usage: codex exec [OPTIONS] [PROMPT]
       codex exec [OPTIONS] <COMMAND> [ARGS]

Commands:
  resume  Resume a previous session by id or pick the most recent with --last
  review  Run a code review against the current repository
  help    Print this message or the help of the given subcommand(s)

Arguments:
  [PROMPT]
          Initial instructions for the agent. If not provided as an argument (or if `-` is used),
          instructions are read from stdin. If stdin is piped and a prompt is also provided, stdin
          is appended as a `<stdin>` block

Options:
  -c, --config <[redacted]=value>
          Override a configuration value that would otherwise be loaded from `~/.codex/config.toml`.
          Use a dotted path (`foo.bar.baz`) to override nested values. The `value` portion is parsed
          as TOML. If it fails to parse as TOML, the raw string is used as a literal.

          Examples: - `-c model="o3"` - `-c 'sandbox_permissions=["di[redacted]full-read-access"]'` - `-c
          shell_environment_policy.inherit=all`

      --enable <FEATURE>
          Enable a feature (repeatable). Equivalent to `-c features.<name>=true`

      --disable <FEATURE>
          Disable a feature (repeatable). Equivalent to `-c features.<name>=false`

      --strict-config
          Error out when config.toml contains fields that are not recognized by this version of
          Codex

  -i, --image <FILE>...
          Optional image(s) to attach to the initial prompt

  -m, --model <MODEL>
          Model the agent should use

      --oss
          Use open-source provider

      --local-provider <OSS_PROVIDER>
          Specify which local provider to use (lmstudio or ollama). If not specified with --oss,
          will use config default or show selection

  -p, --profile <CONFIG_PROFILE_V2>
          Layer $CODEX_HOME/<name>.config.toml on top of the base user config

  -s, --sandbox <SANDBOX_MODE>
          Select the sandbox policy to use when executing model-generated shell commands

          [possible values: read-only, workspace-write, danger-full-access]

      --dangerously-bypass-approvals-and-sandbox
          Skip all confirmation prompts and execute commands without sandboxing. EXTREMELY
          DANGEROUS. Intended solely for running in environments that are externally sandboxed

      --dangerously-bypass-hook-trust
          Run enabled hooks without requiring persisted hook trust for this invocation. DANGEROUS.
          Intended only for automation that already vets hook sources

  -C, --cd <DIR>
          Tell the agent to use the specified directory as its working root

      --add-dir <DIR>
          Additional directories that should be writable alongside the primary workspace

      --skip-git-repo-check
          Allow running Codex outside a Git repository

      --ephemeral
          Run without persisting session files to disk

      --ignore-user-config
          Do not load `$CODEX_HOME/config.toml`; auth still uses `CODEX_HOME`

      --ignore-rules
          Do not load user or project execpolicy `.rules` files

      --output-schema <FILE>
          Path to a JSON Schema file describing the model's final response shape

      --color <COLOR>
          Specifies color settings for use in the output

          [default: auto]
          [possible values: always, never, auto]

      --json
          Print events to stdout as JSONL

  -o, --output-last-message <FILE>
          Specifies file where the last message from the agent should be written

  -h, --help
          Print help (see a summary with '-h')

  -V, --version
          Print version
$echo '<jsonrpc initialize>' | codex mcp-serverreproduced
$ echo '<jsonrpc initialize>' | codex mcp-server
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"codex-mcp-server","title":"Codex","version":"0.142.0","user_agent":"codex_cli_rs/0.142.0 (Mac OS 26.5.2; arm64) xterm-256color (productarena-probe; 1.0)"}}}

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

6 of 22 testable claims verified · 0 contradictedintegrity 27/100

28 distinct capability claims found in Codex’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

6

Verified

16

Unverified

0

Contradicted

33

Undersold

Verified (7)
Unverified (23)
Undersold (33)
Claims outside our story set (2)

Real capability claims found in Codex’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Switch a run to live web search when task needs current external information

    source ↗
  • Generate shell completions, choose a syntax theme, and open long prompts in your configured editor

    source ↗
Suggest a story for these →

Business model

free-tiersubscription-per-seatusage-basedcredits

Free tier, plus ChatGPT Plus/Pro/Business/Enterprise subscriptions include Codex; also API pay-per-token usage and purchasable credits.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score17 (Sep 1 '26)33 (Sep 16 '26)
Agent-ready14 (Aug 28 '26)46 (Sep 16 '26)

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data