Codex vs Claude Code
Claude Code
Anthropic
Claude Code wins · 16–31 (25 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to Claude CodeA dedicated llms.txt file is absent (404 at platform.openai.com/llms.txt), but Codex does publish machine-readable markdown docs (learn.chatgpt.com/docs/codex/cli.md) confirmed reachable by probe, which is an agent-friendly doc format an AI agent could be pointed at. Missing for 10: a standard llms.txt manifest, evidence of agents actually being pointed at these docs, and confirmation across all doc pages (docs/codex.md also 404s).
Claude Code itself ships llms.txt files (docs.claude.com/llms.txt, code.claude.com/docs/llms.txt) confirming it is agent-oriented-docs-aware for its own product, and its agentic search/MCP tooling means it can fetch and consume arbitrary web docs including llms.txt if pointed at them via URL fetch or MCP. However, there is no explicit documented feature or first-party guidance describing 'point Claude Code at llms.txt of a third-party site' as a supported workflow. missing for 10: explicit product feature/docs describing consuming arbitrary llms.txt/agent-oriented docs as a first-class capability, independent hands-on confirmation of this specific use case.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.claude.com/llms.txt # Anthropic Developer Documentation This file provides an overview of the Anth…”
- [probe] “PROBE docs-md: HTTP 200 at https://docs.claude.com/en/docs/claude-code/overview.md > ## Documentation Index > Fetch the complete documentati…”
- [claimed-docs] “Claude Code maps and explains entire codebases in a few seconds. It uses agentic search to understand project structure and dependencies wit…”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to Claude CodeCodex CLI explicitly documents non-interactive execution via `codex exec` for use in repeatable workflows, scripts, and CI/CD pipelines (codex-docs-19, codex-docs-32), and permissions/sandbox controls can be configured for unattended runs (codex-docs-17, codex-docs-39). Missing for 10: no independent case study or CI provider (e.g. GitHub Actions) integration example, and no explicit exit-code/output-format spec for CI parsing.
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
- [claimed-docs] “Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.”
- [claimed-docs] “Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.”
- [claimed-docs] “Set the boundaries for each run — /permissions: Choose when Codex can edit files or run commands without asking, and inspect the active sand…”
Docs explicitly describe running Claude Code in CI (GitHub Actions/GitLab CI/CD) for automated code review and issue triage, piping logs into it, and scheduled/headless runs for repeated automation tasks, plus GitHub Action integration for automatic PR review. This directly matches the headless/CI automation story with strong first-party documentation. Missing for 10: independent/hands-on confirmation of a working CI pipeline (community evidence is silent on CI usage specifically).
- [claimed-docs] “Claude Code is composable and follows the Unix philosophy. Pipe logs into it, run it in CI, or chain it with other tools”
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round drawnCodex CLI explicitly supports adding local/remote MCP servers via `codex mcp add`, inspecting available tools before use, and viewing active servers via `/mcp`; this configuration is shared across ChatGPT desktop app, CLI, and IDE extension. Docs also describe using MCP to connect to third-party tools like browsers or Figma. Missing for 10: independent hands-on verification of MCP tool usage in a real session beyond first-party docs.
- [claimed-docs] “Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.”
- [claimed-docs] “The ChatGPT desktop app, Codex CLI, and IDE extension share this configuration. Once you configure your MCP servers, you can switch among th…”
- [claimed-docs] “Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…”
- [claimed-docs] “Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation, or t…”
- [claimed-docs] “codex mcp add <server-name> --env VAR1=VALUE1 --env VAR2=VALUE2 -- <stdio server-command>”
- [claimed-docs] “In the `codex` TUI, use `/mcp` to see your active MCP servers.”
Claude Code has extensive first-party MCP documentation showing users can add MCP servers (e.g. `claude mcp add --transport http notion ...`), supporting stdio/HTTP transports, connecting to hundreds of external tools like Jira, Slack, Google Drive, Postgres, and even scaffolding new servers via a dev plugin. This is well corroborated across multiple doc pages with concrete CLI examples and use cases. Missing for 10: independent/hands-on community confirmation specifically of MCP tool usage (community evidence covers other topics, not MCP plugging in).
- [claimed-docs] “With MCP, Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom toolin…”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
- [claimed-docs] “Implement features from issue trackers: "Add the feature described in JIRA issue ENG-4521 and create a PR on GitHub."”
- [claimed-docs] “claude mcp add --transport http notion https://mcp.notion.com/mcp”
- [claimed-docs] “Stdio servers run as local processes on your machine. They're ideal for tools that need direct system access or custom scripts.”
- [claimed-docs] “You can also have Claude scaffold a server for you with the official mcp-server-dev plugin”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
ai-native userConnect an agent via an official MCP server
weight 3 · round to Claude CodeCodex explicitly supports running itself as an MCP server (codex mcp-server) so other MCP clients can connect, but OpenAI's own docs mark this interface 'experimental' and now 'deprecated', pointing users to a newer 'Codex app server' as the recommended replacement. This is a genuine server-mode capability (not just Codex-as-MCP-client), but the deprecation and lack of independent hands-on confirmation of the replacement's stability keep it from a full verdict. Missing for 10: independent corroboration that the current 'Codex app server' MCP mode works reliably in production, and clearer first-party documentation of its interface now that the original is deprecated.
- [github] “Codex MCP Server Interface [experimental]: a JSON-RPC API that runs over the Model Context Protocol (MCP) transport to control a local Codex…”
- [claimed-docs] “codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …”
- [claimed-docs] “Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.”
Claude Code documents `claude mcp serve` to run itself as a stdio MCP server that other applications can connect to, in addition to being an MCP client that connects to hundreds of external servers. missing for 10: independent/hands-on third-party confirmation of the `claude mcp serve` server mode in actual use.
- [claimed-docs] “Use Claude Code as an MCP server. You can use Claude Code itself as an MCP server that other applications can connect to: claude mcp serve (…”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
- [claimed-docs] “Stdio servers run as local processes on your machine. They're ideal for tools that need direct system access or custom scripts.”
- [claimed-docs] “claude mcp add --transport http notion https://mcp.notion.com/mcp”
ai-native userUse an official CLI
weight 2 · round drawnCodex ships an official, well-documented CLI (npm install -g @openai/codex) with rich agentic capabilities: local repo editing, exec/non-interactive scripting, MCP support, subagents, image input, sandbox/permissions control, cloud task delegation, and shell completions — all first-party documented and confirmed via GitHub repo and docs. Missing for 10: independent hands-on benchmarking specifically of CLI workflows (community evidence focuses mostly on model quality/UX rather than CLI mechanics) and some Linux-specific gaps noted by users.
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
- [github] “npm install -g @openai/codex”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
- [claimed-docs] “Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.”
- [claimed-docs] “Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…”
- [claimed-docs] “Split up a larger investigation — subagents: Ask Codex to delegate focused work to specialized agents, then bring their findings back into t…”
- [claimed-docs] “Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.”
- [claimed-docs] “Install the Codex CLI with the standalone installer for macOS and Linux.”
- [probe] “official CLI documented at https://learn.chatgpt.com/docs/codex/cli”
Claude Code is itself an official CLI tool with documented install (curl install script), usage (`cd project && claude`), cross-platform support (macOS/Linux/Windows), and deep terminal-native workflows (git, MCP, hooks, CI). GitHub repo and docs confirm first-party CLI status with active community usage corroborating real-world use. Missing for 10: independent benchmarking of CLI robustness/UX beyond mixed community sentiment.
- [claimed-docs] “cd your-project claude”
- [claimed-docs] “curl -fsSL https://claude.ai/install.sh | bash”
- [claimed-docs] “Available for macOS, Linux, and Windows.”
- [github] “Use it in your terminal, IDE, or tag @claude on Github.”
- [github] “helps you code faster by executing routine tasks, explaining complex code, and handling git workflows -- all through natural language comman…”
- [probe] “official CLI documented at https://code.claude.com/docs/en/setup”
ai-native userDrive the product through a documented public API
weight 3 · round to Claude CodeCodex documents multiple programmatic entry points — an MCP server interface for JSON-RPC control (though explicitly marked deprecated/experimental in favor of an undocumented 'app server'), a non-interactive `codex exec` mode for scripts/CI, and 'API key' usage — but these come with real caveats: API-key use 'requires additional setup', the flagship gpt-5.3-codex model was reportedly not yet available via API, and the primary MCP server route is deprecated rather than a stable first-class API. missing for 10: a single stable, non-deprecated documented public API surface, confirmation that the current model is API-accessible, and independent corroboration that third parties successfully drive Codex via this API.
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [github] “Codex MCP Server Interface [experimental]: a JSON-RPC API that runs over the Model Context Protocol (MCP) transport to control a local Codex…”
- [claimed-docs] “codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
- [claimed-docs] “Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.”
- [community] “gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'”
Claude Code exposes multiple documented programmatic surfaces: the Agent SDK for building custom agents with full control over orchestration/tools/permissions, a CLI (claude, claude mcp serve) that can be scripted/piped/run in CI, and ANTHROPIC_API_KEY-based direct API access, all documented in first-party docs. This goes beyond a closed UI and gives AI-native users documented, programmatic control paths. Missing for 10: independent/hands-on validation of the Agent SDK's API surface and no explicit REST/OpenAPI reference beyond the SDK and CLI docs.
- [claimed-docs] “the Agent SDK lets you build your own agents powered by Claude Code's tools and capabilities, with full control over orchestration, tool acc…”
- [claimed-docs] “Use Claude Code as an MCP server. You can use Claude Code itself as an MCP server that other applications can connect to: claude mcp serve (…”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
- [claimed-docs] “Claude Code is composable and follows the Unix philosophy. Pipe logs into it, run it in CI, or chain it with other tools”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round to CodexOpenAI's platform docs describe RBAC and project/org-scoped API keys/custom roles, and Codex can authenticate via an API key (codex-gh-4), so scoped credentials are technically available to a Codex-using account. However, none of the evidence ties this RBAC/API-key scoping specifically to configuring or restricting a Codex agent's own permissions — missing for 10: Codex-specific docs on issuing least-privilege keys for agent sessions, guidance on scoping credentials per-task/per-repo, and independent confirmation that this RBAC applies to Codex's own execution rather than just general API access.
- [claimed-docs] “Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [claimed-docs] “Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…”
Enterprise IAM docs mention role-based permissions, managed policy settings, and SSO/domain capture for org-wide configurations, plus sandboxing controls that restrict file/network access at runtime, suggesting some least-privilege controls exist. However, there is no explicit documentation of issuing scoped or limited-permission API keys/credentials specifically for an agent's use. Missing for 10: explicit scoped API key creation/management flow, granular credential scoping documentation, and independent verification of least-privilege credential issuance.
- [claimed-docs] “Claude for Enterprise: adds SSO, domain capture, role-based permissions, compliance API, and managed policy settings for organization-wide C…”
- [claimed-docs] “Single sign-on (SSO/SAML) and domain capture”
- [claimed-docs] “Learn how Claude Code's sandboxed Bash tool provides filesystem and network isolation for safer, more autonomous agent execution. The Bash s…”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “You can sign in to your Console account without creating an API key, even when your organization doesn't let developers create them.”
ai-native userBuild against official SDKs
weight 2 · round to Claude CodeCodex is a coding agent, but the evidence shows a genuine SDK-adjacent surface: the underlying OpenAI Responses API has an official OpenAPI spec and multi-language code samples (Python, TypeScript, Go, Ruby, Java, HTTP, CLI), and Codex integrates via CLI/MCP for programmatic extension. However, there is no evidence of an official Codex-specific SDK (as opposed to the general OpenAI API SDK), and API access for the Codex model itself is explicitly noted as not yet available. missing for 10: a dedicated Codex SDK/library distinct from the general OpenAI Responses API, confirmation that Codex agent capabilities (not just chat completions) are exposed via SDK, independent developer corroboration of building against these SDKs.
- [claimed-docs] “Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…”
- [github] “A machine-readable description of the OpenAI REST API, authored in OpenAPI 3.1.”
- [community] “gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'”
- [claimed-docs] “Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.”
Claude Code offers the official Agent SDK, letting developers build their own agents with full control over orchestration, tool access, and permissions, on top of Claude Code's tools/capabilities — a direct SDK for AI-native builders. This is backed by first-party docs and complemented by API-key-based programmatic access (ANTHROPIC_API_KEY) for direct integration. Missing for 10: independent/hands-on developer reports building production apps with the Agent SDK, and deeper docs on SDK language coverage/versioning.
- [claimed-docs] “the Agent SDK lets you build your own agents powered by Claude Code's tools and capabilities, with full control over orchestration, tool acc…”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
ai-native userSubscribe to events via webhooks
weight 2 · round to Claude CodeCodexnone0/10No evidence in the pack mentions webhooks or event subscription capabilities for Codex; the product exposes MCP servers, CLI, and cloud task integrations but nothing about outbound webhook events for AI-native consumers.
Claude Code doesn't offer a first-party webhook subscription feature, but docs note that an MCP server can act as a channel pushing events—including webhook events—into a Claude Code session while the user is away, enabling indirect event subscription via custom MCP tooling. Missing for 10: a native/first-party webhook subscription mechanism, official documentation or example of setting up webhook-triggered sessions, and independent confirmation this works in practice.
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to Claude CodeCodex generates AI-driven insights and suggestions specifically about code: it produces prioritized review findings, diffs, and summaries during automated reviews and delegated tasks (codex-docs-5, codex-docs-10, codex-docs-41, codex-docs-45), and can delegate to subagents for deeper investigation (codex-docs-35). However, this is scoped to code/repository data rather than general business or product data insights. Missing for 10: evidence of insight generation over non-code data sources, dashboards, or analytics-style summaries beyond code review findings.
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Split up a larger investigation — subagents: Ask Codex to delegate focused work to specialized agents, then bring their findings back into t…”
Claude Code generates AI-driven insights and suggestions from a user's data: it maps/explains entire codebases automatically, reviews code and PRs for security issues with explanations, and via MCP can query databases (e.g., PostgreSQL) or pull data from Slack/Jira/Google Drive to answer questions and suggest actions. This is all documented first-party capability with concrete examples (codebase mapping, automatic PR/security review, data queries via MCP). missing for 10: independent/hands-on corroboration specifically validating the quality of data-driven insights (community evidence is mostly about coding reliability, not insight generation), and no dedicated analytics/dashboard-style insight feature beyond code/data-source querying.
- [claimed-docs] “Claude Code maps and explains entire codebases in a few seconds. It uses agentic search to understand project structure and dependencies wit…”
- [claimed-docs] “Claude helps security teams and developers by reviewing code for security issues, drafts patches, and explains the risk in language your who…”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
- [claimed-docs] “Find emails of 10 random users who used feature ENG-4521, based on our PostgreSQL database.”
- [claimed-docs] “With MCP, Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom toolin…”
- [claimed-docs] “Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom tooling.”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to Claude CodeCodex cloud supports delegating longer tasks that run in isolated cloud environments in parallel, triggered from GitHub, GitLab, Linear, or Slack, and returning results (diff/PR) when ready — a clear background-automation workflow, and the CLI also supports non-interactive/repeatable workflows for scripted automation. Missing for 10: no documentation of scheduled/cron-style recurring triggers, and no independent/hands-on confirmation that long unattended background runs work reliably (community commentary focuses on interactive model quality/UX rather than background automation specifically).
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
Docs describe explicit background-automation features: scheduled runs for recurring tasks (morning PR reviews, overnight CI analysis, weekly audits), cloud/web sessions for long-running tasks you check back on, GitHub Actions/GitLab CI integration for automated review/triage, and MCP servers that push events (Telegram/Discord/webhooks) into a session while the user is away. Missing for 10: independent/hands-on corroboration that scheduled or background jobs run reliably unattended over time (only first-party docs cited).
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “Step away from your desk and keep working from your phone or any browser with Remote Control”
- [claimed-docs] “Kick off a long-running task on the web or the Claude mobile app, then pull it into your terminal with `claude --teleport`.”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round drawnCodex documents explicit task delegation to its built-in agent, both for long-running cloud tasks ('Delegate a longer task and return when it is ready') and for sub-agent delegation within a session ('Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session'), backed by detailed CLI/cloud docs. Missing for 10: independent hands-on verification specifically of the subagent delegation flow (community evidence discusses general agent quality/UX but not this feature directly).
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.”
- [claimed-docs] “Split up a larger investigation — subagents: Ask Codex to delegate focused work to specialized agents, then bring their findings back into t…”
- [claimed-docs] “Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …”
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
Claude Code's entire premise is delegating tasks to a built-in AI agent: docs describe it planning approaches, writing code across files, running tests, handling git workflows, and autonomously completing multi-step tasks (docs-2, docs-3, docs-20, docs-28, gh-3). This is corroborated by extensive first-party documentation and some community confirmation that it performs well as a coding assistant, though other community reports describe reliability issues and failure modes in autonomous execution. Missing for 10: independent benchmark-level validation of consistent task delegation success and stronger consensus on reliability (community reports show notable failure cases).
- [claimed-docs] “Describe what you want in plain language. Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “Claude Code works directly with git. It stages changes, writes commit messages, creates branches, and opens pull requests.”
- [claimed-docs] “Claude Code handles the tedious tasks that eat up your day: writing tests for untested code, fixing lint errors across a project, resolving …”
- [claimed-docs] “Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [github] “helps you code faster by executing routine tasks, explaining complex code, and handling git workflows -- all through natural language comman…”
- [community] “Claude is significantly better than other models at code assistant tasks, or at least in the way I use it.”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
- [community] “I've tried to use Claude code for a month now. It has a 100% failure rate so far. Comparing that to creating a project and just chatting wit…”
ai-native userOperate the product with natural-language commands
weight 2 · round drawnCodex CLI, IDE extension, cloud, and web surfaces are all operated by natural-language prompts/chats — e.g. starting tasks from prompts, resuming chats, delegating subagents, pasting images into the composer, and non-interactive `codex exec` for scripted natural-language instructions — all documented as the primary interaction mode across surfaces. Community threads corroborate heavy real-world use of this conversational/agentic workflow, even amid quality complaints about model performance. missing for 10: independent benchmarking specifically of natural-language command comprehension/robustness (community evidence is about overall agent quality/speed, not NL parsing specifically).
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
- [claimed-docs] “`codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
- [claimed-docs] “Bring visual context into the prompt — codex --image: Pass an error screenshot, architecture diagram, or design reference with the first pro…”
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
- [community] “Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…”
Claude Code is explicitly designed to be operated via plain-language instructions—describing tasks, git workflows, MCP tool use, and even natural-language chat commands (@claude in Slack, GitHub) all documented as core interaction modes, and GitHub docs explicitly state it works 'all through natural language commands.' missing for 10: independent hands-on benchmarking specifically confirming natural-language command comprehension breadth/accuracy versus slash-command or scripted usage, and some community reports note failure modes/hallucination under natural language instructions reducing reliability.
- [claimed-docs] “Describe what you want in plain language. Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “Claude Code works directly with git. It stages changes, writes commit messages, creates branches, and opens pull requests.”
- [github] “helps you code faster by executing routine tasks, explaining complex code, and handling git workflows -- all through natural language comman…”
- [claimed-docs] “Route tasks from team chat: mention @Claude in Slack with a bug report and get a pull request back”
- [claimed-docs] “cd your-project claude”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round to CodexThe evidence shows OpenAI's general API reference (developers.openai.com) has runnable, per-language code samples with live examples, which an AI-native user could explore. However this is the general OpenAI Responses API reference, not a Codex-specific interactive API reference, and Codex itself is documented as a CLI/agent product rather than an API with its own dedicated reference docs. Missing for 10: a Codex-specific API reference page, evidence of interactivity beyond code-sample selection (e.g., live sandbox execution), and any Codex-specific documentation of this reference.
- [claimed-docs] “Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
Claude Codenone0/10The evidence pack shows standard documentation pages and an Agent SDK reference, but nothing describing an interactive API reference with runnable/executable code examples (e.g., an in-browser sandbox or live API explorer). No such capability is evidenced anywhere in the docs, GitHub, or community items.
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round to CodexOpenAI publishes a machine-readable OpenAPI 3.1 spec for its REST API (codex-gh-9) and Codex can be used via that API (codex-gh-4), but the evidence never confirms this spec explicitly covers or is dedicated to Codex-specific endpoints, nor is there a direct 'download spec' link tied to Codex docs. missing for 10: a Codex-specific OpenAPI/spec file, explicit download instructions, or confirmation the general OpenAI OpenAPI spec includes Codex CLI/agent endpoints.
- [github] “A machine-readable description of the OpenAI REST API, authored in OpenAPI 3.1.”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [claimed-docs] “Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…”
Claude Codenone0/10The evidence pack shows Claude Code as a CLI/agent tool with SDK, MCP, and CI integrations, but no mention of a downloadable OpenAPI or equivalent machine-readable API spec for Claude Code itself. This axis is plausible for a product with an Agent SDK and API-key based access, but the pack contains no such artifact.
ai-native userTest against a sandbox environment without touching production data
weight 1 · round to CodexCodex offers isolated cloud task environments and CLI sandbox controls (writable roots, permission gating) that keep agent actions contained rather than acting directly on a live system, which functions as a sandbox layer for testing changes. However, there's no explicit documentation of test-vs-production data separation, and a community report raises unresolved concerns about the sandbox reading sensitive filesystem data without asking. Missing for 10: explicit production-data isolation guarantees, first-party documentation addressing the raised sandbox-safety concern, and independent verification that isolated environments never touch real prod data.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Configure the dependencies, tools, variables, and setup steps each repository needs.”
- [claimed-docs] “Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.”
- [community] “Does that version of Codex still read sensitive data on your file system without even asking? Just curious. [links to github.com/openai/code…”
Claude Code documents a sandboxed Bash tool that enforces filesystem and network isolation, letting Claude execute commands within OS-enforced boundaries rather than freely touching arbitrary systems — this supports the spirit of testing in isolation, but the docs don't specifically describe spinning up a 'sandbox vs production' environment or protecting production data per se. Missing for 10: explicit documentation of test/staging vs production environment separation, guidance on preventing production data access, and independent/hands-on validation that the sandbox reliably prevents production data exposure.
- [claimed-docs] “Learn how Claude Code's sandboxed Bash tool provides filesystem and network isolation for safer, more autonomous agent execution. The Bash s…”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round to CodexThere is a documented OpenAPI 3.1 spec and API reference (codex-gh-9, codex-docs-29) and one concrete example of a deprecation notice (codex mcp-server deprecated in favor of the Codex app server, codex-docs-23), showing some practice of versioning and deprecation. However, there is no comprehensive, documented deprecation policy (timelines, notice periods, version numbering scheme) covering the Codex/OpenAI API generally. Missing for 10: an explicit deprecation policy document, API version numbering scheme, and independent corroboration of adherence to it.
- [github] “A machine-readable description of the OpenAI REST API, authored in OpenAPI 3.1.”
- [claimed-docs] “Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…”
- [claimed-docs] “codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to CodexCodex supports running multiple cloud tasks in parallel across repos (codex-docs-1, codex-docs-3, codex-docs-6) and delegating focused work to specialized sub-agents within a session (codex-docs-13), which gives some bulk/parallel automation capability. However, there's no explicit evidence of a bulk operation primitive (e.g., batch-apply an action across many files/items/tickets in one command) — the parallelism described is task-level (multiple independent runs) rather than a documented 'operate over N items at once' feature. Missing for 10: explicit bulk/batch API or CLI verb for acting across many items in one invocation, and independent confirmation of large-scale parallel throughput in practice.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
Claude Codedisputedcontradicted6/10Claude Code's docs explicitly support bulk operations — fixing lint errors 'across a project', multi-file writes, spawning multiple agents to work on different parts of a task simultaneously, and running multiple sessions/tasks in parallel or on a schedule — which strongly matches the story. However, a hands-on community report describes a concrete failure mode during a bulk-style replace_all operation that corrupted code (turning a constant into 'GROQ_URL = GROQ_URL'), with the user stating you 'absolutely can't trust it to self-verify' on such operations, directly contradicting reliable execution of bulk changes at scale. Missing for 10: independent corroboration that large-scale bulk operations complete reliably without manual review, and resolution/acknowledgment of the reported failure mode.
- [claimed-docs] “writing tests for untested code, fixing lint errors across a project, resolving merge conflicts, updating dependencies, and writing release …”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously. A lead agent coordinates the work, assigns subtasks…”
- [claimed-docs] “Claude Code handles the tedious tasks that eat up your day: writing tests for untested code, fixing lint errors across a project, resolving …”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously.”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to Claude CodeCodexnone0/10Codex supports triggering tasks from external events (GitHub/GitLab/Linear/Slack) and running non-interactive workflows, but there is no evidence of a user-defined rules engine that lets users specify arbitrary trigger conditions and automated actions (e.g., 'on X event, do Y') — this is closer to integration hooks than a rules/automation framework. missing for 10: evidence of a rules/trigger definition interface, conditional logic configuration, or event-to-action mapping system that users can author themselves.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
Claude Code supports Hooks (shell commands triggered before/after actions like auto-formatting or lint on edits) and scheduled runs plus MCP channels (Telegram/Discord/webhook events) that push messages into a session automatically, which together constitute event-triggered automation rules. However, there's no unified declarative 'rules engine' with conditions/triggers documented — it's a patchwork of hooks, cron-like scheduling, and MCP event channels rather than a first-class rule-definition system. missing for 10: a unified rules/trigger definition UI or config, broader event types beyond hooks/schedule/MCP channels, and independent/hands-on validation of these automation triggers working reliably.
- [claimed-docs] “Hooks let you run shell commands before or after Claude Code actions, like auto-formatting after every file edit or running lint before a co…”
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to Claude CodeCodexnone0/10The evidence shows Codex can run in CI/scripts (codex exec), be triggered from GitHub/GitLab/Slack, and run cloud tasks, but there is no mention of a native recurring/scheduled job or cron-like trigger mechanism within Codex itself. Automation is triggered by external events or manual invocation, not scheduled recurrence.
- [claimed-docs] “Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …”
Docs explicitly describe running Claude Code on a schedule for recurring automation (PR reviews, CI failure analysis, dependency audits, doc syncing) and mention 'schedule recurring tasks' as a feature. Missing for 10: independent/hands-on confirmation of the scheduling mechanism and details on configuration (cron syntax, triggers, reliability).
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
ai-native userVersion, review, and roll back my automations
weight 1 · round to Claude CodeCodex's CLI includes a dedicated review command that inspects diffs/commits without modifying the working tree (codex-docs-10, codex-doces-41/45), and it operates within git repos so changes are inherently versioned and revertible via git; skills/plugins can be packaged as reusable automations (codex-docs-20/42). However, there is no documented mechanism to version, review, or roll back the automations/skills/workflows themselves (e.g., skill version history, rollback of a plugin config, audit trail for automation changes) — only code diffs are reviewed. Missing for 10: explicit versioning of skills/automations, a rollback UI/command for automation configs, and independent evidence of this workflow in practice.
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.”
- [claimed-docs] “Use skills and plugins: Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without l…”
Automations in Claude Code (CLAUDE.md, skills, hooks, slash commands) are plain files that live in the repo, so they inherit git's version history, and Claude Code natively works with git (staging, commits, diffs) and supports visual diff review (claude-code-docs-3, claude-code-docs-13, claude-code-docs-32, claude-code-docs-33, claude-code-docs-22). However, there is no dedicated feature for versioning/rolling back automations themselves (e.g., no automation-specific history log, no built-in 'revert this hook/skill run' or undo mechanism) — reviewers rely entirely on generic git workflows rather than a purpose-built automation-lifecycle tool. missing for 10: a dedicated automation versioning/audit history UI, an explicit rollback/undo command for skills or hooks, and independent hands-on confirmation that rollback of automations works as intended.
- [claimed-docs] “Claude Code works directly with git. It stages changes, writes commit messages, creates branches, and opens pull requests.”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Create skills to package repeatable workflows your team can share, like `/review-pr` or `/deploy-staging`.”
- [claimed-docs] “Hooks let you run shell commands before or after Claude Code actions, like auto-formatting after every file edit or running lint before a co…”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session. Use it to set coding standar…”
Autonomy agents — stories about autonomy agents in this arenaAutonomy agents
Stories about autonomy agents in this arena
Background execution
ai-native userHave a cloud agent build, test, and demo a feature end-to-end for my review
weight 2 · round to CodexCodex cloud lets users delegate tasks that run in isolated cloud environments, inspect summaries/diffs, request follow-ups, and open pull requests for review, effectively building/testing/demoing changes end-to-end for user review (codex-docs-1,5,6,7,37). Community commentary corroborates real-world agentic task completion, though with performance/reliability caveats. Missing for 10: independent hands-on verification specifically of the cloud (not CLI) workflow's demo/test artifacts, and no explicit mention of a 'demo' step (e.g., live preview) beyond diff/PR review.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Start and review work from the web or Codex CLI.”
- [claimed-docs] “Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …”
- [community] “Often Claude Code Opus 4.6, on hard enough problems, can do the impression of acting fast without really making progress. Then you spin the …”
- [community] “Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…”
Docs show Claude Code can run as a cloud/browser session for long-running tasks (web, mobile, remote control, teleport), plan and write code across files, write tests, and open PRs with diff review for others to inspect — covering build, test, and reviewable-artifact steps end-to-end without local setup (claude-code-docs-9,10,13,14,26,28,3,12). However there's no explicit 'demo' feature (e.g., live preview/staging deploy) beyond PR/diff review, and independent hands-on reports raise reliability concerns about self-verification on complex tasks. Missing for 10: dedicated demo/preview-environment tooling, independent corroboration of full cloud build-test-PR pipelines succeeding end-to-end without human intervention.
- [claimed-docs] “Step away from your desk and keep working from your phone or any browser with Remote Control”
- [claimed-docs] “Kick off a long-running task on the web or the Claude mobile app, then pull it into your terminal with `claude --teleport`.”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
- [claimed-docs] “Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “Claude Code works directly with git. It stages changes, writes commit messages, creates branches, and opens pull requests.”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
developerDelegate longer-running coding tasks to run in the background in an isolated cloud environment
weight 3 · round drawnOpenAI's docs describe a dedicated Codex cloud mode that runs tasks in isolated cloud environments, in parallel, triggered from web/GitHub/GitLab/Linear/Slack, with configurable repo setup and a workflow to inspect diffs/PRs on completion, plus a CLI command (`codex cloud`) to submit and later pull results locally — squarely matching the story of delegating longer background tasks to an isolated cloud environment. missing for 10: independent or hands-on community corroboration specifically validating the cloud/background execution feature (community evidence in the pack discusses CLI/app UX and model quality, not the cloud delegation flow itself).
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Configure the dependencies, tools, variables, and setup steps each repository needs.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …”
- [github] “If you are looking for the cloud-based agent from OpenAI, Codex Web, go to chatgpt.com/codex.”
- [github] “If you are looking for the <em>cloud-based agent</em> from OpenAI, <strong>Codex Web</strong>, go to <a href="https://chatgpt.com/codex">cha…”
Docs describe running Claude Code in-browser with no local setup, kicking off long-running tasks and checking back later, working on repos not present locally, running multiple tasks in parallel, and remote control/teleport features to move sessions between web/mobile and terminal — matching the delegate-to-cloud story directly. Missing for 10: independent/hands-on confirmation of the cloud environment's isolation guarantees (the sandboxing docs cited relate to local Bash tool isolation, not the cloud session itself) and details on how isolated/secure the cloud runtime is.
- [claimed-docs] “Step away from your desk and keep working from your phone or any browser with Remote Control”
- [claimed-docs] “Kick off a long-running task on the web or the Claude mobile app, then pull it into your terminal with `claude --teleport`.”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
developerConfigure a reproducible cloud environment with the dependencies and setup steps my repository needs
weight 2 · round to CodexCodex Cloud docs state you can configure the dependencies, tools, variables, and setup steps each repository needs for isolated cloud environments, directly matching the story. However, there is no detail on how reproducibility is guaranteed (e.g., container images, caching, version pinning) or independent hands-on confirmation of this setup workflow. Missing for 10: concrete configuration file/schema details, reproducibility guarantees, and independent verification of the setup working as documented.
- [claimed-docs] “Configure the dependencies, tools, variables, and setup steps each repository needs.”
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
Docs mention running Claude Code in the cloud/browser with no local setup and working on repos you don't have locally, implying some environment is provisioned, but there's no documentation of configuring a reproducible environment (e.g., setup scripts, dependency installation, devcontainer-style config) for cloud sessions. missing for 10: explicit environment/config file for cloud sandboxes, dependency installation steps, reproducibility guarantees across runs.
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
Parallel agents
ai-native userLaunch fleets of autonomous agents that work in parallel on different tasks for hours or days
weight 2 · round to Claude CodeCodex Cloud supports running multiple tasks in parallel in isolated cloud environments, triggered from GitHub/GitLab/Linear/Slack, and delegating longer tasks to return to later, which covers parallel/async agent work. However, there is no explicit evidence of orchestrating large 'fleets' of many simultaneous agents, no stated duration limits confirming multi-day autonomous runs, and community feedback highlights usage-limit throttling that would constrain sustained parallel/long-running fleets. missing for 10: evidence of fleet-scale orchestration (many concurrent agents), confirmed multi-day autonomous run duration, and independent confirmation that parallel tasks aren't throttled by usage limits.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Configure the dependencies, tools, variables, and setup steps each repository needs.”
- [community] “Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…”
- [community] “The main issue I have with Codex is that the best model is insanely slow, except at nights and weekends when Silicon Valley goes to bed... I…”
Docs explicitly describe spawning multiple Claude Code agents with a lead agent coordinating subtasks, running multiple sessions/tasks in parallel in the cloud, scheduling recurring/long-running tasks, and remote/teleport control to check back later — directly matching the fleet/parallel/long-duration story. Missing for 10: independent hands-on verification of multi-day unattended fleet runs and clearer guarantees on stability over very long horizons (community reports note reliability/quality drift over extended sessions).
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously. A lead agent coordinates the work, assigns subtasks…”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously.”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously. A lead agent coordin”
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
- [claimed-docs] “Step away from your desk and keep working from your phone or any browser with Remote Control”
- [claimed-docs] “Kick off a long-running task on the web or the Claude mobile app, then pull it into your terminal with `claude --teleport`.”
developerRun several task attempts in parallel and compare results before choosing one
weight 1 · round drawnDocs confirm Codex cloud can run tasks in parallel in isolated cloud environments without tying up the local machine, and results can be inspected (summary/diff) before choosing to follow up or open a PR — this covers running multiple attempts and reviewing outcomes. However, there's no explicit documentation of a dedicated 'compare multiple attempts side-by-side' UI/workflow, and no independent/community evidence confirming this parallel-comparison workflow works well in practice. missing for 10: explicit side-by-side comparison UI documentation, independent hands-on confirmation of comparing parallel attempts.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
Docs mention running 'multiple sessions side by side' and reviewing diffs visually in the web/cloud interface, plus running multiple tasks in parallel and spawning multiple agents—supporting parallel execution and comparison, though not explicitly framed as multiple attempts at the *same* task with a selection step. Missing for 10: explicit documentation of running several independent attempts at one identical task and a UI/workflow for choosing the best among them, and independent hands-on confirmation of this specific workflow.
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously. A lead agent coordinates the work, assigns subtasks…”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously.”
Scheduled automation
ai-native userSet up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
weight 2 · round to Claude CodeCodex cloud supports starting tasks from external triggers (GitHub/GitLab issues & PRs, Linear issues, Slack messages) and running them in parallel isolated environments, which covers the 'triggers' half of the story, but there's no evidence of a true schedule/cron-based always-on agent that proactively maintains a repo without an external event. Missing for 10: explicit scheduled/cron execution, evidence of continuous unattended monitoring/maintenance loops, and independent confirmation these triggers reliably run autonomous fixes end-to-end.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
First-party docs show robust support for scheduled/triggered automation: 'Run Claude on a schedule' for recurring maintenance tasks (docs-8), 'schedule recurring tasks' in the web UI (docs-13), MCP servers that push Telegram/Discord/webhook events into a session 'while you're away' (docs-31/54), and Slack @mentions triggering PRs (docs-11), plus CI integration for automated review/triage (docs-36). However, community reports raise real concerns about autonomous reliability over sustained/unsupervised runs (e.g. degrading output quality, self-verification failures, 'can't trust it to self-verify' — comm-16, comm-17, comm-19, comm-20), which tempers confidence that always-on autonomous maintenance works robustly in practice. Missing for 10: independent/hands-on validation that scheduled/triggered agents reliably self-maintain software over time without human correction, and no explicit multi-day/continuous 'always-on' uptime evidence beyond scheduled/triggered runs.
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “an MCP server can also act as a channel that pushes messages into your session, so Claude reacts to Telegram messages, Discord chats, or web…”
- [claimed-docs] “Route tasks from team chat: mention @Claude in Slack with a bug report and get a pull request back”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
- [community] “Whenever the phrase 'simplest fix' appears, it's time to pull the emergency break. This has gotten much worse over the past few weeks. It wi…”
- [community] “I've tried to use Claude code for a month now. It has a 100% failure rate so far. Comparing that to creating a project and just chatting wit…”
- [community] “A month ago the agents researched, designed, and implemented a compelling app idea with minimal guidance and felt super human. A month later…”
Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation
Quality of generated code — correctness, style, fit to the codebase
Debugging
developerDebug issues and troubleshoot using natural-language queries
weight 2 · round to Claude CodeCodex CLI docs show clear natural-language debugging workflows: exploring unfamiliar code, running local tools, passing error screenshots for context, and running dedicated code review that reports prioritized findings (codex-docs-8, codex-docs-9, codex-docs-10, codex-docs-12). However, community evidence shows mixed real-world reliability on agentic/coding tasks and no independent confirmation specifically validating debugging accuracy. Missing for 10: hands-on validation of debugging/troubleshooting accuracy, and independent case studies showing successful root-cause diagnosis via NL queries.
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Pass an error screenshot, architecture diagram, or design reference with the first prompt, or paste an image into the interactive composer.”
- [community] “Having used codex a fair bit I find it really struggles with … almost anything. However using the equivalent chat gpt model is fantastic.”
- [community] “Often Claude Code Opus 4.6, on hard enough problems, can do the impression of acting fast without really making progress. Then you spin the …”
Docs explicitly cover debugging: 'Debug live web applications' (Chrome integration), 'overnight CI failure analysis', explaining complex code, and codebase-wide understanding to trace issues via natural-language prompts. This is core positioning ('Build, debug, and ship from your terminal, IDE...'). missing for 10: independent hands-on validation specifically of debugging workflows (community evidence instead highlights reliability issues like self-verification failures and bugs introduced during edits, which are adjacent but not direct proof debugging-via-NL fails).
- [claimed-docs] “Debug live web applications | Chrome”
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “Work with Claude directly in your codebase. Build, debug, and ship from your terminal, IDE, Slack, web, and more.”
- [claimed-docs] “It understands your entire codebase and can work across multiple files and tools to get things done.”
- [github] “helps you code faster by executing routine tasks, explaining complex code, and handling git workflows -- all through natural language comman…”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
Feature implementation
developerTurn a tracked issue into a complete pull request end-to-end
weight 3 · round to CodexCodex explicitly supports starting work from a tracked issue (GitHub, GitLab, Linear) in cloud environments, running the task, inspecting the diff/summary, and opening a pull request when done, covering the full issue-to-PR loop. missing for 10: independent hands-on confirmation of a full issue-to-merged-PR workflow succeeding end-to-end, and detail on how issue context/acceptance criteria are actually parsed.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
Docs explicitly describe the full loop: reading tracked issues (Jira, GitHub, Slack) via MCP, generating code across multiple files, running tests, creating branches, and opening PRs — e.g. 'Add the feature described in JIRA issue ENG-4521 and create a PR on GitHub' and 'reading issues, writing code, running tests, and submitting PRs—all from your terminal.' Community reports corroborate real-world usage but also note reliability issues (self-verification failures, quality degradation over time), so results aren't guaranteed to be flawless end-to-end. Missing for 10: independent case studies quantifying success rate of full issue-to-PR automation, and detail on how failures/test verification are handled when the generated PR doesn't pass CI.
- [claimed-docs] “Claude Code works directly with git. It stages changes, writes commit messages, creates branches, and opens pull requests.”
- [claimed-docs] “Claude Code integrates with GitHub, GitLab, and your command line tools to handle the entire workflow—reading issues, writing code, running …”
- [claimed-docs] “Implement features from issue trackers: "Add the feature described in JIRA issue ENG-4521 and create a PR on GitHub."”
- [claimed-docs] “Add the feature described in JIRA issue ENG-4521 and create a PR on GitHub.”
- [claimed-docs] “Route tasks from team chat: mention @Claude in Slack with a bug report and get a pull request back”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
- [community] “I've tried to use Claude code for a month now. It has a 100% failure rate so far. Comparing that to creating a project and just chatting wit…”
developerDescribe a feature or bug in plain language and have the agent implement or fix it across multiple files
weight 3 · round to CodexCodex CLI and cloud docs describe the core loop of natural-language task description leading to autonomous file inspection, editing, running local tools, and producing a diff/PR (codex-docs-30, codex-docs-9, codex-docs-6, codex-docs-5), and community commentary corroborates it does real multi-file edits ('it edits the files and I can use the proper tooling', 'shockingly good... no worse than average L3-L4 engs') alongside some negative UX complaints that don't dispute the core capability. Missing for 10: independent benchmark/case-study evidence specifically confirming complex multi-file refactors across large codebases, and some community reports of it 'struggling with almost anything' create mild quality tension without rising to a concrete dispute.
- [claimed-docs] “Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
- [community] “Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…”
- [community] “Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…”
- [community] “Having used codex a fair bit I find it really struggles with … almost anything. However using the equivalent chat gpt model is fantastic.”
Claude Codedisputedcontradicted6/10Docs strongly claim the core capability: describe a feature/bug in plain language and Claude Code plans, implements, and verifies code changes across multiple files (claude-code-docs-2/28/51/20, claude-code-gh-3). However, hands-on community reports cite concrete failures undermining reliability of multi-file edits, e.g. a replace_all bug corrupting a constant (GROQ_URL=GROQ_URL) and inability to self-verify, plus a user reporting a '100% failure rate' and quality degradation over time (claude-code-comm-16, claude-code-comm-17, claude-code-comm-19, claude-code-comm-20), balanced against other users praising its code-assistant ability (claude-code-comm-5). missing for 10: consistent independent benchmarks confirming reliability across diverse multi-file tasks, resolution of reported failure modes.
- [claimed-docs] “Describe what you want in plain language. Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “It understands your entire codebase and can work across multiple files and tools to get things done.”
- [claimed-docs] “Claude Code handles the tedious tasks that eat up your day: writing tests for untested code, fixing lint errors across a project, resolving …”
- [github] “helps you code faster by executing routine tasks, explaining complex code, and handling git workflows -- all through natural language comman…”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
- [community] “Whenever the phrase 'simplest fix' appears, it's time to pull the emergency break. This has gotten much worse over the past few weeks. It wi…”
- [community] “I've tried to use Claude code for a month now. It has a 100% failure rate so far. Comparing that to creating a project and just chatting wit…”
- [community] “A month ago the agents researched, designed, and implemented a compelling app idea with minimal guidance and felt super human. A month later…”
- [community] “Claude is significantly better than other models at code assistant tasks, or at least in the way I use it.”
Maintenance automation
developerHave the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
weight 3 · round to Claude CodeCodex CLI/cloud docs describe a general-purpose coding agent that can inspect code, edit files, run local dev tools, automate repeatable work, and review diffs before PRs — capabilities broad enough to plausibly cover writing tests, fixing lint issues, resolving conflicts, and updating dependencies (codex-docs-8, codex-docs-9, codex-docs-30, codex-docs-41). However, none of the docs explicitly name test-writing, lint-fixing, merge-conflict resolution, or dependency updates as supported workflows, and community feedback is mixed on real-world reliability for complex agentic tasks (codex-comm-3, codex-comm-13). missing for 10: explicit documentation/examples of test generation, lint-fix automation, merge-conflict resolution, and dependency-update workflows, plus hands-on confirmation these specific tasks succeed.
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [community] “Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…”
- [community] “Having used codex a fair bit I find it really struggles with … almost anything. However using the equivalent chat gpt model is fantastic.”
First-party docs explicitly list this exact story's capabilities verbatim ('writing tests for untested code, fixing lint errors across a project, resolving merge conflicts, updating dependencies') and Claude Code is broadly documented as an agentic coding assistant that edits files, runs commands, and manages projects end-to-end. Community feedback confirms general coding competence but also raises reliability concerns (e.g., self-verification failures) not specific to these four tasks. Missing for 10: independent hands-on verification specifically for lint-fixing, merge-conflict resolution, and dependency updates rather than general coding tasks.
- [claimed-docs] “writing tests for untested code, fixing lint errors across a project, resolving merge conflicts, updating dependencies, and writing release …”
- [claimed-docs] “Claude Code handles the tedious tasks that eat up your day: writing tests for untested code, fixing lint errors across a project, resolving …”
- [claimed-docs] “Claude Code maps and explains entire codebases in a few seconds. It uses agentic search to understand project structure and dependencies wit…”
- [claimed-docs] “Claude Code integrates with GitHub, GitLab, and your command line tools to handle the entire workflow—reading issues, writing code, running …”
- [community] “Claude is significantly better than other models at code assistant tasks, or at least in the way I use it.”
Multimodal generation
ai-native userGenerate a working app from a sketch, image, or PDF design
weight 2 · round to CodexCodex supports passing images (error screenshots, architecture diagrams, design references) into prompts, which is a partial building block for generating apps from a sketch/image, but there's no evidence of dedicated PDF-to-app workflows, multi-page design ingestion, or documented end-to-end 'sketch/image to working app' generation feature. missing for 10: explicit PDF design ingestion, dedicated image/design-to-app pipeline or template, independent hands-on demonstration of generating a full app from a design artifact.
- [claimed-docs] “Pass an error screenshot, architecture diagram, or design reference with the first prompt, or paste an image into the interactive composer.”
Claude Codenone0/10The evidence pack describes Claude Code's general coding, git, MCP, and automation capabilities but never mentions accepting a sketch, image, or PDF as design input to scaffold or generate an app. The closest reference (claude-code-docs-23) only describes updating an email template from Figma designs shared in Slack, not app generation from visual designs. Missing for 10: any documentation or example of image/PDF/sketch-to-code app generation, multimodal input support in the CLI, or a demonstrated workflow turning a design mockup into a working application.
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding
How deeply the tool maps your repo — cross-file context, architecture awareness, history
Codebase mapping
developerUnderstand how a codebase fits together to find where to start making changes
weight 3 · round to Claude CodeCodex CLI docs explicitly mention exploring unfamiliar code and planning changes within a repository, and it can inspect code, run local dev tools, and review diffs/commits — supporting codebase orientation. However, there's no dedicated codebase-mapping/visualization feature, no evidence of dependency-graph or architecture-summary generation, and community feedback focuses on agentic task execution rather than comprehension aids. Missing for 10: dedicated codebase-map/architecture-overview feature, independent hands-on evidence of effectively onboarding to unfamiliar large codebases, and richer navigation/search tooling beyond terminal chat resume.
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
Docs explicitly claim Claude Code 'maps and explains entire codebases in a few seconds' using agentic search to understand project structure and dependencies without manual context selection, and separately states it 'understands your entire codebase' across files; CLAUDE.md further lets teams encode architecture decisions for onboarding. Missing for 10: independent/hands-on corroboration specifically validating codebase-mapping accuracy, and no benchmark or case study showing it correctly locates the right starting point in a large real-world repo.
- [claimed-docs] “Claude Code maps and explains entire codebases in a few seconds. It uses agentic search to understand project structure and dependencies wit…”
- [claimed-docs] “It understands your entire codebase and can work across multiple files and tools to get things done.”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session.”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session. Use it to set coding standar…”
developerHave the agent map and explain an entire unfamiliar codebase without manually selecting context files
weight 3 · round to Claude CodeCodex CLI docs explicitly state it can be started in a repository 'to explore unfamiliar code, plan a change, edit files, and run your local development tools' (codex-docs-9), implying the agent autonomously navigates the codebase rather than requiring manual file selection, and codex-gh-1 confirms it runs as an autonomous coding agent locally. However, there's no detailed documentation of how it builds a whole-codebase map/summary, no explicit 'explain codebase' feature, and no independent hands-on evidence confirming this works well on large unfamiliar repos. Missing for 10: dedicated codebase-mapping/summarization feature documentation, evidence of handling very large repos, and independent user reports validating this specific capability.
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
Claude Code's own product page explicitly states it 'maps and explains entire codebases in a few seconds' using 'agentic search to understand project structure and dependencies without you having to manually select context files,' directly matching the story, and other docs reinforce that it 'understands your entire codebase' across multiple files. Missing for 10: independent/hands-on evidence specifically corroborating the automatic codebase-mapping claim (community evidence covers general coding quality/trust issues but not this specific feature).
- [claimed-docs] “Claude Code maps and explains entire codebases in a few seconds. It uses agentic search to understand project structure and dependencies wit…”
- [claimed-docs] “It understands your entire codebase and can work across multiple files and tools to get things done.”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session.”
Context management
developerHave the agent build and recall memory automatically across sessions
weight 2 · round to Claude CodeCodex CLI supports `codex resume` to reopen or search past local chats in a repository, giving a limited form of session recall, but this requires manual user action rather than automatic memory building/recall across sessions. Missing for 10: evidence of automatic persistent memory (learned facts, preferences, or context) that Codex builds unprompted and recalls without explicit resume/search commands, and any cross-session synthesis beyond raw chat transcripts.
- [claimed-docs] “Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
- [claimed-docs] “`codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
Claude Code supports persistent project context via CLAUDE.md, which it reads at the start of every session, giving some continuity of 'memory' across sessions, and the VS Code extension keeps conversation history in-editor. However, this is a manually authored/maintained file, not an automatically built or recalled memory system that captures learnings from prior sessions without user intervention. Missing for 10: evidence of automatic memory formation/summarization from past sessions, automatic recall of prior task context without a manually maintained file, and any documentation of a persistent 'agent memory' feature beyond CLAUDE.md.
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session.”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session. Use it to set coding standar…”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
developerInclude multiple project directories in a single session for broader context
weight 2 · round drawnCodexnone0/10No evidence in the pack describes Codex supporting multiple project directories or repositories being combined in a single session/context; documentation focuses on single-repository sessions, cloud tasks, and per-repository setup steps.
Claude Codenone0/10The evidence pack describes Claude Code understanding a single project's entire codebase and working across multiple files within it, but there is no mention of including multiple separate project directories in one session (e.g., an --add-dir style flag or multi-root workspace support).
developerAdd a project instructions file to set coding standards and conventions the agent follows
weight 3 · round to Claude CodeCodexnone0/10The evidence pack covers Codex's CLI, cloud, MCP, and review features but contains no mention of a project-level instructions/config file (e.g., AGENTS.md or similar) for setting coding standards or conventions the agent should follow. This is a plausible and common capability for coding agents, but nothing in the pack documents or demonstrates it.
First-party docs explicitly describe CLAUDE.md as a project-root markdown file read at every session start, used to set coding standards, architecture decisions, preferred libraries, and review checklists (claude-code-docs-5, claude-code-docs-22). Community evidence (claude-code-comm-15) independently confirms real-world use of CLAUDE.md files for guiding the agent, corroborating the feature exists and is actively used. Missing for 10: broader independent/hands-on documentation of best practices or examples beyond a single community mention.
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session.”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session. Use it to set coding standar…”
- [community] “I've found that I have to add more and more CLAUDE.md guide rails, and my CLAUDE.md files have been exploding since around mid-March... I've…”
Issue diagnosis
developerReproduce issues, narrow down root causes, and verify fixes
weight 3 · round to CodexCodex CLI docs explicitly describe exploring unfamiliar code and running local dev tools to investigate issues, passing error screenshots for context, delegating focused investigation to subagents, and running dedicated reviews against uncommitted changes/commits/base branches to verify fixes before committing — covering reproduce, narrow-down, and verify steps. Missing for 10: no explicit 'reproduce a bug' walkthrough or first-hand/independent account of successfully diagnosing and fixing a real bug end-to-end with Codex.
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Pass an error screenshot, architecture diagram, or design reference with the first prompt, or paste an image into the interactive composer.”
- [claimed-docs] “Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.”
- [claimed-docs] “Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
Claude Codedisputedcontradicted5/10Docs claim Claude Code can debug live apps, plan fixes, and 'verifies it works' across multi-file changes (claude-code-docs-2/17/28/51), supporting reproduce/root-cause/verify workflows, but hands-on community reports give a concrete counter-example where self-verification failed (a replace_all bug silently corrupted a constant, 'You absolutely can't trust it to self-verify') and describe recurring low-quality 'simplest fix' patches that break things (claude-code-comm-16, claude-code-comm-17). missing for 10: independent benchmark/case study specifically on bug reproduction and root-cause isolation, and resolution of the self-verification reliability concerns raised by users.
- [claimed-docs] “Describe what you want in plain language. Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “Debug live web applications | Chrome”
- [claimed-docs] “Claude Code plans the approach, writes the code across multiple files, and verifies it works.”
- [claimed-docs] “It understands your entire codebase and can work across multiple files and tools to get things done.”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
- [community] “Whenever the phrase 'simplest fix' appears, it's time to pull the emergency break. This has gotten much worse over the past few weeks. It wi…”
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem
Integrations, plugins, and third-party ecosystem stories
Marketplace
developerEquip the agent with custom skills to perform specialized tasks
weight 1 · round to Claude CodeCodex CLI docs explicitly describe packaging repeatable instructions as "skills" and adding plugins to connect Codex to team tools/data from the CLI, directly matching the custom-skills story. Missing for 10: independent hands-on validation of skill creation/usage, and deeper documentation on skill authoring format/lifecycle beyond a single mention.
- [claimed-docs] “Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.”
Claude Code explicitly supports custom Skills ('Create skills to package repeatable workflows your team can share, like /review-pr or /deploy-staging') plus a scaffolding plugin (mcp-server-dev) for building custom tool integrations, giving developers a documented mechanism to equip the agent with specialized, shareable capabilities. Missing for 10: independent hands-on validation of the skills system's reliability/quality beyond first-party docs.
- [claimed-docs] “Create skills to package repeatable workflows your team can share, like `/review-pr` or `/deploy-staging`.”
- [claimed-docs] “You can also have Claude scaffold a server for you with the official mcp-server-dev plugin”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
engineering-leadIntegrate third-party partner-built agent apps into my workflows
weight 1 · round drawnCodex documents integration points for third-party ecosystem tools — triggering work from GitHub, GitLab, Linear, and Slack (partner platforms), and connecting to third-party MCP servers, plugins, and skills that give access to tools like Figma or a browser — which supports embedding partner-built capabilities into engineering workflows. However, the evidence is framed around Codex consuming tools/data sources rather than a curated marketplace of partner-built 'agent apps,' and there's no independent case study of a partner agent integration working end-to-end. Missing for 10: evidence of a partner/agent-app marketplace or certified third-party agent integrations, and independent verification of such integrations working in practice.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.”
- [claimed-docs] “Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.”
- [claimed-docs] “Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…”
- [claimed-docs] “Use skills and plugins: Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without l…”
- [claimed-docs] “Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation, or t…”
Claude Code supports MCP integration with third-party tools/servers (Notion, Jira, Slack, Google Drive, custom servers) and can be extended via the Agent SDK, plugins, and Slack/GitHub integrations, enabling integration of partner-built apps into workflows. However, there's no explicit evidence of a curated marketplace or formal partner-app ecosystem comparable to a dedicated app store, and integration relies mainly on generic MCP connectors rather than pre-built 'partner agent apps.' Missing for 10: a documented partner/marketplace program for third-party agent apps, independent verification of partner integrations working reliably, and case studies of engineering teams integrating named partner-built agents.
- [claimed-docs] “With MCP, Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom toolin…”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
- [claimed-docs] “claude mcp add --transport http notion https://mcp.notion.com/mcp”
- [claimed-docs] “You can also have Claude scaffold a server for you with the official mcp-server-dev plugin”
- [claimed-docs] “the Agent SDK lets you build your own agents powered by Claude Code's tools and capabilities, with full control over orchestration, tool acc…”
- [claimed-docs] “Route tasks from team chat: mention @Claude in Slack with a bug report and get a pull request back”
Team knowledge
engineering-leadCreate a shared workspace from my docs and repos as a common source of truth for the team
weight 1 · round to Claude CodeCodexnone0/10Codex documents repo-level cloud environments, RBAC, and MCP connections to team tools, but no evidence describes a shared 'workspace' feature that unifies docs and repos into a common source of truth for a team; this is a plausible ask for an engineering tool but Codex's evidence only covers per-task cloud environments and repo configuration, not a persistent shared knowledge/workspace layer.
- [claimed-docs] “Configure the dependencies, tools, variables, and setup steps each repository needs.”
- [claimed-docs] “Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…”
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
CLAUDE.md gives teams a shared, repo-committed markdown file for coding standards, architecture decisions, and review checklists that Claude reads every session, and shareable Skills (e.g. /review-pr, /deploy-staging) let a lead codify team workflows; MCP integrations let Claude also pull in Google Drive docs, Jira tickets, and Slack data as additional context sources. However, this is scattered configuration/context-injection tooling rather than a dedicated 'workspace' or knowledge-base product that unifies docs and repos into one queryable source of truth for the whole team. Missing for 10: a purpose-built shared workspace/knowledge-base UI, cross-repo aggregation, and evidence of team-wide adoption/governance beyond per-repo CLAUDE.md files.
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session.”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session. Use it to set coding standar…”
- [claimed-docs] “Create skills to package repeatable workflows your team can share, like `/review-pr` or `/deploy-staging`.”
- [claimed-docs] “With MCP, Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom toolin…”
- [claimed-docs] “Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom tooling.”
Tool integration
developerConnect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
weight 3 · round to Claude CodeCodex explicitly supports starting work from Slack (and GitHub/GitLab/Linear) and lets users add local or remote MCP servers to connect to third-party tools/docs (e.g. Figma, browser), giving a generic mechanism to extend context to workflow tools. However, there is no explicit documentation of native Jira or Google Drive connectors—only Slack is named among the story's specific tools, with Jira/Google Drive requiring the generic (and for one variant, deprecated/experimental) MCP server pathway. Missing for 10: named Jira integration, named Google Drive integration, and confirmation that the current (non-deprecated) MCP mechanism is broadly used for these specific SaaS tools.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.”
- [claimed-docs] “Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.”
- [claimed-docs] “Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.”
- [claimed-docs] “Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation, or t…”
- [claimed-docs] “Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…”
- [github] “Codex MCP Server Interface [experimental]: a JSON-RPC API that runs over the Model Context Protocol (MCP) transport to control a local Codex…”
- [claimed-docs] “codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …”
Docs explicitly state Claude Code can connect via MCP to Jira, Slack, Google Drive, and other custom tooling, with concrete examples (updating Jira tickets, pulling Slack data, Notion MCP server add command) and multiple transport options. Missing for 10: independent/hands-on third-party confirmation of these specific integrations working in practice beyond vendor docs.
- [claimed-docs] “Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom tooling.”
- [claimed-docs] “With MCP, Claude Code can read your design docs in Google Drive, update tickets in Jira, pull data from Slack, or use your own custom toolin…”
- [claimed-docs] “Update our standard email template based on the new Figma designs that were posted in Slack”
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
- [claimed-docs] “Implement features from issue trackers: "Add the feature described in JIRA issue ENG-4521 and create a PR on GitHub."”
- [claimed-docs] “claude mcp add --transport http notion https://mcp.notion.com/mcp”
- [claimed-docs] “Add the feature described in JIRA issue ENG-4521 and create a PR on GitHub.”
developerKick off agent tasks directly from GitHub, GitLab, Linear, or Slack
weight 2 · round to CodexFirst-party docs explicitly state Codex cloud tasks can be started from GitHub pull requests, GitLab merge requests/issues, Linear issues, or Slack channels/threads, matching the story directly. Missing for 10: independent/hands-on verification of these specific integrations working in practice (community evidence covers CLI/app UX but not the GitHub/GitLab/Linear/Slack kickoff flows specifically).
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
Docs confirm task kickoff from GitHub (@claude mentions, GitHub Code Review, GitHub Actions) and Slack (@Claude mention returns a PR), plus GitLab CI/CD integration, but there is no evidence of Linear integration or a Linear-triggered agent workflow. missing for 10: explicit Linear integration/trigger support, independent/hands-on confirmation of cross-platform task kickoff.
- [github] “Use it in your terminal, IDE, or tag @claude on Github.”
- [claimed-docs] “Route tasks from team chat: mention @Claude in Slack with a bug report and get a pull request back”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
- [claimed-docs] “Claude Code integrates with GitHub, GitLab, and your command line tools to handle the entire workflow—reading issues, writing code, running …”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration
Meeting you in the IDE and terminal — extensions, inline flows, context
Cross device continuity
developerStart a task on one device and continue it later from another device or browser
weight 2 · round drawnCodex supports starting tasks in the cloud from web/GitHub/GitLab/Linear/Slack, working in parallel cloud environments, and later resuming or continuing work from the CLI via 'codex cloud' (browse active/completed chats, submit/apply results) or 'codex resume' to reopen local chats, plus a shared MCP config across ChatGPT desktop, CLI, and IDE extension enabling cross-client continuity. missing for 10: independent hands-on confirmation of seamless state sync across devices/browsers, and no explicit mention of resuming a cloud-started task from a different physical device's browser session.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Start and review work from the web or Codex CLI.”
- [claimed-docs] “Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …”
- [claimed-docs] “`codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
- [claimed-docs] “The ChatGPT desktop app, Codex CLI, and IDE extension share this configuration. Once you configure your MCP servers, you can switch among th…”
Docs explicitly describe cross-device continuity: 'Remote Control' lets you continue work from phone/browser (docs-9), and 'claude --teleport' lets you start a task on web/mobile and pull it into your terminal later (docs-10), backed by browser/cloud session support (docs-13, docs-14, docs-26). missing for 10: independent/hands-on confirmation of teleport and remote-control reliability across devices
- [claimed-docs] “Step away from your desk and keep working from your phone or any browser with Remote Control”
- [claimed-docs] “Kick off a long-running task on the web or the Claude mobile app, then pull it into your terminal with `claude --teleport`.”
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
Ide integration
developerView interactive diffs and share selected code as context from within my JetBrains IDE
weight 1 · round to Claude CodeCodexnone0/10Evidence shows Codex's IDE extension explicitly targets VS Code, Cursor, and Windsurf (codex-gh-2), with no mention of JetBrains IDEs, interactive diff viewing within an IDE, or a 'share selected code as context' feature. The axis (IDE integration) is clearly applicable to Codex as a coding agent, but JetBrains-specific support and the described interactive-diff/context-sharing workflow are simply absent from the evidence pack.
- [github] “If you want Codex in your code editor (VS Code, Cursor, Windsurf), install in your IDE.”
Docs explicitly describe a JetBrains plugin (IntelliJ IDEA, PyCharm, WebStorm, etc.) with interactive diff viewing and selection context sharing, directly matching the story. Missing for 10: independent/hands-on corroboration of the JetBrains plugin specifically (community evidence only covers CLI/terminal experience, not the IDE plugin).
- [claimed-docs] “A plugin for IntelliJ IDEA, PyCharm, WebStorm, and other JetBrains IDEs with interactive diff viewing and selection context sharing.”
developerChat with the coding assistant directly inside my IDE for contextual help
weight 3 · round drawnCodex explicitly offers an IDE extension for VS Code, Cursor, and Windsurf, plus a CLI usable within the terminal in your repo, both providing contextual chat/help with the codebase (edit files, run commands, review diffs). Community evidence confirms real-world usage of Codex CLI/app for editing and testing files in context, though some note UX friction compared to competitors. Missing for 10: deeper first-party documentation/screenshots of the IDE extension's chat UI specifically, and stronger independent hands-on corroboration of in-IDE chat quality.
- [github] “If you want Codex in your code editor (VS Code, Cursor, Windsurf), install in your IDE.”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [community] “Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…”
Official docs confirm dedicated IDE integrations (VS Code extension with inline diffs, @-mentions, plan review, conversation history; JetBrains plugin with diff viewing and selection context sharing), plus terminal-based chat usable from within an IDE, and GitHub explicitly states 'Use it in your terminal, IDE, or tag @claude on Github.' Missing for 10: independent hands-on validation specifically of the IDE chat experience (community evidence is mostly about CLI/terminal use and general quality, not IDE-embedded chat specifically).
- [claimed-docs] “The VS Code extension provides inline diffs, @-mentions, plan review, and conversation history directly in your editor.”
- [claimed-docs] “A plugin for IntelliJ IDEA, PyCharm, WebStorm, and other JetBrains IDEs with interactive diff viewing and selection context sharing.”
- [github] “Use it in your terminal, IDE, or tag @claude on Github.”
- [claimed-docs] “Work with Claude directly in your codebase. Build, debug, and ship from your terminal, IDE, Slack, web, and more.”
Session management
developerReview diffs visually and run multiple sessions side by side in a desktop app
weight 2 · round to Claude CodeCodex ships a desktop app ("codex app"/Codex App page) and documents parallel task execution plus diff/summary inspection before merging, suggesting the underlying pieces exist, but the evidence never shows the desktop app UI actually presenting a visual diff viewer or multiple sessions arranged side by side. Community notes even flag basic desktop-app reliability issues (stuck on 'Loading projects...', Mac-only availability). Missing for 10: concrete documentation/screenshots of the desktop app's diff viewer, explicit multi-session/side-by-side UI description, and independent confirmation it works smoothly.
- [github] “If you want the desktop app experience, run <code>codex app</code> or visit the Codex App page.”
- [github] “If you want the desktop app experience, run <code>codex app</code> or visit <a href="https://chatgpt.com/codex?app-landing-page=true">the Co…”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [community] “Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…”
- [community] “Mac only. Again. Apple is great but this is OpenAI devs showing their disconnect from the mainstream.”
First-party docs explicitly state the capability ('Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions'), closely matching the story, and related IDE integrations (VS Code inline diffs, JetBrains interactive diff viewer) support visual diff review, but this appears to describe a web/desktop companion app rather than a fully detailed, screenshot-documented desktop client, and no independent or hands-on evidence corroborates the side-by-side multi-session desktop UI. Missing for 10: independent/hands-on confirmation of the desktop app's diff viewer and multi-session UI, and richer first-party documentation (screenshots, feature depth) beyond a single summary line.
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “A plugin for IntelliJ IDEA, PyCharm, WebStorm, and other JetBrains IDEs with interactive diff viewing and selection context sharing.”
- [claimed-docs] “The VS Code extension provides inline diffs, @-mentions, plan review, and conversation history directly in your editor.”
- [claimed-docs] “Available for macOS, Linux, and Windows.”
engineering-leadManage multiple agent-driven coding sessions from one unified workspace
weight 2 · round drawnCodex documents cloud parallel task execution across multiple repos/environments (codex-docs-1,3,6), a web/CLI dashboard to browse active and completed chats and apply results locally (codex-docs-15), and resuming/searching across sessions (codex-docs-11,24), which together support managing multiple concurrent agent sessions from a unified interface. However, evidence is vendor-documentation only with no independent hands-on confirmation of a true 'unified workspace' UX for an engineering-lead managing many sessions simultaneously, and some community comments note UX rough edges (codex-comm-9,18). Missing for 10: independent/hands-on verification of multi-session management at scale, and clearer detail on cross-session visibility/coordination for a lead overseeing a team's agents.
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Run tasks in parallel without tying up your local machine.”
- [claimed-docs] “Delegate a longer task and return when it is ready.”
- [claimed-docs] “Browse active and completed chats, submit work to a configured environment, and apply the result to your local repository from the terminal.”
- [claimed-docs] “Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
- [claimed-docs] “`codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.”
- [community] “I wish Codex App was open source. I like it, but there are always a bunch of little paper cuts that, if you were using codex cli, you could …”
Docs describe running multiple sessions side by side, kicking off parallel/cloud sessions from a browser, and spawning multiple coordinated sub-agents under a lead agent, which directly support a lead managing several agent sessions from one workspace (claude-code-docs-13, -14, -26, -6, -34, -44). Missing for 10: independent/hands-on confirmation of the 'unified workspace' UX (no community reports specifically validate multi-session management) and no detail on session-level access control across a team for the lead-agent view.
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “Kick off long-running tasks and check back when they're done, work on repos you don't have locally, or run multiple tasks in parallel.”
- [claimed-docs] “Run Claude Code in your browser with no local setup. Kick off long-running tasks and check back when they're done, work on repos you don't h…”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously. A lead agent coordinates the work, assigns subtasks…”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously.”
- [claimed-docs] “Spawn multiple Claude Code agents that work on different parts of a task simultaneously. A lead agent coordin”
Terminal workflow
developerRun a coding agent locally from my terminal
weight 3 · round drawnCodex CLI is explicitly documented as a coding agent that runs locally in the terminal, with npm/standalone install, working against the local repository, editing files, running commands, and offering interactive TUI plus non-interactive exec mode — well corroborated by first-party docs and GitHub README, with community usage discussion confirming real-world use. Missing for 10: independent hands-on verification specifically of pure local terminal usage (most community commentary discusses model quality/UX rather than the local-run mechanics) and some caveats about performance/limits reported by users.
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
- [github] “npm install -g @openai/codex”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.”
- [claimed-docs] “Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.”
- [claimed-docs] “Install the Codex CLI with the standalone installer for macOS and Linux.”
- [community] “Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…”
- [community] “Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…”
Claude Code is explicitly documented as a terminal-native coding agent: install via curl script, run with cd your-project && claude, available on macOS/Linux/Windows, and GitHub README confirms 'Use it in your terminal, IDE, or tag @claude on Github.' Community posts corroborate hands-on terminal use, noting it's 'implemented as a bash tool and not an editor replacement.' Missing for 10: broader independent benchmark or third-party review confirming consistent reliability of local terminal operation beyond a few anecdotal community posts.
- [claimed-docs] “cd your-project claude”
- [claimed-docs] “curl -fsSL https://claude.ai/install.sh | bash”
- [claimed-docs] “Available for macOS, Linux, and Windows.”
- [github] “Use it in your terminal, IDE, or tag @claude on Github.”
- [community] “The cost is absurd (compared to other LLM providers these days). I asked 3 questions and the cost was ~0.77c. I do like how this is implemen…”
developerRun the agent non-interactively in scripts for workflow automation
weight 2 · round drawnDocs explicitly describe running 'a non-interactive command in a repeatable workflow' and automating repeatable work without leaving the terminal, plus support for submitting work to configured environments from scripts (codex exec-style usage implied). Missing for 10: independent hands-on confirmation of non-interactive/CI usage and detailed exit-code/output-format documentation for scripting.
- [claimed-docs] “Run a non-interactive command in a repeatable workflow.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Browse active and completed chats, submit work to a configured environment, and apply the result to your local repository from the terminal.”
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
Docs explicitly describe non-interactive automation: piping logs, running in CI, scheduling recurring tasks, GitHub Actions/GitLab CI/CD integration for automated code review and issue triage, and headless-style scripting per Unix philosophy. missing for 10: no explicit mention of a documented --print/non-interactive flag or exit-code behavior, and no independent/hands-on report confirming scripted CI usage works as described.
- [claimed-docs] “Claude Code is composable and follows the Unix philosophy. Pipe logs into it, run it in CI, or chain it with other tools”
- [claimed-docs] “Run Claude on a schedule to automate work that repeats: morning PR reviews, overnight CI failure analysis, weekly dependency audits, or sync…”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to Claude CodeCodex ships rich CLI/UI-only capabilities (cloud tasks, resume/review, skills, plugins, MCP client integration) with no evidence these are exposed via a dedicated Codex API, and the general OpenAI API (RBAC, Responses API) is not shown to cover Codex-specific workflows; community evidence even confirms the latest gpt-5.3-codex model 'isn't available on the API yet,' a documented parity gap. Missing for 10: documented API endpoints for cloud task delegation, chat/session resume, MCP tool orchestration, and confirmation that current models/features are API-accessible at parity with CLI/UI.
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [community] “gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'”
- [claimed-docs] “Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…”
- [claimed-docs] “Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…”
- [claimed-docs] “codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …”
Claude Code exposes an Agent SDK for building custom agents with 'full control over orchestration, tool access, and permissions' (docs-18) and supports direct API-key access and CI/headless automation (docs-36, docs-39/40), suggesting core coding capabilities are programmatically accessible. However, evidence doesn't confirm parity for UI-specific features like Remote Control, teleport, mobile app, or Slack routing being fully reachable via the API/SDK. Missing for 10: explicit documentation that all UI-surfaced features (remote control, teleport, IDE-specific interactions) are equally available through the API/SDK, and independent confirmation of this parity.
- [claimed-docs] “the Agent SDK lets you build your own agents powered by Claude Code's tools and capabilities, with full control over orchestration, tool acc…”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
- [claimed-docs] “Use Claude Code as an MCP server. You can use Claude Code itself as an MCP server that other applications can connect to: claude mcp serve (…”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnCodexnone0/10No evidence of any data export feature or open-format export mechanism for chat history, project data, or configurations; Codex works with local files/git repos but there's no documented export/portability capability for user data (e.g., conversation history, settings) to leave the platform. Missing for 10: any documented data export tool, open-format export (JSON/Markdown dump), or data portability statement.
Claude Codenone0/10The evidence pack contains no mention of a data export feature, session/conversation history export, or open-format portability guarantees for Claude Code — nothing addresses a user's ability to extract all their data and leave the platform. While Claude Code operates on local files (inherently open), there is no documented mechanism for exporting session logs, configs, or account data in open formats, so this applicable axis is unsupported.
ai-native userRead the product's source under an open license
weight 2 · round to CodexThe Codex CLI source lives in a public GitHub repo (openai/codex) and a community comment implies its openness lets users 'diagnose and file an issue' the way they can't with the closed-source Codex App, suggesting at least the CLI's code is publicly viewable. However, no evidence pack item states an explicit open-source license, and the App/cloud components are explicitly described as closed. missing for 10: explicit license file/name (MIT, Apache, etc.), confirmation the full product (not just CLI) is open, and independent corroboration beyond one forum remark.
- [github] “Codex CLI is a coding agent from OpenAI that runs locally on your computer.”
- [community] “I wish Codex App was open source. I like it, but there are always a bunch of little paper cuts that, if you were using codex cli, you could …”
Claude Codenone0/10No evidence Claude Code's source is available under an open license; in fact community discussion explicitly contrasts it with an open-source competitor, noting 'Codex CLI is FOSS, unlike Claude Code' — confirming it is closed-source.
- [community] “Codex CLI is FOSS, unlike Claude Code, so Codex is less likely to do things like that, and it's one more reason to avoid Claude Code and Cla…”
ai-native userSelf-host the core product
weight 3 · round drawnCodexnone0/10Codex CLI runs locally but requires signing into a ChatGPT account or OpenAI API key, and the core inference/model and cloud environments are OpenAI-hosted only; there is no self-hosted backend option. A commenter explicitly wishes the Codex App were open source, implying it is not, which forecloses self-hosting the core product.
- [github] “We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [community] “I wish Codex App was open source. I like it, but there are always a bunch of little paper cuts that, if you were using codex cli, you could …”
Claude Codenone0/10Claude Code is a closed-source CLI that requires an Anthropic API key or Claude.ai/Console login to function (docs-37, docs-39, docs-55) — there is no evidence of a self-hostable core model or backend. Community evidence explicitly notes it is not open source, unlike alternatives (comm-4), confirming the product cannot be self-hosted.
- [claimed-docs] “Claude Pro or Max subscription: log in with your Claude.ai account.”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “Individual users can log in with a claude.ai account, while teams can use Claude for Teams or Enterprise, the Claude Console, or a cloud pro…”
- [community] “Codex CLI is FOSS, unlike Claude Code, so Codex is less likely to do things like that, and it's one more reason to avoid Claude Code and Cla…”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Authentication
developerAuthenticate with an API key instead of an account login
weight 2 · round to Claude CodeGitHub docs confirm Codex CLI supports API key authentication as an alternative to ChatGPT account login, but note it 'requires additional setup,' and the account-login flow (Sign in with ChatGPT) is the recommended default. Missing for 10: detailed API-key setup documentation, first-party quickstart parity with account login, and independent confirmation that API-key auth is fully feature-equivalent (e.g. codex-comm-5 shows some newer models aren't even available via API yet).
- [github] “We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [community] “gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'”
Docs explicitly confirm ANTHROPIC_API_KEY env var authentication bypasses the account login prompt, using it for direct API access via X-Api-Key header, as an alternative to Claude.ai account login. missing for 10: independent/hands-on community confirmation of this specific auth flow (only first-party docs cited).
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
- [claimed-docs] “Claude Pro or Max subscription: log in with your Claude.ai account.”
- [claimed-docs] “Individual users can log in with a claude.ai account, while teams can use Claude for Teams or Enterprise, the Claude Console, or a cloud pro…”
engineering-leadAuthenticate through an enterprise identity or cloud platform for compliance and scalability
weight 2 · round to Claude CodeCodex supports signing in with a ChatGPT Business/Enterprise/Edu account (codex-gh-3, codex-gh-7) and OpenAI's platform offers RBAC to scope access at org/project level (codex-docs-28), suggesting enterprise-grade authentication and access control exist. However, there is no explicit documentation of SSO/SAML/OIDC federation with enterprise identity providers (e.g., Okta, Azure AD) specific to Codex, nor details on how ChatGPT Enterprise auth ties into RBAC for Codex usage. Missing for 10: explicit SSO/SAML/OIDC integration docs, enterprise IdP federation details, and independent confirmation of compliance-grade auth flows for Codex specifically.
- [github] “We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.”
- [github] “Run `codex` and select **Sign in with ChatGPT**. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Busi…”
- [claimed-docs] “Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…”
Claude Code documents enterprise authentication via SSO/SAML, domain capture, role-based permissions, compliance API, and managed policy settings under Claude for Enterprise, plus flexible auth options (Console API key, Claude.ai account, Teams/Enterprise, cloud provider) for scaling across org structures. missing for 10: independent/hands-on corroboration of SSO setup working in practice, and no explicit mention of cloud IAM integration (e.g., AWS/GCP native identity federation) beyond 'cloud provider' mention.
- [claimed-docs] “Claude for Enterprise: adds SSO, domain capture, role-based permissions, compliance API, and managed policy settings for organization-wide C…”
- [claimed-docs] “Single sign-on (SSO/SAML) and domain capture”
- [claimed-docs] “Individual users can log in with a claude.ai account, while teams can use Claude for Teams or Enterprise, the Claude Console, or a cloud pro…”
- [claimed-docs] “Claude Pro or Max subscription: log in with your Claude.ai account.”
- [claimed-docs] “You can sign in to your Console account without creating an API key, even when your organization doesn't let developers create them.”
developerSign in with my existing product subscription plan to use the coding agent
weight 2 · round drawnGitHub docs explicitly recommend signing in with ChatGPT to use Codex under existing Plus, Pro, Business, Edu, or Enterprise subscription plans, with API key as an alternative for those without such plans, directly confirming subscription-based sign-in. missing for 10: independent hands-on confirmation of the sign-in flow itself (evidence focuses on capability descriptions rather than a walkthrough).
- [github] “We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.”
- [github] “Run `codex` and select **Sign in with ChatGPT**. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Busi…”
- [github] “Run codex and select Sign in with ChatGPT. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, …”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
Docs explicitly confirm developers can log in with their existing Claude Pro or Max subscription (claude.ai account) instead of needing a separate API key, with API key as an alternative for direct API access. Missing for 10: independent/hands-on confirmation of the subscription login flow working smoothly in practice (community evidence focuses on other topics, not this login flow specifically).
- [claimed-docs] “Claude Pro or Max subscription: log in with your Claude.ai account.”
- [claimed-docs] “Individual users can log in with a claude.ai account, while teams can use Claude for Teams or Enterprise, the Claude Console, or a cloud pro…”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
developerSign in with a personal account to get free-tier access without managing API keys
weight 1 · round to CodexCodex CLI explicitly recommends signing in with a ChatGPT account (Plus/Pro/Business/Edu/Enterprise) to use Codex without an API key, with API key usage noted as an alternative requiring additional setup. This directly matches the story of personal-account sign-in without managing API keys, though the exact free-tier scope/limits aren't detailed. Missing for 10: explicit confirmation of a genuinely free tier (vs. paid ChatGPT plans) and independent corroboration of the login flow's simplicity.
- [github] “We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
- [github] “Run `codex` and select **Sign in with ChatGPT**. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Busi…”
Docs confirm individual developers can log in with a personal claude.ai account (Pro/Max subscription) instead of managing an API key, and that API-key auth is optional/alternate. However, evidence only references Pro/Max subscription login, not an explicit free tier for Claude Code — missing for 10: explicit confirmation that a free/no-cost claude.ai account grants Claude Code access, and independent user corroboration of free-tier login flow.
- [claimed-docs] “Claude Pro or Max subscription: log in with your Claude.ai account.”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
- [claimed-docs] “Individual users can log in with a claude.ai account, while teams can use Claude for Teams or Enterprise, the Claude Console, or a cloud pro…”
Model choice
developerLet the tool automatically pick the best model for each task
weight 1 · round drawnCodexnone0/10Evidence shows Codex lets users manually choose the model and reasoning effort ('Stay in control: Choose the model, reasoning effort, permissions...') rather than any automatic best-model-per-task selection; no docs or community evidence describe an automatic model-routing/selection feature tied to cost or task type.
- [claimed-docs] “Stay in control: Choose the model, reasoning effort, permissions, and commands that fit the task.”
- [community] “The main issue I have with Codex is that the best model is insanely slow, except at nights and weekends when Silicon Valley goes to bed... I…”
- [community] “First thoughts using gpt-5.3-codex-spark in Codex CLI: Blazing fast but it definitely has a small model feel... It has to be prompted to do …”
Claude Codenone0/10No evidence in the pack describes automatic model selection or routing per task; users manually choose models (e.g., Sonnet vs Opus per comm-19) and there's no mention of an auto-select feature. Missing for 10: any docs describing automatic model routing/selection logic based on task complexity or cost.
developerChoose which underlying AI model powers my session from multiple providers
weight 2 · round drawnCodexnone0/10Docs confirm Codex lets users 'Choose the model, reasoning effort, permissions' (codex-docs-31), but this refers to selecting among OpenAI's own Codex/GPT models, not switching between different AI providers (e.g., Anthropic, Google). No evidence shows Codex supports plugging in or selecting non-OpenAI models/providers within a session.
- [claimed-docs] “Stay in control: Choose the model, reasoning effort, permissions, and commands that fit the task.”
- [github] “We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.”
- [github] “You can also use Codex with an API key, but this requires additional setup.”
Claude Codenone0/10Evidence shows Claude Code authentication routes (Claude.ai login, API key, Console, Enterprise SSO) are all tied to Anthropic's own Claude models; there is no mention of selecting GPT, Gemini, or other third-party model providers to power a session. Since comparable coding tools do offer multi-provider model selection, this axis applies but is unevidenced here.
- [claimed-docs] “Claude Pro or Max subscription: log in with your Claude.ai account.”
- [claimed-docs] “If you've set the ANTHROPIC_API_KEY environment variable, Claude Code skips the login prompt and asks you to approve the key instead.”
- [claimed-docs] “ANTHROPIC_API_KEY environment variable. Sent as the X-Api-Key header. Use this for direct Anthropic API access with a key from the Claude Co…”
- [claimed-docs] “Individual users can log in with a claude.ai account, while teams can use Claude for Teams or Enterprise, the Claude Console, or a cloud pro…”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnCodexnone0/10No evidence in the pack mentions data residency, regional storage options, or geographic controls for where Codex data is stored; the pack covers RBAC, MCP, CLI features, and cloud task execution but nothing about choosing a storage region.
Claude Codenone0/10No evidence pack items mention data residency, regional data storage options, or geographic controls for where Claude Code data is processed/stored; only SSO/domain capture/compliance API for enterprise IAM are mentioned. Missing for 10: any documentation of region selection, data residency guarantees, or geo-specific storage controls.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnCodexnone0/10The evidence pack contains no mention of data-training opt-out controls, enterprise data usage policies, or privacy settings for excluding user data from model training; it covers CLI features, MCP, RBAC, and community sentiment but nothing about training-data exclusion.
ai-native userControl data retention and deletion
weight 2 · round drawnCodexnone0/10No evidence pack items address data retention controls, deletion policies, or configurable retention windows for Codex; RBAC docs address access control, not retention/deletion. Missing for 10: any documentation of data retention settings, deletion APIs/workflows, or retention policy configuration.
Claude Codenone0/10The evidence pack shows enterprise features like SSO, domain capture, and a vague 'compliance API' but nothing describing user-controllable data retention settings or deletion of stored conversation/code data. No documentation addresses how users can view, export, or delete retained data.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnCodexnone0/10No evidence in the pack mentions telemetry, usage tracking, data collection settings, or an opt-out mechanism for Codex; the docs and community threads cover features like MCP, CLI usage, and performance but never privacy/telemetry controls.
Claude Codenone0/10The evidence pack contains no documentation or reference to a telemetry/usage-tracking opt-out setting (e.g., no mention of a DISABLE_TELEMETRY flag, privacy settings page, or opt-out toggle) for Claude Code. Community commentary touches on unrelated trust/security concerns (anti-distillation fake tools, undercover mode) but none confirm or deny a telemetry opt-out mechanism.
Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety
Keeping generated changes safe — diffs, approvals, guardrails
Data governance
engineering-leadOpt out of having my code and prompts used for AI model training
weight 1 · round drawnCodexnone0/10No evidence in the pack addresses data usage or training opt-out policies for code/prompts; RBAC and MCP docs are unrelated to this axis. Missing for 10: any enterprise data-usage/training opt-out policy documentation, admin controls for opting out, or third-party confirmation of such a policy.
Claude Codenone0/10The evidence pack contains no documentation or statements about Claude Code's data usage or model-training policies, nor any opt-out mechanism for code/prompt data. Enterprise features mentioned (SSO, compliance API, RBAC) do not address training data usage, and community items are unrelated to this specific concern.
Pr review
developerHave the agent stage changes, write commit messages, create branches, and open pull requests
weight 3 · round to Claude CodeCodex docs explicitly describe inspecting diffs and opening a pull request when cloud work is ready (codex-docs-5), and CLI docs note reviewing changes 'before you commit or open a pull request' (codex-docs-45), implying git workflow integration. However, staging changes, writing commit messages, and creating branches are not explicitly documented as first-class agent actions — they are only implied via general local repo access and command execution (codex-docs-9, codex-docs-30, codex-docs-17). Missing for 10: explicit documentation of commit-message generation, branch creation, and staging as named agent capabilities, plus independent hands-on confirmation of full PR workflow automation.
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.”
- [claimed-docs] “Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.”
First-party docs explicitly state Claude Code 'stages changes, writes commit messages, creates branches, and opens pull requests' and integrates with GitHub/GitLab to handle the entire workflow including submitting PRs, corroborated by the GitHub repo description mentioning it 'handles git workflows'. Missing for 10: independent hands-on verification of a full stage-commit-branch-PR flow (community evidence discusses code quality/trust issues but not this specific git workflow failing).
- [claimed-docs] “Claude Code works directly with git. It stages changes, writes commit messages, creates branches, and opens pull requests.”
- [claimed-docs] “Claude Code integrates with GitHub, GitLab, and your command line tools to handle the entire workflow—reading issues, writing code, running …”
- [github] “helps you code faster by executing routine tasks, explaining complex code, and handling git workflows -- all through natural language comman…”
developerGet automatic code review with contextual feedback on every pull request
weight 3 · round to Claude CodeCodex CLI/cloud ships a dedicated 'review' capability that inspects uncommitted changes, a commit, or a base branch and reports prioritized findings without touching the working tree, and cloud tasks can be kicked off from GitHub PRs and later opened as PRs. However, there is no evidence of an automatic, PR-triggered review bot that comments on every pull request without manual invocation. missing for 10: evidence of automatic triggering on every PR (e.g., GitHub App/webhook auto-review), evidence of inline PR comments, independent confirmation of review quality on real PRs.
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
Docs explicitly advertise 'Get automatic code review on every PR | GitHub Code Review' plus CI-based automated code review/issue triage and enterprise security code review, and CLAUDE.md can encode review checklists; community evidence even notes Claude performs well specifically as a reviewer. missing for 10: independent hands-on validation of the GitHub Code Review integration itself and detail on how contextual feedback is generated/delivered on PRs.
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [claimed-docs] “Claude helps security teams and developers by reviewing code for security issues, drafts patches, and explains the risk in language your who…”
- [claimed-docs] “CLAUDE.md is a markdown file you add to your project root that Claude Code reads at the start of every session. Use it to set coding standar…”
- [community] “I have found that Claude Opus 4.6 is a better reviewer than it is an implementer. When Codex implements and Claude reviews, it's usually jus…”
developerInspect diffs and run checks to catch problems before merging
weight 3 · round to CodexCodex CLI has a dedicated review command that inspects diffs against uncommitted changes, a commit, or a base branch, reporting prioritized findings without modifying the working tree, plus cloud/web flows to inspect summaries and diffs before opening a PR. Missing for 10: independent/hands-on corroboration of the review command's accuracy and any CI-integrated check-running beyond exec/scripts.
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.”
- [claimed-docs] “Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.”
Claude Code supports diff inspection (inline diffs in VS Code/JetBrains, visual diff review in web/desktop UI) and can run tests, lint, and CI checks as part of its workflow, plus automatic PR code review via GitHub integration. However, the story's 'inspect diffs and run checks before merging' as a cohesive reviewer workflow is only partially evidenced — there's no dedicated diff/lint/test-gate UI walkthrough, and community reports raise self-verification concerns (e.g., replace_all bugs going undetected). missing for 10: a dedicated pre-merge review workflow with integrated check-gating (not just individual features), independent hands-on validation of diff-review accuracy, and evidence addressing the self-verification skepticism raised in community reports.
- [claimed-docs] “Review diffs visually, run multiple sessions side by side, schedule recurring tasks, and kick off cloud sessions.”
- [claimed-docs] “The VS Code extension provides inline diffs, @-mentions, plan review, and conversation history directly in your editor.”
- [claimed-docs] “A plugin for IntelliJ IDEA, PyCharm, WebStorm, and other JetBrains IDEs with interactive diff viewing and selection context sharing.”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
- [claimed-docs] “Hooks let you run shell commands before or after Claude Code actions, like auto-formatting after every file edit or running lint before a co…”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
- [community] “I've been using Claude Code daily for months on a project with Elixir, Rust, and Python. The worst failure mode is when it does a replace_al…”
- [community] “I have found that Claude Opus 4.6 is a better reviewer than it is an implementer. When Codex implements and Claude reviews, it's usually jus…”
Safe execution
engineering-leadControl which external tools and integrations the agent is allowed to access
weight 2 · round drawnCodex documents fine-grained control over external tool access at the session/repo level: engineers can add/remove local or remote MCP servers, inspect available tools before they're used, and set permission boundaries for edits/commands via /permissions (codex-docs-16, codex-docs-38, codex-docs-39, codex-docs-46). This gives an engineer meaningful control over which integrations the agent can reach, and RBAC exists for org/project-level API access (codex-docs-28), but that RBAC is about API/dashboard permissions, not specifically about restricting agent tool/integration access org-wide for a lead managing a team's Codex usage. Missing for 10: evidence of centralized, lead-enforced policy that restricts which MCP servers/tools individual developers can enable (vs. per-session self-configuration), and independent confirmation this control actually prevents unauthorized tool access in practice.
- [claimed-docs] “Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.”
- [claimed-docs] “Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…”
- [claimed-docs] “Set the boundaries for each run — /permissions: Choose when Codex can edit files or run commands without asking, and inspect the active sand…”
- [claimed-docs] “In the `codex` TUI, use `/mcp` to see your active MCP servers.”
- [claimed-docs] “Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…”
Claude Code supports MCP server allow-listing via config (claude mcp add), sandboxed Bash tool with filesystem/network domain controls, and Enterprise-tier managed policy settings/SSO/role-based permissions that let an engineering lead govern tool and integration access. However, evidence doesn't show granular per-tool allow/deny lists at a team-policy level outside Enterprise, nor independent confirmation these controls reliably block unauthorized MCP/tool use in practice. missing for 10: fine-grained non-enterprise tool permission controls, independent/hands-on verification that access restrictions are enforced, and centralized audit/reporting of which integrations were actually used.
- [claimed-docs] “Claude Code can connect to hundreds of external tools and data sources through the Model Context Protocol (MCP)”
- [claimed-docs] “claude mcp add --transport http notion https://mcp.notion.com/mcp”
- [claimed-docs] “Stdio servers run as local processes on your machine. They're ideal for tools that need direct system access or custom scripts.”
- [claimed-docs] “Use Claude Code as an MCP server. You can use Claude Code itself as an MCP server that other applications can connect to: claude mcp serve (…”
- [claimed-docs] “Learn how Claude Code's sandboxed Bash tool provides filesystem and network isolation for safer, more autonomous agent execution. The Bash s…”
- [claimed-docs] “Claude for Enterprise: adds SSO, domain capture, role-based permissions, compliance API, and managed policy settings for organization-wide C…”
- [claimed-docs] “Single sign-on (SSO/SAML) and domain capture”
engineering-leadHave the agent operate inside a sandbox when interacting with code, tools, and network resources
weight 2 · round to Claude CodeFirst-party docs explicitly describe sandboxed execution: Codex lets you 'choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots' (codex-docs-17), and cloud tasks run in 'isolated cloud environments' with configurable dependencies/tools (codex-docs-1, codex-docs-4). This directly matches the engineering-lead's need for sandboxed code/tool interaction, though network-resource sandboxing specifics are not spelled out and there's no independent hands-on verification of sandbox robustness (a community comment raises but does not concretely confirm a sandbox-bypass issue). Missing for 10: explicit documentation of network-level sandbox controls, and independent/hands-on confirmation that the sandbox reliably contains tool/network access.
- [claimed-docs] “Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.”
- [claimed-docs] “Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.”
- [claimed-docs] “Configure the dependencies, tools, variables, and setup steps each repository needs.”
- [community] “Do people really want codex to have control over their computer and apps? I'm still paranoid about keeping things securely sandboxed.”
Claude Code documents a dedicated sandboxed Bash tool that enforces filesystem and network isolation via OS-level boundaries, letting the agent run commands autonomously within defined limits rather than requiring per-command approval. missing for 10: independent/hands-on verification of sandbox robustness, and detail on sandboxing coverage for non-Bash tool calls (e.g., MCP tool network access).
- [claimed-docs] “Learn how Claude Code's sandboxed Bash tool provides filesystem and network isolation for safer, more autonomous agent execution. The Bash s…”
Security checks
engineering-leadSee license and public-code matching references for AI-suggested code
weight 1 · round drawnCodexnone0/10No evidence anywhere in the pack mentions license detection, public-code matching, or provenance references for AI-suggested code; Codex's review features (codex-docs-10, -41, -45) only cover code quality/prioritized findings, not license/public-code attribution.
Claude Codenone0/10No evidence anywhere in the pack of license detection, public-code/OSS match references, or provenance attribution for AI-suggested code; Claude Code's documented features focus on code generation, review, MCP integrations, and workflow automation, not license/plagiarism matching.
developerGet contextual explanations and automatic fixes for security vulnerabilities
weight 2 · round to Claude CodeCodex CLI has a dedicated review command that inspects uncommitted changes, commits, or branches and reports 'prioritized findings' (codex-docs-10, codex-docs-41, codex-docs-45), which could surface security issues, and as a general coding agent it can edit files/run commands. However, the review feature explicitly reports findings 'without modifying your working tree,' meaning it does not auto-fix, and no evidence specifically frames this as security-vulnerability detection/explanation with automatic remediation. missing for 10: explicit security-vulnerability scanning/explanation feature, evidence of automatic fix application (vs. just flagging), and any independent confirmation that Codex reliably identifies/fixes security issues.
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…”
- [claimed-docs] “Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…”
- [claimed-docs] “Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.”
Anthropic's enterprise docs explicitly state Claude Code reviews code for security issues, drafts patches, and explains risk in plain language, directly matching the story's contextual-explanation-plus-fix pattern, and this is reinforced by automatic PR code review integration. Missing for 10: independent/hands-on evidence confirming automatic vulnerability fixes work reliably in practice, and more detail on the security-specific workflow beyond a single marketing mention.
- [claimed-docs] “Claude helps security teams and developers by reviewing code for security issues, drafts patches, and explains the risk in language your who…”
- [claimed-docs] “Get automatic code review on every PR | GitHub Code Review”
- [claimed-docs] “In CI, you can automate code review and issue triage with GitHub Actions or GitLab CI/CD.”
Not comparable on these axes
developerReceive inline code completions and next-edit suggestions as I type
weight 3 · not comparableCodexn/aCodex is an agentic coding assistant (CLI, cloud tasks, IDE extension) focused on delegated task completion, code review, and terminal-based editing, not on inline autocomplete-style completions or next-edit suggestions as you type. This story targets IDE-style inline autocomplete tooling, a different axis than Codex's agent-driven workflow model.
Claude Codenone0/10Claude Code's documented interaction model is conversational/agentic (terminal commands, plan-then-execute, PR generation) and its IDE extensions offer inline diffs and @-mentions, not ghost-text style inline completions or next-edit suggestions as the user types. No evidence pack item describes autocomplete-style inline suggestions.
- [claimed-docs] “The VS Code extension provides inline diffs, @-mentions, plan review, and conversation history directly in your editor.”
- [claimed-docs] “A plugin for IntelliJ IDEA, PyCharm, WebStorm, and other JetBrains IDEs with interactive diff viewing and selection context sharing.”
developerDebug a live running web application directly from my coding assistant
weight 1 · not comparableCodexn/aCodex is a coding agent focused on code generation, editing, review, and CLI/cloud task automation; there is no evidence of any capability to attach to or inspect a live running web application (e.g., browser DevTools integration, runtime debugging, log/network inspection of a live app). Debugging a live running app is a different axis (runtime observability/dev-tools) than code editing and static review, which is what this product's evidence covers.
Docs explicitly list a Chrome integration for debugging live web applications, indicating Claude Code can connect to and debug a running app via browser tooling rather than just editing static code. However, evidence is thin — just a single doc title/link with no detail on setup, capabilities (e.g., breakpoints, console/network inspection), or hands-on/community verification of this workflow. missing for 10: detailed documentation of the Chrome debugging workflow, independent/hands-on confirmation it works on real live apps, coverage of non-Chrome runtime debugging scenarios.
- [claimed-docs] “Debug live web applications | Chrome”
- [claimed-docs] “Work with Claude directly in your codebase. Build, debug, and ship from your terminal, IDE, Slack, web, and more.”