Skip to content

Rank #2 of 13 in AI Coding Agents

Claude Code logo

Anthropic · commercial

144.7k93.5k/yr +853

Showcase

Claude Code homepage screenshot
homepage · captured Sep 2026 · view live ↗
Claude Code docs screenshot
docs · captured Sep 2026 · view live ↗

Anthropic ships more than one product — each judged line competes in its own arena on the same stories as everyone else.

Not yet judged (8 — no arena where they compete): Claude Cowork · Claude in Chrome · @Claude (Slack & Teams) · Claude for Microsoft 365 · Claude Science · Claude Security · Managed Agents · Claude Developer Platform

Try itExperimental

See what an agent can do with Claude Code before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$claude --versionrecorded session — replayed, not live
recorded 2026-09-03 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

59.1/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

39.2/100

Autonomy agents — stories about autonomy agents in this arenaAutonomy agentsevidence →

Stories about autonomy agents in this arena

49.2/100

Code generation — quality of generated code — correctness, style, fit to the codebaseCode generationevidence →

Quality of generated code — correctness, style, fit to the codebase

40.0/100

Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understandingevidence →

How deeply the tool maps your repo — cross-file context, architecture awareness, history

48.9/100

Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystemevidence →

Integrations, plugins, and third-party ecosystem stories

62.5/100

Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integrationevidence →

Meeting you in the IDE and terminal — extensions, inline flows, context

74.8/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

6.0/100

Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →

Free-tier ceilings, usage caps, and rate limits before you have to pay

55.6/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safetyevidence →

Keeping generated changes safe — diffs, approvals, guardrails

54.7/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 0 free · 2 paid · 3 enterprise · 50 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 74/74 stories · click a row’s chevron for the rationale and evidence

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full±9/10C

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10C

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full±8/10X

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full±8/10C

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10C

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full±8/10C

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10T

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial±enterprise4/10C

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10C

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial5/10C

Add a project instructions file to set coding standards and conventions the agent follows C

Context management

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3full9/10X

Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context C

Tool integration

developerEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem3full9/10C

Run a coding agent locally from my terminal C

Terminal workflow

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration3full9/10X

Chat with the coding assistant directly inside my IDE for contextual help C

Ide integration

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration3full8/10C

Have the agent stage changes, write commit messages, create branches, and open pull requests C

Pr review

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety3full8/10C

Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me C

Maintenance automation

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3full8/10X

Delegate longer-running coding tasks to run in the background in an isolated cloud environment C

Background execution

developerAutonomy agents — stories about autonomy agents in this arenaAutonomy agents3full7/10C

Get automatic code review with contextual feedback on every pull request C

Pr review

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety3full7/10X

Have the agent map and explain an entire unfamiliar codebase without manually selecting context files C

Codebase mapping

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3full7/10C

Turn a tracked issue into a complete pull request end-to-end C

Feature implementation

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3full7/10X

Understand how a codebase fits together to find where to start making changes C

Codebase mapping

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3full7/10C

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial6/10C

Describe a feature or bug in plain language and have the agent implement or fix it across multiple files C

Feature implementation

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3disputed6/10D

Inspect diffs and run checks to catch problems before merging C

Pr review

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety3partial6/10X

Reproduce issues, narrow down root causes, and verify fixes C

Issue diagnosis

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3disputed5/10D

Receive inline code completions and next-edit suggestions as I type C

Code completion

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation3none0/10

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3noneuntestednone yet

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

Authenticate with an API key instead of an account login G

Authentication

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2full9/10C

Sign in with my existing product subscription plan to use the coding agent C

Authentication

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2fullpaid9/10C

Authenticate through an enterprise identity or cloud platform for compliance and scalability G

Authentication

engineering-leadPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2fullenterprise8/10C

Have the agent operate inside a sandbox when interacting with code, tools, and network resources C

Safe execution

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety2full8/10C

Run the agent non-interactively in scripts for workflow automation C

Terminal workflow

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2full8/10C

Start a task on one device and continue it later from another device or browser C

Cross device continuity

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2full8/10C

Debug issues and troubleshoot using natural-language queries C

Debugging

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation2full7/10X

Get contextual explanations and automatic fixes for security vulnerabilities C

Security checks

developerReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety2full7/10C

Have a cloud agent build, test, and demo a feature end-to-end for my review C

Background execution

ai-native userAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2partial7/10X

Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack C

Tool integration

developerEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem2partial7/10C

Launch fleets of autonomous agents that work in parallel on different tasks for hours or days C

Parallel agents

ai-native userAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2full7/10C

Manage multiple agent-driven coding sessions from one unified workspace C

Session management

engineering-leadIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2full7/10C

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2full7/10C

Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously C

Scheduled automation

ai-native userAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2partial7/10X

Control which external tools and integrations the agent is allowed to access C

Safe execution

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety2partialenterprise6/10C

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2disputed6/10D

Review diffs visually and run multiple sessions side by side in a desktop app C

Session management

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration2partial6/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial5/10C

Have the agent build and recall memory automatically across sessions C

Context management

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding2partial4/10C

Configure a reproducible cloud environment with the dependencies and setup steps my repository needs C

Background execution

developerAutonomy agents — stories about autonomy agents in this arenaAutonomy agents2partial3/10C

Choose which underlying AI model powers my session from multiple providers C

Model choice

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2none0/10

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Generate a working app from a sketch, image, or PDF design C

Multimodal generation

ai-native userCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation2noneuntestednone yet

Include multiple project directories in a single session for broader context C

Context management

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Equip the agent with custom skills to perform specialized tasks C

Marketplace

developerEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem1full8/10C

View interactive diffs and share selected code as context from within my JetBrains IDE C

Ide integration

developerIde terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration1full8/10C

Debug a live running web application directly from my coding assistant C

Debugging

developerCode generation — quality of generated code — correctness, style, fit to the codebaseCode generation1partial6/10C

Integrate third-party partner-built agent apps into my workflows C

Marketplace

engineering-leadEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem1partial6/10C

Run several task attempts in parallel and compare results before choosing one C

Parallel agents

developerAutonomy agents — stories about autonomy agents in this arenaAutonomy agents1partial6/10C

Sign in with a personal account to get free-tier access without managing API keys G

Authentication

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits1partialpaid6/10C

Create a shared workspace from my docs and repos as a common source of truth for the team C

Team knowledge

engineering-leadEcosystem — integrations, plugins, and third-party ecosystem storiesEcosystem1partial5/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial5/10C

Let the tool automatically pick the best model for each task C

Model choice

developerPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits1noneuntestednone yet

Opt out of having my code and prompts used for AI model training C

Data governance

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety1noneuntestednone yet

See license and public-code matching references for AI-suggested code C

Security checks

engineering-leadReview safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety1noneuntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 37 stories with headroom

What would move Claude Code’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Code generation — quality of generated code — correctness, style, fit to the codebaseReceive inline code completions and next-edit suggestions as I type

    nonemoves PA Scoreimpact 30

    Claude Code's documented interaction model is conversational/agentic (terminal commands, plan-then-execute, PR generation) and its IDE extensions offer inline diffs and @-mentions, not ghost-text style inline completions or next-edit suggestions as the user types.

  2. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    The evidence pack contains no mention of a data export feature, session/conversation history export, or open-format portability guarantees for Claude Code — nothing addresses a user's ability to extract all their data and leave the platform.

  3. Openness — open source, data portability, and self-hosting storiesSelf-host the core product

    nonemoves PA Scoreimpact 30

    Claude Code is a closed-source CLI that requires an Anthropic API key or Claude.ai/Console login to function (docs-37, docs-39, docs-55) — there is no evidence of a self-hostable core model or backend.

  4. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    The evidence pack includes enterprise/compliance features (SSO, compliance API, managed policies) but contains no mention of any training-data opt-out, data-usage policy, or explicit statement that user code/conversations are excluded from model training.

  5. Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples

    nonemoves API qualityimpact 30

    The evidence pack shows standard documentation pages and an Agent SDK reference, but nothing describing an interactive API reference with runnable/executable code examples (e.g., an in-browser sandbox or live API explorer).

  6. Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)

    nonemoves API qualityimpact 30

    The evidence pack shows Claude Code as a CLI/agent tool with SDK, MCP, and CI integrations, but no mention of a downloadable OpenAPI or equivalent machine-readable API spec for Claude Code itself.

  7. Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy

    nonemoves API qualityimpact 30

    No evidence pack items mention API versioning schemes, version numbers, or a documented deprecation policy for Claude Code's APIs/CLI/SDK; the pack covers features, integrations, and community sentiment but nothing about API stability or deprecation commitments.

  8. Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyInclude multiple project directories in a single session for broader context

    nonemoves PA Scoreimpact 20

    The evidence pack describes Claude Code understanding a single project's entire codebase and working across multiple files within it, but there is no mention of including multiple separate project directories in one session (e.g., an --add-dir style flag or multi-root workspace support).

Showing the top 8 of 37 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 57 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

En docs47 stories

docs24 stories

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$claude --versionreproduced
$ claude --version
2.1.259 (Claude Code)
proves: Use an official CLIrecorded 2026-09-03
$echo '<jsonrpc initialize>' | claude mcp servereproduced
$ echo '<jsonrpc initialize>' | claude mcp serve
{"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{}},"serverInfo":{"name":"claude/tengu","version":"2.1.259"}},"jsonrpc":"2.0","id":1}

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

5 of 24 testable claims verified · 1 contradictedintegrity 13/100

23 distinct capability claims found in Claude Code’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

5

Verified

18

Unverified

1

Contradicted

31

Undersold

Verified (5)
Unverified (21)
Contradicted (1)
Undersold (31)
Claims outside our story set (1)

Real capability claims found in Claude Code’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Can query a PostgreSQL database to find user emails matching a feature usage criterion

    source ↗
Suggest a story for these →

Business model

subscription-per-seatusage-basedcredits

Requires a paid Claude Pro/Max/Team/Enterprise plan or Console account; API usage billed per token, with a small free credit for new accounts.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score30 (Sep 1 '26)40 (Sep 16 '26)
Agent-ready28 (Aug 28 '26)69 (Sep 16 '26)

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)