Access
Install
npm install -g @anthropic-ai/claude-codeShowcase


Products
Anthropic, product by product →Anthropic ships more than one product — each judged line competes in its own arena on the same stories as everyone else.
| Line | Arena | Rank | PA Score | Agent-ready |
|---|---|---|---|---|
| Claude | AI Assistants | #2/9 | 29/100 | 28/100 |
| Claude Codethis page | AI Coding Agents | #2/13 | 40/100 | 69/100 |
| Claude Design | Design & Prototyping | #8/8 | 18/100 | 16/100 |
| Claude Agent SDK | Agent Frameworks & SDKs | #1/9 | 38/100 | 67/100 |
| Anthropic Skills | Agent Skills & Extensions | #1/5 | 24/100 | 54/100 |
Not yet judged (8 — no arena where they compete): Claude Cowork · Claude in Chrome · @Claude (Slack & Teams) · Claude for Microsoft 365 · Claude Science · Claude Security · Managed Agents · Claude Developer Platform
Try itExperimental
See what an agent can do with Claude Code before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$claude --versionrecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Autonomy agents — stories about autonomy agents in this arenaAutonomy agentsevidence →
Stories about autonomy agents in this arena
Code generation — quality of generated code — correctness, style, fit to the codebaseCode generationevidence →
Quality of generated code — correctness, style, fit to the codebase
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understandingevidence →
How deeply the tool maps your repo — cross-file context, architecture awareness, history
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystemevidence →
Integrations, plugins, and third-party ecosystem stories
Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integrationevidence →
Meeting you in the IDE and terminal — extensions, inline flows, context
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safetyevidence →
Keeping generated changes safe — diffs, approvals, guardrails
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 2 paid · 3 enterprise · 50 not stated in evidence
Follow the green: where the map greys out is where Claude Code stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Machine-readable spec · Versioning policy · Full data export
Subscribe to events via webhooks
~4/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
~4/10
Connect an agent via an official MCP server
✓9/10
Download a machine-readable API spec (OpenAPI or equivalent)
—–
Rely on versioned APIs with a documented deprecation policy
—–
Test against a sandbox environment without touching production data
~5/10
Explore an interactive API reference with runnable examples
—–
Docs for agents
Point an agent at llms.txt or agent-oriented docs
~5/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓8/10
Operate the product with natural-language commands
✓8/10
Plug MCP servers into this product so it can use their tools
✓9/10
Get AI-generated insights and suggestions from my data inside the product
✓7/10
Set up automations that run autonomously in the background
✓8/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Autonomy agents — stories about autonomy agents in this arenaAutonomy agents
Stories about autonomy agents in this arena
Background execution
Parallel agents
Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
~7/10
Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation
Quality of generated code — correctness, style, fit to the codebase
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding
How deeply the tool maps your repo — cross-file context, architecture awareness, history
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem
Integrations, plugins, and third-party ecosystem stories
Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration
Meeting you in the IDE and terminal — extensions, inline flows, context
Start a task on one device and continue it later from another device or browser
✓8/10
Ide integration
Session management
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety
Keeping generated changes safe — diffs, approvals, guardrails
Sorted by importance (agentic first) (high → low) · 74/74 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full± | 9/10 | Cclaimed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Cclaimed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full± | 8/10 | Xcommunity | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full± | 8/10 | Cclaimed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full± | 8/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial±enterprise | 4/10 | Cclaimed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 5/10 | Cclaimed | |
Add a project instructions file to set coding standards and conventions the agent follows C Context management | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | full | 9/10 | Xcommunity | |
Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context C Tool integration | developer | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 3 | full | 9/10 | Cclaimed | |
Run a coding agent locally from my terminal C Terminal workflow | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 3 | full | 9/10 | Xcommunity | |
Chat with the coding assistant directly inside my IDE for contextual help C Ide integration | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 3 | full | 8/10 | Cclaimed | |
Have the agent stage changes, write commit messages, create branches, and open pull requests C Pr review | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 3 | full | 8/10 | Cclaimed | |
Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me C Maintenance automation | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | full | 8/10 | Xcommunity | |
Delegate longer-running coding tasks to run in the background in an isolated cloud environment C Background execution | developer | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 3 | full | 7/10 | Cclaimed | |
Get automatic code review with contextual feedback on every pull request C Pr review | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 3 | full | 7/10 | Xcommunity | |
Have the agent map and explain an entire unfamiliar codebase without manually selecting context files C Codebase mapping | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | full | 7/10 | Cclaimed | |
Turn a tracked issue into a complete pull request end-to-end C Feature implementation | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | full | 7/10 | Xcommunity | |
Understand how a codebase fits together to find where to start making changes C Codebase mapping | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | full | 7/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 6/10 | Cclaimed | |
Describe a feature or bug in plain language and have the agent implement or fix it across multiple files C Feature implementation | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | disputed | 6/10 | Dcontradicted | |
Inspect diffs and run checks to catch problems before merging C Pr review | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 3 | partial | 6/10 | Xcommunity | |
Reproduce issues, narrow down root causes, and verify fixes C Issue diagnosis | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 3 | disputed | 5/10 | Dcontradicted | |
Receive inline code completions and next-edit suggestions as I type C Code completion | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Authenticate with an API key instead of an account login G Authentication | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | full | 9/10 | Cclaimed | |
Sign in with my existing product subscription plan to use the coding agent C Authentication | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | fullpaid | 9/10 | Cclaimed | |
Authenticate through an enterprise identity or cloud platform for compliance and scalability G Authentication | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | fullenterprise | 8/10 | Cclaimed | |
Have the agent operate inside a sandbox when interacting with code, tools, and network resources C Safe execution | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 2 | full | 8/10 | Cclaimed | |
Run the agent non-interactively in scripts for workflow automation C Terminal workflow | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | full | 8/10 | Cclaimed | |
Start a task on one device and continue it later from another device or browser C Cross device continuity | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | full | 8/10 | Cclaimed | |
Debug issues and troubleshoot using natural-language queries C Debugging | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 2 | full | 7/10 | Xcommunity | |
Get contextual explanations and automatic fixes for security vulnerabilities C Security checks | developer | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 2 | full | 7/10 | Cclaimed | |
Have a cloud agent build, test, and demo a feature end-to-end for my review C Background execution | ai-native user | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | partial | 7/10 | Xcommunity | |
Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack C Tool integration | developer | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 2 | partial | 7/10 | Cclaimed | |
Launch fleets of autonomous agents that work in parallel on different tasks for hours or days C Parallel agents | ai-native user | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | full | 7/10 | Cclaimed | |
Manage multiple agent-driven coding sessions from one unified workspace C Session management | engineering-lead | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | full | 7/10 | Cclaimed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | full | 7/10 | Cclaimed | |
Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously C Scheduled automation | ai-native user | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | partial | 7/10 | Xcommunity | |
Control which external tools and integrations the agent is allowed to access C Safe execution | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 2 | partialenterprise | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | disputed | 6/10 | Dcontradicted | |
Review diffs visually and run multiple sessions side by side in a desktop app C Session management | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 2 | partial | 6/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Cclaimed | |
Have the agent build and recall memory automatically across sessions C Context management | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 2 | partial | 4/10 | Cclaimed | |
Configure a reproducible cloud environment with the dependencies and setup steps my repository needs C Background execution | developer | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 2 | partial | 3/10 | Cclaimed | |
Choose which underlying AI model powers my session from multiple providers C Model choice | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Generate a working app from a sketch, image, or PDF design C Multimodal generation | ai-native user | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 2 | none | untested | none yet | |
Include multiple project directories in a single session for broader context C Context management | developer | Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Equip the agent with custom skills to perform specialized tasks C Marketplace | developer | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 1 | full | 8/10 | Cclaimed | |
View interactive diffs and share selected code as context from within my JetBrains IDE C Ide integration | developer | Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration | 1 | full | 8/10 | Cclaimed | |
Debug a live running web application directly from my coding assistant C Debugging | developer | Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation | 1 | partial | 6/10 | Cclaimed | |
Integrate third-party partner-built agent apps into my workflows C Marketplace | engineering-lead | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 1 | partial | 6/10 | Cclaimed | |
Run several task attempts in parallel and compare results before choosing one C Parallel agents | developer | Autonomy agents — stories about autonomy agents in this arenaAutonomy agents | 1 | partial | 6/10 | Cclaimed | |
Sign in with a personal account to get free-tier access without managing API keys G Authentication | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | partialpaid | 6/10 | Cclaimed | |
Create a shared workspace from my docs and repos as a common source of truth for the team C Team knowledge | engineering-lead | Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem | 1 | partial | 5/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Cclaimed | |
Let the tool automatically pick the best model for each task C Model choice | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | none | untested | none yet | |
Opt out of having my code and prompts used for AI model training C Data governance | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 1 | none | untested | none yet | |
See license and public-code matching references for AI-suggested code C Security checks | engineering-lead | Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 37 stories with headroom
What would move Claude Code’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Code generation — quality of generated code — correctness, style, fit to the codebaseReceive inline code completions and next-edit suggestions as I type
nonemoves PA Scoreimpact 30
Claude Code's documented interaction model is conversational/agentic (terminal commands, plan-then-execute, PR generation) and its IDE extensions offer inline diffs and @-mentions, not ghost-text style inline completions or next-edit suggestions as the user types.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
The evidence pack contains no mention of a data export feature, session/conversation history export, or open-format portability guarantees for Claude Code — nothing addresses a user's ability to extract all their data and leave the platform.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Claude Code is a closed-source CLI that requires an Anthropic API key or Claude.ai/Console login to function (docs-37, docs-39, docs-55) — there is no evidence of a self-hostable core model or backend.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
The evidence pack includes enterprise/compliance features (SSO, compliance API, managed policies) but contains no mention of any training-data opt-out, data-usage policy, or explicit statement that user code/conversations are excluded from model training.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
The evidence pack shows standard documentation pages and an Agent SDK reference, but nothing describing an interactive API reference with runnable/executable code examples (e.g., an in-browser sandbox or live API explorer).
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
The evidence pack shows Claude Code as a CLI/agent tool with SDK, MCP, and CI integrations, but no mention of a downloadable OpenAPI or equivalent machine-readable API spec for Claude Code itself.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
No evidence pack items mention API versioning schemes, version numbers, or a documented deprecation policy for Claude Code's APIs/CLI/SDK; the pack covers features, integrations, and community sentiment but nothing about API stability or deprecation commitments.
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyInclude multiple project directories in a single session for broader context
nonemoves PA Scoreimpact 20
The evidence pack describes Claude Code understanding a single project's entire codebase and working across multiple files within it, but there is no mention of including multiple separate project directories in one session (e.g., an --add-dir style flag or multi-root workspace support).
Showing the top 8 of 37 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 57 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
En docs47 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Plug MCP servers into this product so it can use their tools
- Use an official CLI
- Drive the product through a documented public API
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Version, review, and roll back my automations
- Have a cloud agent build, test, and demo a feature end-to-end for my review
- Delegate longer-running coding tasks to run in the background in an isolated cloud environment
- Configure a reproducible cloud environment with the dependencies and setup steps my repository needs
- Launch fleets of autonomous agents that work in parallel on different tasks for hours or days
- Run several task attempts in parallel and compare results before choosing one
- Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
- Debug a live running web application directly from my coding assistant
- Debug issues and troubleshoot using natural-language queries
- Turn a tracked issue into a complete pull request end-to-end
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
- Understand how a codebase fits together to find where to start making changes
- Have the agent map and explain an entire unfamiliar codebase without manually selecting context files
- Have the agent build and recall memory automatically across sessions
- Add a project instructions file to set coding standards and conventions the agent follows
- Reproduce issues, narrow down root causes, and verify fixes
- Equip the agent with custom skills to perform specialized tasks
- Integrate third-party partner-built agent apps into my workflows
- Create a shared workspace from my docs and repos as a common source of truth for the team
- Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
- Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack
- Start a task on one device and continue it later from another device or browser
- View interactive diffs and share selected code as context from within my JetBrains IDE
- Chat with the coding assistant directly inside my IDE for contextual help
- Review diffs visually and run multiple sessions side by side in a desktop app
- Manage multiple agent-driven coding sessions from one unified workspace
- Run a coding agent locally from my terminal
- Run the agent non-interactively in scripts for workflow automation
- Do everything through the API that I can do in the UI
- Have the agent stage changes, write commit messages, create branches, and open pull requests
- Get automatic code review with contextual feedback on every pull request
- Inspect diffs and run checks to catch problems before merging
- Get contextual explanations and automatic fixes for security vulnerabilities
docs24 stories
- Point an agent at llms.txt or agent-oriented docs
- Plug MCP servers into this product so it can use their tools
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Subscribe to events via webhooks
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Test against a sandbox environment without touching production data
- Define rules that trigger actions automatically on events
- Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
- Turn a tracked issue into a complete pull request end-to-end
- Equip the agent with custom skills to perform specialized tasks
- Integrate third-party partner-built agent apps into my workflows
- Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
- Do everything through the API that I can do in the UI
- Authenticate with an API key instead of an account login
- Authenticate through an enterprise identity or cloud platform for compliance and scalability
- Sign in with my existing product subscription plan to use the coding agent
- Sign in with a personal account to get free-tier access without managing API keys
- Control which external tools and integrations the agent is allowed to access
- Have the agent operate inside a sandbox when interacting with code, tools, and network resources
Product docs14 stories
- Point an agent at llms.txt or agent-oriented docs
- Use an official CLI
- Get AI-generated insights and suggestions from my data inside the product
- Debug a live running web application directly from my coding assistant
- Debug issues and troubleshoot using natural-language queries
- Turn a tracked issue into a complete pull request end-to-end
- Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
- Understand how a codebase fits together to find where to start making changes
- Have the agent map and explain an entire unfamiliar codebase without manually selecting context files
- Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack
- Chat with the coding assistant directly inside my IDE for contextual help
- Review diffs visually and run multiple sessions side by side in a desktop app
- Run a coding agent locally from my terminal
- Have the agent stage changes, write commit messages, create branches, and open pull requests
Hacker News13 stories
- Delegate tasks to a built-in AI assistant inside the product
- Perform bulk operations across many items at once
- Have a cloud agent build, test, and demo a feature end-to-end for my review
- Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
- Debug issues and troubleshoot using natural-language queries
- Turn a tracked issue into a complete pull request end-to-end
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
- Add a project instructions file to set coding standards and conventions the agent follows
- Reproduce issues, narrow down root causes, and verify fixes
- Run a coding agent locally from my terminal
- Get automatic code review with contextual feedback on every pull request
- Inspect diffs and run checks to catch problems before merging
GitHub README9 stories
- Use an official CLI
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Debug issues and troubleshoot using natural-language queries
- Describe a feature or bug in plain language and have the agent implement or fix it across multiple files
- Kick off agent tasks directly from GitHub, GitLab, Linear, or Slack
- Chat with the coding assistant directly inside my IDE for contextual help
- Run a coding agent locally from my terminal
- Have the agent stage changes, write commit messages, create branches, and open pull requests
Solutions docs6 stories
- Issue scoped/least-privilege API credentials for an agent
- Get AI-generated insights and suggestions from my data inside the product
- Authenticate through an enterprise identity or cloud platform for compliance and scalability
- Get automatic code review with contextual feedback on every pull request
- Control which external tools and integrations the agent is allowed to access
- Get contextual explanations and automatic fixes for security vulnerabilities
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$claude --versionreproduced$ claude --version 2.1.259 (Claude Code)
$echo '<jsonrpc initialize>' | claude mcp servereproduced$ echo '<jsonrpc initialize>' | claude mcp serve
{"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{}},"serverInfo":{"name":"claude/tengu","version":"2.1.259"}},"jsonrpc":"2.0","id":1}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
5 of 24 testable claims verified · 1 contradicted → integrity 13/100
23 distinct capability claims found in Claude Code’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
5
Verified
18
Unverified
1
Contradicted
31
Undersold
Verified (5)
“Writes tests, fixes lint errors, resolves merge conflicts, updates dependencies, and writes release notes”
Have the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for mefullproof ↗
“Reads a CLAUDE.md project file at session start to apply coding standards, architecture decisions, and review checklists”
Add a project instructions file to set coding standards and conventions the agent followsfullproof ↗
“Can run on a schedule to automate recurring work like PR reviews, CI failure analysis, dependency audits, and doc syncing”
Set up always-on agents that run on schedules or triggers to maintain and fix my software autonomouslypartialproof ↗
“Provides automatic code review on every GitHub pull request”
Get automatic code review with contextual feedback on every pull requestfullproof ↗
“Integrates with GitHub, GitLab, and CLI tools to handle a full workflow: reading issues, writing code, running tests, and submitting PRs”
Turn a tracked issue into a complete pull request end-to-endfullproof ↗
Unverified (21)
“Works directly with git: stages changes, writes commit messages, creates branches, and opens pull requests”
Have the agent stage changes, write commit messages, create branches, and open pull requestsfullproof ↗
“Via MCP, can read design docs in Google Drive, update Jira tickets, pull data from Slack, or use custom tooling”
Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its contextfullproof ↗
“Via MCP, can read design docs in Google Drive, update Jira tickets, pull data from Slack, or use custom tooling”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Can spawn multiple agents to work on different parts of a task at once, with a lead agent coordinating and merging results”
Launch fleets of autonomous agents that work in parallel on different tasks for hours or daysfullproof ↗
“Composable Unix-style CLI: logs can be piped in, run in CI, or chained with other tools”
Run the product headlessly / in CI for automationfullproof ↗
“Composable Unix-style CLI: logs can be piped in, run in CI, or chained with other tools”
Run the agent non-interactively in scripts for workflow automationfullproof ↗
“Can run on a schedule to automate recurring work like PR reviews, CI failure analysis, dependency audits, and doc syncing”
“Remote Control lets you keep working on tasks from your phone or any browser”
Start a task on one device and continue it later from another device or browserfullproof ↗
“Can start a long-running task on the web or mobile app, then pull it into your terminal with a teleport command”
Start a task on one device and continue it later from another device or browserfullproof ↗
“Can be triggered from Slack by mentioning @Claude with a bug report to get back a pull request”
Kick off agent tasks directly from GitHub, GitLab, Linear, or Slackpartialproof ↗
“Desktop app lets you review diffs visually and run multiple sessions side by side”
Review diffs visually and run multiple sessions side by side in a desktop apppartialproof ↗
“Can run long tasks in the cloud with no local setup, work on repos not checked out locally, and run multiple tasks in parallel”
Delegate longer-running coding tasks to run in the background in an isolated cloud environmentfullproof ↗
“Maps and explains entire codebases in seconds using agentic search, without manual context file selection”
Have the agent map and explain an entire unfamiliar codebase without manually selecting context filesfullproof ↗
“Can debug live running web applications via a Chrome integration”
Debug a live running web application directly from my coding assistantpartialproof ↗
“Agent SDK lets developers build their own agents using Claude Code's tools, with control over orchestration, tool access, and permissions”
“Agent SDK lets developers build their own agents using Claude Code's tools, with control over orchestration, tool access, and permissions”
Control which external tools and integrations the agent is allowed to accesspartialproof ↗
“JetBrains plugin for IntelliJ IDEA, PyCharm, WebStorm and others with interactive diff viewing and selection context sharing”
View interactive diffs and share selected code as context from within my JetBrains IDEfullproof ↗
“VS Code extension provides inline diffs, @-mentions, plan review, and conversation history in the editor”
Chat with the coding assistant directly inside my IDE for contextual helpfullproof ↗
“Can update an email template based on new Figma designs posted in Slack, combining design and chat tool context”
Connect the agent to workflow tools like Jira, Slack, and Google Drive to extend its contextfullproof ↗
“Kick off long-running tasks and check back later, work on repos not local, or run multiple tasks in parallel”
Delegate longer-running coding tasks to run in the background in an isolated cloud environmentfullproof ↗
“Kick off long-running tasks and check back later, work on repos not local, or run multiple tasks in parallel”
Run several task attempts in parallel and compare results before choosing onepartialproof ↗
Contradicted (1)
“Takes a plain-language description, plans the approach, and writes code across multiple files, then verifies it works”
Describe a feature or bug in plain language and have the agent implement or fix it across multiple filesdisputedproof ↗
Undersold (31)
Point an agent at llms.txt or agent-oriented docspartialproof ↗
Drive the product through a documented public APIfullproof ↗
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productfullproof ↗
Set up automations that run autonomously in the backgroundfullproof ↗
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
Operate the product with natural-language commandsfullproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Have a cloud agent build, test, and demo a feature end-to-end for my reviewpartialproof ↗
Configure a reproducible cloud environment with the dependencies and setup steps my repository needspartialproof ↗
Debug issues and troubleshoot using natural-language queriesfullproof ↗
Understand how a codebase fits together to find where to start making changesfullproof ↗
Have the agent build and recall memory automatically across sessionspartialproof ↗
Equip the agent with custom skills to perform specialized tasksfullproof ↗
Integrate third-party partner-built agent apps into my workflowspartialproof ↗
Create a shared workspace from my docs and repos as a common source of truth for the teampartialproof ↗
Manage multiple agent-driven coding sessions from one unified workspacefullproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Authenticate with an API key instead of an account loginfullproof ↗
Authenticate through an enterprise identity or cloud platform for compliance and scalabilityfullproof ↗
Sign in with my existing product subscription plan to use the coding agentfullproof ↗
Sign in with a personal account to get free-tier access without managing API keyspartialproof ↗
Inspect diffs and run checks to catch problems before mergingpartialproof ↗
Have the agent operate inside a sandbox when interacting with code, tools, and network resourcesfullproof ↗
Get contextual explanations and automatic fixes for security vulnerabilitiesfullproof ↗
Claims outside our story set (1)
Real capability claims found in Claude Code’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Can query a PostgreSQL database to find user emails matching a feature usage criterion”
source ↗
Business model
Requires a paid Claude Pro/Max/Team/Enterprise plan or Console account; API usage billed per token, with a small free credit for new accounts.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
