AI Code Review arenaAI Code Review
AI code review agents — the bots that review every pull request: summarizing changes, catching real bugs with full-codebase context, enforcing team standards, and increasingly holding the line on the flood of agent-authored code no human team could review alone — judged on review accuracy and noise discipline, codebase understanding and team memory, native GitHub/GitLab integration, in-repo configuration and custom rules, PR chat, agentic autofixes and pre-merge checks, merge gating and analytics, and IDE/CLI review surfaces. The 2026 market moved fast and the arena models it honestly: CodeRabbit raised a $143M Series C at a $1.5B valuation and rebranded around agentic change management; Graphite retired its Diamond branding into Graphite AI reviews; Qodo handed the open-source PR-Agent to a community org (The-PR-Agent/pr-agent now states it is "not the Qodo free tier") and sells credit-metered Qodo Review; Cursor’s Bugbot switched to usage-based billing (~$1.00–$1.50 per review run); cubic (formerly the mrge stacking tool, YC X25) ships custom review agents and Ultrareview. Ellipsis (pivoted to a managed cloud for coding agents) and Bito (pivoted to the Governor model router) were excluded as pivots, not omissions.
53 user stories · 318 judged cells · updated 2026-09-16 · Evidence as of 2026-09-16
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | free-tier vs Qodo ↗ | 35/100 | 78/100 | 26/100 | 25/100 | 29/100 | 14/45 verified · 3 disputed | 13/100 integrity | |||
| 2 | usage-based vs CodeRabbit ↗ | 37/100 | 57/100 | 0/100 | 26/100 | 24/100 | 14/39 verified | 20/100 integrity | |||
| 3 | free-tier vs CodeRabbit ↗ | 36/100 | 77/100 | 0/100 | 6/100 | 28/100 | npm 1.4k/wk | 24/38 verified | 71/100 integrity | ||
| 4 | free-tier vs CodeRabbit ↗ | 39/100 | 31/100 | 9/100 | 27/100 | 23/100 | npm 10.3k/wk | 17/40 verified · 2 disputed | 40/100 integrity | ||
| 5 | free-tier vs CodeRabbit ↗ | 30/100 | 68/100 | 0/100 | 12/100 | 24/100 | npm 49.9k/wk | 16/36 verified | 24/100 integrity | ||
| 6 | usage-based vs CodeRabbit ↗ | 26/100 | 51/100 | untested | 0/100 | 22/100 | 8/30 verified · 2 disputed | 0/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Best for security-engineer
Cursor Bugbot
80/100
Runner-up:
Greptile (60/100)
1 security-engineer story scored
Story matrix — every product × every judged storyStory matrix
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 9/10 | fullT 8/10 | partialT 5/10 | fullT 8/10 | fullT 8/10 | fullT 9/10 |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | partialT 6/10 | fullT 8/10 | partialT 4/10 | fullT 7/10 | partialC 6/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | fullC 7/10 | none 0/10 | none 0/10 | none 0/10 | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | none 0/10 | fullT 8/10 | fullT 6/10 | fullT 7/10 | n/a | fullT 8/10 |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | fullT 8/10 | fullT 8/10 | fullT 9/10 | fullT 8/10 | n/a | fullT 8/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | partialT 6/10 | partialT 6/10 | partialT 5/10 | partialT 6/10 | none 0/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | none 0/10 | none 0/10 | partialT 6/10 | none 0/10 | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Agentic features
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | fullX 9/10 | disputedD 6/10 | fullX 8/10 | fullX 8/10 | fullX 8/10 | fullX 9/10 |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | fullC 7/10 | fullC 7/10 | partialC 6/10 | partialC 6/10 | fullC 7/10 | fullC 8/10 |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | fullX 8/10 | disputedD 5/10 | fullX 8/10 | fullX 7/10 | partialC 5/10 | fullX 7/10 |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | fullC 7/10 | partialX 5/10 | fullC 7/10 | partialC 6/10 | partialC 6/10 | fullC 7/10 |
Api quality
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | fullT 9/10 | none 0/10 | none 0/10 | none 0/10 | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | none 0/10 | fullC 6/10 | none 0/10 | n/a | n/a | n/a |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a | none 0/10 |
Autofix agents — stories about autofix agents in this arenaAutofix agents
Ai authored
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Autofix agents — stories about autofix agents in this arenaThe reviewer holds the line on AI-generated PRs — it verifies agent-authored code at a volume no human team could review | ai-native | disputedD 6/10 | fullX 8/10 | partialX 6/10 | partialX 7/10 | disputedD 6/10 | fullX 8/10 |
Checks
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Autofix agents — stories about autofix agents in this arenaI define custom agentic pre-merge checks in plain language — 'docs updated', 'tests cover new paths' — that run on every PR | ai-native | partialC 6/10 | partialC 6/10 | partialC 5/10 | partialC 5/10 | partialC 6/10 | fullX 7/10 |
Fixes
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Autofix agents — stories about autofix agents in this arenaI turn a review finding into an applied fix — a committed patch or an agent-generated follow-up — without leaving the PR | developer | fullC 8/10 | fullC 8/10 | fullC 7/10 | fullC 7/10 | partialX 6/10 | fullC 8/10 |
Handoff
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Autofix agents — stories about autofix agents in this arenaReview findings hand off cleanly to my coding agent — copyable fix prompts or direct integration with Claude Code, Cursor, or Codex | ai-native | partialC 7/10 | fullC 9/10 | partialT 5/10 | fullT 7/10 | partialC 6/10 | fullT 8/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | partialC 4/10 | partialC 6/10 | partialX 6/10 | partialC 5/10 | partialC 4/10 | partialC 5/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | partialC 7/10 | partialC 5/10 | partialC 5/10 | partialC 6/10 | partialC 6/10 | partialX 5/10 |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | partialC 3/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | partialC 3/10 | partialC 4/10 | partialC 5/10 | partialC 4/10 | n/a | partialC 3/10 |
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding
Context
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyThe reviewer understands changes that span multiple repositories or a large monorepo and reviews them coherently | engineering-lead | partialC 7/10 | none 0/10 | partialC 4/10 | partialC 7/10 | none 0/10 | partialC 7/10 |
| Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyReview comments reflect the whole repository — call sites, related modules, existing conventions — not just the changed hunks | developer | fullC 7/10 | partialX 6/10 | fullC 7/10 | fullC 8/10 | partialC 4/10 | partialC 7/10 |
Memory
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyThe reviewer builds a persistent memory of my team's conventions and past review decisions and applies it to future PRs | ai-native | fullC 8/10 | fullC 7/10 | partialC 4/10 | fullC 7/10 | partialC 4/10 | fullX 8/10 |
Interaction — how you steer it — commands, replies, review conversations, configurability in the loopInteraction
Chat
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Interaction — how you steer it — commands, replies, review conversations, configurability in the loopI reply to the reviewer in the PR thread to ask questions, get explanations, or issue commands — and it answers in context | developer | fullC 8/10 | partialX 5/10 | fullC 7/10 | fullC 7/10 | partialC 4/10 | fullX 8/10 |
Control
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Interaction — how you steer it — commands, replies, review conversations, configurability in the loopI control when reviews run — skip drafts, trigger on demand, filter by branch or label — so the bot shows up only when wanted | developer | fullC 8/10 | partialC 5/10 | partialC 4/10 | partialC 6/10 | partialC 5/10 | partialC 4/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialT 3/10 | partialT 5/10 | partialT 4/10 | partialT 4/10 | none 0/10 | partialT 5/10 |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | partialC 2/10 | none 0/10 | partialX 4/10 | none 0/10 | none 0/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | fullC 6/10 | fullC 7/10 | none 0/10 | fullC 7/10 | n/a | none 0/10 |
Pr integration — stories about pr integration in this arenaPr integration
Platforms
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Pr integration — stories about pr integration in this arenaThe reviewer installs as a GitHub/GitLab app and posts reviews as native inline comments on my pull requests within minutes | developer | fullX 8/10 | partialX 7/10 | partialX 6/10 | fullC 7/10 | fullX 8/10 | fullX 8/10 |
Suggestions
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Pr integration — stories about pr integration in this arenaReview comments include committable suggested diffs I can apply with one click | developer | fullC 9/10 | partialC 5/10 | fullC 7/10 | fullC 7/10 | partialC 5/10 | fullC 7/10 |
Summaries
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Pr integration — stories about pr integration in this arenaEvery PR gets an auto-generated summary and change walkthrough so human reviewers orient fast | developer | fullX 9/10 | fullX 7/10 | partialC 6/10 | fullC 9/10 | none 0/10 | fullX 9/10 |
Updates
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Pr integration — stories about pr integration in this arenaPushing new commits triggers an incremental re-review that tracks what was fixed instead of repeating old comments | developer | fullC 8/10 | partialC 7/10 | none 0/10 | partialC 4/10 | fullT 7/10 | partialC 6/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | partialC 5/10 | partialC 6/10 | none 0/10 | partialC 5/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | fullC 9/10 | partialC 4/10 | fullC 8/10 | fullC 8/10 | partialC 6/10 | fullC 8/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | partialC 6/10 | partialC 5/10 | none 0/10 | partialC 4/10 | partialC 4/10 | partialC 3/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | partialC 5/10 | partialC 4/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 |
Quality gates — stories about quality gates in this arenaQuality gates
Analytics
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Quality gates — stories about quality gates in this arenaI see dashboards of findings, acceptance rates, and review coverage across my org | engineering-lead | fullC 8/10 | none 0/10 | partialC 3/10 | partialC 6/10 | partialC 4/10 | fullC 7/10 |
Gates
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Quality gates — stories about quality gates in this arenaThe reviewer can gate merges — a required status check or blocking review that enforces resolution of critical findings | engineering-lead | fullC 7/10 | none 0/10 | partialC 5/10 | none 0/10 | none 0/10 | none 0/10 |
Review accuracy — stories about review accuracy in this arenaReview accuracy
Detection
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Review accuracy — stories about review accuracy in this arenaThe reviewer catches real bugs in my PR — logic errors, race conditions, broken edge cases — not just style nits | developer | disputedD 6/10 | partialX 6/10 | partialX 4/10 | partialX 5/10 | disputedD 6/10 | fullX 7/10 |
Learning
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Review accuracy — stories about review accuracy in this arenaPush back on a bad review comment and the reviewer learns — it stops repeating the same rejected feedback | developer | partialX 7/10 | partialC 7/10 | partialC 4/10 | none 0/10 | partialX 4/10 | fullX 7/10 |
Noise
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Review accuracy — stories about review accuracy in this arenaThe reviewer keeps noise low — few false positives, deduplicated comments, severity labels — so my team doesn't tune it out | engineering-lead | disputedD 4/10 | partialX 6/10 | partialX 5/10 | partialX 6/10 | partialX 5/10 | partialX 6/10 |
Security
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Review accuracy — stories about review accuracy in this arenaReviews flag security problems in the diff — injection risks, leaked secrets, insecure patterns — alongside functional bugs | security-engineer | partialX 6/10 | fullX 6/10 | none 0/10 | partialC 5/10 | fullT 8/10 | partialX 6/10 |
Surfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksSurfaces
Cli
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Surfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksI run reviews from a CLI against local diffs or in CI scripts, with machine-readable output my tooling can consume | developer | partialT 5/10 | partialT 5/10 | none 0/10 | partialT 5/10 | none 0/10 | partialT 5/10 |
Ide
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Surfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksI get the same review inside my IDE before I push, catching issues while the code is still in my editor | developer | fullT 8/10 | partialT 6/10 | none 0/10 | partialT 6/10 | none 0/10 | fullT 7/10 |
Workflow config — stories about workflow config in this arenaWorkflow config
Config
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Workflow config — stories about workflow config in this arenaI configure the reviewer with a versioned config file in my repo — path filters, per-path instructions, review profiles | engineering-lead | fullC 8/10 | partialC 6/10 | partialC 5/10 | partialC 6/10 | partialC 5/10 | fullX 7/10 |
Governance
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Workflow config — stories about workflow config in this arenaI roll out org-level review defaults across hundreds of repos and manage exceptions centrally | engineering-lead | partialC 7/10 | partialC 7/10 | partialC 5/10 | fullC 8/10 | partialC 5/10 | partialC 5/10 |
Rules
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Workflow config — stories about workflow config in this arenaI encode my team's own review guidelines — natural-language rules, AST patterns, or linked style guides — and the reviewer enforces them | engineering-lead | fullC 9/10 | fullC 8/10 | partialC 6/10 | partialC 7/10 | fullC 7/10 | partialX 7/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.