Skip to content

Rank #6 of 6 in AI Code Review

Cursor Bugbot logo

Anysphere, Inc. · commercial

no public signals

Showcase

Cursor Bugbot homepage screenshot
homepage · captured Sep 2026 · view live ↗
Cursor Bugbot docs screenshot
docs · captured Sep 2026 · view live ↗

Cursor ships more than one product — each judged line competes in its own arena on the same stories as everyone else.

LineArenaRankPA Score
CursorAI Coding Agents#11/1327/100
Cursor Bugbotthis pageAI Code Review#6/624/100

Try itExperimental

See what an agent can do with Cursor Bugbot before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$curl -s https://cursor.com/llms.txt | head -6recorded session — replayed, not live
recorded 2026-09-10 · exit 0 · captured verbatim by our probe harness, secrets redacted · pure-HTTP probe — ▶ run live re-runs it from our edge

Verified integrations

No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

38.6/100

Autofix agents — stories about autofix agents in this arenaAutofix agentsevidence →

Stories about autofix agents in this arena

32.0/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

22.3/100

Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understandingevidence →

How deeply the tool maps your repo — cross-file context, architecture awareness, history

17.1/100

Interaction — how you steer it — commands, replies, review conversations, configurability in the loopInteractionevidence →

How you steer it — commands, replies, review conversations, configurability in the loop

26.0/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

0.0/100

Pr integration — stories about pr integration in this arenaPr integrationevidence →

Stories about pr integration in this arena

47.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

22.7/100

Quality gates — stories about quality gates in this arenaQuality gatesevidence →

Stories about quality gates in this arena

8.0/100

Review accuracy — stories about review accuracy in this arenaReview accuracyevidence →

Stories about review accuracy in this arena

35.2/100

Surfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksSurfacesevidence →

Where it meets your workflow — IDE, CLI, web, PR comments, CI checks

0.0/100

Workflow config — stories about workflow config in this arenaWorkflow configevidence →

Stories about workflow config in this arena

43.3/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 3 free · 3 paid · 0 enterprise · 19 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3partial5/10C

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3none0/10

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10X

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partialpaid6/10C

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10C

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1n/auntestednone yet

The reviewer installs as a GitHub/GitLab app and posts reviews as native inline comments on my pull requests within minutes C

Platforms

developerPr integration — stories about pr integration in this arenaPr integration3full8/10X

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial6/10C

I turn a review finding into an applied fix — a committed patch or an agent-generated follow-up — without leaving the PR C

Fixes

developerAutofix agents — stories about autofix agents in this arenaAutofix agents3partial6/10X

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3partialfree6/10C

The reviewer catches real bugs in my PR — logic errors, race conditions, broken edge cases — not just style nits C

Detection

developerReview accuracy — stories about review accuracy in this arenaReview accuracy3disputed6/10D

I configure the reviewer with a versioned config file in my repo — path filters, per-path instructions, review profiles C

Config

engineering leadWorkflow config — stories about workflow config in this arenaWorkflow config3partialpaid5/10C

Review comments include committable suggested diffs I can apply with one click C

Suggestions

developerPr integration — stories about pr integration in this arenaPr integration3partial5/10C

The reviewer keeps noise low — few false positives, deduplicated comments, severity labels — so my team doesn't tune it out C

Noise

engineering leadReview accuracy — stories about review accuracy in this arenaReview accuracy3partial5/10X

Review comments reflect the whole repository — call sites, related modules, existing conventions — not just the changed hunks C

Context

developerCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding3partial4/10C

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3n/auntestednone yet

Reviews flag security problems in the diff — injection risks, leaked secrets, insecure patterns — alongside functional bugs C

Security

security engineerReview accuracy — stories about review accuracy in this arenaReview accuracy2full8/10T

I encode my team's own review guidelines — natural-language rules, AST patterns, or linked style guides — and the reviewer enforces them C

Rules

engineering leadWorkflow config — stories about workflow config in this arenaWorkflow config2full7/10C

Pushing new commits triggers an incremental re-review that tracks what was fixed instead of repeating old comments C

Updates

developerPr integration — stories about pr integration in this arenaPr integration2full7/10T

I define custom agentic pre-merge checks in plain language — 'docs updated', 'tests cover new paths' — that run on every PR C

Checks

ai-native userAutofix agents — stories about autofix agents in this arenaAutofix agents2partial6/10C

Review findings hand off cleanly to my coding agent — copyable fix prompts or direct integration with Claude Code, Cursor, or Codex C

Handoff

ai-native userAutofix agents — stories about autofix agents in this arenaAutofix agents2partial6/10C

The reviewer holds the line on AI-generated PRs — it verifies agent-authored code at a volume no human team could review C

Ai authored

ai-native userAutofix agents — stories about autofix agents in this arenaAutofix agents2disputed6/10D

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partialfree4/10C

I reply to the reviewer in the PR thread to ask questions, get explanations, or issue commands — and it answers in context C

Chat

developerInteraction — how you steer it — commands, replies, review conversations, configurability in the loopInteraction2partial4/10C

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partialfree4/10C

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial4/10C

Push back on a bad review comment and the reviewer learns — it stops repeating the same rejected feedback C

Learning

developerReview accuracy — stories about review accuracy in this arenaReview accuracy2partial4/10X

The reviewer builds a persistent memory of my team's conventions and past review decisions and applies it to future PRs C

Memory

ai-native userCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding2partial4/10C

Every PR gets an auto-generated summary and change walkthrough so human reviewers orient fast C

Summaries

developerPr integration — stories about pr integration in this arenaPr integration2none0/10

I get the same review inside my IDE before I push, catching issues while the code is still in my editor C

Ide

developerSurfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksSurfaces2none0/10

I run reviews from a CLI against local diffs or in CI scripts, with machine-readable output my tooling can consume C

Cli

developerSurfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksSurfaces2none0/10

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2none0/10

The reviewer can gate merges — a required status check or blocking review that enforces resolution of critical findings C

Gates

engineering leadQuality gates — stories about quality gates in this arenaQuality gates2none0/10

The reviewer understands changes that span multiple repositories or a large monorepo and reviews them coherently C

Context

engineering leadCodebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2noneuntestednone yet

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2noneuntestednone yet

I control when reviews run — skip drafts, trigger on demand, filter by branch or label — so the bot shows up only when wanted C

Control

developerInteraction — how you steer it — commands, replies, review conversations, configurability in the loopInteraction1partialpaid5/10C

I roll out org-level review defaults across hundreds of repos and manage exceptions centrally C

Governance

engineering leadWorkflow config — stories about workflow config in this arenaWorkflow config1partial5/10C

I see dashboards of findings, acceptance rates, and review coverage across my org C

Analytics

engineering leadQuality gates — stories about quality gates in this arenaQuality gates1partial4/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1n/auntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 33 stories with headroom

What would move Cursor Bugbot’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productDrive the product through a documented public API

    nonemoves agent-readyimpact 45

    Evidence shows Bugbot is driven via PR comment triggers (`bugbot run`, `cursor review`) and VCS integrations (GitHub/GitLab/Bitbucket/Azure DevOps), not via any documented public REST/webhook API for programmatic control.

  2. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    Missing: any export tool, data portability feature, or open-format download of Bugbot's findings/config.

  3. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    No evidence anywhere in the pack of webhook subscription capability for Bugbot events; it integrates via PR comments/git provider webhooks internally but exposes no user-facing webhook subscription API.

  4. Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product

    partialq5/10moves Built-in AIimpact 22.5

    Missing: evidence of open-ended task delegation beyond PR review/fix workflows, and a conversational/general-assistant interface within Bugbot itself.

  5. Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows

    nonemoves PA Scoreimpact 20

    Bugbot's docs describe event-triggered reviews (automatic on PR update, or manual comment trigger) and org-wide 'Automations' rules, but there is no evidence of true recurring/scheduled job execution (e.g., cron-like or time-based triggers) as opposed to PR-event triggers.

  6. Quality gates — stories about quality gates in this arenaThe reviewer can gate merges — a required status check or blocking review that enforces resolution of critical findings

    nonemoves PA Scoreimpact 20

    Missing: any mention of CI status check integration, merge-blocking configuration, or required-review enforcement.

  7. Surfaces — where it meets your workflow — IDE, CLI, web, PR comments, CI checksI run reviews from a CLI against local diffs or in CI scripts, with machine-readable output my tooling can consume

    nonemoves PA Scoreimpact 20

    Missing: any CLI binary/command, local diff support, structured/machine-readable output format, or CI-script-oriented API.

  8. Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyThe reviewer understands changes that span multiple repositories or a large monorepo and reviews them coherently

    nonemoves PA Scoreimpact 20

    Missing: any mention of cross-repo dependency awareness, monorepo-scale indexing, or coordinated review across repos in a single PR/change set.

Showing the top 8 of 33 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 30 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Help docs25 stories

docs24 stories

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$curl -s https://cursor.com/llms.txt | head -6reproduced
$ curl -s https://cursor.com/llms.txt | head -6
# Cursor Documentation

## Get Started

- https://cursor.com/docs.md
- https://cursor.com/docs/get-started/quickstart.md
$curl -sL https://cursor.com/docs/bugbot.md | head -6reproduced
$ curl -sL https://cursor.com/docs/bugbot.md | head -6
# Bugbot

Bugbot reviews pull requests and identifies bugs, security issues, and code quality problems.

Configure Bugbot in [Automations](https://cursor.com/automations/from-cursor/bugbot).

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

4 of 12 testable claims verified · 2 contradictedintegrity 0/100

17 distinct capability claims found in Cursor Bugbot’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

4

Verified

6

Unverified

2

Contradicted

18

Undersold

Verified (6)
Unverified (7)
Contradicted (2)
Undersold (18)
Claims outside our story set (4)

Real capability claims found in Cursor Bugbot’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Verbose trigger mode replies with a table listing which rules were applied in that review

    source ↗
  • With usage-based billing, users can configure Bugbot's effort level to run deeper, longer reviews

    source ↗
  • Natural-language instructions can dynamically set Bugbot's review effort level (low/default/high)

    source ↗
  • Usage-based billing can be enabled to run reviews on all PRs

    source ↗
Suggest a story for these →

Business model

usage-basedsubscription-flatenterprise-custom

Requires a paid Cursor plan (Pro $20/mo, Pro+ $60/mo, Ultra $200/mo); since May 2026 Bugbot bills usage-based at roughly $1.00-$1.50 per review run; Enterprise is custom.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score22 (Sep 10 '26)24 (Sep 16 '26)
Agent-ready26 (Sep 10 '26)26 (Sep 16 '26)

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime llms.txt up (tracking since Sep 11 '26)