Cursor vs Devin
Devin wins · 17–23 (29 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to DevinCursornone0/10No evidence pack item mentions llms.txt, agent-oriented documentation ingestion, or a mechanism to point Cursor's agent at such files; only generic doc/MCP/tooling references are present. missing for 10: any mention of llms.txt support, crawling agent-oriented doc formats, or a documented feature for feeding external agent docs to Cursor's agent.
Devin's own docs site serves llms.txt (HTTP 200, confirmed by probe) and per-page .md variants, and also supports AGENTS.md as an agent-oriented instructions standard, directly matching the story. Missing for 10: no independent/community confirmation that external agents have actually consumed llms.txt successfully.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.devin.ai/llms.txt # Devin Docs - [Desktop (100 pages)](https://docs.devin.ai/_llms/en/desktop.md):…”
- [probe] “PROBE docs-md: HTTP 200 at https://docs.devin.ai/get-started/devin-intro.md > ## Documentation Index > Fetch the complete documentation inde…”
- [claimed-docs] “Devin supports AGENTS.md - a simple, open standard for providing context and instructions to AI agents.”
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to DevinCursor ships an official CLI (cursor.com/cli, curl installer) and background/cloud agents that run 'on schedules or triggers' to build and fix software autonomously, which implies non-interactive/headless automation. However, there is no explicit documentation of CI pipeline integration, exit codes, or scripting examples for pipelines. Missing for 10: explicit CI/CD integration docs (e.g., GitHub Actions example), documented headless flags/exit-code behavior, and independent confirmation of CLI use in automated pipelines.
- [probe] “official CLI documented at https://cursor.com/cli”
- [claimed-docs] “curl https://cursor.com/install -fsS | bash”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin offers a full API for creating/managing sessions programmatically (including create_as_user_id for automation on behalf of users), a CLI with a --sandbox flag for OS-level isolated headless runs, and explicit CI/CD pipeline integration for responding to static analysis findings and PR checks, all supporting headless/automated usage without a human in the loop. Missing for 10: no independent/hands-on report specifically validating CI automation workflows end-to-end, and no explicit CI example (e.g., GitHub Actions snippet) beyond doc references.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “The --sandbox flag runs the CLI with OS-level isolation, enforcing writable paths and deny rules at the operating-system level and optionall…”
- [claimed-docs] “a local command-line coding agent with deep Devin Cloud integration”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round to CursorCursor's docs explicitly describe MCP support: connecting to external tools/data sources, marketplace one-click install with OAuth, custom JSON server configuration, toggling servers, and enterprise admin controls over allowed servers. This directly matches the story of plugging in MCP servers so the agent can use their tools. Missing for 10: independent hands-on verification of MCP tool usage in practice and no community corroboration of the feature's reliability.
- [claimed-docs] “Model Context Protocol (MCP) enables Cursor to connect to external tools and data sources.”
- [claimed-docs] “Click "Add to Cursor" on a marketplace entry to install it and authenticate with OAuth.”
- [claimed-docs] “Configure custom MCP servers with a JSON file”
- [claimed-docs] “Enterprise admins can control which MCP servers users may run from the Cursor dashboard.”
- [claimed-docs] “Toggle servers on/off without removing them”
Devinnone0/10The evidence only documents Devin exposing its own MCP server so other agents/IDEs can call Devin's tools (session management, playbooks, knowledge, scheduling) — the reverse direction of this story. There is no evidence that a user can configure Devin itself to consume/plug in external MCP servers so Devin can use their tools.
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [probe] “official MCP server documented at https://docs.devin.ai/work-with-devin/devin-mcp”
ai-native userUse an official CLI
weight 2 · round drawnCursor documents an official CLI with an install command (curl https://cursor.com/install) and a dedicated CLI docs page (cursor.com/cli), confirming a first-party terminal tool for AI-native workflows. Missing for 10: independent/hands-on corroboration of CLI capabilities and depth of documentation beyond install instructions.
- [claimed-docs] “curl https://cursor.com/install -fsS | bash”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [probe] “official CLI documented at https://cursor.com/cli”
Devin has an official CLI ('Devin CLI, a local command-line coding agent with deep Devin Cloud integration') with documented usage examples and a sandbox flag for OS-level isolation, confirmed by both docs and probe. Missing for 10: independent/hands-on community verification of the CLI specifically (community evidence covers the web/session product, not CLI usage).
- [claimed-docs] “Devin CLI, a local command-line coding agent with deep Devin Cloud integration.”
- [claimed-docs] “devin -- check out this code and suggest a feasible, helpful feature”
- [claimed-docs] “The --sandbox flag runs the CLI with OS-level isolation, enforcing writable paths and deny rules at the operating-system level and optionall…”
- [claimed-docs] “a local command-line coding agent with deep Devin Cloud integration”
- [probe] “official CLI documented at https://docs.devin.ai/cli/index”
ai-native userDrive the product through a documented public API
weight 3 · round to DevinEvidence shows an official CLI (cursor.com/cli) that lets users invoke Cursor from scripts, which partially satisfies 'driving the product programmatically,' but there is no documented public REST/SDK API, authentication scheme, or endpoint reference — MCP docs describe Cursor consuming external tools, not exposing itself as an API. Missing for 10: documented REST/GraphQL API, SDK/client libraries, API authentication and rate-limit docs, independent corroboration of programmatic usage.
- [probe] “official CLI documented at https://cursor.com/cli”
- [claimed-docs] “curl https://cursor.com/install -fsS | bash”
Devin ships a documented public API (docs-7) with session creation, org-level features like create_as_user_id (docs-8), plus a CLI and MCP server for programmatic/agentic control (docs-4, docs-6, probe-4, probe-5), directly enabling AI-native users to drive it programmatically. Missing for 10: a discoverable OpenAPI/swagger spec (probe-3 found 404s on all candidate paths) and independent hands-on corroboration of API usage.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [claimed-docs] “Devin CLI, a local command-line coding agent with deep Devin Cloud integration.”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [probe] “official MCP server documented at https://docs.devin.ai/work-with-devin/devin-mcp”
- [probe] “official CLI documented at https://docs.devin.ai/cli/index”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round drawnCursornone0/10No evidence Cursor lets users mint scoped or least-privilege API credentials for agents; docs cover MCP server toggling and enterprise admin control of which servers can run, but nothing about issuing scoped/limited API keys or credentials specifically for agent use.
Devinnone0/10Devin exposes a general API (devin-docs-7) and can act on behalf of a specified user via create_as_user_id (devin-docs-8), but there is no evidence of scoped/least-privilege API key or token issuance, role-based permission scopes, or credential-level restriction mechanisms for agent access.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
ai-native userBuild against official SDKs
weight 2 · round to DevinCursornone0/10Evidence shows Cursor offers a CLI, MCP integration, and marketplace extensions, but there is no mention of any official SDK (e.g., a documented library/API package) for developers to build against Cursor itself.
Devin provides an API for integration (docs-7, docs-8) allowing developers to build applications and automate workflows, but the evidence never mentions dedicated official SDKs/client libraries (e.g., Python/JS packages), and probes for an OpenAPI spec that would back SDK generation all returned 404s (devin-probe-3). Missing for 10: named SDK packages in specific languages, SDK installation/usage docs, and a published OpenAPI/schema artifact confirming SDK-generation support.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
ai-native userSubscribe to events via webhooks
weight 2 · round drawnCursornone0/10Evidence covers MCP integration, background agents, and IDE integrations, but there is no mention of a webhook subscription mechanism for external event notifications.
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round drawnCursor's core value proposition is analyzing the user's codebase to surface AI-generated insights (tracing repo structure, finding root causes, reviewing diffs) and suggestions for next actions, as documented across multiple first-party docs. Missing for 10: independent/hands-on evidence validating the accuracy or depth of these insights, and no detail on insight types beyond code-centric suggestions (e.g., data analytics or business data outside code).
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
- [claimed-docs] “Scope changes, use Plan Mode, and ship bigger work with confidence”
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
Devin generates insights/suggestions from a user's own codebase data via 'Ask Devin' (code structure/dependency Q&A), auto-generated DeepWiki documentation, and Devin Review's automated PR feedback, all built on repository indexing. Missing for 10: independent/hands-on validation of the quality of these AI-generated insights (community evidence covers general task execution issues, not this specific feature) and no benchmark of insight accuracy.
- [claimed-docs] “Ask Devin can answer questions about code structure and dependencies, and help you scope and plan tasks before implementation.”
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “Use DeepWiki to navigate architecture and code with auto-generated documentation.”
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
ai-native userSet up automations that run autonomously in the background
weight 2 · round drawnCursor explicitly documents 'always-on agents that run on schedules or triggers to build, maintain, and fix your software' and 'fleets of agents that work in parallel for hours or days,' directly matching autonomous background automation. This is first-party vendor documentation without independent hands-on corroboration of scheduling/triggers working reliably. Missing for 10: independent/community verification that scheduled/triggered background agents work reliably in practice, and more detail on trigger configuration options.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin explicitly supports background/autonomous execution: cloud sessions run in their own VM and 'keep going after you close your laptop' (devin-docs-13), MCP access includes 'scheduling' (devin-docs-6/23), the API lets you 'automate workflows' and create sessions programmatically (devin-docs-7/8), CI/CD integration triggers Devin on findings (devin-docs-19), and Auto-Fix creates a closed loop that iterates PRs 'without you in the loop' (devin-docs-18). This spans scheduled triggers, API-driven automation, and hands-off background operation. Missing for 10: independent/community verification that scheduled automations run reliably unattended, and more detail on a dedicated 'automation/schedule' UI beyond scattered doc mentions.
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to CursorCursor's docs clearly describe delegating tasks to built-in agents that plan, code, test, and demo work end-to-end while the user focuses on review/decisions, including background/parallel agents and always-on scheduled agents. This is a core, heavily documented capability of the product, though independent hands-on validation of agent task quality is thin (only general community commentary, some critical, exists). Missing for 10: deeper independent verification of agent task success rates beyond vendor docs.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
- [claimed-docs] “Scope changes, use Plan Mode, and ship bigger work with confidence”
Devin's core product is designed for task delegation — via Ask Devin, ticket assignment, Slack/Teams tagging, and a conversational IDE (devin-docs-1, devin-docs-3, devin-docs-24) — so the axis clearly applies and is well documented. However, hands-on community reports show real caveats: Devin can add extraneous unrequested changes it can't undo, gets stuck for long periods without asking for help, and requires active babysitting/session termination to get value (devin-comm-1, devin-comm-2, devin-comm-5), undercutting a fully seamless delegation experience. Missing for 10: independent verification that delegated tasks reliably complete without extraneous side-effects or getting stuck, and stronger corroboration beyond one HN thread.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
ai-native userOperate the product with natural-language commands
weight 2 · round drawnCursor's core interaction model is natural-language driven agents that plan, code, test, and operate across terminal/Slack/GitHub (cursor-docs-2, cursor-docs-8, cursor-docs-9, cursor-docs-10, cursor-docs-11), consistent with an AI-native product. Missing for 10: independent hands-on evidence specifically validating natural-language command reliability/accuracy (community evidence focuses on bugginess/pricing complaints unrelated to NL command capability itself).
- [claimed-docs] “Scope changes, use Plan Mode, and ship bigger work with confidence”
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin is explicitly designed as a conversational agent: users assign tasks via natural language (Slack/Teams tagging, chat interface, CLI prompts like 'devin -- check out this code...'), and it interprets these into autonomous coding/dev actions across IDE, CLI, and cloud sessions. Community evidence corroborates it operates on natural-language task descriptions in practice, though with noted friction around scope creep and knowing when to stop. Missing for 10: independent benchmarking of NL command accuracy/robustness and richer detail on how ambiguous instructions are resolved.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Tagging Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “devin -- check out this code and suggest a feasible, helpful feature”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round drawnCursornone0/10No evidence of an interactive API reference with runnable examples for Cursor; docs entries describe product features and MCP setup but nothing about an API reference or executable code samples.
Devinnone0/10Devin has an API reference overview but the openapi.json probe returned 404 on all candidate paths, and there's no mention of an interactive reference with runnable examples (e.g., 'try it' console) in the docs pack.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnCursornone0/10No evidence of Cursor publishing a downloadable OpenAPI or equivalent machine-readable API spec; docs reference MCP config and CLI but not an API spec.
Devinnone0/10Devin has a documented API (devin-docs-7) but a direct probe for a machine-readable OpenAPI/swagger spec returned 404 on all candidate paths, and no documentation item references a downloadable spec file.
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnCursornone0/10No evidence in the pack mentions API versioning or a deprecation policy for Cursor's APIs (CLI, extensions, or MCP config); docs cover features like MCP setup, agents, and integrations but nothing about version stability guarantees or deprecation timelines.
Devinnone0/10There's an API reference (devin-docs-7) and even an OpenAPI probe, but that probe found no OpenAPI spec (devin-probe-3), and no evidence anywhere mentions API versioning scheme or a documented deprecation policy for the API.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round drawnCursor supports launching 'fleets of agents' in parallel and always-on scheduled/triggered agents, which enables some multi-item automation, but there's no direct evidence of bulk operations across many discrete items (e.g., bulk file edits, batch refactors, or multi-repo operations) as a first-class feature. missing for 10: explicit documentation or hands-on evidence of bulk/batch operations across many items (files, tickets, repos), user-facing UI for selecting many items at once, and independent corroboration of this working in practice.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
Devin's API supports programmatic session creation (including on behalf of other users) and docs explicitly encourage carving out independent tasks to run simultaneously, which together enable bulk-style automation across many tickets/items, but there is no dedicated 'bulk operations' or batch-processing feature documented, and community feedback raises concerns about reliability/oversight needed per session that would complicate true bulk workflows. Missing for 10: an explicit batch/bulk API endpoint or UI for processing many items in one request, and independent evidence of successful large-scale bulk runs.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round drawnCursor docs describe 'always-on agents that run on schedules or triggers' and a way to 'add rules' from one place, matching the idea of rule-based automation triggered by events. However the evidence pack doesn't detail how rules are authored/scoped to specific events beyond the marketing blurb, and there's no independent/hands-on confirmation of this automation working as described. Missing for 10: concrete rule-definition syntax/examples, independent verification that scheduled/triggered agents reliably fire on events, and detail on event types supported.
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Add plugins, skills, MCPs, and rules from one place”
Devin supports several built-in event-triggered automations — Auto-Fix responds automatically to PR review comments and CI failures, CI/CD integration triggers Devin off static-analysis findings (SonarQube/Fortify/Veracode), and MCP exposes 'scheduling' as a capability — but there's no evidence of a general-purpose, user-defined rules/webhook engine for arbitrary custom triggers. Missing for 10: documentation of a configurable custom-rule/webhook trigger system, details on the scheduling feature's flexibility, and independent confirmation that these automations work reliably as event triggers.
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Enable Devin Review with Auto-Fix so Devin automatically responds to code review comments, fixe”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to CursorCursor documents 'always-on agents that run on schedules or triggers to build, maintain, and fix your software,' directly matching recurring scheduled workflow automation, alongside parallel agent fleets for ambitious tasks. Missing for 10: independent hands-on verification of scheduling reliability, details on trigger configuration options, and any community corroboration of this specific feature working in practice.
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
Devin's MCP docs mention 'scheduling' as one of the capabilities exposed to MCP-compatible agents, implying some scheduling functionality exists, but there is no dedicated documentation, UI, or examples describing recurring jobs, cron-like triggers, or workflow automation configuration. missing for 10: dedicated scheduling feature docs, examples of recurring/cron jobs, independent confirmation of scheduled workflows in practice.
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnCursornone0/10Evidence shows Cursor can inspect diffs and review changes before merge, but there is no documented capability to version, review, or roll back the automations themselves (e.g., scheduled/always-on agents, rules, MCP configs) as distinct artifacts with history/rollback support.
Devinnone0/10The evidence shows Devin has knowledge, playbooks, and scheduling features but nothing about versioning, reviewing, or rolling back those automation configurations themselves; Devin Review/Auto-Fix pertains to PR code review, not to the automation definitions. Missing for 10: version history for playbooks/knowledge, a review workflow for automation changes, and a rollback mechanism for automations.
Autonomy agents — stories about autonomy agents in this arenaAutonomy agents
Stories about autonomy agents in this arena
Background execution
ai-native userHave a cloud agent build, test, and demo a feature end-to-end for my review
weight 2 · round to CursorCursor's docs explicitly describe cloud/background agents that 'use their own computers to build, test, and demo features end to end for you to review,' plus the ability to launch fleets of agents working in parallel for hours/days, and always-on scheduled agents — directly matching the story. Corroboration is entirely first-party marketing/docs rather than independent hands-on verification of an actual demo workflow. Missing for 10: independent/hands-on evidence confirming the build-test-demo loop works reliably end-to-end, and detail on what 'demo' concretely produces (e.g., preview links, recordings).
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin's cloud sessions run in dedicated VMs with shell/browser/full repo access, can implement features, run tests, and continue autonomously after handoff, then present PRs for review (docs-1,13,17). However, community hands-on reports describe unreliable autonomy — extraneous breaking changes, inability to self-correct, and needing frequent human monitoring/termination — undercutting the 'build, test, demo end-to-end' promise. missing for 10: reliable independent verification of unattended end-to-end demo quality, and clearer evidence of a built-in demo/walkthrough artifact for reviewers beyond PR creation.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
developerDelegate longer-running coding tasks to run in the background in an isolated cloud environment
weight 3 · round to DevinCursor documents cloud/background agents ('Agents use their own computers to build, test, and demo features end to end', 'Launch fleets of agents that work in parallel on ambitious tasks for hours or days', and hand-off delegation while the developer focuses elsewhere), matching the isolated cloud-background-task story. Missing for 10: independent hands-on verification of the background agent's isolation/reliability and details on session duration limits or failure modes.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
Devin runs tasks in isolated cloud VMs with full shell/browser/repo access that persist after the user disconnects, explicitly supporting long-running background work and parallel independent tasks, corroborated by community reports of multi-day autonomous runs. Missing for 10: independent third-party benchmarking of long-running task reliability/quality beyond anecdotal HN reports.
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
- [claimed-docs] “Configure it once, and every session boots into that known-good state.”
- [claimed-docs] “Hand a task off to a cloud Devin session and keep working locally.”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “You can set a 'max work time' before Devin pauses so it won't go for days endlessly spending your credits. By default it's set to 10 credits…”
developerConfigure a reproducible cloud environment with the dependencies and setup steps my repository needs
weight 2 · round to DevinCursor's docs mention cloud/background agents that 'use their own computers to build, test, and demo features' and can be launched in fleets or run on schedules, implying some cloud execution environment, but there's no evidence pack detail on how a developer configures dependencies, install scripts, or a reproducible environment spec (e.g. Dockerfile/environment.json) for these agents. Missing for 10: explicit documentation of environment configuration format, dependency/setup step definition, and evidence of reproducibility across runs.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin explicitly supports configuring environment 'blueprints' that specify tools, runtimes, and dependencies so 'every session boots into that known-good state,' with auto-detection of requirements from the repo (docs-31, docs-32), plus indexing (docs-27), knowledge/AGENTS.md context files (docs-29, docs-30), and VPN access for internal dependencies (docs-28), all running in isolated cloud VMs (docs-13). Missing for 10: independent/hands-on confirmation that blueprint-based environments reliably reproduce across sessions in practice.
- [claimed-docs] “Configure it once, and every session boots into that known-good state.”
- [claimed-docs] “Devin inspects your repository, figures out which tools, runtimes, and dependencies are needed, and generates the blueprint for you.”
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
- [claimed-docs] “Knowledge is a collection of instructions and advice that Devin can reference in all sessions.”
- [claimed-docs] “Devin supports AGENTS.md - a simple, open standard for providing context and instructions to AI agents.”
- [claimed-docs] “Devin can connect to a VPN from inside its workspace, so sessions can reach internal services such as package registries, databases, and int…”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
Parallel agents
ai-native userLaunch fleets of autonomous agents that work in parallel on different tasks for hours or days
weight 2 · round to CursorFirst-party marketing/docs explicitly state the exact capability: "Launch fleets of agents that work in parallel on ambitious tasks for hours or days," plus supporting evidence of background/always-on agents and agents using their own compute to build/test/demo. No independent hands-on verification of multi-day parallel fleet execution is present, and no community corroboration confirms this specific feature works at scale. Missing for 10: independent/hands-on validation of parallel agent fleets running for hours/days, details on concurrency limits or reliability over long runs.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
Docs confirm parallel task execution ('Carve out independent tasks and run them simultaneously'), cloud sessions that persist after closing the laptop, and an API to spin up multiple sessions programmatically (including on behalf of other users), which together support a 'fleet of parallel long-running agents' story. However, community hands-on reports show real friction with the 'hours/days autonomous' claim: sessions can get stuck without an escape hatch, users must babysit and manually terminate sessions every 10-15 minutes, and there's a default max-work-time cap limiting unsupervised runtime. Missing for 10: independent verification of successful multi-day/multi-task fleets running unattended, and evidence addressing the reported lack of a reliable 'ask for help' escalation during long runs.
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “Devin does ask for help when it can't do something, but it really hates asking for help if it's a skill issue - it would prefer running in c…”
- [community] “You can set a 'max work time' before Devin pauses so it won't go for days endlessly spending your credits. By default it's set to 10 credits…”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
developerRun several task attempts in parallel and compare results before choosing one
weight 1 · round to CursorCursor's docs describe launching 'fleets of agents that work in parallel on ambitious tasks for hours or days,' directly supporting parallel task execution, and agents run in isolated environments for review before merging changes. However, there is no explicit documentation of a UI/workflow for comparing multiple parallel attempts side-by-side before choosing one, and no independent/hands-on evidence corroborating this specific comparison workflow. Missing for 10: dedicated compare/diff-across-attempts feature documentation, independent verification of parallel-agent comparison in practice.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devinnone0/10Devin's docs mention running multiple independent tasks simultaneously (devin-docs-25) but this describes parallelizing different tasks, not running several parallel attempts of the SAME task to compare and choose the best result. No evidence describes a compare/choose-best-attempt workflow.
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
Scheduled automation
ai-native userSet up always-on agents that run on schedules or triggers to maintain and fix my software autonomously
weight 2 · round to CursorCursor's own site directly states the capability: "Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software," plus related background-agent features (parallel fleets, agents running on their own machines) that support this workflow. Missing for 10: independent/hands-on confirmation of scheduled/triggered agents actually running reliably in practice, and more detail on trigger types/configuration.
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin supports trigger-based autonomous work (Slack/Teams tags, PR review comments, CI/CD/static-analysis findings) and Auto-Fix creates a closed loop that iterates on CI failures without a human in the loop, and MCP exposes 'scheduling' as a session capability, all suggesting some always-on/triggered agent operation. However there's no dedicated docs for cron-like recurring schedules, and community reports describe sessions needing frequent human monitoring/termination rather than fully unattended long-running maintenance. Missing for 10: explicit scheduling/cron configuration docs, independent evidence of reliable unattended multi-day maintenance loops, and confirmation that Auto-Fix/CI triggers work without human oversight in practice.
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “You can set a 'max work time' before Devin pauses so it won't go for days endlessly spending your credits. By default it's set to 10 credits…”
Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation
Quality of generated code — correctness, style, fit to the codebase
Debugging
developerDebug a live running web application directly from my coding assistant
weight 1 · round to DevinCursornone0/10No evidence pack item describes attaching a debugger, inspecting runtime state, or interacting with a live running web app from Cursor; docs mention reproducing issues and root-causing bugs conceptually, but not live-app debugging integration (e.g., breakpoints, browser dev tools, runtime inspection). missing for 10: evidence of live debugger attach/breakpoints, browser/runtime inspection tooling, or integration with running app state.
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
Devin's computer-use/desktop mode gives it a full browser and desktop environment (mouse, keyboard, screenshots) plus VPN access to internal services, and docs explicitly mention reproducing and fixing bugs, which together support interacting with and debugging a live running app. However, there's no explicit documentation of dev-tools-style debugging features (breakpoints, console/log inspection, network tracing) or a dedicated 'live app debugging' workflow, and community evidence doesn't corroborate this specific use case. Missing for 10: explicit live-debugging tooling (breakpoints/console/log inspection), independent hands-on confirmation of debugging a running app.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Devin has access to a full desktop environment — not just a browser. It can move the mouse, click on UI elements, type on the keyboard, take…”
- [claimed-docs] “Devin can connect to a VPN from inside its workspace, so sessions can reach internal services such as package registries, databases, and int…”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
developerDebug issues and troubleshoot using natural-language queries
weight 2 · round to Devincursor-docs-3 directly claims support for reproducing issues, narrowing root cause, and verifying fixes via natural-language-driven agent workflows, and docs-1 supports tracing how a repo fits together to find bug locations. However, there's no independent/hands-on evidence corroborating debugging quality, and community evidence highlights buginess and unreliability concerns (cursor-comm-2, cursor-comm-8) that add caveats without directly contradicting the specific debugging workflow claim. Missing for 10: independent verification of debugging accuracy, concrete examples of NL-driven troubleshooting sessions, and resolution of buggy-product complaints.
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [community] “"Cursor is weird. They have a basically unused GitHub with a thousand unanswered Issues. It's so buggy in ways that VSCode isn't. I hate it.…”
- [community] “"That's a lot of money for a buggy product that is at best slightly better than its competitors."”
Docs clearly support NL-driven debugging: 'repro and fix bugs', 'Ask Devin can answer questions about code structure...help you scope and plan tasks', and tagging Devin in Slack/Teams about a bug thread. However, hands-on community reports describe practical caveats—Devin adding extraneous changes it can't undo, getting stuck without escalating, requiring manual babysitting—that temper reliability for troubleshooting workflows. Missing for 10: independent benchmark/case study specifically on debugging accuracy, and resolution of the 'getting stuck on bugs' community complaint.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Tagging Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “Ask Devin can answer questions about code structure and dependencies, and help you scope and plan tasks before implementation.”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “Devin does ask for help when it can't do something, but it really hates asking for help if it's a skill issue - it would prefer running in c…”
Feature implementation
developerTurn a tracked issue into a complete pull request end-to-end
weight 3 · round to CursorCursor's docs describe agents that trace repos, plan changes, reproduce issues, inspect diffs/run checks, and integrate with issue trackers (GitHub, Linear) and PR review, which together support a full issue-to-PR workflow (cursor-docs-1 through cursor-docs-4, cursor-docs-6, cursor-docs-8–cursor-docs-12). However, there's no explicit first-party or independent case study showing a single tracked issue being turned into a merged PR end-to-end without manual intervention, and community evidence focuses on unrelated bugs/pricing complaints rather than this workflow. Missing for 10: a concrete end-to-end example/case study of issue→PR automation and independent verification that the full pipeline works reliably.
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
- [claimed-docs] “Scope changes, use Plan Mode, and ship bigger work with confidence”
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin's docs explicitly describe taking Linear/Jira tickets and implementing full features, with Devin Review/Auto-Fix looping PRs toward merge-ready status without human involvement, covering the issue-to-PR pipeline end-to-end. However, hands-on community reports describe practical friction — extraneous unrelated changes that can break things, inability to easily undo them, and agents getting stuck for days rather than asking for help — casting doubt on how cleanly the 'complete' PR is delivered without oversight. Missing for 10: independent verification of a clean ticket→merged-PR flow without manual intervention, and resolution of the reported reliability/quality issues.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Code migrations, refactors, and modernization”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “Devin does ask for help when it can't do something, but it really hates asking for help if it's a skill issue - it would prefer running in c…”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
developerDescribe a feature or bug in plain language and have the agent implement or fix it across multiple files
weight 3 · round to CursorCursor's docs describe an agent that traces repo structure, plans and scopes multi-file changes, implements features/fixes end-to-end, runs checks, and produces diffs for review — directly matching plain-language feature/bug requests across multiple files. Community evidence corroborates the product is used daily for this purpose (albeit with complaints about bugginess), without disputing the core multi-file agentic editing capability. Missing for 10: independent hands-on benchmarks showing successful multi-file fixes, and no first-party demo/case study detailing a concrete before/after example.
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
- [claimed-docs] “Scope changes, use Plan Mode, and ship bigger work with confidence”
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
- [community] “"Cursor is weird. They have a basically unused GitHub with a thousand unanswered Issues. It's so buggy in ways that VSCode isn't. I hate it.…”
- [community] “"That's a lot of money for a buggy product that is at best slightly better than its competitors."”
Devindisputedcontradicted5/10Devin's docs strongly claim end-to-end feature/bug implementation across a full repo (Linear/Jira tickets, multi-file fixes, code migrations) with full workspace/VM access [devin-docs-1, devin-docs-13, devin-docs-21, devin-docs-27], but hands-on community reports concretely contradict smooth delivery: it 'always adds extraneous changes beyond the task that can break other things, and can't undo those changes if asked' and required constant supervision/termination to get real work done [devin-comm-1, devin-comm-5], with another user noting it can run for days without an escape hatch when stuck [devin-comm-2]. Missing for 10: independent benchmark data on multi-file correctness, and resolution of the extraneous-change/undo failure mode reported by users.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “Code migrations, refactors, and modernization”
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
Maintenance automation
developerHave the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me
weight 3 · round drawnCursor's docs describe agents that write code, run tests/checks, and 'build, maintain, and fix' software autonomously (cursor-docs-3, cursor-docs-4, cursor-docs-9, cursor-docs-12), which implies test-writing and general maintenance tasks, but there is no explicit documentation of lint-error fixing, merge-conflict resolution, or dependency-update workflows specifically. missing for 10: explicit lint-fixing examples, explicit merge-conflict-resolution examples, explicit dependency-update examples, independent hands-on verification of these specific tasks.
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Docs show Devin is a general-purpose coding agent that can fix bugs, implement features, iterate on CI failures, respond to review comments (Auto-Fix), and handle code migrations/refactors/modernization, which plausibly covers lint fixes and CI-related work, but none of the docs explicitly mention writing tests, resolving merge conflicts, or updating dependencies as named capabilities. Community reports (devin-comm-1) also note Devin can introduce extraneous changes and struggles to cleanly undo them, tempering confidence in reliably delivering these specific maintenance tasks. Missing for 10: explicit documentation/evidence of test-writing, lint-fixing, merge-conflict resolution, and dependency-update workflows, plus independent hands-on confirmation of these specific tasks succeeding.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Code migrations, refactors, and modernization”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
Multimodal generation
ai-native userGenerate a working app from a sketch, image, or PDF design
weight 2 · round drawnCursornone0/10No evidence in the pack describes image/sketch/PDF-to-app generation, multimodal design input, or any UI-from-design workflow; the docs snippets cover repo navigation, plan mode, agents, MCP, and integrations but nothing about visual design inputs.
Devinnone0/10No evidence anywhere in the pack that Devin accepts a sketch/image/PDF design as input and generates a working app from it; documentation focuses on text-based tasks, tickets, Slack threads, code review, and CLI/desktop environment features with no multimodal design-to-app capability mentioned.
Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding
How deeply the tool maps your repo — cross-file context, architecture awareness, history
Codebase mapping
developerUnderstand how a codebase fits together to find where to start making changes
weight 3 · round to DevinCursor's docs explicitly claim the ability to 'trace how a repo fits together and find the right places to start,' directly matching the story, but this is a single marketing-style doc line with no detailed walkthrough, feature docs (e.g., codebase indexing/@codebase chat), or independent corroboration of how it actually surfaces architecture understanding. Missing for 10: detailed documentation of the codebase-mapping/indexing feature itself, concrete examples of it locating relevant code, and independent/hands-on validation of accuracy.
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
Devin explicitly indexes repositories to power 'Ask Devin' and 'DeepWiki', which answer questions about code structure and dependencies and help developers scope/plan where to start making changes, directly matching the story. missing for 10: independent/hands-on evidence validating DeepWiki/Ask Devin's accuracy on real codebases (community evidence only covers task execution, not codebase-understanding features).
- [claimed-docs] “Use DeepWiki to navigate architecture and code with auto-generated documentation.”
- [claimed-docs] “Ask Devin can answer questions about code structure and dependencies, and help you scope and plan tasks before implementation.”
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
developerHave the agent map and explain an entire unfamiliar codebase without manually selecting context files
weight 3 · round to DevinFirst-party docs claim Cursor can 'trace how a repo fits together and find the right places to start' (cursor-docs-1), implying automatic codebase mapping, but there's no detail on how context is auto-gathered (e.g., codebase indexing/@codebase) nor any independent/hands-on confirmation that it explains an unfamiliar codebase without manual file selection. Missing for 10: technical explanation of automatic context retrieval, independent user validation of whole-codebase explanation, and comparison to manual context selection workflows.
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
Docs describe repo indexing that lets Devin understand the codebase and power 'Ask Devin' and DeepWiki auto-generated architecture docs, directly enabling exploration/explanation of an unfamiliar codebase without manual file selection (devin-docs-27, devin-docs-15, devin-docs-16). No community evidence contradicts this specific capability. Missing for 10: independent/hands-on verification of codebase-mapping accuracy and no concrete example of DeepWiki output quality.
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
- [claimed-docs] “Use DeepWiki to navigate architecture and code with auto-generated documentation.”
- [claimed-docs] “Ask Devin can answer questions about code structure and dependencies, and help you scope and plan tasks before implementation.”
Context management
developerHave the agent build and recall memory automatically across sessions
weight 2 · round to DevinCursornone0/10No evidence describes persistent memory that automatically builds and recalls context across sessions; docs mention repo tracing, plan mode, and MCP integrations but nothing about cross-session memory recall.
Devin supports persistent cross-session context via "Knowledge" (instructions referenced in all sessions), AGENTS.md, and environment blueprints that let every session boot into a known-good state, which enables some recall across sessions. However, these mechanisms are largely user-configured/onboarded rather than autonomously built by the agent from its own experience, and there's no evidence of automatic memory creation or recall behavior demonstrated in practice. Missing for 10: evidence the agent automatically extracts/updates memory from its own task experience without manual setup, and independent confirmation that recalled memory improves subsequent session performance.
- [claimed-docs] “Knowledge is a collection of instructions and advice that Devin can reference in all sessions.”
- [claimed-docs] “Devin supports AGENTS.md - a simple, open standard for providing context and instructions to AI agents.”
- [claimed-docs] “Configure it once, and every session boots into that known-good state.”
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
developerInclude multiple project directories in a single session for broader context
weight 2 · round drawnCursornone0/10No evidence in the pack mentions multi-root workspaces or including multiple project directories in a single Cursor session; docs cover repo navigation, MCP, agents, and integrations but not multi-directory context.
developerAdd a project instructions file to set coding standards and conventions the agent follows
weight 3 · round to DevinCursor's docs mention adding 'rules' as one of its features (alongside plugins, skills, MCPs) which aligns with the project-instructions concept, but the evidence pack gives no detail on how project rule files work, their scope, or how the agent applies them to enforce coding standards. missing for 10: documentation of the rules file format/location, examples of coding standards enforcement, independent confirmation the agent actually follows these instructions consistently.
- [claimed-docs] “Add plugins, skills, MCPs, and rules from one place”
Devin explicitly supports AGENTS.md, an open standard for providing context and instructions to AI agents, plus a separate 'Knowledge' feature for instructions/advice referenced across all sessions, directly covering project-level coding standards/conventions. Missing for 10: independent/hands-on confirmation that these instructions are reliably followed in practice.
- [claimed-docs] “Devin supports AGENTS.md - a simple, open standard for providing context and instructions to AI agents.”
- [claimed-docs] “Knowledge is a collection of instructions and advice that Devin can reference in all sessions.”
Issue diagnosis
developerReproduce issues, narrow down root causes, and verify fixes
weight 3 · round drawncursor-docs-3 directly claims the exact capability ('Reproduce issues, narrow the root cause, and verify the fix'), and supporting docs on codebase tracing, diffs/checks, and agents running their own environments (cursor-docs-1, cursor-docs-4, cursor-docs-12) plausibly back this workflow. However, this is a first-party marketing/docs claim only, with no independent or hands-on corroboration of actual debugging workflows, and community evidence highlights general bugginess/quality concerns rather than validating this specific capability. Missing for 10: independent verification or hands-on case studies of reproduce/root-cause/verify-fix workflows, more detail on how reproduction (e.g., test running, log inspection) is concretely supported.
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Docs explicitly claim Devin can 'repro and fix bugs' (devin-docs-1), use Ask Devin/DeepWiki to narrow root causes via codebase understanding (devin-docs-15, devin-docs-16, devin-docs-27), and verify fixes through CI iteration/Auto-Fix loops (devin-docs-18). However, hands-on community testimony reports Devin often adds extraneous changes beyond the task scope and cannot reliably undo them when asked, undermining clean verification of fixes (devin-comm-1), and lacks an escape hatch when stuck on root-cause diagnosis (devin-comm-2, devin-comm-3). Missing for 10: independent verification of successful bug reproduction/root-cause narrowing at scale, and resolution of the reported inability to cleanly revert unwanted changes during fix verification.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Use DeepWiki to navigate architecture and code with auto-generated documentation.”
- [claimed-docs] “Ask Devin can answer questions about code structure and dependencies, and help you scope and plan tasks before implementation.”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “One thing that surprised me is there doesn't seem to be an 'ask for help' escape hatch in Devin - it would work away for literally days on a…”
- [community] “Devin does ask for help when it can't do something, but it really hates asking for help if it's a skill issue - it would prefer running in c…”
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem
Integrations, plugins, and third-party ecosystem stories
Marketplace
developerEquip the agent with custom skills to perform specialized tasks
weight 1 · round drawnCursor's docs mention a marketplace to 'Add plugins, skills, MCPs, and rules from one place' and detailed MCP support (custom servers, marketplace install, enterprise controls), enabling developers to extend the agent with specialized tool integrations. However, there's no dedicated documentation on a 'skills' framework distinct from MCP/rules, no examples of custom skill creation workflow, and no independent/community corroboration of this specific capability. Missing for 10: detailed skills documentation/tutorial, examples of custom skill authoring, independent hands-on validation.
- [claimed-docs] “Add plugins, skills, MCPs, and rules from one place”
- [claimed-docs] “Model Context Protocol (MCP) enables Cursor to connect to external tools and data sources.”
- [claimed-docs] “Click "Add to Cursor" on a marketplace entry to install it and authenticate with OAuth.”
- [claimed-docs] “Configure custom MCP servers with a JSON file”
- [claimed-docs] “Enterprise admins can control which MCP servers users may run from the Cursor dashboard.”
Devin supports several skill-like extension mechanisms — 'Knowledge' (persistent instructions/advice for all sessions), AGENTS.md for structured agent instructions, and 'Playbooks' exposed via MCP — plus an open-source 'Devin Handoff' explicitly described as a 'plugin and skill'. This gives developers real levers to encode specialized, reusable task behavior, but the docs don't show a dedicated skill-authoring UI/marketplace or detailed examples of building a complex custom skill, and there is no independent/community evidence confirming this works well in practice. Missing for 10: concrete examples/tutorials of authoring a non-trivial custom skill, a discoverable skills registry, and independent hands-on corroboration.
- [claimed-docs] “Knowledge is a collection of instructions and advice that Devin can reference in all sessions.”
- [claimed-docs] “Devin supports AGENTS.md - a simple, open standard for providing context and instructions to AI agents.”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “Devin Handoff is an open-source plugin and skill that brings the same handoff workflow to any coding agent — Claude Code, Codex, Cursor, and…”
engineering-leadIntegrate third-party partner-built agent apps into my workflows
weight 1 · round to CursorCursor documents a marketplace for adding third-party plugins, skills, and MCP servers with OAuth authentication, plus native integrations with GitHub, GitLab, Slack, Linear, and more, letting teams plug partner-built tools/agents into their workflows, with enterprise admin controls over which servers are allowed. Missing for 10: independent/hands-on corroboration of using specific partner-built agent apps (vs. generic tool connectors) and clearer distinction between simple MCP data-tools and full third-party 'agent apps'.
- [claimed-docs] “Add plugins, skills, MCPs, and rules from one place”
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
- [claimed-docs] “Model Context Protocol (MCP) enables Cursor to connect to external tools and data sources.”
- [claimed-docs] “Click "Add to Cursor" on a marketplace entry to install it and authenticate with OAuth.”
- [claimed-docs] “Configure custom MCP servers with a JSON file”
- [claimed-docs] “Enterprise admins can control which MCP servers users may run from the Cursor dashboard.”
Devin exposes an MCP server so external MCP-compatible agents/IDEs can access its session management and tools, and its open-source 'Devin Handoff' plugin interoperates with other coding agents (Claude Code, Codex, Cursor), showing some cross-agent workflow integration. However there is no evidence of a partner/marketplace ecosystem of third-party agent apps being integrated into Devin's own workflows. Missing for 10: a documented partner-app marketplace or catalog, evidence of installing/configuring third-party agent apps within Devin, and any case study of an engineering-lead orchestrating partner-built agents through Devin.
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “Devin Handoff is an open-source plugin and skill that brings the same handoff workflow to any coding agent — Claude Code, Codex, Cursor, and…”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [probe] “official MCP server documented at https://docs.devin.ai/work-with-devin/devin-mcp”
Team knowledge
engineering-leadCreate a shared workspace from my docs and repos as a common source of truth for the team
weight 1 · round to DevinCursornone0/10Evidence shows integrations (GitHub, Slack, Linear), MCP/plugins, and rules configuration, but nothing describes a dedicated 'shared workspace' feature that unifies docs and repos into a common team source of truth — this is a fair ask for a team-oriented dev tool but unaddressed in the pack.
Devin offers building blocks for a team-wide source of truth — repo indexing that powers 'Ask Devin' and DeepWiki architecture docs, org-wide 'Knowledge' referenced in all sessions, and AGENTS.md support for shared context — which collectively let a lead centralize docs/repo knowledge for the team. However, there's no explicit product feature framed as a 'shared workspace' UI for team-wide docs/repo browsing outside of Devin's own agent sessions. Missing for 10: a dedicated shared workspace/knowledge-base product surface for humans to browse, independent evidence of teams using it as a collaborative source of truth, and clarity on cross-repo doc aggregation beyond per-session knowledge.
- [claimed-docs] “Indexing your repositories allows Devin to understand your codebase and enables powerful features like Ask Devin and DeepWiki”
- [claimed-docs] “Knowledge is a collection of instructions and advice that Devin can reference in all sessions.”
- [claimed-docs] “Devin supports AGENTS.md - a simple, open standard for providing context and instructions to AI agents.”
- [claimed-docs] “Use DeepWiki to navigate architecture and code with auto-generated documentation.”
- [claimed-docs] “Devin inspects your repository, figures out which tools, runtimes, and dependencies are needed, and generates the blueprint for you.”
Tool integration
developerConnect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context
weight 3 · round to CursorCursor documents MCP support that connects to external tools/data sources, an MCP marketplace with OAuth install, and explicit integration with Slack alongside GitHub/GitLab/Linear/Jira-style trackers, plus Slack-based agent collaboration—covering the story's workflow-tool extension use case. Missing for 10: explicit first-party Jira/Google Drive connector documentation and independent hands-on verification of these integrations working end-to-end.
- [claimed-docs] “Model Context Protocol (MCP) enables Cursor to connect to external tools and data sources.”
- [claimed-docs] “Click "Add to Cursor" on a marketplace entry to install it and authenticate with OAuth.”
- [claimed-docs] “Configure custom MCP servers with a JSON file”
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
Devin can be tagged in Slack/Teams threads and work Jira/Linear tickets, giving it direct workflow-tool integration, and its MCP server plus API enable further extension to other tools. However there is no explicit documented Google Drive integration, and no independent/hands-on corroboration of these integrations actually working in practice. missing for 10: Google Drive connector evidence, independent verification of Jira/Slack integration reliability.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Tagging Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
developerKick off agent tasks directly from GitHub, GitLab, Linear, or Slack
weight 2 · round to CursorCursor's docs explicitly list integrations with GitHub, GitLab, Slack, and Linear, and describe agents that run on triggers/schedules and collaborate in Slack or review PRs in GitHub, supporting the story's core claim. However, there's no detailed first-party documentation of the exact trigger mechanics per platform (e.g., a Linear ticket auto-spawning an agent) nor independent/hands-on confirmation that this works reliably. Missing for 10: platform-specific trigger documentation for each of GitHub/GitLab/Linear/Slack, and independent verification of the workflow in practice.
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
Docs confirm Devin can be invoked from Linear/Jira tickets and Slack/Teams threads, plus API integration for building custom workflow triggers, but there's no explicit mention of GitHub or GitLab issue/PR-based task kickoff in the evidence pack. Community evidence doesn't directly contradict the integration claims, only general effectiveness concerns. Missing for 10: explicit GitHub/GitLab-triggered task creation documentation, independent hands-on confirmation of these specific integrations working.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Tagging Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration
Meeting you in the IDE and terminal — extensions, inline flows, context
Cross device continuity
developerStart a task on one device and continue it later from another device or browser
weight 2 · round to DevinCursor's Background Agents run remotely and can be monitored/interacted with via terminal, Slack, and GitHub PRs, implying a task could be checked or continued from different surfaces, but there is no explicit documentation of resuming a specific in-progress task from a different device or browser session. Missing for 10: explicit cross-device/browser session handoff documentation, hands-on confirmation of resuming a task started elsewhere, and details on state syncing across clients.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin's cloud sessions run in a dedicated VM that persists independent of the local device (docs-13), can be started from Slack/Teams, IDE, CLI, or API and continued/taken over in the embedded IDE or web UI (docs-3, docs-4, docs-7, docs-24), and Devin Handoff explicitly lets you start work locally and continue in a cloud session accessible from a browser (docs-38). missing for 10: no explicit first-party walkthrough of resuming the same session from a different browser/device login, and no independent/community confirmation of cross-device continuity.
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “Devin CLI, a local command-line coding agent with deep Devin Cloud integration.”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “Tag Devin on a Slack or Teams thread about a bug you're discussing with coworkers”
- [claimed-docs] “Hand a task off to a cloud Devin session and keep working locally.”
Ide integration
developerChat with the coding assistant directly inside my IDE for contextual help
weight 3 · round to CursorCursor's docs describe an IDE-integrated assistant that traces repo structure, scopes changes via Plan Mode, reproduces issues, and hands off tasks while the developer reviews — all consistent with in-IDE contextual chat, and community commentary confirms it functions as a VS Code-based assistant with prompts/harness. missing for 10: no explicit citation naming a dedicated 'chat panel' UI or independent praise of chat quality/context-awareness specifically.
- [claimed-docs] “Trace how a repo fits together and find the right places to start”
- [claimed-docs] “Scope changes, use Plan Mode, and ship bigger work with confidence”
- [claimed-docs] “Reproduce issues, narrow the root cause, and verify the fix”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
- [community] “"Cursor is an extension for VS Code, a harness and a bunch of prompts. They have their own model (Composer 2) which is based on Kimi K2.5, b…”
Devin exposes a conversational interface via its own embedded IDE within cloud sessions (devin-docs-3) and its Desktop app that imports VS Code/Cursor settings (devin-docs-9), plus an MCP server letting 'any MCP-compatible AI agent or IDE' access sessions (devin-docs-6/23). However there's no evidence of a native extension that lets a developer chat with Devin directly inside their own existing IDE (e.g., a VS Code/JetBrains plugin) — Devin's model is its own IDE/Desktop environment or MCP bridging rather than embedding in the user's IDE. Missing for 10: a first-party IDE extension for VS Code/JetBrains enabling in-IDE chat, and independent confirmation of this workflow working well in practice.
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “Import VS Code or Cursor settings, configure themes, and start coding with AI-powered assistance.”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
Session management
developerReview diffs visually and run multiple sessions side by side in a desktop app
weight 2 · round to CursorCursor's docs explicitly describe inspecting diffs before merge and launching fleets of agents to work in parallel, both core to a desktop IDE experience with visual diff review and concurrent sessions. Missing for 10: independent/hands-on confirmation of the side-by-side session UI and a detailed walkthrough of the diff viewer beyond marketing copy.
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin has a documented Desktop app (docs-9) and supports running multiple independent sessions in parallel (docs-25), but there is no evidence of a visual diff review feature or explicit side-by-side session UI within the desktop app itself. Missing for 10: explicit documentation of an in-app diff viewer, UI showing multiple sessions simultaneously in one window, and any hands-on/community confirmation of this desktop workflow.
- [claimed-docs] “Import VS Code or Cursor settings, configure themes, and start coding with AI-powered assistance.”
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
engineering-leadManage multiple agent-driven coding sessions from one unified workspace
weight 2 · round to CursorCursor's docs explicitly describe launching 'fleets of agents that work in parallel on ambitious tasks for hours or days' and setting up always-on agents on schedules/triggers, all accessible from Cursor's interface spanning terminal, Slack, and GitHub — directly matching a unified multi-session agent workspace for a lead overseeing parallel work. Missing for 10: independent/hands-on corroboration of the multi-agent dashboard UX, and no detail on cross-session visibility/coordination features specifically framed for engineering-lead oversight.
- [claimed-docs] “Launch fleets of agents that work in parallel on ambitious tasks for hours or days.”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Accelerate development by handing off tasks to Cursor, while you focus on making decisions.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin supports running multiple parallel sessions ('carve out independent tasks and run them simultaneously'), an API to create sessions on behalf of users, and an embedded IDE/CLI/desktop app to interact with sessions, which together enable a lead-like workspace for managing several agent sessions. However, there's no dedicated 'unified workspace' dashboard evidence for an engineering-lead specifically monitoring/managing a team's multiple concurrent sessions, and community feedback highlights session reliability issues that would complicate multi-session oversight. missing for 10: explicit multi-session dashboard/UI for a lead role, team-level session oversight features, independent corroboration of smooth multi-session management at scale.
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “Devin CLI, a local command-line coding agent with deep Devin Cloud integration.”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
- [community] “I've learned to just click 'Terminate Session' immediately after spotting Devin doing something hopeless. I've managed to get real work done…”
Terminal workflow
developerRun a coding agent locally from my terminal
weight 3 · round to DevinCursor ships an official CLI (cursor.com/cli) with a documented install command (curl ... | bash) and docs explicitly state 'Cursor runs in your terminal', confirming a local terminal-based agent capability alongside its IDE. Missing for 10: independent/hands-on verification of terminal agent usage and deeper CLI usage documentation beyond the install step.
- [probe] “official CLI documented at https://cursor.com/cli”
- [claimed-docs] “curl https://cursor.com/install -fsS | bash”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
Devin CLI is explicitly documented as a local command-line coding agent that can be invoked from terminal (e.g. `devin -- check out this code...`), with local sandboxing (--sandbox flag) and deep integration with Devin Cloud for handoff. Missing for 10: independent/hands-on community verification specifically of the CLI experience (community evidence only covers the cloud/browser Devin product, not the local CLI).
- [claimed-docs] “Devin CLI, a local command-line coding agent with deep Devin Cloud integration.”
- [claimed-docs] “devin -- check out this code and suggest a feasible, helpful feature”
- [claimed-docs] “a local command-line coding agent with deep Devin Cloud integration”
- [claimed-docs] “The --sandbox flag runs the CLI with OS-level isolation, enforcing writable paths and deny rules at the operating-system level and optionall…”
- [claimed-docs] “The `--sandbox` flag runs the CLI with OS-level isolation, enforcing writable paths and `deny` rules at the operating-system level”
- [probe] “official CLI documented at https://docs.devin.ai/cli/index”
developerRun the agent non-interactively in scripts for workflow automation
weight 2 · round to DevinCursor ships an official CLI (cursor-probe-1, cursor-docs-14) and documents 'always-on agents that run on schedules or triggers to build, maintain, and fix your software' (cursor-docs-9), which implies non-interactive/automated agent execution suitable for scripts/CI. However, there is no concrete documentation of CLI flags, headless/print modes, exit codes, or scripting examples, nor independent hands-on confirmation of this workflow. Missing for 10: explicit CLI non-interactive flag/usage docs, examples of piping/scripting the agent, and independent verification that scheduled/triggered agents work as scripted automation.
- [probe] “official CLI documented at https://cursor.com/cli”
- [claimed-docs] “curl https://cursor.com/install -fsS | bash”
- [claimed-docs] “Set up always-on agents that run on schedules or triggers to build, maintain, and fix your software.”
Devin exposes a documented API for creating sessions programmatically (including on behalf of users) and explicit CI/CD pipeline integration to auto-respond to static-analysis findings, plus a CLI (`devin -- <prompt>`) that can be invoked headlessly, all pointing to non-interactive, scriptable automation. Missing for 10: independent/hands-on confirmation of headless CLI scripting in real CI pipelines and more detail on CLI exit codes/output for scripting.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Devin CLI, a local command-line coding agent with deep Devin Cloud integration.”
- [claimed-docs] “devin -- check out this code and suggest a feasible, helpful feature”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “a local command-line coding agent with deep Devin Cloud integration”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to DevinCursornone0/10The evidence pack shows no public API for Cursor; it mentions an official CLI and MCP (for connecting external tools INTO Cursor), but nothing about a programmatic interface exposing Cursor's own UI capabilities (agents, plan mode, review, etc.) for external control.
Devin exposes a documented API (session creation, create_as_user_id) and an MCP server giving 'full access to session management, playbooks, knowledge, and scheduling', suggesting broad programmatic parity, but there's no explicit claim of full UI/API feature parity, and probes for an OpenAPI spec all 404'd, indicating the full API surface isn't transparently documented. Missing for 10: explicit parity statement covering UI-only features like Devin Review/Auto-Fix/Computer Use/desktop app settings, a discoverable OpenAPI schema, and independent confirmation that all UI actions are API-reachable.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnCursornone0/10No evidence in the pack addresses data export, portability, or open-format data extraction from Cursor; being built on VS Code (a fork) doesn't itself guarantee user data/settings export tooling is documented anywhere here.
ai-native userRead the product's source under an open license
weight 2 · round drawnCursornone0/10Cursor is known to be closed-source (a proprietary VS Code fork), and the evidence pack contains no mention of any open-source license or public repository for Cursor's source code; all references are to product docs, features, or community complaints unrelated to licensing.
Devinnone0/10Devin is a closed, proprietary commercial product; no evidence indicates its source code is available under an open license. The only mention of 'open-source' is for the separate Devin Handoff plugin/skill, not Devin itself.
- [claimed-docs] “Devin Handoff is an open-source plugin and skill that brings the same handoff workflow to any coding agent — Claude Code, Codex, Cursor, and…”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Authentication
developerAuthenticate with an API key instead of an account login
weight 2 · round to DevinCursornone0/10No evidence in the pack describes API key authentication as an alternative to account login for Cursor; docs reference OAuth-based MCP authentication and account-based login flows only.
Devin has a documented API (devin-docs-7) intended for programmatic integration, implying API key auth as an alternative to account login, and even supports creating sessions on behalf of other users (devin-docs-8), suggesting a service-level credential model. However, no explicit documentation of API key generation/management or authentication mechanics is present in the evidence pack. Missing for 10: explicit API key creation/management docs, confirmation that API key auth fully replaces login flows, and independent/community verification of this workflow.
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
- [claimed-docs] “you can create sessions on behalf of any user in your organization using the create_as_user_id parameter”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.devin.ai/openapi.json, https://docs.devin.ai/swagger.json, https://docs.devin.ai/api/op…”
engineering-leadAuthenticate through an enterprise identity or cloud platform for compliance and scalability
weight 2 · round drawnCursornone0/10The evidence pack shows Cursor has enterprise admin controls for MCP servers but contains no evidence of SSO/SAML integration, enterprise identity provider authentication (e.g., Okta, Azure AD, Google Workspace), or cloud platform login for compliance purposes. This is a fair and applicable axis for a widely-adopted dev tool sold to enterprises, so absence of evidence yields 'none' rather than 'na'.
- [claimed-docs] “Enterprise admins can control which MCP servers users may run from the Cursor dashboard.”
developerSign in with my existing product subscription plan to use the coding agent
weight 2 · round drawnCursornone0/10The evidence pack describes Cursor's agent features, MCP integrations, and installation steps, but contains no documentation or confirmation that users authenticate with an existing subscription plan to access the coding agent (only tangential community chatter about login policy hallucinations). Missing for 10: explicit account/subscription sign-in flow docs, plan-tier access confirmation, and any first-party statement linking subscription plan to agent usage.
- [community] “Cursor's AI support agent hallucinated a single-device login policy, telling a user this was intentional. A Cursor developer later clarified…”
Devinnone0/10The evidence describes Devin's own subscription tiers (Pro, Max, Core/Team/Enterprise) but nothing about letting a developer sign in using an existing external subscription (e.g., an existing LLM provider or IDE subscription) to access the agent. No mention of SSO-linked billing or bring-your-own-subscription support.
- [claimed-docs] “Max is for individual users who consistently exceed the Pro quota. It includes everything in Pro, plus a significantly larger weekly usage q…”
- [claimed-docs] “Power users who need more quota”
developerSign in with a personal account to get free-tier access without managing API keys
weight 1 · round drawnCursornone0/10The evidence pack contains no documentation or first-party description of a sign-in flow that grants free-tier access without requiring API keys; only tangential community mentions of account workarounds for usage limits exist. Missing for 10: any docs on account creation/sign-in, free-tier terms, or explicit no-API-key requirement.
- [community] “Cursor is caught in a cat-and-mouse game against workarounds where users create new accounts to get unlimited use; a repo enabling this (cur…”
Devinnone0/10Evidence only describes paid tiers (Pro, Max, 'power users who need more quota') and API key based integration; there's no mention of a free tier accessible via personal account sign-in without API key management.
- [claimed-docs] “Max is for individual users who consistently exceed the Pro quota. It includes everything in Pro, plus a significantly larger weekly usage q…”
- [claimed-docs] “Power users who need more quota”
- [claimed-docs] “The Devin API enables you to integrate Devin into your applications, automate workflows, and build powerful tools.”
Model choice
developerLet the tool automatically pick the best model for each task
weight 1 · round drawnCursornone0/10The evidence shows Cursor lets developers manually choose among multiple models (OpenAI, Anthropic, Gemini, etc.) but nothing indicates an automatic 'best model for the task' selection feature. missing for 10: any documentation or claim of an auto-select/router feature that picks models per task, evidence of cost/performance-based automatic routing.
- [claimed-docs] “Choose between every cutting-edge model from OpenAI, Anthropic, Gemini, SpaceXAI, and Cursor.”
developerChoose which underlying AI model powers my session from multiple providers
weight 2 · round to Cursorcursor-docs-7 confirms Cursor lets developers choose between models from multiple providers (OpenAI, Anthropic, Gemini, and Cursor's own), directly matching the story. Missing for 10: independent hands-on verification of per-session model switching UI/behavior and pricing implications tied to model choice.
- [claimed-docs] “Choose between every cutting-edge model from OpenAI, Anthropic, Gemini, SpaceXAI, and Cursor.”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnCursornone0/10No evidence in the pack mentions data residency, region selection, or storage location controls for Cursor; the docs snippets cover agents, MCP, and integrations but nothing about choosing data storage region. Missing for 10: any mention of regional data residency options, enterprise data location controls, or compliance documentation addressing storage jurisdiction.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnCursornone0/10The evidence pack contains no documentation of a privacy mode, opt-out of training, or data-retention controls for Cursor; all cited docs cover unrelated features (agents, MCP, integrations) and community threads are unrelated to training-data privacy.
ai-native userControl data retention and deletion
weight 2 · round drawnCursornone0/10The evidence pack contains no documentation of data retention settings, deletion controls, privacy dashboard, or data handling policies for Cursor; only unrelated docs on features (MCP, agents, integrations) and community complaints about bugs/pricing are present. Missing for 10: any first-party privacy policy docs, retention period settings, data deletion request mechanism, or enterprise data controls.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnCursornone0/10The evidence pack contains no mention of telemetry settings, privacy controls, or usage-tracking opt-out mechanisms; docs only cover unrelated features like MCP, agents, and integrations. Missing for 10: any privacy policy or settings documentation, telemetry opt-out toggle, or usage data collection disclosure.
Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety
Keeping generated changes safe — diffs, approvals, guardrails
Data governance
engineering-leadOpt out of having my code and prompts used for AI model training
weight 1 · round drawnCursornone0/10The evidence pack contains no mention of privacy settings, opt-out of training, or data usage policies for Cursor; all docs entries relate to unrelated features (agents, MCP, integrations) and community items focus on bugs/pricing/model sourcing, not training data controls.
Pr review
developerHave the agent stage changes, write commit messages, create branches, and open pull requests
weight 3 · round to DevinDocs show GitHub/GitLab integration and agents that build/test/demo work end-to-end for review (cursor-docs-6, cursor-docs-10, cursor-docs-12), implying some git-workflow automation, but there's no explicit documentation of the agent staging changes, writing commit messages, creating branches, or opening pull requests. missing for 10: explicit commit-message generation, branch creation, PR-opening workflow documentation, and any hands-on confirmation these steps work end-to-end.
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
Devin's docs describe it implementing features/fixing bugs and producing pull requests that get automated review and iteration (devin-docs-17, devin-docs-18), implying it handles the full git workflow (branch, commit, PR) autonomously, but no doc explicitly details staging, commit-message generation, or branch creation as discrete steps. Community feedback (devin-comm-1) also notes it can add extraneous changes it can't cleanly undo, a real caveat on commit hygiene. Missing for 10: explicit documentation of commit/staging/branch mechanics and independent confirmation that generated commits/PRs are clean and reviewable.
- [claimed-docs] “Ask Devin to tackle Linear/Jira tickets, implement entirely new features, repro and fix bugs, build internal tools, and more!”
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Carve out independent tasks and run them simultaneously.”
- [community] “Trialed Devin, it's quite impressive when it understands code formatting and local test setup, but it always adds extraneous changes beyond …”
developerGet automatic code review with contextual feedback on every pull request
weight 3 · round to DevinCursor's docs explicitly claim it 'reviews PRs in GitHub' and can 'inspect diffs, run checks, and catch problems before you merge,' directly matching automated PR review with contextual feedback, backed by GitHub/GitLab/Bitbucket integration claims. missing for 10: independent/hands-on verification of review quality, details on triggering on every PR automatically, and no community corroboration of this specific feature.
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
Devin Review is explicitly documented as an automated first-pass PR reviewer checking correctness and org best-practice conformance, with Auto-Fix closing the loop by responding to review comments and CI failures. This directly matches the story of automatic contextual code review on PRs, though evidence is vendor-documentation only with no independent hands-on validation of review quality/contextual accuracy. Missing for 10: independent/community corroboration of review quality, and detail on how 'contextual feedback' is surfaced per-PR beyond docs description.
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Enable Devin Review with Auto-Fix so Devin automatically responds to code review comments, fixe”
developerInspect diffs and run checks to catch problems before merging
weight 3 · round to Devincursor-docs-4 explicitly claims the capability ('Inspect diffs, run checks, and catch problems before you merge') and cursor-docs-10/12 support a broader PR review workflow, but there is no independent or hands-on corroboration of diff inspection or check-running in practice, and community evidence focuses on unrelated bugs/pricing rather than this feature. missing for 10: independent verification of diff review UI, details on what 'checks' run (tests/linters/CI), and hands-on confirmation of pre-merge workflow.
- [claimed-docs] “Inspect diffs, run checks, and catch problems before you merge”
- [claimed-docs] “Cursor runs in your terminal, collaborates in Slack, and reviews PRs in GitHub.”
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin's docs describe Devin Review, which performs automated first-pass PR reviews checking correctness and org conformance, plus CI/CD integration to respond to static-analysis findings (SonarQube, Fortify, Veracode) and Auto-Fix iterating on CI failures — directly enabling developers to inspect diffs and run checks before merge. This is corroborated by explicit SDLC integration workflow docs, not just a single mention. missing for 10: independent/hands-on evidence confirming Devin Review's diff-inspection quality in practice, and detail on how diffs are surfaced/inspected by the developer (UI specifics) beyond docs claims.
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Enable Devin Review with Auto-Fix so Devin automatically responds to code review comments, fixe”
Safe execution
engineering-leadControl which external tools and integrations the agent is allowed to access
weight 2 · round to CursorDocs show enterprise admins can restrict which MCP servers users may run from the Cursor dashboard, and users can toggle individual servers on/off, giving engineering leads direct control over external tool/integration access. Missing for 10: independent/hands-on corroboration of the admin dashboard controls and finer-grained per-tool permission examples beyond MCP servers.
- [claimed-docs] “Enterprise admins can control which MCP servers users may run from the Cursor dashboard.”
- [claimed-docs] “Toggle servers on/off without removing them”
- [claimed-docs] “Model Context Protocol (MCP) enables Cursor to connect to external tools and data sources.”
- [claimed-docs] “Configure custom MCP servers with a JSON file”
Devin exposes some admin-level controls over its access—an org-wide 'Enable desktop mode' toggle for Computer Use, a CLI --sandbox flag enforcing OS-level writable-path/deny rules and network restriction, and Outposts for running sessions in infra you control—giving leads levers to constrain what Devin can reach or do. However there's no documented centralized policy/allowlist for specific external integrations (e.g., disabling Slack, Jira, GitHub, VPN access per-tool) or granular permission/audit management for engineering leads. missing for 10: a unified integration-permission/allowlist admin panel, per-tool enable/disable controls beyond desktop mode, and independent verification of these controls in practice.
- [claimed-docs] “Computer Use is controlled by the Enable desktop mode toggle in your organization's customization options.”
- [claimed-docs] “The --sandbox flag runs the CLI with OS-level isolation, enforcing writable paths and deny rules at the operating-system level and optionall…”
- [claimed-docs] “The `--sandbox` flag runs the CLI with OS-level isolation, enforcing writable paths and `deny` rules at the operating-system level”
- [claimed-docs] “Outposts lets you run Devin sessions inside infrastructure you control — your own VMs, containers, Kubernetes clusters, or even a Mac Mini o…”
- [claimed-docs] “Devin can connect to a VPN from inside its workspace, so sessions can reach internal services such as package registries, databases, and int…”
engineering-leadHave the agent operate inside a sandbox when interacting with code, tools, and network resources
weight 2 · round to DevinCursornone0/10The evidence pack contains no mention of sandboxing, isolated execution environments, or network/tool restriction controls for the agent; docs describe agents using 'their own computers' but give no detail on containment/sandboxing mechanisms. Missing for 10: any documentation of a sandbox/isolation feature, network egress controls, or filesystem restriction for agent actions.
- [claimed-docs] “Agents use their own computers to build, test, and demo features end to end for you to review.”
Devin runs cloud sessions in isolated VMs (devin-docs-13) and provides an explicit CLI --sandbox flag enforcing OS-level isolation, writable path restrictions, deny rules, and optional network restriction (devin-docs-12, devin-docs-37), directly matching the sandboxed code/tool/network isolation story. Missing for 10: independent/hands-on verification of sandbox robustness and more detail on network isolation guarantees beyond docs claims.
- [claimed-docs] “The --sandbox flag runs the CLI with OS-level isolation, enforcing writable paths and deny rules at the operating-system level and optionall…”
- [claimed-docs] “The `--sandbox` flag runs the CLI with OS-level isolation, enforcing writable paths and `deny` rules at the operating-system level”
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “Devin can connect to a VPN from inside its workspace, so sessions can reach internal services such as package registries, databases, and int…”
Security checks
engineering-leadSee license and public-code matching references for AI-suggested code
weight 1 · round drawnCursornone0/10No evidence anywhere in the pack mentions license detection, public code matching, provenance references, or IP attribution for AI-suggested code; docs focus on repo navigation, diffs, agents, and integrations, none of which addresses license/code-match transparency.
developerGet contextual explanations and automatic fixes for security vulnerabilities
weight 2 · round to DevinCursornone0/10The evidence pack shows general code review/diff-inspection features (cursor-docs-4) and broad agent capabilities, but nothing specifically documents contextual security vulnerability explanations or automated security fixes. Missing for 10: any mention of vulnerability detection, security scanning integration, or CVE/security-specific fix suggestions.
Devin's docs describe Devin Review giving automated PR reviews with explanations for correctness/best-practice issues, Auto-Fix automatically responding to review comments and fixing flagged bugs/CI failures, and CI/CD integration to respond to findings from security scanners like SonarQube, Fortify, and Veracode — directly matching contextual explanation plus automatic fixing of vulnerabilities. Missing for 10: independent/hands-on evidence specifically validating security-vulnerability fixes (community evidence only discusses general reliability/scope-creep issues, not security-fix accuracy).
- [claimed-docs] “Devin Review provides automated first-pass reviews on pull requests, checking for correctness and conformance with organizational best pract…”
- [claimed-docs] “With Auto-Fix enabled, Devin automatically responds to code review comments, fixes flagged bugs, and iterates on CI failures — creating a cl…”
- [claimed-docs] “Integrate Devin into your CI/CD pipeline to respond to findings from static analysis tools like SonarQube, Fortify, or Veracode.”
- [claimed-docs] “Ask Devin can answer questions about code structure and dependencies, and help you scope and plan tasks before implementation.”
Not comparable on these axes
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableCursorn/aCursor is itself an AI coding agent; the evidence (cursor-docs-15 to cursor-docs-19) shows Cursor acting as an MCP client that connects to external MCP servers, not Cursor exposing an official MCP server for other agents to connect to. Per the agent-role exception, client-side MCP support does not make this server-side story applicable.
Devin ships an official documented MCP server (devin-mcp) that gives any MCP-compatible agent or IDE full access to session management, playbooks, knowledge, and scheduling, confirmed both in docs and via probe. Missing for 10: independent/hands-on third-party corroboration of the MCP server working in practice, and detail on setup/auth specifics.
- [claimed-docs] “it gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [claimed-docs] “gives any MCP-compatible AI agent or IDE full access to session management, playbooks, knowledge, and scheduling”
- [probe] “official MCP server documented at https://docs.devin.ai/work-with-devin/devin-mcp”
ai-native userTest against a sandbox environment without touching production data
weight 1 · not comparableCursorn/aSandbox testing environments vs production data isolation is a data/infrastructure axis relevant to backend/platform products, not to an AI coding assistant like Cursor, which operates on local/repo code rather than managing production data environments.
Devin sessions run in isolated cloud VMs with their own shell/browser/full repo access, configurable environment blueprints for a 'known-good state,' and a CLI --sandbox flag enforcing OS-level write/network isolation, all of which support testing in isolated environments away from live infrastructure. However, there is no explicit documentation addressing production-data isolation or masking, or confirmation that these sandboxes are guaranteed free of production data. missing for 10: explicit statement on production-data separation/masking, independent hands-on confirmation that sandbox testing never touches production data.
- [claimed-docs] “The cloud session gets its own VM with a shell, browser, and full repo access, so it can keep going after you close your laptop.”
- [claimed-docs] “Configure it once, and every session boots into that known-good state.”
- [claimed-docs] “The --sandbox flag runs the CLI with OS-level isolation, enforcing writable paths and deny rules at the operating-system level and optionall…”
- [claimed-docs] “The `--sandbox` flag runs the CLI with OS-level isolation, enforcing writable paths and `deny` rules at the operating-system level”
- [claimed-docs] “Outposts lets you run Devin sessions inside infrastructure you control — your own VMs, containers, Kubernetes clusters, or even a Mac Mini o…”
- [claimed-docs] “Devin inspects your repository, figures out which tools, runtimes, and dependencies are needed, and generates the blueprint for you.”
developerReceive inline code completions and next-edit suggestions as I type
weight 3 · not comparableCursornone0/10The evidence pack contains no first-party documentation or hands-on account describing Cursor's own inline code completion or next-edit suggestion feature; only tangential community references compare competitors' tab-completion tools (e.g., Continue, SuperMaven) without confirming or detailing Cursor's implementation. Missing for 10: any first-party doc on Cursor's Tab/inline completion feature, hands-on confirmation it works as typed, and mention of 'next-edit' suggestion behavior.
Devinn/aDevin is an autonomous agentic coding product that works via task delegation (sessions, tickets, Slack), IDE handoff, and CLI/API integration rather than an inline editor completion tool; there is no evidence of an inline-completion or next-edit-suggestion feature as you type, and this axis is a different product category (IDE autocomplete tooling) than Devin's agent model.
developerView interactive diffs and share selected code as context from within my JetBrains IDE
weight 1 · not comparableCursornone0/10The only evidence touching JetBrains is a single line listing JetBrains among integrations (cursor-docs-6), with no detail on interactive diffs or context-sharing features within a JetBrains IDE specifically. No documentation, screenshots, or community reports confirm this JetBrains-specific capability.
- [claimed-docs] “Work with GitHub, GitLab, Azure DevOps, Bitbucket, JetBrains, Slack, Linear, and more”
Devinn/aThere is no evidence Devin ships a JetBrains IDE plugin; Devin's IDE integration is its own embedded/desktop IDE (imports VS Code/Cursor settings) rather than a JetBrains plugin, making this a category mismatch for how Devin operates.
- [claimed-docs] “Devin is designed to be a conversational user interface, and allows you to follow and take over Devin's development process in the embedded …”
- [claimed-docs] “Import VS Code or Cursor settings, configure themes, and start coding with AI-powered assistance.”
ai-native userSelf-host the core product
weight 3 · not comparableCursorn/aCursor is a proprietary AI coding assistant/IDE fork product, not an open-source or self-hostable platform; self-hosting the core product is a category error for this type of closed commercial tool.
Devinnone0/10Devin is a cloud-based SaaS agent; Outposts lets you run sessions on your own infrastructure but the core Devin model/orchestration itself remains Cognition-hosted, and there's no evidence of a self-hostable core product/model package. No mention of on-prem/self-hosted deployment of the core Devin engine.
- [claimed-docs] “Outposts lets you run Devin sessions inside infrastructure you control — your own VMs, containers, Kubernetes clusters, or even a Mac Mini o…”