OpenHands vs Jules
open-source · free-tier · usage-based
·free-tier · subscription-per-seat
OpenHands wins · 27–19 (26 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to OpenHandsA probe confirms OpenHands hosts a live llms.txt at docs.openhands.dev/llms.txt returning HTTP 200 with an explicit LLM-friendly documentation index, directly satisfying the story. Missing for 10: no independent third-party confirmation that agents actually consume/parse this file successfully in practice.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.openhands.dev/llms.txt # OpenHands Docs > LLM-friendly index of OpenHands documentation (V1). Lega…”
Julesnone0/10Probes show no llms.txt, docs.md, or machine-readable API spec exist at expected paths, and no evidence Jules can be pointed at such agent-oriented documentation formats; Jules does support AGENTS.md for repo context but that's a different mechanism than consuming llms.txt-style docs.
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to OpenHandsOpenHands supports CLI mode with auto-approve for non-interactive runs, API support for automation/scripting, webhook/schedule-triggered automations (Slack, GitHub, Linear), and can run headless in Docker/VMs/servers — all consistent with CI/automation usage. Missing for 10: an explicit first-party CI pipeline example (e.g., GitHub Actions config) or independent hands-on confirmation of headless CI runs.
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [claimed-docs] “API support for automation and scripting”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules provides a public API with API keys for creating custom workflows, sending tasks programmatically, and an optional automationMode to auto-create PRs, plus a community example of building an MCP server to dispatch tasks from another tool — this directly supports headless/CI-style automation. Missing for 10: no official CI/CD integration examples (e.g. GitHub Actions), no documented webhook/polling pattern for task completion, and no independent verification of reliability at scale in automated pipelines.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round drawnOpenHandsnone0/10OpenHands is an AI coding agent, so the axis of consuming MCP servers as a client is applicable (unlike serving MCP, which would be na for an agent), but no evidence in the pack mentions MCP integration, configuration, or tool-plugging capability at all.
Julesnone0/10No evidence that Jules can connect to or consume external MCP servers as tool providers; the only MCP-related evidence (jules-comm-19) describes someone building an MCP server that calls INTO the Jules API, which is the reverse integration direction, not Jules plugging in MCP servers for its own tool use.
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
ai-native userUse an official CLI
weight 2 · round to OpenHandsOpenHands documents an official CLI mode with natural language task execution, conversation resumption, and auto-approve controls, confirmed by dedicated docs pages and a probe verifying the CLI documentation page exists. Missing for 10: independent/hands-on third-party corroboration beyond vendor docs, and more detail on CLI installation/distribution mechanics.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Resume previous conversations”
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Julesnone0/10Evidence shows Jules offers a web app, GitHub integration, and a REST API for custom workflows, but no official CLI tool is documented or mentioned anywhere in the evidence pack.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
ai-native userDrive the product through a documented public API
weight 3 · round drawnOpenHands exposes a documented OpenAPI spec (openapi.json) plus an llms.txt LLM-friendly docs index, and pricing docs explicitly mention 'API support for automation and scripting,' confirming a public, documented API for programmatic/agentic control. missing for 10: independent third-party corroboration of API usage/reliability and more detailed API reference docs beyond the openapi.json probe.
- [probe] “PROBE openapi: HTTP 200 at https://docs.openhands.dev/openapi.json — contains "openapi" key”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.openhands.dev/llms.txt # OpenHands Docs > LLM-friendly index of OpenHands documentation (V1). Lega…”
- [claimed-docs] “API support for automation and scripting”
Jules ships a documented public API (developers.google.com/jules/api) with API key management, custom workflows, sending messages to the agent, and automationMode config; a community user independently built an MCP server on top of the API confirming real-world programmatic access. Missing for 10: no discoverable OpenAPI/swagger spec or llms.txt (probes returned 404s), reducing machine-readability/self-service tooling confidence.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round drawnOpenHandsnone0/10Evidence shows OpenHands supports bringing your own LLM keys, SAML/SSO, and general API access, but nothing describes issuing scoped or least-privilege API credentials specifically for an agent's actions/tool access.
Julesnone0/10Jules documents a basic API key creation flow (max 3 keys) but provides no evidence of scoped or least-privilege permission controls — no mention of scopes, roles, or restricted-access tokens for the agent.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
ai-native userBuild against official SDKs
weight 2 · round to OpenHandsOpenHands documents an official Software Agent SDK ('a composable Python library for building agents that work with code') plus a public OpenAPI spec and llms.txt index, giving AI-native users concrete official interfaces to build against. Missing for 10: independent/third-party corroboration of SDK usage, and richer SDK-specific docs (examples, API reference depth) beyond the single description.
- [claimed-docs] “The Software Agent SDK is a composable Python library for building agents that work with code.”
- [probe] “PROBE openapi: HTTP 200 at https://docs.openhands.dev/openapi.json — contains "openapi" key”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.openhands.dev/llms.txt # OpenHands Docs > LLM-friendly index of OpenHands documentation (V1). Lega…”
Jules ships an official public API with documented endpoints, API key management, and use cases for building custom workflows/integrations, and a community member confirms building a personal MCP server against the Jules API. missing for 10: no official OpenAPI/SDK spec discoverable (openapi probes 404), no first-party language SDKs mentioned, and no independent SDK-quality corroboration beyond one community integration example.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userSubscribe to events via webhooks
weight 2 · round to OpenHandsGitHub README explicitly states OpenHands can 'run on a schedule or in response to webhook events' for automations, indicating webhook-triggered event subscription, but there is no dedicated documentation page detailing webhook setup, payload schema, or event types. missing for 10: dedicated webhook docs/config guide, independent/hands-on confirmation, and detail on which events can be subscribed to.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Julesnone0/10Evidence shows Jules has an API for creating tasks/sending messages, notifications for task completion, and GitHub integration, but there is no mention of webhooks or event subscription mechanisms anywhere in the docs or community evidence. Missing for 10: any documentation of a webhook endpoint, event subscription API, or push-based notification mechanism to external systems.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to OpenHandsOpenHands ships automation agents that inspect real data (incident logs, PR diffs, workflow failures, security alerts) and generate AI insights/suggestions such as severity assessment, review comments, root-cause analysis, and remediation PRs, directly matching the story's intent for a coding-agent product. Evidence is vendor-documented only, with no independent/hands-on corroboration of the quality of these insights. Missing for 10: independent validation of suggestion quality, a unified 'insights' UI/dashboard, and evidence of insights beyond code/security/ops contexts.
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
Jules generates AI-derived plans, diffs, and PR suggestions based on analysis of the user's repository data, which is the coding-agent analog of 'AI-generated insights/suggestions from your data.' Community evidence is mixed on the quality of these suggestions, with some praising usable PRs and others criticizing low-quality output on complex codebases. Missing for 10: no evidence of broader analytics-style insights beyond code-change suggestions, and no independent benchmarking confirming insight quality across use cases.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to OpenHandsOpenHands supports scheduled/webhook-triggered automations integrating with Slack, GitHub, Linear, etc., and ships prebuilt autonomous workflows (bug triage, PR review, CI failure fixing, security remediation, Slack mention handling) that run without human intervention, plus API support for scripting automations. Missing for 10: independent/hands-on verification of the scheduling/webhook trigger reliability and no detailed docs excerpt on configuring schedules beyond marketing copy.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
- [claimed-docs] “API support for automation and scripting”
Jules supports background/autonomous execution via GitHub label-triggered tasks and a public API for building custom automations (e.g., automated bug-fixing, code review) with an optional automationMode for auto-PR creation, and it runs tasks in a cloud VM without needing the user present. However, docs also show a human-approval gate (plan approval, PR review) rather than fully hands-off automation, and community reports note tasks getting stuck, hitting limits, or needing babysitting, undercutting reliability of unattended runs. Missing for 10: evidence of true scheduled/cron-style recurring automations, and independent confirmation that automations run to completion without manual intervention.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round drawnOpenHands is itself the built-in AI assistant/agent: CLI mode lets users type natural language tasks and get instant feedback, resume conversations, and auto-approve actions, and it can run locally, remote, or in the cloud with any LLM. This directly matches delegating tasks to a built-in assistant within the product. Missing for 10: independent/hands-on user reports validating the delegation experience beyond vendor docs, and more detail on task delegation depth (e.g., multi-step autonomy limits).
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Resume previous conversations”
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [github] “Switch between local, remote, and cloud agents without losing focus”
- [github] “Use with any LLM”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules is exactly this kind of built-in agentic assistant: users delegate coding tasks via GitHub issue labels, the web app, or API, and Jules autonomously plans, clones repos, runs in a VM, and returns diffs/PRs for approval (jules-docs-2,3,4,7,8,9). Community evidence corroborates real-world task delegation and usable output (jules-comm-2, jules-comm-11, jules-comm-15, jules-comm-16), though mixed reports of loops, babysitting needs, and reliability issues (jules-comm-14, jules-comm-18) temper quality. Missing for 10: independent benchmarking of task success rates and evidence of consistent reliability at scale.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
ai-native userOperate the product with natural-language commands
weight 2 · round drawnOpenHands' CLI mode explicitly lets users type natural language tasks and get instant feedback, resume conversations, and control approvals, directly matching the story, and this is corroborated by an official documented CLI probe. Missing for 10: independent/hands-on user reports validating the natural-language interaction quality beyond vendor docs.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Resume previous conversations”
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules is designed around natural-language task submission and messaging: users submit a task description, Jules generates a plan, and you can send follow-up messages to the agent via API or web app (jules-docs-4, jules-docs-12). Community reports confirm real-world use of NL prompts to drive coding tasks, including iterative feedback and error-message follow-ups (jules-comm-2, jules-comm-15, jules-comm-16). Missing for 10: no detailed documentation of the full range/complexity of natural-language commands supported (e.g., multi-step conversational control, command reference) and no independent benchmark of NL command robustness beyond anecdotal reports.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “To send a message to the agent:”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round to OpenHandsEvidence shows an OpenAPI spec is served (openhands-probe-2) and API support is advertised for automation/scripting (openhands-docs-13), implying some API reference exists, but there is no evidence of an interactive documentation UI (e.g., Swagger/Redoc 'try it' console) or runnable code examples tied to that spec. missing for 10: interactive API explorer UI, runnable/try-it code examples, confirmation the openapi.json is rendered as browsable docs.
- [probe] “PROBE openapi: HTTP 200 at https://docs.openhands.dev/openapi.json — contains "openapi" key”
- [claimed-docs] “API support for automation and scripting”
Julesnone0/10Jules documents an API (create tasks, send messages, API keys) but there is no evidence of an interactive API reference or runnable code examples; probes for openapi/swagger specs all returned 404s, suggesting no such interactive reference exists.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “To send a message to the agent:”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round to OpenHandsA probe confirms an OpenAPI spec is served at docs.openhands.dev/openapi.json with a valid 'openapi' key, directly satisfying the machine-readable API spec requirement, and API support for automation/scripting is also documented in pricing. Missing for 10: independent third-party corroboration or detailed docs describing spec coverage/versioning.
- [probe] “PROBE openapi: HTTP 200 at https://docs.openhands.dev/openapi.json — contains "openapi" key”
- [claimed-docs] “API support for automation and scripting”
Julesnone0/10Jules ships a documented REST API (jules-docs-9..12) but there is no evidence of a downloadable OpenAPI/Swagger spec; explicit probes for llms.txt, docs.md, and standard OpenAPI paths all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), confirming no machine-readable spec is published.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userTest against a sandbox environment without touching production data
weight 1 · round to JulesOpenHands can run agents in Docker/VMs (openhands-gh-5), which implies isolated execution rather than direct production access, but the evidence pack never explicitly describes a sandbox environment for testing safely against non-production data. Missing for 10: explicit documentation of sandbox/isolation guarantees, workspace-vs-production data separation, and any hands-on confirmation that production systems are protected.
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
Jules runs tasks in an isolated cloud VM that clones the repo, meaning code changes happen in a sandboxed environment rather than directly on production, and PRs must be reviewed/approved before merging to the real branch. However, this is a code-sandbox for making changes, not a dedicated 'test against sandbox data/environment' feature, and there's no evidence of test-data isolation, staging environment provisioning, or explicit protection of production data/services beyond the VM/PR review flow. missing for 10: explicit sandbox test-data isolation, staging/production separation guarantees, independent verification that production systems are never touched.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnOpenHandsnone0/10Evidence shows OpenHands exposes an OpenAPI spec and general 'API support for automation and scripting,' and docs reference a 'V1' with legacy V0 pages excluded, but there is no documented API versioning scheme or deprecation policy anywhere in the pack.
- [probe] “PROBE openapi: HTTP 200 at https://docs.openhands.dev/openapi.json — contains "openapi" key”
- [claimed-docs] “API support for automation and scripting”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.openhands.dev/llms.txt # OpenHands Docs > LLM-friendly index of OpenHands documentation (V1). Lega…”
Julesnone0/10Jules has a public API (jules-docs-9, jules-docs-10) but there is no evidence of API versioning scheme or a documented deprecation policy; probes for OpenAPI spec/docs.md/llms.txt all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), suggesting no formal machine-readable API contract or lifecycle documentation is available.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round drawnOpenHands supports automations/workflows triggered by events (Slack, GitHub, Linear, webhooks) and API access for scripting, which enables some multi-item automation, but there's no explicit evidence of a bulk-operation feature (e.g., processing a batch list of items/tasks in one command or UI action). missing for 10: explicit bulk/batch operation feature, evidence of processing multiple items in a single invocation, UI/CLI support for batch task lists, independent confirmation of scale.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
Jules exposes an API that lets users script custom workflows and dispatch tasks programmatically (jules-docs-9, jules-docs-12), and usage-limit tiers explicitly target 'power users & agent-heavy workflows' with up to 300 tasks (jules-docs-13), while community reports confirm it 'handles large quantities of tasks very well' and users built custom dispatch tooling via the API/MCP bridge (jules-comm-11, jules-comm-19). However, there's no documented native UI or endpoint for a single bulk action across many items (e.g., batch-applying one task to many repos/issues) — each task still appears to be created and reviewed individually via GitHub labels or API calls. Missing for 10: a documented batch/bulk endpoint or UI feature, first-party bulk-operation examples, and independent verification of true parallel bulk execution rather than just high per-day task volume.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to OpenHandsOpenHands ships documented automation triggers—scheduled runs, webhook events, PR label events, Slack mentions, failed-workflow detection, and security-alert triage—each automatically invoking an agent action, matching the 'rules trigger actions on events' story (openhands-gh-4, openhands-docs-7 to openhands-docs-11). missing for 10: independent/hands-on verification that users can define fully custom rule logic (vs. fixed preset automations), and no evidence of a general-purpose rule-authoring UI or DSL for arbitrary event/action pairing
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
Jules supports some event-driven automation — e.g., assigning a 'jules' label to a GitHub issue automatically triggers a task, and the API lets developers build custom workflows/automations around Jules (jules-docs-2, jules-docs-9). However, there's no evidence of a general rules/conditions engine where users define arbitrary triggers and actions beyond this specific label mechanism and API calls. Missing for 10: a configurable rules interface (conditions + multiple trigger types), documentation of scheduled/webhook-based triggers, and independent confirmation that automated rule-based triggering works reliably in practice.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to OpenHandsOpenHands explicitly supports creating automations/workflows that run on a schedule or in response to webhook events, integrating with Slack, GitHub, Linear, etc., which directly matches recurring job scheduling. Missing for 10: independent/hands-on corroboration of the scheduling UI/config and details on job management (pause/edit/monitor recurring jobs).
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
Julesnone0/10Evidence covers task assignment via GitHub labels, API-driven custom workflows, and one-off task automation, but nothing describes recurring/scheduled jobs (e.g., cron-like triggers) — the API docs only mention creating tasks and sending messages, not recurrence.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnOpenHandsnone0/10Evidence shows OpenHands can create automations/workflows (Slack, GitHub, Linear integrations, scheduled/webhook triggers) but nothing in the pack describes version history, review, or rollback mechanisms specifically for these automations themselves — no changelog, diff view, or revert feature is documented.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
Julesnone0/10Jules produces PRs and diffs via GitHub but there is no evidence of a mechanism to version, review history of, or roll back automations/tasks themselves (as opposed to code changes tracked by git/GitHub). missing for 10: automation versioning/history feature, rollback/undo of Jules tasks, changelog or audit trail for automations.
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation
End-to-end implementation by the agent — multi-file changes, task completion
End to end feature delivery
ai-native userHave an agent automatically generate and run tests to validate its own code changes before proposing them
weight 2 · round to JulesOpenHandsnone0/10No evidence in the pack describes OpenHands agents autonomously generating or running tests to validate their own code changes before proposing them; the docs/GitHub excerpts cover CLI usage, multi-agent backends, automations, and pricing, but not test-generation/self-validation workflows.
Jules runs code changes in an isolated VM and generates diffs/PRs for review, and one community report notes it is 'good at making tests for very specific/isolated functions,' implying some test generation capability, but there is no first-party documentation describing an explicit test-generate-and-run validation loop before proposing changes. missing for 10: official docs on automated test generation/execution as a validation step, evidence of test results being surfaced to users before PR creation, and independent confirmation across varied codebases.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
developerHave an agent autonomously diagnose and fix a reported bug
weight 3 · round drawnOpenHands ships explicit automation for bug diagnosis and fixing: watching for new bugs/incidents and gathering logs (openhands-docs-7), detecting failed workflows/inspecting logs/identifying cause and opening a PR fix (openhands-docs-9), and general natural-language task execution via CLI (openhands-docs-1). This directly matches autonomous bug diagnosis-and-fix, though missing for 10: independent hands-on verification of fix success rate and end-to-end case studies beyond vendor docs.
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “The Software Agent SDK is a composable Python library for building agents that work with code.”
Jules is designed to autonomously diagnose and fix issues: it clones the repo into a VM, generates a plan, makes code changes, and opens a PR for review, and can be triggered directly from a GitHub issue via the 'jules' label (a common bug-report workflow). Community evidence corroborates real-world bug-fixing use, including a case where it failed then successfully fixed the issue after being given the error message. Missing for 10: independent benchmarking on bug-fix success rate and more consistent evidence across complex codebases (some reports of failures/loops on harder tasks).
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
product-managerGo from a mockup or design to a working implementation without an engineering handoff
weight 2 · round drawnOpenHandsnone0/10The evidence shows OpenHands executes natural-language coding tasks and automations, but nothing in the pack addresses ingesting mockups/design files or a PM-oriented, engineer-free workflow from design to implementation. missing for 10: mockup/design ingestion capability, no-code PM-facing workflow evidence, any example of design-to-code handoff elimination.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Julesnone0/10Evidence shows Jules operates on code repositories via GitHub issues/tasks, plans, diffs, and PRs, but there is no mention of accepting mockups/designs as input or any PM-oriented no-code workflow — usage still requires repo access, issue creation, and reviewing technical diffs/plans. Nothing in the pack demonstrates a design-to-implementation path bypassing engineering.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
developerHave an agent implement a requested feature end-to-end, including writing tests
weight 3 · round to JulesOpenHands is positioned as an autonomous coding agent that takes natural language tasks and works across CLI, cloud, and automation workflows (e.g., PR review, bug-fixing, incident response), implying it can implement features end-to-end, but the evidence pack lacks a concrete first-party or hands-on example of the agent writing a feature plus tests. Missing for 10: explicit documentation/demo of full feature-implementation-with-tests workflow, independent benchmark or hands-on verification of test-writing capability, and confirmation of end-to-end PR creation including tests.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [github] “Run OpenHands, Claude Code, Codex, Gemini, or any ACP-compatible agent across local, remote, and cloud backends.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
Jules is explicitly built for autonomous end-to-end coding: it clones the repo, generates a plan, modifies files, produces a diff, and opens a PR (jules-docs-3/4/7/8), and one user specifically notes it's 'good at making tests for very specific/isolated functions' (jules-comm-11), with others confirming usable PRs and time savings (jules-comm-6, jules-comm-16). However, community evidence shows significant caveats: it struggles on complex/monorepo codebases, gets stuck in loops without a stop button, needs heavy babysitting, and often fails to complete tasks unassisted (jules-comm-3, jules-comm-9, jules-comm-10, jules-comm-13, jules-comm-14, jules-comm-18). Missing for 10: consistent reliability across complex real-world features, broader evidence of comprehensive test coverage (not just isolated functions), and independent benchmarks confirming end-to-end success rate.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “I've been playing with it, and I've been generally not impressed. There are obvious annoying UI bugs and the output isn't very good for anyt…”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
Environment setup
developerHave an agent automatically clone the repo, install dependencies, and configure its own working environment
weight 2 · round to JulesEvidence shows OpenHands runs agents in configurable sandboxed backends (Docker, VMs, local/remote/cloud) and supports natural-language task execution with auto-approve, which implies some autonomous environment setup, but no citation explicitly confirms the agent auto-clones repos or installs dependencies on its own. Missing for 10: explicit documentation of repo cloning, dependency installation, and end-to-end environment bootstrap steps.
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Auto-approve all actions (use with caution)”
Jules' core workflow is documented: it runs in a cloud VM, clones the repo, installs dependencies, and configures its environment automatically before making changes, corroborated by community usage reports of automated PR generation. missing for 10: independent technical verification of dependency-install robustness across complex/monorepo setups, and some community reports note environment/config confusion in bespoke or monorepo projects.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
Interactive takeover
developerTake over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
weight 2 · round drawnOpenHands supports resuming conversations and switching between local/remote/cloud agents 'without losing focus,' implying some cross-surface continuity, but there's no explicit documentation of a developer taking over a live in-progress task from editor/terminal/browser mid-run with redirect capability. missing for 10: explicit IDE/editor integration for live takeover, concrete UI/UX description of mid-task redirect, and independent hands-on confirmation of seamless handoff across all three surfaces.
- [github] “Switch between local, remote, and cloud agents without losing focus”
- [github] “Run OpenHands, Claude Code, Codex, Gemini, or any ACP-compatible agent across local, remote, and cloud backends.”
- [claimed-docs] “Resume previous conversations”
Jules supports reviewing plans, diffs, and sending follow-up messages to redirect a task (jules-docs-4, jules-docs-6, jules-docs-7, jules-docs-12), and third parties have wired the API into VS Code/Copilot Chat to dispatch tasks (jules-comm-19), but this is dispatch, not mid-task takeover. A hands-on report explicitly notes there's no 'STOP' button to interrupt a looping task (jules-comm-14), undercutting redirect control, and all takeover is browser/API-based with no native terminal or editor 'take over' UI documented. Missing for 10: documented in-editor/terminal takeover UI, ability to pause/interrupt an in-progress run, and evidence the API-based messaging genuinely redirects rather than just appends instructions.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
developerSend follow-up instructions to an active agent session to steer its work without restarting
weight 2 · round to JulesCLI mode docs mention typing natural language tasks with instant feedback and resuming previous conversations, which implies interactive follow-up steering, but there is no explicit evidence describing sending new instructions mid-task to an already-running/active agent session without restarting it. missing for 10: explicit documentation of mid-session steering/interrupt-and-redirect behavior while an agent is actively working, and independent/hands-on confirmation that follow-up messages are incorporated without restarting the session.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Resume previous conversations”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules's API docs explicitly describe sending a message to an active agent session (jules-docs-12), and community reports confirm the workflow of reviewing partial work and providing feedback for Jules to continue rather than restarting (jules-comm-3). This shows the steering capability works in practice, though documentation on how follow-ups affect an in-progress plan/execution is thin. Missing for 10: detailed first-party documentation of mid-task message handling/UI chat thread, and independent hands-on confirmation of steering effectiveness beyond one anecdote.
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
Sandbox execution
developerHave an agent safely execute code and install dependencies inside an isolated sandbox
weight 3 · round to JulesEvidence confirms OpenHands supports running agents in isolated environments (Docker/VMs) and explicitly references 'sandbox infrastructure' as part of its backend, implying code execution is sandboxed. However, there is no detailed documentation on dependency installation within the sandbox, security guarantees, or isolation mechanics beyond these brief mentions. missing for 10: explicit sandbox architecture docs, dependency-installation workflow details, isolation/security guarantees, independent verification of sandbox safety.
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [claimed-docs] “OpenHands Cloud is the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure.”
Jules explicitly runs in a cloud VM that clones code and installs dependencies, isolated from the user's local machine, with a plan-review step before code changes are made — directly matching the sandboxed execution story. Community reports corroborate it actually running tasks end-to-end (producing PRs) in this VM environment, though some criticize reliability/loops on complex codebases. Missing for 10: detailed docs on sandbox security boundaries (network isolation, resource limits) and independent security audit of the VM isolation.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight
Keeping a human in the loop — approvals, checkpoints, interrupts
Approval controls
developerConfigure an agent to auto-approve all its actions instead of confirming each one
weight 2 · round to OpenHandsDocs explicitly list an 'Auto-approve all actions (use with caution)' setting for CLI mode, directly matching the story. Missing for 10: independent/hands-on confirmation and details on scope (per-action vs global) or how to configure it beyond CLI mode.
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules's default workflow requires manual approval at multiple stages (plan approval, diff review, PR approval), but the API docs mention an optional `automationMode` field that changes default behavior around automatic PR creation, suggesting some configurable auto-approve path exists via API rather than the standard UI flow. Missing for 10: explicit documentation of a setting that suppresses plan-approval and diff-review confirmations entirely, and any hands-on/community confirmation that this automation mode actually skips human checkpoints.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
product-managerApprove key agent decisions from my phone while agents continue working
weight 1 · round to JulesOpenHands Cloud offers 'hosted cloud access from desktop and mobile' and Slack-based interaction (@mentions, replies) that could let a PM check in remotely, but there is no documented feature for approving specific in-progress agent actions/decisions via a mobile interface while the agent continues autonomously — the only approval-related control mentioned is a blanket 'auto-approve all actions' CLI flag, not selective human-in-the-loop approval. Missing for 10: explicit mobile approval/confirmation UI, human-in-the-loop decision gating documentation, evidence of push notifications or approval prompts reaching a phone.
- [claimed-docs] “Hosted cloud access from desktop and mobile”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
- [claimed-docs] “Auto-approve all actions (use with caution)”
Jules is a cloud-based agent with a web app and notification system, plan approval, diff review, and PR approval steps, and community evidence confirms users review/approve PRs from a browser including 'coding on my phone' via mobile web. However, there is no dedicated native mobile app or explicit mobile-optimized approval UI, and no evidence of push-notification-driven approval flows tailored to on-the-go PM decision-making. missing for 10: dedicated mobile app or mobile-specific UI, evidence of push notifications enabling quick phone-based approvals, PM-specific (non-developer) approval workflow.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
engineering-leadSet tiered autonomy levels controlling what an agent can do without manual confirmation
weight 3 · round drawnEvidence shows only a binary confirmation toggle (auto-approve all actions vs manual confirmation) in CLI mode, not a tiered/granular autonomy system with multiple configurable levels for an engineering lead to set. Missing for 10: documented multi-tier permission/autonomy levels, admin controls to enforce team-wide policies, and per-action or per-risk-category confirmation settings.
- [claimed-docs] “Auto-approve all actions (use with caution)”
Jules has a binary plan-approval workflow (review/approve plan, review diffs, approve PRs) and an API `automationMode` field that toggles whether PRs are auto-created, but there's no evidence of configurable tiered autonomy levels (e.g., low/medium/high trust settings, granular permission scopes, or per-action confirmation thresholds) that an engineering-lead could set. Missing for 10: explicit multi-tier autonomy/permission settings, admin-configurable trust levels, and any org-wide policy controls beyond the single automationMode toggle and default plan-approval gate.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
Model control
ai-native userHave each task prompt automatically routed to the most suitable underlying model
weight 2 · round drawnOpenHandsnone0/10Evidence shows OpenHands supports using any LLM and switching between agent backends manually, but there is no mention of automatic routing of prompts to the most suitable model based on task characteristics. Missing for 10: any model-routing/selection logic, per-task model suitability heuristics, or documentation of automatic model selection.
- [github] “Use with any LLM”
- [github] “Switch between local, remote, and cloud agents without losing focus”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Julesnone0/10No evidence Jules performs automatic model routing per task; docs mention only a single flagship model (Gemini 3 Pro) as 'priority access,' with no mention of routing prompts to different models based on suitability. Missing for 10: any documentation of multi-model routing logic, model-selection criteria, or automatic switching between models.
- [claimed-docs] “Priority access to the latest models, starting with Gemini 3 Pro”
engineering-leadSwitch away from automatic model selection to a specific model of my choice
weight 1 · round to OpenHandsEvidence confirms OpenHands can be used with any LLM and lets users bring their own LLM keys, implying model choice is configurable, but there is no explicit documentation of an 'automatic model selection' default or a UI/CLI toggle to override it. Missing for 10: explicit docs on default/automatic model selection behavior, step-by-step instructions for switching models, and independent confirmation of the switch working in practice.
- [github] “Use with any LLM”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Visibility monitoring
developerWatch what a running agent is doing in real time, including its current status
weight 3 · round to JulesOpenHands CLI mode offers instant feedback on tasks and resumable conversations, implying some real-time interaction, but there's no explicit evidence of a live status dashboard, streaming action log, or step-by-step progress view while an agent runs. missing for 10: explicit real-time status/progress UI documentation, evidence of live action streaming or step visibility, independent hands-on confirmation of watching an agent live.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Resume previous conversations”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules provides notifications on task completion/input-needed and a plan-approval step, implying some status visibility, plus a diff/PR review flow, but there is no evidence of a live real-time activity feed or streaming log of the agent's current actions while running; community feedback even notes the absence of a stop/interrupt control mid-run, suggesting limited real-time observability. Missing for 10: evidence of a real-time execution log/console view, granular step-by-step status updates, and independent confirmation of live monitoring UX.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
developerGet notified when an agent completes a task or needs my input
weight 2 · round to JulesOpenHands supports Slack integration where it 'replies when the agent finishes' and workflow automations tied to webhook events (Slack, GitHub, Linear), which implies notification-like behavior on task completion; CLI mode also gives instant feedback in interactive sessions. However, there's no explicit evidence of a dedicated notification system for 'needs my input' events or push/desktop alerts outside Slack context. missing for 10: explicit documentation of notifications when agent needs human input/approval, native mobile/desktop push alerts, and independent/hands-on confirmation of notification reliability.
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Type natural language tasks and receive instant feedback”
First-party docs explicitly state users are notified when a task completes or needs input, and the workflow (plan approval, diff review, PR creation) reinforces oversight checkpoints where notifications matter. Community evidence corroborates the async, check-back-later usage pattern (e.g. reviewing PRs after time away), consistent with notification-driven workflows. Missing for 10: independent verification of notification channels (email/push/Slack), and no detail on notification reliability or configurability.
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
Intent to spec — stories about intent to spec in this arenaIntent to spec
Stories about intent to spec in this arena
Natural language task intake
developerDescribe a feature or bug in plain language and have it automatically turned into a scoped implementation task
weight 3 · round to JulesOpenHands lets users type natural-language tasks directly in the CLI and receive feedback, which is the core mechanism for turning plain-language input into agent-executed work, and automation triggers (Slack mentions, PR labels, failed workflows) show it converting informal signals into concrete PRs/tasks. However, there is no evidence of an explicit 'spec' or scoped task artifact (e.g., a generated plan, ticket, or requirements doc) being produced before implementation — missing for 10: explicit task-scoping/spec generation step, evidence of a structured plan artifact, and independent confirmation that vague bug/feature descriptions reliably become well-scoped tasks rather than direct code edits.
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
Jules docs clearly show a plain-language task submission flow that generates a reviewable/approvable plan before code changes, then produces a diff and PR (jules-docs-4, jules-docs-7, jules-docs-8), directly matching the intent-to-spec story. However community reports show inconsistent scoping quality — confusion in monorepos, endless loops, and needing heavy babysitting — indicating the automatic scoping doesn't always hold up in practice. Missing for 10: independent verification of plan/spec quality on complex codebases, and consistent evidence the generated plan reliably matches developer intent without back-and-forth correction.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
product-managerConvert user feedback submissions into structured tasks with proposed scope
weight 2 · round drawnOpenHandsnone0/10OpenHands integrates with Linear/GitHub/Slack for automations and can process natural-language tasks, so the general axis of turning input into work items is plausible, but there is no evidence of a feature that ingests user feedback and outputs a structured task with proposed scope for PM workflows.
Julesnone0/10Jules's documented workflow is a coding-agent pipeline (clone repo, generate a plan, produce diffs/PRs) driven by developer-specified tasks or GitHub issues, not a mechanism for ingesting unstructured user-feedback and converting it into a structured task with proposed scope for a PM. No evidence shows feedback intake, requirement structuring, or scope proposal features aimed at product managers.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
developerAttach a marked-up screenshot or mockup to a task so the agent implements the correct visual change
weight 2 · round drawnOpenHandsnone0/10No evidence anywhere in the pack that OpenHands supports attaching images, screenshots, or mockups to a task, or that the agent can interpret visual markup to drive implementation; documentation focuses on text-based CLI tasks, integrations, and automations.
Julesnone0/10No evidence Jules supports attaching images, screenshots, or mockups to a task — task submission is described only via text prompts, GitHub issue labels, or API messages. Missing for 10: any mention of image/screenshot attachment, multimodal input handling, or visual mockup interpretation.
Plan approval
developerReview and approve an agent's implementation plan before any code changes are made
weight 3 · round to JulesOpenHandsnone0/10The evidence pack mentions auto-approve controls and CLI usage but contains no mention of a plan-review/approval step before code changes are made; no planning-phase or approval-gate feature is documented. missing for 10: any documentation of a plan-generation step, an approval/confirmation gate prior to code edits, or user testimony confirming such a workflow exists.
Jules docs explicitly state the agent generates a plan upon task submission that users can review and approve before any code changes are made, directly matching the story. No independent/hands-on evidence specifically confirms or contradicts this plan-approval step, though other workflow claims (diffs, PRs) are corroborated by community mentions. Missing for 10: independent/hands-on confirmation of the plan-review step specifically, and more detail on what the plan interface looks like.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
engineering-leadApprove a task's scope and contract before an agent is allowed to modify the repository
weight 2 · round to JulesOpenHandsnone0/10No evidence of a scope/contract approval workflow gating agent repository modifications; only auto-approve settings and general automation features are documented, not a pre-execution scope-approval gate for engineering leads.
- [claimed-docs] “Auto-approve all actions (use with caution)”
Jules generates a plan that the user can review and approve before any code changes are made, giving an engineering-lead a gate before modification, and it also produces diffs/PRs requiring approval before merge. However, this is a general 'plan approval' UX rather than a formal scope/contract sign-off workflow tailored for engineering-lead governance (e.g., no evidence of role-based approval gates, policy enforcement, or blocking unauthorized starts). missing for 10: explicit lead/role-based approval gating before task execution starts, evidence of enforceable contract/scope definitions beyond a plan preview, independent confirmation that the plan-approval step reliably blocks modification.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
Ticket driven tasking
developerAssign a coding task to an agent directly from an existing issue or ticket
weight 3 · round to JulesOpenHands supports automations that integrate with GitHub and Linear and can respond to webhook events, plus label-triggered PR review and bug-watching automations, implying issue/ticket-triggered agent runs are possible. However there's no explicit documentation of a workflow where a developer directly assigns a specific issue/ticket to an agent (e.g. via an 'assign to OpenHands' button or issue-comment trigger) as opposed to general automation setup. Missing for 10: explicit documentation of issue-to-agent assignment UX (e.g., GitHub issue comment/label triggering agent to pick up that specific ticket), independent/hands-on confirmation of this workflow, and ticketing system coverage beyond GitHub/Linear mentions.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
Jules explicitly supports assigning tasks directly from GitHub issues via a 'jules' label, and API/community evidence corroborates task dispatch and PR creation workflows. Missing for 10: evidence of ticket-tracker integrations beyond GitHub (e.g., Jira/Linear) and independent hands-on confirmation of the label-to-task flow itself.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to OpenHandsOpenHands exposes a documented OpenAPI spec and dedicated 'API support for automation and scripting' tier, plus CLI mode with resume/auto-approve that mirrors UI conversation flows, suggesting broad API parity. However, no evidence explicitly confirms that every UI feature (e.g., cloud dashboard views, integrations config, enterprise SSO management) is fully API-accessible. Missing for 10: explicit parity documentation enumerating UI vs API feature coverage, and independent/hands-on confirmation that API can replicate all UI workflows.
- [probe] “PROBE openapi: HTTP 200 at https://docs.openhands.dev/openapi.json — contains "openapi" key”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “Type natural language tasks and receive instant feedback”
- [claimed-docs] “Resume previous conversations”
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
Jules exposes a public API for creating tasks, sending messages, and enabling automation (PR creation) — core building blocks of the UI workflow — and a community user confirms dispatching tasks via the API from an external MCP client, showing real usage beyond the UI. However, there's no documented API coverage for plan approval, diff browsing/approval, or settings/API-key management, and no OpenAPI spec is discoverable (probes 404), so full UI/API parity is unconfirmed. Missing for 10: API endpoints for plan review/approval, diff/PR review parity, settings management, and a public API spec confirming full feature coverage.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnOpenHandsnone0/10No evidence in the pack describes an explicit data export feature or open-format data portability; only conversation resume and self-hosting options are mentioned, not a way to export/leave with all user data. missing for 10: explicit export/download feature, documented open data formats, statement on data portability upon leaving the platform.
Julesnone0/10No evidence of any data export feature, open-format export, or account/data portability mechanism for Jules; the product works via GitHub repos/PRs but nothing indicates users can export Jules-specific data (task history, configs, etc.) and leave. missing for 10: any documentation of data export, open format support, or account portability/deletion workflow.
ai-native userRead the product's source under an open license
weight 2 · round to OpenHandsThe GitHub repository (openhands-gh-1..5) confirms the source code is publicly hosted and readable, and openhands-docs-15 explicitly references an 'OpenHands Open Source' plan/tier, implying the core project is open-licensed. However, no evidence pack item names the specific license (e.g., MIT/Apache) or points to a LICENSE file, so full open-license confirmation is unverified. Missing for 10: explicit license name/file citation, independent confirmation of license terms.
- [github] “Switch between local, remote, and cloud agents without losing focus”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Julesnone0/10Jules is a closed, proprietary Google product (cloud VM, API keys, usage limits); no evidence anywhere of source code being published or licensed openly, and probes for docs/openapi artifacts return 404s, further suggesting no open publishing.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userSelf-host the core product
weight 3 · round to OpenHandsEvidence shows OpenHands can be run locally/self-hosted (Docker, VMs, or any agent server backend) as opposed to only using the managed Cloud/Enterprise offerings, and it is open-source with an SDK for building on it. Missing for 10: independent hands-on confirmation of a full self-hosted setup (e.g., a third-party report of running the entire stack without cloud dependencies) and detailed self-hosting infra requirements/documentation.
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [claimed-docs] “OpenHands Cloud is the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure.”
- [claimed-docs] “The Software Agent SDK is a composable Python library for building agents that work with code.”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Julesnone0/10Jules is explicitly a cloud-based VM service (Google-hosted) with no evidence of any self-hosting option; community even notes cloud-only design as a drawback compared to local tools like Claude Code.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “It's a shame Google picked the wrong system design for Jules. Claude Code's system design is clearly superior at this point.”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Enterprise licensing
engineering-leadLicense an enterprise deployment with SSO and commercial support for organization-wide rollout
weight 2 · round to OpenHandsOpenHands explicitly markets an Enterprise tier with licensed self-hosting/managed deployment and commercial support (openhands-docs-5), and the pricing page lists Enterprise SAML/SSO (openhands-docs-14). However, there is no detail on contract terms, SLA specifics, or independent confirmation of enterprise rollouts. Missing for 10: concrete SLA/support-tier documentation, case studies or third-party validation of enterprise SSO rollout, and clarity on org-wide admin/governance controls.
- [claimed-docs] “OpenHands Enterprise provides commercial capabilities and support for organizations that need licensed self-hosting or managed deployment op…”
- [claimed-docs] “Enterprise SAML / SSO”
- [claimed-docs] “OpenHands Cloud is the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure.”
Julesnone0/10No evidence of enterprise licensing, SSO, or commercial support offerings; docs mention only individual API keys and usage-tier limits, and community discussion references a free beta plan, not enterprise deployment. missing for 10: SSO integration, enterprise licensing/contract terms, commercial support SLAs, org-wide admin/rollout tooling.
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “Is Jules free of charge? Yes, for now, Jules is free of charge. Jules is in beta and available without payment while we learn from usage.”
Model flexibility
engineering-leadBring my own LLM or API key so agents run on the model of my choice
weight 2 · round to OpenHandsGitHub docs explicitly state OpenHands can be used with any LLM, and pricing docs confirm both Open Source and Cloud plans support bringing your own LLM keys, directly matching the story. Missing for 10: independent/hands-on verification of BYO-key setup and any model-specific limitations or edge cases.
- [github] “Use with any LLM”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Julesnone0/10No evidence that Jules supports bringing a custom LLM or third-party API key; it is tied to Google's own models (Gemini), with documentation only mentioning Jules's own API key for accessing Jules itself, not for configuring underlying model providers. missing for 10: any mention of BYO-LLM support, model selection options, or third-party API key configuration.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “Priority access to the latest models, starting with Gemini 3 Pro”
Usage quotas
engineering-leadSee and manage plan-based daily task and concurrency limits for agent workflows
weight 2 · round to JulesOpenHandsnone0/10Evidence shows pricing page mentions plan features like SSO, API access, and BYO LLM keys, but nothing about daily task limits, concurrency limits, or any management dashboard for such limits.
- [claimed-docs] “Hosted cloud access from desktop and mobile”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “Enterprise SAML / SSO”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Jules docs reference a usage-limits page with plan-tiered daily task numbers (e.g., 300/60) and community confirms daily task limits exist and change between releases (60→15 on free plan), showing plan-based limits are real and documented. However there's no evidence of an engineering-lead-facing dashboard or admin console to monitor team usage, set/adjust concurrency, or manage limits across a team — only a static limits reference page and anecdotal user experience of hitting caps. Missing for 10: team/org usage dashboard, per-user concurrency visibility, ability to configure or request limit changes, and any admin/lead-specific management UI.
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnOpenHandsnone0/10No evidence of data residency/region selection controls anywhere in the pack; only self-hosting options (local/Docker/VM) are mentioned, which is a workaround, not a documented region-choice feature for the managed/cloud offering.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnOpenHandsnone0/10No evidence in the pack addresses data usage for AI model training, opt-out policies, or any privacy commitments regarding training data; the pack only covers CLI usage, deployment options, and integrations.
ai-native userControl data retention and deletion
weight 2 · round drawnOpenHandsnone0/10No evidence pack items describe data retention policies, deletion controls, or user-facing settings for managing stored conversation/data lifecycle; while OpenHands is open-source and self-hostable (implying some inherent control), no explicit retention/deletion feature or documentation is cited.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnOpenHandsnone0/10No evidence in the pack mentions telemetry, usage tracking, analytics, or opt-out settings; the pack covers CLI features, agent orchestration, and pricing tiers only.
Repo integration — stories about repo integration in this arenaRepo integration
Stories about repo integration in this arena
Chat integration
developerTag an agent in a chat thread to discuss and delegate a bug or task
weight 2 · round to OpenHandsOpenHands documents a Slack integration where the agent watches for @openhands mentions, opens a conversation with the message context, and replies when done (openhands-docs-11), which matches tagging an agent in a chat thread to delegate a task. However this is limited to Slack specifically (not other chat platforms), appears tied to the Cloud/automation feature set rather than the core product, and lacks independent/hands-on corroboration. Missing for 10: support for other chat platforms (e.g., Teams, Discord), independent verification of the Slack flow working in practice, and detail on how delegated context/threading is preserved during the exchange.
- [claimed-docs] “Watch Slack channels for @openhands mentions, open a conversation with the message context, and reply when the agent finishes.”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Jules supports delegating tasks via a GitHub issue label ('jules') and via API messages to the agent, which is a form of task delegation but not the classic 'tag an agent in a chat thread to discuss' interaction pattern; one community user built a custom MCP server to dispatch Jules tasks from VS Code Copilot Chat, showing it's possible but not a native feature. missing for 10: native chat-thread tagging/mention UI, evidence of back-and-forth discussion in a thread before delegation, first-party support for Slack/Teams-style @mentions.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
Knowledge context
developerAdd a context file describing my codebase conventions so agents generate more relevant plans and code
weight 3 · round to JulesOpenHandsnone0/10No evidence in this pack mentions a repo-level context/convention file (e.g., microagents, custom instructions, or similar) that developers can add to guide agent behavior; the pack only covers CLI usage, deployment, integrations, and pricing.
Jules automatically reads an AGENTS.md file in the repo root to guide its plans and code generation, directly matching the story of adding a context file describing codebase conventions. Missing for 10: no independent/hands-on confirmation of how well conventions in AGENTS.md actually improve plan/code relevance, and no detail on supported format/scope beyond the single doc mention.
- [claimed-docs] “Jules now automatically looks for a file named AGENTS.md in the root of your repository.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
developerQuery generated documentation for any public or private repository
weight 1 · round drawnOpenHandsnone0/10No evidence of a documentation-generation or repo-doc-querying feature; OpenHands is a coding agent focused on tasks, automations, and integrations, not a repo documentation search/query tool.
Julesnone0/10Evidence describes Jules as a task-based coding agent that clones repos, edits code, and creates PRs, but there is no mention of generating or letting users query documentation for a repository (public or private). missing for 10: any docs-generation feature, a documentation query/search interface, evidence of indexing repo docs for Q&A.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
Project management integration
product-managerConnect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly
weight 3 · round to OpenHandsOpenHands documents automation workflows that integrate with Linear (and Slack/GitHub) via webhooks/schedules, showing some issue-tracker connectivity, but Jira, ClickUp, and Monday.com are never mentioned anywhere in the evidence pack — only vague 'and more' language covers them. missing for 10: explicit Jira/ClickUp/Monday.com integrations, docs on ticket management workflows beyond Linear, evidence of two-way ticket manipulation (create/update/close) rather than just webhook triggers.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Version control integration
developerConnect a GitHub repository so an agent can access the code and open pull requests against it
weight 3 · round to JulesGitHub is explicitly listed as an integration target for OpenHands automations, and multiple documented workflows show the agent opening pull requests (fixing failed CI, remediating security alerts, responding to PR-review label events), which implies repo access and PR creation. However, there is no first-party documentation of the actual repo-connection/auth flow (e.g., installing a GitHub App, granting repo scopes) or hands-on confirmation that this works end-to-end. Missing for 10: explicit repo-connection setup docs, evidence of PR creation permissions/scopes, and independent verification of successful PRs opened.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
Jules is built around GitHub repo access: docs describe cloning repos, GitHub issue label assignment, generating diffs/plans, and creating PRs that can be merged on GitHub, corroborated by community reports of receiving usable PRs from Jules-driven tasks. Missing for 10: independent verification of the full connect-repo setup flow and edge-case reliability across repo types.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
developerGrant an agent access to my repositories with a one-click install, without complex setup
weight 2 · round to JulesOpenHands Cloud offers hosted access and integrations (Slack, GitHub, Linear, webhooks) suggesting some repo connection flow, but there is no concrete evidence of a one-click repo install/auth flow — no screenshots, GitHub App install steps, or onboarding walkthrough. missing for 10: documented one-click GitHub/GitLab App install flow, evidence of minimal setup steps, independent confirmation of ease of onboarding.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “OpenHands Cloud is the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure.”
- [claimed-docs] “Hosted cloud access from desktop and mobile”
Jules integrates with GitHub (label-based task assignment, PR creation) and needs repo access to work, implying a connect/authorize flow, and community evidence confirms GitHub integration works well for many users. However, there's no explicit documentation of a literal 'one-click install' onboarding flow or app installation process described in the evidence pack. missing for 10: explicit one-click GitHub App install/authorization flow documentation, independent confirmation of setup simplicity.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates
Quality gates on changes — review flow, required checks, merge protection
Ci remediation
engineering-leadHave failed CI workflows automatically diagnosed and fixed with a proposed pull request
weight 3 · round to OpenHandsOpenHands documents a dedicated automation that detects failed CI workflows, inspects logs, identifies the likely cause, and opens a pull request with a proposed fix — matching the story directly. Missing for 10: independent/hands-on verification of this automation working in practice, and detail on configuration/setup beyond the marketing description.
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Jules offers building blocks—an API for custom automation that can 'automate tasks like bug fixing' and auto-create PRs via automationMode—but there is no direct evidence of a native integration that detects a failed CI workflow and automatically diagnoses/fixes it with a PR; this would require custom API wiring, not an out-of-box feature. Missing for 10: documented CI-failure trigger/integration, evidence of automatic diagnosis of CI logs, and case studies of this specific end-to-end flow.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
developerTrigger an agent from CI/CD pipelines to fix a broken build or failing test
weight 2 · round to OpenHandsOpenHands advertises a workflow that watches for failed CI/CD workflows, inspects logs, identifies causes, and opens a PR with a fix (openhands-docs-9), plus webhook/schedule-triggered automations (openhands-gh-4) and API support for scripting (openhands-docs-13), which together support triggering an agent from CI/CD to fix broken builds. However, there is no concrete example of GitHub Actions/CI pipeline configuration, no evidence of test-failure-specific triggers, and no independent/hands-on confirmation that this works as described. Missing for 10: explicit CI pipeline integration docs/examples, test-failure-specific triggers, third-party verification.
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “The Software Agent SDK is a composable Python library for building agents that work with code.”
Jules exposes a documented API for creating custom workflows and automating tasks like bug fixing (jules-docs-9, jules-docs-10, jules-docs-12), which could be wired into a CI/CD pipeline to trigger a fix, and community evidence shows users building custom integrations to dispatch Jules tasks programmatically (jules-comm-19). However, there is no first-party CI/CD-specific integration (e.g. GitHub Actions step, build-failure webhook) or documented example of triggering Jules to fix a broken build/failing test directly from a pipeline. missing for 10: explicit CI/CD pipeline integration/example, evidence of triggering on build/test failure events, hands-on confirmation of this specific workflow succeeding.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
Diff review
developerConfigure an agent to automatically open a pull request when its task completes
weight 2 · round to JulesOpenHands documents workflow automations that open pull requests automatically for specific triggers (failed CI, security alerts) via [openhands-docs-9] and [openhands-docs-10], showing the underlying capability exists. However, there's no direct documentation of configuring a general coding-task agent to auto-open a PR upon arbitrary task completion — the evidence only covers specific automation templates (bug-fix, security remediation) rather than a generic 'open PR on task completion' setting. Missing for 10: explicit config option/flag for auto-PR-on-completion in standard task workflows, independent/hands-on confirmation of this behavior.
- [claimed-docs] “Detect failed workflows, inspect logs, identify the likely cause, and open a pull request with a proposed fix.”
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Jules natively creates a PR when a task completes (jules-docs-8), and the API exposes an `automationMode` field to configure automatic PR creation, defaulting to off unless explicitly configured (jules-docs-11), directly matching the story of configuring automatic PR-on-completion. Community reports corroborate that Jules produces usable PRs and users regularly review PRs it opens (jules-comm-6, jules-comm-12, jules-comm-16). missing for 10: independent hands-on verification of the automationMode toggle specifically, and more detail on configuration options/edge cases (e.g., partial completions, failed tasks) affecting whether a PR is always opened.
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
developerReview a diff of an agent's changes and approve it before it becomes a pull request
weight 3 · round to JulesOpenHands has an auto-approve/manual-approve action mode (openhands-docs-3 implies a default confirmation step exists before auto-approve is enabled) and can open PRs after agent work, suggesting some human-in-the-loop gating exists, but there's no explicit documentation of a diff review UI or an approval gate specifically before PR creation. missing for 10: explicit diff-review interface, documented approve/reject step tied to PR creation, evidence of a review-before-merge workflow.
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
Jules explicitly provides a diff of changes for review and approval before it creates/publishes a PR, with plan approval and diff approval steps documented; community mentions confirm PRs are generated for review. Missing for 10: independent hands-on confirmation of the diff-approval UI flow itself (comments focus on PR quality/output rather than the diff-review step specifically).
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
Pr review automation
ai-native userHave incoming issues automatically triaged with severity suggested and routed to the right owner
weight 2 · round to OpenHandsopenhands-docs-7 directly describes an automation that watches for new bugs/incidents, gathers logs and recent changes, suggests severity, and routes the issue to the right owner — matching the story closely, backed by GitHub-documented webhook/scheduled automation infrastructure (openhands-gh-4). Missing for 10: independent/hands-on verification of triage accuracy and routing correctness beyond vendor's own site copy.
- [claimed-docs] “Watch for new bugs or incidents, gather logs and recent changes, suggest severity, and route the issue to the right owner.”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
Julesnone0/10Jules is a coding agent that executes tasks assigned via label or API, but there is no evidence of automatic issue triage, severity classification, or routing to owners — the closest feature is manual label-based task assignment, not automated triage. Missing for 10: any mention of severity scoring, triage logic, or owner-routing automation.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
engineering-leadHave every pull request automatically reviewed with AI-generated inline comments
weight 3 · round to OpenHandsOpenHands documents a specific automation that watches for a configurable PR label, inspects full PR context, and posts an AI review comment, which directly matches automated PR review with AI-generated comments. However, it is label-triggered rather than automatic on every PR, and posts once per label event rather than full inline (line-by-line) comments, so it's a partial match to 'every pull request' with 'inline comments'. missing for 10: evidence of automatic triggering on all PRs without manual labeling, confirmation of true inline (line-level) code comments vs a single summary comment, and independent/hands-on verification of this workflow in practice.
- [claimed-docs] “Watch for a configurable label on pull requests, inspect full PR context, and post an AI review comment once per label event.”
Julesnone0/10Jules is documented as a task-execution agent that clones repos, makes code changes, and creates its own PRs for approval — not as a bot that automatically reviews incoming pull requests with inline comments. The only tangential mention is jules-docs-9's passing reference to 'automate tasks like ... code reviews' via API, but there is no documentation of an automatic inline-comment review gate applied to every PR.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
Readiness checks
engineering-leadAutomatically fix failing agent-readiness criteria in my repository
weight 1 · round drawnOpenHandsnone0/10The evidence shows OpenHands can fix failing CI workflows, remediate security alerts, and review PRs, but nothing ties this to a defined 'agent-readiness' criteria/checklist concept that an engineering lead could gate on and auto-remediate. Missing for 10: any mention of agent-readiness scoring, a checklist/criteria framework, or evidence that OpenHands detects and fixes failures against such a standard.
Julesnone0/10Jules can be assigned generic bug-fix/coding tasks and reads an AGENTS.md file if present, but there is no evidence of it detecting or automatically remediating 'agent-readiness' quality-gate criteria (e.g., missing/invalid AGENTS.md, agent-compatibility checks) as a review gate — it only consumes such files, it doesn't audit or fix them as a compliance gate.
- [claimed-docs] “Jules now automatically looks for a file named AGENTS.md in the root of your repository.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
engineering-leadRun a readiness report that evaluates how ready my repository is for autonomous agents
weight 2 · round drawnOpenHandsnone0/10No evidence of a repository readiness assessment/report feature for autonomous agents; OpenHands' evidence covers task execution, automations, and integrations but nothing about scoring or auditing repo readiness.
Julesnone0/10No evidence of a readiness report, scorecard, or repository-assessment feature for autonomous agent suitability; evidence only covers task execution, PR creation, and API usage. Missing for 10: any readiness-scoring feature, repo audit/checklist, or report output evaluating agent-readiness.
Security remediation
engineering-leadHave security alerts automatically validated and remediated with an opened pull request
weight 2 · round to OpenHandsopenhands-docs-10 explicitly describes an automation that reviews security alerts, validates the finding, updates affected code, and opens a pull request with the remediation, directly matching the story. This is a first-party vendor claim without independent/hands-on corroboration or detail on validation accuracy. Missing for 10: independent verification of remediation quality, and detail on how 'validated' findings are confirmed before code changes.
- [claimed-docs] “Review security alerts, validate the finding, update affected code, and open a pull request with the remediation.”
Julesnone0/10Jules is a general coding agent that can be assigned tasks (via GitHub label or API) and opens PRs with diffs for review, but there is no evidence of any security-alert scanning, vulnerability detection/validation, or a workflow that automatically triggers remediation from a security alert (e.g., Dependabot/CVE integration). The evidence pack shows generic task-to-PR flow, not a security-alert-specific pipeline.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
Concurrent execution
engineering-leadRun many agent tasks concurrently to scale delivery throughput
weight 3 · round drawnOpenHands supports running agents across many backends (local, Docker, VMs, cloud) and building automations triggered by schedules or webhooks (Slack, GitHub, Linear), plus API access for scripting — all of which enable running multiple tasks in parallel across infrastructure. However, there is no explicit documentation of a dashboard or orchestration layer for managing many concurrent agent runs, concurrency limits, or throughput metrics for an engineering lead. Missing for 10: explicit multi-task concurrency/orchestration UI, stated concurrency limits or scaling guarantees, and independent evidence of teams running many parallel agents successfully.
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “OpenHands Cloud is the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure.”
Jules architecture (cloud VMs per task, API for creating tasks, GitHub issue labeling) supports dispatching many tasks in parallel, and usage-limit docs explicitly reference 'agent-heavy workflows' with numeric caps (e.g. 300/60), and community reports (jules-comm-11, jules-comm-16) describe handling large task volumes successfully. However, other evidence shows real caveats: daily task limits were cut from 60 to 15 on the free plan (jules-comm-8), tasks can get stuck in loops with no stop button (jules-comm-14), and users report needing heavy babysitting (jules-comm-18), undercutting smooth high-throughput scaling. Missing for 10: dedicated documentation of a concurrency/queue dashboard, enterprise-tier concurrency guarantees, and independent benchmarks confirming reliable parallel execution at scale.
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
engineering-leadCreate agent sessions on behalf of other users in my organization
weight 2 · round drawnOpenHandsnone0/10Evidence shows enterprise features like SSO/SAML and API access, which imply some multi-user org management, but there is no evidence of an engineering-lead being able to create or launch agent sessions on behalf of other users. Missing for 10: any documentation of delegated session creation, admin-initiated sessions, or role-based session provisioning for other org members.
- [claimed-docs] “Enterprise SAML / SSO”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “OpenHands Enterprise provides commercial capabilities and support for organizations that need licensed self-hosting or managed deployment op…”
Julesnone0/10No evidence of any organization, team, or admin/delegation features allowing an engineering-lead to create sessions on behalf of other users; Jules docs only describe individual API keys, per-user GitHub connections, and personal task creation. Missing for 10: org/team management, delegated session creation, role-based admin controls, multi-user account provisioning.
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Jules needs access to your repositories in order to work.”
Deployment flexibility
developerUse a managed cloud offering to run agents without operating my own backend infrastructure
weight 2 · round drawnOpenHands Cloud is explicitly documented as the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure, with hosted access, API support, and bring-your-own-LLM-key options detailed on the pricing page. Missing for 10: independent/hands-on corroboration of the cloud offering's reliability and scale, and more detail on infrastructure limits or SLAs.
- [claimed-docs] “OpenHands Cloud is the managed commercial service for running OpenHands without operating your own backend and sandbox infrastructure.”
- [claimed-docs] “Hosted cloud access from desktop and mobile”
- [claimed-docs] “API support for automation and scripting”
- [claimed-docs] “OpenHands Open Source and OpenHands Cloud plans offer support to bring your own LLM keys.”
Jules is explicitly a cloud-hosted agent service: tasks run in Google-managed VMs, integrated with GitHub, with an API and usage limits/plans, requiring no self-hosted backend. Community evidence corroborates real-world usage at scale (75% success on customer PRs, teams dispatching many tasks) though some users report reliability/UI issues. Missing for 10: no independent infra/SLA details, no discoverable OpenAPI/docs.md confirming API completeness, and mixed community reliability reports.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
engineering-leadSelf-host agent infrastructure locally, in containers, or on my own VMs
weight 2 · round to OpenHandsGitHub docs explicitly state agents can be run locally, in Docker, on VMs, or any agent server backend, and Enterprise offers licensed self-hosting for organizations. This directly matches the engineering-lead's need for flexible self-hosted deployment. Missing for 10: independent/hands-on verification of self-hosting setup complexity and no detailed self-host deployment guide excerpt in evidence.
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [claimed-docs] “OpenHands Enterprise provides commercial capabilities and support for organizations that need licensed self-hosting or managed deployment op…”
- [github] “Switch between local, remote, and cloud agents without losing focus”
- [github] “Run OpenHands, Claude Code, Codex, Gemini, or any ACP-compatible agent across local, remote, and cloud backends.”
Julesnone0/10Jules is explicitly cloud-hosted: it runs tasks in Google's own VMs, and community feedback confirms it is not local (jules-comm-1). There is no evidence of any self-hosting, on-prem, container, or private-VM deployment option.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
Headless automation
developerRun an agent headlessly inside CI/CD pipelines and shell scripts
weight 2 · round to OpenHandsDocs show a CLI mode with an 'auto-approve all actions' flag for non-interactive runs, and GitHub evidence describes running agents locally/Docker/VMs and triggering automations 'on a schedule or in response to webhook events,' plus API support for automation/scripting — all consistent with headless CI/CD use. Missing for 10: an explicit CI/CD pipeline example (e.g., GitHub Actions YAML snippet) or independent hands-on confirmation of headless exit-code/scripting behavior.
- [claimed-docs] “Auto-approve all actions (use with caution)”
- [probe] “official CLI documented at https://docs.openhands.dev/usage/how-to/cli-mode”
- [github] “Create automations and workflows that integrate with Slack, GitHub, Linear, and more. Run on a schedule or in response to webhook events”
- [github] “Run agents locally, in Docker, on VMs, or anywhere you can run an agent server backend”
- [claimed-docs] “API support for automation and scripting”
Jules exposes a public API (create tasks, send messages, optional automationMode) explicitly pitched for 'custom workflows' and embedding into other tools, which can be scripted/curled headlessly, and a community member built an MCP server hitting the API to dispatch tasks programmatically. However there's no documented CLI, no explicit CI/CD pipeline examples (e.g., GitHub Actions), and Jules always runs its own cloud VM rather than a lightweight headless process invokable inline in a shell script/pipeline step. missing for 10: official CLI or CI/CD pipeline integration docs (e.g. GitHub Actions step), examples of shell-script invocation, confirmation that API calls run synchronously enough for pipeline gating.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
Not comparable on these axes
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableOpenHandsn/aOpenHands is itself an AI coding agent, so serving as an MCP server for other agents is a different product role (agent-as-client, not server) — this axis is out of scope per the agent-role exception. No evidence shows OpenHands running as an MCP server endpoint.
Julesn/aJules is itself a coding agent (client role), and the evidence only shows a REST API plus a third-party/community-built MCP server (jules-comm-19), not an official first-party MCP server exposing Jules as a tool endpoint. Per the agent-role rule, this axis is out of category rather than a failed capability.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”