Skip to content

Software Factory Arena

Factory vs Jules

Jules wins · 1822 (31 drawn)

Agenticness — how well agents can access and operate the productAgenticness

How well agents can access and operate the product

Agent access

  1. ai-native userPoint an agent at llms.txt or agent-oriented docs

    weight 2 · round to Factory
    Factoryfullprobed8/10

    Factory hosts an actual llms.txt at docs.factory.ai/llms.txt (HTTP 200) confirmed by direct probe, and its docs describe an agent-native platform with structured agent-oriented documentation (droid-cli, agent-readiness, missions) that an agent could be pointed at. missing for 10: independent/hands-on confirmation that an agent successfully consumes llms.txt in practice, and no explicit vendor statement encouraging users to point agents at llms.txt.

    • [probe] PROBE llms.txt: HTTP 200 at https://docs.factory.ai/llms.txt # Factory Documentation > Documentation for Factory, the agent-native software…
    • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
    • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
    Julesnone0/10

    Probes show no llms.txt, docs.md, or machine-readable API spec exist at expected paths, and no evidence Jules can be pointed at such agent-oriented documentation formats; Jules does support AGENTS.md for repo context but that's a different mechanism than consuming llms.txt-style docs.

    • [probe] PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt
    • [probe] PROBE docs-md: HTTP 404 at https://jules.google/docs.md
    • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
  2. ai-native userRun the product headlessly / in CI for automation

    weight 2 · round to Factory
    Factoryfullprobed8/10

    Factory documents droid exec as a one-shot CLI command explicitly designed for CI/CD pipelines, shell scripts, and batch processing, with tiered autonomy controls for unattended operation. This directly satisfies headless/CI automation. Missing for 10: independent/hands-on third-party verification of CI usage and more detail on exit codes/output formats for pipeline integration.

    • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
    • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
    • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
    • [probe] official CLI documented at https://docs.factory.ai/droid-cli/quickstart

    Jules provides a public API with API keys for creating custom workflows, sending tasks programmatically, and an optional automationMode to auto-create PRs, plus a community example of building an MCP server to dispatch tasks from another tool — this directly supports headless/CI-style automation. Missing for 10: no official CI/CD integration examples (e.g. GitHub Actions), no documented webhook/polling pattern for task completion, and no independent verification of reliability at scale in automated pipelines.

    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
    • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
    • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
    • [claimed-docs] To send a message to the agent:
    • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
  3. ai-native userPlug MCP servers into this product so it can use their tools

    weight 3 · round to Factory
    Factorypartialclaimed6/10

    Docs explicitly state Droid CLI can connect MCP tools alongside Jira/Notion/Slack/Linear/PagerDuty integrations, confirming MCP client support. However, there is no detail on setup/configuration process, supported transport types, or independent/hands-on confirmation of MCP tool usage in practice. missing for 10: configuration details for adding MCP servers, examples of MCP tools being invoked, independent verification of functionality.

    • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems
    • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
    Julesnone0/10

    No evidence that Jules can connect to or consume external MCP servers as tool providers; the only MCP-related evidence (jules-comm-19) describes someone building an MCP server that calls INTO the Jules API, which is the reverse integration direction, not Jules plugging in MCP servers for its own tool use.

    • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
  4. ai-native userUse an official CLI

    weight 2 · round to Factory
    Factoryfullprobed9/10

    Factory ships an official Droid CLI with documented quickstart, terminal/editor/Git integration, bash mode, one-shot exec for CI/CD, and slash-command features—clearly an official CLI for AI-native workflows. Missing for 10: independent/hands-on third-party corroboration beyond vendor docs.

    • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
    • [claimed-docs] Press ! to toggle bash mode and run shell commands directly without AI interpretation.
    • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
    • [probe] official CLI documented at https://docs.factory.ai/droid-cli/quickstart
    Julesnone0/10

    Evidence shows Jules offers a web app, GitHub integration, and a REST API for custom workflows, but no official CLI tool is documented or mentioned anywhere in the evidence pack.

    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
    • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
  5. ai-native userDrive the product through a documented public API

    weight 3 · round to Jules
    Factorypartialprobed6/10

    Factory documents a public API for creating and driving Droid sessions (factory-docs-13) plus CLI/exec automation paths (droid exec, CI/CD) that give programmatic control, but there's no discoverable OpenAPI/swagger spec (probe found only 404s), suggesting the API reference is thinner or less formally specified than a full REST API story would require. missing for 10: a published OpenAPI/swagger schema, broader API endpoint coverage beyond sessions, independent third-party confirmation of API usage.

    • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
    • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
    • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
    • [probe] PROBE openapi: all candidate paths 404 (https://docs.factory.ai/openapi.json, https://docs.factory.ai/swagger.json, https://docs.factory.ai/…
    Julesfullprobed8/10

    Jules ships a documented public API (developers.google.com/jules/api) with API key management, custom workflows, sending messages to the agent, and automationMode config; a community user independently built an MCP server on top of the API confirming real-world programmatic access. Missing for 10: no discoverable OpenAPI/swagger spec or llms.txt (probes returned 404s), reducing machine-readability/self-service tooling confidence.

    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
    • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
    • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
    • [claimed-docs] To send a message to the agent:
    • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
    • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
  6. ai-native userIssue scoped/least-privilege API credentials for an agent

    weight 2 · round drawn
    Factorynone0/10

    No evidence of scoped or least-privilege API credential/token issuance for agents; docs mention API sessions and integrations (Jira, Slack, MCP) but nothing about credential scoping, permission tiers for API keys, or least-privilege access control. Tiered autonomy (factory-docs-7) governs action confirmation, not credential scoping.

      Julesnone0/10

      Jules documents a basic API key creation flow (max 3 keys) but provides no evidence of scoped or least-privilege permission controls — no mention of scopes, roles, or restricted-access tokens for the agent.

      • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
      • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
      • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
    • ai-native userBuild against official SDKs

      weight 2 · round to Jules
      Factorypartialprobed4/10

      Factory documents an API reference for creating/driving Droid sessions (factory-docs-13) and MCP/CLI integration points, suggesting some programmatic access, but there is no evidence of an official language SDK (Python/JS/etc.) and the OpenAPI spec probe returned 404s, indicating thin or unpublished API tooling. missing for 10: no named official SDK packages, no OpenAPI/schema availability, no code samples showing SDK usage.

      • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
      • [probe] PROBE openapi: all candidate paths 404 (https://docs.factory.ai/openapi.json, https://docs.factory.ai/swagger.json, https://docs.factory.ai/…
      • [probe] official MCP server documented at https://docs.factory.ai/harness/mcp
      Julesfullprobed7/10

      Jules ships an official public API with documented endpoints, API key management, and use cases for building custom workflows/integrations, and a community member confirms building a personal MCP server against the Jules API. missing for 10: no official OpenAPI/SDK spec discoverable (openapi probes 404), no first-party language SDKs mentioned, and no independent SDK-quality corroboration beyond one community integration example.

      • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
      • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
      • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
      • [claimed-docs] To send a message to the agent:
      • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
      • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
    • ai-native userSubscribe to events via webhooks

      weight 2 · round drawn
      Factorynone0/10

      No evidence of webhook subscription support; docs mention API sessions, MCP tools, and integrations (Jira, Slack, etc.) but nothing about webhooks for event notifications. missing for 10: any webhook documentation, event types, subscription endpoints, or third-party confirmation of webhook support.

        Julesnone0/10

        Evidence shows Jules has an API for creating tasks/sending messages, notifications for task completion, and GitHub integration, but there is no mention of webhooks or event subscription mechanisms anywhere in the docs or community evidence. Missing for 10: any documentation of a webhook endpoint, event subscription API, or push-based notification mechanism to external systems.

        • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
        • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
        • [claimed-docs] To send a message to the agent:
        • [claimed-docs] You’ll be notified when the task completes or needs your input.

      Agentic features

      1. ai-native userSet up automations that run autonomously in the background

        weight 2 · round drawn
        Factorypartialclaimed6/10

        droid exec supports one-shot autonomous runs with tiered autonomy suited for CI/CD, scripts, and batch automation, and the API supports creating/driving Droid sessions programmatically, which enables background automation setups. However, there's no evidence of a scheduling/trigger system (e.g., cron-like or event-driven automations) or a dedicated 'automations' dashboard for persistent background jobs. Missing for 10: native scheduling/triggers for autonomous background runs, independent hands-on confirmation of unattended long-running automations, and a dedicated automations management UI.

        • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
        • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
        • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
        • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.

        Jules supports background/autonomous execution via GitHub label-triggered tasks and a public API for building custom automations (e.g., automated bug-fixing, code review) with an optional automationMode for auto-PR creation, and it runs tasks in a cloud VM without needing the user present. However, docs also show a human-approval gate (plan approval, PR review) rather than fully hands-off automation, and community reports note tasks getting stuck, hitting limits, or needing babysitting, undercutting reliability of unattended runs. Missing for 10: evidence of true scheduled/cron-style recurring automations, and independent confirmation that automations run to completion without manual intervention.

        • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
        • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
        • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
        • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
        • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
        • [community] Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.
        • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
      2. ai-native userDelegate tasks to a built-in AI assistant inside the product

        weight 3 · round drawn
        Factoryfullclaimed8/10

        Factory's core product is built around delegating tasks to Droid agents via App, CLI, or exec, with clear docs on task delegation, custom droids, and session management. missing for 10: independent/hands-on corroboration beyond vendor docs, and quantified reliability of delegated task completion.

        • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
        • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
        • [claimed-docs] delegate scoped tasks to Custom Droids via /droids, and package reusable procedures as Skills with /skills or /create-skill
        • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
        • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.

        Jules is exactly this kind of built-in agentic assistant: users delegate coding tasks via GitHub issue labels, the web app, or API, and Jules autonomously plans, clones repos, runs in a VM, and returns diffs/PRs for approval (jules-docs-2,3,4,7,8,9). Community evidence corroborates real-world task delegation and usable output (jules-comm-2, jules-comm-11, jules-comm-15, jules-comm-16), though mixed reports of loops, babysitting needs, and reliability issues (jules-comm-14, jules-comm-18) temper quality. Missing for 10: independent benchmarking of task success rates and evidence of consistent reliability at scale.

        • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
        • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
        • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
        • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
        • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
        • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
        • [community] I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.
        • [community] I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…
        • [community] I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…
        • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
        • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
        • [community] Do people really find Jules useful? I find it needs babysitting much more than Cursor.
      3. ai-native userOperate the product with natural-language commands

        weight 2 · round drawn
        Factoryfullprobed8/10

        Factory's Droid CLI and App let users delegate tasks, run commands, and manage workflows via natural-language prompts, with bash-mode fallback and slash commands (/droids, /skills, /readiness-report) layered on top of NL interaction, showing deep agentic natural-language operation across CLI, exec, and Missions. missing for 10: independent/hands-on user reports confirming NL command quality and reliability outside vendor docs.

        • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
        • [claimed-docs] Press ! to toggle bash mode and run shell commands directly without AI interpretation.
        • [claimed-docs] delegate scoped tasks to Custom Droids via /droids, and package reusable procedures as Skills with /skills or /create-skill
        • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
        • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.
        • [probe] official CLI documented at https://docs.factory.ai/droid-cli/quickstart

        Jules is designed around natural-language task submission and messaging: users submit a task description, Jules generates a plan, and you can send follow-up messages to the agent via API or web app (jules-docs-4, jules-docs-12). Community reports confirm real-world use of NL prompts to drive coding tasks, including iterative feedback and error-message follow-ups (jules-comm-2, jules-comm-15, jules-comm-16). Missing for 10: no detailed documentation of the full range/complexity of natural-language commands supported (e.g., multi-step conversational control, command reference) and no independent benchmark of NL command robustness beyond anecdotal reports.

        • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
        • [claimed-docs] To send a message to the agent:
        • [community] I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.
        • [community] I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…
        • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…

      Api quality

      1. ai-native userExplore an interactive API reference with runnable examples

        weight 2 · round drawn
        Factorynone0/10

        Factory does have an API reference section (docs.factory.ai/api-reference/sessions) but there's no evidence of an interactive, runnable-example explorer — the probe for standard OpenAPI/Swagger interactive docs (openapi.json, swagger.json, etc.) returned 404 at all candidate paths, suggesting no live 'try it' interface exists.

        • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
        • [probe] PROBE openapi: all candidate paths 404 (https://docs.factory.ai/openapi.json, https://docs.factory.ai/swagger.json, https://docs.factory.ai/…
        Julesnone0/10

        Jules documents an API (create tasks, send messages, API keys) but there is no evidence of an interactive API reference or runnable code examples; probes for openapi/swagger specs all returned 404s, suggesting no such interactive reference exists.

        • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
        • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
        • [claimed-docs] To send a message to the agent:
        • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
      2. ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

        weight 2 · round drawn
        Factorynone0/10

        Factory has an API reference (sessions endpoints) but probes for standard OpenAPI/swagger spec locations all returned 404, indicating no downloadable machine-readable spec is published; missing for 10: any accessible OpenAPI/swagger JSON file or equivalent machine-readable spec.

        • [probe] PROBE openapi: all candidate paths 404 (https://docs.factory.ai/openapi.json, https://docs.factory.ai/swagger.json, https://docs.factory.ai/…
        • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
        Julesnone0/10

        Jules ships a documented REST API (jules-docs-9..12) but there is no evidence of a downloadable OpenAPI/Swagger spec; explicit probes for llms.txt, docs.md, and standard OpenAPI paths all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), confirming no machine-readable spec is published.

        • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
        • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
        • [probe] PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt
        • [probe] PROBE docs-md: HTTP 404 at https://jules.google/docs.md
        • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
      3. ai-native userTest against a sandbox environment without touching production data

        weight 1 · round to Jules
        Factorynone0/10

        Factory's evidence covers coding agent workflows (CLI, sessions, MCP, readiness reports) but nothing addresses a sandbox environment for testing separate from production data. No mention of sandbox mode, staging environment, or data isolation guarantees.

          Jules runs tasks in an isolated cloud VM that clones the repo, meaning code changes happen in a sandboxed environment rather than directly on production, and PRs must be reviewed/approved before merging to the real branch. However, this is a code-sandbox for making changes, not a dedicated 'test against sandbox data/environment' feature, and there's no evidence of test-data isolation, staging environment provisioning, or explicit protection of production data/services beyond the VM/PR review flow. missing for 10: explicit sandbox test-data isolation, staging/production separation guarantees, independent verification that production systems are never touched.

          • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
          • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
          • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
          • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
        • ai-native userRely on versioned APIs with a documented deprecation policy

          weight 2 · round drawn
          Factorynone0/10

          There is an API reference (sessions) but no evidence of API versioning scheme or a documented deprecation policy; the OpenAPI spec probe even returned 404s across candidate paths, suggesting no formal versioned spec is exposed.

          • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
          • [probe] PROBE openapi: all candidate paths 404 (https://docs.factory.ai/openapi.json, https://docs.factory.ai/swagger.json, https://docs.factory.ai/…
          Julesnone0/10

          Jules has a public API (jules-docs-9, jules-docs-10) but there is no evidence of API versioning scheme or a documented deprecation policy; probes for OpenAPI spec/docs.md/llms.txt all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), suggesting no formal machine-readable API contract or lifecycle documentation is available.

          • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
          • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
          • [probe] PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt
          • [probe] PROBE docs-md: HTTP 404 at https://jules.google/docs.md
          • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…

        Automation depth — how much of the product can run unattendedAutomation depth

        How much of the product can run unattended

        1. ai-native userPerform bulk operations across many items at once

          weight 2 · round to Factory
          Factorypartialclaimed6/10

          droid exec is explicitly documented as a one-shot command 'ideal for CI/CD pipelines, shell scripts, and batch processing,' and the API lets users programmatically create/drive many Droid sessions, both enabling bulk automation across items. However, there's no explicit example, docs, or tooling demonstrating a native 'run across many files/repos/tickets at once' bulk operation feature — it relies on the user scripting droid exec in loops rather than a built-in bulk-operation UI/command. Missing for 10: a dedicated bulk-operation command or documented multi-item batch workflow example, and independent/hands-on evidence of it working at scale.

          • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
          • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
          • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.

          Jules exposes an API that lets users script custom workflows and dispatch tasks programmatically (jules-docs-9, jules-docs-12), and usage-limit tiers explicitly target 'power users & agent-heavy workflows' with up to 300 tasks (jules-docs-13), while community reports confirm it 'handles large quantities of tasks very well' and users built custom dispatch tooling via the API/MCP bridge (jules-comm-11, jules-comm-19). However, there's no documented native UI or endpoint for a single bulk action across many items (e.g., batch-applying one task to many repos/issues) — each task still appears to be created and reviewed individually via GitHub labels or API calls. Missing for 10: a documented batch/bulk endpoint or UI feature, first-party bulk-operation examples, and independent verification of true parallel bulk execution rather than just high per-day task volume.

          • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
          • [claimed-docs] To send a message to the agent:
          • [claimed-docs] Power users & agent-heavy workflows ... 300 ... 60
          • [community] I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…
          • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
        2. ai-native userDefine rules that trigger actions automatically on events

          weight 3 · round to Jules
          Factorypartialclaimed4/10

          Factory's droid exec is explicitly designed to run as a one-shot command in CI/CD pipelines, shell scripts, and batch jobs, which implies it can be wired to external events (e.g., git hooks, CI triggers) with tiered autonomy controlling what runs unattended — but this is an execution mode, not a native rule/trigger definition system where a user declares 'on event X, do Y'. Missing for 10: explicit rule/trigger authoring UI or config, built-in event listeners (e.g., webhook triggers, issue-created triggers), and any documented automation-rules engine beyond CI invocation.

          • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
          • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
          • [claimed-docs] Open Software Factory to see your delivery lifecycle as an automation coverage map.

          Jules supports some event-driven automation — e.g., assigning a 'jules' label to a GitHub issue automatically triggers a task, and the API lets developers build custom workflows/automations around Jules (jules-docs-2, jules-docs-9). However, there's no evidence of a general rules/conditions engine where users define arbitrary triggers and actions beyond this specific label mechanism and API calls. Missing for 10: a configurable rules interface (conditions + multiple trigger types), documentation of scheduled/webhook-based triggers, and independent confirmation that automated rule-based triggering works reliably in practice.

          • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
          • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
          • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
        3. ai-native userSchedule recurring jobs or workflows

          weight 2 · round drawn
          Factorynone0/10

          Evidence shows droid exec for CI/CD one-shot automation and API session management, but no documentation of scheduling or recurring/cron-style job execution exists anywhere in the pack. missing for 10: any mention of scheduling, cron, recurring triggers, or timed/repeated workflow execution.

            Julesnone0/10

            Evidence covers task assignment via GitHub labels, API-driven custom workflows, and one-off task automation, but nothing describes recurring/scheduled jobs (e.g., cron-like triggers) — the API docs only mention creating tasks and sending messages, not recurrence.

            • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
            • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
            • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
            • [claimed-docs] To send a message to the agent:
          • ai-native userVersion, review, and roll back my automations

            weight 1 · round to Factory
            Factorypartialclaimed5/10

            Factory supports reviewing diffs and merging via git workflow (factory-docs-1, factory-docs-2), and packages automations as reusable Custom Droids/Skills (factory-docs-5), which implies some git-based versioning, but there is no explicit documentation of a version-history feature for Droids/Skills nor an explicit rollback mechanism for automations themselves. missing for 10: explicit versioning UI/history for Skills/Droids, dedicated rollback command or feature distinct from generic git revert.

            • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
            • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
            • [claimed-docs] delegate scoped tasks to Custom Droids via /droids, and package reusable procedures as Skills with /skills or /create-skill
            Julesnone0/10

            Jules produces PRs and diffs via GitHub but there is no evidence of a mechanism to version, review history of, or roll back automations/tasks themselves (as opposed to code changes tracked by git/GitHub). missing for 10: automation versioning/history feature, rollback/undo of Jules tasks, changelog or audit trail for automations.

            Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation

            End-to-end implementation by the agent — multi-file changes, task completion

            End to end feature delivery

            1. ai-native userHave an agent automatically generate and run tests to validate its own code changes before proposing them

              weight 2 · round to Jules
              Factorynone0/10

              Docs mention integration with 'tests' as part of Git workflow and bash-mode shell execution, plus readiness-report/readiness-fix commands, but none of this describes an agent autonomously generating and running tests to validate its own code changes before proposing a diff. Missing for 10: explicit documentation of automated test generation, self-validation loop, or evidence droid runs tests as a pre-proposal gate.

              • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
              • [claimed-docs] Press ! to toggle bash mode and run shell commands directly without AI interpretation.
              • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
              • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI

              Jules runs code changes in an isolated VM and generates diffs/PRs for review, and one community report notes it is 'good at making tests for very specific/isolated functions,' implying some test generation capability, but there is no first-party documentation describing an explicit test-generate-and-run validation loop before proposing changes. missing for 10: official docs on automated test generation/execution as a validation step, evidence of test results being surfaced to users before PR creation, and independent confirmation across varied codebases.

              • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
              • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
              • [community] I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…
            2. developerHave an agent autonomously diagnose and fix a reported bug

              weight 3 · round to Jules
              Factorypartialclaimed6/10

              Factory's Droid CLI/exec and delegated task workflow support autonomous code changes (diff review and merge) and integrations like Jira/Linear/PagerDuty for bug tracking, plus tiered autonomy for unattended operation, which together plausibly support autonomous bug diagnosis and fixing. However, no evidence explicitly describes an end-to-end 'diagnose a reported bug from ticket to verified fix' workflow or hands-on validation of bug-fixing accuracy. Missing for 10: explicit bug-diagnosis workflow documentation, independent/hands-on evidence of successful autonomous bug fixes, and details on root-cause diagnosis capability.

              • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
              • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
              • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
              • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems
              • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI

              Jules is designed to autonomously diagnose and fix issues: it clones the repo into a VM, generates a plan, makes code changes, and opens a PR for review, and can be triggered directly from a GitHub issue via the 'jules' label (a common bug-report workflow). Community evidence corroborates real-world bug-fixing use, including a case where it failed then successfully fixed the issue after being given the error message. Missing for 10: independent benchmarking on bug-fix success rate and more consistent evidence across complex codebases (some reports of failures/loops on harder tasks).

              • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
              • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
              • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
              • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
              • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
              • [community] I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…
              • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
            3. product-managerGo from a mockup or design to a working implementation without an engineering handoff

              weight 2 · round to Factory
              Factoryfullclaimed7/10

              Factory's agent-readiness docs explicitly describe the exact scenario: "A designer shares a mockup, and the system implements it without handoffs" (factory-docs-8), backed by delegation/review/merge workflow (factory-docs-1) and readiness tooling (factory-docs-9, factory-docs-10) that lets non-engineers trigger and validate implementation. Missing for 10: independent or hands-on corroboration of a PM-specific end-to-end mockup-to-merge case study, and detail on how a non-technical PM reviews/approves the diff without engineering involvement.

              • [claimed-docs] A designer shares a mockup, and the system implements it without handoffs.
              • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
              • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
              • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI
              Julesnone0/10

              Evidence shows Jules operates on code repositories via GitHub issues/tasks, plans, diffs, and PRs, but there is no mention of accepting mockups/designs as input or any PM-oriented no-code workflow — usage still requires repo access, issue creation, and reviewing technical diffs/plans. Nothing in the pack demonstrates a design-to-implementation path bypassing engineering.

              • [claimed-docs] Jules needs access to your repositories in order to work.
              • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
              • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
              • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
              • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
            4. developerHave an agent implement a requested feature end-to-end, including writing tests

              weight 3 · round to Factory
              Factorypartialclaimed7/10

              Factory's docs show agents can be delegated end-to-end feature tasks (delegate, diff review, merge), run in terminal/CI with test execution and git workflow, and orchestrate multi-feature projects via Missions, but no explicit first-party evidence confirms the agent autonomously writes tests as part of implementation. missing for 10: explicit documentation of test-writing behavior, independent/hands-on verification of end-to-end feature delivery.

              • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
              • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
              • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
              • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.

              Jules is explicitly built for autonomous end-to-end coding: it clones the repo, generates a plan, modifies files, produces a diff, and opens a PR (jules-docs-3/4/7/8), and one user specifically notes it's 'good at making tests for very specific/isolated functions' (jules-comm-11), with others confirming usable PRs and time savings (jules-comm-6, jules-comm-16). However, community evidence shows significant caveats: it struggles on complex/monorepo codebases, gets stuck in loops without a stop button, needs heavy babysitting, and often fails to complete tasks unassisted (jules-comm-3, jules-comm-9, jules-comm-10, jules-comm-13, jules-comm-14, jules-comm-18). Missing for 10: consistent reliability across complex real-world features, broader evidence of comprehensive test coverage (not just isolated functions), and independent benchmarks confirming end-to-end success rate.

              • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
              • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
              • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
              • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
              • [community] I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…
              • [community] its way better than the github thing in my experience it produces usable PRs
              • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
              • [community] I've been playing with it, and I've been generally not impressed. There are obvious annoying UI bugs and the output isn't very good for anyt…
              • [community] I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…
              • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
              • [community] Do people really find Jules useful? I find it needs babysitting much more than Cursor.

            Environment setup

            1. developerHave an agent automatically clone the repo, install dependencies, and configure its own working environment

              weight 2 · round to Jules
              Factorypartialclaimed4/10

              Factory's Droid CLI/exec can run shell commands autonomously (bash mode, tiered autonomy for unconfirmed operations) and operates within a repo's terminal/Git workflow, which implies it could run clone/install commands, but no evidence explicitly describes an agent autonomously cloning a repo or bootstrapping its own dev environment from scratch. missing for 10: explicit documentation of automated repo cloning, dependency installation, or environment provisioning as a first-class capability, and any hands-on example showing this workflow.

              • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
              • [claimed-docs] Press ! to toggle bash mode and run shell commands directly without AI interpretation.
              • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
              • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.

              Jules' core workflow is documented: it runs in a cloud VM, clones the repo, installs dependencies, and configures its environment automatically before making changes, corroborated by community usage reports of automated PR generation. missing for 10: independent technical verification of dependency-install robustness across complex/monorepo setups, and some community reports note environment/config confusion in bespoke or monorepo projects.

              • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
              • [claimed-docs] Jules needs access to your repositories in order to work.
              • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
              • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
              • [community] Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…
              • [community] I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…

            Interactive takeover

            1. developerTake over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work

              weight 2 · round to Factory
              Factorypartialclaimed6/10

              Factory explicitly supports multi-surface access (App/web, terminal via Droid CLI, browser) and delegating tasks, reviewing diffs, and merging from any of these surfaces, which implies continuity across surfaces. However, there's no explicit documentation of a 'takeover mid-task' handoff flow (e.g., pausing an in-progress session in one surface and resuming/redirecting it live in another) — the evidence shows task delegation and review/merge but not explicit interactive takeover semantics. Missing for 10: explicit documentation of resuming/redirecting an in-progress session across surfaces, and independent/hands-on confirmation of this handoff working smoothly.

              • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
              • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
              • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.

              Jules supports reviewing plans, diffs, and sending follow-up messages to redirect a task (jules-docs-4, jules-docs-6, jules-docs-7, jules-docs-12), and third parties have wired the API into VS Code/Copilot Chat to dispatch tasks (jules-comm-19), but this is dispatch, not mid-task takeover. A hands-on report explicitly notes there's no 'STOP' button to interrupt a looping task (jules-comm-14), undercutting redirect control, and all takeover is browser/API-based with no native terminal or editor 'take over' UI documented. Missing for 10: documented in-editor/terminal takeover UI, ability to pause/interrupt an in-progress run, and evidence the API-based messaging genuinely redirects rather than just appends instructions.

              • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
              • [claimed-docs] You’ll be notified when the task completes or needs your input.
              • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
              • [claimed-docs] To send a message to the agent:
              • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
              • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
            2. developerSend follow-up instructions to an active agent session to steer its work without restarting

              weight 2 · round to Jules
              Factorypartialclaimed6/10

              Factory's API reference explicitly supports creating and driving Droid sessions including sending messages within an active session (factory-docs-13), and the CLI is interactive by nature, implying follow-up prompts are possible without restarting. However, there is no explicit documentation describing mid-task interruption/steering while the agent is actively executing a long-running task (e.g., Missions or droid exec), so the steering-while-running behavior is not directly evidenced. Missing for 10: explicit docs on interrupting/redirecting an in-progress autonomous run, and independent/hands-on confirmation that follow-up messages actually steer ongoing work rather than queue for the next turn.

              • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
              • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
              • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.

              Jules's API docs explicitly describe sending a message to an active agent session (jules-docs-12), and community reports confirm the workflow of reviewing partial work and providing feedback for Jules to continue rather than restarting (jules-comm-3). This shows the steering capability works in practice, though documentation on how follow-ups affect an in-progress plan/execution is thin. Missing for 10: detailed first-party documentation of mid-task message handling/UI chat thread, and independent hands-on confirmation of steering effectiveness beyond one anecdote.

              • [claimed-docs] To send a message to the agent:
              • [claimed-docs] You’ll be notified when the task completes or needs your input.
              • [community] Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.

            Sandbox execution

            1. developerHave an agent safely execute code and install dependencies inside an isolated sandbox

              weight 3 · round to Jules
              Factorynone0/10

              Evidence describes tiered autonomy, bash mode, and CI/CD execution (droid exec) but never mentions an isolated sandbox environment for code execution or dependency installation; no container/VM isolation is documented.

                Jules explicitly runs in a cloud VM that clones code and installs dependencies, isolated from the user's local machine, with a plan-review step before code changes are made — directly matching the sandboxed execution story. Community reports corroborate it actually running tasks end-to-end (producing PRs) in this VM environment, though some criticize reliability/loops on complex codebases. Missing for 10: detailed docs on sandbox security boundaries (network isolation, resource limits) and independent security audit of the VM isolation.

                • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                • [community] Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…
                • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…

              Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight

              Keeping a human in the loop — approvals, checkpoints, interrupts

              Approval controls

              1. developerConfigure an agent to auto-approve all its actions instead of confirming each one

                weight 2 · round to Factory
                Factorypartialclaimed6/10

                Factory's docs confirm 'tiered autonomy' in droid exec that controls what operations run without manual confirmation, implying a configurable auto-approve mode for CI/batch use, but there's no explicit documentation of a full 'auto-approve all actions' toggle or its exact configuration options/flags. missing for 10: explicit config syntax/flag for full auto-approval, independent confirmation of behavior, coverage of auto-approve in interactive (non-exec) sessions.

                • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.

                Jules's default workflow requires manual approval at multiple stages (plan approval, diff review, PR approval), but the API docs mention an optional `automationMode` field that changes default behavior around automatic PR creation, suggesting some configurable auto-approve path exists via API rather than the standard UI flow. Missing for 10: explicit documentation of a setting that suppresses plan-approval and diff-review confirmations entirely, and any hands-on/community confirmation that this automation mode actually skips human checkpoints.

                • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
              2. product-managerApprove key agent decisions from my phone while agents continue working

                weight 1 · round to Jules
                Factorypartialclaimed3/10

                Docs mention delegating tasks, reviewing diffs, and merging 'from the App or your terminal' and tiered autonomy that gates operations needing confirmation, implying some human-in-the-loop review outside the terminal, but there is no explicit evidence of a mobile/phone app or of approving in-flight agent decisions remotely while agents keep working. Missing for 10: explicit mobile app/phone interface, evidence of asynchronous approval while agent continues running, and independent confirmation of this workflow.

                • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.

                Jules is a cloud-based agent with a web app and notification system, plan approval, diff review, and PR approval steps, and community evidence confirms users review/approve PRs from a browser including 'coding on my phone' via mobile web. However, there is no dedicated native mobile app or explicit mobile-optimized approval UI, and no evidence of push-notification-driven approval flows tailored to on-the-go PM decision-making. missing for 10: dedicated mobile app or mobile-specific UI, evidence of push notifications enabling quick phone-based approvals, PM-specific (non-developer) approval workflow.

                • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                • [claimed-docs] You’ll be notified when the task completes or needs your input.
                • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                • [community] I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…
              3. engineering-leadSet tiered autonomy levels controlling what an agent can do without manual confirmation

                weight 3 · round to Factory
                Factorypartialclaimed6/10

                Factory explicitly documents 'tiered autonomy' in droid exec to control which operations run without manual confirmation, directly matching the story, but this is scoped to the CI/CD-oriented droid exec mode rather than a broader, configurable set of autonomy tiers across all agent surfaces. missing for 10: detail on specific tier levels/permissions, configuration UI or granular controls, and evidence this applies uniformly across App/CLI sessions, not just droid exec.

                • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.

                Jules has a binary plan-approval workflow (review/approve plan, review diffs, approve PRs) and an API `automationMode` field that toggles whether PRs are auto-created, but there's no evidence of configurable tiered autonomy levels (e.g., low/medium/high trust settings, granular permission scopes, or per-action confirmation thresholds) that an engineering-lead could set. Missing for 10: explicit multi-tier autonomy/permission settings, admin-configurable trust levels, and any org-wide policy controls beyond the single automationMode toggle and default plan-approval gate.

                • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.

              Model control

              1. ai-native userHave each task prompt automatically routed to the most suitable underlying model

                weight 2 · round drawn
                Factorynone0/10

                No evidence describes automatic routing of prompts to the most suitable underlying model; docs mention model-agnostic droid workflows, custom droids, and orchestration but nothing about auto-selecting models per task.

                  Julesnone0/10

                  No evidence Jules performs automatic model routing per task; docs mention only a single flagship model (Gemini 3 Pro) as 'priority access,' with no mention of routing prompts to different models based on suitability. Missing for 10: any documentation of multi-model routing logic, model-selection criteria, or automatic switching between models.

                  • [claimed-docs] Priority access to the latest models, starting with Gemini 3 Pro
                • engineering-leadSwitch away from automatic model selection to a specific model of my choice

                  weight 1 · round drawn
                  Factorynone0/10

                  No evidence in the pack mentions model selection, automatic model routing, or the ability to choose a specific model over an automatic default; the docs cover CLI usage, integrations, missions, and readiness reports but nothing about model choice controls.

                    Julesnone0/10

                    No evidence pack item mentions model selection settings or the ability to choose a specific underlying model; the only related mention is 'Priority access to the latest models, starting with Gemini 3 Pro' which describes access tiers, not user-controlled model switching.

                    Visibility monitoring

                    1. developerWatch what a running agent is doing in real time, including its current status

                      weight 3 · round drawn
                      Factorypartialclaimed5/10

                      Factory's session API lets you create and manage Droid sessions—including lifecycle, settings, and messages—implying some ability to track a running agent's state, and the App/CLI let you review diffs as work progresses. However, there is no explicit documentation of a live status dashboard, streaming logs, or real-time progress view of an in-flight agent. Missing for 10: dedicated real-time monitoring UI/stream, explicit 'live status' feature documentation, independent confirmation of live tracking.

                      • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
                      • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                      • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.

                      Jules provides notifications on task completion/input-needed and a plan-approval step, implying some status visibility, plus a diff/PR review flow, but there is no evidence of a live real-time activity feed or streaming log of the agent's current actions while running; community feedback even notes the absence of a stop/interrupt control mid-run, suggesting limited real-time observability. Missing for 10: evidence of a real-time execution log/console view, granular step-by-step status updates, and independent confirmation of live monitoring UX.

                      • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                      • [claimed-docs] You’ll be notified when the task completes or needs your input.
                      • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                      • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
                    2. developerGet notified when an agent completes a task or needs my input

                      weight 2 · round to Jules
                      Factorypartialclaimed4/10

                      Factory integrates with Slack and PagerDuty and supports tiered autonomy that pauses for manual confirmation, implying some mechanism for alerting developers when input is needed, but there's no explicit documentation of a notification feature for task completion or input requests. missing for 10: explicit notification/alert documentation, evidence of completion pings, confirmation of Slack/PagerDuty being used specifically for task-status alerts.

                      • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems
                      • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
                      • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.

                      First-party docs explicitly state users are notified when a task completes or needs input, and the workflow (plan approval, diff review, PR creation) reinforces oversight checkpoints where notifications matter. Community evidence corroborates the async, check-back-later usage pattern (e.g. reviewing PRs after time away), consistent with notification-driven workflows. Missing for 10: independent verification of notification channels (email/push/Slack), and no detail on notification reliability or configurability.

                      • [claimed-docs] You’ll be notified when the task completes or needs your input.
                      • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                      • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                      • [community] I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…
                      • [community] Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.

                    Intent to spec — stories about intent to spec in this arenaIntent to spec

                    Stories about intent to spec in this arena

                    Natural language task intake

                    1. developerDescribe a feature or bug in plain language and have it automatically turned into a scoped implementation task

                      weight 3 · round drawn
                      Factorypartialclaimed6/10

                      Factory supports delegating tasks described in plain language (via droid CLI, App, or Missions) which get turned into scoped, executable work with diffs to review and merge, and readiness tooling helps scope repo work automatically. However, there's no explicit documented workflow for turning a raw bug/feature description into a formal 'scoped implementation task' spec artifact (e.g., structured ticket/spec generation before execution) distinct from just running the agent. missing for 10: explicit intent-to-spec artifact generation/preview step, independent/hands-on evidence of accurate scoping from vague input, and detail on how ambiguous requests are clarified before execution.

                      • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                      • [claimed-docs] delegate scoped tasks to Custom Droids via /droids, and package reusable procedures as Skills with /skills or /create-skill
                      • [claimed-docs] A designer shares a mockup, and the system implements it without handoffs.
                      • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
                      • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.

                      Jules docs clearly show a plain-language task submission flow that generates a reviewable/approvable plan before code changes, then produces a diff and PR (jules-docs-4, jules-docs-7, jules-docs-8), directly matching the intent-to-spec story. However community reports show inconsistent scoping quality — confusion in monorepos, endless loops, and needing heavy babysitting — indicating the automatic scoping doesn't always hold up in practice. Missing for 10: independent verification of plan/spec quality on complex codebases, and consistent evidence the generated plan reliably matches developer intent without back-and-forth correction.

                      • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                      • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                      • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                      • [community] I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…
                      • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
                      • [community] Do people really find Jules useful? I find it needs babysitting much more than Cursor.
                    2. product-managerConvert user feedback submissions into structured tasks with proposed scope

                      weight 2 · round drawn
                      Factorynone0/10

                      Factory documents task delegation, Missions planning, and integrations with Jira/Linear/Notion, but nothing in the evidence describes ingesting user feedback submissions and outputting a structured task with proposed scope — this specific PM workflow is not documented.

                        Julesnone0/10

                        Jules's documented workflow is a coding-agent pipeline (clone repo, generate a plan, produce diffs/PRs) driven by developer-specified tasks or GitHub issues, not a mechanism for ingesting unstructured user-feedback and converting it into a structured task with proposed scope for a PM. No evidence shows feedback intake, requirement structuring, or scope proposal features aimed at product managers.

                        • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                        • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                        • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                      • developerAttach a marked-up screenshot or mockup to a task so the agent implements the correct visual change

                        weight 2 · round drawn
                        Factorynone0/10

                        Evidence only mentions a designer sharing a mockup conceptually (factory-docs-8), with no documented capability to attach an image/screenshot with markup annotations to a task or session, nor any multimodal image-input feature described.

                        • [claimed-docs] A designer shares a mockup, and the system implements it without handoffs.
                        Julesnone0/10

                        No evidence Jules supports attaching images, screenshots, or mockups to a task — task submission is described only via text prompts, GitHub issue labels, or API messages. Missing for 10: any mention of image/screenshot attachment, multimodal input handling, or visual mockup interpretation.

                        Plan approval

                        1. developerReview and approve an agent's implementation plan before any code changes are made

                          weight 3 · round to Jules
                          Factorypartialclaimed4/10

                          Factory's tiered autonomy in droid exec implies some operations require manual confirmation before running, and Missions offer 'plan and execute' orchestration, suggesting a planning phase exists, but there is no explicit documentation of a discrete implementation-plan artifact that a developer reviews and approves before any code changes are made — the described workflow (docs-1) instead centers on reviewing the diff/output after changes. missing for 10: explicit plan-approval UI/step description, evidence of a pre-execution plan artifact, confirmation that no code is touched until plan is approved.

                          • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
                          • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.
                          • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.

                          Jules docs explicitly state the agent generates a plan upon task submission that users can review and approve before any code changes are made, directly matching the story. No independent/hands-on evidence specifically confirms or contradicts this plan-approval step, though other workflow claims (diffs, PRs) are corroborated by community mentions. Missing for 10: independent/hands-on confirmation of the plan-review step specifically, and more detail on what the plan interface looks like.

                          • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                          • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                          • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                        2. engineering-leadApprove a task's scope and contract before an agent is allowed to modify the repository

                          weight 2 · round to Jules
                          Factorypartialclaimed4/10

                          Factory documents tiered autonomy that limits what operations run without manual confirmation and diff review before merge, which implies some human-gate mechanics, but there's no explicit evidence of an engineering-lead approving a task's defined scope/contract *before* the agent is permitted to touch the repository — reviews are framed as post-hoc diff review rather than pre-execution scope sign-off. Missing for 10: explicit scope/contract definition step, an approval gate that blocks agent execution until lead sign-off, and evidence this is lead-specific rather than generic autonomy tiering.

                          • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                          • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.

                          Jules generates a plan that the user can review and approve before any code changes are made, giving an engineering-lead a gate before modification, and it also produces diffs/PRs requiring approval before merge. However, this is a general 'plan approval' UX rather than a formal scope/contract sign-off workflow tailored for engineering-lead governance (e.g., no evidence of role-based approval gates, policy enforcement, or blocking unauthorized starts). missing for 10: explicit lead/role-based approval gating before task execution starts, evidence of enforceable contract/scope definitions beyond a plan preview, independent confirmation that the plan-approval step reliably blocks modification.

                          • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                          • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                          • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.

                        Ticket driven tasking

                        1. developerAssign a coding task to an agent directly from an existing issue or ticket

                          weight 3 · round to Jules
                          Factorypartialclaimed4/10

                          Factory documents integrations with Jira and Linear ('Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems'), implying tickets can feed into agent workflows, but there is no explicit documentation of assigning/triggering a droid task directly from an issue or ticket UI. missing for 10: explicit workflow/screenshot showing task creation from a Jira/Linear ticket, confirmation that ticket assignment auto-triggers an agent session, independent corroboration.

                          • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems

                          Jules explicitly supports assigning tasks directly from GitHub issues via a 'jules' label, and API/community evidence corroborates task dispatch and PR creation workflows. Missing for 10: evidence of ticket-tracker integrations beyond GitHub (e.g., Jira/Linear) and independent hands-on confirmation of the label-to-task flow itself.

                          • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                          • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                          • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                          • [community] its way better than the github thing in my experience it produces usable PRs

                        Openness — open source, data portability, and self-hosting storiesOpenness

                        Open source, data portability, and self-hosting stories

                        1. ai-native userDo everything through the API that I can do in the UI

                          weight 2 · round drawn
                          Factorypartialprobed5/10

                          Factory documents a real API (session creation/lifecycle) alongside its CLI and UI, showing some functional overlap, but there's no OpenAPI spec (404s on all candidate URLs) and no evidence that UI-exclusive features like Missions, Software Factory's coverage map, or readiness-report/fix slash commands are exposed via the API. missing for 10: documented API parity for Missions, Software Factory, and readiness commands; a public OpenAPI/schema confirming full surface coverage.

                          • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
                          • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.
                          • [claimed-docs] Open Software Factory to see your delivery lifecycle as an automation coverage map.
                          • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
                          • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI
                          • [probe] PROBE openapi: all candidate paths 404 (https://docs.factory.ai/openapi.json, https://docs.factory.ai/swagger.json, https://docs.factory.ai/…

                          Jules exposes a public API for creating tasks, sending messages, and enabling automation (PR creation) — core building blocks of the UI workflow — and a community user confirms dispatching tasks via the API from an external MCP client, showing real usage beyond the UI. However, there's no documented API coverage for plan approval, diff browsing/approval, or settings/API-key management, and no OpenAPI spec is discoverable (probes 404), so full UI/API parity is unconfirmed. Missing for 10: API endpoints for plan review/approval, diff/PR review parity, settings management, and a public API spec confirming full feature coverage.

                          • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                          • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
                          • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
                          • [claimed-docs] To send a message to the agent:
                          • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
                          • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
                        2. ai-native userExport all of my data in open formats and leave

                          weight 3 · round drawn
                          Factorynone0/10

                          No evidence addresses data export or portability in open formats, nor any account-deletion/data-takeout mechanism; the docs focus on session management, CLI, and integrations, not exporting user data to leave the platform.

                            Julesnone0/10

                            No evidence of any data export feature, open-format export, or account/data portability mechanism for Jules; the product works via GitHub repos/PRs but nothing indicates users can export Jules-specific data (task history, configs, etc.) and leave. missing for 10: any documentation of data export, open format support, or account portability/deletion workflow.

                            • ai-native userRead the product's source under an open license

                              weight 2 · round drawn
                              Factorynone0/10

                              No evidence of an open-source license or publicly available source code for Factory/Droid; all evidence points to closed docs and a proprietary CLI/platform. missing for 10: any license file, GitHub repo, or open-source statement covering the product's source code.

                                Julesnone0/10

                                Jules is a closed, proprietary Google product (cloud VM, API keys, usage limits); no evidence anywhere of source code being published or licensed openly, and probes for docs/openapi artifacts return 404s, further suggesting no open publishing.

                                • [claimed-docs] Jules needs access to your repositories in order to work.
                                • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                • [probe] PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt
                                • [probe] PROBE docs-md: HTTP 404 at https://jules.google/docs.md
                                • [probe] PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…
                              • ai-native userSelf-host the core product

                                weight 3 · round drawn
                                Factorynone0/10

                                Factory is presented as a cloud-hosted platform (Factory App, Droid CLI connecting to hosted services, API sessions) with no evidence of a self-hostable core server or on-prem deployment option anywhere in the docs or probes.

                                  Julesnone0/10

                                  Jules is explicitly a cloud-based VM service (Google-hosted) with no evidence of any self-hosting option; community even notes cloud-only design as a drawback compared to local tools like Claude Code.

                                  • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                  • [community] Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…
                                  • [community] It's a shame Google picked the wrong system design for Jules. Claude Code's system design is clearly superior at this point.

                                Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits

                                Free-tier ceilings, usage caps, and rate limits before you have to pay

                                Enterprise licensing

                                1. engineering-leadLicense an enterprise deployment with SSO and commercial support for organization-wide rollout

                                  weight 2 · round drawn
                                  Factorynone0/10

                                  No evidence pack items mention enterprise licensing, SSO, or commercial support offerings; documentation only covers product features like Droid CLI, MCP integration, and readiness reports. Missing for 10: any mention of SSO/SAML, enterprise tier, commercial support SLAs, or org-wide licensing terms.

                                    Julesnone0/10

                                    No evidence of enterprise licensing, SSO, or commercial support offerings; docs mention only individual API keys and usage-tier limits, and community discussion references a free beta plan, not enterprise deployment. missing for 10: SSO integration, enterprise licensing/contract terms, commercial support SLAs, org-wide admin/rollout tooling.

                                    • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
                                    • [claimed-docs] Power users & agent-heavy workflows ... 300 ... 60
                                    • [community] Is Jules free of charge? Yes, for now, Jules is free of charge. Jules is in beta and available without payment while we learn from usage.

                                  Model flexibility

                                  1. engineering-leadBring my own LLM or API key so agents run on the model of my choice

                                    weight 2 · round drawn
                                    Factorynone0/10

                                    No evidence pack item mentions bringing your own LLM, custom API keys, or model selection/configuration options; all docs focus on CLI, integrations, and workflow features.

                                      Julesnone0/10

                                      No evidence that Jules supports bringing a custom LLM or third-party API key; it is tied to Google's own models (Gemini), with documentation only mentioning Jules's own API key for accessing Jules itself, not for configuring underlying model providers. missing for 10: any mention of BYO-LLM support, model selection options, or third-party API key configuration.

                                      • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                      • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
                                      • [claimed-docs] Priority access to the latest models, starting with Gemini 3 Pro

                                    Usage quotas

                                    1. engineering-leadSee and manage plan-based daily task and concurrency limits for agent workflows

                                      weight 2 · round to Jules
                                      Factorynone0/10

                                      No evidence of plan-based daily task limits, concurrency caps, or admin controls for managing such limits; docs cover CLI, sessions, missions, and integrations but nothing on usage/plan limit visibility or management.

                                        Jules docs reference a usage-limits page with plan-tiered daily task numbers (e.g., 300/60) and community confirms daily task limits exist and change between releases (60→15 on free plan), showing plan-based limits are real and documented. However there's no evidence of an engineering-lead-facing dashboard or admin console to monitor team usage, set/adjust concurrency, or manage limits across a team — only a static limits reference page and anecdotal user experience of hitting caps. Missing for 10: team/org usage dashboard, per-user concurrency visibility, ability to configure or request limit changes, and any admin/lead-specific management UI.

                                        • [claimed-docs] Power users & agent-heavy workflows ... 300 ... 60
                                        • [community] The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…

                                      Privacy posture — data-handling and privacy storiesPrivacy posture

                                      Data-handling and privacy stories

                                      1. ai-native userChoose where my data is stored (region/residency)

                                        weight 2 · round drawn
                                        Factorynone0/10

                                        No evidence in the pack addresses data residency, regional storage options, or compliance/data-location controls; all evidence covers agent workflows, CLI, and integrations unrelated to data storage location.

                                          Julesnone0/10

                                          No evidence in the pack mentions data residency, regional storage options, or any control over where Jules stores/processes data; docs only describe VM execution and GitHub integration without residency settings.

                                          • ai-native userPrevent my data from being used to train AI models

                                            weight 3 · round drawn
                                            Factorynone0/10

                                            No evidence in the pack addresses data usage/training opt-out, privacy policy, or data retention controls; the docs focus entirely on product features like CLI, missions, and integrations.

                                              Julesnone0/10

                                              No evidence in the pack addresses data usage/training opt-out policies, privacy controls, or any statement about whether user code/data is used to train AI models; the docs focus on functionality (VM execution, PR creation, API) rather than data governance.

                                              • ai-native userControl data retention and deletion

                                                weight 2 · round drawn
                                                Factorynone0/10

                                                No evidence pack items address data retention policies, deletion controls, or privacy settings for user data; all citations relate to product features like CLI, MCP integrations, and agent workflows rather than privacy posture.

                                                  Julesnone0/10

                                                  No evidence pack item mentions data retention policies, data deletion controls, or privacy settings for repositories/code processed by Jules; docs cover workflow, VM execution, and API usage but not retention/deletion controls.

                                                  • ai-native userOpt out of telemetry and usage tracking

                                                    weight 2 · round drawn
                                                    Factorynone0/10

                                                    No evidence in the pack mentions telemetry, usage tracking, data collection, or any opt-out/privacy settings for Factory or the Droid CLI.

                                                      Julesnone0/10

                                                      No evidence in the pack mentions telemetry, usage tracking, or an opt-out mechanism; the docs focus on repo access, VM execution, PR workflow, and API usage without addressing privacy/telemetry controls.

                                                      Repo integration — stories about repo integration in this arenaRepo integration

                                                      Stories about repo integration in this arena

                                                      Chat integration

                                                      1. developerTag an agent in a chat thread to discuss and delegate a bug or task

                                                        weight 2 · round to Jules
                                                        Factorypartialclaimed3/10

                                                        Factory's docs mention connecting Slack as an integration to 'keep development synchronized with team systems,' implying some chat-based interaction, but there is no explicit evidence of an @-mention/tagging mechanism in a chat thread to discuss or delegate a specific bug/task to the agent. Missing for 10: explicit Slack @droid tagging workflow, thread-based task delegation UI, and confirmation that discussion happens inline in chat rather than just triggering external actions.

                                                        • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems

                                                        Jules supports delegating tasks via a GitHub issue label ('jules') and via API messages to the agent, which is a form of task delegation but not the classic 'tag an agent in a chat thread to discuss' interaction pattern; one community user built a custom MCP server to dispatch Jules tasks from VS Code Copilot Chat, showing it's possible but not a native feature. missing for 10: native chat-thread tagging/mention UI, evidence of back-and-forth discussion in a thread before delegation, first-party support for Slack/Teams-style @mentions.

                                                        • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                                                        • [claimed-docs] To send a message to the agent:
                                                        • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.

                                                      Knowledge context

                                                      1. developerAdd a context file describing my codebase conventions so agents generate more relevant plans and code

                                                        weight 3 · round to Jules
                                                        Factorynone0/10

                                                        The evidence pack does not mention any context file mechanism (e.g., AGENTS.md, .factory config, or similar) for describing codebase conventions to guide agent behavior; it covers CLI usage, integrations, readiness reports, and missions but nothing about persistent repo-convention context files.

                                                          Jules automatically reads an AGENTS.md file in the repo root to guide its plans and code generation, directly matching the story of adding a context file describing codebase conventions. Missing for 10: no independent/hands-on confirmation of how well conventions in AGENTS.md actually improve plan/code relevance, and no detail on supported format/scope beyond the single doc mention.

                                                          • [claimed-docs] Jules now automatically looks for a file named AGENTS.md in the root of your repository.
                                                          • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                                                        • developerQuery generated documentation for any public or private repository

                                                          weight 1 · round drawn
                                                          Factorynone0/10

                                                          The evidence pack covers Factory's CLI, agent-readiness reports, missions, and integrations, but nothing describes generating or querying documentation for a repository's codebase (e.g., an auto-generated repo wiki/docs feature). This is a plausible capability for a repo-integrated dev tool, so the axis applies, but no supporting evidence exists.

                                                            Julesnone0/10

                                                            Evidence describes Jules as a task-based coding agent that clones repos, edits code, and creates PRs, but there is no mention of generating or letting users query documentation for a repository (public or private). missing for 10: any docs-generation feature, a documentation query/search interface, evidence of indexing repo docs for Q&A.

                                                            • [claimed-docs] Jules needs access to your repositories in order to work.
                                                            • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                                            • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.

                                                          Project management integration

                                                          1. product-managerConnect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly

                                                            weight 3 · round to Factory
                                                            Factorypartialclaimed5/10

                                                            Docs confirm Jira and Linear integration explicitly (factory-docs-4), but ClickUp and Monday.com are not mentioned anywhere in the evidence, and there's no detail on ticket management workflows (creating/updating tickets) beyond 'connecting' systems to sync development. Missing for 10: ClickUp integration, Monday.com integration, concrete ticket-management/CRUD workflows via these connectors, independent corroboration.

                                                            • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems
                                                            Julesnone0/10

                                                            Evidence only shows Jules integrating with GitHub (repos, issues, PRs) and provides an API for custom workflows, but there is no mention of Jira, Linear, ClickUp, Monday.com, or any issue-tracker integration beyond GitHub.

                                                            Version control integration

                                                            1. developerConnect a GitHub repository so an agent can access the code and open pull requests against it

                                                              weight 3 · round to Jules
                                                              Factorypartialclaimed5/10

                                                              Docs indicate Factory works within Git workflows (delegate tasks, review diffs, merge from App/terminal) and can run in CI/CD via droid exec, implying repo access and PR-opening capability, but there is no explicit documentation describing connecting/authorizing a GitHub repository or an explicit PR-creation feature. Missing for 10: explicit GitHub repo connection/auth flow documentation, explicit 'open pull request' feature description, and independent/hands-on confirmation of PR creation.

                                                              • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                                                              • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
                                                              • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…

                                                              Jules is built around GitHub repo access: docs describe cloning repos, GitHub issue label assignment, generating diffs/plans, and creating PRs that can be merged on GitHub, corroborated by community reports of receiving usable PRs from Jules-driven tasks. Missing for 10: independent verification of the full connect-repo setup flow and edge-case reliability across repo types.

                                                              • [claimed-docs] Jules needs access to your repositories in order to work.
                                                              • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                                                              • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                                              • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                                                              • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                              • [community] its way better than the github thing in my experience it produces usable PRs
                                                              • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
                                                            2. developerGrant an agent access to my repositories with a one-click install, without complex setup

                                                              weight 2 · round to Jules
                                                              Factorynone0/10

                                                              Evidence shows Factory integrates with Git workflows, CLI, and external tools like Jira/Slack/MCP, but there is no mention of a one-click GitHub/repo install or simplified OAuth-based repo authorization flow. Setup appears to require CLI installation and configuration rather than a one-click grant.

                                                                Jules integrates with GitHub (label-based task assignment, PR creation) and needs repo access to work, implying a connect/authorize flow, and community evidence confirms GitHub integration works well for many users. However, there's no explicit documentation of a literal 'one-click install' onboarding flow or app installation process described in the evidence pack. missing for 10: explicit one-click GitHub App install/authorization flow documentation, independent confirmation of setup simplicity.

                                                                • [claimed-docs] Jules needs access to your repositories in order to work.
                                                                • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                                                                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                • [community] I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.
                                                                • [community] its way better than the github thing in my experience it produces usable PRs

                                                              Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates

                                                              Quality gates on changes — review flow, required checks, merge protection

                                                              Ci remediation

                                                              1. engineering-leadHave failed CI workflows automatically diagnosed and fixed with a proposed pull request

                                                                weight 3 · round to Factory
                                                                Factorypartialclaimed4/10

                                                                Factory's droid exec is explicitly designed to run in CI/CD pipelines as a one-shot task with tiered autonomy, and Droid CLI integrates with Git workflows, suggesting the building blocks exist for automating CI fixes. However, there is no direct evidence of a dedicated feature that detects a failed CI workflow, diagnoses the failure, and automatically opens a proposed pull request end-to-end. missing for 10: explicit CI-failure detection/trigger integration, automatic diagnosis-to-PR workflow documentation, and any hands-on/independent proof of this specific use case.

                                                                • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                                                                • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
                                                                • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.

                                                                Jules offers building blocks—an API for custom automation that can 'automate tasks like bug fixing' and auto-create PRs via automationMode—but there is no direct evidence of a native integration that detects a failed CI workflow and automatically diagnoses/fixes it with a PR; this would require custom API wiring, not an out-of-box feature. Missing for 10: documented CI-failure trigger/integration, evidence of automatic diagnosis of CI logs, and case studies of this specific end-to-end flow.

                                                                • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
                                                                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                                                              2. developerTrigger an agent from CI/CD pipelines to fix a broken build or failing test

                                                                weight 2 · round to Factory
                                                                Factoryfullprobed8/10

                                                                Factory explicitly documents `droid exec` as a one-shot CLI command designed for CI/CD pipelines, shell scripts, and batch processing, with tiered autonomy controls for unattended operation — directly enabling triggering an agent from CI to fix builds/tests. Missing for 10: a concrete worked example of a CI pipeline invoking droid exec on a failing test/build, and independent/hands-on corroboration beyond vendor docs.

                                                                • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                                                                • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
                                                                • [probe] official CLI documented at https://docs.factory.ai/droid-cli/quickstart

                                                                Jules exposes a documented API for creating custom workflows and automating tasks like bug fixing (jules-docs-9, jules-docs-10, jules-docs-12), which could be wired into a CI/CD pipeline to trigger a fix, and community evidence shows users building custom integrations to dispatch Jules tasks programmatically (jules-comm-19). However, there is no first-party CI/CD-specific integration (e.g. GitHub Actions step, build-failure webhook) or documented example of triggering Jules to fix a broken build/failing test directly from a pipeline. missing for 10: explicit CI/CD pipeline integration/example, evidence of triggering on build/test failure events, hands-on confirmation of this specific workflow succeeding.

                                                                • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
                                                                • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
                                                                • [claimed-docs] To send a message to the agent:
                                                                • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.

                                                              Diff review

                                                              1. developerConfigure an agent to automatically open a pull request when its task completes

                                                                weight 2 · round to Jules
                                                                Factorypartialclaimed3/10

                                                                Docs show git workflow integration (droid-cli git workflow, delegate task then review diff and merge, droid exec for CI/CD pipelines) which implies PR-related automation, but there is no explicit documentation of a feature to automatically open a pull request when a task completes. Missing for 10: explicit config/flag for auto-PR creation, first-party example of a droid opening a PR on completion, independent/hands-on confirmation.

                                                                • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                                                                • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                                                                • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.

                                                                Jules natively creates a PR when a task completes (jules-docs-8), and the API exposes an `automationMode` field to configure automatic PR creation, defaulting to off unless explicitly configured (jules-docs-11), directly matching the story of configuring automatic PR-on-completion. Community reports corroborate that Jules produces usable PRs and users regularly review PRs it opens (jules-comm-6, jules-comm-12, jules-comm-16). missing for 10: independent hands-on verification of the automationMode toggle specifically, and more detail on configuration options/edge cases (e.g., partial completions, failed tasks) affecting whether a PR is always opened.

                                                                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
                                                                • [community] its way better than the github thing in my experience it produces usable PRs
                                                                • [community] I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…
                                                                • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
                                                              2. developerReview a diff of an agent's changes and approve it before it becomes a pull request

                                                                weight 3 · round to Jules
                                                                Factorypartialclaimed6/10

                                                                Docs explicitly describe the delegate-review-merge workflow ('Delegate a task, review the diff, and merge from the App or your terminal') and tiered autonomy controls that gate what runs without confirmation, supporting a review-before-PR gate. However, there is no detailed documentation of the diff review UI itself, approval mechanics, or independent/hands-on confirmation of this exact flow. Missing for 10: dedicated diff-review UI documentation, explicit approval-step mechanics, and independent verification of the review-before-PR gate.

                                                                • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                                                                • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.

                                                                Jules explicitly provides a diff of changes for review and approval before it creates/publishes a PR, with plan approval and diff approval steps documented; community mentions confirm PRs are generated for review. Missing for 10: independent hands-on confirmation of the diff-approval UI flow itself (comments focus on PR quality/output rather than the diff-review step specifically).

                                                                • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                                                                • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                                                                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                • [community] I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…

                                                              Pr review automation

                                                              1. ai-native userHave incoming issues automatically triaged with severity suggested and routed to the right owner

                                                                weight 2 · round drawn
                                                                Factorynone0/10

                                                                Factory documents integrations with issue trackers like Jira, Linear, and PagerDuty (factory-docs-4) and generic custom droid/skill automation (factory-docs-5), but there is no evidence of an automated triage workflow that suggests severity or routes issues to an owner. missing for 10: any documented triage feature, severity classification logic, or owner-routing capability.

                                                                • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems
                                                                • [claimed-docs] delegate scoped tasks to Custom Droids via /droids, and package reusable procedures as Skills with /skills or /create-skill
                                                                Julesnone0/10

                                                                Jules is a coding agent that executes tasks assigned via label or API, but there is no evidence of automatic issue triage, severity classification, or routing to owners — the closest feature is manual label-based task assignment, not automated triage. Missing for 10: any mention of severity scoring, triage logic, or owner-routing automation.

                                                                • [claimed-docs] Use the "jules" label in an issue to assign a task directly in GitHub.
                                                                • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                              2. engineering-leadHave every pull request automatically reviewed with AI-generated inline comments

                                                                weight 3 · round drawn
                                                                Factorynone0/10

                                                                Evidence shows Factory's droid CLI/exec can review diffs, be triggered in CI/CD, and connect to Git workflows, but nothing in the pack describes automatic PR review with AI-generated inline comments posted to pull requests. Missing for 10: any documentation of automated PR-triggered review, inline comment generation on PRs, or GitHub/GitLab PR integration specifics.

                                                                • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                                                                • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                                                                • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.
                                                                Julesnone0/10

                                                                Jules is documented as a task-execution agent that clones repos, makes code changes, and creates its own PRs for approval — not as a bot that automatically reviews incoming pull requests with inline comments. The only tangential mention is jules-docs-9's passing reference to 'automate tasks like ... code reviews' via API, but there is no documentation of an automatic inline-comment review gate applied to every PR.

                                                                • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.

                                                              Readiness checks

                                                              1. engineering-leadAutomatically fix failing agent-readiness criteria in my repository

                                                                weight 1 · round to Factory
                                                                Factoryfullclaimed8/10

                                                                Factory explicitly documents a /readiness-fix slash command that automatically fixes failing agent-readiness criteria, usable from the Factory App or Droid CLI, complementing the /readiness-report diagnostic command. This directly matches the story's ask, though only first-party docs support it with no independent/hands-on corroboration. Missing for 10: independent or hands-on verification that /readiness-fix reliably resolves criteria, and detail on which criteria types it can/cannot auto-fix.

                                                                • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
                                                                • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI
                                                                Julesnone0/10

                                                                Jules can be assigned generic bug-fix/coding tasks and reads an AGENTS.md file if present, but there is no evidence of it detecting or automatically remediating 'agent-readiness' quality-gate criteria (e.g., missing/invalid AGENTS.md, agent-compatibility checks) as a review gate — it only consumes such files, it doesn't audit or fix them as a compliance gate.

                                                                • [claimed-docs] Jules now automatically looks for a file named AGENTS.md in the root of your repository.
                                                                • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                                                • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                                                                • [community] I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…
                                                              2. engineering-leadRun a readiness report that evaluates how ready my repository is for autonomous agents

                                                                weight 2 · round to Factory
                                                                Factoryfullclaimed8/10

                                                                Factory explicitly documents a `/readiness-report` slash command (in App or Droid CLI) that evaluates a repository's agent-readiness level, plus a companion `/readiness-fix` command to auto-remediate failing criteria, directly matching the engineering-lead's story. Missing for 10: independent/hands-on validation of report contents and no sample report output shown.

                                                                • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
                                                                • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI
                                                                • [claimed-docs] A designer shares a mockup, and the system implements it without handoffs.
                                                                Julesnone0/10

                                                                No evidence of a readiness report, scorecard, or repository-assessment feature for autonomous agent suitability; evidence only covers task execution, PR creation, and API usage. Missing for 10: any readiness-scoring feature, repo audit/checklist, or report output evaluating agent-readiness.

                                                                Security remediation

                                                                1. engineering-leadHave security alerts automatically validated and remediated with an opened pull request

                                                                  weight 2 · round drawn
                                                                  Factorynone0/10

                                                                  Evidence covers task delegation, CLI automation, readiness reports/fixes, and CI/CD execution, but nothing describes ingesting security alerts (e.g., from a scanner like Dependabot/Snyk) and automatically validating and remediating them via an opened PR. No security-alert-specific workflow or integration is documented.

                                                                    Julesnone0/10

                                                                    Jules is a general coding agent that can be assigned tasks (via GitHub label or API) and opens PRs with diffs for review, but there is no evidence of any security-alert scanning, vulnerability detection/validation, or a workflow that automatically triggers remediation from a security alert (e.g., Dependabot/CVE integration). The evidence pack shows generic task-to-PR flow, not a security-alert-specific pipeline.

                                                                    • [claimed-docs] Jules needs access to your repositories in order to work.
                                                                    • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                                                    • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                                                                    • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                                                                    • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…

                                                                  Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism

                                                                  Running many jobs at once — concurrency, fleets, queueing

                                                                  Concurrent execution

                                                                  1. engineering-leadRun many agent tasks concurrently to scale delivery throughput

                                                                    weight 3 · round to Factory
                                                                    Factorypartialclaimed7/10

                                                                    Factory provides several building blocks for scaling agent work: droid exec is designed for one-shot, CI/CD, and batch processing use cases, the Sessions API lets you create and manage multiple Droid sessions programmatically, and Missions support orchestrating large multi-feature projects with structured coordination. Together these imply the ability to run many concurrent tasks, but no evidence explicitly states a documented concurrency limit, dashboard for tracking many simultaneous droids, or independent case study proving throughput scaling. Missing for 10: explicit concurrency/parallelism guarantees or limits, a multi-task monitoring UI description, and independent/hands-on validation of running many tasks simultaneously.

                                                                    • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                                                                    • [claimed-docs] Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration.
                                                                    • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
                                                                    • [claimed-docs] Open Software Factory to see your delivery lifecycle as an automation coverage map.

                                                                    Jules architecture (cloud VMs per task, API for creating tasks, GitHub issue labeling) supports dispatching many tasks in parallel, and usage-limit docs explicitly reference 'agent-heavy workflows' with numeric caps (e.g. 300/60), and community reports (jules-comm-11, jules-comm-16) describe handling large task volumes successfully. However, other evidence shows real caveats: daily task limits were cut from 60 to 15 on the free plan (jules-comm-8), tasks can get stuck in loops with no stop button (jules-comm-14), and users report needing heavy babysitting (jules-comm-18), undercutting smooth high-throughput scaling. Missing for 10: dedicated documentation of a concurrency/queue dashboard, enterprise-tier concurrency guarantees, and independent benchmarks confirming reliable parallel execution at scale.

                                                                    • [claimed-docs] Power users & agent-heavy workflows ... 300 ... 60
                                                                    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                    • [community] I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…
                                                                    • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
                                                                    • [community] The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…
                                                                    • [community] Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…
                                                                    • [community] Do people really find Jules useful? I find it needs babysitting much more than Cursor.
                                                                  2. engineering-leadCreate agent sessions on behalf of other users in my organization

                                                                    weight 2 · round drawn
                                                                    Factorynone0/10

                                                                    The API reference (factory-docs-13) shows session creation/management exists, but nothing in the evidence indicates an org-admin or lead can create/manage sessions on behalf of other specific users within an organization. missing for 10: evidence of org-level user impersonation, admin controls for delegating sessions to teammates, or role-based session ownership across users.

                                                                    • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.
                                                                    Julesnone0/10

                                                                    No evidence of any organization, team, or admin/delegation features allowing an engineering-lead to create sessions on behalf of other users; Jules docs only describe individual API keys, per-user GitHub connections, and personal task creation. Missing for 10: org/team management, delegated session creation, role-based admin controls, multi-user account provisioning.

                                                                    • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
                                                                    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                    • [claimed-docs] Jules needs access to your repositories in order to work.

                                                                  Deployment flexibility

                                                                  1. developerUse a managed cloud offering to run agents without operating my own backend infrastructure

                                                                    weight 2 · round to Jules
                                                                    Factoryfullclaimed7/10

                                                                    Factory's docs describe a hosted 'Factory App' (web/desktop) where users delegate tasks and review diffs without local setup, plus an API for creating and managing Droid sessions server-side, indicating a managed backend rather than self-hosted infra. Missing for 10: independent/hands-on confirmation of the cloud service's scalability and infrastructure details beyond first-party docs.

                                                                    • [claimed-docs] Delegate a task, review the diff, and merge from the App or your terminal.
                                                                    • [claimed-docs] Run the /readiness-report slash command in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness l…
                                                                    • [claimed-docs] Automatically fix failing criteria with the /readiness-fix slash command in a Factory App web or desktop session, or in the Droid CLI
                                                                    • [claimed-docs] Create and drive Droid sessions: manage their lifecycle, settings, and messages.

                                                                    Jules is explicitly a cloud-hosted agent service: tasks run in Google-managed VMs, integrated with GitHub, with an API and usage limits/plans, requiring no self-hosted backend. Community evidence corroborates real-world usage at scale (75% success on customer PRs, teams dispatching many tasks) though some users report reliability/UI issues. Missing for 10: no independent infra/SLA details, no discoverable OpenAPI/docs.md confirming API completeness, and mixed community reliability reports.

                                                                    • [claimed-docs] Jules needs access to your repositories in order to work.
                                                                    • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                                                    • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                    • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                    • [claimed-docs] Power users & agent-heavy workflows ... 300 ... 60
                                                                    • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…
                                                                    • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
                                                                  2. engineering-leadSelf-host agent infrastructure locally, in containers, or on my own VMs

                                                                    weight 2 · round drawn
                                                                    Factorynone0/10

                                                                    No evidence describes self-hosting Factory's agent infrastructure locally, in containers, or on customer-owned VMs; all evidence points to Factory's own cloud-hosted App, CLI, and API rather than a deployable/self-hosted backend.

                                                                      Julesnone0/10

                                                                      Jules is explicitly cloud-hosted: it runs tasks in Google's own VMs, and community feedback confirms it is not local (jules-comm-1). There is no evidence of any self-hosting, on-prem, container, or private-VM deployment option.

                                                                      • [claimed-docs] Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.
                                                                      • [community] Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…

                                                                    Headless automation

                                                                    1. developerRun an agent headlessly inside CI/CD pipelines and shell scripts

                                                                      weight 2 · round to Factory
                                                                      Factoryfullclaimed9/10

                                                                      Factory explicitly documents droid exec as a one-shot headless command designed for CI/CD pipelines, shell scripts, and batch processing, with tiered autonomy controls for unattended operation. Missing for 10: independent/hands-on confirmation of real-world CI pipeline usage beyond first-party docs.

                                                                      • [claimed-docs] droid exec runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch process…
                                                                      • [claimed-docs] Droid exec uses tiered autonomy to control what operations can run without manual confirmation.
                                                                      • [claimed-docs] Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow.

                                                                      Jules exposes a public API (create tasks, send messages, optional automationMode) explicitly pitched for 'custom workflows' and embedding into other tools, which can be scripted/curled headlessly, and a community member built an MCP server hitting the API to dispatch tasks programmatically. However there's no documented CLI, no explicit CI/CD pipeline examples (e.g., GitHub Actions), and Jules always runs its own cloud VM rather than a lightweight headless process invokable inline in a shell script/pipeline step. missing for 10: official CLI or CI/CD pipeline integration docs (e.g. GitHub Actions step), examples of shell-script invocation, confirmation that API calls run synchronously enough for pipeline gating.

                                                                      • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                      • [claimed-docs] In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.
                                                                      • [claimed-docs] The `automationMode` field is optional. By default, no PR will be automatically created.
                                                                      • [claimed-docs] To send a message to the agent:
                                                                      • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
                                                                      • [community] I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…

                                                                    Not comparable on these axes

                                                                    1. ai-native userConnect an agent via an official MCP server

                                                                      weight 3 · not comparable
                                                                      Factorypartialprobed5/10

                                                                      Factory is primarily a coding agent (Droid), and factory-docs-4 shows it can connect OUT to MCP tools (client-side), which wouldn't count under the na rule — but factory-probe-3 explicitly indicates an official MCP server is documented at docs.factory.ai/harness/mcp, flipping this to the applicable server-side capability. However, the evidence pack only confirms the page's existence via a probe hit with no detail on what the server exposes or how other agents connect to it. Missing for 10: detailed documentation of the MCP server's tool surface, setup instructions for external agents, and independent confirmation of a working connection.

                                                                      • [probe] official MCP server documented at https://docs.factory.ai/harness/mcp
                                                                      • [claimed-docs] Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems
                                                                      Julesn/a

                                                                      Jules is itself a coding agent (client role), and the evidence only shows a REST API plus a third-party/community-built MCP server (jules-comm-19), not an official first-party MCP server exposing Jules as a tool endpoint. Per the agent-role rule, this axis is out of category rather than a failed capability.

                                                                      • [claimed-docs] You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…
                                                                      • [community] Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.
                                                                    2. ai-native userGet AI-generated insights and suggestions from my data inside the product

                                                                      weight 2 · not comparable
                                                                      Factoryn/a

                                                                      Factory is an agentic coding platform focused on delegating software development tasks (code diffs, PRs, readiness reports), not a data analytics or BI product that surfaces AI-generated insights/suggestions from a user's own data. This story is a category error for this product type.

                                                                        Jules generates AI-derived plans, diffs, and PR suggestions based on analysis of the user's repository data, which is the coding-agent analog of 'AI-generated insights/suggestions from your data.' Community evidence is mixed on the quality of these suggestions, with some praising usable PRs and others criticizing low-quality output on complex codebases. Missing for 10: no evidence of broader analytics-style insights beyond code-change suggestions, and no independent benchmarking confirming insight quality across use cases.

                                                                        • [claimed-docs] Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.
                                                                        • [claimed-docs] Jules provides a diff of the changes. Quickly browse and approve code edits.
                                                                        • [claimed-docs] Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.
                                                                        • [community] its way better than the github thing in my experience it produces usable PRs
                                                                        • [community] I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…