Jules vs YYLO
Jules wins · 27–14 (20 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round drawnJulesnone0/10Probes show no llms.txt, docs.md, or machine-readable API spec exist at expected paths, and no evidence Jules can be pointed at such agent-oriented documentation formats; Jules does support AGENTS.md for repo context but that's a different mechanism than consuming llms.txt-style docs.
YYLOnone0/10Direct probes show no llms.txt (404), no docs.md machine-readable variant (404), and no OpenAPI spec (404), meaning there is no agent-oriented docs endpoint to point an agent at; the only docs are standard human-facing HTML pages.
ai-native userRun the product headlessly / in CI for automation
weight 2 · round drawnJules provides a public API with API keys for creating custom workflows, sending tasks programmatically, and an optional automationMode to auto-create PRs, plus a community example of building an MCP server to dispatch tasks from another tool — this directly supports headless/CI-style automation. Missing for 10: no official CI/CD integration examples (e.g. GitHub Actions), no documented webhook/polling pattern for task completion, and no independent verification of reliability at scale in automated pipelines.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLO is a CLI orchestrator with scriptable commands (init, task start, ledger, loop, parallel-runner, run-until-completion) that produce structured watch receipts with exit codes and JSON state, and it's installable via npm as a public package with documented bins (yylo, yy, ypl) suitable for CI invocation. missing for 10: no explicit CI pipeline example (e.g., GitHub Actions config), no independent hands-on report of running it headlessly in CI, and llms.txt/openapi probes returned 404 suggesting thinner machine-readable integration docs.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “`yy loop` repeats arbitrary shell commands sequentially.”
- [claimed-docs] “Bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item.”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [github] “A successful run prints the installed YYLO version, initializes `.juno_task/`, then emits a watch receipt with `"state":"COMPLETED"`, `"exit…”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round drawnJulesnone0/10No evidence that Jules can connect to or consume external MCP servers as tool providers; the only MCP-related evidence (jules-comm-19) describes someone building an MCP server that calls INTO the Jules API, which is the reverse integration direction, not Jules plugging in MCP servers for its own tool use.
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
ai-native userUse an official CLI
weight 2 · round to YYLOJulesnone0/10Evidence shows Jules offers a web app, GitHub integration, and a REST API for custom workflows, but no official CLI tool is documented or mentioned anywhere in the evidence pack.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
YYLO ships an official CLI (@yylo/cli) with documented commands (init, ledger, loop, doctor, task start, merge land) and is verifiably published on npm plus MIT-licensed on GitHub, matching an AI-native agentic workflow tool. missing for 10: independent third-party usage reports/reviews beyond vendor docs and registry probes, and some llms.txt/docs-md/openapi endpoints 404 suggesting incomplete machine-readable doc surface.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
ai-native userDrive the product through a documented public API
weight 3 · round to JulesJules ships a documented public API (developers.google.com/jules/api) with API key management, custom workflows, sending messages to the agent, and automationMode config; a community user independently built an MCP server on top of the API confirming real-world programmatic access. Missing for 10: no discoverable OpenAPI/swagger spec or llms.txt (probes returned 404s), reducing machine-readability/self-service tooling confidence.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
YYLO exposes a well-documented CLI (`yylo`, `yy`, `ypl`) with structured JSON output, a ledger/task API, YAML workflow contracts, and templating for chaining steps, which an AI-native user could script against — supported by first-party docs and a live public npm package. However, there is no true public HTTP/OpenAPI-style API: explicit probes for llms.txt, docs.md, and OpenAPI specs all returned 404, meaning the only 'API' is the CLI surface, not a documented network API a remote agent could call directly. Missing for 10: an OpenAPI/REST API spec, an llms.txt or machine-readable API manifest, and evidence of remote/programmatic (non-CLI) invocation.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “`yy loop` repeats arbitrary shell commands sequentially.”
- [claimed-docs] “Save the same contract as YAML for a reusable workflow.”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
- [probe] “PROBE llms.txt: HTTP 404 at https://yylo.dev/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://yylo.dev/docs/yylo.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://yylo.dev/openapi.json, https://yylo.dev/swagger.json, https://yylo.dev/api/openapi.json, htt…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
ai-native userBuild against official SDKs
weight 2 · round to JulesJules ships an official public API with documented endpoints, API key management, and use cases for building custom workflows/integrations, and a community member confirms building a personal MCP server against the Jules API. missing for 10: no official OpenAPI/SDK spec discoverable (openapi probes 404), no first-party language SDKs mentioned, and no independent SDK-quality corroboration beyond one community integration example.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
YYLOnone0/10YYLO ships an official CLI (@yylo/cli) and open MIT-licensed repo, but there is no evidence of an SDK (client library/API) to build against — OpenAPI/spec probes and llms.txt/docs-md endpoints all return 404, and no SDK docs are mentioned anywhere in the pack.
- [probe] “PROBE llms.txt: HTTP 404 at https://yylo.dev/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://yylo.dev/docs/yylo.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://yylo.dev/openapi.json, https://yylo.dev/swagger.json, https://yylo.dev/api/openapi.json, htt…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
ai-native userSubscribe to events via webhooks
weight 2 · round drawnJulesnone0/10Evidence shows Jules has an API for creating tasks/sending messages, notifications for task completion, and GitHub integration, but there is no mention of webhooks or event subscription mechanisms anywhere in the docs or community evidence. Missing for 10: any documentation of a webhook endpoint, event subscription API, or push-based notification mechanism to external systems.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
YYLOnone0/10No evidence of any webhook subscription mechanism; YYLO documents Slack/GitHub integrations pulling into kanban but nothing about outbound event webhooks, and API/openapi probes returned 404s.
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [probe] “PROBE openapi: all candidate paths 404 (https://yylo.dev/openapi.json, https://yylo.dev/swagger.json, https://yylo.dev/api/openapi.json, htt…”
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to JulesJules generates AI-derived plans, diffs, and PR suggestions based on analysis of the user's repository data, which is the coding-agent analog of 'AI-generated insights/suggestions from your data.' Community evidence is mixed on the quality of these suggestions, with some praising usable PRs and others criticizing low-quality output on complex codebases. Missing for 10: no evidence of broader analytics-style insights beyond code-change suggestions, and no independent benchmarking confirming insight quality across use cases.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
YYLOnone0/10YYLO is an orchestration/CLI tool for running coding agents, workflows, and task ledgers; it launches subagents to perform tasks but nothing in the evidence describes it generating analytic insights or suggestions from a user's own data. Merge/validation explicitly avoids invoking models (yylo-gh-6), and no docs mention dashboards, analytics, or AI-generated insight surfacing.
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to YYLOJules supports background/autonomous execution via GitHub label-triggered tasks and a public API for building custom automations (e.g., automated bug-fixing, code review) with an optional automationMode for auto-PR creation, and it runs tasks in a cloud VM without needing the user present. However, docs also show a human-approval gate (plan approval, PR review) rather than fully hands-off automation, and community reports note tasks getting stuck, hitting limits, or needing babysitting, undercutting reliability of unattended runs. Missing for 10: evidence of true scheduled/cron-style recurring automations, and independent confirmation that automations run to completion without manual intervention.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
Docs describe multiple mechanisms for autonomous, repeated execution — `yy loop` for repeating shell commands, run-until-completion scripts that iterate until kanban work is done, a workflow-runner for chained multi-step automations, a bounded parallel-runner for concurrent fan-out, and integrations that pull Slack/GitHub work into a kanban queue and post results back. This supports background-style autonomous task execution, and sessions can be resumed via `yy continue SESSION_ID` rather than reconstructed from a terminal. However, there's no evidence of a persistent scheduler/cron-like trigger or a hosted background service — everything appears CLI/session-driven, and there is no independent or hands-on confirmation that these loops truly run unattended over long periods. Missing for 10: evidence of scheduled/triggered automations independent of an active session, and third-party corroboration of long-running unattended execution.
- [claimed-docs] “`yy loop` repeats arbitrary shell commands sequentially.”
- [claimed-docs] “Bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item.”
- [claimed-docs] “Continue a captured session with `yy continue SESSION_ID`; do not reconstruct work from tmux scrollback.”
- [claimed-docs] “Choose Workflow Runner when a step consumes `{{ steps.<id>.response }}`, a generated file, or a session from an earlier step.”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to JulesJules is exactly this kind of built-in agentic assistant: users delegate coding tasks via GitHub issue labels, the web app, or API, and Jules autonomously plans, clones repos, runs in a VM, and returns diffs/PRs for approval (jules-docs-2,3,4,7,8,9). Community evidence corroborates real-world task delegation and usable output (jules-comm-2, jules-comm-11, jules-comm-15, jules-comm-16), though mixed reports of loops, babysitting needs, and reliability issues (jules-comm-14, jules-comm-18) temper quality. Missing for 10: independent benchmarking of task success rates and evidence of consistent reliability at scale.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
YYLO's CLI lets a user delegate a task to an AI agent via `--subagent claude` and manage it through kanban/ledger workflows, so task delegation to an AI is documented, but the AI capability is an external subagent (e.g., Claude) invoked by the orchestrator rather than an assistant built into YYLO itself. missing for 10: evidence of an assistant embedded in the product (not requiring an external model/agent), and any first-party assistant UI/API rather than orchestration of third-party agents.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
ai-native userOperate the product with natural-language commands
weight 2 · round to JulesJules is designed around natural-language task submission and messaging: users submit a task description, Jules generates a plan, and you can send follow-up messages to the agent via API or web app (jules-docs-4, jules-docs-12). Community reports confirm real-world use of NL prompts to drive coding tasks, including iterative feedback and error-message follow-ups (jules-comm-2, jules-comm-15, jules-comm-16). Missing for 10: no detailed documentation of the full range/complexity of natural-language commands supported (e.g., multi-step conversational control, command reference) and no independent benchmark of NL command robustness beyond anecdotal reports.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “To send a message to the agent:”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
YYLOnone0/10YYLO's interface is a structured CLI (yylo/yy commands with explicit flags like --task, --subagent) rather than a natural-language command interface; the --task string is passed through to a subagent, not parsed as an operator instruction to YYLO itself. No evidence shows a chat-like or NL command surface for driving YYLO's own operations (init, start, finish, ledger, merge, etc.).
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round drawnJulesnone0/10Jules documents an API (create tasks, send messages, API keys) but there is no evidence of an interactive API reference or runnable code examples; probes for openapi/swagger specs all returned 404s, suggesting no such interactive reference exists.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “To send a message to the agent:”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
YYLOnone0/10YYLO is a CLI orchestrator for coding agents, not an API product; there is no evidence of an interactive API reference, and explicit probes for openapi.json/swagger.json and llms.txt all return 404, indicating no such reference exists.
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnJulesnone0/10Jules ships a documented REST API (jules-docs-9..12) but there is no evidence of a downloadable OpenAPI/Swagger spec; explicit probes for llms.txt, docs.md, and standard OpenAPI paths all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), confirming no machine-readable spec is published.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
YYLOnone0/10Explicit probes for llms.txt, docs-md, and OpenAPI/swagger endpoints all returned 404, and no evidence shows any downloadable machine-readable API spec; YYLO is a CLI orchestrator without a documented API surface to spec.
ai-native userTest against a sandbox environment without touching production data
weight 1 · round to YYLOJules runs tasks in an isolated cloud VM that clones the repo, meaning code changes happen in a sandboxed environment rather than directly on production, and PRs must be reviewed/approved before merging to the real branch. However, this is a code-sandbox for making changes, not a dedicated 'test against sandbox data/environment' feature, and there's no evidence of test-data isolation, staging environment provisioning, or explicit protection of production data/services beyond the VM/PR review flow. missing for 10: explicit sandbox test-data isolation, staging/production separation guarantees, independent verification that production systems are never touched.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
YYLO's task worktrees isolate work from the protected target SHA, preflight checks are documented as read-only, and merges compose changes into a private detached candidate rather than touching the live branch directly, while the benchmark tool explicitly runs 'isolated attempts' with 'recover safely' semantics—together these describe a sandboxed testing flow that avoids touching the protected/production state. Missing for 10: explicit 'production data' terminology or a dedicated staging/prod environment concept, and independent (non-vendor) confirmation that isolation holds up in practice.
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`doctor workspace` is intentionally nonzero when it finds an actionable topology problem; it never fetches or changes the workspace.”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [claimed-docs] “Plan immutable experiments, execute isolated attempts, retain evaluator provenance, recover safely, and produce bounded reports.”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnJulesnone0/10Jules has a public API (jules-docs-9, jules-docs-10) but there is no evidence of API versioning scheme or a documented deprecation policy; probes for OpenAPI spec/docs.md/llms.txt all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), suggesting no formal machine-readable API contract or lifecycle documentation is available.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
YYLOnone0/10No evidence of any versioned API contract or documented deprecation policy; the product is a CLI orchestrator with a version like 0.2.1rc6, but no API versioning scheme or deprecation guarantees are mentioned, and openapi/llms.txt probes returned 404.
- [probe] “PROBE openapi: all candidate paths 404 (https://yylo.dev/openapi.json, https://yylo.dev/swagger.json, https://yylo.dev/api/openapi.json, htt…”
- [probe] “PROBE llms.txt: HTTP 404 at https://yylo.dev/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://yylo.dev/docs/yylo.md”
- [claimed-docs] “The 0.2.1rc6 channel adds ID-first general Records and typed task, wiki, workflow, and artifact profiles.”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to YYLOJules exposes an API that lets users script custom workflows and dispatch tasks programmatically (jules-docs-9, jules-docs-12), and usage-limit tiers explicitly target 'power users & agent-heavy workflows' with up to 300 tasks (jules-docs-13), while community reports confirm it 'handles large quantities of tasks very well' and users built custom dispatch tooling via the API/MCP bridge (jules-comm-11, jules-comm-19). However, there's no documented native UI or endpoint for a single bulk action across many items (e.g., batch-applying one task to many repos/issues) — each task still appears to be created and reviewed individually via GitHub labels or API calls. Missing for 10: a documented batch/bulk endpoint or UI feature, first-party bulk-operation examples, and independent verification of true parallel bulk execution rather than just high per-day task volume.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLO's parallel-runner explicitly supports 'bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item,' and run-until-completion repeats iterations until all kanban work is done, directly enabling bulk operations across many items with automated evidence capture. This is corroborated by a live public CLI (npm registry, MIT-licensed source), though there's no independent hands-on report of large-scale fan-out in practice. Missing for 10: independent/third-party validation of bulk-scale runs and concrete concurrency limits or throughput numbers.
- [claimed-docs] “Bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item.”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [claimed-docs] “Continue a captured session with `yy continue SESSION_ID`; do not reconstruct work from tmux scrollback.”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round drawnJules supports some event-driven automation — e.g., assigning a 'jules' label to a GitHub issue automatically triggers a task, and the API lets developers build custom workflows/automations around Jules (jules-docs-2, jules-docs-9). However, there's no evidence of a general rules/conditions engine where users define arbitrary triggers and actions beyond this specific label mechanism and API calls. Missing for 10: a configurable rules interface (conditions + multiple trigger types), documentation of scheduled/webhook-based triggers, and independent confirmation that automated rule-based triggering works reliably in practice.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
YYLO's integrations pull Slack messages and GitHub issues into its kanban system and return responses to source threads, which is a form of event-triggered automation, and run-until-completion/workflow-runner allow chained/looping actions. However there's no evidence of a general user-defined rule engine (conditions + custom triggers) — the event handling is limited to specific hardcoded integrations rather than an open rule-definition system. Missing for 10: a documented rules/conditions DSL, support for arbitrary custom event sources/triggers, and evidence of user-authored trigger logic beyond the built-in Slack/GitHub integrations.
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [claimed-docs] “Choose Workflow Runner when a step consumes `{{ steps.<id>.response }}`, a generated file, or a session from an earlier step.”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to YYLOJulesnone0/10Evidence covers task assignment via GitHub labels, API-driven custom workflows, and one-off task automation, but nothing describes recurring/scheduled jobs (e.g., cron-like triggers) — the API docs only mention creating tasks and sending messages, not recurrence.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
YYLO supports repeatable workflows (YAML-saved contracts, `yy loop` shell repetition, run-until-completion looping, workflow-runner chaining) which cover automation-depth, but there is no evidence of a scheduler, cron-like trigger, or time-based recurring job execution — workflows appear to be manually invoked or run-to-completion rather than scheduled on a recurring basis. missing for 10: cron/schedule syntax or trigger mechanism, evidence of unattended recurring execution (e.g., daily/hourly runs), and independent confirmation of scheduling in production use.
- [claimed-docs] “`yy loop` repeats arbitrary shell commands sequentially.”
- [claimed-docs] “Save the same contract as YAML for a reusable workflow.”
- [claimed-docs] “Choose Workflow Runner when a step consumes `{{ steps.<id>.response }}`, a generated file, or a session from an earlier step.”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
ai-native userVersion, review, and roll back my automations
weight 1 · round to YYLOJulesnone0/10Jules produces PRs and diffs via GitHub but there is no evidence of a mechanism to version, review history of, or roll back automations/tasks themselves (as opposed to code changes tracked by git/GitHub). missing for 10: automation versioning/history feature, rollback/undo of Jules tasks, changelog or audit trail for automations.
YYLO ledger provides task versioning-like constructs (immutable task source, worktree branching, candidate composition, expected-old ref protection for merges) and workflow YAML contracts can be saved/reused, giving some review/rollback semantics via git-based merge protections. However there is no explicit 'automation version history', diff/review UI, or a documented rollback command for automations themselves beyond git branch/worktree mechanics. missing for 10: dedicated automation versioning/history view, explicit rollback command for a run/automation (not just git ref protection), review workflow UI, and independent/hands-on confirmation of these rollback flows.
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [claimed-docs] “Save the same contract as YAML for a reusable workflow.”
- [claimed-docs] “The 0.2.1rc6 channel adds ID-first general Records and typed task, wiki, workflow, and artifact profiles.”
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation
End-to-end implementation by the agent — multi-file changes, task completion
End to end feature delivery
ai-native userHave an agent automatically generate and run tests to validate its own code changes before proposing them
weight 2 · round to JulesJules runs code changes in an isolated VM and generates diffs/PRs for review, and one community report notes it is 'good at making tests for very specific/isolated functions,' implying some test generation capability, but there is no first-party documentation describing an explicit test-generate-and-run validation loop before proposing changes. missing for 10: official docs on automated test generation/execution as a validation step, evidence of test results being surfaced to users before PR creation, and independent confirmation across varied codebases.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
YYLOnone0/10The evidence explicitly states that tests and semantic reviews are 'explicit project checks outside merge' and that YYLO's merge step 'launches no models, chooses no reviewers, schedules no suites, and maintains no validation cache' (yylo-gh-6). While `yy loop` can repeat arbitrary shell commands, there is no evidence of an agent autonomously generating tests or validating its own changes before proposing them.
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [claimed-docs] “`yy loop` repeats arbitrary shell commands sequentially.”
developerHave an agent autonomously diagnose and fix a reported bug
weight 3 · round to JulesJules is designed to autonomously diagnose and fix issues: it clones the repo into a VM, generates a plan, makes code changes, and opens a PR for review, and can be triggered directly from a GitHub issue via the 'jules' label (a common bug-report workflow). Community evidence corroborates real-world bug-fixing use, including a case where it failed then successfully fixed the issue after being given the error message. Missing for 10: independent benchmarking on bug-fix success rate and more consistent evidence across complex codebases (some reports of failures/loops on harder tasks).
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
YYLO orchestrates coding agents through task worktrees, kanban-driven work items, and iteration loops (run-until-completion, subagent tasking) that could be pointed at a bug-fix task, and it supports ingesting GitHub issues into kanban as a trigger. However there is no evidence of autonomous bug diagnosis (root-causing, log/trace analysis) as a distinct capability — the docs describe generic task/workflow orchestration and merge/validation boundaries rather than an end-to-end 'diagnose then fix' agent behavior. missing for 10: explicit diagnosis/root-cause capability, an end-to-end bug-fix case study or hands-on validation, evidence the agent itself (vs. the orchestrator) performs debugging reasoning.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
product-managerGo from a mockup or design to a working implementation without an engineering handoff
weight 2 · round drawnJulesnone0/10Evidence shows Jules operates on code repositories via GitHub issues/tasks, plans, diffs, and PRs, but there is no mention of accepting mockups/designs as input or any PM-oriented no-code workflow — usage still requires repo access, issue creation, and reviewing technical diffs/plans. Nothing in the pack demonstrates a design-to-implementation path bypassing engineering.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
YYLOnone0/10YYLO's evidence describes a CLI orchestrator for coding agents (task/worktree/merge/ledger workflows) aimed at developers and project operators, with no mention of ingesting mockups/designs or enabling a non-technical PM to go from a design to working code without engineering involvement.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
developerHave an agent implement a requested feature end-to-end, including writing tests
weight 3 · round to JulesJules is explicitly built for autonomous end-to-end coding: it clones the repo, generates a plan, modifies files, produces a diff, and opens a PR (jules-docs-3/4/7/8), and one user specifically notes it's 'good at making tests for very specific/isolated functions' (jules-comm-11), with others confirming usable PRs and time savings (jules-comm-6, jules-comm-16). However, community evidence shows significant caveats: it struggles on complex/monorepo codebases, gets stuck in loops without a stop button, needs heavy babysitting, and often fails to complete tasks unassisted (jules-comm-3, jules-comm-9, jules-comm-10, jules-comm-13, jules-comm-14, jules-comm-18). Missing for 10: consistent reliability across complex real-world features, broader evidence of comprehensive test coverage (not just isolated functions), and independent benchmarks confirming end-to-end success rate.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “I've been playing with it, and I've been generally not impressed. There are obvious annoying UI bugs and the output isn't very good for anyt…”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
YYLO documents an end-to-end task loop (yylo init --task ... --subagent claude, worktree creation, commit-and-queue candidate flow) showing an agent can implement a described feature autonomously, but the evidence explicitly states 'tests and semantic reviews are explicit project checks outside merge' — meaning YYLO's own agent loop does not itself guarantee test-writing as part of implementation, only that separate check scripts exist for validation. missing for 10: explicit evidence the invoked subagent writes/adds tests as part of a task, and any example showing test-authoring within the init/finish workflow.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
Environment setup
developerHave an agent automatically clone the repo, install dependencies, and configure its own working environment
weight 2 · round to JulesJules' core workflow is documented: it runs in a cloud VM, clones the repo, installs dependencies, and configures its environment automatically before making changes, corroborated by community usage reports of automated PR generation. missing for 10: independent technical verification of dependency-install robustness across complex/monorepo setups, and some community reports note environment/config confusion in bespoke or monorepo projects.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
Docs and GitHub README describe `task start`/`yy init` hydrating a dedicated worktree from a protected base SHA and completing 'configured dependency hydration' before reporting WORKING, which covers automated environment setup and dependency install; the CLI is confirmed live on npm and the repo public. Missing for 10: explicit description of cloning an arbitrary remote repo (vs. hydrating a pre-defined workspace), and independent/hands-on confirmation that dependency install works end-to-end.
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [github] “A successful run prints the installed YYLO version, initializes `.juno_task/`, then emits a watch receipt with `"state":"COMPLETED"`, `"exit…”
Interactive takeover
developerTake over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
weight 2 · round to JulesJules supports reviewing plans, diffs, and sending follow-up messages to redirect a task (jules-docs-4, jules-docs-6, jules-docs-7, jules-docs-12), and third parties have wired the API into VS Code/Copilot Chat to dispatch tasks (jules-comm-19), but this is dispatch, not mid-task takeover. A hands-on report explicitly notes there's no 'STOP' button to interrupt a looping task (jules-comm-14), undercutting redirect control, and all takeover is browser/API-based with no native terminal or editor 'take over' UI documented. Missing for 10: documented in-editor/terminal takeover UI, ability to pause/interrupt an in-progress run, and evidence the API-based messaging genuinely redirects rather than just appends instructions.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLO offers terminal-based session continuation (`yy continue SESSION_ID`) that lets a developer pick back up captured agent work rather than reconstructing it from scrollback, and task start/finish flows expose worktree/branch state that could be inspected or redirected from the CLI. However, there is no evidence of any editor integration or browser UI for taking over tasks — YYLO is documented purely as a CLI/orchestrator tool. Missing for 10: editor plugin/extension support, browser-based task takeover UI, and explicit interactive 'redirect mid-task' semantics beyond resuming a session.
- [claimed-docs] “Continue a captured session with `yy continue SESSION_ID`; do not reconstruct work from tmux scrollback.”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
developerSend follow-up instructions to an active agent session to steer its work without restarting
weight 2 · round to JulesJules's API docs explicitly describe sending a message to an active agent session (jules-docs-12), and community reports confirm the workflow of reviewing partial work and providing feedback for Jules to continue rather than restarting (jules-comm-3). This shows the steering capability works in practice, though documentation on how follow-ups affect an in-progress plan/execution is thin. Missing for 10: detailed first-party documentation of mid-task message handling/UI chat thread, and independent hands-on confirmation of steering effectiveness beyond one anecdote.
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
YYLOnone0/10YYLO's docs describe task lifecycle (init/start/preflight/finish) and resuming a captured session via `yy continue SESSION_ID`, but there is no evidence of sending new instructions to an already-running agent session to redirect its work mid-flight; `continue` appears to resume/reattach rather than inject steering input. missing for 10: any documented mechanism for mid-session instruction injection or steering, evidence that an active agent process accepts new prompts without restart.
- [claimed-docs] “Continue a captured session with `yy continue SESSION_ID`; do not reconstruct work from tmux scrollback.”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
Sandbox execution
developerHave an agent safely execute code and install dependencies inside an isolated sandbox
weight 3 · round to JulesJules explicitly runs in a cloud VM that clones code and installs dependencies, isolated from the user's local machine, with a plan-review step before code changes are made — directly matching the sandboxed execution story. Community reports corroborate it actually running tasks end-to-end (producing PRs) in this VM environment, though some criticize reliability/loops on complex codebases. Missing for 10: detailed docs on sandbox security boundaries (network isolation, resource limits) and independent security audit of the VM isolation.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
YYLOnone0/10YYLO's docs describe git-worktree/branch isolation for tasks and merge safety, but none of the evidence mentions a sandboxed execution environment (container/VM) for running agent code or installing dependencies safely. Worktree isolation protects git state, not runtime/process isolation.
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight
Keeping a human in the loop — approvals, checkpoints, interrupts
Approval controls
developerConfigure an agent to auto-approve all its actions instead of confirming each one
weight 2 · round to JulesJules's default workflow requires manual approval at multiple stages (plan approval, diff review, PR approval), but the API docs mention an optional `automationMode` field that changes default behavior around automatic PR creation, suggesting some configurable auto-approve path exists via API rather than the standard UI flow. Missing for 10: explicit documentation of a setting that suppresses plan-approval and diff-review confirmations entirely, and any hands-on/community confirmation that this automation mode actually skips human checkpoints.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
YYLOnone0/10The evidence describes YYLO's orchestration architecture (task worktrees, ledger, merge protections) but nothing addresses a configurable auto-approve/no-confirmation mode for agent actions; in fact merge land explicitly requires checks and human-oversight-style gating rather than blanket auto-approval. Missing for 10: any documented auto-approve flag/setting, evidence of confirmation prompts being bypassable, or explicit human-oversight configuration options.
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
product-managerApprove key agent decisions from my phone while agents continue working
weight 1 · round to JulesJules is a cloud-based agent with a web app and notification system, plan approval, diff review, and PR approval steps, and community evidence confirms users review/approve PRs from a browser including 'coding on my phone' via mobile web. However, there is no dedicated native mobile app or explicit mobile-optimized approval UI, and no evidence of push-notification-driven approval flows tailored to on-the-go PM decision-making. missing for 10: dedicated mobile app or mobile-specific UI, evidence of push notifications enabling quick phone-based approvals, PM-specific (non-developer) approval workflow.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
YYLOnone0/10YYLO is documented purely as a CLI/terminal orchestrator (yylo/yy commands, ledger, worktrees, merge gating) with no mention of any mobile app, phone notification, or remote-approval interface for product managers. The axis is plausible for an agent-orchestration tool but no evidence supports it.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
engineering-leadSet tiered autonomy levels controlling what an agent can do without manual confirmation
weight 3 · round to JulesJules has a binary plan-approval workflow (review/approve plan, review diffs, approve PRs) and an API `automationMode` field that toggles whether PRs are auto-created, but there's no evidence of configurable tiered autonomy levels (e.g., low/medium/high trust settings, granular permission scopes, or per-action confirmation thresholds) that an engineering-lead could set. Missing for 10: explicit multi-tier autonomy/permission settings, admin-configurable trust levels, and any org-wide policy controls beyond the single automationMode toggle and default plan-approval gate.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
YYLOnone0/10Evidence shows preflight/read-only checks, merge protections, and validation boundaries, but there is no mention of configurable tiered autonomy levels or a settings mechanism letting an engineering-lead define graduated confirmation thresholds for agent actions. missing for 10: explicit autonomy-level configuration, tiered permission settings, evidence of user-controlled confirmation thresholds.
Model control
ai-native userHave each task prompt automatically routed to the most suitable underlying model
weight 2 · round drawnJulesnone0/10No evidence Jules performs automatic model routing per task; docs mention only a single flagship model (Gemini 3 Pro) as 'priority access,' with no mention of routing prompts to different models based on suitability. Missing for 10: any documentation of multi-model routing logic, model-selection criteria, or automatic switching between models.
- [claimed-docs] “Priority access to the latest models, starting with Gemini 3 Pro”
YYLOnone0/10Evidence shows YYLO lets users manually specify a subagent/model (e.g. `--subagent claude`) but nothing describes automatic routing of a task prompt to the 'most suitable' underlying model based on task characteristics. missing for 10: any evidence of automatic model-selection logic, routing criteria, or multi-model comparison/selection mechanism.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
engineering-leadSwitch away from automatic model selection to a specific model of my choice
weight 1 · round to YYLOJulesnone0/10No evidence pack item mentions model selection settings or the ability to choose a specific underlying model; the only related mention is 'Priority access to the latest models, starting with Gemini 3 Pro' which describes access tiers, not user-controlled model switching.
The CLI init command shows a --subagent flag (e.g. 'claude') letting a user specify which model/agent to use instead of relying on defaults, implying manual model selection is possible; however there is no documentation of an explicit 'automatic model selection' mode being overridden, no list of supported models, and no engineering-lead-oriented control/settings UI shown. missing for 10: explicit documentation of an automatic/default model-selection mode, a full list of selectable models, and confirmation that this override is persistent/configurable at a project or team level.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
Visibility monitoring
developerWatch what a running agent is doing in real time, including its current status
weight 3 · round to JulesJules provides notifications on task completion/input-needed and a plan-approval step, implying some status visibility, plus a diff/PR review flow, but there is no evidence of a live real-time activity feed or streaming log of the agent's current actions while running; community feedback even notes the absence of a stop/interrupt control mid-run, suggesting limited real-time observability. Missing for 10: evidence of a real-time execution log/console view, granular step-by-step status updates, and independent confirmation of live monitoring UX.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
The GitHub docs mention a successful run ending in a 'watch receipt' with a state field (e.g. COMPLETED, exit_code, log_bytes), implying some status-tracking mechanism exists, and 'task start' reports a WORKING state before completion. However there is no dedicated 'watch' command, live dashboard, or streaming log evidence showing real-time observation of an in-progress agent's actions. missing for 10: explicit real-time streaming/monitoring command or UI, documentation of live status polling during execution, independent confirmation of watch behavior.
- [github] “A successful run prints the installed YYLO version, initializes `.juno_task/`, then emits a watch receipt with `"state":"COMPLETED"`, `"exit…”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [claimed-docs] “Continue a captured session with `yy continue SESSION_ID`; do not reconstruct work from tmux scrollback.”
developerGet notified when an agent completes a task or needs my input
weight 2 · round to JulesFirst-party docs explicitly state users are notified when a task completes or needs input, and the workflow (plan approval, diff review, PR creation) reinforces oversight checkpoints where notifications matter. Community evidence corroborates the async, check-back-later usage pattern (e.g. reviewing PRs after time away), consistent with notification-driven workflows. Missing for 10: independent verification of notification channels (email/push/Slack), and no detail on notification reliability or configurability.
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
YYLO emits watch receipts with a COMPLETED state/exit_code after a run finishes, and its Slack/GitHub integrations can return completed responses to the originating thread, giving developers a way to learn a task finished. However there's no evidence of a proactive notification for when an agent 'needs input' mid-task, nor any push/alert channel beyond polling receipts or the source-thread reply. Missing for 10: explicit 'needs input' alerting, a dedicated notification/webhook system, and independent confirmation the integration loop works end-to-end.
- [github] “A successful run prints the installed YYLO version, initializes `.juno_task/`, then emits a watch receipt with `"state":"COMPLETED"`, `"exit…”
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
Intent to spec — stories about intent to spec in this arenaIntent to spec
Stories about intent to spec in this arena
Natural language task intake
developerDescribe a feature or bug in plain language and have it automatically turned into a scoped implementation task
weight 3 · round drawnJules docs clearly show a plain-language task submission flow that generates a reviewable/approvable plan before code changes, then produces a diff and PR (jules-docs-4, jules-docs-7, jules-docs-8), directly matching the intent-to-spec story. However community reports show inconsistent scoping quality — confusion in monorepos, endless loops, and needing heavy babysitting — indicating the automatic scoping doesn't always hold up in practice. Missing for 10: independent verification of plan/spec quality on complex codebases, and consistent evidence the generated plan reliably matches developer intent without back-and-forth correction.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
YYLO's CLI lets a developer pass a plain-language task string (e.g. `yylo init --task "Describe one verifiable outcome"` or `yy ledger create "Validate recovery"`) which is then hydrated into a dedicated branch/worktree and handed to a subagent (yylo-docs-1, yylo-docs-2, yylo-gh-4, yylo-docs-11). This covers the 'turned into a scoped implementation task' half of the story, but there's no evidence of NLP-based scoping/decomposition logic — the description appears passed through largely as-is rather than analyzed/refined into a structured spec. Missing for 10: evidence of automatic task decomposition or requirement extraction from free-text input, and independent/hands-on confirmation that vague feature/bug descriptions produce well-scoped tasks.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
product-managerConvert user feedback submissions into structured tasks with proposed scope
weight 2 · round to YYLOJulesnone0/10Jules's documented workflow is a coding-agent pipeline (clone repo, generate a plan, produce diffs/PRs) driven by developer-specified tasks or GitHub issues, not a mechanism for ingesting unstructured user-feedback and converting it into a structured task with proposed scope for a PM. No evidence shows feedback intake, requirement structuring, or scope proposal features aimed at product managers.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
YYLO can ingest external feedback (Slack messages, GitHub issues) directly into its kanban/ledger as structured tasks (yylo-docs-10), and a dedicated `feedback-yylo` CLI binary is confirmed live on npm (yylo-probe-rt-1), suggesting a feedback-to-task pathway exists. However, there is no documented mechanism for generating a 'proposed scope' alongside the task — no scope estimation, sizing, or planning artifact is described anywhere in the docs or GitHub evidence. Missing for 10: explicit scope-proposal output tied to feedback conversion, documentation of what the feedback-yylo binary actually produces, and any PM-facing evidence of structured scoping fields.
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
Plan approval
developerReview and approve an agent's implementation plan before any code changes are made
weight 3 · round to JulesJules docs explicitly state the agent generates a plan upon task submission that users can review and approve before any code changes are made, directly matching the story. No independent/hands-on evidence specifically confirms or contradicts this plan-approval step, though other workflow claims (diffs, PRs) are corroborated by community mentions. Missing for 10: independent/hands-on confirmation of the plan-review step specifically, and more detail on what the plan interface looks like.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
YYLOnone0/10Evidence describes YYLO's task lifecycle (init, start, preflight read-only, finish queuing a candidate, merge land) but nothing indicates the agent produces an implementation plan that a developer reviews and approves before any code is written — preflight/checks occur on already-produced work, not a pre-code plan gate.
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
engineering-leadApprove a task's scope and contract before an agent is allowed to modify the repository
weight 2 · round to JulesJules generates a plan that the user can review and approve before any code changes are made, giving an engineering-lead a gate before modification, and it also produces diffs/PRs requiring approval before merge. However, this is a general 'plan approval' UX rather than a formal scope/contract sign-off workflow tailored for engineering-lead governance (e.g., no evidence of role-based approval gates, policy enforcement, or blocking unauthorized starts). missing for 10: explicit lead/role-based approval gating before task execution starts, evidence of enforceable contract/scope definitions beyond a plan preview, independent confirmation that the plan-approval step reliably blocks modification.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
YYLO documents a task/contract concept (YAML contracts, ledger tasks) and isolates agent work in dedicated branches/worktrees with a read-only preflight before any commit is queued (yylo-docs-2, yylo-docs-4, yylo-gh-4), which implies some spec-first gating. However there is no documented human approval/sign-off step where an engineering-lead explicitly reviews and approves scope/contract before the agent is permitted to start modifying the repo—task start appears automatic once invoked. Missing for 10: explicit lead-approval gate/workflow, evidence of a review UI or command requiring human sign-off, and confirmation that agent modification is blocked pending that approval.
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
- [claimed-docs] “Save the same contract as YAML for a reusable workflow.”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
Ticket driven tasking
developerAssign a coding task to an agent directly from an existing issue or ticket
weight 3 · round to JulesJules explicitly supports assigning tasks directly from GitHub issues via a 'jules' label, and API/community evidence corroborates task dispatch and PR creation workflows. Missing for 10: evidence of ticket-tracker integrations beyond GitHub (e.g., Jira/Linear) and independent hands-on confirmation of the label-to-task flow itself.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
YYLO integrations pull GitHub issues into its kanban ledger and return completed responses to the source thread, and ledger tasks can then be started with a subagent via task start or yylo init, supporting the flow of turning an issue into an agent task. Missing for 10: a single documented command that directly converts one specific issue into an agent task in one step, and hands-on confirmation the GitHub-issue import works end-to-end.
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “yy ledger create "Validate recovery" --status todo --tags auth yy ledger ready --sort asc yy ledger get TASK_ID”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userExport all of my data in open formats and leave
weight 3 · round to YYLOJulesnone0/10No evidence of any data export feature, open-format export, or account/data portability mechanism for Jules; the product works via GitHub repos/PRs but nothing indicates users can export Jules-specific data (task history, configs, etc.) and leave. missing for 10: any documentation of data export, open format support, or account portability/deletion workflow.
Workflow contracts can be saved as YAML (yylo-docs-4) and the CLI/ledger source is fully open-source under MIT (yylo-probe-rt-2), suggesting no vendor lock-in, but there is no documented single 'export all data' command covering ledger records, task history, sessions, and artifacts in open formats. Missing for 10: an explicit data-export command/feature, documentation of export formats for ledger/kanban/session data, and confirmation that all state (not just workflow YAML) is portable.
- [claimed-docs] “Save the same contract as YAML for a reusable workflow.”
- [claimed-docs] “The 0.2.1rc6 channel adds ID-first general Records and typed task, wiki, workflow, and artifact profiles.”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
ai-native userRead the product's source under an open license
weight 2 · round to YYLOJulesnone0/10Jules is a closed, proprietary Google product (cloud VM, API keys, usage limits); no evidence anywhere of source code being published or licensed openly, and probes for docs/openapi artifacts return 404s, further suggesting no open publishing.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
The GitHub repo is public and licensed under MIT, confirmed by a runtime probe reading the LICENSE file directly, and the repo (yylo-dev/yylo) is documented as the CLI orchestrator's source. missing for 10: no independent third-party audit or community commentary confirming completeness of the published source beyond the license file check.
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
ai-native userSelf-host the core product
weight 3 · round to YYLOJulesnone0/10Jules is explicitly a cloud-based VM service (Google-hosted) with no evidence of any self-hosting option; community even notes cloud-only design as a drawback compared to local tools like Claude Code.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “It's a shame Google picked the wrong system design for Jules. Claude Code's system design is clearly superior at this point.”
YYLO is an open-source, MIT-licensed CLI orchestrator (installable via npm, source on GitHub) that runs locally against a user's own repo/agents, so self-hosting the core product is inherently satisfied — confirmed by the public npm registry listing and the MIT LICENSE in the public repo. missing for 10: no dedicated self-hosting/deployment guide (e.g. server install, Docker, or persistent-service setup instructions) and no independent third-party report of someone self-hosting it in production.
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Enterprise licensing
engineering-leadLicense an enterprise deployment with SSO and commercial support for organization-wide rollout
weight 2 · round drawnJulesnone0/10No evidence of enterprise licensing, SSO, or commercial support offerings; docs mention only individual API keys and usage-tier limits, and community discussion references a free beta plan, not enterprise deployment. missing for 10: SSO integration, enterprise licensing/contract terms, commercial support SLAs, org-wide admin/rollout tooling.
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “Is Jules free of charge? Yes, for now, Jules is free of charge. Jules is in beta and available without payment while we learn from usage.”
YYLOnone0/10YYLO's evidence pack shows only an open-source CLI orchestrator (MIT-licensed, npm package) with no mention of enterprise licensing tiers, SSO integration, or commercial support offerings; there is no pricing/plans page or enterprise sales material in evidence. missing for 10: enterprise/SSO licensing tier, commercial support plans, organization-wide deployment documentation.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
Model flexibility
engineering-leadBring my own LLM or API key so agents run on the model of my choice
weight 2 · round drawnJulesnone0/10No evidence that Jules supports bringing a custom LLM or third-party API key; it is tied to Google's own models (Gemini), with documentation only mentioning Jules's own API key for accessing Jules itself, not for configuring underlying model providers. missing for 10: any mention of BYO-LLM support, model selection options, or third-party API key configuration.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “Priority access to the latest models, starting with Gemini 3 Pro”
YYLOnone0/10Evidence shows a `--subagent claude` CLI flag implying some agent selection, but there is no documentation of configuring API keys, choosing alternate LLM providers, or any pricing/billing control for engineering leads. Missing for 10: explicit BYO-API-key setup, multi-provider/model configuration docs, and any pricing-limits guidance.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userControl data retention and deletion
weight 2 · round drawnJulesnone0/10No evidence pack item mentions data retention policies, data deletion controls, or privacy settings for repositories/code processed by Jules; docs cover workflow, VM execution, and API usage but not retention/deletion controls.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnJulesnone0/10No evidence in the pack mentions telemetry, usage tracking, or an opt-out mechanism; the docs focus on repo access, VM execution, PR workflow, and API usage without addressing privacy/telemetry controls.
Repo integration — stories about repo integration in this arenaRepo integration
Stories about repo integration in this arena
Knowledge context
developerAdd a context file describing my codebase conventions so agents generate more relevant plans and code
weight 3 · round to JulesJules automatically reads an AGENTS.md file in the repo root to guide its plans and code generation, directly matching the story of adding a context file describing codebase conventions. Missing for 10: no independent/hands-on confirmation of how well conventions in AGENTS.md actually improve plan/code relevance, and no detail on supported format/scope beyond the single doc mention.
- [claimed-docs] “Jules now automatically looks for a file named AGENTS.md in the root of your repository.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
YYLOnone0/10The evidence pack covers task orchestration, kanban ledgers, merge protections, and workflow runners, but nowhere describes a context/conventions file that agents read to generate more relevant plans or code. No mention of AGENTS.md, project instructions, or codebase-convention configuration exists in the docs or GitHub items provided.
Project management integration
product-managerConnect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly
weight 3 · round drawnJulesnone0/10Evidence only shows Jules integrating with GitHub (repos, issues, PRs) and provides an API for custom workflows, but there is no mention of Jira, Linear, ClickUp, Monday.com, or any issue-tracker integration beyond GitHub.
YYLOnone0/10The integrations doc only mentions bringing Slack messages or GitHub issues into kanban (yylo-docs-10); there is no mention of Jira, Linear, ClickUp, or Monday.com anywhere in the evidence pack. Missing for 10: any documented connector or API integration for Jira, Linear, ClickUp, or Monday.com.
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
Version control integration
developerConnect a GitHub repository so an agent can access the code and open pull requests against it
weight 3 · round to JulesJules is built around GitHub repo access: docs describe cloning repos, GitHub issue label assignment, generating diffs/plans, and creating PRs that can be merged on GitHub, corroborated by community reports of receiving usable PRs from Jules-driven tasks. Missing for 10: independent verification of the full connect-repo setup flow and edge-case reliability across repo types.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
YYLOnone0/10Evidence shows YYLO operates on local git worktrees/branches and has an internal 'merge land' step, and can pull GitHub issues into its kanban, but there is no evidence of connecting a GitHub repository as a remote and having the agent open pull requests against it — the merge feature explicitly stays local/internal with no GitHub PR API integration mentioned.
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
developerGrant an agent access to my repositories with a one-click install, without complex setup
weight 2 · round to JulesJules integrates with GitHub (label-based task assignment, PR creation) and needs repo access to work, implying a connect/authorize flow, and community evidence confirms GitHub integration works well for many users. However, there's no explicit documentation of a literal 'one-click install' onboarding flow or app installation process described in the evidence pack. missing for 10: explicit one-click GitHub App install/authorization flow documentation, independent confirmation of setup simplicity.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
YYLOnone0/10YYLO is documented as a CLI orchestrator requiring `npm install -g @yylo/cli` plus explicit `yylo init`/`task start` workflows that freeze SHAs, create worktrees, and hydrate dependencies — this is CLI-based setup, not a one-click repo-access grant. No evidence describes a GitHub App-style one-click install or OAuth flow for repo access. Missing for 10: any one-click install/authorization mechanism, evidence of simplified non-CLI onboarding, or a hosted install button.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates
Quality gates on changes — review flow, required checks, merge protection
Ci remediation
engineering-leadHave failed CI workflows automatically diagnosed and fixed with a proposed pull request
weight 3 · round to JulesJules offers building blocks—an API for custom automation that can 'automate tasks like bug fixing' and auto-create PRs via automationMode—but there is no direct evidence of a native integration that detects a failed CI workflow and automatically diagnoses/fixes it with a PR; this would require custom API wiring, not an out-of-box feature. Missing for 10: documented CI-failure trigger/integration, evidence of automatic diagnosis of CI logs, and case studies of this specific end-to-end flow.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
YYLOnone0/10YYLO is a CLI orchestrator for coding agents/workflows with kanban, ledger, and merge tooling, but nothing in the evidence pack mentions CI workflow failure detection, diagnosis, or auto-generating a fix PR from a failing CI run. Merge/land explicitly excludes running tests or validation ('Tests and semantic reviews are explicit project checks outside merge'), which is the opposite of an integrated CI-fix loop.
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
developerTrigger an agent from CI/CD pipelines to fix a broken build or failing test
weight 2 · round to JulesJules exposes a documented API for creating custom workflows and automating tasks like bug fixing (jules-docs-9, jules-docs-10, jules-docs-12), which could be wired into a CI/CD pipeline to trigger a fix, and community evidence shows users building custom integrations to dispatch Jules tasks programmatically (jules-comm-19). However, there is no first-party CI/CD-specific integration (e.g. GitHub Actions step, build-failure webhook) or documented example of triggering Jules to fix a broken build/failing test directly from a pipeline. missing for 10: explicit CI/CD pipeline integration/example, evidence of triggering on build/test failure events, hands-on confirmation of this specific workflow succeeding.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLOnone0/10YYLO is documented as a CLI orchestrator with kanban tasks, workflow/parallel runners, and Slack/GitHub-issue integrations, but nothing in the evidence describes triggering it from a CI/CD pipeline or having it react to a failing build/test. Since it's a scriptable CLI, this axis plausibly applies, but there is no documented CI hook, GitHub Actions example, or build-failure-triggered workflow.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
Diff review
developerConfigure an agent to automatically open a pull request when its task completes
weight 2 · round to JulesJules natively creates a PR when a task completes (jules-docs-8), and the API exposes an `automationMode` field to configure automatic PR creation, defaulting to off unless explicitly configured (jules-docs-11), directly matching the story of configuring automatic PR-on-completion. Community reports corroborate that Jules produces usable PRs and users regularly review PRs it opens (jules-comm-6, jules-comm-12, jules-comm-16). missing for 10: independent hands-on verification of the automationMode toggle specifically, and more detail on configuration options/edge cases (e.g., partial completions, failed tasks) affecting whether a PR is always opened.
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
YYLOnone0/10YYLO's docs describe task branches, worktrees, and an internal 'merge land' step that composes a candidate directly, explicitly stating merge 'launches no models, chooses no reviewers'—there is no mention anywhere of opening a GitHub pull request on task completion, only importing issues into kanban and returning responses to source threads. No CLI flag, workflow step, or integration doc references PR creation.
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
developerReview a diff of an agent's changes and approve it before it becomes a pull request
weight 3 · round to JulesJules explicitly provides a diff of changes for review and approval before it creates/publishes a PR, with plan approval and diff approval steps documented; community mentions confirm PRs are generated for review. Missing for 10: independent hands-on confirmation of the diff-approval UI flow itself (comments focus on PR quality/output rather than the diff-review step specifically).
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
YYLOnone0/10Docs describe worktrees, candidate branches, and a 'merge land' step, but nowhere is there evidence of a diff-review UI or an explicit developer approval gate before a pull request is opened; in fact merge is described as launching 'no models, no reviewers' and reviews are called 'explicit project checks outside merge', with no PR-creation flow documented at all.
- [github] “`task start` freezes the protected target SHA, creates a dedicated branch/worktree, and completes configured dependency hydration before rep…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [claimed-docs] “Start hydrates an exact-base task worktree. Preflight is read-only, finish queues the committed candidate”
Pr review automation
ai-native userHave incoming issues automatically triaged with severity suggested and routed to the right owner
weight 2 · round drawnJulesnone0/10Jules is a coding agent that executes tasks assigned via label or API, but there is no evidence of automatic issue triage, severity classification, or routing to owners — the closest feature is manual label-based task assignment, not automated triage. Missing for 10: any mention of severity scoring, triage logic, or owner-routing automation.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
YYLOnone0/10YYLO's integrations feature only pulls GitHub issues/Slack messages into a kanban board and returns responses to the source thread (yylo-docs-10); there is no evidence of automated severity classification or owner-based routing logic anywhere in the pack.
- [claimed-docs] “Bring Slack messages or GitHub issues into kanban and return completed responses to their source thread.”
engineering-leadHave every pull request automatically reviewed with AI-generated inline comments
weight 3 · round drawnJulesnone0/10Jules is documented as a task-execution agent that clones repos, makes code changes, and creates its own PRs for approval — not as a bot that automatically reviews incoming pull requests with inline comments. The only tangential mention is jules-docs-9's passing reference to 'automate tasks like ... code reviews' via API, but there is no documentation of an automatic inline-comment review gate applied to every PR.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
YYLOnone0/10YYLO's own docs describe it as a CLI orchestrator for coding-agent tasks, workflows, and receipt-backed merges — not a PR-review tool. Evidence explicitly states 'Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers' (yylo-gh-6), meaning there is no automated AI-generated inline PR review capability in the product.
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
- [github] “`merge land` selects one immutable task source, composes in a private detached candidate, and uses Git expected-old ref protection. A moved …”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
Readiness checks
engineering-leadAutomatically fix failing agent-readiness criteria in my repository
weight 1 · round drawnJulesnone0/10Jules can be assigned generic bug-fix/coding tasks and reads an AGENTS.md file if present, but there is no evidence of it detecting or automatically remediating 'agent-readiness' quality-gate criteria (e.g., missing/invalid AGENTS.md, agent-compatibility checks) as a review gate — it only consumes such files, it doesn't audit or fix them as a compliance gate.
- [claimed-docs] “Jules now automatically looks for a file named AGENTS.md in the root of your repository.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
YYLOnone0/10YYLO's diagnostic tool (`doctor workspace`) explicitly never fetches or changes the workspace, and merge/tests are described as explicit checks with no automated remediation; there is no documented feature that automatically fixes failing agent-readiness criteria.
engineering-leadRun a readiness report that evaluates how ready my repository is for autonomous agents
weight 2 · round to YYLOJulesnone0/10No evidence of a readiness report, scorecard, or repository-assessment feature for autonomous agent suitability; evidence only covers task execution, PR creation, and API usage. Missing for 10: any readiness-scoring feature, repo audit/checklist, or report output evaluating agent-readiness.
YYLO ships a `doctor workspace` check that flags actionable topology problems without mutating the repo, and its GitHub description references 'release-readiness boundaries,' which gesture at repo-readiness diagnostics, but there is no documented dedicated report scoring or evaluating overall repository readiness for autonomous agents. missing for 10: a named readiness-report command/output, criteria for 'agent readiness' beyond topology checks, and any sample report artifact.
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
Concurrent execution
engineering-leadRun many agent tasks concurrently to scale delivery throughput
weight 3 · round drawnJules architecture (cloud VMs per task, API for creating tasks, GitHub issue labeling) supports dispatching many tasks in parallel, and usage-limit docs explicitly reference 'agent-heavy workflows' with numeric caps (e.g. 300/60), and community reports (jules-comm-11, jules-comm-16) describe handling large task volumes successfully. However, other evidence shows real caveats: daily task limits were cut from 60 to 15 on the free plan (jules-comm-8), tasks can get stuck in loops with no stop button (jules-comm-14), and users report needing heavy babysitting (jules-comm-18), undercutting smooth high-throughput scaling. Missing for 10: dedicated documentation of a concurrency/queue dashboard, enterprise-tier concurrency guarantees, and independent benchmarks confirming reliable parallel execution at scale.
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
YYLO documents a 'bounded concurrent fan-out' parallel-runner for independent kanban tasks with structured evidence per item, plus run-until-completion looping, which supports running multiple agent tasks concurrently. However, evidence lacks details on concurrency limits, scaling numbers, resource isolation at scale, or independent/hands-on validation of throughput gains. missing for 10: concrete concurrency limits/benchmarks, independent third-party validation of parallel scaling, evidence of large-scale (10s-100s of tasks) usage in production.
- [claimed-docs] “Bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item.”
- [claimed-docs] “Repeat bounded YYLO iterations until no open kanban work remains.”
- [claimed-docs] “Continue a captured session with `yy continue SESSION_ID`; do not reconstruct work from tmux scrollback.”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
Deployment flexibility
engineering-leadSelf-host agent infrastructure locally, in containers, or on my own VMs
weight 2 · round to YYLOJulesnone0/10Jules is explicitly cloud-hosted: it runs tasks in Google's own VMs, and community feedback confirms it is not local (jules-comm-1). There is no evidence of any self-hosting, on-prem, container, or private-VM deployment option.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
YYLO is an open-source, MIT-licensed CLI (npm package + public GitHub repo) that runs locally as an orchestrator, implying it can be run on a developer's own machine, in containers, or VMs since it's just a Node CLI operating on a local git worktree. However, there is no explicit documentation of container/VM deployment, Docker images, self-hosting guides, or infrastructure requirements for running at scale. missing for 10: explicit self-hosting/deployment docs (Docker/container images, VM setup guides), infrastructure/scaling guidance, and confirmation of statelessness or multi-instance operation for parallel agent infrastructure.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
- [github] “A successful run prints the installed YYLO version, initializes `.juno_task/`, then emits a watch receipt with `"state":"COMPLETED"`, `"exit…”
Headless automation
developerRun an agent headlessly inside CI/CD pipelines and shell scripts
weight 2 · round to YYLOJules exposes a public API (create tasks, send messages, optional automationMode) explicitly pitched for 'custom workflows' and embedding into other tools, which can be scripted/curled headlessly, and a community member built an MCP server hitting the API to dispatch tasks programmatically. However there's no documented CLI, no explicit CI/CD pipeline examples (e.g., GitHub Actions), and Jules always runs its own cloud VM rather than a lightweight headless process invokable inline in a shell script/pipeline step. missing for 10: official CLI or CI/CD pipeline integration docs (e.g. GitHub Actions step), examples of shell-script invocation, confirmation that API calls run synchronously enough for pipeline gating.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
YYLO is a CLI-first orchestrator with commands (init, task start, ledger, loop, workflow-runner, parallel-runner) that are inherently scriptable and non-interactive, and its runtime bins (yylo, yy, ypl) are published on npm confirming CLI availability. However, there's no explicit CI/CD documentation, no exit-code/error-handling guarantance for pipeline use beyond doctor workspace, and no CI examples (GitHub Actions, GitLab CI, etc.) or headless/no-TTY confirmation. missing for 10: explicit CI/CD pipeline examples or docs, confirmed non-interactive/headless mode guarantees, exit-code contract documentation for scripting, independent hands-on CI usage reports.
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [claimed-docs] “`yy loop` repeats arbitrary shell commands sequentially.”
- [claimed-docs] “Bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item.”
- [github] “`doctor workspace` is intentionally nonzero when it finds an actionable topology problem; it never fetches or changes the workspace.”
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
Not comparable on these axes
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableJulesn/aJules is itself a coding agent (client role), and the evidence only shows a REST API plus a third-party/community-built MCP server (jules-comm-19), not an official first-party MCP server exposing Jules as a tool endpoint. Per the agent-role rule, this axis is out of category rather than a failed capability.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLOnone0/10YYLO is a CLI orchestrator for coding agents/workflows, and as such platform-type product it could plausibly ship an official MCP server for other agents to connect to, but no evidence pack item mentions MCP at all (only 'open-standard' skills installation, ledger, workflow-runner, etc.). No official MCP server documentation, endpoint, or announcement exists in the evidence.
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · not comparableJulesnone0/10Jules documents a basic API key creation flow (max 3 keys) but provides no evidence of scoped or least-privilege permission controls — no mention of scopes, roles, or restricted-access tokens for the agent.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
developerAttach a marked-up screenshot or mockup to a task so the agent implements the correct visual change
weight 2 · not comparableJulesnone0/10No evidence Jules supports attaching images, screenshots, or mockups to a task — task submission is described only via text prompts, GitHub issue labels, or API messages. Missing for 10: any mention of image/screenshot attachment, multimodal input handling, or visual mockup interpretation.
YYLOn/aYYLO is a CLI orchestrator for text-based task/workflow management around coding agents; there is no evidence of image/screenshot attachment, mockup annotation, or any visual-input handling in its task creation or ledger features. Attaching marked-up visual mockups to drive implementation is a UI/multimodal-input capability outside this CLI orchestration product's category.
ai-native userDo everything through the API that I can do in the UI
weight 2 · not comparableJules exposes a public API for creating tasks, sending messages, and enabling automation (PR creation) — core building blocks of the UI workflow — and a community user confirms dispatching tasks via the API from an external MCP client, showing real usage beyond the UI. However, there's no documented API coverage for plan approval, diff browsing/approval, or settings/API-key management, and no OpenAPI spec is discoverable (probes 404), so full UI/API parity is unconfirmed. Missing for 10: API endpoints for plan review/approval, diff/PR review parity, settings management, and a public API spec confirming full feature coverage.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
YYLOn/aYYLO is a CLI/command-line orchestrator (yy/yylo commands) with no evidence of any graphical UI to compare against; the story presumes a UI+API product with parity concerns, which doesn't fit a CLI-first tool where the CLI itself is the sole interface.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
- [probe] “PROBE runtime (recorded 2026-09-14): `npm view @yylo/cli name version bin` — the official CLI is live on the public npm registry (@yylo/cli …”
engineering-leadSee and manage plan-based daily task and concurrency limits for agent workflows
weight 2 · not comparableJules docs reference a usage-limits page with plan-tiered daily task numbers (e.g., 300/60) and community confirms daily task limits exist and change between releases (60→15 on free plan), showing plan-based limits are real and documented. However there's no evidence of an engineering-lead-facing dashboard or admin console to monitor team usage, set/adjust concurrency, or manage limits across a team — only a static limits reference page and anecdotal user experience of hitting caps. Missing for 10: team/org usage dashboard, per-user concurrency visibility, ability to configure or request limit changes, and any admin/lead-specific management UI.
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…”
YYLOn/aYYLO is a self-hosted, open-source CLI orchestrator with no evidence of a hosted plan/pricing model or subscription tiers; concepts like 'plan-based daily task and concurrency limits' apply to SaaS pricing tiers, not to a locally-run open-source tool where users control their own concurrency via config (e.g., parallel-runner). This story's axis (plan/subscription-based usage limits) does not fit this product's category.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “Bounded concurrent fan-out for independent kanban tasks, data items, or complete commands—with structured evidence for every item.”
- [probe] “PROBE runtime (recorded 2026-09-14): `curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3` returns 'MIT License /…”
ai-native userChoose where my data is stored (region/residency)
weight 2 · not comparableJulesnone0/10No evidence in the pack mentions data residency, regional storage options, or any control over where Jules stores/processes data; docs only describe VM execution and GitHub integration without residency settings.
YYLOn/aYYLO is a CLI orchestrator that runs locally on the user's own machine/repo, coordinating coding agents and git workflows—it does not store user data in a hosted service where region/residency would be a choice. Data residency is a category error for a local CLI tool rather than an unmet capability.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
ai-native userPrevent my data from being used to train AI models
weight 3 · not comparableJulesnone0/10No evidence in the pack addresses data usage/training opt-out policies, privacy controls, or any statement about whether user code/data is used to train AI models; the docs focus on functionality (VM execution, PR creation, API) rather than data governance.
YYLOn/aYYLO is a CLI orchestrator for coding agents and repo workflows, not a data-processing/AI training service; the axis of preventing data from being used for AI model training does not apply to this kind of product, and nothing in the evidence pack even implies YYLO handles user data for model training purposes.
developerTag an agent in a chat thread to discuss and delegate a bug or task
weight 2 · not comparableJules supports delegating tasks via a GitHub issue label ('jules') and via API messages to the agent, which is a form of task delegation but not the classic 'tag an agent in a chat thread to discuss' interaction pattern; one community user built a custom MCP server to dispatch Jules tasks from VS Code Copilot Chat, showing it's possible but not a native feature. missing for 10: native chat-thread tagging/mention UI, evidence of back-and-forth discussion in a thread before delegation, first-party support for Slack/Teams-style @mentions.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLOn/aYYLO is a CLI orchestrator for coding agents and repo workflows, not a chat/messaging interface where users tag agents in threads; its integrations (yylo-docs-10) pull Slack/GitHub items into a kanban board rather than supporting in-thread @-mention delegation. This is a category mismatch, not a missing feature.
developerQuery generated documentation for any public or private repository
weight 1 · not comparableJulesnone0/10Evidence describes Jules as a task-based coding agent that clones repos, edits code, and creates PRs, but there is no mention of generating or letting users query documentation for a repository (public or private). missing for 10: any docs-generation feature, a documentation query/search interface, evidence of indexing repo docs for Q&A.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
engineering-leadHave security alerts automatically validated and remediated with an opened pull request
weight 2 · not comparableJulesnone0/10Jules is a general coding agent that can be assigned tasks (via GitHub label or API) and opens PRs with diffs for review, but there is no evidence of any security-alert scanning, vulnerability detection/validation, or a workflow that automatically triggers remediation from a security alert (e.g., Dependabot/CVE integration). The evidence pack shows generic task-to-PR flow, not a security-alert-specific pipeline.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
YYLOn/aYYLO is a CLI orchestrator for coding-agent task workflows, kanban tasks, and merge/candidate management, not a security-scanning or SCA/dependency-alert tool; there is no concept of security alerts to validate. This is a category mismatch rather than an unmet capability.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [github] “Tests and semantic reviews are explicit project checks outside merge. Merge launches no models, chooses no reviewers, schedules no suites, a…”
engineering-leadCreate agent sessions on behalf of other users in my organization
weight 2 · not comparableJulesnone0/10No evidence of any organization, team, or admin/delegation features allowing an engineering-lead to create sessions on behalf of other users; Jules docs only describe individual API keys, per-user GitHub connections, and personal task creation. Missing for 10: org/team management, delegated session creation, role-based admin controls, multi-user account provisioning.
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Jules needs access to your repositories in order to work.”
YYLOn/aYYLO is a single-user CLI orchestrator for coding agents run locally; it has no organization/user-management, multi-tenant, or delegated-access model referenced anywhere in the evidence. Creating sessions 'on behalf of other users in an org' is a category mismatch for a local CLI tool rather than a missing feature.
- [github] “YYLO is a command-line orchestrator for coding agents, repeatable workflows, and receipt-backed repository changes. It is for developers who…”
- [claimed-docs] “npm install -g @yylo/cli@0.2.2 yylo init --task "Describe one verifiable outcome" --subagent claude”
developerUse a managed cloud offering to run agents without operating my own backend infrastructure
weight 2 · not comparableJules is explicitly a cloud-hosted agent service: tasks run in Google-managed VMs, integrated with GitHub, with an API and usage limits/plans, requiring no self-hosted backend. Community evidence corroborates real-world usage at scale (75% success on customer PRs, teams dispatching many tasks) though some users report reliability/UI issues. Missing for 10: no independent infra/SLA details, no discoverable OpenAPI/docs.md confirming API completeness, and mixed community reliability reports.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
YYLOn/aYYLO is a local CLI orchestrator that runs on a developer's own machine/repo (npm-installed, git worktrees, local kanban/ledger) rather than a managed cloud service; there is no evidence of a hosted runtime, cloud dashboard, or backend YYLO operates on the user's behalf. This story asks about offloading backend infra to a vendor-run cloud, which is a different product category than a CLI tool.