Jules vs HumanLayer
Jules wins · 25–20 (25 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round drawnJulesnone0/10Probes show no llms.txt, docs.md, or machine-readable API spec exist at expected paths, and no evidence Jules can be pointed at such agent-oriented documentation formats; Jules does support AGENTS.md for repo context but that's a different mechanism than consuming llms.txt-style docs.
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to HumanLayerJules provides a public API with API keys for creating custom workflows, sending tasks programmatically, and an optional automationMode to auto-create PRs, plus a community example of building an MCP server to dispatch tasks from another tool — this directly supports headless/CI-style automation. Missing for 10: no official CI/CD integration examples (e.g. GitHub Actions), no documented webhook/polling pattern for task completion, and no independent verification of reliability at scale in automated pipelines.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
Docs explicitly describe running HumanLayer headlessly via `humanlayer automation run` from CI jobs, cron machines, or scripts, plus launch tokens for non-interactive/non-PTY environments, and remote daemon hosts on cloud VMs or servers, directly matching the CI/automation story. Missing for 10: independent/hands-on confirmation of CI usage and more detail on auth/config specifics for automated pipelines.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [probe] “official CLI documented at https://docs.humanlayer.com/guide/remote-daemons”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round drawnJulesnone0/10No evidence that Jules can connect to or consume external MCP servers as tool providers; the only MCP-related evidence (jules-comm-19) describes someone building an MCP server that calls INTO the Jules API, which is the reverse integration direction, not Jules plugging in MCP servers for its own tool use.
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
HumanLayernone0/10No evidence anywhere in the pack mentions MCP servers or the ability to plug external tool servers into HumanLayer; integrations mentioned (Jira, Slack, GitHub, Linear) are task-source connectors, not MCP tool servers. Missing for 10: any mention of MCP protocol support, MCP server configuration, or tool-plugin mechanism.
ai-native userUse an official CLI
weight 2 · round to HumanLayerJulesnone0/10Evidence shows Jules offers a web app, GitHub integration, and a REST API for custom workflows, but no official CLI tool is documented or mentioned anywhere in the evidence pack.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
HumanLayer documents an official CLI (e.g. `humanlayer automation run`, launch tokens, remote-daemon control) used for CI, cron, and scripted agentic workflows, confirmed by a dedicated CLI doc page. missing for 10: no independent/hands-on verification of the CLI, no full command reference, and no evidence of broader CLI feature parity with the app.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [probe] “official CLI documented at https://docs.humanlayer.com/guide/remote-daemons”
ai-native userDrive the product through a documented public API
weight 3 · round to JulesJules ships a documented public API (developers.google.com/jules/api) with API key management, custom workflows, sending messages to the agent, and automationMode config; a community user independently built an MCP server on top of the API confirming real-world programmatic access. Missing for 10: no discoverable OpenAPI/swagger spec or llms.txt (probes returned 404s), reducing machine-readability/self-service tooling confidence.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
HumanLayer documents a CLI (`humanlayer automation run`, launch tokens, remote daemons) that lets automation environments drive sessions programmatically, which is a form of documented programmatic control, but there is no evidence of a documented public REST/OpenAPI API — probes explicitly found openapi.json/swagger.json/llms.txt all 404. missing for 10: a documented HTTP/OpenAPI public API spec, SDK/client library docs, and independent confirmation of API usage.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.humanlayer.com/openapi.json, https://docs.humanlayer.com/swagger.json, https://docs.hum…”
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.humanlayer.com/llms.txt”
- [probe] “official CLI documented at https://docs.humanlayer.com/guide/remote-daemons”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round to HumanLayerJulesnone0/10Jules documents a basic API key creation flow (max 3 keys) but provides no evidence of scoped or least-privilege permission controls — no mention of scopes, roles, or restricted-access tokens for the agent.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
The only relevant evidence is a mention of a 'launch token' scoped to a single non-interactive command, which implies some least-privilege token issuance, but there is no documentation of a broader credential/permission model, scopes, or API key management for agents. missing for 10: explicit least-privilege credential scoping model, permission granularity, revocation/rotation mechanisms, and any independent corroboration.
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
ai-native userBuild against official SDKs
weight 2 · round to JulesJules ships an official public API with documented endpoints, API key management, and use cases for building custom workflows/integrations, and a community member confirms building a personal MCP server against the Jules API. missing for 10: no official OpenAPI/SDK spec discoverable (openapi probes 404), no first-party language SDKs mentioned, and no independent SDK-quality corroboration beyond one community integration example.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
HumanLayernone0/10The evidence pack covers HumanLayer's CLI, workspace config, and third-party integrations (Jira, Slack, GitHub, Linear) but contains no mention of an official SDK (Python, TypeScript, etc.) for building against HumanLayer programmatically. Probes for openapi.json and llms.txt both returned 404, further suggesting no discoverable API/SDK surface.
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.humanlayer.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.humanlayer.com/openapi.json, https://docs.humanlayer.com/swagger.json, https://docs.hum…”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
ai-native userSubscribe to events via webhooks
weight 2 · round drawnJulesnone0/10Evidence shows Jules has an API for creating tasks/sending messages, notifications for task completion, and GitHub integration, but there is no mention of webhooks or event subscription mechanisms anywhere in the docs or community evidence. Missing for 10: any documentation of a webhook endpoint, event subscription API, or push-based notification mechanism to external systems.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
HumanLayernone0/10No evidence pack item mentions webhooks or event subscription mechanisms; integrations described (Slack, Jira, Linear, GitHub) are inbound task-creation connectors, not outbound webhook events, and API/OpenAPI probes returned 404s.
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to JulesJules generates AI-derived plans, diffs, and PR suggestions based on analysis of the user's repository data, which is the coding-agent analog of 'AI-generated insights/suggestions from your data.' Community evidence is mixed on the quality of these suggestions, with some praising usable PRs and others criticizing low-quality output on complex codebases. Missing for 10: no evidence of broader analytics-style insights beyond code-change suggestions, and no independent benchmarking confirming insight quality across use cases.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
HumanLayer's agents do generate task artifacts, draft PRs, and comments derived from a user's codebase/tickets, and 'Advanced Metrics' surfaces usage/cost/productivity data, which loosely resembles data-derived insight. However there is no evidence of dedicated AI-generated analytical insights or proactive suggestions distinct from executing assigned coding tasks. Missing for 10: explicit insight/analytics generation from user data, proactive suggestion features beyond task execution, and any first-party or community evidence of an 'insights' capability.
- [claimed-docs] “Advanced Metrics for all paid plans — View usage, cost, and productivity metrics, with optional access for every organization member.”
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to HumanLayerJules supports background/autonomous execution via GitHub label-triggered tasks and a public API for building custom automations (e.g., automated bug-fixing, code review) with an optional automationMode for auto-PR creation, and it runs tasks in a cloud VM without needing the user present. However, docs also show a human-approval gate (plan approval, PR review) rather than fully hands-off automation, and community reports note tasks getting stuck, hitting limits, or needing babysitting, undercutting reliability of unattended runs. Missing for 10: evidence of true scheduled/cron-style recurring automations, and independent confirmation that automations run to completion without manual intervention.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
Docs explicitly describe `humanlayer automation run` for running Cloud-visible coding sessions from CI jobs, cron machines, or scripts, plus remote daemons and launch tokens for non-interactive/background execution contexts, directly supporting autonomous background automations. Missing for 10: independent/hands-on verification of long-running background automations, native scheduling UI, and clarity on how human-approval gates interact with continuous autonomous runs.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use `.humanlayer/workspace.json` for shared repository or team configuration. Use `.humanlayer/workspace.local.json` for optional user or ma…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to JulesJules is exactly this kind of built-in agentic assistant: users delegate coding tasks via GitHub issue labels, the web app, or API, and Jules autonomously plans, clones repos, runs in a VM, and returns diffs/PRs for approval (jules-docs-2,3,4,7,8,9). Community evidence corroborates real-world task delegation and usable output (jules-comm-2, jules-comm-11, jules-comm-15, jules-comm-16), though mixed reports of loops, babysitting needs, and reliability issues (jules-comm-14, jules-comm-18) temper quality. Missing for 10: independent benchmarking of task success rates and evidence of consistent reliability at scale.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
HumanLayer's core product model is task delegation to AI coding agents (Claude Code, Codex, Bedrock-backed Claude) via 'sessions', with workflow modes (Oneshot, RPI, PRD-Oriented, Freeform) and automation runs from CI/cron, directly matching 'delegate tasks to a built-in AI assistant'. missing for 10: independent/hands-on verification of the delegation experience beyond vendor docs, and clarity on how autonomous vs supervised the assistant is in practice.
- [claimed-docs] “This tutorial teaches you how to create and run one small task in HumanLayer. You will use the macOS app and Claude Code.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “This guide shows you how to install, authenticate, select, and check Codex for HumanLayer sessions.”
ai-native userOperate the product with natural-language commands
weight 2 · round to JulesJules is designed around natural-language task submission and messaging: users submit a task description, Jules generates a plan, and you can send follow-up messages to the agent via API or web app (jules-docs-4, jules-docs-12). Community reports confirm real-world use of NL prompts to drive coding tasks, including iterative feedback and error-message follow-ups (jules-comm-2, jules-comm-15, jules-comm-16). Missing for 10: no detailed documentation of the full range/complexity of natural-language commands supported (e.g., multi-step conversational control, command reference) and no independent benchmark of NL command robustness beyond anecdotal reports.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “To send a message to the agent:”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
HumanLayer's core interaction model is giving natural-language instructions to agent sessions (Claude Code, Codex) to create tasks, configure workspaces, and choose workflow phases, as shown by the example NL workspace-config prompt and workflow-selection docs. Missing for 10: independent/hands-on corroboration of NL command robustness and no evidence of a broader NL command surface beyond task/workflow setup.
- [claimed-docs] “Configure a multi-repository workspace with this repository, ../api, and ../web. Make ../web the primary repository. Ask me before you choos…”
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
- [claimed-docs] “This tutorial teaches you how to create and run one small task in HumanLayer. You will use the macOS app and Claude Code.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round drawnJulesnone0/10Jules documents an API (create tasks, send messages, API keys) but there is no evidence of an interactive API reference or runnable code examples; probes for openapi/swagger specs all returned 404s, suggesting no such interactive reference exists.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “To send a message to the agent:”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
HumanLayernone0/10No evidence of an interactive API reference or runnable examples; probes explicitly show no OpenAPI/swagger spec and no llms.txt found, and docs are guide/tutorial style rather than an API reference sandbox.
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnJulesnone0/10Jules ships a documented REST API (jules-docs-9..12) but there is no evidence of a downloadable OpenAPI/Swagger spec; explicit probes for llms.txt, docs.md, and standard OpenAPI paths all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), confirming no machine-readable spec is published.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
HumanLayernone0/10A direct probe for OpenAPI/swagger specs at all standard locations returned 404s, and no evidence pack item shows a downloadable machine-readable API spec being offered.
ai-native userTest against a sandbox environment without touching production data
weight 1 · round to JulesJules runs tasks in an isolated cloud VM that clones the repo, meaning code changes happen in a sandboxed environment rather than directly on production, and PRs must be reviewed/approved before merging to the real branch. However, this is a code-sandbox for making changes, not a dedicated 'test against sandbox data/environment' feature, and there's no evidence of test-data isolation, staging environment provisioning, or explicit protection of production data/services beyond the VM/PR review flow. missing for 10: explicit sandbox test-data isolation, staging/production separation guarantees, independent verification that production systems are never touched.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
HumanLayernone0/10HumanLayer's docs describe remote daemons, workspaces, and automation sessions, but there is no mention of a sandbox environment, staging/test data isolation, or any mechanism to separate test runs from production data. missing for 10: explicit sandbox/staging environment documentation, data isolation guarantees, evidence of test-vs-production separation.
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnJulesnone0/10Jules has a public API (jules-docs-9, jules-docs-10) but there is no evidence of API versioning scheme or a documented deprecation policy; probes for OpenAPI spec/docs.md/llms.txt all returned 404 (jules-probe-1, jules-probe-2, jules-probe-3), suggesting no formal machine-readable API contract or lifecycle documentation is available.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
HumanLayernone0/10No evidence of API versioning scheme or a documented deprecation policy; probes for openapi.json/llms.txt returned 404s and no API reference or changelog covering versioning/deprecation is present.
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to JulesJules exposes an API that lets users script custom workflows and dispatch tasks programmatically (jules-docs-9, jules-docs-12), and usage-limit tiers explicitly target 'power users & agent-heavy workflows' with up to 300 tasks (jules-docs-13), while community reports confirm it 'handles large quantities of tasks very well' and users built custom dispatch tooling via the API/MCP bridge (jules-comm-11, jules-comm-19). However, there's no documented native UI or endpoint for a single bulk action across many items (e.g., batch-applying one task to many repos/issues) — each task still appears to be created and reviewed individually via GitHub labels or API calls. Missing for 10: a documented batch/bulk endpoint or UI feature, first-party bulk-operation examples, and independent verification of true parallel bulk execution rather than just high per-day task volume.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
HumanLayernone0/10HumanLayer's documentation consistently frames work as single tasks/sessions ('one small task', 'one task on a remote machine', 'one Cloud-visible coding session') with per-task review and approval workflows; there is no mention of batch/bulk operations spanning many items at once. missing for 10: any documented bulk-action API/CLI flag, batch approval mechanism, or multi-item automation workflow.
- [claimed-docs] “This tutorial teaches you how to create and run one small task in HumanLayer. You will use the macOS app and Claude Code.”
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to HumanLayerJules supports some event-driven automation — e.g., assigning a 'jules' label to a GitHub issue automatically triggers a task, and the API lets developers build custom workflows/automations around Jules (jules-docs-2, jules-docs-9). However, there's no evidence of a general rules/conditions engine where users define arbitrary triggers and actions beyond this specific label mechanism and API calls. Missing for 10: a configurable rules interface (conditions + multiple trigger types), documentation of scheduled/webhook-based triggers, and independent confirmation that automated rule-based triggering works reliably in practice.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
HumanLayer integrations (GitHub, Jira, Linear) create tasks automatically from external events like new issues/tickets, and `humanlayer automation run` lets sessions be triggered from CI jobs, cron, or scripts — both are forms of event-driven automation. However, there's no evidence of a general-purpose rules/conditions engine letting users define arbitrary trigger-condition-action logic; the automation is limited to fixed integration hooks and script-based invocation. Missing for 10: a documented rule-definition interface (conditions, filters, custom triggers) beyond fixed platform integrations, and independent confirmation these event-triggers work reliably in practice.
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to HumanLayerJulesnone0/10Evidence covers task assignment via GitHub labels, API-driven custom workflows, and one-off task automation, but nothing describes recurring/scheduled jobs (e.g., cron-like triggers) — the API docs only mention creating tasks and sending messages, not recurrence.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
HumanLayer's `automation run` CLI lets you trigger a Cloud-visible coding session from a cron machine or CI job, implying recurring runs are possible via external schedulers, but there is no documented native scheduling/recurrence feature (no cron syntax, interval config, or job queue) inside HumanLayer itself. missing for 10: built-in recurring scheduler, interval/cron configuration options, evidence of persistent recurring workflow management.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
ai-native userVersion, review, and roll back my automations
weight 1 · round to HumanLayerJulesnone0/10Jules produces PRs and diffs via GitHub but there is no evidence of a mechanism to version, review history of, or roll back automations/tasks themselves (as opposed to code changes tracked by git/GitHub). missing for 10: automation versioning/history feature, rollback/undo of Jules tasks, changelog or audit trail for automations.
HumanLayer's task model provides review (comments, PR draft creation) and history that persists across sessions, giving some review/audit capability, but there is no explicit documentation of versioning workflow definitions or rolling back an automation to a prior version. missing for 10: explicit automation versioning/history diffing, a documented rollback mechanism, and independent confirmation these review features extend to full automation lifecycle management.
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation
End-to-end implementation by the agent — multi-file changes, task completion
End to end feature delivery
ai-native userHave an agent automatically generate and run tests to validate its own code changes before proposing them
weight 2 · round to JulesJules runs code changes in an isolated VM and generates diffs/PRs for review, and one community report notes it is 'good at making tests for very specific/isolated functions,' implying some test generation capability, but there is no first-party documentation describing an explicit test-generate-and-run validation loop before proposing changes. missing for 10: official docs on automated test generation/execution as a validation step, evidence of test results being surfaced to users before PR creation, and independent confirmation across varied codebases.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
developerHave an agent autonomously diagnose and fix a reported bug
weight 3 · round to JulesJules is designed to autonomously diagnose and fix issues: it clones the repo into a VM, generates a plan, makes code changes, and opens a PR for review, and can be triggered directly from a GitHub issue via the 'jules' label (a common bug-report workflow). Community evidence corroborates real-world bug-fixing use, including a case where it failed then successfully fixed the issue after being given the error message. Missing for 10: independent benchmarking on bug-fix success rate and more consistent evidence across complex codebases (some reports of failures/loops on harder tasks).
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
HumanLayer supports creating tasks directly from GitHub/Jira/Linear issues and then running an agent session (Claude Code/Codex) against the linked repo, with an 'Oneshot' workflow phase designed for small, clear changes — a plausible bug-fix pipeline. However, the product's core design is human-in-the-loop with approval gates rather than fully autonomous action, and there's no end-to-end documented example of an agent independently diagnosing a bug from a ticket and shipping a fix without human review. Missing for 10: a concrete autonomous bug-diagnosis-and-fix walkthrough, and clarity on how much human approval is required mid-flow.
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [community] “"I feel much more comfortable with senior developers approving AI changes to our codebase, then letting loose an autonomous agent with no hu…”
product-managerGo from a mockup or design to a working implementation without an engineering handoff
weight 2 · round drawnJulesnone0/10Evidence shows Jules operates on code repositories via GitHub issues/tasks, plans, diffs, and PRs, but there is no mention of accepting mockups/designs as input or any PM-oriented no-code workflow — usage still requires repo access, issue creation, and reviewing technical diffs/plans. Nothing in the pack demonstrates a design-to-implementation path bypassing engineering.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
HumanLayernone0/10HumanLayer's evidence is entirely about developer-facing workflows: running coding sessions, connecting Jira/Linear/GitHub/Slack, managing remote daemons, and CLI automation for engineers overseeing coding agents. Nothing in the pack shows a mockup/design import capability, a no-code interface, or any path for a non-engineer product manager to turn a design into a working implementation without engineering involvement — in fact the workflow language (RPI, PRD-oriented, Oneshot) and workspace configs assume an engineering operator. Missing for evidence of delivering this story: mockup/design ingestion, PM-oriented no-code UI, and any case study of a non-engineer shipping code end-to-end.
- [claimed-docs] “This tutorial teaches you how to create and run one small task in HumanLayer. You will use the macOS app and Claude Code.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
developerHave an agent implement a requested feature end-to-end, including writing tests
weight 3 · round to JulesJules is explicitly built for autonomous end-to-end coding: it clones the repo, generates a plan, modifies files, produces a diff, and opens a PR (jules-docs-3/4/7/8), and one user specifically notes it's 'good at making tests for very specific/isolated functions' (jules-comm-11), with others confirming usable PRs and time savings (jules-comm-6, jules-comm-16). However, community evidence shows significant caveats: it struggles on complex/monorepo codebases, gets stuck in loops without a stop button, needs heavy babysitting, and often fails to complete tasks unassisted (jules-comm-3, jules-comm-9, jules-comm-10, jules-comm-13, jules-comm-14, jules-comm-18). Missing for 10: consistent reliability across complex real-world features, broader evidence of comprehensive test coverage (not just isolated functions), and independent benchmarks confirming end-to-end success rate.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “I've been playing with it, and I've been generally not impressed. There are obvious annoying UI bugs and the output isn't very good for anyt…”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
HumanLayer clearly supports end-to-end autonomous coding sessions (Oneshot, RPI, PRD-Oriented workflows) that implement tasks using agents like Claude Code and Codex, including structured phases and sub-agents for research/plan/implement, which implies substantial feature work can be delegated (humanlayer-docs-6, humanlayer-docs-9, humanlayer-docs-10, humanlayer-docs-14). However, no evidence explicitly confirms the agent writes or runs tests as part of the workflow, and no hands-on account demonstrates a full feature-plus-tests delivery. Missing for 10: explicit documentation or example showing test generation/execution as part of the implementation flow, and independent verification of end-to-end feature completion including tests.
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
- [claimed-docs] “This guide shows you how to install, authenticate, select, and check Codex for HumanLayer sessions.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
Environment setup
developerHave an agent automatically clone the repo, install dependencies, and configure its own working environment
weight 2 · round to JulesJules' core workflow is documented: it runs in a cloud VM, clones the repo, installs dependencies, and configures its environment automatically before making changes, corroborated by community usage reports of automated PR generation. missing for 10: independent technical verification of dependency-install robustness across complex/monorepo setups, and some community reports note environment/config confusion in bespoke or monorepo projects.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
Docs describe workspace configuration (workspace.json, multi-repo setups) and remote hosts that must have access to code/tools/credentials, and one example prompt asks the agent about 'setup commands or local files to copy,' implying some environment configuration ability. However, there is no explicit description of the agent autonomously cloning a repo or installing dependencies end-to-end without human setup of the host/workspace first. Missing for 10: explicit documentation of automatic repo cloning, dependency installation steps, and end-to-end environment bootstrap without prior manual host/workspace configuration.
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use `.humanlayer/workspace.json` for shared repository or team configuration. Use `.humanlayer/workspace.local.json` for optional user or ma…”
- [claimed-docs] “Configure a multi-repository workspace with this repository, ../api, and ../web. Make ../web the primary repository. Ask me before you choos…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
Interactive takeover
developerTake over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
weight 2 · round to HumanLayerJules supports reviewing plans, diffs, and sending follow-up messages to redirect a task (jules-docs-4, jules-docs-6, jules-docs-7, jules-docs-12), and third parties have wired the API into VS Code/Copilot Chat to dispatch tasks (jules-comm-19), but this is dispatch, not mid-task takeover. A hands-on report explicitly notes there's no 'STOP' button to interrupt a looping task (jules-comm-14), undercutting redirect control, and all takeover is browser/API-based with no native terminal or editor 'take over' UI documented. Missing for 10: documented in-editor/terminal takeover UI, ability to pause/interrupt an in-progress run, and evidence the API-based messaging genuinely redirects rather than just appends instructions.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
Docs describe tasks with session/history continuity 'across agents and workstations', remote control from app.humanlayer.com on any machine or phone, and CLI-driven remote daemons for terminal/server contexts, all pointing to genuine hand-off of in-progress work between editor (Claude Code), terminal (remote daemon/CLI), and browser (web app). Live multiplayer draft recovery further shows shared/continuable session state. missing for 10: independent/hands-on confirmation of a live takeover mid-task, and explicit description of an in-editor (IDE plugin) takeover UI rather than just CLI/app/web.
- [claimed-docs] “This tutorial teaches you how to create and run one small task in HumanLayer. You will use the macOS app and Claude Code.”
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “Live multiplayer response drafts — Write prompts together with shared text, named cursors, presence, read-only viewing, and draft recovery.”
developerSend follow-up instructions to an active agent session to steer its work without restarting
weight 2 · round to JulesJules's API docs explicitly describe sending a message to an active agent session (jules-docs-12), and community reports confirm the workflow of reviewing partial work and providing feedback for Jules to continue rather than restarting (jules-comm-3). This shows the steering capability works in practice, though documentation on how follow-ups affect an in-progress plan/execution is thin. Missing for 10: detailed first-party documentation of mid-task message handling/UI chat thread, and independent hands-on confirmation of steering effectiveness beyond one anecdote.
- [claimed-docs] “To send a message to the agent:”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
HumanLayer's task/session model implies ongoing interaction with agents across a task's lifetime (docs-5) and release notes mention live prompt collaboration during sessions (docs-18), suggesting some capacity to interact with an active session, but there is no explicit documentation of sending mid-session follow-up instructions to steer a running agent without restarting it. missing for 10: explicit docs on injecting new instructions into a live/running session, confirmation the agent incorporates such input without restart, and independent/hands-on verification of this steering behavior.
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “Live multiplayer response drafts — Write prompts together with shared text, named cursors, presence, read-only viewing, and draft recovery.”
Sandbox execution
developerHave an agent safely execute code and install dependencies inside an isolated sandbox
weight 3 · round to JulesJules explicitly runs in a cloud VM that clones code and installs dependencies, isolated from the user's local machine, with a plan-review step before code changes are made — directly matching the sandboxed execution story. Community reports corroborate it actually running tasks end-to-end (producing PRs) in this VM environment, though some criticize reliability/loops on complex codebases. Missing for 10: detailed docs on sandbox security boundaries (network isolation, resource limits) and independent security audit of the VM isolation.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
HumanLayernone0/10HumanLayer docs describe running sessions on remote hosts (cloud VM, workstation, private-network machine) and automation environments, but there is no mention of an isolated/sandboxed execution environment for running code or installing dependencies safely — the host selection is about access/credentials, not isolation guarantees. Missing for full/partial: any explicit sandbox, container, or isolation mechanism; no evidence of dependency-install safety controls.
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight
Keeping a human in the loop — approvals, checkpoints, interrupts
Approval controls
developerConfigure an agent to auto-approve all its actions instead of confirming each one
weight 2 · round drawnJules's default workflow requires manual approval at multiple stages (plan approval, diff review, PR approval), but the API docs mention an optional `automationMode` field that changes default behavior around automatic PR creation, suggesting some configurable auto-approve path exists via API rather than the standard UI flow. Missing for 10: explicit documentation of a setting that suppresses plan-approval and diff-review confirmations entirely, and any hands-on/community confirmation that this automation mode actually skips human checkpoints.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
HumanLayer's core premise is human-in-the-loop approval, and docs mention workflow phases like 'Oneshot' for low-risk changes and automation sessions (humanlayer automation run) that run non-interactively without confirmation, implying some auto-approve capability exists, but no explicit documentation of a configurable 'auto-approve all actions' toggle or setting is shown. missing for 10: explicit config/flag to disable per-action confirmation entirely, documentation confirming automation sessions skip all human review rather than just running unattended, and independent confirmation this works as intended.
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
product-managerApprove key agent decisions from my phone while agents continue working
weight 1 · round to HumanLayerJules is a cloud-based agent with a web app and notification system, plan approval, diff review, and PR approval steps, and community evidence confirms users review/approve PRs from a browser including 'coding on my phone' via mobile web. However, there is no dedicated native mobile app or explicit mobile-optimized approval UI, and no evidence of push-notification-driven approval flows tailored to on-the-go PM decision-making. missing for 10: dedicated mobile app or mobile-specific UI, evidence of push notifications enabling quick phone-based approvals, PM-specific (non-developer) approval workflow.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
Docs explicitly describe controlling and continuing a running agent task from a phone via app.humanlayer.com, with tasks providing a shared review/comment history across devices, directly matching the phone-approval-while-agent-continues story; community sentiment corroborates the human-approval-of-agent-actions use case. Missing for 10: a hands-on/independent account specifically confirming the mobile approval UI in practice, and explicit documentation of an 'approve/deny decision' action (vs. general task control) on mobile.
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [community] “"I feel much more comfortable with senior developers approving AI changes to our codebase, then letting loose an autonomous agent with no hu…”
engineering-leadSet tiered autonomy levels controlling what an agent can do without manual confirmation
weight 3 · round to HumanLayerJules has a binary plan-approval workflow (review/approve plan, review diffs, approve PRs) and an API `automationMode` field that toggles whether PRs are auto-created, but there's no evidence of configurable tiered autonomy levels (e.g., low/medium/high trust settings, granular permission scopes, or per-action confirmation thresholds) that an engineering-lead could set. Missing for 10: explicit multi-tier autonomy/permission settings, admin-configurable trust levels, and any org-wide policy controls beyond the single automationMode toggle and default plan-approval gate.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
HumanLayer's workflow-phases doc shows tiered approaches (Oneshot for low-risk changes vs. more review for unclear/risky changes) and skills-workflows lets users select Oneshot/RPI/PRD-Oriented/Freeform modes, which map to different levels of autonomy vs. oversight. However, there's no explicit documentation of a formal 'autonomy level' setting per agent/task with configurable confirmation thresholds, and no independent evidence confirming this tiered control works as an oversight mechanism in practice. missing for 10: explicit named autonomy-tier configuration (e.g., low/medium/high) tied to confirmation gating, evidence of engineering-lead-level policy controls across a team, and independent/hands-on validation that these workflow phases actually reduce unnecessary confirmations without sacrificing safety.
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
Model control
ai-native userHave each task prompt automatically routed to the most suitable underlying model
weight 2 · round drawnJulesnone0/10No evidence Jules performs automatic model routing per task; docs mention only a single flagship model (Gemini 3 Pro) as 'priority access,' with no mention of routing prompts to different models based on suitability. Missing for 10: any documentation of multi-model routing logic, model-selection criteria, or automatic switching between models.
- [claimed-docs] “Priority access to the latest models, starting with Gemini 3 Pro”
HumanLayernone0/10Evidence shows HumanLayer lets users manually select or configure which model/backend to use (Claude via Bedrock, Codex, RPI sub-agents) but there is no evidence of automatic routing of a task prompt to the 'most suitable' model based on task characteristics.
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
- [claimed-docs] “This guide shows you how to install, authenticate, select, and check Codex for HumanLayer sessions.”
- [claimed-docs] “HumanLayer sessions can run Claude through Amazon Bedrock instead of the Anthropic API.”
engineering-leadSwitch away from automatic model selection to a specific model of my choice
weight 1 · round to HumanLayerJulesnone0/10No evidence pack item mentions model selection settings or the ability to choose a specific underlying model; the only related mention is 'Priority access to the latest models, starting with Gemini 3 Pro' which describes access tiers, not user-controlled model switching.
Docs show explicit model/backend selection — choosing Codex ("install, authenticate, select, and check Codex") or running Claude via Amazon Bedrock instead of the Anthropic API, plus subagent model registration — indicating an engineering lead can pick a specific model rather than a default. However, there is no explicit documentation of an 'automatic' default-selection mode being toggled off, so the framing of 'switching away from automatic' isn't directly evidenced. Missing for 10: explicit mention of an automatic/default model-selection setting and a documented UI/CLI flag to override it, independent confirmation of model-switching behavior.
- [claimed-docs] “This guide shows you how to install, authenticate, select, and check Codex for HumanLayer sessions.”
- [claimed-docs] “HumanLayer sessions can run Claude through Amazon Bedrock instead of the Anthropic API.”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
Visibility monitoring
developerWatch what a running agent is doing in real time, including its current status
weight 3 · round to HumanLayerJules provides notifications on task completion/input-needed and a plan-approval step, implying some status visibility, plus a diff/PR review flow, but there is no evidence of a live real-time activity feed or streaming log of the agent's current actions while running; community feedback even notes the absence of a stop/interrupt control mid-run, suggesting limited real-time observability. Missing for 10: evidence of a real-time execution log/console view, granular step-by-step status updates, and independent confirmation of live monitoring UX.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
HumanLayer's docs describe remote-daemon control from app.humanlayer.com (including from a phone), live multiplayer session viewing with presence/read-only viewing, and task/session history — all suggesting real-time visibility into agent activity. However, there's no explicit documentation of a dedicated live status/log stream or dashboard showing granular agent state (e.g., current tool call, progress bar) beyond session/task views. missing for 10: explicit real-time status/log streaming documentation, independent hands-on confirmation of live monitoring UX.
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “Live multiplayer response drafts — Write prompts together with shared text, named cursors, presence, read-only viewing, and draft recovery.”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
developerGet notified when an agent completes a task or needs my input
weight 2 · round to JulesFirst-party docs explicitly state users are notified when a task completes or needs input, and the workflow (plan approval, diff review, PR creation) reinforces oversight checkpoints where notifications matter. Community evidence corroborates the async, check-back-later usage pattern (e.g. reviewing PRs after time away), consistent with notification-driven workflows. Missing for 10: independent verification of notification channels (email/push/Slack), and no detail on notification reliability or configurability.
- [claimed-docs] “You’ll be notified when the task completes or needs your input.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
- [community] “Jules was unable to complete the task in time. Please review the work done so far and provide feedback for Jules to continue.”
HumanLayer supports Slack/GitHub/Jira/Linear integrations that push task artifact updates and human-in-the-loop approvals, implying notification when tasks progress or need input, and its core design centers on human oversight of agent work. However, there is no explicit documentation of a dedicated 'task complete' or 'needs input' notification/alert mechanism (e.g., push notification, email, or webhook triggered specifically on completion/approval-required events) beyond generic artifact updates in Slack. missing for 10: explicit completion/needs-input notification docs, evidence of notification latency/reliability, independent/hands-on confirmation of notification delivery.
- [claimed-docs] “Connect Slack to send HumanLayer task artifact updates into the channels where your team already works.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [community] “"I feel much more comfortable with senior developers approving AI changes to our codebase, then letting loose an autonomous agent with no hu…”
Intent to spec — stories about intent to spec in this arenaIntent to spec
Stories about intent to spec in this arena
Natural language task intake
developerDescribe a feature or bug in plain language and have it automatically turned into a scoped implementation task
weight 3 · round drawnJules docs clearly show a plain-language task submission flow that generates a reviewable/approvable plan before code changes, then produces a diff and PR (jules-docs-4, jules-docs-7, jules-docs-8), directly matching the intent-to-spec story. However community reports show inconsistent scoping quality — confusion in monorepos, endless loops, and needing heavy babysitting — indicating the automatic scoping doesn't always hold up in practice. Missing for 10: independent verification of plan/spec quality on complex codebases, and consistent evidence the generated plan reliably matches developer intent without back-and-forth correction.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've tried using Jules for a side project, and the code quality it emits is much worse than GH Copilot, Gemini CLI, and Claude Code. It also…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
HumanLayer's task model (docs-5) and workflow-phase selection (docs-6, docs-10) show that a task is created and can be routed through 'Oneshot' for small clear changes or heavier RPI/PRD-oriented flows for ambiguous work, and RPI sub-agents (docs-9) imply a research→plan→implement pipeline that turns a description into a scoped plan. However, there is no explicit walkthrough showing a raw plain-language bug/feature description being automatically parsed into a scoped implementation task end-to-end, and integrations (Jira/Linear/GitHub) mostly create tasks from existing tickets rather than free-form language input. Missing for 10: a concrete example or tutorial of plain-language-to-scoped-task conversion, and independent/hands-on confirmation that this pipeline works as described.
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
product-managerConvert user feedback submissions into structured tasks with proposed scope
weight 2 · round to HumanLayerJulesnone0/10Jules's documented workflow is a coding-agent pipeline (clone repo, generate a plan, produce diffs/PRs) driven by developer-specified tasks or GitHub issues, not a mechanism for ingesting unstructured user-feedback and converting it into a structured task with proposed scope for a PM. No evidence shows feedback intake, requirement structuring, or scope proposal features aimed at product managers.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
HumanLayer supports creating tasks from external issue trackers (Jira, Linear, GitHub) which could serve as a proxy for user feedback submissions, and tasks include shared files/scope info, but there's no evidence of a dedicated feature for ingesting raw user feedback (e.g., support tickets, survey responses) and auto-structuring it into a task with a proposed scope specifically tailored for PM workflows. missing for 10: dedicated feedback-ingestion mechanism, explicit 'proposed scope' generation from unstructured feedback text, PM-specific workflow templates, and any hands-on/community validation of this specific use case.
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
developerAttach a marked-up screenshot or mockup to a task so the agent implements the correct visual change
weight 2 · round to HumanLayerJulesnone0/10No evidence Jules supports attaching images, screenshots, or mockups to a task — task submission is described only via text prompts, GitHub issue labels, or API messages. Missing for 10: any mention of image/screenshot attachment, multimodal input handling, or visual mockup interpretation.
Docs confirm images can be pasted into the new task composer as attachments (humanlayer-docs-22), which supports attaching a screenshot to a task, but there is no evidence of markup/annotation tooling or of the agent parsing visual annotations to implement a corresponding UI change. Missing for 10: annotation/markup capability for screenshots, evidence the agent interprets visual markup into a specific implementation, and any hands-on example of this workflow succeeding.
- [claimed-docs] “Image paste in new tasks — Paste images straight into the new task composer as attachments.”
Plan approval
developerReview and approve an agent's implementation plan before any code changes are made
weight 3 · round to JulesJules docs explicitly state the agent generates a plan upon task submission that users can review and approve before any code changes are made, directly matching the story. No independent/hands-on evidence specifically confirms or contradicts this plan-approval step, though other workflow claims (diffs, PRs) are corroborated by community mentions. Missing for 10: independent/hands-on confirmation of the plan-review step specifically, and more detail on what the plan interface looks like.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
HumanLayer's workflow-phases and RPI sub-agent docs describe planning phases (e.g., 'a change with unclear behavior or code shape needs more review before implementation') and PRD-Oriented/RPI workflows imply a plan stage before code changes, with tasks providing 'one place for comments and review.' However, no evidence explicitly shows a dedicated plan-approval gate/UI step where a developer reviews and approves a plan artifact before implementation begins. missing for 10: explicit documentation of a plan-approval step/UI, first-party example of blocking implementation until plan is approved, independent/hands-on confirmation of this specific gate.
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
engineering-leadApprove a task's scope and contract before an agent is allowed to modify the repository
weight 2 · round drawnJules generates a plan that the user can review and approve before any code changes are made, giving an engineering-lead a gate before modification, and it also produces diffs/PRs requiring approval before merge. However, this is a general 'plan approval' UX rather than a formal scope/contract sign-off workflow tailored for engineering-lead governance (e.g., no evidence of role-based approval gates, policy enforcement, or blocking unauthorized starts). missing for 10: explicit lead/role-based approval gating before task execution starts, evidence of enforceable contract/scope definitions beyond a plan preview, independent confirmation that the plan-approval step reliably blocks modification.
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
HumanLayer's workflow-phases doc explicitly supports scoping review before implementation (e.g., 'a change with unclear behavior or code shape needs more review before implementation'), and tasks/sessions provide a structured place for comments and review prior to agent execution, plus human-in-the-loop approval is core to the product's value prop per community discussion. However, there's no explicit documented feature for an engineering-lead specifically approving a 'scope and contract' artifact as a gating step before repo modification — it's inferred from general workflow-phase and review mechanics rather than a dedicated scope-approval gate. missing for 10: explicit documentation of a formal scope/contract approval step tied to lead sign-off, evidence of blocking repo writes until such approval, and independent/hands-on confirmation this gate works as intended.
- [claimed-docs] “Use the shortest path that controls the main risk. A clear and small change can use Oneshot. A change with unclear behavior or code shape ne…”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “Select **Oneshot**, **RPI**, **PRD-Oriented**, or **Freeform**.”
- [community] “"I feel much more comfortable with senior developers approving AI changes to our codebase, then letting loose an autonomous agent with no hu…”
Ticket driven tasking
developerAssign a coding task to an agent directly from an existing issue or ticket
weight 3 · round drawnJules explicitly supports assigning tasks directly from GitHub issues via a 'jules' label, and API/community evidence corroborates task dispatch and PR creation workflows. Missing for 10: evidence of ticket-tracker integrations beyond GitHub (e.g., Jira/Linear) and independent hands-on confirmation of the label-to-task flow itself.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
Docs explicitly describe connecting GitHub, Jira, and Linear so HumanLayer creates tasks directly from issues/tickets and links artifacts back to the source, directly matching the story of assigning agent work from an existing ticket. Missing for 10: independent/hands-on confirmation that this ticket-to-task flow works reliably in practice, and more detail on the actual assignment UX.
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to JulesJules exposes a public API for creating tasks, sending messages, and enabling automation (PR creation) — core building blocks of the UI workflow — and a community user confirms dispatching tasks via the API from an external MCP client, showing real usage beyond the UI. However, there's no documented API coverage for plan approval, diff browsing/approval, or settings/API-key management, and no OpenAPI spec is discoverable (probes 404), so full UI/API parity is unconfirmed. Missing for 10: API endpoints for plan review/approval, diff/PR review parity, settings management, and a public API spec confirming full feature coverage.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
HumanLayer offers a CLI (`humanlayer automation run`) and remote daemon controls that let you launch and manage sessions outside the UI, but there is no documented public API/OpenAPI spec (both openapi.json and llms.txt probes 404), and UI-only features like live multiplayer drafts, keyboard navigation, and image paste have no CLI/API equivalent documented. missing for 10: a documented REST/GraphQL API or OpenAPI spec, confirmation that all UI actions (draft PRs, multiplayer editing, metrics views) are exposed programmatically, and independent verification of API-UI parity.
- [probe] “official CLI documented at https://docs.humanlayer.com/guide/remote-daemons”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.humanlayer.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.humanlayer.com/openapi.json, https://docs.humanlayer.com/swagger.json, https://docs.hum…”
- [claimed-docs] “Live multiplayer response drafts — Write prompts together with shared text, named cursors, presence, read-only viewing, and draft recovery.”
- [claimed-docs] “Keyboard navigation for changed files — Move through the PR changes tree with J/K, N/P, G shortcuts, and Enter.”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnJulesnone0/10No evidence of any data export feature, open-format export, or account/data portability mechanism for Jules; the product works via GitHub repos/PRs but nothing indicates users can export Jules-specific data (task history, configs, etc.) and leave. missing for 10: any documentation of data export, open format support, or account portability/deletion workflow.
ai-native userRead the product's source under an open license
weight 2 · round drawnJulesnone0/10Jules is a closed, proprietary Google product (cloud VM, API keys, usage limits); no evidence anywhere of source code being published or licensed openly, and probes for docs/openapi artifacts return 404s, further suggesting no open publishing.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [probe] “PROBE llms.txt: HTTP 404 at https://jules.google/llms.txt”
- [probe] “PROBE docs-md: HTTP 404 at https://jules.google/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://jules.google/openapi.json, https://jules.google/swagger.json, https://jules.google/api/opena…”
ai-native userSelf-host the core product
weight 3 · round drawnJulesnone0/10Jules is explicitly a cloud-based VM service (Google-hosted) with no evidence of any self-hosting option; community even notes cloud-only design as a drawback compared to local tools like Claude Code.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
- [community] “It's a shame Google picked the wrong system design for Jules. Claude Code's system design is clearly superior at this point.”
HumanLayernone0/10Evidence shows HumanLayer's daemon/agent execution can run on a user-controlled host (cloud VM, workstation, private network), but the core control plane is explicitly tied to the hosted app.humanlayer.com service — no docs describe self-hosting that core product. missing for 10: no self-hosted control-plane/server option, no on-prem deployment guide, no Docker/Helm chart or license for running the full stack independently.
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Enterprise licensing
engineering-leadLicense an enterprise deployment with SSO and commercial support for organization-wide rollout
weight 2 · round drawnJulesnone0/10No evidence of enterprise licensing, SSO, or commercial support offerings; docs mention only individual API keys and usage-tier limits, and community discussion references a free beta plan, not enterprise deployment. missing for 10: SSO integration, enterprise licensing/contract terms, commercial support SLAs, org-wide admin/rollout tooling.
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “Is Jules free of charge? Yes, for now, Jules is free of charge. Jules is in beta and available without payment while we learn from usage.”
Model flexibility
engineering-leadBring my own LLM or API key so agents run on the model of my choice
weight 2 · round to HumanLayerJulesnone0/10No evidence that Jules supports bringing a custom LLM or third-party API key; it is tied to Google's own models (Gemini), with documentation only mentioning Jules's own API key for accessing Jules itself, not for configuring underlying model providers. missing for 10: any mention of BYO-LLM support, model selection options, or third-party API key configuration.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “Priority access to the latest models, starting with Gemini 3 Pro”
Docs show HumanLayer sessions can use different backends/models — Claude Code, OpenAI Codex, and Claude via Amazon Bedrock instead of the Anthropic API — indicating some flexibility in model/provider choice, which implies bringing your own credentials for these paths. However, there is no explicit doc describing a generic 'bring your own API key' mechanism for arbitrary LLMs or a pricing-tier note tying this to cost savings for engineering leads. Missing for 10: explicit BYO-API-key configuration docs, support for arbitrary/third-party model providers beyond Claude/Codex/Bedrock, and any pricing-related messaging about cost control via own keys.
- [claimed-docs] “This guide shows you how to install, authenticate, select, and check Codex for HumanLayer sessions.”
- [claimed-docs] “HumanLayer sessions can run Claude through Amazon Bedrock instead of the Anthropic API.”
- [claimed-docs] “HumanLayer registers these RPI sub-agents for Claude Code sessions”
Usage quotas
engineering-leadSee and manage plan-based daily task and concurrency limits for agent workflows
weight 2 · round to JulesJules docs reference a usage-limits page with plan-tiered daily task numbers (e.g., 300/60) and community confirms daily task limits exist and change between releases (60→15 on free plan), showing plan-based limits are real and documented. However there's no evidence of an engineering-lead-facing dashboard or admin console to monitor team usage, set/adjust concurrency, or manage limits across a team — only a static limits reference page and anecdotal user experience of hitting caps. Missing for 10: team/org usage dashboard, per-user concurrency visibility, ability to configure or request limit changes, and any admin/lead-specific management UI.
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…”
HumanLayernone0/10The evidence pack has no mention of plan-based daily task/concurrency limits or any admin controls for managing such limits; only a vague reference to 'Advanced Metrics for all paid plans' which covers usage/cost/productivity viewing, not concurrency or daily task limits management.
- [claimed-docs] “Advanced Metrics for all paid plans — View usage, cost, and productivity metrics, with optional access for every organization member.”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnJulesnone0/10No evidence in the pack mentions data residency, regional storage options, or any control over where Jules stores/processes data; docs only describe VM execution and GitHub integration without residency settings.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnJulesnone0/10No evidence in the pack addresses data usage/training opt-out policies, privacy controls, or any statement about whether user code/data is used to train AI models; the docs focus on functionality (VM execution, PR creation, API) rather than data governance.
ai-native userControl data retention and deletion
weight 2 · round drawnJulesnone0/10No evidence pack item mentions data retention policies, data deletion controls, or privacy settings for repositories/code processed by Jules; docs cover workflow, VM execution, and API usage but not retention/deletion controls.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnJulesnone0/10No evidence in the pack mentions telemetry, usage tracking, or an opt-out mechanism; the docs focus on repo access, VM execution, PR workflow, and API usage without addressing privacy/telemetry controls.
HumanLayernone0/10No evidence in the pack addresses telemetry, usage tracking, or opt-out controls for HumanLayer; docs cover integrations, workflows, and CLI usage but nothing about privacy/telemetry settings. Missing for 10: any mention of telemetry collection, opt-out mechanism, privacy policy, or data-handling documentation.
Repo integration — stories about repo integration in this arenaRepo integration
Stories about repo integration in this arena
Chat integration
developerTag an agent in a chat thread to discuss and delegate a bug or task
weight 2 · round to JulesJules supports delegating tasks via a GitHub issue label ('jules') and via API messages to the agent, which is a form of task delegation but not the classic 'tag an agent in a chat thread to discuss' interaction pattern; one community user built a custom MCP server to dispatch Jules tasks from VS Code Copilot Chat, showing it's possible but not a native feature. missing for 10: native chat-thread tagging/mention UI, evidence of back-and-forth discussion in a thread before delegation, first-party support for Slack/Teams-style @mentions.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
HumanLayernone0/10HumanLayer's Slack integration only pushes task-artifact updates into channels (docs-4) and other integrations (GitHub, Jira, Linear) create tasks from tickets/issues, not from tagging an agent inside a chat thread. There is no evidence of a chat-native @mention or in-thread delegation workflow for discussing/assigning tasks to an agent.
- [claimed-docs] “Connect Slack to send HumanLayer task artifact updates into the channels where your team already works.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
Knowledge context
developerAdd a context file describing my codebase conventions so agents generate more relevant plans and code
weight 3 · round to JulesJules automatically reads an AGENTS.md file in the repo root to guide its plans and code generation, directly matching the story of adding a context file describing codebase conventions. Missing for 10: no independent/hands-on confirmation of how well conventions in AGENTS.md actually improve plan/code relevance, and no detail on supported format/scope beyond the single doc mention.
- [claimed-docs] “Jules now automatically looks for a file named AGENTS.md in the root of your repository.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
HumanLayernone0/10The docs describe workspace-level config files (workspace.json/workspace.local.json) for team/machine settings and multi-repo setup, but there is no evidence of a dedicated context file for describing codebase conventions to improve agent-generated plans/code.
- [claimed-docs] “Use `.humanlayer/workspace.json` for shared repository or team configuration. Use `.humanlayer/workspace.local.json` for optional user or ma…”
- [claimed-docs] “Configure a multi-repository workspace with this repository, ../api, and ../web. Make ../web the primary repository. Ask me before you choos…”
Project management integration
product-managerConnect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly
weight 3 · round to HumanLayerJulesnone0/10Evidence only shows Jules integrating with GitHub (repos, issues, PRs) and provides an API for custom workflows, but there is no mention of Jira, Linear, ClickUp, Monday.com, or any issue-tracker integration beyond GitHub.
HumanLayer documents native Jira Cloud and Linear integrations that create tasks from tickets and sync/link status back to source issues, directly matching the story for those trackers. However, there is no evidence of ClickUp or Monday.com integrations, so the story is only partially delivered. Missing for 10: ClickUp integration docs, Monday.com integration docs, independent/hands-on verification of ticket sync working in practice.
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
Version control integration
developerConnect a GitHub repository so an agent can access the code and open pull requests against it
weight 3 · round to JulesJules is built around GitHub repo access: docs describe cloning repos, GitHub issue label assignment, generating diffs/plans, and creating PRs that can be merged on GitHub, corroborated by community reports of receiving usable PRs from Jules-driven tasks. Missing for 10: independent verification of the full connect-repo setup flow and edge-case reliability across repo types.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
Docs describe connecting GitHub to create tasks from issues and link artifacts back (humanlayer-docs-16), plus agent sessions can access code via configured workspaces/repos (humanlayer-docs-8, humanlayer-docs-11) and open draft PRs directly from the session UI (humanlayer-docs-19). missing for 10: independent/hands-on corroboration of the GitHub connection flow and PR-opening working end-to-end beyond first-party docs.
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “Use `.humanlayer/workspace.json` for shared repository or team configuration. Use `.humanlayer/workspace.local.json` for optional user or ma…”
- [claimed-docs] “Configure a multi-repository workspace with this repository, ../api, and ../web. Make ../web the primary repository. Ask me before you choos…”
developerGrant an agent access to my repositories with a one-click install, without complex setup
weight 2 · round to JulesJules integrates with GitHub (label-based task assignment, PR creation) and needs repo access to work, implying a connect/authorize flow, and community evidence confirms GitHub integration works well for many users. However, there's no explicit documentation of a literal 'one-click install' onboarding flow or app installation process described in the evidence pack. missing for 10: explicit one-click GitHub App install/authorization flow documentation, independent confirmation of setup simplicity.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I used Jules three times today, very impressive! It also handles coding-adjacent work. Good github integrations.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
HumanLayernone0/10Docs describe GitHub/Jira/Linear integrations for creating tasks from issues, but connecting a repo requires selecting a host, configuring workspace.json/workspace.local.json, and setting up remote daemons or multi-repo workspaces — none of this is framed as a one-click, no-setup install. No evidence pack item claims or demonstrates a one-click repo access flow.
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use `.humanlayer/workspace.json` for shared repository or team configuration. Use `.humanlayer/workspace.local.json` for optional user or ma…”
- [claimed-docs] “Configure a multi-repository workspace with this repository, ../api, and ../web. Make ../web the primary repository. Ask me before you choos…”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates
Quality gates on changes — review flow, required checks, merge protection
Ci remediation
engineering-leadHave failed CI workflows automatically diagnosed and fixed with a proposed pull request
weight 3 · round to HumanLayerJules offers building blocks—an API for custom automation that can 'automate tasks like bug fixing' and auto-create PRs via automationMode—but there is no direct evidence of a native integration that detects a failed CI workflow and automatically diagnoses/fixes it with a PR; this would require custom API wiring, not an out-of-box feature. Missing for 10: documented CI-failure trigger/integration, evidence of automatic diagnosis of CI logs, and case studies of this specific end-to-end flow.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
HumanLayer supports running automation sessions from CI (`humanlayer automation run`) and can open draft PRs from a session's diff/GitHub tab, and GitHub integration links tasks to issues—so the building blocks for a CI-triggered fix-and-PR flow exist. However there is no direct evidence of automatic diagnosis of failed CI logs/errors or a documented end-to-end 'CI failure → agent diagnosis → PR' pipeline. Missing for 10: explicit CI-failure-detection/diagnosis workflow docs, example of a failing pipeline auto-triggering a session, and confirmation the resulting PR addresses the CI failure specifically.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
developerTrigger an agent from CI/CD pipelines to fix a broken build or failing test
weight 2 · round to HumanLayerJules exposes a documented API for creating custom workflows and automating tasks like bug fixing (jules-docs-9, jules-docs-10, jules-docs-12), which could be wired into a CI/CD pipeline to trigger a fix, and community evidence shows users building custom integrations to dispatch Jules tasks programmatically (jules-comm-19). However, there is no first-party CI/CD-specific integration (e.g. GitHub Actions step, build-failure webhook) or documented example of triggering Jules to fix a broken build/failing test directly from a pipeline. missing for 10: explicit CI/CD pipeline integration/example, evidence of triggering on build/test failure events, hands-on confirmation of this specific workflow succeeding.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
HumanLayer docs explicitly describe `humanlayer automation run` for running a Cloud-visible coding session 'from any automation environment — a CI job, a cron machine, or a script on a server,' plus launch tokens for non-interactive command execution, directly supporting CI/CD-triggered agent runs. However, there is no explicit example or integration guide showing a CI failure (broken build/failing test) triggering the agent to diagnose and fix it, nor independent/hands-on confirmation of this workflow. missing for 10: a concrete CI/CD pipeline example tied to build/test failures, evidence of automatic failure detection triggering the agent, and independent verification of this automation flow working in practice.
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
Diff review
developerConfigure an agent to automatically open a pull request when its task completes
weight 2 · round to JulesJules natively creates a PR when a task completes (jules-docs-8), and the API exposes an `automationMode` field to configure automatic PR creation, defaulting to off unless explicitly configured (jules-docs-11), directly matching the story of configuring automatic PR-on-completion. Community reports corroborate that Jules produces usable PRs and users regularly review PRs it opens (jules-comm-6, jules-comm-12, jules-comm-16). missing for 10: independent hands-on verification of the automationMode toggle specifically, and more detail on configuration options/edge cases (e.g., partial completions, failed tasks) affecting whether a PR is always opened.
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [community] “its way better than the github thing in my experience it produces usable PRs”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
Docs confirm HumanLayer can open a draft pull request from a session (humanlayer-docs-19) and integrates with GitHub for issue-linked tasks (humanlayer-docs-16), but the evidence describes PR creation as a manual 'ask the session agent' action from the UI, not an automatic trigger fired upon task completion. Missing for 10: explicit configuration option/workflow setting for automatic PR creation on task completion, and any evidence of it happening without a manual prompt.
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
developerReview a diff of an agent's changes and approve it before it becomes a pull request
weight 3 · round to JulesJules explicitly provides a diff of changes for review and approval before it creates/publishes a PR, with plan approval and diff approval steps documented; community mentions confirm PRs are generated for review. Missing for 10: independent hands-on confirmation of the diff-approval UI flow itself (comments focus on PR quality/output rather than the diff-review step specifically).
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [community] “I've been actually kind-of enjoying using Jules as a way of 'coding' my side project using my phone. By the time I get home in the evening I…”
Release notes explicitly reference a diff view and 'draft PR creation' workflow (docs-19, docs-21), and tasks include 'One place for comments and review' (docs-5), indicating a review-before-PR mechanism. However, there's no detailed documentation of an explicit approve/reject gate tied specifically to diff review prior to PR creation, and community discussion raises concerns about automation bias in approvals (comm-3) without disputing the core capability. Missing for 10: explicit step-by-step approval workflow docs, independent hands-on verification of the diff-review-then-approve flow, and clarity on how rejection/edits are handled before PR creation.
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “Keyboard navigation for changed files — Move through the PR changes tree with J/K, N/P, G shortcuts, and Enter.”
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [community] “"I feel much more comfortable with senior developers approving AI changes to our codebase, then letting loose an autonomous agent with no hu…”
- [community] “User raised concern about automation bias: once an agent proves reliable, humans may rubber-stamp approvals without careful review, letting …”
Pr review automation
ai-native userHave incoming issues automatically triaged with severity suggested and routed to the right owner
weight 2 · round drawnJulesnone0/10Jules is a coding agent that executes tasks assigned via label or API, but there is no evidence of automatic issue triage, severity classification, or routing to owners — the closest feature is manual label-based task assignment, not automated triage. Missing for 10: any mention of severity scoring, triage logic, or owner-routing automation.
- [claimed-docs] “Use the "jules" label in an issue to assign a task directly in GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
HumanLayernone0/10HumanLayer connects issue trackers (Jira, GitHub, Linear) to create tasks from tickets, but there is no evidence of automatic triage, severity classification, or routing to an owner — integrations only create/link tasks, not assess or assign severity/ownership.
- [claimed-docs] “Connect Jira Cloud so HumanLayer can create tasks from Jira tickets and keep ticket-driven agent work connected to its source.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
- [claimed-docs] “Connect Linear so HumanLayer can create tasks from Linear issues, sync issue status, and link HumanLayer artifacts back to the source issue.”
engineering-leadHave every pull request automatically reviewed with AI-generated inline comments
weight 3 · round drawnJulesnone0/10Jules is documented as a task-execution agent that clones repos, makes code changes, and creates its own PRs for approval — not as a bot that automatically reviews incoming pull requests with inline comments. The only tangential mention is jules-docs-9's passing reference to 'automate tasks like ... code reviews' via API, but there is no documentation of an automatic inline-comment review gate applied to every PR.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
HumanLayernone0/10HumanLayer's evidence covers task/session management, human-in-the-loop approvals, draft PR creation, and a PR diff-viewing UI, but there is no mention of an automated review process that generates inline review comments on every pull request. Missing for 10: no evidence of automatic PR review triggers, no mention of AI-generated inline comments, no review-quality-gate CI integration for PRs.
- [claimed-docs] “Draft PR creation — Ask the session agent to open a draft pull request from the GitHub tab or diff view.”
- [claimed-docs] “Keyboard navigation for changed files — Move through the PR changes tree with J/K, N/P, G shortcuts, and Enter.”
- [claimed-docs] “Connect GitHub to create HumanLayer tasks from issues and link task artifacts back to the source issue.”
Readiness checks
engineering-leadRun a readiness report that evaluates how ready my repository is for autonomous agents
weight 2 · round drawnJulesnone0/10No evidence of a readiness report, scorecard, or repository-assessment feature for autonomous agent suitability; evidence only covers task execution, PR creation, and API usage. Missing for 10: any readiness-scoring feature, repo audit/checklist, or report output evaluating agent-readiness.
Security remediation
engineering-leadHave security alerts automatically validated and remediated with an opened pull request
weight 2 · round drawnJulesnone0/10Jules is a general coding agent that can be assigned tasks (via GitHub label or API) and opens PRs with diffs for review, but there is no evidence of any security-alert scanning, vulnerability detection/validation, or a workflow that automatically triggers remediation from a security alert (e.g., Dependabot/CVE integration). The evidence pack shows generic task-to-PR flow, not a security-alert-specific pipeline.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [claimed-docs] “Jules provides a diff of the changes. Quickly browse and approve code edits.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
HumanLayernone0/10HumanLayer's docs show generic task creation from GitHub/Jira/Linear issues and draft PR creation from agent sessions, but there is no evidence of any security-alert-specific validation or automated vulnerability remediation workflow (no CVE, dependency-alert, or security-scanner integration mentioned).
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
Concurrent execution
engineering-leadRun many agent tasks concurrently to scale delivery throughput
weight 3 · round to JulesJules architecture (cloud VMs per task, API for creating tasks, GitHub issue labeling) supports dispatching many tasks in parallel, and usage-limit docs explicitly reference 'agent-heavy workflows' with numeric caps (e.g. 300/60), and community reports (jules-comm-11, jules-comm-16) describe handling large task volumes successfully. However, other evidence shows real caveats: daily task limits were cut from 60 to 15 on the free plan (jules-comm-8), tasks can get stuck in loops with no stop button (jules-comm-14), and users report needing heavy babysitting (jules-comm-18), undercutting smooth high-throughput scaling. Missing for 10: dedicated documentation of a concurrency/queue dashboard, enterprise-tier concurrency guarantees, and independent benchmarks confirming reliable parallel execution at scale.
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “I've used this tool for a few months now and have been pretty impressed by it. It handles large quantities of tasks very well and is good at…”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “The daily task limit went down from 60 to 15 (on the free plan) with this release. Personally I wasn't close to exhausting the limit because…”
- [community] “Jules quickly after a few messages/prompts just gets stuck in an endless loop like Gemini-cli does. The worst part is that there is no 'STOP…”
- [community] “Do people really find Jules useful? I find it needs babysitting much more than Cursor.”
Docs describe an architecture (tasks/sessions, multi-repository workspaces, remote daemons on cloud VMs, and a CLI 'automation run' for CI/cron/scripts) that could support running many agent tasks in parallel, and 'Advanced Metrics' track usage/productivity across an org. However, tutorials and guides are framed around running one task/session at a time, and there is no explicit documentation or evidence of concurrent multi-task orchestration, throughput dashboards, or scaling guarantees for many simultaneous agents. Missing for 10: explicit multi-session concurrency docs, evidence of parallel task orchestration at scale, and independent confirmation of throughput gains.
- [claimed-docs] “A HumanLayer task gives one piece of work: A group of related sessions, A shared set of task files, One place for comments and review, A his…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Configure a multi-repository workspace with this repository, ../api, and ../web. Make ../web the primary repository. Ask me before you choos…”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Advanced Metrics for all paid plans — View usage, cost, and productivity metrics, with optional access for every organization member.”
engineering-leadCreate agent sessions on behalf of other users in my organization
weight 2 · round drawnJulesnone0/10No evidence of any organization, team, or admin/delegation features allowing an engineering-lead to create sessions on behalf of other users; Jules docs only describe individual API keys, per-user GitHub connections, and personal task creation. Missing for 10: org/team management, delegated session creation, role-based admin controls, multi-user account provisioning.
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Jules needs access to your repositories in order to work.”
HumanLayernone0/10The evidence describes tasks, sessions, integrations, and org-wide metrics access but never mentions an engineering-lead or admin creating/assigning agent sessions on behalf of another named user in the organization. No account-delegation, impersonation, or 'assign session to teammate' capability is documented.
Deployment flexibility
developerUse a managed cloud offering to run agents without operating my own backend infrastructure
weight 2 · round to JulesJules is explicitly a cloud-hosted agent service: tasks run in Google-managed VMs, integrated with GitHub, with an API and usage limits/plans, requiring no self-hosted backend. Community evidence corroborates real-world usage at scale (75% success on customer PRs, teams dispatching many tasks) though some users report reliability/UI issues. Missing for 10: no independent infra/SLA details, no discoverable OpenAPI/docs.md confirming API completeness, and mixed community reliability reports.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “Power users & agent-heavy workflows ... 300 ... 60”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
HumanLayer offers a cloud control plane (app.humanlayer.com, automation sessions marked 'Cloud-visible', advanced metrics for paid plans) that lets users monitor and manage agent runs remotely, suggesting a managed service layer. However, docs explicitly state the actual agent execution still runs on a host that the user must select and operate—'a cloud VM, workstation, or private-network machine'—meaning the developer still needs to provision/manage backend compute for the daemon rather than HumanLayer fully hosting execution. Missing for 10: evidence of a fully HumanLayer-operated compute backend (no user-managed VM/daemon required), and independent confirmation of a true zero-ops managed runtime.
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “Advanced Metrics for all paid plans — View usage, cost, and productivity metrics, with optional access for every organization member.”
engineering-leadSelf-host agent infrastructure locally, in containers, or on my own VMs
weight 2 · round to HumanLayerJulesnone0/10Jules is explicitly cloud-hosted: it runs tasks in Google's own VMs, and community feedback confirms it is not local (jules-comm-1). There is no evidence of any self-hosting, on-prem, container, or private-VM deployment option.
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [community] “Jules runs in cloud-based VMs instead of on my local machine, making it much less useful than Claude Code. My projects have bespoke build sc…”
Docs describe running the remote daemon on a cloud VM, workstation, or private-network machine that you control (humanlayer-docs-7), plus automation sessions from CI/cron/server environments (humanlayer-docs-12) and launch tokens for bootstrap scripts (humanlayer-docs-13), showing self-hosted deployment flexibility. However there is no explicit mention of container/Docker deployment or an official container image, and no independent verification of self-hosted setups at scale. Missing for 10: explicit container/Docker packaging docs, independent hands-on confirmation of self-hosted deployments.
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
- [claimed-docs] “This tutorial teaches you how to run one task on a remote machine. You will control the task from app.humanlayer.com on a machine or phone.”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
Headless automation
developerRun an agent headlessly inside CI/CD pipelines and shell scripts
weight 2 · round to HumanLayerJules exposes a public API (create tasks, send messages, optional automationMode) explicitly pitched for 'custom workflows' and embedding into other tools, which can be scripted/curled headlessly, and a community member built an MCP server hitting the API to dispatch tasks programmatically. However there's no documented CLI, no explicit CI/CD pipeline examples (e.g., GitHub Actions), and Jules always runs its own cloud VM rather than a lightweight headless process invokable inline in a shell script/pipeline step. missing for 10: official CLI or CI/CD pipeline integration docs (e.g. GitHub Actions step), examples of shell-script invocation, confirmation that API calls run synchronously enough for pipeline gating.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [claimed-docs] “In the Jules web app, go to the Settings page to create a new API key. You can have at most 3 API keys at a time.”
- [claimed-docs] “The `automationMode` field is optional. By default, no PR will be automatically created.”
- [claimed-docs] “To send a message to the agent:”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
- [community] “I've moved infrastructure to Railway and taught the customer to use Jules. They make their own PRs now, Railway spins up an environment with…”
Docs explicitly describe `humanlayer automation run` for running Cloud-visible coding sessions from CI jobs, cron machines, or scripts, plus launch tokens for non-interactive/non-PTY execution suited to headless environments. This directly matches running an agent headlessly in CI/CD and shell scripts. Missing for 10: independent/hands-on verification of CI usage and concrete pipeline examples (e.g. GitHub Actions config).
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
- [claimed-docs] “Use a launch token for one non-interactive command. Examples include a shell without an interactive PTY and a one-time remote bootstrap scri…”
- [claimed-docs] “The host can be a cloud VM, workstation, or private-network machine. Select a host that can access the code, tools, credentials, and private…”
Not comparable on these axes
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableJulesn/aJules is itself a coding agent (client role), and the evidence only shows a REST API plus a third-party/community-built MCP server (jules-comm-19), not an official first-party MCP server exposing Jules as a tool endpoint. Per the agent-role rule, this axis is out of category rather than a failed capability.
- [claimed-docs] “You can use the API to create custom workflows, automate tasks like bug fixing and code reviews, and embed Jules's intelligence directly int…”
- [community] “Was able to build a personal MCP server that connects to the Jules API, letting me dispatch tasks to Jules, from Copilot Chat in VS Code.”
HumanLayernone0/10HumanLayer is a platform/control-plane for running and overseeing coding-agent sessions (Claude Code, Codex) rather than itself being an agent, so an official MCP server is a fair, applicable axis. The evidence pack documents many integrations (Jira, Slack, GitHub, Linear), a CLI, and remote daemons, but no MCP server offering is mentioned anywhere, and API/OpenAPI probes returned 404s. missing for 10: any first-party MCP server documentation, endpoint, or 'mcp serve' style capability.
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.humanlayer.com/openapi.json, https://docs.humanlayer.com/swagger.json, https://docs.hum…”
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.humanlayer.com/llms.txt”
- [claimed-docs] “Use `humanlayer automation run` to run one Cloud-visible coding session from any automation environment — a CI job, a cron machine, or a scr…”
developerQuery generated documentation for any public or private repository
weight 1 · not comparableJulesnone0/10Evidence describes Jules as a task-based coding agent that clones repos, edits code, and creates PRs, but there is no mention of generating or letting users query documentation for a repository (public or private). missing for 10: any docs-generation feature, a documentation query/search interface, evidence of indexing repo docs for Q&A.
- [claimed-docs] “Jules needs access to your repositories in order to work.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub.”
engineering-leadAutomatically fix failing agent-readiness criteria in my repository
weight 1 · not comparableJulesnone0/10Jules can be assigned generic bug-fix/coding tasks and reads an AGENTS.md file if present, but there is no evidence of it detecting or automatically remediating 'agent-readiness' quality-gate criteria (e.g., missing/invalid AGENTS.md, agent-compatibility checks) as a review gate — it only consumes such files, it doesn't audit or fix them as a compliance gate.
- [claimed-docs] “Jules now automatically looks for a file named AGENTS.md in the root of your repository.”
- [claimed-docs] “Jules runs in a virtual machine where it clones your code, installs dependencies, and modifies files.”
- [claimed-docs] “Once you submit a task, Jules will generate a plan. You can review and approve it before any code changes are made.”
- [community] “I gave Jules a try on a new, very unorganized Python project. It failed the first time, I gave it the error message and it fixed it. I was s…”
HumanLayern/aHumanLayer is a human-in-the-loop agent orchestration/approval tool for running coding tasks, not a repository readiness/compliance scanner with auto-remediation of 'agent-readiness criteria'. This axis is a category error for this product type — no evidence pack content relates to detecting or auto-fixing repo readiness criteria.