Access
Verified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementationevidence →
End-to-end implementation by the agent — multi-file changes, task completion
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversightevidence →
Keeping a human in the loop — approvals, checkpoints, interrupts
Intent to spec — stories about intent to spec in this arenaIntent to specevidence →
Stories about intent to spec in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Repo integration — stories about repo integration in this arenaRepo integrationevidence →
Stories about repo integration in this arena
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gatesevidence →
Quality gates on changes — review flow, required checks, merge protection
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelismevidence →
Running many jobs at once — concurrency, fleets, queueing
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 1 paid · 0 enterprise · 39 not stated in evidence
Follow the green: where the map greys out is where HumanLayer stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
~4/10
unlocks → Webhooks · Official SDKs · MCP server · Machine-readable spec · Versioning policy · API sandbox · Full data export · Have an agent automatically generate and run tests to validate its own code changes before proposing them · Go from a mockup or design to a working implementation without an engineering handoff · Have an agent safely execute code and install dependencies inside an isolated sandbox · Tag an agent in a chat thread to discuss and delegate a bug or task · Add a context file describing my codebase conventions so agents generate more relevant plans and code · Grant an agent access to my repositories with a one-click install, without complex setup
Subscribe to events via webhooks
—0/10
Build against official SDKs
—0/10
Issue scoped/least-privilege API credentials for an agent
~3/10
Connect an agent via an official MCP server
—0/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—–
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓7/10
unlocks → MCP client
Operate the product with natural-language commands
✓7/10
Plug MCP servers into this product so it can use their tools
—–
Get AI-generated insights and suggestions from my data inside the product
~4/10
Set up automations that run autonomously in the background
✓7/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation
End-to-end implementation by the agent — multi-file changes, task completion
End to end feature delivery
Have an agent automatically clone the repo, install dependencies, and configure its own working environment
~5/10
Interactive takeover
Have an agent safely execute code and install dependencies inside an isolated sandbox
—0/10
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight
Keeping a human in the loop — approvals, checkpoints, interrupts
Approval controls
Model control
Intent to spec — stories about intent to spec in this arenaIntent to spec
Stories about intent to spec in this arena
Natural language task intake
Plan approval
Assign a coding task to an agent directly from an existing issue or ticket
✓8/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Repo integration — stories about repo integration in this arenaRepo integration
Stories about repo integration in this arena
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates
Quality gates on changes — review flow, required checks, merge protection
Ci remediation
Diff review
Pr review automation
Readiness checks
Have security alerts automatically validated and remediated with an opened pull request
—–
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
Concurrent execution
Deployment flexibility
Run an agent headlessly inside CI/CD pipelines and shell scripts
✓8/10
Sorted by importance (agentic first) (high → low) · 73/73 stories · click a row’s chevron for the rationale and evidence
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 7/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 4/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialpaid | 4/10 | Cclaimed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 3/10 | Cclaimed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | untested | none yet | |
Assign a coding task to an agent directly from an existing issue or ticket C Ticket driven tasking | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 3 | full | 8/10 | Cclaimed | |
Connect a GitHub repository so an agent can access the code and open pull requests against it C Version control integration | developer | Repo integration — stories about repo integration in this arenaRepo integration | 3 | full | 8/10 | Cclaimed | |
Review a diff of an agent's changes and approve it before it becomes a pull request C Diff review | developer | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 3 | partial | 7/10 | Xcommunity | |
Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly C Project management integration | product-manager | Repo integration — stories about repo integration in this arenaRepo integration | 3 | partial | 6/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 6/10 | Cclaimed | |
Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task C Natural language task intake | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 3 | partial | 6/10 | Cclaimed | |
Have an agent autonomously diagnose and fix a reported bug C End to end feature delivery | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 3 | partial | 6/10 | Xcommunity | |
Review and approve an agent's implementation plan before any code changes are made C Plan approval | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 3 | partial | 6/10 | Cclaimed | |
Watch what a running agent is doing in real time, including its current status C Visibility monitoring | developer | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 3 | partial | 6/10 | Cclaimed | |
Have an agent implement a requested feature end-to-end, including writing tests C End to end feature delivery | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 3 | partial | 5/10 | Cclaimed | |
Have failed CI workflows automatically diagnosed and fixed with a proposed pull request C Ci remediation | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 3 | partial | 5/10 | Cclaimed | |
Run many agent tasks concurrently to scale delivery throughput C Concurrent execution | engineering-lead | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 3 | partial | 5/10 | Cclaimed | |
Set tiered autonomy levels controlling what an agent can do without manual confirmation C Approval controls | engineering-lead | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 3 | partial | 5/10 | Cclaimed | |
Add a context file describing my codebase conventions so agents generate more relevant plans and code C Knowledge context | developer | Repo integration — stories about repo integration in this arenaRepo integration | 3 | none | 0/10 | ||
Have an agent safely execute code and install dependencies inside an isolated sandbox C Sandbox execution | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 3 | none | 0/10 | ||
Have every pull request automatically reviewed with AI-generated inline comments C Pr review automation | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Run an agent headlessly inside CI/CD pipelines and shell scripts C Headless automation | developer | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | full | 8/10 | Cclaimed | |
Self-host agent infrastructure locally, in containers, or on my own VMs C Deployment flexibility | engineering-lead | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | partial | 7/10 | Cclaimed | |
Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work C Interactive takeover | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | full | 7/10 | Cclaimed | |
Approve a task's scope and contract before an agent is allowed to modify the repository C Plan approval | engineering-lead | Intent to spec — stories about intent to spec in this arenaIntent to spec | 2 | partial | 6/10 | Xcommunity | |
Get notified when an agent completes a task or needs my input C Visibility monitoring | developer | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 2 | partial | 6/10 | Xcommunity | |
Trigger an agent from CI/CD pipelines to fix a broken build or failing test C Ci remediation | developer | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | partial | 6/10 | Cclaimed | |
Bring my own LLM or API key so agents run on the model of my choice C Model flexibility | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | partial | 5/10 | Cclaimed | |
Have an agent automatically clone the repo, install dependencies, and configure its own working environment C Environment setup | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | partial | 5/10 | Cclaimed | |
Use a managed cloud offering to run agents without operating my own backend infrastructure C Deployment flexibility | developer | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | partial | 5/10 | Cclaimed | |
Attach a marked-up screenshot or mockup to a task so the agent implements the correct visual change C Natural language task intake | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 2 | partial | 4/10 | Cclaimed | |
Configure an agent to auto-approve all its actions instead of confirming each one C Approval controls | developer | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 2 | partial | 4/10 | Cclaimed | |
Configure an agent to automatically open a pull request when its task completes C Diff review | developer | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | partial | 4/10 | Cclaimed | |
Convert user feedback submissions into structured tasks with proposed scope C Natural language task intake | product-manager | Intent to spec — stories about intent to spec in this arenaIntent to spec | 2 | partial | 4/10 | Cclaimed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 4/10 | Cclaimed | |
Send follow-up instructions to an active agent session to steer its work without restarting C Interactive takeover | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | partial | 4/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 3/10 | Tprobed | |
Go from a mockup or design to a working implementation without an engineering handoff C End to end feature delivery | product-manager | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | none | 0/10 | ||
Grant an agent access to my repositories with a one-click install, without complex setup C Version control integration | developer | Repo integration — stories about repo integration in this arenaRepo integration | 2 | none | 0/10 | ||
Have each task prompt automatically routed to the most suitable underlying model C Model control | ai-native user | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 2 | none | 0/10 | ||
Have incoming issues automatically triaged with severity suggested and routed to the right owner C Pr review automation | ai-native user | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | none | 0/10 | ||
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
See and manage plan-based daily task and concurrency limits for agent workflows G Usage quotas | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | 0/10 | ||
Tag an agent in a chat thread to discuss and delegate a bug or task C Chat integration | developer | Repo integration — stories about repo integration in this arenaRepo integration | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Create agent sessions on behalf of other users in my organization C Concurrent execution | engineering-lead | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | none | untested | none yet | |
Have an agent automatically generate and run tests to validate its own code changes before proposing them C End to end feature delivery | ai-native user | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | none | untested | none yet | |
Have security alerts automatically validated and remediated with an opened pull request C Security remediation | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | none | untested | none yet | |
License an enterprise deployment with SSO and commercial support for organization-wide rollout G Enterprise licensing | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Run a readiness report that evaluates how ready my repository is for autonomous agents C Readiness checks | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | none | untested | none yet | |
Approve key agent decisions from my phone while agents continue working C Approval controls | product-manager | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 1 | full | 8/10 | Xcommunity | |
Switch away from automatic model selection to a specific model of my choice C Model control | engineering-lead | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 1 | partial | 6/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Cclaimed | |
Automatically fix failing agent-readiness criteria in my repository C Readiness checks | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 1 | n/a | untested | none yet | |
Query generated documentation for any public or private repository C Knowledge context | developer | Repo integration — stories about repo integration in this arenaRepo integration | 1 | n/a | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 61 stories with headroom
What would move HumanLayer’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
Missing: any mention of MCP protocol support, MCP server configuration, or tool-plugin mechanism.
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
Missing: any first-party MCP server documentation, endpoint, or 'mcp serve' style capability.
Review quality gates — quality gates on changes — review flow, required checks, merge protectionHave every pull request automatically reviewed with AI-generated inline comments
nonemoves PA Scoreimpact 30
Missing: no evidence of automatic PR review triggers, no mention of AI-generated inline comments, no review-quality-gate CI integration for PRs.
Repo integration — stories about repo integration in this arenaAdd a context file describing my codebase conventions so agents generate more relevant plans and code
nonemoves PA Scoreimpact 30
The docs describe workspace-level config files (workspace.json/workspace.local.json) for team/machine settings and multi-repo setup, but there is no evidence of a dedicated context file for describing codebase conventions to improve agent-generated plans/code.
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionHave an agent safely execute code and install dependencies inside an isolated sandbox
nonemoves PA Scoreimpact 30
HumanLayer docs describe running sessions on remote hosts (cloud VM, workstation, private-network machine) and automation environments, but there is no mention of an isolated/sandboxed execution environment for running code or installing dependencies safely — the host selection is about access/credentials, not isolation guarantees.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
No evidence of data export functionality or open-format data portability; docs focus on tasks, workflows, and integrations, with no mention of exporting user data or account deletion/leaving.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Missing: no self-hosted control-plane/server option, no on-prem deployment guide, no Docker/Helm chart or license for running the full stack independently.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence in the pack addresses data usage for AI model training, opt-out policies, or privacy commitments regarding training data; the docs focus on task workflows, integrations, and remote sessions.
Showing the top 8 of 61 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 40 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Guide docs33 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Have an agent autonomously diagnose and fix a reported bug
- Have an agent implement a requested feature end-to-end, including writing tests
- Have an agent automatically clone the repo, install dependencies, and configure its own working environment
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Configure an agent to auto-approve all its actions instead of confirming each one
- Set tiered autonomy levels controlling what an agent can do without manual confirmation
- Switch away from automatic model selection to a specific model of my choice
- Get notified when an agent completes a task or needs my input
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Convert user feedback submissions into structured tasks with proposed scope
- Review and approve an agent's implementation plan before any code changes are made
- Approve a task's scope and contract before an agent is allowed to modify the repository
- Assign a coding task to an agent directly from an existing issue or ticket
- Do everything through the API that I can do in the UI
- Bring my own LLM or API key so agents run on the model of my choice
- Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly
- Connect a GitHub repository so an agent can access the code and open pull requests against it
- Have failed CI workflows automatically diagnosed and fixed with a proposed pull request
- Trigger an agent from CI/CD pipelines to fix a broken build or failing test
- Configure an agent to automatically open a pull request when its task completes
- Run many agent tasks concurrently to scale delivery throughput
- Use a managed cloud offering to run agents without operating my own backend infrastructure
- Self-host agent infrastructure locally, in containers, or on my own VMs
- Run an agent headlessly inside CI/CD pipelines and shell scripts
Explanation docs29 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Version, review, and roll back my automations
- Have an agent autonomously diagnose and fix a reported bug
- Have an agent implement a requested feature end-to-end, including writing tests
- Have an agent automatically clone the repo, install dependencies, and configure its own working environment
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Send follow-up instructions to an active agent session to steer its work without restarting
- Configure an agent to auto-approve all its actions instead of confirming each one
- Approve key agent decisions from my phone while agents continue working
- Set tiered autonomy levels controlling what an agent can do without manual confirmation
- Watch what a running agent is doing in real time, including its current status
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Convert user feedback submissions into structured tasks with proposed scope
- Review and approve an agent's implementation plan before any code changes are made
- Approve a task's scope and contract before an agent is allowed to modify the repository
- Assign a coding task to an agent directly from an existing issue or ticket
- Trigger an agent from CI/CD pipelines to fix a broken build or failing test
- Review a diff of an agent's changes and approve it before it becomes a pull request
- Run many agent tasks concurrently to scale delivery throughput
- Use a managed cloud offering to run agents without operating my own backend infrastructure
- Self-host agent infrastructure locally, in containers, or on my own VMs
- Run an agent headlessly inside CI/CD pipelines and shell scripts
Release notes docs13 stories
- Get AI-generated insights and suggestions from my data inside the product
- Version, review, and roll back my automations
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Send follow-up instructions to an active agent session to steer its work without restarting
- Watch what a running agent is doing in real time, including its current status
- Attach a marked-up screenshot or mockup to a task so the agent implements the correct visual change
- Do everything through the API that I can do in the UI
- Connect a GitHub repository so an agent can access the code and open pull requests against it
- Have failed CI workflows automatically diagnosed and fixed with a proposed pull request
- Configure an agent to automatically open a pull request when its task completes
- Review a diff of an agent's changes and approve it before it becomes a pull request
- Run many agent tasks concurrently to scale delivery throughput
- Use a managed cloud offering to run agents without operating my own backend infrastructure
API reference10 stories
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Have an agent implement a requested feature end-to-end, including writing tests
- Have an agent automatically clone the repo, install dependencies, and configure its own working environment
- Set tiered autonomy levels controlling what an agent can do without manual confirmation
- Switch away from automatic model selection to a specific model of my choice
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Review and approve an agent's implementation plan before any code changes are made
- Bring my own LLM or API key so agents run on the model of my choice
- Connect a GitHub repository so an agent can access the code and open pull requests against it
Tutorials docs7 stories
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Approve key agent decisions from my phone while agents continue working
- Watch what a running agent is doing in real time, including its current status
- Use a managed cloud offering to run agents without operating my own backend infrastructure
- Self-host agent infrastructure locally, in containers, or on my own VMs
Hacker News5 stories
- Have an agent autonomously diagnose and fix a reported bug
- Approve key agent decisions from my phone while agents continue working
- Get notified when an agent completes a task or needs my input
- Approve a task's scope and contract before an agent is allowed to modify the repository
- Review a diff of an agent's changes and approve it before it becomes a pull request
llms.txt2 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
3 of 17 testable claims verified · 1 contradicted → integrity 6/100
20 distinct capability claims found in HumanLayer’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
3
Verified
13
Unverified
1
Contradicted
24
Undersold
Verified (3)
“Run tasks from a phone or any machine via app.humanlayer.com for remote control”
Approve key agent decisions from my phone while agents continue workingfullproof ↗
“Run a Cloud-visible coding session from CI jobs, cron machines, or scripts via the automation CLI”
Run the product headlessly / in CI for automationfullproof ↗
“Ask the session agent to open a draft pull request from the GitHub tab or diff view”
Review a diff of an agent's changes and approve it before it becomes a pull requestpartialproof ↗
Unverified (16)
“Connect Jira Cloud so tasks are created automatically from Jira tickets”
Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directlypartialproof ↗
“Connect Linear so tasks are created from Linear issues and status stays synced”
Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directlypartialproof ↗
“Connect GitHub so tasks can be created from issues and linked back to them”
Connect a GitHub repository so an agent can access the code and open pull requests against itfullproof ↗
“Connect GitHub so tasks can be created from issues and linked back to them”
Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directlypartialproof ↗
“Choose the shortest workflow path (e.g. Oneshot) matched to the risk of the change”
Set tiered autonomy levels controlling what an agent can do without manual confirmationpartialproof ↗
“Select between Oneshot, RPI, PRD-Oriented, or Freeform task workflow modes”
Review and approve an agent's implementation plan before any code changes are madepartialproof ↗
“Choose a host (cloud VM, workstation, or private-network machine) with access to needed code, tools, and credentials”
Self-host agent infrastructure locally, in containers, or on my own VMspartialproof ↗
“Configure a multi-repository workspace with a primary repo and setup preferences”
Have an agent automatically clone the repo, install dependencies, and configure its own working environmentpartialproof ↗
“Run a Cloud-visible coding session from CI jobs, cron machines, or scripts via the automation CLI”
Run an agent headlessly inside CI/CD pipelines and shell scriptsfullproof ↗
“Use a launch token to authorize a single non-interactive automation command”
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
“Install, authenticate, and select Codex as the model for HumanLayer sessions”
Bring my own LLM or API key so agents run on the model of my choicepartialproof ↗
“Install, authenticate, and select Codex as the model for HumanLayer sessions”
Switch away from automatic model selection to a specific model of my choicepartialproof ↗
“Run Claude sessions through Amazon Bedrock instead of the Anthropic API”
Bring my own LLM or API key so agents run on the model of my choicepartialproof ↗
“Ask the session agent to open a draft pull request from the GitHub tab or diff view”
Configure an agent to automatically open a pull request when its task completespartialproof ↗
“Paste images directly into a new task as attachments for visual context”
Attach a marked-up screenshot or mockup to a task so the agent implements the correct visual changepartialproof ↗
“Tasks maintain a shared history that can continue across agents and workstations”
Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the workfullproof ↗
Contradicted (1)
“Use workspace.json and workspace.local.json for shared and machine-specific configuration”
Add a context file describing my codebase conventions so agents generate more relevant plans and codenoneproof ↗
Undersold (24)
Drive the product through a documented public APIpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Set up automations that run autonomously in the backgroundfullproof ↗
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
Operate the product with natural-language commandsfullproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Have an agent autonomously diagnose and fix a reported bugpartialproof ↗
Have an agent implement a requested feature end-to-end, including writing testspartialproof ↗
Send follow-up instructions to an active agent session to steer its work without restartingpartialproof ↗
Configure an agent to auto-approve all its actions instead of confirming each onepartialproof ↗
Watch what a running agent is doing in real time, including its current statuspartialproof ↗
Get notified when an agent completes a task or needs my inputpartialproof ↗
Describe a feature or bug in plain language and have it automatically turned into a scoped implementation taskpartialproof ↗
Convert user feedback submissions into structured tasks with proposed scopepartialproof ↗
Approve a task's scope and contract before an agent is allowed to modify the repositorypartialproof ↗
Assign a coding task to an agent directly from an existing issue or ticketfullproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Have failed CI workflows automatically diagnosed and fixed with a proposed pull requestpartialproof ↗
Trigger an agent from CI/CD pipelines to fix a broken build or failing testpartialproof ↗
Run many agent tasks concurrently to scale delivery throughputpartialproof ↗
Use a managed cloud offering to run agents without operating my own backend infrastructurepartialproof ↗
Claims outside our story set (4)
Real capability claims found in HumanLayer’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Send task artifact updates into Slack channels”
source ↗“Live multiplayer editing of response drafts with shared cursors and presence”
source ↗“View usage, cost, and productivity metrics for paid plans”
source ↗“HumanLayer registers RPI sub-agents for Claude Code sessions”
source ↗
Business model
Starter free for teams up to 3 (200 sessions/mo); Pro $100/user/mo with BYOK for Claude/Codex subscriptions or keys and unlimited sessions; Enterprise adds SSO/SAML, audit logs, on-prem/VPC.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
