Install
Showcase


Verified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementationevidence →
End-to-end implementation by the agent — multi-file changes, task completion
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversightevidence →
Keeping a human in the loop — approvals, checkpoints, interrupts
Intent to spec — stories about intent to spec in this arenaIntent to specevidence →
Stories about intent to spec in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Repo integration — stories about repo integration in this arenaRepo integrationevidence →
Stories about repo integration in this arena
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gatesevidence →
Quality gates on changes — review flow, required checks, merge protection
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelismevidence →
Running many jobs at once — concurrency, fleets, queueing
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Factory stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
~6/10
unlocks → Webhooks · Scoped API keys · Machine-readable spec · Versioning policy · API sandbox · Full data export · Have an agent automatically generate and run tests to validate its own code changes before proposing them · Have an agent safely execute code and install dependencies inside an isolated sandbox · Add a context file describing my codebase conventions so agents generate more relevant plans and code · Query generated documentation for any public or private repository · Grant an agent access to my repositories with a one-click install, without complex setup
Subscribe to events via webhooks
—–
Build against official SDKs
~4/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
~5/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—–
Explore an interactive API reference with runnable examples
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓8/10
Operate the product with natural-language commands
✓8/10
Plug MCP servers into this product so it can use their tools
~6/10
Get AI-generated insights and suggestions from my data inside the product
n/an/a
Set up automations that run autonomously in the background
~6/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation
End-to-end implementation by the agent — multi-file changes, task completion
End to end feature delivery
Have an agent automatically clone the repo, install dependencies, and configure its own working environment
~4/10
Interactive takeover
Have an agent safely execute code and install dependencies inside an isolated sandbox
—–
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight
Keeping a human in the loop — approvals, checkpoints, interrupts
Approval controls
Model control
Intent to spec — stories about intent to spec in this arenaIntent to spec
Stories about intent to spec in this arena
Natural language task intake
Plan approval
Assign a coding task to an agent directly from an existing issue or ticket
~4/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Repo integration — stories about repo integration in this arenaRepo integration
Stories about repo integration in this arena
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates
Quality gates on changes — review flow, required checks, merge protection
Ci remediation
Diff review
Pr review automation
Readiness checks
Have security alerts automatically validated and remediated with an opened pull request
—–
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
Concurrent execution
Deployment flexibility
Run an agent headlessly inside CI/CD pipelines and shell scripts
✓9/10
Sorted by importance (agentic first) (high → low) · 73/73 stories · click a row’s chevron for the rationale and evidence
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 6/10 | Tprobed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 6/10 | Cclaimed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 5/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | untested | none yet | |
Have an agent implement a requested feature end-to-end, including writing tests C End to end feature delivery | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 3 | partial | 7/10 | Cclaimed | |
Run many agent tasks concurrently to scale delivery throughput C Concurrent execution | engineering-lead | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 3 | partial | 7/10 | Cclaimed | |
Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task C Natural language task intake | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 3 | partial | 6/10 | Cclaimed | |
Have an agent autonomously diagnose and fix a reported bug C End to end feature delivery | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 3 | partial | 6/10 | Cclaimed | |
Review a diff of an agent's changes and approve it before it becomes a pull request C Diff review | developer | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 3 | partial | 6/10 | Cclaimed | |
Set tiered autonomy levels controlling what an agent can do without manual confirmation C Approval controls | engineering-lead | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 3 | partial | 6/10 | Cclaimed | |
Connect a GitHub repository so an agent can access the code and open pull requests against it C Version control integration | developer | Repo integration — stories about repo integration in this arenaRepo integration | 3 | partial | 5/10 | Cclaimed | |
Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly C Project management integration | product-manager | Repo integration — stories about repo integration in this arenaRepo integration | 3 | partial | 5/10 | Cclaimed | |
Watch what a running agent is doing in real time, including its current status C Visibility monitoring | developer | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 3 | partial | 5/10 | Cclaimed | |
Assign a coding task to an agent directly from an existing issue or ticket C Ticket driven tasking | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 3 | partial | 4/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Have failed CI workflows automatically diagnosed and fixed with a proposed pull request C Ci remediation | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 3 | partial | 4/10 | Cclaimed | |
Review and approve an agent's implementation plan before any code changes are made C Plan approval | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 3 | partial | 4/10 | Cclaimed | |
Have every pull request automatically reviewed with AI-generated inline comments C Pr review automation | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 3 | none | 0/10 | ||
Add a context file describing my codebase conventions so agents generate more relevant plans and code C Knowledge context | developer | Repo integration — stories about repo integration in this arenaRepo integration | 3 | none | untested | none yet | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Have an agent safely execute code and install dependencies inside an isolated sandbox C Sandbox execution | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Run an agent headlessly inside CI/CD pipelines and shell scripts C Headless automation | developer | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | full | 9/10 | Cclaimed | |
Run a readiness report that evaluates how ready my repository is for autonomous agents C Readiness checks | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | full | 8/10 | Cclaimed | |
Trigger an agent from CI/CD pipelines to fix a broken build or failing test C Ci remediation | developer | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | full | 8/10 | Tprobed | |
Go from a mockup or design to a working implementation without an engineering handoff C End to end feature delivery | product-manager | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | full | 7/10 | Cclaimed | |
Use a managed cloud offering to run agents without operating my own backend infrastructure C Deployment flexibility | developer | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | full | 7/10 | Cclaimed | |
Configure an agent to auto-approve all its actions instead of confirming each one C Approval controls | developer | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Send follow-up instructions to an active agent session to steer its work without restarting C Interactive takeover | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | partial | 6/10 | Cclaimed | |
Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work C Interactive takeover | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | partial | 6/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Tprobed | |
Approve a task's scope and contract before an agent is allowed to modify the repository C Plan approval | engineering-lead | Intent to spec — stories about intent to spec in this arenaIntent to spec | 2 | partial | 4/10 | Cclaimed | |
Get notified when an agent completes a task or needs my input C Visibility monitoring | developer | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 2 | partial | 4/10 | Cclaimed | |
Have an agent automatically clone the repo, install dependencies, and configure its own working environment C Environment setup | developer | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | partial | 4/10 | Cclaimed | |
Configure an agent to automatically open a pull request when its task completes C Diff review | developer | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | partial | 3/10 | Cclaimed | |
Tag an agent in a chat thread to discuss and delegate a bug or task C Chat integration | developer | Repo integration — stories about repo integration in this arenaRepo integration | 2 | partial | 3/10 | Cclaimed | |
Attach a marked-up screenshot or mockup to a task so the agent implements the correct visual change C Natural language task intake | developer | Intent to spec — stories about intent to spec in this arenaIntent to spec | 2 | none | 0/10 | ||
Create agent sessions on behalf of other users in my organization C Concurrent execution | engineering-lead | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | none | 0/10 | ||
Have an agent automatically generate and run tests to validate its own code changes before proposing them C End to end feature delivery | ai-native user | Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation | 2 | none | 0/10 | ||
Have incoming issues automatically triaged with severity suggested and routed to the right owner C Pr review automation | ai-native user | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | none | 0/10 | ||
Bring my own LLM or API key so agents run on the model of my choice C Model flexibility | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | untested | none yet | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Convert user feedback submissions into structured tasks with proposed scope C Natural language task intake | product-manager | Intent to spec — stories about intent to spec in this arenaIntent to spec | 2 | none | untested | none yet | |
Grant an agent access to my repositories with a one-click install, without complex setup C Version control integration | developer | Repo integration — stories about repo integration in this arenaRepo integration | 2 | none | untested | none yet | |
Have each task prompt automatically routed to the most suitable underlying model C Model control | ai-native user | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 2 | none | untested | none yet | |
Have security alerts automatically validated and remediated with an opened pull request C Security remediation | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 2 | none | untested | none yet | |
License an enterprise deployment with SSO and commercial support for organization-wide rollout G Enterprise licensing | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
See and manage plan-based daily task and concurrency limits for agent workflows G Usage quotas | engineering-lead | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | untested | none yet | |
Self-host agent infrastructure locally, in containers, or on my own VMs C Deployment flexibility | engineering-lead | Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism | 2 | none | untested | none yet | |
Automatically fix failing agent-readiness criteria in my repository C Readiness checks | engineering-lead | Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates | 1 | full | 8/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Cclaimed | |
Approve key agent decisions from my phone while agents continue working C Approval controls | product-manager | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 1 | partial | 3/10 | Cclaimed | |
Query generated documentation for any public or private repository C Knowledge context | developer | Repo integration — stories about repo integration in this arenaRepo integration | 1 | none | untested | none yet | |
Switch away from automatic model selection to a specific model of my choice C Model control | engineering-lead | Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 61 stories with headroom
What would move Factory’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Review quality gates — quality gates on changes — review flow, required checks, merge protectionHave every pull request automatically reviewed with AI-generated inline comments
nonemoves PA Scoreimpact 30
Missing: any documentation of automated PR-triggered review, inline comment generation on PRs, or GitHub/GitLab PR integration specifics.
Repo integration — stories about repo integration in this arenaAdd a context file describing my codebase conventions so agents generate more relevant plans and code
nonemoves PA Scoreimpact 30
The evidence pack does not mention any context file mechanism (e.g., AGENTS.md, .factory config, or similar) for describing codebase conventions to guide agent behavior; it covers CLI usage, integrations, readiness reports, and missions but nothing about persistent repo-convention context files.
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionHave an agent safely execute code and install dependencies inside an isolated sandbox
nonemoves PA Scoreimpact 30
Evidence describes tiered autonomy, bash mode, and CI/CD execution (droid exec) but never mentions an isolated sandbox environment for code execution or dependency installation; no container/VM isolation is documented.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
No evidence addresses data export or portability in open formats, nor any account-deletion/data-takeout mechanism; the docs focus on session management, CLI, and integrations, not exporting user data to leave the platform.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Factory is presented as a cloud-hosted platform (Factory App, Droid CLI connecting to hosted services, API sessions) with no evidence of a self-hostable core server or on-prem deployment option anywhere in the docs or probes.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence in the pack addresses data usage/training opt-out, privacy policy, or data retention controls; the docs focus entirely on product features like CLI, missions, and integrations.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
No evidence of scoped or least-privilege API credential/token issuance for agents; docs mention API sessions and integrations (Jira, Slack, MCP) but nothing about credential scoping, permission tiers for API keys, or least-privilege access control.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
Missing: any webhook documentation, event types, subscription endpoints, or third-party confirmation of webhook support.
Showing the top 8 of 61 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map10 surfaces · 41 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Droid CLI docs24 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Plug MCP servers into this product so it can use their tools
- Connect an agent via an official MCP server
- Use an official CLI
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Version, review, and roll back my automations
- Have an agent autonomously diagnose and fix a reported bug
- Have an agent implement a requested feature end-to-end, including writing tests
- Have an agent automatically clone the repo, install dependencies, and configure its own working environment
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Send follow-up instructions to an active agent session to steer its work without restarting
- Watch what a running agent is doing in real time, including its current status
- Get notified when an agent completes a task or needs my input
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Assign a coding task to an agent directly from an existing issue or ticket
- Tag an agent in a chat thread to discuss and delegate a bug or task
- Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directly
- Connect a GitHub repository so an agent can access the code and open pull requests against it
- Have failed CI workflows automatically diagnosed and fixed with a proposed pull request
- Trigger an agent from CI/CD pipelines to fix a broken build or failing test
- Configure an agent to automatically open a pull request when its task completes
- Run an agent headlessly inside CI/CD pipelines and shell scripts
Droid exec docs24 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Have an agent autonomously diagnose and fix a reported bug
- Have an agent implement a requested feature end-to-end, including writing tests
- Have an agent automatically clone the repo, install dependencies, and configure its own working environment
- Configure an agent to auto-approve all its actions instead of confirming each one
- Approve key agent decisions from my phone while agents continue working
- Set tiered autonomy levels controlling what an agent can do without manual confirmation
- Get notified when an agent completes a task or needs my input
- Review and approve an agent's implementation plan before any code changes are made
- Approve a task's scope and contract before an agent is allowed to modify the repository
- Connect a GitHub repository so an agent can access the code and open pull requests against it
- Have failed CI workflows automatically diagnosed and fixed with a proposed pull request
- Trigger an agent from CI/CD pipelines to fix a broken build or failing test
- Configure an agent to automatically open a pull request when its task completes
- Review a diff of an agent's changes and approve it before it becomes a pull request
- Run many agent tasks concurrently to scale delivery throughput
- Run an agent headlessly inside CI/CD pipelines and shell scripts
docs.factory.ai16 stories
- Delegate tasks to a built-in AI assistant inside the product
- Version, review, and roll back my automations
- Have an agent autonomously diagnose and fix a reported bug
- Go from a mockup or design to a working implementation without an engineering handoff
- Have an agent implement a requested feature end-to-end, including writing tests
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Approve key agent decisions from my phone while agents continue working
- Watch what a running agent is doing in real time, including its current status
- Get notified when an agent completes a task or needs my input
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Review and approve an agent's implementation plan before any code changes are made
- Approve a task's scope and contract before an agent is allowed to modify the repository
- Connect a GitHub repository so an agent can access the code and open pull requests against it
- Configure an agent to automatically open a pull request when its task completes
- Review a diff of an agent's changes and approve it before it becomes a pull request
- Use a managed cloud offering to run agents without operating my own backend infrastructure
API reference11 stories
- Drive the product through a documented public API
- Build against official SDKs
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Perform bulk operations across many items at once
- Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the work
- Send follow-up instructions to an active agent session to steer its work without restarting
- Watch what a running agent is doing in real time, including its current status
- Do everything through the API that I can do in the UI
- Run many agent tasks concurrently to scale delivery throughput
- Use a managed cloud offering to run agents without operating my own backend infrastructure
Missions docs9 stories
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Have an agent implement a requested feature end-to-end, including writing tests
- Send follow-up instructions to an active agent session to steer its work without restarting
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Review and approve an agent's implementation plan before any code changes are made
- Do everything through the API that I can do in the UI
- Run many agent tasks concurrently to scale delivery throughput
Agent readiness docs8 stories
- Point an agent at llms.txt or agent-oriented docs
- Have an agent autonomously diagnose and fix a reported bug
- Go from a mockup or design to a working implementation without an engineering handoff
- Describe a feature or bug in plain language and have it automatically turned into a scoped implementation task
- Do everything through the API that I can do in the UI
- Automatically fix failing agent-readiness criteria in my repository
- Run a readiness report that evaluates how ready my repository is for autonomous agents
- Use a managed cloud offering to run agents without operating my own backend infrastructure
OpenAPI spec3 stories
Software factory docs3 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
2 of 10 testable claims verified · 0 contradicted → integrity 20/100
15 distinct capability claims found in Factory’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
2
Verified
8
Unverified
0
Contradicted
31
Undersold
Verified (2)
“Official Droid CLI runs from terminal, editor, tests, and Git workflow”
“API/SDK support for creating and managing Droid sessions including lifecycle, settings, and messages”
Drive the product through a documented public APIpartialproof ↗
Unverified (8)
“Delegate a task, review the diff, and merge from the app or terminal”
Review a diff of an agent's changes and approve it before it becomes a pull requestpartialproof ↗
“Connects to Jira, Notion, Slack, Linear, PagerDuty and MCP tools to sync with team systems”
Connect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directlypartialproof ↗
“Supports plugging in MCP tools alongside team systems for agent use”
Plug MCP servers into this product so it can use their toolspartialproof ↗
“droid exec runs as a one-shot command ideal for CI/CD pipelines, shell scripts, and batch processing”
Run an agent headlessly inside CI/CD pipelines and shell scriptsfullproof ↗
“droid exec uses tiered autonomy to control which operations run without manual confirmation”
Set tiered autonomy levels controlling what an agent can do without manual confirmationpartialproof ↗
“A designer can share a mockup and have the system implement it without an engineering handoff”
Go from a mockup or design to a working implementation without an engineering handofffullproof ↗
“Run a /readiness-report command to evaluate a repository's readiness level for agents”
Run a readiness report that evaluates how ready my repository is for autonomous agentsfullproof ↗
“Automatically fix failing readiness criteria with the /readiness-fix command”
Automatically fix failing agent-readiness criteria in my repositoryfullproof ↗
Undersold (31)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
Operate the product with natural-language commandsfullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Have an agent autonomously diagnose and fix a reported bugpartialproof ↗
Have an agent implement a requested feature end-to-end, including writing testspartialproof ↗
Have an agent automatically clone the repo, install dependencies, and configure its own working environmentpartialproof ↗
Take over an in-progress agent task in my editor, terminal, or browser to finish or redirect the workpartialproof ↗
Send follow-up instructions to an active agent session to steer its work without restartingpartialproof ↗
Configure an agent to auto-approve all its actions instead of confirming each onepartialproof ↗
Approve key agent decisions from my phone while agents continue workingpartialproof ↗
Watch what a running agent is doing in real time, including its current statuspartialproof ↗
Get notified when an agent completes a task or needs my inputpartialproof ↗
Describe a feature or bug in plain language and have it automatically turned into a scoped implementation taskpartialproof ↗
Review and approve an agent's implementation plan before any code changes are madepartialproof ↗
Approve a task's scope and contract before an agent is allowed to modify the repositorypartialproof ↗
Assign a coding task to an agent directly from an existing issue or ticketpartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Tag an agent in a chat thread to discuss and delegate a bug or taskpartialproof ↗
Connect a GitHub repository so an agent can access the code and open pull requests against itpartialproof ↗
Have failed CI workflows automatically diagnosed and fixed with a proposed pull requestpartialproof ↗
Trigger an agent from CI/CD pipelines to fix a broken build or failing testfullproof ↗
Configure an agent to automatically open a pull request when its task completespartialproof ↗
Run many agent tasks concurrently to scale delivery throughputpartialproof ↗
Use a managed cloud offering to run agents without operating my own backend infrastructurefullproof ↗
Claims outside our story set (5)
Real capability claims found in Factory’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Toggle bash mode with ! to run raw shell commands directly, bypassing AI interpretation”
source ↗“Delegate scoped tasks to Custom Droids via /droids command”
source ↗“Package reusable procedures as Skills via /skills or /create-skill”
source ↗“Factory Missions allow planning and executing large, multi-feature projects with structured orchestration”
source ↗“Software Factory view shows the delivery lifecycle as an automation coverage map”
source ↗
Business model
Free trial credits, then Pro/Plus/Max individual plans or Teams/Business/Enterprise org plans, plus usage-based model spend.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
