Software Factory — procurement report
ProductArena · rankings as of 2026-09-16 · evidence as of 2026-09-15 · 9 products · 73 judged requirements · 657 judged cells
Methodology: Every product is judged against a shared taxonomy of user stories using cited evidence — hands-on probes > repository code > independent community sources > vendor claims — never opinion. Full writeup: https://ultrametric.ai/productarena/methodology
Leaderboard
| # | Product | PA Score | Coverage score | Applicable cells | Confidence |
|---|---|---|---|---|---|
| 1 | OpenHands | 50.3 | 36.8 | 72/73 | B |
| 2 | Omnara | 41.4 | 29.1 | 63/73 | B |
| 3 | Devin | 39.1 | 37.9 | 73/73 | C |
| 4 | Codegen | 31.7 | 32.7 | 72/73 | C |
| 5 | YYLO | 28.9 | 17.7 | 62/73 | A |
| 6 | Factory | 26.8 | 25.1 | 72/73 | C |
| 7 | Jules | 22.6 | 30.5 | 72/73 | C |
| 8 | HumanLayer | 19.0 | 23.7 | 71/73 | C |
| 9 | Foreloop | 18.5 | 21.9 | 66/73 | C |
PA Score = agent-readiness blend (see methodology). Coverage score = weighted share of judged requirements met. Confidence = how much of the score rests on tested vs claimed evidence (A–D).
Uncertainty note
The current #1/#2 gap in this arena is not close enough to qualify for the multi-judge uncertainty pass (or the pass has not covered it yet) — no extra caveat applies beyond the per-product confidence grades above.
Buyer checklist (RFP)
The arena's 73 judged user stories as requirements, grouped by theme. Priorities mirror the story weights our scoring uses (3 = must-have, 2 = should-have, 1 = nice-to-have). Interactive version with per-requirement verdicts for the top products: /arena/software-factory/checklist
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
- ai-native userPlug MCP servers into this product so it can use their toolsmust-have
- ai-native userConnect an agent via an official MCP servermust-have
- ai-native userDrive the product through a documented public APImust-have
- ai-native userDelegate tasks to a built-in AI assistant inside the productmust-have
- ai-native userPoint an agent at llms.txt or agent-oriented docsshould-have
- ai-native userRun the product headlessly / in CI for automationshould-have
- ai-native userUse an official CLIshould-have
- ai-native userIssue scoped/least-privilege API credentials for an agentshould-have
- ai-native userBuild against official SDKsshould-have
- ai-native userSubscribe to events via webhooksshould-have
- ai-native userGet AI-generated insights and suggestions from my data inside the productshould-have
- ai-native userSet up automations that run autonomously in the backgroundshould-have
- ai-native userOperate the product with natural-language commandsshould-have
- ai-native userExplore an interactive API reference with runnable examplesshould-have
- ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)should-have
- ai-native userRely on versioned APIs with a documented deprecation policyshould-have
- ai-native userTest against a sandbox environment without touching production datanice-to-have
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
- ai-native userDefine rules that trigger actions automatically on eventsmust-have
- ai-native userPerform bulk operations across many items at onceshould-have
- ai-native userSchedule recurring jobs or workflowsshould-have
- ai-native userVersion, review, and roll back my automationsnice-to-have
Autonomous implementation — end-to-end implementation by the agent — multi-file changes, task completionAutonomous implementation
End-to-end implementation by the agent — multi-file changes, task completion
- developerHave an agent autonomously diagnose and fix a reported bugmust-have
- developerHave an agent implement a requested feature end-to-end, including writing testsmust-have
- developerHave an agent safely execute code and install dependencies inside an isolated sandboxmust-have
- ai-native userHave an agent automatically generate and run tests to validate its own code changes before proposing themshould-have
- product-managerGo from a mockup or design to a working implementation without an engineering handoffshould-have
- developerHave an agent automatically clone the repo, install dependencies, and configure its own working environmentshould-have
- developerTake over an in-progress agent task in my editor, terminal, or browser to finish or redirect the workshould-have
- developerSend follow-up instructions to an active agent session to steer its work without restartingshould-have
Human oversight — keeping a human in the loop — approvals, checkpoints, interruptsHuman oversight
Keeping a human in the loop — approvals, checkpoints, interrupts
- engineering-leadSet tiered autonomy levels controlling what an agent can do without manual confirmationmust-have
- developerWatch what a running agent is doing in real time, including its current statusmust-have
- developerConfigure an agent to auto-approve all its actions instead of confirming each oneshould-have
- ai-native userHave each task prompt automatically routed to the most suitable underlying modelshould-have
- developerGet notified when an agent completes a task or needs my inputshould-have
- product-managerApprove key agent decisions from my phone while agents continue workingnice-to-have
- engineering-leadSwitch away from automatic model selection to a specific model of my choicenice-to-have
Intent to spec — stories about intent to spec in this arenaIntent to spec
Stories about intent to spec in this arena
- developerDescribe a feature or bug in plain language and have it automatically turned into a scoped implementation taskmust-have
- developerReview and approve an agent's implementation plan before any code changes are mademust-have
- developerAssign a coding task to an agent directly from an existing issue or ticketmust-have
- product-managerConvert user feedback submissions into structured tasks with proposed scopeshould-have
- developerAttach a marked-up screenshot or mockup to a task so the agent implements the correct visual changeshould-have
- engineering-leadApprove a task's scope and contract before an agent is allowed to modify the repositoryshould-have
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
- ai-native userExport all of my data in open formats and leavemust-have
- ai-native userSelf-host the core productmust-have
- ai-native userDo everything through the API that I can do in the UIshould-have
- ai-native userRead the product's source under an open licenseshould-have
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
- engineering-leadLicense an enterprise deployment with SSO and commercial support for organization-wide rolloutshould-have
- engineering-leadBring my own LLM or API key so agents run on the model of my choiceshould-have
- engineering-leadSee and manage plan-based daily task and concurrency limits for agent workflowsshould-have
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
- ai-native userPrevent my data from being used to train AI modelsmust-have
- ai-native userChoose where my data is stored (region/residency)should-have
- ai-native userControl data retention and deletionshould-have
- ai-native userOpt out of telemetry and usage trackingshould-have
Repo integration — stories about repo integration in this arenaRepo integration
Stories about repo integration in this arena
- developerAdd a context file describing my codebase conventions so agents generate more relevant plans and codemust-have
- product-managerConnect issue trackers like Jira, Linear, ClickUp, or Monday.com so agents can manage tickets directlymust-have
- developerConnect a GitHub repository so an agent can access the code and open pull requests against itmust-have
- developerTag an agent in a chat thread to discuss and delegate a bug or taskshould-have
- developerGrant an agent access to my repositories with a one-click install, without complex setupshould-have
- developerQuery generated documentation for any public or private repositorynice-to-have
Review quality gates — quality gates on changes — review flow, required checks, merge protectionReview quality gates
Quality gates on changes — review flow, required checks, merge protection
- engineering-leadHave failed CI workflows automatically diagnosed and fixed with a proposed pull requestmust-have
- developerReview a diff of an agent's changes and approve it before it becomes a pull requestmust-have
- engineering-leadHave every pull request automatically reviewed with AI-generated inline commentsmust-have
- developerTrigger an agent from CI/CD pipelines to fix a broken build or failing testshould-have
- developerConfigure an agent to automatically open a pull request when its task completesshould-have
- ai-native userHave incoming issues automatically triaged with severity suggested and routed to the right ownershould-have
- engineering-leadRun a readiness report that evaluates how ready my repository is for autonomous agentsshould-have
- engineering-leadHave security alerts automatically validated and remediated with an opened pull requestshould-have
- engineering-leadAutomatically fix failing agent-readiness criteria in my repositorynice-to-have
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
- engineering-leadRun many agent tasks concurrently to scale delivery throughputmust-have
- engineering-leadCreate agent sessions on behalf of other users in my organizationshould-have
- developerUse a managed cloud offering to run agents without operating my own backend infrastructureshould-have
- engineering-leadSelf-host agent infrastructure locally, in containers, or on my own VMsshould-have
- developerRun an agent headlessly inside CI/CD pipelines and shell scriptsshould-have
Appendix: recorded probes
Hands-on probe recordings — transcripts/videos a human can replay, the strongest evidence tier. Watch them at https://ultrametric.ai/productarena/proofs
- Omnara
curl -s https://docs.omnara.com/introduction.md | head -6terminal · recorded 2026-09-10 · exit 0 - Omnara
curl -sL https://docs.omnara.com/llms.txt | head -6terminal · recorded 2026-09-10 · exit 0 - YYLO
npm view @yylo/cli name version binterminal · recorded 2026-09-14 · exit 0 - YYLO
curl -s https://raw.githubusercontent.com/yylo-dev/yylo/HEAD/LICENSE | head -3terminal · recorded 2026-09-14 · exit 0
Cite as: ProductArena by Ultrametric Inc, Software Factory arena, rankings as of 2026-09-16 — https://ultrametric.ai/productarena/arena/software-factory
License: © 2026 Ultrametric Inc. Brief quotation of individual verdicts, scores, or evidence excerpts is permitted with attribution to "ProductArena by Ultrametric Inc (ultrametric.ai/productarena)", as is use of the data to evaluate, contest, or contribute corrections. Bulk copying, redistribution, or use to build competing datasets requires prior written permission (see DATA-LICENSE in the repository).
No liability: rankings, verdicts, and scores are research outputs derived from the cited evidence at a point in time, provided "as is", without warranties. Ultrametric Inc accepts no responsibility for procurement, purchasing, or other decisions made in reliance on them — verify against the cited evidence before acting (https://ultrametric.ai/productarena/terms).