Testing pipeline — coverage and gapsTesting pipeline
What we have not tested
Every ranking on this site is built from per-cell verdicts — and 8,330 of 29,133 cells (28.6%) are still untested: a zero-evidence none/na where we found nothing pro or con and never probed it. Those cells can’t score — they read as unknown, never as 0 — and this page is the standing list of them. 14.1% of all cells are backed by a hands-on probe.
The flip side of this page — which products’ verdicts rest on the MOST tested evidence — is its own global ranking: most tested →
cells judgedcells judged
29,133
83 arenas
still untestedstill untested
28.6%
8,330 zero-evidence cells
probed hands-onprobed hands-on
14.1%
4,114 cells cite a probe
Pricing coverage: 41 of 63 products in the 9 pricing-covered arenas have unit pricing extracted verbatim from the vendor’s own pricing page (7 record an honest “pricing unclear” — JS-shell or quote-only pages we refuse to guess at). Every figure carries its source URL, exact quote, and fetch date; we never compute a price we didn’t extract.
Most-wanted untested cells — the highest-impact gapsMost-wanted untested cells
The ten untested (product, story) pairs whose testing would move the most-read scores the most — heaviest stories on the most-watched products (capped at two per product so one giant can’t fill the board). Have first-hand evidence for one? Send it in.
| Product | Untested story | Arena | Weight | GitHub ★ |
|---|---|---|---|---|
| Connect an agent via an official MCP server | Agent Skills & Extensions | ×3 | 286,309 | |
| Drive the product through a documented public API | Agent Skills & Extensions | ×3 | 286,309 | |
| Delegate tasks to a built-in AI assistant inside the product | Agent Skills & Extensions | ×3 | 261,468 | |
| Drive the product through a documented public API | Agent Skills & Extensions | ×3 | 261,468 | |
| Delegate tasks to a built-in AI assistant inside the product | Frontend Frameworks | ×3 | 250,422 | |
| Plug MCP servers into this product so it can use their tools | Frontend Frameworks | ×3 | 250,422 | |
| Delegate tasks to a built-in AI assistant inside the product | Workflow Automation | ×3 | 204,216 | |
| Define rules that trigger actions automatically on events | Local LLM Runtimes | ×3 | 180,849 | |
| The documented maximum concurrent requests or connections the local server can handle before throughput degrades | Local LLM Runtimes | ×3 | 180,849 | |
| Whether exceeding my plan's monthly credit or request quota triggers overage charges or a hard cutoff | Web Scraping APIs | ×3 | 180,429 |
Coverage per arena — untested and probed sharesCoverage per arena
Sorted worst-first: the arenas with the largest untested share are where the rankings deserve the most skepticism — and the most contributed evidence.
Agent surface health — 6-hourly keyless uptime checksAgent surface health
Every 6 hours we keylessly ping each product’s documented agent surfaces — its llms.txt, remote MCP endpoint, and openapi.json where we previously found one. An auth-gated MCP endpoint answering 401 counts as up; only timeouts, 404/410 and 5xx count as down. Currently monitoring 387 surfaces across 261 products (248 llms.txt · 98 MCP · 41 openapi.json), tracking since Sep 8 '26 — uptime percentages appear on product pages after a week of history.
| Product | Arena | Surface | URL | Down since | Last status |
|---|---|---|---|---|---|
| API platforms | llms.txt | https://kong.com/llms.txt | Sep 11 '26 | timeout |
Next up — tier-1 arenas awaiting the pipelineNext up
Tier-1 arenas on the roadmap that haven’t been through the evidence pipeline yet — the categories we think matter most in an agent-first world, in no particular order.
Autonomous Coding Agents
Unsupervised task completion rate, PR quality, and cost-per-merged-change.
Devin · OpenAI Codex cloud · Claude Code on the web · Google Jules · GitHub Copilot coding agent
Frontier Model APIs
Tool-use reliability, context economics, caching, rate-limit reality vs published numbers.
Anthropic Claude API · OpenAI API · Google Gemini API · xAI Grok API · Mistral La Plateforme
Open-Weight LLMs
License openness, tool-calling quality, quantization ecosystem, hosting breadth.
Llama 4 · DeepSeek R1 · Qwen3 · Kimi K2 · GLM-4.6
Enterprise AI Search & Work Assistants
Connector breadth, permission-aware retrieval, and whether external agents can query it.
Glean · Dust · Onyx · Microsoft 365 Copilot · Notion AI
AI Search Visibility (GEO/AEO)
Measurement methodology transparency in a category about being seen by AI.
Profound · Peec AI · Scrunch AI · Otterly.AI · AthenaHQ
Browsers (AI Era)
Built-in agents, page-context APIs, and the privacy cost of an AI that sees every tab.
Chrome · Comet · Dia · Brave · Edge
Search Engines
AI-answer quality vs link quality, index independence, API access for agents.
Google · Bing · Kagi · DuckDuckGo · Brave Search
Managed Postgres
Instant branch-per-agent databases and provisioning APIs built for generated apps.
Neon · Supabase · PlanetScale for Postgres · AWS Aurora · TigerData
Untested is a per-cell status: a none/na verdict that cites zero evidence (the same definition the tables use to render “untested” instead of 0 — see methodology). “Probed” is stricter than “evidenced”: only verdicts citing a hands-on probe count.