Skip to content

Testing pipeline — coverage and gapsTesting pipeline

What we have not tested

Every ranking on this site is built from per-cell verdicts — and 8,330 of 29,133 cells (28.6%) are still untested: a zero-evidence none/na where we found nothing pro or con and never probed it. Those cells can’t score — they read as unknown, never as 0 — and this page is the standing list of them. 14.1% of all cells are backed by a hands-on probe.

The flip side of this page — which products’ verdicts rest on the MOST tested evidence — is its own global ranking: most tested →

cells judgedcells judged

29,133

83 arenas

still untestedstill untested

28.6%

8,330 zero-evidence cells

probed hands-onprobed hands-on

14.1%

4,114 cells cite a probe

Pricing coverage: 41 of 63 products in the 9 pricing-covered arenas have unit pricing extracted verbatim from the vendor’s own pricing page (7 record an honest “pricing unclear” — JS-shell or quote-only pages we refuse to guess at). Every figure carries its source URL, exact quote, and fetch date; we never compute a price we didn’t extract.

Most-wanted untested cells — the highest-impact gapsMost-wanted untested cells

The ten untested (product, story) pairs whose testing would move the most-read scores the most — heaviest stories on the most-watched products (capped at two per product so one giant can’t fill the board). Have first-hand evidence for one? Send it in.

ProductUntested storyWeight
Superpowers logoSuperpowersConnect an agent via an official MCP server×3
Superpowers logoSuperpowersDrive the product through a documented public API×3
Skills for Real Engineers logoSkills for Real EngineersDelegate tasks to a built-in AI assistant inside the product×3
Skills for Real Engineers logoSkills for Real EngineersDrive the product through a documented public API×3
React logoReactDelegate tasks to a built-in AI assistant inside the product×3
React logoReactPlug MCP servers into this product so it can use their tools×3
n8n logon8nDelegate tasks to a built-in AI assistant inside the product×3
Ollama logoOllamaDefine rules that trigger actions automatically on events×3
Ollama logoOllamaThe documented maximum concurrent requests or connections the local server can handle before throughput degrades×3
Firecrawl logoFirecrawlWhether exceeding my plan's monthly credit or request quota triggers overage charges or a hard cutoff×3

Coverage per arena — untested and probed sharesCoverage per arena

Sorted worst-first: the arenas with the largest untested share are where the rankings deserve the most skepticism — and the most contributed evidence.

ArenaCellsUntestedProbed
Processors352232 (65.9%)4.8%
GPUs & AI Accelerators344212 (61.6%)8.7%
Mobile AI Dev Tools414244 (58.9%)2.2%
Frontend Frameworks425235 (55.3%)2.4%
Desktop OS390197 (50.5%)2.6%
Startup Legal & Incorporation378178 (47.1%)8.2%
Banking as a Service318143 (45%)13.8%
Payment Fraud Prevention270119 (44.1%)8.9%
Hardware Security Keys324143 (44.1%)16%
Package & Toolchain Managers324140 (43.2%)5.2%
Marketplace & Platform Payments318127 (39.9%)13.2%
Web Scraping APIs768303 (39.5%)9.9%
Robotics Software Platforms270102 (37.8%)5.9%
Terminals324122 (37.7%)12.3%
Authenticator Apps440163 (37%)13.9%
Project Management498183 (36.7%)7.8%
Identity Verification & KYC318116 (36.5%)11.9%
Startup Banking370130 (35.1%)13.8%
Local LLM Runtimes644219 (34%)4%
AI Research Agents25887 (33.7%)6.6%
Game Engines448150 (33.5%)26.8%
Stablecoin Payments318104 (32.7%)18.9%
Card Issuing Platforms26586 (32.5%)16.2%
Cap Table & Equity336109 (32.4%)6.8%
Product Feedback & Intent19562 (31.8%)13.3%
Team Chat25580 (31.4%)16.5%
Payroll & HR Ops21664 (29.6%)8.3%
AI Search APIs25074 (29.6%)16.8%
AI Inference Providers371110 (29.6%)19.4%
Email Clients & Email AI32496 (29.6%)12%
AI Assistants468138 (29.5%)7.3%
MCP Infrastructure & Registries31292 (29.5%)27.9%
Mobile & In-Person Payments20861 (29.3%)15.4%
Agent Skills & Extensions26076 (29.2%)4.2%
Model Gateways & Routers35098 (28%)14%
Banking Data APIs371103 (27.8%)16.2%
Sales Tax Automation31888 (27.7%)17%
Edge & App Platforms576156 (27.1%)10.6%
Billing & Subscriptions371100 (27%)18.9%
Agentic Commerce392106 (27%)26.3%
Voice Agent Platforms549146 (26.6%)21.3%
Email Marketing34892 (26.4%)16.7%
Vector Databases & Memory Stores37197 (26.1%)14.8%
Security Scanners33085 (25.8%)14.2%
Infrastructure as Code21654 (25%)28.2%
Agent Sandboxes & Code Execution408100 (24.5%)16.9%
Design & Prototyping424102 (24.1%)15.3%
Expense Management26563 (23.8%)15.5%
Online Payments540128 (23.7%)12.2%
Code Hosting28867 (23.3%)6.9%
Meeting AI & Notetakers27062 (23%)26.7%
Vibe-Coding App Builders31271 (22.8%)11.2%
Developer Docs Platforms26560 (22.6%)16.2%
Scheduling & Calendar27562 (22.5%)9.1%
Software Factory657147 (22.4%)11%
Accounting & Bookkeeping583129 (22.1%)10.6%
Search Infrastructure26558 (21.9%)15.1%
GPU Clouds27058 (21.5%)24.4%
Browser Automation for Agents36478 (21.4%)17.6%
Serverless & Developer Databases28057 (20.4%)30%
Agent Frameworks & SDKs45991 (19.8%)8.3%
API platforms28556 (19.6%)15.1%
Workflow Automation33666 (19.6%)18.2%
E-commerce Platforms27053 (19.6%)15.9%
Auth & Identity28054 (19.3%)13.6%
AI Memory Layers34266 (19.3%)19.9%
AI Coding Agents962183 (19%)7.7%
AI Customer Support Agents31860 (18.9%)10.7%
Error Tracking33060 (18.2%)20.6%
Incident Management & On-call27048 (17.8%)18.1%
LLM Evals & Observability41673 (17.5%)17.1%
Product Analytics21236 (17%)18.4%
CRM20434 (16.7%)17.6%
Document Extraction APIs31852 (16.4%)11%
Customer Data Platforms25541 (16.1%)26.3%
Backend as a Service20431 (15.2%)23%
Observability & Monitoring32448 (14.8%)24.1%
Feature Flags & Experimentation27037 (13.7%)31.5%
Notes & Knowledge Bases33045 (13.6%)19.4%
Data Warehouses & Lakehouses21628 (13%)24.5%
Data Pipelines & ELT26532 (12.1%)31.7%
Durable Execution Engines31836 (11.3%)11.3%
AI Code Review31836 (11.3%)14.2%

Agent surface health — 6-hourly keyless uptime checksAgent surface health

Every 6 hours we keylessly ping each product’s documented agent surfaces — its llms.txt, remote MCP endpoint, and openapi.json where we previously found one. An auth-gated MCP endpoint answering 401 counts as up; only timeouts, 404/410 and 5xx count as down. Currently monitoring 387 surfaces across 261 products (248 llms.txt · 98 MCP · 41 openapi.json), tracking since Sep 8 '26 — uptime percentages appear on product pages after a week of history.

ProductSurfaceDown sinceLast status
Kong logoKongllms.txtSep 11 '26timeout

Next up — tier-1 arenas awaiting the pipelineNext up

Tier-1 arenas on the roadmap that haven’t been through the evidence pipeline yet — the categories we think matter most in an agent-first world, in no particular order.

Autonomous Coding Agents

Unsupervised task completion rate, PR quality, and cost-per-merged-change.

Devin · OpenAI Codex cloud · Claude Code on the web · Google Jules · GitHub Copilot coding agent

Frontier Model APIs

Tool-use reliability, context economics, caching, rate-limit reality vs published numbers.

Anthropic Claude API · OpenAI API · Google Gemini API · xAI Grok API · Mistral La Plateforme

Open-Weight LLMs

License openness, tool-calling quality, quantization ecosystem, hosting breadth.

Llama 4 · DeepSeek R1 · Qwen3 · Kimi K2 · GLM-4.6

Enterprise AI Search & Work Assistants

Connector breadth, permission-aware retrieval, and whether external agents can query it.

Glean · Dust · Onyx · Microsoft 365 Copilot · Notion AI

AI Search Visibility (GEO/AEO)

Measurement methodology transparency in a category about being seen by AI.

Profound · Peec AI · Scrunch AI · Otterly.AI · AthenaHQ

Browsers (AI Era)

Built-in agents, page-context APIs, and the privacy cost of an AI that sees every tab.

Chrome · Comet · Dia · Brave · Edge

Search Engines

AI-answer quality vs link quality, index independence, API access for agents.

Google · Bing · Kagi · DuckDuckGo · Brave Search

Managed Postgres

Instant branch-per-agent databases and provisioning APIs built for generated apps.

Neon · Supabase · PlanetScale for Postgres · AWS Aurora · TigerData

Untested is a per-cell status: a none/na verdict that cites zero evidence (the same definition the tables use to render “untested” instead of 0 — see methodology). “Probed” is stricter than “evidenced”: only verdicts citing a hands-on probe count.