AI Inference Providers arenaAI Inference Providers
Hosted inference clouds serving open-weight and frontier models over OpenAI-compatible APIs, judged on model catalog breadth, measured speed claims, structured-output and tool-calling fidelity, batch and fine-tune serving, dedicated capacity, pricing and rate-limit clarity, and how completely an AI agent can discover models and drive inference end to end.
53 user stories · 371 judged cells · updated 2026-09-14 · Evidence as of 2026-09-14
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | free-tier vs Fireworks AI ↗ | 52/100 | 10/100 | 7/100 | 10/100 | 70/100 | ★ 265▲ 103/yrnpm 673.9k/wk | unclear | 16/28 verified · 3 disputed | 54/100 integrity | ||
| 2 | usage-based vs Groq ↗ | 33/100 | 13/100 | 0/100 | 21/100 | 70/100 | pypi 267.8k/wk | $0.008/ 1M input tokens | 9/29 verified | 6/100 integrity | ||
| 3 | free-tier vs Groq ↗ | 36/100 | 5/100 | 12/100 | 10/100 | 70/100 | ★ 87▲ 41/yrnpm 108k/wk | free tier | 19/29 verified | 57/100 integrity | ||
| 4 | usage-based vs Groq ↗ | 54/100 | 10/100 | 9/100 | 26/100 | 9/100 | ★ 1.2k▲ 287/yr | $0.1/ 1M input tokens | 14/35 verified | 25/100 integrity | ||
| 5 | usage-based vs Groq ↗ | 50/100 | untested | 0/100 | 12/100 | 35/100 | ★ 71▲ 30/yrnpm 71.5k/wk | $0.3/ 1M input tokens | 12/27 verified · 1 disputed | 28/100 integrity | ||
| 6 | Morph usage-based vs Groq ↗ | 43/100 | 0/100 | 0/100 | 5/100 | 15/100 | ★ 6▲ 5/yr | 15/22 verified | 40/100 integrity | |||
| 7 | usage-based vs Groq ↗ | 26/100 | 0/100 | 0/100 | 10/100 | 0/100 | npm 1.6k/wk | $0.005/ 1M input tokens | 12/21 verified | 33/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Story matrix — every product × every judged storyStory matrix
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 9/10 | fullT 9/10 | fullT 9/10 | fullT 9/10 | none 0/10 | fullT 9/10 | fullT 9/10 |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | fullT⚿ 8/10 | fullT⚿ 7/10 | fullT⚿ 7/10 | fullX 7/10 | fullT 7/10 | fullC 7/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | fullT 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | partialT 4/10 | fullT 8/10 | none 0/10 | none 0/10 | none 0/10 | fullT 8/10 | fullT 7/10 |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | none 0/10 | partialC 5/10 | partialC 4/10 | none 0/10 | none 0/10 | fullC 7/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | fullT⚿ 9/10 | fullT⚿ 9/10 | fullT⚿ 9/10 | fullT⚿ 9/10 | fullT 8/10 | fullT⚿ 8/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | none 0/10 | none⚿ 0/10 | none⚿ 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | fullC 8/10 | fullC 8/10 | partialC 5/10 | fullX 8/10 | fullT 8/10 | fullC 7/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 |
Agentic features
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | n/a | n/a | n/a | n/a | n/a | n/a | n/a |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | none 0/10 | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | partialC 6/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | partialT 6/10 | n/a |
Api quality
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | partialT 4/10 | none 0/10 | none⚿ 0/10 | partialT 5/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 | partialC 4/10 | n/a |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialT 3/10 | none 0/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | fullC 7/10 | fullC 7/10 | fullC 7/10 | fullC 7/10 | none 0/10 | partialC 5/10 | partialC 4/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | n/a | n/a | n/a | n/a | n/a | none 0/10 | partialC 4/10 |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | n/a | none 0/10 | n/a | n/a | n/a | none 0/10 | none 0/10 |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | n/a | n/a | n/a | n/a | n/a | n/a | none 0/10 |
Batch async — stories about batch async in this arenaBatch async
Batch
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Batch async — stories about batch async in this arenaSubmit asynchronous batch inference jobs at a documented discount versus real-time pricing | ml-engineer | fullC 8/10 | fullC 8/10 | fullC 8/10 | partialC 4/10 | partialC 4/10 | partialC 4/10 | partialC 6/10 |
Dedicated capacity — stories about dedicated capacity in this arenaDedicated capacity
Dedicated
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Dedicated capacity — stories about dedicated capacity in this arenaDeploy a model on dedicated GPU capacity with autoscaling so my traffic is isolated from the shared serverless pool | ml-engineer | none 0/10 | fullC 8/10 | fullC 8/10 | partialC 5/10 | fullC 7/10 | partialC 5/10 | partialC 4/10 |
Fine tune serving — stories about fine tune serving in this arenaFine tune serving
Fine tune
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Fine tune serving — stories about fine tune serving in this arenaFine-tune a supported base model on my own data and serve the result on the same platform | ml-engineer | none 0/10 | fullC 9/10 | fullC 8/10 | partialC 4/10 | partialC 4/10 | fullC 8/10 | partialC 4/10 |
| Fine tune serving — stories about fine tune serving in this arenaUpload and serve my own custom model weights or LoRA adapters | ml-engineer | partialC 6/10 | partialC 7/10 | fullC 9/10 | partialC 4/10 | partialC 6/10 | fullC 8/10 | none 0/10 |
Model catalog — stories about model catalog in this arenaModel catalog
Catalog
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Model catalog — stories about model catalog in this arenaGet newly released open-weight models on the platform quickly after their public release | ml-engineer | partialX 5/10 | none 0/10 | partialC 5/10 | partialX 3/10 | none 0/10 | none 0/10 | none 0/10 |
| Model catalog — stories about model catalog in this arenaRely on a documented deprecation policy with advance notice before a hosted model is removed | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 |
| Model catalog — stories about model catalog in this arenaHave an agent enumerate the live model catalog programmatically via a documented GET /v1/models-style endpoint | ai-native | fullT⚿ 8/10 | fullT⚿ 8/10 | fullT⚿ 7/10 | partialT⚿ 7/10 | fullT 8/10 | fullT⚿ 8/10 | none 0/10 |
| Model catalog — stories about model catalog in this arenaChoose among a broad catalog of current open-weight model families (Llama, Qwen, DeepSeek, GPT-OSS and peers) on shared serverless endpoints | developer | partialT⚿ 6/10 | fullT⚿ 7/10 | fullT⚿ 8/10 | partialX 7/10 | partialT 7/10 | partialT⚿ 4/10 | partialC 5/10 |
Multimodal — stories about multimodal in this arenaMultimodal
Modalities
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Multimodal — stories about multimodal in this arenaGenerate embeddings (and rerank results) for retrieval pipelines without a second vendor | developer | none 0/10 | partialC 5/10 | fullC 7/10 | none 0/10 | fullC 8/10 | partialX 5/10 | none 0/10 |
| Multimodal — stories about multimodal in this arenaCall vision, audio, or image-generation models beyond text chat on the same platform | developer | partialC 6/10 | fullC 8/10 | fullC 8/10 | partialC 3/10 | fullC 8/10 | partialC 5/10 | none 0/10 |
Openai compat — stories about openai compat in this arenaOpenai compat
Compat
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Openai compat — stories about openai compat in this arenaPlug the provider into coding agents and agent frameworks through documented, first-party integration guides | ai-native | partialT 7/10 | fullT 9/10 | partialC 4/10 | partialX 6/10 | partialT 5/10 | fullT 9/10 | fullT 8/10 |
| Openai compat — stories about openai compat in this arenaHave an agent switch to or away from this provider mid-workflow because it speaks the standard chat-completions API without provider-specific code changes | ai-native | fullT⚿ 8/10 | fullT⚿ 8/10 | fullT⚿ 8/10 | partialX 7/10 | fullT 9/10 | fullC 8/10 | fullT 8/10 |
| Openai compat — stories about openai compat in this arenaPoint an existing OpenAI SDK client at the provider by changing only the base URL and API key | developer | fullT⚿ 9/10 | fullT⚿ 8/10 | fullT⚿ 9/10 | fullT⚿ 9/10 | fullT 9/10 | fullT⚿ 9/10 | fullT 8/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialT 6/10 | partialT 7/10 | partialC 6/10 | partialX 6/10 | partialT 6/10 | partialT 6/10 | partialT 4/10 |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | n/a | n/a | n/a | n/a | n/a | n/a | none 0/10 |
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Limits
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payRead documented rate limits and how they scale across usage tiers before I hit them in production | developer | partialC 5/10 | partialC 6/10 | none 0/10 | partialX 6/10 | none 0/10 | partialC 6/10 | none 0/10 |
Pricing
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to paySet spending caps or budget alerts so a runaway workload cannot generate an unbounded bill | founder | fullC 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 5/10 | none 0/10 |
| Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to paySee public per-token prices for every hosted model without talking to sales | founder | disputedD 4/10 | partialT⚿ 6/10 | partialX 5/10 | partialX 4/10 | partialT 5/10 | partialT⚿ 3/10 | none 0/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 | partialX 5/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 5/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialX 3/10 |
Reliability status — stories about reliability status in this arenaReliability status
Reliability
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Reliability status — stories about reliability status in this arenaGet a stated availability SLA on paid or enterprise tiers | founder | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Reliability status — stories about reliability status in this arenaCheck a public status page with incident history before betting production traffic on the platform | founder | partialT 5/10 | fullT 8/10 | partialT 6/10 | partialT 6/10 | partialT 6/10 | partialT 6/10 | none 0/10 |
Speed latency — stories about speed latency in this arenaSpeed latency
Benchmarks
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Speed latency — stories about speed latency in this arenaSee published tokens-per-second or latency numbers, benchmarks, or load-testing guides backing the provider's speed claims | ml-engineer | disputedD 5/10 | none 0/10 | partialC 4/10 | partialX 5/10 | none 0/10 | none 0/10 | partialX 5/10 |
Serving
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Speed latency — stories about speed latency in this arenaServe latency-sensitive workloads with fast time-to-first-token and high-throughput generation | developer | disputedD 6/10 | disputedD 5/10 | fullC 8/10 | fullX 9/10 | fullX 8/10 | fullX 8/10 | fullX 8/10 |
| Speed latency — stories about speed latency in this arenaBenefit from prompt/prefix caching that reduces latency or cost on repeated context | ml-engineer | none 0/10 | fullC 8/10 | partialC 6/10 | partialC 6/10 | none 0/10 | fullC 7/10 | fullC 8/10 |
| Speed latency — stories about speed latency in this arenaStream completions token by token over SSE for responsive user experiences | developer | fullX 9/10 | fullC 8/10 | none 0/10 | fullX 8/10 | partialC 4/10 | fullC 8/10 | partialT 3/10 |
Structured tool calling — stories about structured tool calling in this arenaStructured tool calling
Structured
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Structured tool calling — stories about structured tool calling in this arenaEnforce structured outputs against a JSON schema (or grammar) so model responses parse reliably | developer | fullC 8/10 | fullC 8/10 | fullC 8/10 | fullC 7/10 | none 0/10 | fullC 8/10 | none 0/10 |
Tools
| Story | Persona | Morph | ||||||
|---|---|---|---|---|---|---|---|---|
| Structured tool calling — stories about structured tool calling in this arenaRely on faithful function/tool calling — including parallel and multi-step tool use — so agent loops run on open models without breaking | ai-native | partialX 5/10 | partialC 6/10 | partialC 5/10 | partialX 5/10 | none 0/10 | partialC 4/10 | partialX 5/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.