AI Inference Providers — procurement report
ProductArena · rankings as of 2026-09-14 · evidence as of 2026-09-14 · 7 products · 53 judged requirements · 371 judged cells
Methodology: Every product is judged against a shared taxonomy of user stories using cited evidence — hands-on probes > repository code > independent community sources > vendor claims — never opinion. Full writeup: https://ultrametric.ai/productarena/methodology
Leaderboard
| # | Product | PA Score | Coverage score | Applicable cells | Confidence |
|---|---|---|---|---|---|
| 1 | Groq | 30.9 | 34.2 | 48/53 | B |
| 2 | Fireworks AI | 26.5 | 35.9 | 48/53 | B |
| 3 | Cerebras Inference | 26.4 | 30.5 | 48/53 | B |
| 4 | Baseten | 25.8 | 37.3 | 50/53 | B |
| 5 | Together AI | 22.7 | 38.6 | 49/53 | B |
| 6 | Morph | 16.1 | 24.2 | 48/53 | B |
| 7 | DeepInfra | 9.8 | 26.6 | 48/53 | B |
PA Score = agent-readiness blend (see methodology). Coverage score = weighted share of judged requirements met. Confidence = how much of the score rests on tested vs claimed evidence (A–D).
Uncertainty note
The current #1/#2 gap in this arena is not close enough to qualify for the multi-judge uncertainty pass (or the pass has not covered it yet) — no extra caveat applies beyond the per-product confidence grades above.
Buyer checklist (RFP)
The arena's 53 judged user stories as requirements, grouped by theme. Priorities mirror the story weights our scoring uses (3 = must-have, 2 = should-have, 1 = nice-to-have). Interactive version with per-requirement verdicts for the top products: /arena/inference-providers/checklist
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
- ai-native userPlug MCP servers into this product so it can use their toolsmust-have
- ai-native userConnect an agent via an official MCP servermust-have
- ai-native userDrive the product through a documented public APImust-have
- ai-native userDelegate tasks to a built-in AI assistant inside the productmust-have
- ai-native userPoint an agent at llms.txt or agent-oriented docsshould-have
- ai-native userRun the product headlessly / in CI for automationshould-have
- ai-native userUse an official CLIshould-have
- ai-native userIssue scoped/least-privilege API credentials for an agentshould-have
- ai-native userBuild against official SDKsshould-have
- ai-native userSubscribe to events via webhooksshould-have
- ai-native userGet AI-generated insights and suggestions from my data inside the productshould-have
- ai-native userSet up automations that run autonomously in the backgroundshould-have
- ai-native userOperate the product with natural-language commandsshould-have
- ai-native userExplore an interactive API reference with runnable examplesshould-have
- ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)should-have
- ai-native userRely on versioned APIs with a documented deprecation policyshould-have
- ai-native userTest against a sandbox environment without touching production datanice-to-have
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
- ai-native userDefine rules that trigger actions automatically on eventsmust-have
- ai-native userPerform bulk operations across many items at onceshould-have
- ai-native userSchedule recurring jobs or workflowsshould-have
- ai-native userVersion, review, and roll back my automationsnice-to-have
Batch async — stories about batch async in this arenaBatch async
Stories about batch async in this arena
- ml-engineerSubmit asynchronous batch inference jobs at a documented discount versus real-time pricingshould-have
Dedicated capacity — stories about dedicated capacity in this arenaDedicated capacity
Stories about dedicated capacity in this arena
- ml-engineerDeploy a model on dedicated GPU capacity with autoscaling so my traffic is isolated from the shared serverless poolshould-have
Fine tune serving — stories about fine tune serving in this arenaFine tune serving
Stories about fine tune serving in this arena
- ml-engineerFine-tune a supported base model on my own data and serve the result on the same platformshould-have
- ml-engineerUpload and serve my own custom model weights or LoRA adaptersshould-have
Model catalog — stories about model catalog in this arenaModel catalog
Stories about model catalog in this arena
- developerChoose among a broad catalog of current open-weight model families (Llama, Qwen, DeepSeek, GPT-OSS and peers) on shared serverless endpointsmust-have
- ml-engineerGet newly released open-weight models on the platform quickly after their public releaseshould-have
- ai-native userHave an agent enumerate the live model catalog programmatically via a documented GET /v1/models-style endpointshould-have
- developerRely on a documented deprecation policy with advance notice before a hosted model is removednice-to-have
Multimodal — stories about multimodal in this arenaMultimodal
Stories about multimodal in this arena
- developerGenerate embeddings (and rerank results) for retrieval pipelines without a second vendornice-to-have
- developerCall vision, audio, or image-generation models beyond text chat on the same platformnice-to-have
Openai compat — stories about openai compat in this arenaOpenai compat
Stories about openai compat in this arena
- ai-native userHave an agent switch to or away from this provider mid-workflow because it speaks the standard chat-completions API without provider-specific code changesmust-have
- developerPoint an existing OpenAI SDK client at the provider by changing only the base URL and API keymust-have
- ai-native userPlug the provider into coding agents and agent frameworks through documented, first-party integration guidesshould-have
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
- ai-native userExport all of my data in open formats and leavemust-have
- ai-native userSelf-host the core productmust-have
- ai-native userDo everything through the API that I can do in the UIshould-have
- ai-native userRead the product's source under an open licenseshould-have
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
- founderSee public per-token prices for every hosted model without talking to salesmust-have
- developerRead documented rate limits and how they scale across usage tiers before I hit them in productionshould-have
- founderSet spending caps or budget alerts so a runaway workload cannot generate an unbounded billnice-to-have
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
- ai-native userPrevent my data from being used to train AI modelsmust-have
- ai-native userChoose where my data is stored (region/residency)should-have
- ai-native userControl data retention and deletionshould-have
- ai-native userOpt out of telemetry and usage trackingshould-have
Reliability status — stories about reliability status in this arenaReliability status
Stories about reliability status in this arena
- founderCheck a public status page with incident history before betting production traffic on the platformshould-have
- founderGet a stated availability SLA on paid or enterprise tiersnice-to-have
Speed latency — stories about speed latency in this arenaSpeed latency
Stories about speed latency in this arena
- developerServe latency-sensitive workloads with fast time-to-first-token and high-throughput generationmust-have
- developerStream completions token by token over SSE for responsive user experiencesmust-have
- ml-engineerSee published tokens-per-second or latency numbers, benchmarks, or load-testing guides backing the provider's speed claimsshould-have
- ml-engineerBenefit from prompt/prefix caching that reduces latency or cost on repeated contextnice-to-have
Structured tool calling — stories about structured tool calling in this arenaStructured tool calling
Stories about structured tool calling in this arena
- developerEnforce structured outputs against a JSON schema (or grammar) so model responses parse reliablymust-have
- ai-native userRely on faithful function/tool calling — including parallel and multi-step tool use — so agent loops run on open models without breakingmust-have
Pricing signals
Extracted verbatim from each vendor's own pricing page — never converted, averaged, or derived. Products whose page prints no unit price are recorded as unclear, honestly.
| Product | Headline price | Unit | As of |
|---|---|---|---|
| Groq | pricing unclear — The page is a marketing/landing page with no printed pricing figures for any model, token rate, or plan. | 2026-09-07 | |
| Fireworks AI | $0.008 | 1M input tokens | 2026-09-07 |
| Cerebras Inference | free tier | 1M tokens | 2026-09-07 |
| Baseten | $0.1 | 1M input tokens | 2026-09-07 |
| Together AI | $0.3 | 1M input tokens | 2026-09-07 |
| DeepInfra | $0.005 | 1M input tokens | 2026-09-07 |
Appendix: recorded probes
No replayable probe recordings exist for this arena yet. Probe-tier evidence (hands-on checks) still backs verdicts where cited — see each product page for the evidence trail.
Cite as: ProductArena by Ultrametric Inc, AI Inference Providers arena, rankings as of 2026-09-14 — https://ultrametric.ai/productarena/arena/inference-providers
License: © 2026 Ultrametric Inc. Brief quotation of individual verdicts, scores, or evidence excerpts is permitted with attribution to "ProductArena by Ultrametric Inc (ultrametric.ai/productarena)", as is use of the data to evaluate, contest, or contribute corrections. Bulk copying, redistribution, or use to build competing datasets requires prior written permission (see DATA-LICENSE in the repository).
No liability: rankings, verdicts, and scores are research outputs derived from the cited evidence at a point in time, provided "as is", without warranties. Ultrametric Inc accepts no responsibility for procurement, purchasing, or other decisions made in reliance on them — verify against the cited evidence before acting (https://ultrametric.ai/productarena/terms).