Rank #7 of 7 in AI Inference Providers
Access
Install
npm install deepinfraShowcase


Verified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Batch async — stories about batch async in this arenaBatch asyncevidence →
Stories about batch async in this arena
Dedicated capacity — stories about dedicated capacity in this arenaDedicated capacityevidence →
Stories about dedicated capacity in this arena
Fine tune serving — stories about fine tune serving in this arenaFine tune servingevidence →
Stories about fine tune serving in this arena
Model catalog — stories about model catalog in this arenaModel catalogevidence →
Stories about model catalog in this arena
Multimodal — stories about multimodal in this arenaMultimodalevidence →
Stories about multimodal in this arena
Openai compat — stories about openai compat in this arenaOpenai compatevidence →
Stories about openai compat in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Reliability status — stories about reliability status in this arenaReliability statusevidence →
Stories about reliability status in this arena
Speed latency — stories about speed latency in this arenaSpeed latencyevidence →
Stories about speed latency in this arena
Structured tool calling — stories about structured tool calling in this arenaStructured tool callingevidence →
Stories about structured tool calling in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 1 free · 14 paid · 0 enterprise · 6 not stated in evidence
Follow the green: where the map greys out is where DeepInfra stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Webhooks · Scoped API keys · MCP server · Machine-readable spec · Versioning policy · API sandbox · Official CLI · Full data export
Subscribe to events via webhooks
—–
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
—0/10
Connect an agent via an official MCP server
—–
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—0/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—–
Operate the product with natural-language commands
—–
Plug MCP servers into this product so it can use their tools
—–
Get AI-generated insights and suggestions from my data inside the product
n/an/a
Set up automations that run autonomously in the background
—–
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Batch async — stories about batch async in this arenaBatch async
Stories about batch async in this arena
Dedicated capacity — stories about dedicated capacity in this arenaDedicated capacity
Stories about dedicated capacity in this arena
Fine tune serving — stories about fine tune serving in this arenaFine tune serving
Stories about fine tune serving in this arena
Model catalog — stories about model catalog in this arenaModel catalog
Stories about model catalog in this arena
Catalog
Multimodal — stories about multimodal in this arenaMultimodal
Stories about multimodal in this arena
Openai compat — stories about openai compat in this arenaOpenai compat
Stories about openai compat in this arena
Compat
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Reliability status — stories about reliability status in this arenaReliability status
Stories about reliability status in this arena
Speed latency — stories about speed latency in this arenaSpeed latency
Stories about speed latency in this arena
Structured tool calling — stories about structured tool calling in this arenaStructured tool calling
Stories about structured tool calling in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | fullpaid | 8/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullpaid | 8/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullpaid | 7/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | 0/10 | ||
Have an agent switch to or away from this provider mid-workflow because it speaks the standard chat-completions API without provider-specific code changes C Compat | ai-native user | Openai compat — stories about openai compat in this arenaOpenai compat | 3 | fullpaid | 9/10 | Tprobed | |
Point an existing OpenAI SDK client at the provider by changing only the base URL and API key C Compat | developer | Openai compat — stories about openai compat in this arenaOpenai compat | 3 | fullpaid | 9/10 | Tprobed | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | full | 8/10 | Cclaimed | |
Serve latency-sensitive workloads with fast time-to-first-token and high-throughput generation C Serving | developer | Speed latency — stories about speed latency in this arenaSpeed latency | 3 | fullpaid | 8/10 | Xcommunity | |
Choose among a broad catalog of current open-weight model families (Llama, Qwen, DeepSeek, GPT-OSS and peers) on shared serverless endpoints C Catalog | developer | Model catalog — stories about model catalog in this arenaModel catalog | 3 | partialpaid | 7/10 | Tprobed | |
See public per-token prices for every hosted model without talking to sales G Pricing | founder | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 3 | partialpaid | 5/10 | Tprobed | |
Stream completions token by token over SSE for responsive user experiences C Serving | developer | Speed latency — stories about speed latency in this arenaSpeed latency | 3 | partialpaid | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | n/a | untested | none yet | |
Enforce structured outputs against a JSON schema (or grammar) so model responses parse reliably C Structured | developer | Structured tool calling — stories about structured tool calling in this arenaStructured tool calling | 3 | none | untested | none yet | |
Rely on faithful function/tool calling — including parallel and multi-step tool use — so agent loops run on open models without breaking C Tools | ai-native user | Structured tool calling — stories about structured tool calling in this arenaStructured tool calling | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | n/a | untested | none yet | |
Have an agent enumerate the live model catalog programmatically via a documented GET /v1/models-style endpoint C Catalog | ai-native user | Model catalog — stories about model catalog in this arenaModel catalog | 2 | fullfree | 8/10 | Tprobed | |
Deploy a model on dedicated GPU capacity with autoscaling so my traffic is isolated from the shared serverless pool C Dedicated | ml-engineer | Dedicated capacity — stories about dedicated capacity in this arenaDedicated capacity | 2 | full | 7/10 | Cclaimed | |
Check a public status page with incident history before betting production traffic on the platform G Reliability | founder | Reliability status — stories about reliability status in this arenaReliability status | 2 | partial | 6/10 | Tprobed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partialpaid | 6/10 | Tprobed | |
Upload and serve my own custom model weights or LoRA adapters C Fine tune | ml-engineer | Fine tune serving — stories about fine tune serving in this arenaFine tune serving | 2 | partial | 6/10 | Cclaimed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 5/10 | Cclaimed | |
Plug the provider into coding agents and agent frameworks through documented, first-party integration guides C Compat | ai-native user | Openai compat — stories about openai compat in this arenaOpenai compat | 2 | partialpaid | 5/10 | Tprobed | |
Fine-tune a supported base model on my own data and serve the result on the same platform C Fine tune | ml-engineer | Fine tune serving — stories about fine tune serving in this arenaFine tune serving | 2 | partialpaid | 4/10 | Cclaimed | |
Submit asynchronous batch inference jobs at a documented discount versus real-time pricing G Batch | ml-engineer | Batch async — stories about batch async in this arenaBatch async | 2 | partialpaid | 4/10 | Cclaimed | |
Get newly released open-weight models on the platform quickly after their public release C Catalog | ml-engineer | Model catalog — stories about model catalog in this arenaModel catalog | 2 | none | 0/10 | ||
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Read documented rate limits and how they scale across usage tiers before I hit them in production G Limits | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | 0/10 | ||
See published tokens-per-second or latency numbers, benchmarks, or load-testing guides backing the provider's speed claims C Benchmarks | ml-engineer | Speed latency — stories about speed latency in this arenaSpeed latency | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | n/a | untested | none yet | |
Call vision, audio, or image-generation models beyond text chat on the same platform C Modalities | developer | Multimodal — stories about multimodal in this arenaMultimodal | 1 | full | 8/10 | Cclaimed | |
Generate embeddings (and rerank results) for retrieval pipelines without a second vendor C Modalities | developer | Multimodal — stories about multimodal in this arenaMultimodal | 1 | fullpaid | 8/10 | Cclaimed | |
Get a stated availability SLA on paid or enterprise tiers C Reliability | founder | Reliability status — stories about reliability status in this arenaReliability status | 1 | none | 0/10 | ||
Set spending caps or budget alerts so a runaway workload cannot generate an unbounded bill G Pricing | founder | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | none | 0/10 | ||
Benefit from prompt/prefix caching that reduces latency or cost on repeated context C Serving | ml-engineer | Speed latency — stories about speed latency in this arenaSpeed latency | 1 | none | untested | none yet | |
Rely on a documented deprecation policy with advance notice before a hosted model is removed C Catalog | developer | Model catalog — stories about model catalog in this arenaModel catalog | 1 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | n/a | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 37 stories with headroom
What would move DeepInfra’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
DeepInfra is an inference API/platform (not an agent), so an official MCP server is a fair axis to expect, but no evidence in the pack mentions MCP support at all—only OpenAI-compatible REST API docs.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
DeepInfra's evidence pack focuses entirely on its OpenAI-compatible inference API, model catalog, and infrastructure offerings; there is no mention of any user data export feature, open-format data portability, or account data download capability.
Structured tool calling — stories about structured tool calling in this arenaEnforce structured outputs against a JSON schema (or grammar) so model responses parse reliably
nonemoves PA Scoreimpact 30
Missing: any documentation of `response_format`/json_schema support, grammar-based constrained generation, or examples showing reliable JSON parsing guarantees.
Structured tool calling — stories about structured tool calling in this arenaRely on faithful function/tool calling — including parallel and multi-step tool use — so agent loops run on open models without breaking
nonemoves PA Scoreimpact 30
Missing: explicit documentation of a `tools`/`function_call` parameter, evidence of parallel tool call support, and any hands-on or benchmark confirmation that tool calling works reliably in agent loops.
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs
nonemoves agent-readyimpact 30
Direct probes show llms.txt, docs.md, and openapi.json all return 404 on deepinfra.com, and no evidence pack item claims an agent-oriented docs format exists; the only machine-readable endpoint found is the models list, not documentation.
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background
nonemoves Built-in AIimpact 30
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Showing the top 8 of 37 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 21 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs18 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Submit asynchronous batch inference jobs at a documented discount versus real-time pricing
- Deploy a model on dedicated GPU capacity with autoscaling so my traffic is isolated from the shared serverless pool
- Fine-tune a supported base model on my own data and serve the result on the same platform
- Upload and serve my own custom model weights or LoRA adapters
- Have an agent enumerate the live model catalog programmatically via a documented GET /v1/models-style endpoint
- Choose among a broad catalog of current open-weight model families (Llama, Qwen, DeepSeek, GPT-OSS and peers) on shared serverless endpoints
- Generate embeddings (and rerank results) for retrieval pipelines without a second vendor
- Call vision, audio, or image-generation models beyond text chat on the same platform
- Plug the provider into coding agents and agent frameworks through documented, first-party integration guides
- Have an agent switch to or away from this provider mid-workflow because it speaks the standard chat-completions API without provider-specific code changes
- Point an existing OpenAI SDK client at the provider by changing only the base URL and API key
- Do everything through the API that I can do in the UI
- See public per-token prices for every hosted model without talking to sales
- Serve latency-sensitive workloads with fast time-to-first-token and high-throughput generation
- Stream completions token by token over SSE for responsive user experiences
V1 docs10 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Have an agent enumerate the live model catalog programmatically via a documented GET /v1/models-style endpoint
- Choose among a broad catalog of current open-weight model families (Llama, Qwen, DeepSeek, GPT-OSS and peers) on shared serverless endpoints
- Plug the provider into coding agents and agent frameworks through documented, first-party integration guides
- Have an agent switch to or away from this provider mid-workflow because it speaks the standard chat-completions API without provider-specific code changes
- Point an existing OpenAI SDK client at the provider by changing only the base URL and API key
- Do everything through the API that I can do in the UI
- See public per-token prices for every hosted model without talking to sales
deepinfra.com3 stories
llms.txt2 stories
OpenAPI spec2 stories
status.deepinfra.com1 story
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
4 of 12 testable claims verified · 0 contradicted → integrity 33/100
18 distinct capability claims found in DeepInfra’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
4
Verified
8
Unverified
0
Contradicted
9
Undersold
Verified (7)
“Existing OpenAI SDK code works unchanged by just pointing base_url to DeepInfra's endpoint”
Point an existing OpenAI SDK client at the provider by changing only the base URL and API keyfullproof ↗
“Existing OpenAI SDK code works unchanged by just pointing base_url to DeepInfra's endpoint”
Have an agent switch to or away from this provider mid-workflow because it speaks the standard chat-completions API without provider-specific code changesfullproof ↗
“OpenAI-compatible chat completions API — swap base URL and model name only”
Point an existing OpenAI SDK client at the provider by changing only the base URL and API keyfullproof ↗
“Three-step migration: set base_url, api_key, and model to use existing OpenAI code”
Point an existing OpenAI SDK client at the provider by changing only the base URL and API keyfullproof ↗
“Priority service tier gives faster time-to-first-token and higher throughput under load”
Serve latency-sensitive workloads with fast time-to-first-token and high-throughput generationfullproof ↗
“Pay-per-token pricing with no idle GPU time, minimums, or seat fees”
See public per-token prices for every hosted model without talking to salespartialproof ↗
“Low pay-as-you-go pricing with no long-term contracts or hidden fees, scaling from startup to enterprise”
See public per-token prices for every hosted model without talking to salespartialproof ↗
Unverified (11)
“Flex tier offers lower-cost inference for non-production workloads like evals, enrichment, async jobs”
Submit asynchronous batch inference jobs at a documented discount versus real-time pricingpartialproof ↗
“Flex tier is billed at a 20% discount off standard per-token pricing”
Submit asynchronous batch inference jobs at a documented discount versus real-time pricingpartialproof ↗
“Deploy your own fine-tuned LLM on dedicated GPUs (A100/H100/H200/B200/B300) with autoscaling and a private endpoint”
Deploy a model on dedicated GPU capacity with autoscaling so my traffic is isolated from the shared serverless poolfullproof ↗
“Deploy your own fine-tuned LLM on dedicated GPUs (A100/H100/H200/B200/B300) with autoscaling and a private endpoint”
Fine-tune a supported base model on my own data and serve the result on the same platformpartialproof ↗
“Deploy your own fine-tuned LLM on dedicated GPUs (A100/H100/H200/B200/B300) with autoscaling and a private endpoint”
Upload and serve my own custom model weights or LoRA adapterspartialproof ↗
“Offers image and video generation models such as FLUX, Stable Diffusion, and text-to-video”
Call vision, audio, or image-generation models beyond text chat on the same platformfullproof ↗
“Provides state-of-the-art embedding and reranker models for search and RAG”
Generate embeddings (and rerank results) for retrieval pipelines without a second vendorfullproof ↗
“Offers speech recognition (Whisper) and text-to-speech models”
Call vision, audio, or image-generation models beyond text chat on the same platformfullproof ↗
“Provides multimodal vision models for visual understanding and document OCR”
Call vision, audio, or image-generation models beyond text chat on the same platformfullproof ↗
“Zero retention policy keeps inputs/outputs/user data private; SOC 2 and ISO 27001 certified”
Prevent my data from being used to train AI modelsfullproof ↗
“Zero retention policy keeps inputs/outputs/user data private; SOC 2 and ISO 27001 certified”
Undersold (9)
Run the product headlessly / in CI for automationfullproof ↗
Drive the product through a documented public APIfullproof ↗
Have an agent enumerate the live model catalog programmatically via a documented GET /v1/models-style endpointfullproof ↗
Choose among a broad catalog of current open-weight model families (Llama, Qwen, DeepSeek, GPT-OSS and peers) on shared serverless endpointspartialproof ↗
Plug the provider into coding agents and agent frameworks through documented, first-party integration guidespartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Check a public status page with incident history before betting production traffic on the platformpartialproof ↗
Stream completions token by token over SSE for responsive user experiencespartialproof ↗
Claims outside our story set (4)
Real capability claims found in DeepInfra’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Supports multi-turn conversations by including full message history in each request”
source ↗“fail_fast parameter returns immediate HTTP 429 instead of queuing the request”
source ↗“reasoning_effort parameter controls reasoning depth for reasoning models”
source ↗“Rent bare-metal B200/B300 GPU clusters with SSH access for training and full control”
source ↗
Pricing signals
- $0.09per 1M input tokenspay-as-you-goDeepSeek-V4-Flash-0731 input pricing, representative per-model ratesource ↗as of 2026-09-07
- $0.18per 1M output tokenspay-as-you-goDeepSeek-V4-Flash-0731 output pricing, representative per-model ratesource ↗as of 2026-09-07
- $0.02per 1M input tokenspay-as-you-goMeta-Llama-3.1-8B-Instruct-Turbo, cheapest listed input ratesource ↗as of 2026-09-07
- $0.04per 1M output tokenspay-as-you-goMeta-Llama-3.1-8B-Instruct-Turbo output ratesource ↗as of 2026-09-07
- $0.005per 1M input tokenspay-as-you-goEmbeddings model bge-base-en-v1.5, representative embeddings ratesource ↗as of 2026-09-07
Extracted verbatim from the vendor’s own pricing page — hover a figure for the exact quote.
Business model
Pay-as-you-go per-token pricing per model (or per-second for some workloads) with no long-term contracts, plus hourly-billed dedicated GPU instances for custom deployments.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
