Voice Agent Platforms — procurement report
ProductArena · rankings as of 2026-09-15 · evidence as of 2026-09-15 · 9 products · 61 judged requirements · 549 judged cells
Methodology: Every product is judged against a shared taxonomy of user stories using cited evidence — hands-on probes > repository code > independent community sources > vendor claims — never opinion. Full writeup: https://ultrametric.ai/productarena/methodology
Leaderboard
| # | Product | PA Score | Coverage score | Applicable cells | Confidence |
|---|---|---|---|---|---|
| 1 | telli | 37.8 | 27.1 | 61/61 | B |
| 2 | Bland | 37.6 | 31.8 | 61/61 | B |
| 3 | Retell AI | 35.2 | 36.3 | 61/61 | B |
| 4 | LiveKit Agents | 32.9 | 32.0 | 61/61 | B |
| 5 | Bolna | 29.4 | 26.0 | 61/61 | B |
| 6 | Pipecat | 28.8 | 29.3 | 61/61 | B |
| 7 | ElevenLabs Agents | 28.4 | 31.3 | 61/61 | C |
| 8 | Vapi | 28.4 | 23.9 | 61/61 | C |
| 9 | Deepgram Voice Agent | 25.6 | 20.1 | 61/61 | C |
PA Score = agent-readiness blend (see methodology). Coverage score = weighted share of judged requirements met. Confidence = how much of the score rests on tested vs claimed evidence (A–D).
Uncertainty note
This arena is currently a close race: telli (37.8) vs Bland (37.6), a gap of 0.2 PA Score. The ordering was re-checked with extra judge samples: 34 decisive cells were triple-judged and 8 came back unstable. Treat the #1/#2 ordering as contested — shortlist both.
Buyer checklist (RFP)
The arena's 61 judged user stories as requirements, grouped by theme. Priorities mirror the story weights our scoring uses (3 = must-have, 2 = should-have, 1 = nice-to-have). Interactive version with per-requirement verdicts for the top products: /arena/voice-agents/checklist
Agent building — building agents — abstractions, tool wiring, control flowAgent building
Building agents — abstractions, tool wiring, control flow
- ai-native userMy coding agent can provision a complete voice agent end to end — create the agent, attach a number, and place a call — through the API, CLI, or MCP without touching the dashboardmust-have
- developerBuild a working phone voice agent — prompt, voice, and phone number — and take my first live call within an hourmust-have
- developerRun conversations in multiple languages, including detecting and switching language mid-callshould-have
- founderDesign multi-step conversation flows in a visual builder with branching, states, and handoffs without writing codeshould-have
- developerInject dynamic variables and per-caller context at call time so each conversation is personalizedshould-have
- developerGround the agent on my documents with a built-in knowledge base or RAG so it answers from my contentshould-have
- ai-native userThe platform's own AI helps me author agents — generating or improving prompts, flows, and test cases from a descriptionnice-to-have
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
- ai-native userPlug MCP servers into this product so it can use their toolsmust-have
- ai-native userConnect an agent via an official MCP servermust-have
- ai-native userDrive the product through a documented public APImust-have
- ai-native userDelegate tasks to a built-in AI assistant inside the productmust-have
- ai-native userPoint an agent at llms.txt or agent-oriented docsshould-have
- ai-native userRun the product headlessly / in CI for automationshould-have
- ai-native userUse an official CLIshould-have
- ai-native userIssue scoped/least-privilege API credentials for an agentshould-have
- ai-native userBuild against official SDKsshould-have
- ai-native userSubscribe to events via webhooksshould-have
- ai-native userGet AI-generated insights and suggestions from my data inside the productshould-have
- ai-native userSet up automations that run autonomously in the backgroundshould-have
- ai-native userOperate the product with natural-language commandsshould-have
- ai-native userExplore an interactive API reference with runnable examplesshould-have
- ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)should-have
- ai-native userRely on versioned APIs with a documented deprecation policyshould-have
- ai-native userTest against a sandbox environment without touching production datanice-to-have
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
- ai-native userDefine rules that trigger actions automatically on eventsmust-have
- ai-native userPerform bulk operations across many items at onceshould-have
- ai-native userSchedule recurring jobs or workflowsshould-have
- ai-native userVersion, review, and roll back my automationsnice-to-have
Compliance trust — stories about compliance trust in this arenaCompliance trust
Stories about compliance trust in this arena
- founderMeet call-recording consent and disclosure obligations with per-call recording controls and configurable data retentionshould-have
- platform-engineerRun regulated workloads with HIPAA/BAA support, SOC 2, and data-residency optionsshould-have
Deployment scale — stories about deployment scale in this arenaDeployment scale
Stories about deployment scale in this arena
- platform-engineerSelf-host the voice agent runtime from open-source code on my own infrastructuremust-have
- platform-engineerSee documented concurrency limits and scale to many simultaneous calls without manual capacity beggingshould-have
Latency turntaking — stories about latency turntaking in this arenaLatency turntaking
Stories about latency turntaking in this arena
- platform-engineerSee documented end-to-end voice latency numbers or tuning guidance backing the platform's speed claimsmust-have
- developerRely on the agent to handle interruptions (barge-in) gracefully — stopping speech, updating context, and recovering the turnmust-have
- developerUse model-based end-of-turn detection beyond simple VAD silence timeouts so the agent doesn't talk over slow speakersshould-have
- developerEnable noise suppression or audio filtering so the agent stays coherent on noisy real-world callsnice-to-have
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
- ai-native userExport all of my data in open formats and leavemust-have
- ai-native userSelf-host the core productmust-have
- ai-native userDo everything through the API that I can do in the UIshould-have
- ai-native userRead the product's source under an open licenseshould-have
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans
Plan structure and value — what each tier costs and what it unlocks
- founderSee published per-minute or usage pricing and estimate cost per call before committingshould-have
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
- ai-native userPrevent my data from being used to train AI modelsmust-have
- ai-native userChoose where my data is stored (region/residency)should-have
- ai-native userControl data retention and deletionshould-have
- ai-native userOpt out of telemetry and usage trackingshould-have
Telephony — stories about telephony in this arenaTelephony
Stories about telephony in this arena
- developerProvision phone numbers and run both inbound and outbound calls through the platform's APImust-have
- developerEscalate a live call to a human with warm or blind transfer, passing context alongshould-have
- founderRun batch outbound call campaigns with scheduling and throughput controlsshould-have
- platform-engineerConnect my own carrier or PBX via SIP trunking (or import Twilio/Telnyx numbers) instead of being locked to bundled telephonyshould-have
- developerMy agent can send DTMF keypresses, navigate IVR menus, and detect or leave voicemailnice-to-have
Testing analytics — stories about testing analytics in this arenaTesting analytics
Stories about testing analytics in this arena
- ai-native userThe platform's AI reviews my calls for me — scoring quality, flagging failures, and analyzing resolution automaticallyshould-have
- founderSee call analytics — success rates, durations, outcomes, sentiment — in dashboards without building my ownshould-have
- developerTest agents with simulated conversations or evals before putting them on real phone callsshould-have
- platform-engineerMonitor live calls in production and get alerts when agents misbehave or error rates spikenice-to-have
Tools function calling — stories about tools function calling in this arenaTools function calling
Stories about tools function calling in this arena
- developerMy agent can call external APIs and custom functions mid-conversation and speak the result without awkward dead airmust-have
- developerExtract structured data from every call — outcomes, entities, dispositions — delivered via API or webhook after the callshould-have
- ai-native userMy voice agent can plug in MCP servers as tool sources so one integration grants it whole toolsets mid-callshould-have
Transcription recording — stories about transcription recording in this arenaTranscription recording
Stories about transcription recording in this arena
- platform-engineerRetrieve full call recordings and transcripts programmatically for every callshould-have
- developerGet accurate real-time transcription with control over the STT provider, language models, or key termsshould-have
Voices tts — stories about voices tts in this arenaVoices tts
Stories about voices tts in this arena
- founderClone a custom brand voice and use it for my agents, with a documented consent processshould-have
- developerChoose from a broad voice library or plug in multiple TTS providers to get the voice I wantshould-have
Pricing signals
Extracted verbatim from each vendor's own pricing page — never converted, averaged, or derived. Products whose page prints no unit price are recorded as unclear, honestly.
| Product | Headline price | Unit | As of |
|---|---|---|---|
| Bland | $0.14 | minute | 2026-09-07 |
| Retell AI | $0.015 | minute | 2026-09-07 |
| LiveKit Agents | $0.01 | minute | 2026-09-07 |
| Pipecat | $0.01 | minute | 2026-09-07 |
| ElevenLabs Agents | $0.2 | minute | 2026-09-07 |
| Vapi | $0.05 | minute | 2026-09-07 |
Appendix: recorded probes
Hands-on probe recordings — transcripts/videos a human can replay, the strongest evidence tier. Watch them at https://ultrametric.ai/productarena/proofs
- Bland
npx -y bland-cli --versionterminal · recorded 2026-09-05 · exit 0 - Bland
curl -si -X POST https://api.bland.ai/v1/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-05 · exit 0 - Bolna
curl -sL https://docs.bolna.ai/llms.txt | head -6terminal · recorded 2026-09-10 · exit 0 - Bolna
curl -si -X POST https://mcp.bolna.ai/api/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-10 · exit 0 - Deepgram Voice Agent
curl -si https://api.deepgram.com/v1/projectsterminal · recorded 2026-09-15 · exit 0 - Deepgram Voice Agent
curl -si -X POST https://developers.deepgram.com/_mcp/server -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-15 · exit 0 - Deepgram Voice Agent
curl -s https://developers.deepgram.com/llms.txt | head -8terminal · recorded 2026-09-15 · exit 0 - ElevenLabs Agents
npx -y @elevenlabs/cli --versionterminal · recorded 2026-09-05 · exit 0 - ElevenLabs Agents
curl -si -X POST https://api.elevenlabs.io/v1/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-05 · exit 0 - LiveKit Agents
lk --versionterminal · recorded 2026-09-05 · exit 0 - LiveKit Agents
livekit-server --dev # keyless boot, curl :7880, then killterminal · recorded 2026-09-05 · exit 0 - LiveKit Agents
uv run --with livekit-agents python3 -c 'import livekit.agents as a; print("PA_PROBE_OK livekit-agents", a.__version__)'terminal · recorded 2026-09-05 · exit 0 - Pipecat
uvx --from "pipecat-ai[cli]" pipecat --helpterminal · recorded 2026-09-05 · exit 0 - Pipecat
uv run --with pipecat-ai python3 -c 'import pipecat; print("PA_PROBE_OK pipecat-ai imported")'terminal · recorded 2026-09-05 · exit 0 - Retell AI
npx -y @retell-ai/retell-cli --versionterminal · recorded 2026-09-05 · exit 0 - Retell AI
curl -si -X POST https://mcp.retellai.com -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-05 · exit 0 - telli
curl -sL https://docs.telli.com/llms.txt | head -6terminal · recorded 2026-09-14 · exit 0 - telli
curl -si -X POST https://mcp.telli.com/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-14 · exit 0 - telli
curl -sL https://docs.telli.com/openapi-v2.json | head -c 300terminal · recorded 2026-09-14 · exit 0 - Vapi
vapi --version # installed via `curl -sSL https://vapi.ai/install.sh | bash`terminal · recorded 2026-09-05 · exit 0 - Vapi
curl -si -X POST https://mcp.vapi.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-05 · exit 0
Cite as: ProductArena by Ultrametric Inc, Voice Agent Platforms arena, rankings as of 2026-09-15 — https://ultrametric.ai/productarena/arena/voice-agents
License: © 2026 Ultrametric Inc. Brief quotation of individual verdicts, scores, or evidence excerpts is permitted with attribution to "ProductArena by Ultrametric Inc (ultrametric.ai/productarena)", as is use of the data to evaluate, contest, or contribute corrections. Bulk copying, redistribution, or use to build competing datasets requires prior written permission (see DATA-LICENSE in the repository).
No liability: rankings, verdicts, and scores are research outputs derived from the cited evidence at a point in time, provided "as is", without warranties. Ultrametric Inc accepts no responsibility for procurement, purchasing, or other decisions made in reliance on them — verify against the cited evidence before acting (https://ultrametric.ai/productarena/terms).