AI Customer Support Agents Arena
AI Customer Support Agents — procurement report
ProductArena · rankings as of 2026-09-16 · evidence as of 2026-09-16 · 6 products · 53 judged requirements · 318 judged cells
Methodology: Every product is judged against a shared taxonomy of user stories using cited evidence — hands-on probes > repository code > independent community sources > vendor claims — never opinion. Full writeup: https://ultrametric.ai/productarena/methodology
Leaderboard
| # | Product | PA Score | Coverage score | Applicable cells | Confidence |
|---|---|---|---|---|---|
| 1 | Lorikeet | 31.2 | 37.6 | 50/53 | C |
| 2 | Pylon | 30.9 | 27.0 | 52/53 | B |
| 3 | Parahelp | 29.4 | 30.3 | 49/53 | C |
| 4 | Fin | 29.2 | 38.5 | 51/53 | B |
| 5 | Sierra | 14.1 | 22.6 | 52/53 | B |
| 6 | Decagon | 12.3 | 20.1 | 50/53 | C |
PA Score = agent-readiness blend (see methodology). Coverage score = weighted share of judged requirements met. Confidence = how much of the score rests on tested vs claimed evidence (A–D).
Uncertainty note
The current #1/#2 gap in this arena is not close enough to qualify for the multi-judge uncertainty pass (or the pass has not covered it yet) — no extra caveat applies beyond the per-product confidence grades above.
Buyer checklist (RFP)
The arena's 53 judged user stories as requirements, grouped by theme. Priorities mirror the story weights our scoring uses (3 = must-have, 2 = should-have, 1 = nice-to-have). Interactive version with per-requirement verdicts for the top products: /arena/ai-support-agents/checklist
Agent actions — stories about agent actions in this arenaAgent actions
Stories about agent actions in this arena
- developerThe agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per actionmust-have
- support ops leadI encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branchingshould-have
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
- ai-native userPlug MCP servers into this product so it can use their toolsmust-have
- ai-native userConnect an agent via an official MCP servermust-have
- ai-native userDrive the product through a documented public APImust-have
- ai-native userDelegate tasks to a built-in AI assistant inside the productmust-have
- ai-native userPoint an agent at llms.txt or agent-oriented docsshould-have
- ai-native userRun the product headlessly / in CI for automationshould-have
- ai-native userUse an official CLIshould-have
- ai-native userIssue scoped/least-privilege API credentials for an agentshould-have
- ai-native userBuild against official SDKsshould-have
- ai-native userSubscribe to events via webhooksshould-have
- ai-native userGet AI-generated insights and suggestions from my data inside the productshould-have
- ai-native userSet up automations that run autonomously in the backgroundshould-have
- ai-native userOperate the product with natural-language commandsshould-have
- ai-native userExplore an interactive API reference with runnable examplesshould-have
- ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)should-have
- ai-native userRely on versioned APIs with a documented deprecation policyshould-have
- ai-native userTest against a sandbox environment without touching production datanice-to-have
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
- ai-native userDefine rules that trigger actions automatically on eventsmust-have
- ai-native userPerform bulk operations across many items at onceshould-have
- ai-native userSchedule recurring jobs or workflowsshould-have
- ai-native userVersion, review, and roll back my automationsnice-to-have
Channels languages — stories about channels languages in this arenaChannels languages
Stories about channels languages in this arena
- support leaderOne agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, socialshould-have
- support leaderThe agent supports customers in many languages, even where my knowledge base exists only in Englishshould-have
- support leaderThe agent handles phone calls — speech in, speech out — with the same knowledge and actions as chatshould-have
Escalation handoff — stories about escalation handoff in this arenaEscalation handoff
Stories about escalation handoff in this arena
- support leaderWhen the agent escalates, the human gets the full conversation, a summary, and collected details — the customer never repeats themselvesmust-have
- support ops leadI configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeysshould-have
Guardrails safety — stories about guardrails safety in this arenaGuardrails safety
Stories about guardrails safety in this arena
- ai-native userGuardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guessmust-have
- support ops leadLaunch in a supervised mode where the agent drafts replies for human approval before anything reaches a customershould-have
- support ops leadI mark topics as human-only — legal threats, cancellations, security — and the agent never freelances on themshould-have
Insights analytics — stories about insights analytics in this arenaInsights analytics
Stories about insights analytics in this arena
- support leaderDashboards show resolution rate, CSAT, handoff rate, and cost per resolution — the numbers I report to my exec teammust-have
- support leaderThe platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volumenice-to-have
Integrations platform — stories about integrations platform in this arenaIntegrations platform
Stories about integrations platform in this arena
- developerThe agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both waysmust-have
Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding
Stories about knowledge grounding in this arena
- ai-native userEvery answer is grounded in my own content and shows which article or source it drew frommust-have
- support ops leadThe agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoringmust-have
- support ops leadKnowledge stays current automatically — the agent re-syncs sources on a schedule or on change, not via manual re-uploadsshould-have
- support ops leadThe platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questionsnice-to-have
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
- ai-native userExport all of my data in open formats and leavemust-have
- ai-native userSelf-host the core productmust-have
- ai-native userDo everything through the API that I can do in the UIshould-have
- ai-native userRead the product's source under an open licenseshould-have
Pricing economics — stories about pricing economics in this arenaPricing economics
Stories about pricing economics in this arena
- support leaderPricing is outcome-based and published — I pay per resolution with caps and controls, not an opaque enterprise quoteshould-have
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
- ai-native userPrevent my data from being used to train AI modelsmust-have
- ai-native userChoose where my data is stored (region/residency)should-have
- ai-native userControl data retention and deletionshould-have
- ai-native userOpt out of telemetry and usage trackingshould-have
Resolution quality — stories about resolution quality in this arenaResolution quality
Stories about resolution quality in this arena
- support leaderThe agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bouncesmust-have
- support leaderAnswers use the customer's live data — plan, order status, account history — not just generic help articlesshould-have
- support leaderThe agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answershould-have
- support leaderI control the agent's tone and brand voice, and it stays consistent across topics and languagesnice-to-have
Testing qa — stories about testing qa in this arenaTesting qa
Stories about testing qa in this arena
- support ops leadI test the agent against historical tickets or simulated conversations before it faces real customersshould-have
- support ops leadAI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agentnice-to-have
Appendix: recorded probes
Hands-on probe recordings — transcripts/videos a human can replay, the strongest evidence tier. Watch them at https://ultrametric.ai/productarena/proofs
- Decagon
curl -si https://docs.decagon.ai/ | head -12 # finding: technical docs are login-gatedterminal · recorded 2026-09-10 · exit 0 - Decagon
curl -s https://decagon.ai/sitemap.xml | grep -o 'https://decagon.ai/product/[^<]*' | sort | head -12terminal · recorded 2026-09-10 · exit 0 - Decagon
curl -s https://decagon.ai/llms.txt | head -8terminal · recorded 2026-09-10 · exit 0 - Fin
curl -si https://api.intercom.io/meterminal · recorded 2026-09-10 · exit 0 - Fin
curl -s https://developers.intercom.com/llms.txt | head -6terminal · recorded 2026-09-10 · exit 0 - Fin
curl -si -X POST https://mcp.intercom.com/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-10 · exit 0 - Fin
curl -s https://fin.ai/llms.txt | head -6terminal · recorded 2026-09-10 · exit 0 - Fin
npm view intercom-client name versionterminal · recorded 2026-09-10 · exit 0 - Lorikeet
curl -si https://api.lorikeetcx.ai/v1/customerterminal · recorded 2026-09-10 · exit 0 - Lorikeet
curl -s https://docs.lorikeetcx.ai/llms.txt | head -8terminal · recorded 2026-09-10 · exit 0 - Lorikeet
curl -s https://docs.lorikeetcx.ai/mcp/mcp-server.md | head -8terminal · recorded 2026-09-10 · exit 0 - Lorikeet
curl -si https://mcp.lorikeetcx.ai | head -8terminal · recorded 2026-09-10 · exit 0 - Parahelp
curl -s https://docs.parahelp.com/llms.txt | head -6terminal · recorded 2026-09-10 · exit 0 - Parahelp
curl -s https://docs.parahelp.com/customer-agent/api.md | head -8terminal · recorded 2026-09-10 · exit 0 - Parahelp
curl -sL https://app.parahelp.com/api/docs | head -c 400terminal · recorded 2026-09-10 · exit 0 - Pylon
curl -si https://api.usepylon.com/meterminal · recorded 2026-09-10 · exit 0 - Pylon
curl -sL https://docs.usepylon.com/pylon-docs/llms.txt | head -6terminal · recorded 2026-09-10 · exit 0 - Pylon
curl -sL https://docs.usepylon.com/pylon-docs/integrations/pylon-mcp.md | head -8terminal · recorded 2026-09-10 · exit 0 - Pylon
curl -si -X POST https://mcp.usepylon.com -H 'Content-Type: application/json' -d '<jsonrpc initialize>'terminal · recorded 2026-09-10 · exit 0 - Sierra
curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}' # finding: resolves to the login SPA's HTML shell, not an agent-legible indexterminal · recorded 2026-09-10 · exit 0 - Sierra
curl -s https://sierra.ai/sitemap.xml | grep -o '<loc>https://sierra.ai/product/[^<]*' | sort | head -12terminal · recorded 2026-09-10 · exit 0 - Sierra
curl -sL https://raw.githubusercontent.com/sierra-research/tau-bench/HEAD/README.md | head -8terminal · recorded 2026-09-10 · exit 0
Cite as: ProductArena by Ultrametric Inc, AI Customer Support Agents arena, rankings as of 2026-09-16 — https://ultrametric.ai/productarena/arena/ai-support-agents
License: © 2026 Ultrametric Inc. Brief quotation of individual verdicts, scores, or evidence excerpts is permitted with attribution to "ProductArena by Ultrametric Inc (ultrametric.ai/productarena)", as is use of the data to evaluate, contest, or contribute corrections. Bulk copying, redistribution, or use to build competing datasets requires prior written permission (see DATA-LICENSE in the repository).
No liability: rankings, verdicts, and scores are research outputs derived from the cited evidence at a point in time, provided "as is", without warranties. Ultrametric Inc accepts no responsibility for procurement, purchasing, or other decisions made in reliance on them — verify against the cited evidence before acting (https://ultrametric.ai/productarena/terms).