Access
Try itExperimental
See what an agent can do with Sierra before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}' # finding: resolves to the login SPA's HTML shell, not an agent-legible indexrecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agent actions — stories about agent actions in this arenaAgent actionsevidence →
Stories about agent actions in this arena
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Channels languages — stories about channels languages in this arenaChannels languagesevidence →
Stories about channels languages in this arena
Escalation handoff — stories about escalation handoff in this arenaEscalation handoffevidence →
Stories about escalation handoff in this arena
Guardrails safety — stories about guardrails safety in this arenaGuardrails safetyevidence →
Stories about guardrails safety in this arena
Insights analytics — stories about insights analytics in this arenaInsights analyticsevidence →
Stories about insights analytics in this arena
Integrations platform — stories about integrations platform in this arenaIntegrations platformevidence →
Stories about integrations platform in this arena
Knowledge grounding — stories about knowledge grounding in this arenaKnowledge groundingevidence →
Stories about knowledge grounding in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing economics — stories about pricing economics in this arenaPricing economicsevidence →
Stories about pricing economics in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Resolution quality — stories about resolution quality in this arenaResolution qualityevidence →
Stories about resolution quality in this arena
Testing qa — stories about testing qa in this arenaTesting qaevidence →
Stories about testing qa in this arena
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Sierra stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agent actions — stories about agent actions in this arenaAgent actions
Stories about agent actions in this arena
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
~4/10
unlocks → Webhooks · Scoped API keys · Machine-readable spec · Versioning policy · Official CLI · API/UI parity · Full data export · The agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both ways
Subscribe to events via webhooks
—0/10
Build against official SDKs
~5/10
Issue scoped/least-privilege API credentials for an agent
—0/10
Connect an agent via an official MCP server
n/an/a
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
~5/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
~6/10
unlocks → MCP client
Operate the product with natural-language commands
✓7/10
Plug MCP servers into this product so it can use their tools
—–
Get AI-generated insights and suggestions from my data inside the product
✓8/10
Set up automations that run autonomously in the background
~6/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Channels languages — stories about channels languages in this arenaChannels languages
Stories about channels languages in this arena
One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, social
~5/10
The agent supports customers in many languages, even where my knowledge base exists only in English
~6/10
The agent handles phone calls — speech in, speech out — with the same knowledge and actions as chat
✓8/10
Escalation handoff — stories about escalation handoff in this arenaEscalation handoff
Stories about escalation handoff in this arena
Guardrails safety — stories about guardrails safety in this arenaGuardrails safety
Stories about guardrails safety in this arena
Guardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guess
~4/10
Launch in a supervised mode where the agent drafts replies for human approval before anything reaches a customer
—0/10
I mark topics as human-only — legal threats, cancellations, security — and the agent never freelances on them
—0/10
Insights analytics — stories about insights analytics in this arenaInsights analytics
Stories about insights analytics in this arena
Integrations platform — stories about integrations platform in this arenaIntegrations platform
Stories about integrations platform in this arena
Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding
Stories about knowledge grounding in this arena
Knowledge stays current automatically — the agent re-syncs sources on a schedule or on change, not via manual re-uploads
—0/10
The platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questions
~5/10
Every answer is grounded in my own content and shows which article or source it drew from
✓7/10
The agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoring
~6/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing economics — stories about pricing economics in this arenaPricing economics
Stories about pricing economics in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Resolution quality — stories about resolution quality in this arenaResolution quality
Stories about resolution quality in this arena
Answers use the customer's live data — plan, order status, account history — not just generic help articles
~7/10
The agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answer
~6/10
The agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bounces
~6/10
I control the agent's tone and brand voice, and it stays consistent across topics and languages
~7/10
Testing qa — stories about testing qa in this arenaTesting qa
Stories about testing qa in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 6/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 4/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Xcommunity | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Xcommunity | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Xcommunity | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 5/10 | Cclaimed | |
Every answer is grounded in my own content and shows which article or source it drew from C Grounding | ai-native user | Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding | 3 | full | 7/10 | Cclaimed | |
The agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bounces C Resolution | support leader | Resolution quality — stories about resolution quality in this arenaResolution quality | 3 | partial | 6/10 | Tprobed | |
The agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoring C Ingestion | support ops lead | Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding | 3 | partial | 6/10 | Tprobed | |
When the agent escalates, the human gets the full conversation, a summary, and collected details — the customer never repeats themselves C Handoff | support leader | Escalation handoff — stories about escalation handoff in this arenaEscalation handoff | 3 | partial | 6/10 | Cclaimed | |
Dashboards show resolution rate, CSAT, handoff rate, and cost per resolution — the numbers I report to my exec team C Analytics | support leader | Insights analytics — stories about insights analytics in this arenaInsights analytics | 3 | partial | 5/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 5/10 | Cclaimed | |
The agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per action C Actions | developer | Agent actions — stories about agent actions in this arenaAgent actions | 3 | partial | 5/10 | Xcommunity | |
Guardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guess C Hallucination | ai-native user | Guardrails safety — stories about guardrails safety in this arenaGuardrails safety | 3 | partial | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
The agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both ways C Helpdesk | developer | Integrations platform — stories about integrations platform in this arenaIntegrations platform | 3 | none | untested | none yet | |
I encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branching C Procedures | support ops lead | Agent actions — stories about agent actions in this arenaAgent actions | 2 | full | 8/10 | Xcommunity | |
I test the agent against historical tickets or simulated conversations before it faces real customers C Simulation | support ops lead | Testing qa — stories about testing qa in this arenaTesting qa | 2 | full | 8/10 | Tprobed | |
The agent handles phone calls — speech in, speech out — with the same knowledge and actions as chat C Voice | support leader | Channels languages — stories about channels languages in this arenaChannels languages | 2 | full | 8/10 | Cclaimed | |
Answers use the customer's live data — plan, order status, account history — not just generic help articles C Personalization | support leader | Resolution quality — stories about resolution quality in this arenaResolution quality | 2 | partial | 7/10 | Xcommunity | |
The agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answer C Reasoning | support leader | Resolution quality — stories about resolution quality in this arenaResolution quality | 2 | partial | 6/10 | Xcommunity | |
The agent supports customers in many languages, even where my knowledge base exists only in English C Languages | support leader | Channels languages — stories about channels languages in this arenaChannels languages | 2 | partial | 6/10 | Cclaimed | |
One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, social C Channels | support leader | Channels languages — stories about channels languages in this arenaChannels languages | 2 | partial | 5/10 | Cclaimed | |
I configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeys C Rules | support ops lead | Escalation handoff — stories about escalation handoff in this arenaEscalation handoff | 2 | partial | 3/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
I mark topics as human-only — legal threats, cancellations, security — and the agent never freelances on them C Topic controls | support ops lead | Guardrails safety — stories about guardrails safety in this arenaGuardrails safety | 2 | none | 0/10 | ||
Knowledge stays current automatically — the agent re-syncs sources on a schedule or on change, not via manual re-uploads C Freshness | support ops lead | Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding | 2 | none | 0/10 | ||
Launch in a supervised mode where the agent drafts replies for human approval before anything reaches a customer C Supervision | support ops lead | Guardrails safety — stories about guardrails safety in this arenaGuardrails safety | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
Pricing is outcome-based and published — I pay per resolution with caps and controls, not an opaque enterprise quote G Pricing | support leader | Pricing economics — stories about pricing economics in this arenaPricing economics | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
AI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agent C Qa | support ops lead | Testing qa — stories about testing qa in this arenaTesting qa | 1 | partial | 7/10 | Cclaimed | |
I control the agent's tone and brand voice, and it stays consistent across topics and languages C Voice | support leader | Resolution quality — stories about resolution quality in this arenaResolution quality | 1 | partial | 7/10 | Cclaimed | |
The platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volume C Insights | support leader | Insights analytics — stories about insights analytics in this arenaInsights analytics | 1 | partial | 6/10 | Cclaimed | |
The platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questions C Gaps | support ops lead | Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding | 1 | partial | 5/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Tprobed |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 46 stories with headroom
What would move Sierra’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
Missing: any mention of MCP protocol support, MCP client configuration, or third-party tool server integration via MCP.
Integrations platform — stories about integrations platform in this arenaThe agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both ways
nonemoves PA Scoreimpact 30
No evidence pack citation mentions Zendesk, Salesforce, Intercom, or bidirectional ticket/context syncing with existing helpdesk platforms; only generic 'systems integrations' and 'internal APIs' are referenced without naming any helpdesk system or describing two-way ticket sync.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
Missing: any export functionality, data portability documentation, or open-format data dump capability.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Missing: any documentation of self-hosting, on-prem deployment, or open-source release of the core agent platform.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence in the pack addresses data-training opt-out or AI model training policies; Sierra's docs focus on product features (agent building, analytics, channels) with no privacy/data-use policy statements provided.
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs
nonemoves agent-readyimpact 30
Direct probes show no llms.txt at sierra.ai (404) and docs.sierra.ai/llms.txt merely resolves to the login SPA HTML shell rather than an actual plain-text agent-oriented index; the real docs are login-gated to contracted customers, so an agent cannot be pointed at a genuine llms.txt or open agent-oriented docs.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
No evidence of an official Sierra CLI tool; the Agent SDK mentions a code-based development workflow but nothing describing a CLI, and probes for llms.txt/openapi return 404s with no CLI reference anywhere in the pack.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
No evidence anywhere in the pack of scoped or least-privilege API credential/token issuance for agents; docs describe agent building, workflows, channels, and analytics but nothing about credential scoping, permissions, or API key management.
Showing the top 8 of 46 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 29 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Product docs29 stories
- The agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per action
- I encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branching
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, social
- The agent supports customers in many languages, even where my knowledge base exists only in English
- The agent handles phone calls — speech in, speech out — with the same knowledge and actions as chat
- When the agent escalates, the human gets the full conversation, a summary, and collected details — the customer never repeats themselves
- I configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeys
- Guardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guess
- Dashboards show resolution rate, CSAT, handoff rate, and cost per resolution — the numbers I report to my exec team
- The platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volume
- The platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questions
- Every answer is grounded in my own content and shows which article or source it drew from
- The agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoring
- Answers use the customer's live data — plan, order status, account history — not just generic help articles
- The agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answer
- The agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bounces
- I control the agent's tone and brand voice, and it stays consistent across topics and languages
- AI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agent
- I test the agent against historical tickets or simulated conversations before it faces real customers
Blog docs11 stories
- The agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per action
- Run the product headlessly / in CI for automation
- Test against a sandbox environment without touching production data
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, social
- I configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeys
- Guardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guess
- I control the agent's tone and brand voice, and it stays consistent across topics and languages
- AI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agent
- I test the agent against historical tickets or simulated conversations before it faces real customers
Hacker News8 stories
- The agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per action
- I encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branching
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Answers use the customer's live data — plan, order status, account history — not just generic help articles
- The agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answer
- The agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bounces
docs.sierra.ai5 stories
OpenAPI spec3 stories
llms.txt2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}' # finding: resolves to the login SPA's HTML shell, not an agent-legible indexreproduced$ curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}' # finding: resolves to the login SPA's HTML shell, not an agent-legible index
llms.txt: HTTP 200 text/html; charset=utf-8
$curl -s https://sierra.ai/sitemap.xml | grep -o '<loc>https://sierra.ai/product/[^<]*' | sort | head -12reproduced$ curl -s https://sierra.ai/sitemap.xml | grep -o '<loc>https://sierra.ai/product/[^<]*' | sort | head -12 <loc>https://sierra.ai/product/agent-data-platform <loc>https://sierra.ai/product/agent-sdk <loc>https://sierra.ai/product/agent-studio <loc>https://sierra.ai/product/channels <loc>https://sierra.ai/product/explorer <loc>https://sierra.ai/product/ghostwriter <loc>https://sierra.ai/product/horizon <loc>https://sierra.ai/product/insights <loc>https://sierra.ai/product/live-assist <loc>https://sierra.ai/product/meet-your-agent <loc>https://sierra.ai/product/trust-and-reliability <loc>https://sierra.ai/product/voice
$curl -sL https://raw.githubusercontent.com/sierra-research/tau-bench/HEAD/README.md | head -8reproduced$ curl -sL https://raw.githubusercontent.com/sierra-research/tau-bench/HEAD/README.md | head -8 # τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains **⚠️ WARNING: The tasks in this repo are not updated.** This repository contains outdated versions of the airline and retail tasks. Please use [τ³-bench](https://github.com/sierra-research/tau2-bench) for the latest fixed tasks and new domains. **❗News**: The [τ²-bench](https://github.com/sierra-research/tau2-bench) repository has been updated to [τ³-bench](https://github.com/sierra-research/tau2-bench), which includes a new `banking` domain, a `voice` evaluation modality, as well as fixes to the `airline` and `retail` domain tasks. Please navigate to the [τ³-bench repository](https://github.com/sierra-research/tau2-bench) to use the latest version of this benchmark. ---
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
7 of 14 testable claims verified · 1 contradicted → integrity 36/100
21 distinct capability claims found in Sierra’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
7
Verified
6
Unverified
1
Contradicted
16
Undersold
Verified (10)
“Journeys can be written and versioned as code, fitting normal dev workflows”
“Run scenario tests to verify agent behavior and catch regressions before release”
I test the agent against historical tickets or simulated conversations before it faces real customersfullproof ↗
“Define step-by-step agent workflows manually or auto-generate them with AI”
I encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branchingfullproof ↗
“Manage and edit knowledge sources like Help Center content, FAQs, and policies that ground the agent”
The agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoringpartialproof ↗
“Modify agent behavior, integrations, guardrails, and tone using natural-language prompts”
Operate the product with natural-language commandsfullproof ↗
“Create customer journeys automatically from uploaded SOPs, transcripts, or audio interviews”
I encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branchingfullproof ↗
“Ask natural-language questions about customer experience and get answers pulled from real conversations”
Get AI-generated insights and suggestions from my data inside the productfullproof ↗
“Agent Checks and Simulations proactively catch problems before release”
I test the agent against historical tickets or simulated conversations before it faces real customersfullproof ↗
“Merge approval workflows and gradual split-traffic releases control how agent changes roll out”
“Agent can collect card and ACH payments entirely over the phone without IVR handoff”
The agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per actionpartialproof ↗
Unverified (9)
“Build an agent once and deploy it across chat, phone, email, SMS, and messaging channels”
One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, socialpartialproof ↗
“Automatically generates a weekly briefing on trends, issues, and recommendations”
The platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volumepartialproof ↗
“Click any report data point to drill into the conversations driving that trend”
The platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volumepartialproof ↗
“Every agent action or answer shows its reasoning, knowledge sources, and systems accessed”
Every answer is grounded in my own content and shows which article or source it drew fromfullproof ↗
“Guides human reps through next steps and auto-captures customer context during chat or calls”
When the agent escalates, the human gets the full conversation, a summary, and collected details — the customer never repeats themselvespartialproof ↗
“Voice Personas let you customize how the voice agent sounds and speaks, across 59 languages”
The agent handles phone calls — speech in, speech out — with the same knowledge and actions as chatfullproof ↗
“Voice Personas let you customize how the voice agent sounds and speaks, across 59 languages”
The agent supports customers in many languages, even where my knowledge base exists only in Englishpartialproof ↗
“Publish an agent to ChatGPT with one click or via CI/CD pipeline”
One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, socialpartialproof ↗
“Voice agent replaces rigid IVR menus with an empathetic, context-aware phone experience”
The agent handles phone calls — speech in, speech out — with the same knowledge and actions as chatfullproof ↗
Contradicted (1)
“Reps can trigger workflows directly from a conversation without switching tabs or losing context”
The agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both waysnone
Undersold (16)
Run the product headlessly / in CI for automationpartialproof ↗
Drive the product through a documented public APIpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Delegate tasks to a built-in AI assistant inside the productpartialproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
I configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeyspartialproof ↗
Guardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guesspartialproof ↗
Dashboards show resolution rate, CSAT, handoff rate, and cost per resolution — the numbers I report to my exec teampartialproof ↗
The platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questionspartialproof ↗
Answers use the customer's live data — plan, order status, account history — not just generic help articlespartialproof ↗
The agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answerpartialproof ↗
The agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bouncespartialproof ↗
I control the agent's tone and brand voice, and it stays consistent across topics and languagespartialproof ↗
AI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agentpartialproof ↗
Claims outside our story set (2)
Real capability claims found in Sierra’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Lets you inspect API calls and logic traces to debug and adjust agent behavior”
source ↗“Non-technical teams can build and manage agents with no code required”
source ↗
Business model
Enterprise sales only, no public price list. Sierra champions outcome-based pricing — pay per resolved conversation (typically nothing if unresolved), blended with consumption pricing for routing-style interactions.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents