Skip to content

Rank #5 of 6 in AI Customer Support Agents

Sierra Technologies, Inc. · commercial

no public signals

Try itExperimental

See what an agent can do with Sierra before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}' # finding: resolves to the login SPA's HTML shell, not an agent-legible indexrecorded session — replayed, not live
recorded 2026-09-10 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agent actions — stories about agent actions in this arenaAgent actionsevidence →

Stories about agent actions in this arena

50.0/100

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

20.3/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

15.0/100

Channels languages — stories about channels languages in this arenaChannels languagesevidence →

Stories about channels languages in this arena

48.7/100

Escalation handoff — stories about escalation handoff in this arenaEscalation handoffevidence →

Stories about escalation handoff in this arena

28.8/100

Guardrails safety — stories about guardrails safety in this arenaGuardrails safetyevidence →

Stories about guardrails safety in this arena

10.3/100

Insights analytics — stories about insights analytics in this arenaInsights analyticsevidence →

Stories about insights analytics in this arena

31.5/100

Integrations platform — stories about integrations platform in this arenaIntegrations platformevidence →

Stories about integrations platform in this arena

0.0/100

Knowledge grounding — stories about knowledge grounding in this arenaKnowledge groundingevidence →

Stories about knowledge grounding in this arena

38.7/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

0.0/100

Pricing economics — stories about pricing economics in this arenaPricing economicsevidence →

Stories about pricing economics in this arena

0.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Resolution quality — stories about resolution quality in this arenaResolution qualityevidence →

Stories about resolution quality in this arena

38.2/100

Testing qa — stories about testing qa in this arenaTesting qaevidence →

Stories about testing qa in this arena

67.3/100

Story verdicts — every judged story with its evidenceStory verdicts

?

Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3partial6/10C

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3partial4/10T

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3noneuntestednone yet

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10X

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10X

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10X

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10T

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10T

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial5/10C

Every answer is grounded in my own content and shows which article or source it drew from C

Grounding

ai-native userKnowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding3full7/10C

The agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bounces C

Resolution

support leaderResolution quality — stories about resolution quality in this arenaResolution quality3partial6/10T

The agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoring C

Ingestion

support ops leadKnowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding3partial6/10T

When the agent escalates, the human gets the full conversation, a summary, and collected details — the customer never repeats themselves C

Handoff

support leaderEscalation handoff — stories about escalation handoff in this arenaEscalation handoff3partial6/10C

Dashboards show resolution rate, CSAT, handoff rate, and cost per resolution — the numbers I report to my exec team C

Analytics

support leaderInsights analytics — stories about insights analytics in this arenaInsights analytics3partial5/10C

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial5/10C

The agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per action C

Actions

developerAgent actions — stories about agent actions in this arenaAgent actions3partial5/10X

Guardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guess C

Hallucination

ai-native userGuardrails safety — stories about guardrails safety in this arenaGuardrails safety3partial4/10C

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

The agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both ways C

Helpdesk

developerIntegrations platform — stories about integrations platform in this arenaIntegrations platform3noneuntestednone yet

I encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branching C

Procedures

support ops leadAgent actions — stories about agent actions in this arenaAgent actions2full8/10X

I test the agent against historical tickets or simulated conversations before it faces real customers C

Simulation

support ops leadTesting qa — stories about testing qa in this arenaTesting qa2full8/10T

The agent handles phone calls — speech in, speech out — with the same knowledge and actions as chat C

Voice

support leaderChannels languages — stories about channels languages in this arenaChannels languages2full8/10C

Answers use the customer's live data — plan, order status, account history — not just generic help articles C

Personalization

support leaderResolution quality — stories about resolution quality in this arenaResolution quality2partial7/10X

The agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answer C

Reasoning

support leaderResolution quality — stories about resolution quality in this arenaResolution quality2partial6/10X

The agent supports customers in many languages, even where my knowledge base exists only in English C

Languages

support leaderChannels languages — stories about channels languages in this arenaChannels languages2partial6/10C

One agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, social C

Channels

support leaderChannels languages — stories about channels languages in this arenaChannels languages2partial5/10C

I configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeys C

Rules

support ops leadEscalation handoff — stories about escalation handoff in this arenaEscalation handoff2partial3/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2none0/10

I mark topics as human-only — legal threats, cancellations, security — and the agent never freelances on them C

Topic controls

support ops leadGuardrails safety — stories about guardrails safety in this arenaGuardrails safety2none0/10

Knowledge stays current automatically — the agent re-syncs sources on a schedule or on change, not via manual re-uploads C

Freshness

support ops leadKnowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding2none0/10

Launch in a supervised mode where the agent drafts replies for human approval before anything reaches a customer C

Supervision

support ops leadGuardrails safety — stories about guardrails safety in this arenaGuardrails safety2none0/10

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2noneuntestednone yet

Pricing is outcome-based and published — I pay per resolution with caps and controls, not an opaque enterprise quote G

Pricing

support leaderPricing economics — stories about pricing economics in this arenaPricing economics2noneuntestednone yet

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2noneuntestednone yet

AI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agent C

Qa

support ops leadTesting qa — stories about testing qa in this arenaTesting qa1partial7/10C

I control the agent's tone and brand voice, and it stays consistent across topics and languages C

Voice

support leaderResolution quality — stories about resolution quality in this arenaResolution quality1partial7/10C

The platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volume C

Insights

support leaderInsights analytics — stories about insights analytics in this arenaInsights analytics1partial6/10C

The platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questions C

Gaps

support ops leadKnowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding1partial5/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial5/10T

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 46 stories with headroom

What would move Sierra’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools

    nonemoves agent-readyimpact 45

    Missing: any mention of MCP protocol support, MCP client configuration, or third-party tool server integration via MCP.

  2. Integrations platform — stories about integrations platform in this arenaThe agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both ways

    nonemoves PA Scoreimpact 30

    No evidence pack citation mentions Zendesk, Salesforce, Intercom, or bidirectional ticket/context syncing with existing helpdesk platforms; only generic 'systems integrations' and 'internal APIs' are referenced without naming any helpdesk system or describing two-way ticket sync.

  3. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    Missing: any export functionality, data portability documentation, or open-format data dump capability.

  4. Openness — open source, data portability, and self-hosting storiesSelf-host the core product

    nonemoves PA Scoreimpact 30

    Missing: any documentation of self-hosting, on-prem deployment, or open-source release of the core agent platform.

  5. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    No evidence in the pack addresses data-training opt-out or AI model training policies; Sierra's docs focus on product features (agent building, analytics, channels) with no privacy/data-use policy statements provided.

  6. Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs

    nonemoves agent-readyimpact 30

    Direct probes show no llms.txt at sierra.ai (404) and docs.sierra.ai/llms.txt merely resolves to the login SPA HTML shell rather than an actual plain-text agent-oriented index; the real docs are login-gated to contracted customers, so an agent cannot be pointed at a genuine llms.txt or open agent-oriented docs.

  7. Agenticness — how well agents can access and operate the productUse an official CLI

    nonemoves agent-readyimpact 30

    No evidence of an official Sierra CLI tool; the Agent SDK mentions a code-based development workflow but nothing describing a CLI, and probes for llms.txt/openapi return 404s with no CLI reference anywhere in the pack.

  8. Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent

    nonemoves agent-readyimpact 30

    No evidence anywhere in the pack of scoped or least-privilege API credential/token issuance for agents; docs describe agent building, workflows, channels, and analytics but nothing about credential scoping, permissions, or API key management.

Showing the top 8 of 46 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 29 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Product docs29 stories

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}' # finding: resolves to the login SPA's HTML shell, not an agent-legible indexreproduced
$ curl -sL https://docs.sierra.ai/llms.txt -o /dev/null -w 'llms.txt: HTTP %{http_code} %{content_type}'  # finding: resolves to the login SPA's HTML shell, not an agent-legible index
llms.txt: HTTP 200 text/html; charset=utf-8
$curl -s https://sierra.ai/sitemap.xml | grep -o '<loc>https://sierra.ai/product/[^<]*' | sort | head -12reproduced
$ curl -s https://sierra.ai/sitemap.xml | grep -o '<loc>https://sierra.ai/product/[^<]*' | sort | head -12
<loc>https://sierra.ai/product/agent-data-platform
<loc>https://sierra.ai/product/agent-sdk
<loc>https://sierra.ai/product/agent-studio
<loc>https://sierra.ai/product/channels
<loc>https://sierra.ai/product/explorer
<loc>https://sierra.ai/product/ghostwriter
<loc>https://sierra.ai/product/horizon
<loc>https://sierra.ai/product/insights
<loc>https://sierra.ai/product/live-assist
<loc>https://sierra.ai/product/meet-your-agent
<loc>https://sierra.ai/product/trust-and-reliability
<loc>https://sierra.ai/product/voice
$curl -sL https://raw.githubusercontent.com/sierra-research/tau-bench/HEAD/README.md | head -8reproduced
$ curl -sL https://raw.githubusercontent.com/sierra-research/tau-bench/HEAD/README.md | head -8
# τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

**⚠️ WARNING: The tasks in this repo are not updated.** This repository contains outdated versions of the airline and retail tasks. Please use [τ³-bench](https://github.com/sierra-research/tau2-bench) for the latest fixed tasks and new domains.

**❗News**: The [τ²-bench](https://github.com/sierra-research/tau2-bench) repository has been updated to [τ³-bench](https://github.com/sierra-research/tau2-bench), which includes a new `banking` domain, a `voice` evaluation modality, as well as fixes to the `airline` and `retail` domain tasks. Please navigate to the [τ³-bench repository](https://github.com/sierra-research/tau2-bench) to use the latest version of this benchmark.

---

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

7 of 14 testable claims verified · 1 contradictedintegrity 36/100

21 distinct capability claims found in Sierra’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

7

Verified

6

Unverified

1

Contradicted

16

Undersold

Verified (10)
Unverified (9)
Contradicted (1)
Undersold (16)
Claims outside our story set (2)

Real capability claims found in Sierra’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Lets you inspect API calls and logic traces to debug and adjust agent behavior

    source ↗
  • Non-technical teams can build and manage agents with no code required

    source ↗
Suggest a story for these →

Business model

usage-basedenterprise-custom

Enterprise sales only, no public price list. Sierra champions outcome-based pricing — pay per resolved conversation (typically nothing if unresolved), blended with consumption pricing for routing-style interactions.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Scoretracked since Sep 10 '26 — no movement recorded yet
Agent-readytracked since Sep 10 '26 — no movement recorded yet

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data