Sierra vs Lorikeet
usage-based · enterprise-custom
·usage-based · subscription-flat · enterprise-custom
Lorikeet wins · 8–25 (16 drawn)
Agent actions — stories about agent actions in this arenaAgent actions
Stories about agent actions in this arena
Actions
developerThe agent takes real actions through my APIs — refunds, order changes, subscription updates — with scoped auth per action
weight 3 · round to LorikeetSierra's Agent SDK/Studio docs show agents can call customer APIs to execute actions (order changes, refunds implied by 'internal APIs' use), with guardrails, release governance, and human-in-the-loop approval, and a HN commenter confirms it wires directly into a customer's internal APIs. However there's no explicit documentation of per-action scoped auth/permissioning model for API calls. missing for 10: explicit scoped-auth/permission model per action, concrete refund/subscription action examples, independent confirmation of granular auth scoping.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Understand and rapidly adjust agent behavior by inspecting API calls, logic traces, and more.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
Docs explicitly state the agent 'takes action through your APIs and MCP servers' (docs-28), give concrete examples like refund workflows and order-status tool tests (docs-19, docs-24, docs-43), and describe workflow-scoped tool access, server-side identity validation, and hard execution caps enforced in code rather than prompts (docs-34). This directly matches the story of scoped, real-world API actions like refunds/order changes/subscriptions. Missing for 10: independent/hands-on verification of the auth-scoping mechanism, and more granular detail on how 'per action' scopes are configured/enforced beyond high-level guardrails language.
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
- [claimed-docs] “Customer isolation, server-side identity validation, workflow-scoped tool access and hard execution caps are enforced in code rather than pr…”
- [claimed-docs] “Test the 'get-order-status' tool with order ID 98765 and check if the response matches what we expect.”
- [claimed-docs] “/lorikeet:create-simulations build simulations for the refund workflow”
- [claimed-docs] “Automate disputes, loan inquiries, payments and compliance cases across every customer channel”
- [claimed-docs] “Connect Lorikeet in seconds to your ticketing system, knowledge base, and internal tools to seamlessly ingest data and take action for your …”
Procedures
support ops leadI encode standard operating procedures the agent follows step-by-step for known issue types, with deterministic branching
weight 2 · round to SierraSierra's Agent SDK/Studio explicitly supports authoring step-by-step journeys ('workflows') as code or no-code, with branching logic ('sophisticated logic'), SOP ingestion via Ghostwriter, and simulation/testing to verify deterministic behavior across scenarios. This directly matches encoding SOPs with deterministic branching for known issue types. missing for 10: independent/hands-on verification of deterministic branching behavior specifically (community evidence is generic, not focused on SOP branching), and docs are login-gated so full workflow-editor detail isn't independently viewable.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Define step-by-step workflows for your agent from scratch, or use AI to instantly generate journeys from your existing operating procedures.”
- [claimed-docs] “Upload SOPs, raw transcripts, or even audio interviews of subject matter experts to create customer journeys from scratch.”
- [claimed-docs] “Verify your agent performs as expected across a wide range of scenarios and avoid regressions.”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
Lorikeet explicitly supports training agents on SOPs (docs-25) and building workflows with 'natural-language and deterministic workflows' plus 'pockets of determinism for regulated steps' (docs-28, docs-32), with workflow-scoped tool access and hard execution caps enforced in code (docs-34), directly matching the deterministic-branching SOP story. Missing for 10: independent/hands-on validation of branching logic in practice and more detail on how branching conditions are authored beyond natural-language workflow builder claims.
- [claimed-docs] “train the agent on your business context, brand guidelines, help docs and standard operating procedures”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
- [claimed-docs] “Lorikeet coordinates a team of specialist agents, with pockets of determinism for regulated steps, to handle multi-party, multi-system workf…”
- [claimed-docs] “Customer isolation, server-side identity validation, workflow-scoped tool access and hard execution caps are enforced in code rather than pr…”
- [claimed-docs] “Build workflows - create and iterate on workflows using natural language”
- [claimed-docs] “Build, edit, and deploy workflows using natural language”
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to LorikeetSierranone0/10Direct probes show no llms.txt at sierra.ai (404) and docs.sierra.ai/llms.txt merely resolves to the login SPA HTML shell rather than an actual plain-text agent-oriented index; the real docs are login-gated to contracted customers, so an agent cannot be pointed at a genuine llms.txt or open agent-oriented docs.
Direct probe evidence confirms Lorikeet serves an llms.txt file at docs.lorikeetcx.ai/llms.txt returning HTTP 200 with structured agent-oriented reference links, and documentation is further organized around MCP/agent access. Missing for 10: no independent third-party report of an agent successfully consuming this llms.txt in practice.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.lorikeetcx.ai/llms.txt # Reference - [Lorikeet MCP Server](https://docs.lorikeetcx.ai/mcp/mcp-serv…”
- [probe] “official MCP server documented at https://docs.lorikeetcx.ai/mcp/mcp-server”
- [claimed-docs] “Connect to your Lorikeet account from ChatGPT, Claude, Claude Code, and MintMCP using the Model Context Protocol.”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round to LorikeetSierranone0/10No evidence anywhere in the pack mentions MCP servers or the ability to plug external MCP tool servers into Sierra's agents; Sierra's integration mentions are about internal APIs and custom systems, not MCP. missing for 10: any mention of MCP protocol support, MCP client configuration, or third-party tool server integration via MCP.
Lorikeet's docs state the agent 'takes action through your APIs and MCP servers inside natural-language and deterministic workflows' (lorikeet-docs-28), indicating it can consume external MCP servers as tool sources, but the bulk of the MCP evidence pack actually describes the reverse direction — Lorikeet exposing its own MCP server for external clients like Claude/ChatGPT to connect to (lorikeet-docs-1, lorikeet-docs-46, lorikeet-probe-1/2). Missing for 10: dedicated documentation on how a user configures/adds third-party MCP servers into Lorikeet, a list of supported MCP integrations, and independent confirmation of this client-side tool-use capability.
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
- [claimed-docs] “Connect Lorikeet in seconds to your ticketing system, knowledge base, and internal tools to seamlessly ingest data and take action for your …”
ai-native userUse an official CLI
weight 2 · round drawnSierranone0/10No evidence of an official Sierra CLI tool; the Agent SDK mentions a code-based development workflow but nothing describing a CLI, and probes for llms.txt/openapi return 404s with no CLI reference anywhere in the pack.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
Lorikeetnone0/10Lorikeet is a customer-support AI platform; the story asks for an official CLI for AI-native workflows. Evidence shows an MCP server and Claude-code skills/slash commands but no standalone CLI tool is documented anywhere in the pack. missing for 10: any mention of a CLI binary/tool, installation instructions, or CLI command reference.
- [claimed-docs] “The Lorikeet MCP server lets you interact with your Lorikeet account directly from Claude Code, Claude.ai, ChatGPT, and Codex.”
- [claimed-docs] “/lorikeet:create-simulations build simulations for the refund workflow”
- [probe] “official MCP server documented at https://docs.lorikeetcx.ai/mcp/mcp-server”
ai-native userDrive the product through a documented public API
weight 3 · round to LorikeetSierra markets an Agent SDK for building 'customer journeys as code' with API-call inspection, implying a programmatic interface exists, but there's no publicly discoverable OpenAPI spec, llms.txt, or open API reference — docs.sierra.ai is login-gated to contracted customers rather than a documented public API. Missing for 10: publicly accessible API reference/OpenAPI spec, evidence of self-serve API keys or open documentation, independent developer confirmation of using the API without a sales contract.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Understand and rapidly adjust agent behavior by inspecting API calls, logic traces, and more.”
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeet publishes a documented MCP server (docs.lorikeetcx.ai/mcp/mcp-server) that lets AI-native users drive the product directly from Claude, ChatGPT, Codex, and Claude Code — diagnosing tickets, auditing knowledge bases, building workflows, testing tools, and running simulations, all documented with concrete examples and even slash-command skills. This is a genuine documented programmatic interface built for AI agents, not just human UI docs. Missing for 10: no separate traditional REST/GraphQL API reference beyond MCP, and no independent third-party corroboration of the API's reliability.
- [claimed-docs] “The Lorikeet MCP server lets you interact with your Lorikeet account directly from Claude Code, Claude.ai, ChatGPT, and Codex.”
- [claimed-docs] “Diagnose tickets - trace workflow execution and identify root causes”
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Build workflows - create and iterate on workflows using natural language”
- [claimed-docs] “Test tools - run and validate tool configurations directly from your AI assistant”
- [claimed-docs] “Test the 'get-order-status' tool with order ID 98765 and check if the response matches what we expect.”
- [claimed-docs] “Explore your setup - inspect workflows, tools, and integrations”
- [claimed-docs] “Connect to your Lorikeet account from ChatGPT, Claude, Claude Code, and MintMCP using the Model Context Protocol.”
- [claimed-docs] “Run simulations - test workflows against different customer scenarios”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.lorikeetcx.ai/llms.txt # Reference - [Lorikeet MCP Server](https://docs.lorikeetcx.ai/mcp/mcp-serv…”
- [probe] “official MCP server documented at https://docs.lorikeetcx.ai/mcp/mcp-server”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round to LorikeetSierranone0/10No evidence anywhere in the pack of scoped or least-privilege API credential/token issuance for agents; docs describe agent building, workflows, channels, and analytics but nothing about credential scoping, permissions, or API key management. Probes even show no public OpenAPI/API docs are accessible (sierra-probe-2, sierra-probe-rt-3), reinforcing the absence of evidence.
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeet's guardrails page mentions 'workflow-scoped tool access' and 'server-side identity validation' enforced in code, suggesting some access scoping, but there is no explicit documentation of issuing or managing scoped/least-privilege API credentials or tokens for agents. Missing for 10: explicit credential/token issuance mechanism, documentation of API key scoping or permission granularity, and any user-facing controls for creating least-privilege credentials.
- [claimed-docs] “Customer isolation, server-side identity validation, workflow-scoped tool access and hard execution caps are enforced in code rather than pr…”
ai-native userBuild against official SDKs
weight 2 · round to SierraSierra advertises an official "Agent SDK" for writing customer journeys as code with logic traces and testing (sierra-docs-1/2/4), but runtime probes show no public OpenAPI/swagger spec and the SDK/docs reference is login-gated to contracted customers rather than openly available to any AI-native developer (sierra-probe-2, sierra-probe-rt-3). Missing for 10: publicly accessible API reference/OpenAPI spec, evidence of open sign-up or trial SDK access, and independent developer accounts of building against it outside a paid contract.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Understand and rapidly adjust agent behavior by inspecting API calls, logic traces, and more.”
- [claimed-docs] “Verify your agent performs as expected across a wide range of scenarios and avoid regressions.”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeetnone0/10Evidence only documents an MCP server and integrations/APIs for connecting tools, but there is no mention of an official SDK (e.g., Python/JS client library) for developers to build against. Missing for 10: any documented official SDK, its language support, or developer-facing library docs.
ai-native userSubscribe to events via webhooks
weight 2 · round drawnSierranone0/10No evidence in the pack mentions webhooks or event subscription capabilities for Sierra; docs pages describe SDK, studio, insights, voice, and channels features but nothing about outbound event notifications. Probes further show no public API spec or open documentation to confirm such a mechanism exists.
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round drawnSierra's Explorer and Insights products explicitly deliver AI-generated insights: natural-language querying across conversations, automatic weekly briefings on trends/issues with recommendations, and click-to-investigate drill-downs on report data, plus explainability of agent reasoning. This is a first-party documented feature set directly matching the story, with some independent corroboration of Sierra's data-driven production use (sierra-comm-2/3). Missing for 10: independent hands-on validation of the insights/briefing feature specifically (community evidence is about agent setup, not analytics quality), and no detail on data freshness/accuracy limits.
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
- [claimed-docs] “Explorer automatically delivers a weekly briefing on key trends, emerging issues, and recommendations — without you having to ask.”
- [claimed-docs] “Click any data point on a report to launch Explorer and instantly investigate what's driving that trend.”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
- [community] “The magic isn't in a new LLM technology, it is in reliably productionizing a solution for real-world problems... fill in the gap of missing …”
Lorikeet's Coach and MCP-server capabilities generate AI-driven insights directly from customer data: ticket quality scoring across 100% of conversations, knowledge-base gap/quality audits, root-cause diagnosis of tickets, and analytics on resolution quality, satisfaction, and revenue impact, with Coach able to 'implement improvements... or make suggestions for you to action yourself.' This is all first-party documentation without independent hands-on corroboration. Missing for 10: independent/third-party validation of insight quality and real-world usage examples beyond vendor docs.
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
- [claimed-docs] “Diagnose tickets by tracing workflow execution and identifying root causes”
- [claimed-docs] “Track resolution quality, customer satisfaction, revenue impact, and operational efficiency with industry-leading analytics.”
- [claimed-docs] “Coach can implement improvements on your behalf, or make suggestions for you to action yourself.”
- [claimed-docs] “Coach reviews 100% of tickets against your quality standards and turns findings into fixes, so your AI agent improves every week instead of …”
- [claimed-docs] “Scale and optimize with conversational insights and analytics from Lorikeet Coach”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to LorikeetSierra's agents are built to operate autonomously across channels (chat, phone, email, SMS) once deployed, and Explorer explicitly runs in the background to deliver automatic weekly briefings and trend detection 'without you having to ask,' which is genuine unprompted automation. However, there's no documentation of user-configurable scheduled tasks/triggers beyond the always-on conversational agent and the one named automatic-briefing feature. Missing for 10: explicit support for user-defined scheduled/triggered background jobs beyond conversation handling and briefings, and independent/hands-on confirmation of autonomous background execution reliability.
- [claimed-docs] “Explorer automatically delivers a weekly briefing on key trends, emerging issues, and recommendations — without you having to ask.”
- [claimed-docs] “Build once and deploy across any channel—chat, phone, email, SMS, and messaging.”
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
Lorikeet's core product is an autonomous agent that resolves tickets end-to-end across channels and coordinates specialist agents for multi-step workflows, and Outbound campaigns run on scheduled cadences without human intervention, while Coach can autonomously implement improvements. However, the evidence centers on the vendor's own agent running in background rather than a user-configurable 'automation' builder with triggers/schedules exposed as a general-purpose feature. Missing for 10: explicit user-facing scheduling/trigger configuration UI or API for arbitrary automations, and independent/hands-on confirmation that these automations run unattended reliably.
- [claimed-docs] “one agent that resolves issues end-to-end across chat, email, voice and SMS”
- [claimed-docs] “Lorikeet coordinates a team of specialist agents, with pockets of determinism for regulated steps, to handle multi-party, multi-system workf…”
- [claimed-docs] “campaign cadences with scheduling windows control when outreach happens”
- [claimed-docs] “Coach can implement improvements on your behalf, or make suggestions for you to action yourself.”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to LorikeetSierra's Ghostwriter feature lets users delegate agent-building tasks to a built-in AI assistant via natural-language prompts (build/modify agents, generate journeys from SOPs/transcripts), fitting the 'delegate tasks to built-in AI assistant' story. However, this is vendor-only documentation with no independent/hands-on corroboration of Ghostwriter specifically; community evidence instead describes manual point-and-click agent setup, not AI-assistant delegation. Missing for 10: independent/hands-on validation of Ghostwriter's delegation capability, and detail on scope/limits of what can be delegated.
- [claimed-docs] “Build or modify agents by describing how you want them to behave. Update workflows, systems integrations, guardrails, tone, and style with s…”
- [claimed-docs] “Upload SOPs, raw transcripts, or even audio interviews of subject matter experts to create customer journeys from scratch.”
- [claimed-docs] “Define step-by-step workflows for your agent from scratch, or use AI to instantly generate journeys from your existing operating procedures.”
Lorikeet's 'Coach' is a built-in AI assistant accessible directly inside the Lorikeet platform (as well as via Slack/Claude/ChatGPT/MCP) that users can delegate tasks to — diagnosing tickets, auditing knowledge bases, building workflows via natural language, running simulations, and even implementing improvements automatically on the user's behalf. This is well documented across multiple first-party pages describing concrete delegated actions (e.g. doc-48 'Coach can implement improvements on your behalf'). missing for 10: independent/hands-on verification of Coach's assistant behavior, and clearer distinction between autonomous action vs. suggestion-only mode.
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Talk to Coach wherever you work, whether in Lorikeet, Slack, Claude, ChatGPT, or via MCP.”
- [claimed-docs] “Diagnose tickets by tracing workflow execution and identifying root causes”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
- [claimed-docs] “Build, edit, and deploy workflows using natural language”
- [claimed-docs] “Coach reviews 100% of tickets against your quality standards and turns findings into fixes, so your AI agent improves every week instead of …”
- [claimed-docs] “Coach can implement improvements on your behalf, or make suggestions for you to action yourself.”
ai-native userOperate the product with natural-language commands
weight 2 · round to LorikeetSierra provides explicit natural-language control surfaces: Ghostwriter lets users 'build or modify agents by describing how you want them to behave' with 'simple prompts' for workflows, integrations, guardrails, tone and style, and Explorer lets users 'ask any question about your customer experience in natural language.' This directly matches an AI-native user operating the product via NL commands rather than only clicking through UI. Missing for 10: independent hands-on corroboration of NL command reliability/scope beyond vendor docs, and no evidence of NL control over the entire platform (e.g., release governance, channel deployment) rather than just Ghostwriter/Explorer.
- [claimed-docs] “Build or modify agents by describing how you want them to behave. Update workflows, systems integrations, guardrails, tone, and style with s…”
- [claimed-docs] “Upload SOPs, raw transcripts, or even audio interviews of subject matter experts to create customer journeys from scratch.”
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
Lorikeet ships an MCP server plus Claude/ChatGPT/Codex integrations that let users build workflows, diagnose tickets, audit knowledge bases, and run simulations using natural-language commands (e.g. 'Build workflows - create and iterate on workflows using natural language', example prompts like 'Test the get-order-status tool...'), and even exposes slash-command skills like /lorikeet:create-simulations. This is first-party documentation only, with no independent/hands-on corroboration of the NL command experience. Missing for 10: independent user reports or demos validating the natural-language MCP workflow in practice.
- [claimed-docs] “The Lorikeet MCP server lets you interact with your Lorikeet account directly from Claude Code, Claude.ai, ChatGPT, and Codex.”
- [claimed-docs] “Build workflows - create and iterate on workflows using natural language”
- [claimed-docs] “Build, edit, and deploy workflows using natural language”
- [claimed-docs] “Test the 'get-order-status' tool with order ID 98765 and check if the response matches what we expect.”
- [claimed-docs] “/lorikeet:create-simulations build simulations for the refund workflow”
- [claimed-docs] “Connect to your Lorikeet account from ChatGPT, Claude, Claude Code, and MintMCP using the Model Context Protocol.”
- [probe] “official MCP server documented at https://docs.lorikeetcx.ai/mcp/mcp-server”
Api quality
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnSierranone0/10Direct probes show no OpenAPI/Swagger spec at common paths (all 404) and no machine-readable llms.txt index; docs are login-gated rather than publicly exposing a spec. missing for 10: publicly downloadable OpenAPI/Swagger file, any machine-readable API spec endpoint.
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
ai-native userTest against a sandbox environment without touching production data
weight 1 · round drawnSierra's docs mention scenario testing and regression avoidance (sierra-docs-4) and 'Agent Checks and Simulations' for proactive problem catching (sierra-docs-18), implying some pre-production testing capability, but there is no explicit mention of a dedicated sandbox environment isolated from production data. Missing for 10: explicit sandbox/staging environment documentation, confirmation that test runs don't touch production data, and independent/hands-on verification of this separation.
- [claimed-docs] “Verify your agent performs as expected across a wide range of scenarios and avoid regressions.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
Evidence shows robust simulation/testing tooling (replay historical tickets, synthetic scenarios, guardrail adversarial tests) that approximates sandbox-style testing, but no explicit claim of an isolated sandbox environment distinct from production. missing for 10: explicit documentation of a dedicated sandbox/staging environment, confirmation that simulations do not touch or affect live production data/systems, and independent verification of this isolation.
- [claimed-docs] “Generate simulations straight from your actual tickets and run them in bulk batches, so every workflow change is tested against the conversa…”
- [claimed-docs] “Run simulation batches to test workflows against different customer scenarios”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps, then deploy on the topics …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps”
- [claimed-docs] “Author adversarial scenarios such as false authority claims, mid-conversation goal switches and prompt injection attempts, then run them as …”
- [claimed-docs] “Run simulations - test workflows against different customer scenarios”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnSierranone0/10No evidence of a public versioned API reference or a documented deprecation policy; probes show no OpenAPI spec and docs are login-gated, and llms.txt/openapi.json all 404 or resolve to a login shell rather than API docs.
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeetnone0/10No evidence pack item mentions API versioning, version numbers, or a deprecation policy for Lorikeet's APIs or MCP server; the docs discuss features and integrations but not API lifecycle/versioning commitments. Missing for 10: any documented API version scheme, changelog of breaking changes, or stated deprecation/support timeline.
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to LorikeetSierranone0/10No evidence describes bulk operations across many items (e.g., batch editing knowledge entries, mass workflow updates, or bulk conversation actions). The docs describe individual agent building, knowledge editing, and analytics tools, but nothing about performing actions at scale across many items simultaneously.
Lorikeet documents multiple bulk operations available to AI-native users via its MCP server and product surfaces: running simulations in bulk batches across hundreds of scenarios, auditing entire knowledge bases at scale, and reviewing 100% of conversations for quality (not manual spot checks). These are explicitly framed as batch/bulk actions accessible through natural-language or MCP-driven workflows. Missing for 10: independent/hands-on verification of bulk operation scale and performance, and no explicit example of bulk edits/updates to many tickets or records simultaneously (only bulk testing/auditing/review are documented).
- [claimed-docs] “Generate simulations straight from your actual tickets and run them in bulk batches, so every workflow change is tested against the conversa…”
- [claimed-docs] “Side-by-side batch comparisons show how a workflow edit changed outcomes across hundreds of scenarios, with per-conversation drill-downs whe…”
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Coach’s Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps, then deploy on the topics …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps”
- [claimed-docs] “Side-by-side batch comparisons show how a workflow edit changed outcomes across hundreds of scenarios, with per-conversation drill-downs”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to LorikeetSierra supports building workflows/journeys and guardrail-based release governance (e.g., agent checks, split traffic, merge approval) which implies some rule-based triggering, and Ghostwriter/Agent Studio let users define step-by-step logic, but there's no explicit documentation of an event-driven 'if X happens, trigger Y automatically' rules engine for AI-native users to configure independently. missing for 10: explicit event-trigger/rules-engine documentation, API/SDK examples of automated action-on-event configuration, independent verification of this specific automation capability.
- [claimed-docs] “Define step-by-step workflows for your agent from scratch, or use AI to instantly generate journeys from your existing operating procedures.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [claimed-docs] “Initiate workflows directly from a conversation—no tab-switching, no referencing instructions, no lost context.”
Lorikeet's workflows and guardrails encode conditional, event-triggered actions (e.g., prompt-injection detection triggers block/rewrite/escalate, bad QA scores trigger refunds, campaign cadences control scheduled outreach, escalation triggers hand off to humans), and workflows can be built/edited via natural language including deterministic steps. However this is more built-in platform logic than a general-purpose rule-definition interface for arbitrary custom events an AI-native user could freely wire up. Missing for 10: a documented general rules/automation engine or API letting users define arbitrary trigger-condition-action rules beyond the platform's fixed guardrail/workflow/outbound features, and independent confirmation of this working in practice.
- [claimed-docs] “Customer isolation, server-side identity validation, workflow-scoped tool access and hard execution caps are enforced in code rather than pr…”
- [claimed-docs] “Incoming messages pass a prompt-injection classifier and bad-actor checks, while response guardrails screen what the agent says for groundin…”
- [claimed-docs] “When Coach gives a conversation a bad score, we refund the AI portion of that interaction.”
- [claimed-docs] “campaign cadences with scheduling windows control when outreach happens”
- [claimed-docs] “When a case requires a specialist or regulatory review, it escalates with full interaction history so your team picks up mid-conversation.”
- [claimed-docs] “Lorikeet coordinates a team of specialist agents, with pockets of determinism for regulated steps, to handle multi-party, multi-system workf…”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to LorikeetSierranone0/10Sierra's evidence covers agent building, workflows, channels, analytics, and release governance, but nothing describes scheduling recurring jobs/workflows (e.g., cron-like triggers or timed automation runs) for AI-native users. Missing for 10: any documentation of scheduled/recurring job execution, trigger-based automation, or timer-based workflow runs.
The only relevant evidence is outbound campaign cadences with 'scheduling windows' controlling when outreach happens, which implies some recurring/scheduled automation but is narrowly scoped to outbound messaging rather than general recurring jobs or workflow runs. Missing for 10: explicit cron-like or recurring workflow scheduling for MCP-driven tasks (simulations, audits, diagnostics), documentation of scheduling frequency/config options, and any independent confirmation of recurring job execution.
- [claimed-docs] “campaign cadences with scheduling windows control when outreach happens”
ai-native userVersion, review, and roll back my automations
weight 1 · round to SierraSierra's Agent SDK is described as 'journeys as code' with change tracking, and release-governance docs mention merge approval, agent checks/simulations, and split-traffic gradual rollouts — pointing to versioned, reviewable release workflows. However, no explicit rollback/revert mechanism is documented, and docs are login-gated so no independent confirmation exists. missing for 10: explicit rollback capability, independent/hands-on confirmation of version history and revert function.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeet's docs show workflow building/iteration via natural language and strong review tooling (simulations, batch comparisons showing how an edit changed outcomes across scenarios), which covers the 'review' part of the story. However, there is no explicit mention of a version history or a rollback/undo mechanism for automations. missing for 10: explicit versioning/change-history feature, explicit rollback/undo capability, evidence of restoring a prior workflow state.
- [claimed-docs] “Build workflows - create and iterate on workflows using natural language”
- [claimed-docs] “Build, edit, and deploy workflows using natural language”
- [claimed-docs] “Side-by-side batch comparisons show how a workflow edit changed outcomes across hundreds of scenarios, with per-conversation drill-downs whe…”
- [claimed-docs] “Side-by-side batch comparisons show how a workflow edit changed outcomes across hundreds of scenarios, with per-conversation drill-downs”
- [claimed-docs] “Generate simulations straight from your actual tickets and run them in bulk batches, so every workflow change is tested against the conversa…”
Channels languages — stories about channels languages in this arenaChannels languages
Stories about channels languages in this arena
Channels
support leaderOne agent covers chat, email, and in-app, plus the channels my customers actually use — Slack, WhatsApp, social
weight 2 · round to LorikeetSierra's docs confirm one agent can be built once and deployed across chat, phone, email, SMS, and messaging channels (sierra-docs-3), with channel-specific tuning like Voice Personas (sierra-docs-16) and even ChatGPT publishing (sierra-docs-17), showing broad omnichannel intent. However, none of the evidence explicitly names Slack, WhatsApp, or social media as supported channels, leaving the specific channels support leaders care about unconfirmed. Missing for 10: explicit documentation or independent confirmation that Slack, WhatsApp, and social platforms are supported channels, plus hands-on evidence of a single agent operating consistently across these specific channels.
- [claimed-docs] “Build once and deploy across any channel—chat, phone, email, SMS, and messaging.”
- [claimed-docs] “Use Voice Personas to design how your agent sounds, speaks, and connects, powered by a constellation of models tuned across 59 languages.”
- [claimed-docs] “Publish to ChatGPT with one click: Make your agent available on ChatGPT with the push of a button, or via CI/CD.”
Lorikeet documents one agent resolving issues across chat, email, voice, SMS and WhatsApp (docs-31, docs-49), and mentions Slack as a place to interact with Coach (docs-10), but this is Coach access, not evidence that the customer-facing agent itself operates in Slack or social channels. missing for 10: explicit confirmation that the customer-facing support agent (not just Coach) handles Slack and social media channels, and independent/hands-on validation of omnichannel deployment.
- [claimed-docs] “one agent that resolves issues end-to-end across chat, email, voice and SMS”
- [claimed-docs] “resolve customer problems end-to-end across their lifecycle - via phone, SMS, chat, email, and WhatsApp”
- [claimed-docs] “Talk to Coach wherever you work, whether in Lorikeet, Slack, Claude, ChatGPT, or via MCP.”
- [claimed-docs] “When the AI concierge can't resolve a conversation, a human agent steps in directly inside the same platform”
Languages
support leaderThe agent supports customers in many languages, even where my knowledge base exists only in English
weight 2 · round to SierraSierra documents multilingual voice capability explicitly ('Voice Personas... tuned across 59 languages') and general multi-channel deployment, implying broad language coverage for customer-facing conversations. However, there is no direct evidence about handling a knowledge base that exists only in English—no mention of automatic translation/grounding of English-only content into other languages, so the specific claim in the story (KB gap bridging) is unevidenced. missing for 10: explicit documentation on cross-language grounding from English-only knowledge base, independent verification of multilingual quality beyond voice.
- [claimed-docs] “Use Voice Personas to design how your agent sounds, speaks, and connects, powered by a constellation of models tuned across 59 languages.”
- [claimed-docs] “Build once and deploy across any channel—chat, phone, email, SMS, and messaging.”
Voice
support leaderThe agent handles phone calls — speech in, speech out — with the same knowledge and actions as chat
weight 2 · round to SierraSierra explicitly markets voice as a first-class channel with the same agent logic/knowledge as chat ('Build once and deploy across any channel—chat, phone, email, SMS'), plus voice-specific features like Voice Personas across 59 languages, replacing IVR with empathetic voice agents, and phone payments without IVR handoff. Missing for 10: independent/hands-on verification that voice calls actually share identical knowledge/actions with chat in production (only vendor docs, no community confirmation of phone-specific parity).
- [claimed-docs] “Build once and deploy across any channel—chat, phone, email, SMS, and messaging.”
- [claimed-docs] “Use Voice Personas to design how your agent sounds, speaks, and connects, powered by a constellation of models tuned across 59 languages.”
- [claimed-docs] “Collect card and ACH payments entirely over the phone with no IVR handoff.”
- [claimed-docs] “Replace rigid IVR menus with an empathetic voice agent that understands customer context”
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
Lorikeet's marketing claims 'one agent that resolves issues end-to-end across chat, email, voice and SMS' and lists phone/voice among supported channels, implying the same agent handles calls. However, there is no detail on speech-in/speech-out mechanics, telephony integration, or evidence that voice interactions carry the same knowledge/actions/guardrails as chat beyond a generic channel list. missing for 10: specifics on speech recognition/TTS, call-handling architecture, latency/quality benchmarks, and confirmation that voice shares the same knowledge base and action set as chat.
- [claimed-docs] “one agent that resolves issues end-to-end across chat, email, voice and SMS”
- [claimed-docs] “resolve customer problems end-to-end across their lifecycle - via phone, SMS, chat, email, and WhatsApp”
Escalation handoff — stories about escalation handoff in this arenaEscalation handoff
Stories about escalation handoff in this arena
Handoff
support leaderWhen the agent escalates, the human gets the full conversation, a summary, and collected details — the customer never repeats themselves
weight 3 · round drawnSierra's Live Assist product explicitly targets escalation handoff, claiming reps are guided through next steps with automatically captured and updated customer context and can pick up conversations 'no tab-switching, no referencing instructions, no lost context' — directly addressing the no-repeat-yourself goal. However, there's no explicit mention of a generated conversation summary artifact or hands-on/independent confirmation that reps actually receive full transcript + summary + collected details in practice. Missing for 10: explicit summary-generation evidence, independent/hands-on validation of the handoff experience, confirmation reps see full conversation history alongside context.
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
- [claimed-docs] “Initiate workflows directly from a conversation—no tab-switching, no referencing instructions, no lost context.”
Lorikeet documents human handoff happening 'inside the same platform' when AI can't resolve, and specifically claims escalations carry 'full interaction history so your team picks up mid-conversation' (financial-services vertical) — directly supporting no-repeat handoff. However, evidence doesn't explicitly confirm a generated summary or structured 'collected details' package accompanying every escalation across all verticals, only interaction history. Missing for 10: explicit documentation of an auto-generated conversation summary at escalation, structured collected-details extraction, and independent/customer verification that customers never repeat themselves in practice.
- [claimed-docs] “When the AI concierge can't resolve a conversation, a human agent steps in directly inside the same platform”
- [claimed-docs] “When a case requires a specialist or regulatory review, it escalates with full interaction history so your team picks up mid-conversation.”
- [claimed-docs] “Enforce required disclosures and data handling standards through built-in guardrails that balance conversational flexibility with structured…”
Rules
support ops leadI configure when the agent must hand off — by topic, sentiment, customer tier, or explicit request — and it reliably obeys
weight 2 · round to LorikeetSierra's docs mention configurable guardrails, workflows, and a 'Live Assist' human-handoff product, implying some escalation mechanism exists, but no evidence describes explicit configuration of handoff triggers by topic, sentiment, customer tier, or explicit customer request, nor proof of reliable adherence to such rules. Missing for 10: documentation of specific trigger types (topic/sentiment/tier/request), configuration UI/API for these rules, and evidence (first-party or independent) that the agent reliably obeys them in production.
- [claimed-docs] “Build or modify agents by describing how you want them to behave. Update workflows, systems integrations, guardrails, tone, and style with s…”
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [claimed-docs] “Define step-by-step workflows for your agent from scratch, or use AI to instantly generate journeys from your existing operating procedures.”
Docs show human handoff when the AI can't resolve a ticket, escalation for regulatory/specialist cases with full context, guardrails that can 'escalate' on policy violations, and deployment scoped to trained topics — supporting topic-based and inability/explicit-need escalation. However there is no explicit mention of configuring handoff by sentiment or customer tier, nor independent verification that these rules are 'reliably obeyed'. Missing for 10: explicit sentiment-based trigger config, explicit customer-tier-based trigger config, and independent/hands-on evidence of reliability.
- [claimed-docs] “When the AI concierge can't resolve a conversation, a human agent steps in directly inside the same platform”
- [claimed-docs] “deploy on the topics the agent is trained for and leave the rest with your team”
- [claimed-docs] “Incoming messages pass a prompt-injection classifier and bad-actor checks, while response guardrails screen what the agent says for groundin…”
- [claimed-docs] “When a case requires a specialist or regulatory review, it escalates with full interaction history so your team picks up mid-conversation.”
- [claimed-docs] “Outbound runs on recorded consent that is enforced when messages are sent, with opt-out handling and escalation paths for sensitive situatio…”
Guardrails safety — stories about guardrails safety in this arenaGuardrails safety
Stories about guardrails safety in this arena
Hallucination
ai-native userGuardrails stop the agent from inventing policies, prices, or promises — off-knowledge questions get a safe decline, not a guess
weight 3 · round to LorikeetSierra's docs mention grounding agents in knowledge/FAQs/policies (sierra-docs-7) and general 'guardrails' as a configurable behavior via Ghostwriter (sierra-docs-8), plus release-governance guardrails like Agent Checks/Simulations (sierra-docs-18) and visibility into reasoning/knowledge sources (sierra-docs-13). However, no evidence explicitly describes a 'safe decline' mechanism for off-knowledge questions or confirms the agent won't invent prices/policies rather than guess. Missing for 10: explicit documentation or hands-on proof of decline-on-unknown behavior, third-party validation that hallucinated policies/prices are prevented, and detail on the unexplored 'trust-and-reliability' product page.
- [claimed-docs] “View, manage, and edit knowledge data such as Help Center content, FAQs, and policies that ground your agent.”
- [claimed-docs] “Build or modify agents by describing how you want them to behave. Update workflows, systems integrations, guardrails, tone, and style with s…”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
Lorikeet documents explicit guardrails that enforce grounding and policy compliance in code rather than prompts, including response screening for grounding/policy with block/rewrite/escalate actions, and adversarial simulation testing for false authority claims and prompt injection to validate safe-decline behavior before deployment. This directly targets stopping invented policies/prices/promises via server-side enforcement rather than relying on model honesty. Missing for 10: independent/third-party verification or a concrete hands-on example showing an off-knowledge question actually triggering a safe decline rather than a hallucinated answer.
- [claimed-docs] “Customer isolation, server-side identity validation, workflow-scoped tool access and hard execution caps are enforced in code rather than pr…”
- [claimed-docs] “Incoming messages pass a prompt-injection classifier and bad-actor checks, while response guardrails screen what the agent says for groundin…”
- [claimed-docs] “every event lands in your analytics as a tracked outcome your QA team can review”
- [claimed-docs] “Author adversarial scenarios such as false authority claims, mid-conversation goal switches and prompt injection attempts, then run them as …”
- [claimed-docs] “Enforce required disclosures and data handling standards through built-in guardrails that balance conversational flexibility with structured…”
Supervision
support ops leadLaunch in a supervised mode where the agent drafts replies for human approval before anything reaches a customer
weight 2 · round drawnSierranone0/10Sierra's evidence covers release governance (merge-approval for deploying agent changes) and Live Assist (guiding human reps in real time), but neither describes a mode where the agent drafts individual customer replies that a human must approve before they are sent. No docs or community evidence mention message-level human-in-the-loop approval for live customer conversations.
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
- [claimed-docs] “Initiate workflows directly from a conversation—no tab-switching, no referencing instructions, no lost context.”
Lorikeetnone0/10Lorikeet's documented model is autonomous resolution with post-hoc QA (Coach reviews 100% of tickets after the fact) and escalation to humans only when the AI can't resolve a case, not a pre-send draft-for-approval workflow. No evidence describes a mode where every agent reply is queued for human sign-off before reaching a customer. Missing for 10: any mention of a draft/approve workflow, human-in-the-loop send gating, or a configurable 'supervised' launch mode.
- [claimed-docs] “When the AI concierge can't resolve a conversation, a human agent steps in directly inside the same platform”
- [claimed-docs] “Incoming messages pass a prompt-injection classifier and bad-actor checks, while response guardrails screen what the agent says for groundin…”
- [claimed-docs] “Coach’s Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “When a case requires a specialist or regulatory review, it escalates with full interaction history so your team picks up mid-conversation.”
Topic controls
support ops leadI mark topics as human-only — legal threats, cancellations, security — and the agent never freelances on them
weight 2 · round to LorikeetSierranone0/10Sierra's guardrail evidence covers release governance (Agent Checks, Simulations, merge approval, split traffic) and general knowledge/workflow editing, but nothing describes topic-level human-only flags for categories like legal threats, cancellations, or security that the agent is barred from handling. Live-assist shows human+AI collaboration but not a mechanism to designate certain topics as strictly human-only with enforced escalation.
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
Lorikeet documents topic-scoped deployment ("deploy on the topics the agent is trained for and leave the rest with your team"), workflow-scoped tool access with hard execution caps enforced in code, and escalation paths for regulated/specialist cases with full history handoff — all consistent with restricting the agent from acting on sensitive topics. However, there's no explicit documentation of a UI/config for support-ops to designate specific topics (e.g., legal threats, cancellations, security) as strictly human-only with enforced non-freelancing. Missing for 10: explicit topic-tagging/human-only designation feature, confirmation that agent cannot even attempt those topics (vs. escalating mid-conversation), and independent verification of this behavior in production.
- [claimed-docs] “deploy on the topics the agent is trained for and leave the rest with your team”
- [claimed-docs] “Lorikeet coordinates a team of specialist agents, with pockets of determinism for regulated steps, to handle multi-party, multi-system workf…”
- [claimed-docs] “Customer isolation, server-side identity validation, workflow-scoped tool access and hard execution caps are enforced in code rather than pr…”
- [claimed-docs] “When a case requires a specialist or regulatory review, it escalates with full interaction history so your team picks up mid-conversation.”
- [claimed-docs] “Enforce required disclosures and data handling standards through built-in guardrails that balance conversational flexibility with structured…”
Insights analytics — stories about insights analytics in this arenaInsights analytics
Stories about insights analytics in this arena
Analytics
support leaderDashboards show resolution rate, CSAT, handoff rate, and cost per resolution — the numbers I report to my exec team
weight 3 · round drawnSierra's Insights and Explorer products provide dashboards/reports with drill-down and natural-language trend analysis (sierra-docs-10, sierra-docs-11, sierra-docs-12, sierra-docs-13), showing the platform surfaces analytics for support leaders, but no evidence specifies the exact metrics named in the story (resolution rate, CSAT, handoff rate, cost per resolution). Missing for 10: explicit documentation naming these specific KPIs, screenshots/examples of the actual dashboard metrics, and independent confirmation that leaders use it for exec-level reporting.
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
- [claimed-docs] “Explorer automatically delivers a weekly briefing on key trends, emerging issues, and recommendations — without you having to ask.”
- [claimed-docs] “Click any data point on a report to launch Explorer and instantly investigate what's driving that trend.”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
Lorikeet claims analytics tracking 'resolution quality, customer satisfaction, revenue impact, and operational efficiency' and that 'every event lands in your analytics as a tracked outcome,' which loosely maps to resolution rate, CSAT, and cost metrics, but there is no explicit mention of a handoff-rate metric or a dedicated exec-facing dashboard view combining these four numbers. missing for 10: explicit handoff-rate metric, dedicated dashboard UI/screenshot evidence, cost-per-resolution specifics, and independent corroboration of dashboard usability.
- [claimed-docs] “Track resolution quality, customer satisfaction, revenue impact, and operational efficiency with industry-leading analytics.”
- [claimed-docs] “Scale and optimize with conversational insights and analytics from Lorikeet Coach”
- [claimed-docs] “every event lands in your analytics as a tracked outcome your QA team can review”
Insights
support leaderThe platform clusters conversations by topic and surfaces emerging product issues before they spike ticket volume
weight 1 · round to SierraSierra's Explorer product explicitly surfaces trends and emerging issues via natural-language querying of conversations and proactive weekly briefings, and Insights lets users drill into data points—directly matching the 'surface emerging issues before they spike' theme. However, evidence doesn't explicitly confirm topic clustering methodology, predictive spike detection, or independent/hands-on validation of this specific capability. Missing for 10: explicit description of topic-clustering mechanics, evidence of predictive alerting before ticket-volume spikes (vs. reactive weekly summaries), and independent/customer corroboration of this specific insights capability.
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
- [claimed-docs] “Explorer automatically delivers a weekly briefing on key trends, emerging issues, and recommendations — without you having to ask.”
- [claimed-docs] “Click any data point on a report to launch Explorer and instantly investigate what's driving that trend.”
Lorikeetnone0/10Evidence shows general analytics and quality-review features (e.g., 'Track resolution quality, customer satisfaction... with industry-leading analytics', Coach reviewing 100% of conversations) but no mention of topic clustering of conversations or proactive surfacing of emerging product issues before ticket-volume spikes. missing for 10: topic clustering of conversations, trend/anomaly detection for emerging issues, ticket-volume spike prediction or alerting.
- [claimed-docs] “Track resolution quality, customer satisfaction, revenue impact, and operational efficiency with industry-leading analytics.”
- [claimed-docs] “Scale and optimize with conversational insights and analytics from Lorikeet Coach”
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
Integrations platform — stories about integrations platform in this arenaIntegrations platform
Stories about integrations platform in this arena
Helpdesk
developerThe agent runs inside my existing helpdesk — Zendesk, Salesforce, Intercom — or standalone, syncing tickets and context both ways
weight 3 · round to LorikeetSierranone0/10No evidence pack citation mentions Zendesk, Salesforce, Intercom, or bidirectional ticket/context syncing with existing helpdesk platforms; only generic 'systems integrations' and 'internal APIs' are referenced without naming any helpdesk system or describing two-way ticket sync.
Docs claim Lorikeet works 'alongside your existing tools' and connects 'in seconds to your ticketing system' via integrations, implying helpdesk co-existence, but there is no specific evidence of named Zendesk/Salesforce/Intercom connectors or two-way ticket/context sync — the pack focuses on MCP client integrations (Claude, ChatGPT) rather than helpdesk platforms. missing for 10: named Zendesk/Salesforce/Intercom integration docs, evidence of bidirectional ticket sync, standalone-mode confirmation, independent corroboration of integration reliability.
- [claimed-docs] “No migration required: Lorikeet works alongside your existing tools.”
- [claimed-docs] “Connect Lorikeet in seconds to your ticketing system, knowledge base, and internal tools to seamlessly ingest data and take action for your …”
- [claimed-docs] “When the AI concierge can't resolve a conversation, a human agent steps in directly inside the same platform”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
Knowledge grounding — stories about knowledge grounding in this arenaKnowledge grounding
Stories about knowledge grounding in this arena
Freshness
support ops leadKnowledge stays current automatically — the agent re-syncs sources on a schedule or on change, not via manual re-uploads
weight 2 · round drawnSierranone0/10Evidence shows Sierra lets teams manually 'view, manage, and edit knowledge data' (sierra-docs-7) but nothing describes automatic re-syncing of knowledge sources on a schedule or on-change detection; docs are login-gated so no further detail is visible.
- [claimed-docs] “View, manage, and edit knowledge data such as Help Center content, FAQs, and policies that ground your agent.”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeetnone0/10Evidence shows Lorikeet ingests knowledge bases and can audit them for gaps/outdated articles, and connects to ticketing/knowledge sources 'in seconds', but there is no mention of scheduled or change-triggered re-syncing of sources without manual re-upload. missing for 10: explicit scheduled/webhook-based re-sync mechanism, evidence of automatic detection of source updates, any documentation of sync cadence or on-change triggers.
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Connect Lorikeet in seconds to your ticketing system, knowledge base, and internal tools to seamlessly ingest data and take action for your …”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
- [claimed-docs] “train the agent on your business context, brand guidelines, help docs and standard operating procedures”
Gaps
support ops leadThe platform surfaces knowledge gaps and conflicting content that cause the agent to miss or fumble questions
weight 1 · round to LorikeetSierra's Explorer and Insights products surface emerging issues, trends, and the reasoning/knowledge sources behind agent answers, and Agent Studio lets teams view/manage knowledge content (FAQs, policies), which could help surface gaps — but no evidence explicitly describes detecting knowledge gaps or conflicting content as a distinct feature. Missing for 10: explicit conflicting-content/contradiction detection, explicit 'knowledge gap' flagging, and independent/hands-on confirmation that Explorer or Insights actually pinpoints such gaps rather than general conversation trends.
- [claimed-docs] “View, manage, and edit knowledge data such as Help Center content, FAQs, and policies that ground your agent.”
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
- [claimed-docs] “Explorer automatically delivers a weekly briefing on key trends, emerging issues, and recommendations — without you having to ask.”
- [claimed-docs] “Click any data point on a report to launch Explorer and instantly investigate what's driving that trend.”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
Lorikeet explicitly documents auditing knowledge bases at scale to find gaps and quality issues, and its simulation/replay tooling surfaces knowledge gaps by replaying historical tickets and synthetic scenarios to project resolution quality before deployment. Coach's Ticket Quality Score also reviews 100% of conversations to catch fumbled answers and feed fixes back. Missing for 10: independent/hands-on validation of the gap-detection accuracy and no explicit mention of detecting 'conflicting' content specifically (only gaps/outdated/quality issues).
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps, then deploy on the topics …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps”
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Coach’s Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Coach reviews 100% of tickets against your quality standards and turns findings into fixes, so your AI agent improves every week instead of …”
Grounding
ai-native userEvery answer is grounded in my own content and shows which article or source it drew from
weight 3 · round to SierraSierra explicitly supports grounding agent answers in customer-owned knowledge (Help Center content, FAQs, policies) via Agent Studio, and Insights lets teams 'understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed' which directly maps to showing which source an answer drew from. Missing for 10: independent/hands-on verification that end-user-facing answers visibly cite specific articles, and detail on citation UX rather than just admin-side reasoning traces.
- [claimed-docs] “View, manage, and edit knowledge data such as Help Center content, FAQs, and policies that ground your agent.”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
Lorikeetnone0/10Lorikeet documents training its agent on 'business context, brand guidelines, help docs and standard operating procedures' (docs-25) and mentions 'grounding' as a guardrail check (docs-35), but there is no evidence that end-user-facing answers cite or display the specific article/source used to generate a response. Missing for 10: any documented citation/source-attribution UI or API in agent responses, and independent confirmation that answers reference specific knowledge-base articles.
- [claimed-docs] “train the agent on your business context, brand guidelines, help docs and standard operating procedures”
- [claimed-docs] “Incoming messages pass a prompt-injection classifier and bad-actor checks, while response guardrails screen what the agent says for groundin…”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
Ingestion
support ops leadThe agent ingests my help center, docs, past tickets, and internal wikis as knowledge sources without manual re-authoring
weight 3 · round to LorikeetSierra's docs explicitly claim agent-studio lets teams 'view, manage, and edit knowledge data such as Help Center content, FAQs, and policies' and Ghostwriter can ingest SOPs, transcripts, and audio interviews to build journeys, supporting knowledge ingestion beyond just help center docs. However, there's no explicit evidence of ingesting past tickets or internal wikis specifically, and docs/reference material is login-gated so independent verification of breadth of source-type ingestion is limited. Missing for 10: documented support for tickets/wiki ingestion specifically, and independent/hands-on confirmation of the ingestion workflow beyond marketing copy.
- [claimed-docs] “View, manage, and edit knowledge data such as Help Center content, FAQs, and policies that ground your agent.”
- [claimed-docs] “Upload SOPs, raw transcripts, or even audio interviews of subject matter experts to create customer journeys from scratch.”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeet explicitly claims to 'seamlessly ingest data' from ticketing systems and knowledge bases (docs-13), to train the agent on 'business context, brand guidelines, help docs and standard operating procedures' (docs-25), to audit the entire knowledge base for gaps/outdated content (docs-3, docs-16), and to replay historical tickets to surface knowledge gaps (docs-21, docs-26) — covering help center, docs, and past tickets without manual re-authoring. Internal wikis are not explicitly named as a source type, and all evidence is first-party vendor documentation with no independent/hands-on corroboration of ingestion working end-to-end. Missing for 10: explicit mention of internal wiki ingestion, and independent verification of the ingestion pipeline's fidelity/accuracy.
- [claimed-docs] “Connect Lorikeet in seconds to your ticketing system, knowledge base, and internal tools to seamlessly ingest data and take action for your …”
- [claimed-docs] “train the agent on your business context, brand guidelines, help docs and standard operating procedures”
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Audit your entire knowledge base for gaps, outdated articles, and quality issues”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps, then deploy on the topics …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to LorikeetSierranone0/10Sierra markets an Agent SDK for code-based building alongside a separate no-code Agent Studio, but there is no evidence of API/UI feature parity — no public OpenAPI/swagger spec was found (404s), and the SDK reference docs themselves are login-gated to contracted customers, meaning even documented API scope can't be verified as matching UI capabilities.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Empower teams to build and manage agents—no code required”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeet's MCP server exposes a broad set of actions also available in the UI—diagnosing tickets, auditing knowledge bases, building/testing workflows, running simulations—letting AI-native users act via Claude/ChatGPT/Codex instead of the dashboard. However, there's no evidence of a general-purpose REST/GraphQL API or explicit claim of full UI/API parity (e.g., outbound campaign management, guardrail configuration, analytics dashboards aren't shown as MCP-accessible). Missing for 10: explicit full API parity statement, broader API coverage beyond MCP-listed actions, independent confirmation of completeness.
- [claimed-docs] “The Lorikeet MCP server lets you interact with your Lorikeet account directly from Claude Code, Claude.ai, ChatGPT, and Codex.”
- [claimed-docs] “Diagnose tickets - trace workflow execution and identify root causes”
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Build workflows - create and iterate on workflows using natural language”
- [claimed-docs] “Test tools - run and validate tool configurations directly from your AI assistant”
- [claimed-docs] “Explore your setup - inspect workflows, tools, and integrations”
- [claimed-docs] “Run simulations - test workflows against different customer scenarios”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
- [probe] “official MCP server documented at https://docs.lorikeetcx.ai/mcp/mcp-server”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnSierranone0/10No evidence of any data export or portability feature for customer journeys, knowledge data, or conversation logs in open formats; probes even show docs/API surfaces are gated or unavailable (404s, login-gated /llms.txt). Missing for 10: any export functionality, data portability documentation, or open-format data dump capability.
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
Lorikeetnone0/10No evidence pack items mention data export, open data formats, or account/data portability for leaving the platform; the docs cover MCP integrations, workflows, simulations, and guardrails but nothing about exporting or migrating data out. Missing for 10: any documentation of export functionality, supported open formats (CSV/JSON), or data portability/exit process.
ai-native userRead the product's source under an open license
weight 2 · round drawnSierranone0/10Sierra is presented purely as a closed commercial SaaS platform (Agent SDK, Studio, Ghostwriter, Voice, etc.) with no evidence of an open-source license for the product itself; the only open artifact found is the unrelated tau-bench research benchmark repo, not Sierra's product source code, and docs/API surfaces are login-gated or 404. Missing for 10: any open-license repository for Sierra's actual product code, license file, or public source release.
- [probe] “PROBE runtime (recorded 2026-09-10): Sierra's agent benchmark is genuinely open — the official sierra-research/tau-bench repository README f…”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
Pricing economics — stories about pricing economics in this arenaPricing economics
Stories about pricing economics in this arena
Pricing
support leaderPricing is outcome-based and published — I pay per resolution with caps and controls, not an opaque enterprise quote
weight 2 · round to LorikeetSierranone0/10No evidence pack item mentions pricing, resolution-based billing, caps, or published rates; docs and probes cover product features and access gating only. Sierra is widely known anecdotally for outcome-based pricing but nothing in this evidence pack substantiates published, capped, self-serve pricing terms.
The only evidence touching pricing-economics is the refund policy tying billing to Coach's quality score ('When Coach gives a conversation a bad score, we refund the AI portion of that interaction'), which shows some outcome-linkage in billing, but there is no published price list, per-resolution rate, cap structure, or self-serve pricing page. Missing for 10: a published price sheet or rate card, explicit per-resolution cost, spend/volume caps, and any evidence pricing is transparent versus custom enterprise quoting.
- [claimed-docs] “When Coach gives a conversation a bad score, we refund the AI portion of that interaction.”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnSierranone0/10No evidence pack item addresses data residency, region selection, or storage location controls; the docs cover agent building, workflows, and channels but not data governance/residency options. missing for 10: any mention of data region selection, residency guarantees, or storage location controls.
Lorikeetnone0/10The evidence covers compliance/trust items (Vanta certifications, zero-data-retention with model vendors) but no mention of data residency or region-selection options for storage. Missing for 10: any documentation of regional data storage choices, residency guarantees, or data localization controls.
ai-native userPrevent my data from being used to train AI models
weight 3 · round to LorikeetSierranone0/10No evidence in the pack addresses data-training opt-out or AI model training policies; Sierra's docs focus on product features (agent building, analytics, channels) with no privacy/data-use policy statements provided.
Lorikeet explicitly states zero-data-retention agreements with all model vendors and no fine-tuning on customer data, directly addressing the story's request to prevent data from being used for AI training. This is backed by independently verified trust/compliance reports on their Vanta Trust Center. Missing for 10: independent hands-on verification or third-party audit confirmation of this specific claim, and no detail on user-level opt-out controls or granularity of enforcement.
- [claimed-docs] “Zero-data-retention agreements with all model vendors and no fine-tuning on customer data.”
- [claimed-docs] “All four are independently verified, published on our public Vanta Trust Center with reports downloadable under NDA, and refreshed annually.”
ai-native userControl data retention and deletion
weight 2 · round to LorikeetSierranone0/10No evidence pack item addresses data retention controls, deletion policies, or user-facing data lifecycle management; docs cover agent building, analytics, and channels but never mention retention/deletion settings, and docs.sierra.ai is login-gated so no public verification exists.
Lorikeet mentions zero-data-retention agreements with model vendors and no fine-tuning on customer data, plus SOC2-style independently verified reports on a trust center, which touches data retention posture at the vendor-model level. However, there is no evidence of user-facing controls letting an AI-native user configure or request deletion/retention of their own conversation or account data within Lorikeet itself. Missing for 10: explicit customer-data deletion/export controls, retention period configuration, and user-initiated deletion workflows.
- [claimed-docs] “All four are independently verified, published on our public Vanta Trust Center with reports downloadable under NDA, and refreshed annually.”
- [claimed-docs] “Zero-data-retention agreements with all model vendors and no fine-tuning on customer data.”
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnSierranone0/10No evidence pack item discusses telemetry, usage tracking, opt-out controls, or privacy settings for AI-native/developer users; docs focus on product features and docs are gated behind login. Missing for 10: any documentation of telemetry collection practices, opt-out mechanism, or privacy controls.
Lorikeetnone0/10Lorikeet is a customer-support AI platform, and telemetry opt-out for the product itself is a fair privacy-posture question, but none of the evidence mentions any telemetry/usage-tracking opt-out mechanism for users of the product; it only discusses data retention with model vendors and consent handling for outbound customer messaging, which is unrelated to product telemetry opt-out.
- [claimed-docs] “All four are independently verified, published on our public Vanta Trust Center with reports downloadable under NDA, and refreshed annually.”
- [claimed-docs] “Zero-data-retention agreements with all model vendors and no fine-tuning on customer data.”
Resolution quality — stories about resolution quality in this arenaResolution quality
Stories about resolution quality in this arena
Personalization
support leaderAnswers use the customer's live data — plan, order status, account history — not just generic help articles
weight 2 · round to LorikeetSierra's docs and community evidence show the agent integrates with customers' internal APIs and systems (not just static knowledge), with Insights explicitly noting 'systems accessed' during agent actions, and Live Assist capturing live customer context. However, no evidence explicitly confirms real-time retrieval of specifics like plan, order status, or account history. missing for 10: explicit documented example of pulling live order/plan/account data, independent verification beyond one HN anecdote.
- [claimed-docs] “Understand and rapidly adjust agent behavior by inspecting API calls, logic traces, and more.”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
- [claimed-docs] “Guide care reps through next steps—whether in chat or on calls—automatically capturing and updating customer context as work progresses.”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
Lorikeet's docs describe connecting to ticketing systems, knowledge bases, and internal tools/APIs to 'ingest data and take action for your customers' (lorikeet-docs-13, lorikeet-docs-28, lorikeet-docs-29), with concrete examples like testing a 'get-order-status' tool with a real order ID (lorikeet-docs-19) and financial-services use cases like disputes/loan inquiries requiring account-specific data (lorikeet-docs-43, lorikeet-docs-44). This shows the agent is designed to pull and act on live customer data rather than just static help content. missing for 10: independent/hands-on verification that responses actually reflect real-time account state in production, and more detail on latency/freshness guarantees for live data lookups.
- [claimed-docs] “Connect Lorikeet in seconds to your ticketing system, knowledge base, and internal tools to seamlessly ingest data and take action for your …”
- [claimed-docs] “Test the 'get-order-status' tool with order ID 98765 and check if the response matches what we expect.”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
- [claimed-docs] “Build your customer Concierge to take actions inside the tools your team already uses”
- [claimed-docs] “Automate disputes, loan inquiries, payments and compliance cases across every customer channel”
- [claimed-docs] “When a case requires a specialist or regulatory review, it escalates with full interaction history so your team picks up mid-conversation.”
Reasoning
support leaderThe agent asks clarifying questions and works through multi-step troubleshooting instead of dumping one canned answer
weight 2 · round drawnSierra's docs describe building step-by-step, multi-step 'customer journeys' and workflows (sierra-docs-6, sierra-docs-9) rather than single canned replies, and community commentary confirms agents are configured to work through processes with internal APIs (sierra-comm-2). However, no evidence explicitly demonstrates the agent proactively asking clarifying questions or a hands-on troubleshooting transcript showing multi-turn dialogue in practice. Missing for 10: concrete transcript/demo of clarifying-question behavior, independent evaluation of troubleshooting depth, and explicit mention of clarification-seeking as a designed capability.
- [claimed-docs] “Define step-by-step workflows for your agent from scratch, or use AI to instantly generate journeys from your existing operating procedures.”
- [claimed-docs] “Upload SOPs, raw transcripts, or even audio interviews of subject matter experts to create customer journeys from scratch.”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
Lorikeet's docs describe workflows that coordinate multi-step, multi-system troubleshooting (e.g. 'coordinates a team of specialist agents... to handle multi-party, multi-system workflows end-to-end' and 'diagnose tickets by tracing workflow execution') rather than single canned replies, and its agent resolves issues 'end-to-end' across channels. However, there is no explicit documentation showing the agent proactively asking clarifying questions mid-conversation or examples of dialogue-based troubleshooting turns. Missing for 10: explicit examples/transcripts of clarifying-question behavior, documentation describing conversational back-and-forth troubleshooting logic rather than just workflow/tool orchestration.
- [claimed-docs] “Lorikeet coordinates a team of specialist agents, with pockets of determinism for regulated steps, to handle multi-party, multi-system workf…”
- [claimed-docs] “one agent that resolves issues end-to-end across chat, email, voice and SMS”
- [claimed-docs] “Diagnose tickets by tracing workflow execution and identifying root causes”
- [claimed-docs] “Build, edit, and deploy workflows using natural language”
- [claimed-docs] “The agent takes action through your APIs and MCP servers inside natural-language and deterministic workflows”
Resolution
support leaderThe agent fully resolves a meaningful share of conversations end-to-end — measured as resolutions, not mere deflections or bounces
weight 3 · round drawnSierra's marketing emphasizes agent-driven end-to-end workflow execution (journeys, integrations, payments, guardrails, testing) implying resolution rather than mere deflection, and community evidence corroborates real production use handling customer processes end-to-end for at least one customer. However, no quantitative resolution-rate metrics, benchmarks, or third-party validation of 'meaningful share resolved end-to-end' are provided — tau-bench is a research benchmark, not a customer outcome metric. Missing for 10: published resolution-rate statistics, case studies with concrete resolution percentages, independent audits distinguishing resolution from deflection.
- [claimed-docs] “Define step-by-step workflows for your agent from scratch, or use AI to instantly generate journeys from your existing operating procedures.”
- [claimed-docs] “Upload SOPs, raw transcripts, or even audio interviews of subject matter experts to create customer journeys from scratch.”
- [claimed-docs] “Collect card and ACH payments entirely over the phone with no IVR handoff.”
- [claimed-docs] “Replace rigid IVR menus with an empathetic voice agent that understands customer context”
- [community] “Having seen this set up for a customer, what is interesting is how it works with internal APIs for customers. It's all point and click in a …”
- [community] “The magic isn't in a new LLM technology, it is in reliably productionizing a solution for real-world problems... fill in the gap of missing …”
- [probe] “PROBE runtime (recorded 2026-09-10): Sierra's agent benchmark is genuinely open — the official sierra-research/tau-bench repository README f…”
Lorikeet's docs repeatedly claim end-to-end resolution (chat, email, voice, SMS), with escalation to humans when it can't resolve, plus QA scoring (Ticket Quality Score) and analytics tracking 'resolution quality' as an outcome metric, and even a refund-on-bad-score mechanism tied to quality. However, all evidence is vendor-authored marketing/docs; there are no independent benchmarks, customer case studies, or hard resolution-rate numbers (e.g., % of conversations fully resolved) to substantiate the claims. Missing for 10: independent/third-party resolution-rate data, customer-reported metrics, and clear definition/measurement methodology distinguishing true resolution from deflection.
- [claimed-docs] “one agent that resolves issues end-to-end across chat, email, voice and SMS”
- [claimed-docs] “resolve customer problems end-to-end across their lifecycle - via phone, SMS, chat, email, and WhatsApp”
- [claimed-docs] “When the AI concierge can't resolve a conversation, a human agent steps in directly inside the same platform”
- [claimed-docs] “Track resolution quality, customer satisfaction, revenue impact, and operational efficiency with industry-leading analytics.”
- [claimed-docs] “Coach’s Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “When Coach gives a conversation a bad score, we refund the AI portion of that interaction.”
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
Voice
support leaderI control the agent's tone and brand voice, and it stays consistent across topics and languages
weight 1 · round to SierraSierra provides explicit tone/brand-voice controls (Ghostwriter prompts for tone and style, Voice Personas tuned across 59 languages) plus testing/simulation tools to verify agent behavior across scenarios and avoid regressions, which supports consistency claims. However, all evidence is vendor-authored with no independent or hands-on verification that voice/tone actually stays consistent across topics and languages in production. Missing for 10: independent case studies or benchmarks confirming cross-topic/cross-language tone consistency, and detail on how brand-voice guardrails are enforced at scale.
- [claimed-docs] “Build or modify agents by describing how you want them to behave. Update workflows, systems integrations, guardrails, tone, and style with s…”
- [claimed-docs] “Use Voice Personas to design how your agent sounds, speaks, and connects, powered by a constellation of models tuned across 59 languages.”
- [claimed-docs] “Verify your agent performs as expected across a wide range of scenarios and avoid regressions.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
Lorikeet docs show the agent can be trained on 'business context, brand guidelines, help docs and standard operating procedures' (lorikeet-docs-25), and Coach's Ticket Quality Score reviews 100% of conversations against quality standards to catch drift (lorikeet-docs-9, lorikeet-docs-38), supporting brand-voice control and consistency monitoring. However, there is no explicit evidence of multi-language tone consistency or dedicated brand-voice/style configuration tooling beyond general training inputs. Missing for 10: explicit multilingual consistency support, dedicated tone/voice configuration UI, and independent evidence of voice consistency across topics/languages.
- [claimed-docs] “train the agent on your business context, brand guidelines, help docs and standard operating procedures”
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Coach’s Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “deploy on the topics the agent is trained for and leave the rest with your team”
Testing qa — stories about testing qa in this arenaTesting qa
Stories about testing qa in this arena
Qa
support ops leadAI conversations get ongoing QA — scored samples, flagged failures, and a review loop that feeds fixes back into the agent
weight 1 · round to LorikeetSierra documents Agent Checks and Simulations that proactively catch problems, merge-approval workflows for human review, and split-traffic releases (sierra-docs-18), plus regression testing (sierra-docs-4) and Explorer/Insights tools that surface conversation trends, flag emerging issues, and explain agent reasoning (sierra-docs-10, -11, -13). Ghostwriter lets teams feed fixes back by updating workflows/guardrails via prompts (sierra-docs-8). Missing for 10: explicit description of scored/sampled QA reviews with quantitative scoring rubrics, and independent evidence of the review loop actually closing the gap between flagged failures and shipped fixes.
- [claimed-docs] “Verify your agent performs as expected across a wide range of scenarios and avoid regressions.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [claimed-docs] “Ask any question about your customer experience in natural language, and Explorer identifies the answer across thousands of real conversatio…”
- [claimed-docs] “Explorer automatically delivers a weekly briefing on key trends, emerging issues, and recommendations — without you having to ask.”
- [claimed-docs] “Understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed, and more.”
- [claimed-docs] “Build or modify agents by describing how you want them to behave. Update workflows, systems integrations, guardrails, tone, and style with s…”
Lorikeet's Coach product directly addresses this story: it reviews 100% of conversations against quality standards (not just samples), scores them via a Ticket Quality Score, flags failures (bad scores trigger refunds), and turns findings into fixes that feed back into the agent so it 'improves every week instead of drifting.' Simulations complement this with batch testing and adversarial scenario scoring tied to guardrails and analytics tracking. Missing for 10: independent/hands-on verification of the review loop in practice and more detail on how flagged failures are triaged/assigned to human reviewers.
- [claimed-docs] “Coach's Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “Coach reviews 100% of tickets against your quality standards and turns findings into fixes, so your AI agent improves every week instead of …”
- [claimed-docs] “Coach’s Ticket Quality Score reviews 100% of conversations against your quality standards, AI and human alike, replacing manual spot checks …”
- [claimed-docs] “When Coach gives a conversation a bad score, we refund the AI portion of that interaction.”
- [claimed-docs] “Coach can implement improvements on your behalf, or make suggestions for you to action yourself.”
- [claimed-docs] “every event lands in your analytics as a tracked outcome your QA team can review”
- [claimed-docs] “Generate simulations straight from your actual tickets and run them in bulk batches, so every workflow change is tested against the conversa…”
- [claimed-docs] “Author adversarial scenarios such as false authority claims, mid-conversation goal switches and prompt injection attempts, then run them as …”
Simulation
support ops leadI test the agent against historical tickets or simulated conversations before it faces real customers
weight 2 · round to LorikeetSierra explicitly documents 'Agent Checks and Simulations' and 'verify your agent performs as expected across a wide range of scenarios and avoid regressions' (sierra-docs-4, sierra-docs-18), directly matching pre-release testing against scenarios/regressions. This is further corroborated by Sierra's public tau-bench research benchmark for evaluating conversational agents on simulated user interactions (sierra-probe-rt-1), showing real investment in simulation-based testing methodology. Missing for 10: no explicit mention of testing against historical/real ticket transcripts specifically (only 'scenarios' and simulations), and no independent hands-on customer account of the simulation workflow in practice.
- [claimed-docs] “Verify your agent performs as expected across a wide range of scenarios and avoid regressions.”
- [claimed-docs] “Agent Checks and Simulations catching problems proactively, merge approval workflows putting a person in the loop, and split traffic release…”
- [probe] “PROBE runtime (recorded 2026-09-10): Sierra's agent benchmark is genuinely open — the official sierra-research/tau-bench repository README f…”
Lorikeet's Simulations product directly supports this story: it generates simulations from actual historical tickets and runs them in bulk batches before workflow changes go live, with side-by-side batch comparisons and per-conversation drill-downs, plus authored adversarial/guardrail scenarios to pre-test against tricky real-world behavior. Docs also describe replaying historical tickets and synthetic scenarios in bulk to project resolution quality before deploying on trained topics. Missing for 10: independent/hands-on corroboration beyond first-party docs.
- [claimed-docs] “Generate simulations straight from your actual tickets and run them in bulk batches, so every workflow change is tested against the conversa…”
- [claimed-docs] “Side-by-side batch comparisons show how a workflow edit changed outcomes across hundreds of scenarios, with per-conversation drill-downs whe…”
- [claimed-docs] “Author adversarial scenarios such as false authority claims, mid-conversation goal switches and prompt injection attempts, then run them as …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps, then deploy on the topics …”
- [claimed-docs] “Replay historical tickets and synthetic scenarios in bulk to see projected resolution quality and knowledge gaps”
- [claimed-docs] “Run simulations - test workflows against different customer scenarios”
Not comparable on these axes
ai-native userRun the product headlessly / in CI for automation
weight 2 · not comparableSierra's Agent SDK is described as code-based with dev workflows retained, and one blog post explicitly notes publishing to ChatGPT can be done 'via CI/CD', implying some automation/pipeline support. However there is no dedicated CLI, headless runtime docs, or CI-specific guidance, and docs/API references are login-gated (probe shows /llms.txt is a login SPA and no public OpenAPI spec), so full headless/CI operation is unconfirmed. Missing for 10: explicit CLI/headless execution docs, public API/OpenAPI spec, independent CI usage reports.
- [claimed-docs] “write customer journeys as code, track changes, and build sophisticated logic without giving up our development workflows”
- [claimed-docs] “Publish to ChatGPT with one click: Make your agent available on ChatGPT with the push of a button, or via CI/CD.”
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableSierran/aSierra is itself an AI agent platform (the agent role), not a service being connected to by external agents via MCP; this axis is about serving as an MCP server for other agents, which is a category mismatch for a product that is the agent itself. No evidence shows Sierra exposing an official MCP server, and the story's axis is more appropriate for Sierra being a client integrating others' tools than serving as one.
Lorikeet publishes an official MCP server (docs.lorikeetcx.ai/mcp/mcp-server) that lets Claude, ChatGPT, Claude Code, Codex, and MintMCP connect directly to a Lorikeet account, with documented capabilities like diagnosing tickets, auditing knowledge bases, building workflows, testing tools, and running simulations via MCP. This is confirmed live (HTTP 200) via probe evidence, not just marketing copy. Missing for 10: independent/hands-on third-party verification of the MCP server working end-to-end, and no community reports corroborating reliability.
- [claimed-docs] “The Lorikeet MCP server lets you interact with your Lorikeet account directly from Claude Code, Claude.ai, ChatGPT, and Codex.”
- [claimed-docs] “Diagnose tickets - trace workflow execution and identify root causes”
- [claimed-docs] “Audit knowledge bases - analyze articles at scale, find gaps, and spot quality issues”
- [claimed-docs] “Build workflows - create and iterate on workflows using natural language”
- [claimed-docs] “Test tools - run and validate tool configurations directly from your AI assistant”
- [claimed-docs] “Connect to your Lorikeet account from ChatGPT, Claude, Claude Code, and MintMCP using the Model Context Protocol.”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.lorikeetcx.ai/llms.txt # Reference - [Lorikeet MCP Server](https://docs.lorikeetcx.ai/mcp/mcp-serv…”
- [probe] “official MCP server documented at https://docs.lorikeetcx.ai/mcp/mcp-server”
ai-native userExplore an interactive API reference with runnable examples
weight 2 · not comparableSierranone0/10Probes show no public OpenAPI/interactive reference (404s at openapi.json paths, llms.txt returns login SPA shell), and docs are login-gated to contracted customers rather than an open interactive API reference with runnable examples.
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
ai-native userSelf-host the core product
weight 3 · not comparableSierranone0/10Sierra is offered as a hosted SaaS platform with no evidence of any self-hosted or on-premise deployment option; docs are login-gated to contracted customers rather than exposing an installable/self-hostable core product. missing for 10: any documentation of self-hosting, on-prem deployment, or open-source release of the core agent platform.
- [probe] “PROBE runtime finding (recorded 2026-09-10): docs.sierra.ai has no agent-legible surface — /llms.txt resolves (HTTP 200) to the login SPA's …”
- [probe] “PROBE llms.txt: HTTP 404 at https://sierra.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sierra.ai/openapi.json, https://sierra.ai/swagger.json, https://sierra.ai/api/openapi.json, …”