[
  {
    "productId": "decagon",
    "storyId": "agentic-agent-docs",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "The probe confirms Decagon serves a working llms.txt file at decagon.ai/llms.txt (HTTP 200) with structured product description, showing the product is discoverable by AI agents pointed at agent-oriented docs. However, there's no evidence of broader agent-oriented documentation (e.g., structured API docs, OpenAPI spec which 404'd) or first-party guidance encouraging users to point agents at this file. missing for 10: OpenAPI/API-level machine-readable docs (404s confirmed), first-party documentation explicitly promoting llms.txt usage for AI agents, and independent corroboration of an agent successfully consuming the file.",
    "evidenceIds": [
      "decagon-probe-1",
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-ai-insights",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Decagon's Insights & Reporting and Suggestions features let users ask open-ended natural-language questions (e.g. 'Why are customers requesting refunds?') and get AI-driven analysis, heatmaps, and auto-generated content drafts based on detected knowledge gaps, directly surfacing AI-generated insights from customer data. Duet further auto-generates Agent Operating Procedures from past interactions and iterates based on conversation patterns. Missing for 10: independent/hands-on validation of insight quality and no detail on underlying analytics accuracy or limitations.",
    "evidenceIds": [
      "decagon-docs-9",
      "decagon-docs-10",
      "decagon-docs-17",
      "decagon-docs-18",
      "decagon-docs-30",
      "decagon-docs-6",
      "decagon-docs-24"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-autonomous-automation",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Decagon's Proactive Agents can carry context across sessions and autonomously initiate calls/recommendations based on customer signals, and Watchtower continuously monitors every interaction against custom criteria without manual triggering — both indicate background/autonomous operation. However, there's no explicit documentation of a general-purpose automation/scheduling framework (e.g., triggers, cron-like workflows, or arbitrary background tasks) beyond these two specific features. Missing for 10: a dedicated automation/scheduler product surface, independent/hands-on verification of autonomous behavior, and broader configurability beyond proactive outreach and monitoring.",
    "evidenceIds": [
      "decagon-docs-22",
      "decagon-docs-23",
      "decagon-docs-28",
      "decagon-docs-31"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-builtin-assistant",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Decagon's Duet acts as an in-product AI helper that lets (non-technical) users delegate specific tasks — auto-generating Agent Operating Procedures from past conversations, drafting tests, and producing content suggestions — which is a form of delegating work to a built-in assistant, but it's narrowly scoped to agent-configuration/QA tasks rather than a general-purpose assistant for arbitrary user tasks. missing for 10: evidence of a general-purpose in-product assistant handling open-ended user requests beyond AOP/test/content generation, independent/hands-on validation of Duet's delegation quality, and detail on how broadly tasks can be delegated versus templated workflows.",
    "evidenceIds": [
      "decagon-docs-5",
      "decagon-docs-6",
      "decagon-docs-11",
      "decagon-docs-10"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-headless",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Decagon is a SaaS conversational-AI platform for customer support with no-code configuration (AOPs, Duet, integrations 'no custom code required'); there is no documented CLI, headless mode, or CI-automation interface, and the openapi probe returned 404s, indicating no public API spec for automated/headless invocation. missing for 10: any CLI/SDK for headless execution, CI-pipeline integration docs, or public API reference enabling automation.",
    "evidenceIds": [
      "decagon-docs-19",
      "decagon-docs-26",
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-mcp-client",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Decagon has a blog post explicitly about MCP (\"getting the most out of MCP\") describing an infrastructure layer to curate, scope, and evaluate tool usage, indicating some MCP integration capability, but there is no concrete documentation of how to actually plug an MCP server into an agent, no config steps, and no independent corroboration of it working. missing for 10: technical setup docs for adding an MCP server, list of supported MCP servers/tools, hands-on or independent verification that agents actually invoke MCP tools.",
    "evidenceIds": [
      "decagon-docs-3"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-mcp-server",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Decagon's MCP blog post (decagon-docs-3) discusses using MCP to curate/scope tool access for its own agents (i.e., Decagon as an MCP client consuming external tools), not exposing an official MCP server that lets an external AI agent connect into Decagon. No documentation, endpoint, or announcement of a first-party Decagon MCP server was found, and the OpenAPI/spec probes returned 404s with no MCP-specific server mentioned.",
    "evidenceIds": [
      "decagon-docs-3",
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-nl-commands",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Decagon lets operators define agent behavior, flagging criteria, and analytics queries via natural language (AOPs, Watchtower, Ask AI), which supports natural-language operation of the product for configuration/analysis purposes. However, this is primarily aimed at business/support-ops users configuring an agent, not a general 'AI-native user' issuing arbitrary natural-language commands to control the whole product end-to-end. missing for 10: evidence of natural-language command interface for broader product operations (e.g., admin tasks, integrations setup, deployment) beyond AOPs/Watchtower/Insights, and independent/hands-on verification of this capability.",
    "evidenceIds": [
      "decagon-docs-8",
      "decagon-docs-16",
      "decagon-docs-20",
      "decagon-docs-30",
      "decagon-docs-9"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-official-cli",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Decagon is a customer-support AI agent platform (chat, voice, email, AOPs), not a developer tool or coding-agent product where an official CLI for AI-native workflows would be a relevant axis; no evidence pack content even gestures at a CLI.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "agentic-public-api",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence of a documented public API; the openapi probe returned 404s across all candidate paths and no docs reference an API reference, SDK, or programmatic endpoint. Integrations are described as 'no custom code required' (docs-19, docs-26), suggesting no-code/UI-driven configuration rather than a documented API for AI-native control.",
    "evidenceIds": [
      "decagon-probe-2",
      "decagon-docs-19",
      "decagon-docs-26"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-scoped-keys",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon documents short-lived, scoped JWT tokens for agent access to customer systems and identity-provider (Okta/Entra) integration, which shows some least-privilege credential mechanics, but this is about the agent's own runtime access rather than a user-facing capability to explicitly issue/manage scoped API credentials for an agent. There's no evidence of an API/console feature letting an AI-native user provision, scope, or revoke discrete credentials themselves. missing for 10: user-facing credential issuance/management UI or API, granular scoping controls exposed to users, independent verification of the JWT scoping claims.",
    "evidenceIds": [
      "decagon-docs-2",
      "decagon-docs-21"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-sdks",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence of official SDKs for developers; the pack shows only no-code integrations, MCP blog commentary, and a failed openapi probe (404s), with no SDK documentation or API libraries surfaced.",
    "evidenceIds": [
      "decagon-docs-3",
      "decagon-docs-19",
      "decagon-docs-26",
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "agentic-webhooks",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence in the pack mentions webhooks, event subscriptions, or any push-notification mechanism; the OpenAPI probe returned 404s and no API reference documenting webhook endpoints was found. This axis applies (Decagon integrates with external systems and could plausibly offer webhooks) but there is no supporting evidence.",
    "evidenceIds": [
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "api-interactive-docs",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence of an interactive API reference or runnable examples; the openapi probe returned 404 across all candidate paths and no docs mention API documentation with runnable examples. Missing for 10: any API reference page, SDK/runnable code samples, interactive docs like Swagger/Redoc, or developer sandbox.",
    "evidenceIds": [
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "api-machine-spec",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Probe explicitly checked common OpenAPI/swagger endpoints and all returned 404, and no documentation references a downloadable machine-readable API spec; only an llms.txt file was found, which is not an API spec.",
    "evidenceIds": [
      "decagon-probe-2",
      "decagon-probe-1"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "api-sandbox",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Decagon's Testing & QA product describes 'Simulations' that validate agent behavior 'before deploying to production and with every subsequent update' (decagon-docs-29) and Duet-generated test suites (decagon-docs-11), implying a pre-production testing environment. However, there is no explicit documentation of a dedicated sandbox with isolated/non-production data, and the Experiments feature explicitly runs live in production (decagon-docs-7), which cuts against a clear sandbox-vs-production separation. Missing for 10: explicit description of sandbox data isolation, confirmation that test/simulation environments don't touch live customer data, and independent verification of this claim.",
    "evidenceIds": [
      "decagon-docs-29",
      "decagon-docs-11",
      "decagon-docs-7"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "api-versioning-policy",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence of a public/versioned API, API changelog, or deprecation policy; the openapi probe returned 404s across all candidate paths and no docs mention API versioning or deprecation practices.",
    "evidenceIds": [
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "automation-bulk-operations",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence describes a bulk-operations feature (e.g., batch editing, bulk tagging, bulk export/import of tickets or conversations) for AI-native users; the closest items describe aggregate analysis (Watchtower reviewing every interaction, Insights analyzing many conversations) rather than user-driven bulk actions across items.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "automation-rules-engine",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Decagon offers Watchtower, which monitors conversations against natural-language criteria and can flag/surface events (compliance risks, sentiment, etc.), and Proactive Agents that act on customer signals (e.g., initiating calls, recommendations) — both function as event-triggered automation. However, there is no explicit documentation of a general-purpose rule-builder (if-event-then-action) framework that an AI-native user could configure directly; the evidence describes narrower, product-specific triggering mechanisms rather than a flexible automation-rules engine. Missing for 10: explicit rule/condition builder UI or API, documentation of arbitrary event types and action bindings, and independent verification of how these triggers are configured.",
    "evidenceIds": [
      "decagon-docs-8",
      "decagon-docs-20",
      "decagon-docs-31",
      "decagon-docs-22",
      "decagon-docs-23",
      "decagon-docs-16"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "automation-scheduled-jobs",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Decagon's evidence covers customer-support agent features (chat, voice, email, analytics, testing, proactive outreach) but nothing addresses scheduling recurring jobs or workflows in the AI-native/automation-depth sense; proactive agents mention initiating calls at 'right moments' but no scheduling/cron-like mechanism is documented. missing for 10: any scheduler, cron/recurring trigger config, or workflow automation timing controls.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "automation-versioned-workflows",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon offers some adjacent controls—technical teams retain 'visibility and control over guardrails, integrations, and versioning' and Simulations/testing validate agent behavior 'before deploying to production and with every subsequent update'—suggesting some change-management workflow exists, but there's no explicit documentation of version history browsing, diffing, review/approval workflows, or a rollback mechanism to a prior automation state. missing for 10: explicit versioning UI/history, review/approval workflow for changes, documented rollback mechanism, independent confirmation of these capabilities in use.",
    "evidenceIds": [
      "decagon-docs-1",
      "decagon-docs-25",
      "decagon-docs-29"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "brand-voice-tone",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Decagon's docs show brand-voice consistency built into multiple surfaces: chat, voice (multilingual), and email are all described as 'on-brand', and AOPs let support leaders define agent behavior/tone in natural language like SOPs (docs-13,14,15,16). Testing/QA and Experiments explicitly validate tone/brand across pathways and let leaders 'refine tone' safely before and after deployment (docs-11,29,32). Missing for 10: independent/hands-on evidence that tone actually stays consistent across many topics and languages in production, and no explicit mention of cross-language consistency for chat/email (only voice is called multilingual).",
    "evidenceIds": [
      "decagon-docs-13",
      "decagon-docs-14",
      "decagon-docs-15",
      "decagon-docs-16",
      "decagon-docs-11",
      "decagon-docs-29",
      "decagon-docs-32"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "configurable-escalation-rules",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon's AOPs let teams define agent behavior in natural language (like SOPs), which could include escalation logic, and Watchtower can flag sentiment/compliance signals, but no evidence explicitly describes configuring handoff triggers by topic, sentiment, customer tier, or explicit request, nor confirms reliable adherence to such rules. Simulations/testing-QA claims validating 'policies' are generic and don't specifically address handoff reliability. Missing for 10: explicit documentation of handoff/escalation configuration options, tier-based routing, and evidence of reliable handoff obedience (e.g., test results or case studies on escalation accuracy).",
    "evidenceIds": [
      "decagon-docs-16",
      "decagon-docs-8",
      "decagon-docs-20",
      "decagon-docs-31",
      "decagon-docs-25",
      "decagon-docs-29"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "content-gap-detection",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Decagon's Suggestions product explicitly automates detection of knowledge gaps and drafts content based on how top human agents resolved similar issues, directly addressing the story's core need, and Insights/Watchtower let ops teams query conversations to surface where the agent struggles or conflicts arise. Missing for 10: independent/hands-on verification of gap detection accuracy, and no explicit mention of surfacing 'conflicting content' across knowledge sources rather than just gaps.",
    "evidenceIds": [
      "decagon-docs-10",
      "decagon-docs-18",
      "decagon-docs-9",
      "decagon-docs-8",
      "decagon-docs-17"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "context-rich-handoff",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence pack contains no mention of escalation-to-human handoff, conversation summaries handed to agents, or collected-details transfer preventing repetition. Related items about cross-session context (decagon-docs-22, decagon-docs-28) address agent-to-customer continuity, not human-agent handoff, so the specific story is unevidenced despite being a plausible capability for a support AI platform.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "end-to-end-resolution",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon's docs emphasize automation infrastructure (AOPs, Duet, Browser Actions, integrations) and analytics (Watchtower, Insights) but provide no quantified resolution-rate metrics, no third-party benchmark, and the one customer quote (Duolingo) speaks to maintenance effort, not resolution outcomes. Missing for 10: independent or customer-reported resolution-rate figures, a clear definition/measurement of 'resolution' vs deflection, and case studies quantifying end-to-end conversation completion rather than agent capability lists.",
    "evidenceIds": [
      "decagon-docs-4",
      "decagon-docs-27",
      "decagon-docs-12",
      "decagon-docs-8",
      "decagon-docs-31"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "grounded-cited-answers",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Decagon's evidence describes knowledge-gap detection, AOP generation, and content suggestions (docs-10, docs-18) but nothing shows that end-user-facing answers actually cite or display the specific article/source they were grounded in. Missing for 10: any documentation of inline citations, source attribution UI, or 'view source' feature in chat/voice/email responses.",
    "evidenceIds": [
      "decagon-docs-10",
      "decagon-docs-18",
      "decagon-docs-13"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "hallucination-guardrails",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon provides AOPs to define agent behavior/policies in natural language and a testing suite (Simulations/Duet) that verifies agents 'respond accurately, follow policies' plus Watchtower monitoring for compliance violations, which are the building blocks for guardrails — but no evidence explicitly describes a safe-decline behavior for off-knowledge questions versus fabricating answers. missing for 10: explicit documentation or examples of the agent refusing/declining out-of-scope questions rather than hallucinating, and independent verification that guardrails actually prevent invented prices/policies in practice.",
    "evidenceIds": [
      "decagon-docs-16",
      "decagon-docs-11",
      "decagon-docs-29",
      "decagon-docs-20",
      "decagon-docs-31"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "helpdesk-platform-integrations",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon's integrations page claims no-code integrations and system-of-record data portability, and Browser Actions/MCP posts describe connecting to arbitrary systems including ones without native integrations, implying support for helpdesk platforms like Zendesk/Salesforce/Intercom, but no evidence explicitly names these tools or describes ticket-level two-way sync. Missing for 10: named connectors/case studies for Zendesk, Salesforce, or Intercom, explicit description of bidirectional ticket-context syncing, and confirmation of standalone deployment mode.",
    "evidenceIds": [
      "decagon-docs-19",
      "decagon-docs-26",
      "decagon-docs-21",
      "decagon-docs-28",
      "decagon-docs-4",
      "decagon-docs-27",
      "decagon-docs-3"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "knowledge-auto-sync",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence pack covers knowledge gap detection and content suggestions (docs-10/18), integrations without custom code, and agent iteration via Duet, but nowhere mentions automatic re-syncing of knowledge sources on a schedule or on change — the closest analog (Suggestions) generates draft content for humans to review/publish, not an automated re-sync pipeline. Missing for 10: any mention of scheduled/triggered re-ingestion of source documents, sync frequency, or change-detection on connected knowledge bases.",
    "evidenceIds": [
      "decagon-docs-10",
      "decagon-docs-18"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "knowledge-source-ingestion",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Decagon's integrations page claims 'no custom code required' connections (docs-19/26) and its Suggestions product analyzes help-center content and past agent resolutions to fill knowledge gaps (docs-10/18), implying some ingestion of existing content, but there is no explicit documentation naming help center, internal wikis, docs, and past tickets as ingestible knowledge sources without re-authoring. missing for 10: explicit list/documentation of supported knowledge source connectors (help center, wikis, ticket systems), evidence of automatic re-indexing without manual content rewriting, and independent confirmation of successful multi-source ingestion.",
    "evidenceIds": [
      "decagon-docs-10",
      "decagon-docs-18",
      "decagon-docs-19",
      "decagon-docs-26"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "live-api-actions",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Decagon documents scoped, short-lived JWT access for agent actions across customer systems and integrations 'with no custom code required,' plus explicit mention of agents adjusting 'refund logic,' supporting real action-taking with scoped auth. However there's no detailed documentation of per-action granularity (e.g., separate refund vs subscription-update scopes), no audit trail examples, and no independent/hands-on verification of actual API action execution — missing for 10: per-action scope definitions, action audit logging, and third-party verification of real transactional actions.",
    "evidenceIds": [
      "decagon-docs-2",
      "decagon-docs-19",
      "decagon-docs-26",
      "decagon-docs-32",
      "decagon-docs-27"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "multi-turn-troubleshooting",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Decagon docs describe AOPs that define multi-step agent behavior like human SOPs and mention agents carrying context across sessions and taking multi-step actions (browser actions, integrations), which implies some troubleshooting flow, but there is no direct documentation or example showing the agent asking clarifying questions or walking through iterative multi-step troubleshooting dialogues. missing for 10: concrete example/transcript of clarifying-question behavior, documentation explicitly describing multi-turn troubleshooting logic, independent/hands-on evidence confirming this behavior in practice.",
    "evidenceIds": [
      "decagon-docs-16",
      "decagon-docs-22",
      "decagon-docs-27"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "multilingual-answers",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Only the Voice product page explicitly claims multilingual capability ('built for natural, multilingual conversations'); there is no evidence that Chat or Email channels support multiple languages, nor any documentation of translating or drawing from an English-only knowledge base to serve other languages. Missing for 10: explicit multilingual support claims for chat/email channels, description of how KB content is translated/localized, and independent verification of multilingual quality.",
    "evidenceIds": [
      "decagon-docs-14",
      "decagon-docs-13",
      "decagon-docs-15"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "omnichannel-coverage",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Decagon clearly offers a unified agent across chat (web, mobile, messaging platforms), email, and voice, suggesting broad channel reach [decagon-docs-13][decagon-docs-14][decagon-docs-15], with proactive/session-continuity features tying channels together [decagon-docs-22][decagon-docs-28]. However, the evidence never names specific channels like Slack, WhatsApp, or social media explicitly — only generic 'messaging platforms' is mentioned. Missing for 10: explicit documentation naming Slack, WhatsApp, and social media integrations as supported channels, plus any customer proof of omnichannel handoff across these specific channels.",
    "evidenceIds": [
      "decagon-docs-13",
      "decagon-docs-14",
      "decagon-docs-15",
      "decagon-docs-22",
      "decagon-docs-28"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "ongoing-qa-review",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Decagon documents a dedicated Testing & QA suite (Simulations) that validates agent behavior pre-deploy and on updates, plus Watchtower which reviews every live interaction against custom flagging criteria (compliance, sentiment, frustration) and surfaces failures, and Duet which auto-generates tests/AOPs and iterates on the agent based on conversation gaps — together covering scored/flagged QA and a feedback loop into agent fixes. missing for 10: no independent/hands-on corroboration of scoring accuracy or the closed-loop fix cycle, and no explicit description of 'sampling' methodology for QA review.",
    "evidenceIds": [
      "decagon-docs-8",
      "decagon-docs-11",
      "decagon-docs-20",
      "decagon-docs-29",
      "decagon-docs-24",
      "decagon-docs-31"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "openness-api-parity",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence of a public API at all — the openapi probe returned 404 across all candidate paths, and none of the docs describe an API that mirrors UI capabilities like AOP configuration, Duet, Watchtower, or experiments. Absence of evidence for this applicable capability yields none.",
    "evidenceIds": [
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "openness-full-export",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Decagon's evidence pack contains no documented data-export feature, open-format export tooling, or account-deletion/portability workflow; the only tangential mention (\"data portability and control\" in decagon-docs-28) is vague marketing language about maintaining conversation context, not a concrete export/exit mechanism, and the probe shows no public API/OpenAPI spec that could support programmatic data extraction.",
    "evidenceIds": [
      "decagon-docs-28",
      "decagon-probe-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "openness-open-license",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Decagon is a closed-source enterprise SaaS product; there is no indication it is or ever claims to be open-source software with source code available under a license. This axis applies to open-source projects, not proprietary commercial platforms like Decagon.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "openness-self-host",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Decagon is a fully-hosted SaaS customer support platform with no evidence of any self-hostable core product; self-hosting is a category mismatch for this SaaS offering rather than a missing feature.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "personalized-account-answers",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Decagon's docs describe real-time, scoped access to customer systems via short-lived JWTs, MCP-based tool integrations, and Browser Actions that let agents log into and pull data from any system (including those without native integrations), which supports live-data-driven answers rather than static help content. However, there is no concrete example or case study showing an actual answer surfacing a customer's plan, order status, or account history — the evidence is architectural/capability-level rather than demonstrated output. Missing for 10: a documented example or case study showing live plan/order/account data appearing in an actual customer-facing answer, and independent verification that this works in production.",
    "evidenceIds": [
      "decagon-docs-2",
      "decagon-docs-3",
      "decagon-docs-4",
      "decagon-docs-27",
      "decagon-docs-19",
      "decagon-docs-21"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "pre-launch-simulation",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Decagon's Simulations product is explicitly described as an integrated testing suite that validates agent behavior across channels before production deployment and on every update, and Duet can generate diverse test pathways to check accuracy, policy adherence, and brand voice. Missing for 10: no explicit mention of testing against historical ticket logs specifically, and no independent/hands-on evidence corroborating the testing suite's effectiveness beyond vendor docs.",
    "evidenceIds": [
      "decagon-docs-29",
      "decagon-docs-11",
      "decagon-docs-6"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "privacy-data-residency",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence pack item mentions data residency, region selection, or storage location controls; security page content covers JWTs and SSO but not data residency options. Missing for 10: any mention of regional data storage, residency compliance (e.g., EU/US hosting options), or customer-facing controls to choose storage location.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "privacy-no-training",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence in the pack addresses data usage for AI model training or an opt-out/no-training policy; the security page covers JWT tokens and SSO but not training data practices.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "privacy-retention-controls",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Decagon only vaguely references 'data portability and control' alongside customer context retention, and short-lived JWT tokens discarded after each session, but there is no explicit documentation of configurable data retention periods, deletion APIs, or user-facing controls to purge stored customer data. missing for 10: explicit retention policy/settings, a documented deletion mechanism or API, and independent verification that deletion requests are honored.",
    "evidenceIds": [
      "decagon-docs-28",
      "decagon-docs-2"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "privacy-telemetry-optout",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence in the pack addresses telemetry/usage-tracking opt-out controls; the security and product pages cover access control, JWT tokens, and SSO but never mention analytics/telemetry opt-out settings for end users. Missing for 10: any documentation of a telemetry toggle, data-collection opt-out mechanism, or privacy settings page.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "procedure-sop-builder",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Decagon explicitly ships 'Agent Operating Procedures (AOPs)' described as letting teams 'define agent behavior in natural language, the same way you train human agents with SOPs,' with Duet auto-generating and iterating these procedures from real interactions. This directly matches encoding step-by-step SOPs for known issue types. Missing for 10: concrete technical detail on how branching logic/decision trees are structured and enforced deterministically, and independent/hands-on corroboration beyond vendor docs.",
    "evidenceIds": [
      "decagon-docs-16",
      "decagon-docs-6",
      "decagon-docs-5",
      "decagon-docs-25",
      "decagon-docs-24"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "resolution-analytics",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Decagon's Insights & Reporting product includes customizable dashboards, visual heatmaps for key metrics, and 'Ask AI' conversational analytics for querying conversation trends, which supports general exec-reporting use cases. However, the evidence never explicitly names resolution rate, CSAT, handoff rate, or cost-per-resolution as tracked/reported metrics. missing for 10: explicit confirmation that CSAT, resolution rate, handoff rate, and cost-per-resolution are specific dashboard metrics; independent/customer corroboration of these exact KPIs being reported to execs.",
    "evidenceIds": [
      "decagon-docs-9",
      "decagon-docs-17",
      "decagon-docs-30"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "restricted-topic-controls",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Decagon's AOPs let teams define agent behavior/escalation rules in natural language, and 'guardrails' are described as retained under technical team control, which could support marking certain topics as human-only; Watchtower can flag legal/compliance/sentiment topics for review. However, there's no explicit documentation of a dedicated 'human-only topic' or hard-escalation feature, nor any evidence (hands-on or audit) confirming the agent never engages on flagged categories. Missing for 10: explicit human-only/topic-exclusion configuration feature, proof/testing that the agent reliably refuses or escalates on those topics, and independent verification of enforcement.",
    "evidenceIds": [
      "decagon-docs-1",
      "decagon-docs-16",
      "decagon-docs-25",
      "decagon-docs-8",
      "decagon-docs-20",
      "decagon-docs-31"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "supervised-draft-mode",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "Decagon's evidence covers testing/simulation (Simulations, Duet), experiments, guardrails, and monitoring (Watchtower), but nothing describes a supervised/human-in-the-loop 'draft-for-approval' mode where a human must approve each reply before it reaches a customer. Testing/QA and experiments happen pre-production or on traffic splits, not as a live approval queue for individual replies.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "topic-trend-insights",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Decagon's Insights & Reporting and Watchtower products support natural-language querying of conversations (e.g., 'Why are customers requesting refunds?'), heatmaps to spot spikes/dips in key metrics, and custom flagging criteria across every interaction, which together enable topic-level pattern detection and early issue surfacing. However, the evidence never explicitly describes automated topic clustering or a proactive 'before ticket volume spikes' alerting mechanism—these are inferred from adjacent features. Missing for 10: explicit description of automatic conversation clustering by topic, independent/hands-on evidence of early trend detection preventing ticket spikes, and case-study proof of this specific insight workflow in production.",
    "evidenceIds": [
      "decagon-docs-9",
      "decagon-docs-17",
      "decagon-docs-30",
      "decagon-docs-31",
      "decagon-docs-20",
      "decagon-docs-10"
    ]
  },
  {
    "productId": "decagon",
    "storyId": "transparent-per-resolution-pricing",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence pack item mentions pricing, resolution-based billing, caps, controls, or published pricing pages — all evidence concerns product features (AOPs, Duet, Watchtower, integrations, security) rather than pricing model. Absence of evidence for this applicable axis yields 'none'. Missing for 10: published pricing page, resolution-based billing structure, caps/controls documentation, any pricing transparency claims.",
    "evidenceIds": []
  },
  {
    "productId": "decagon",
    "storyId": "voice-phone-support",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Decagon explicitly offers a Voice product for 'natural, multilingual conversations' and can 'initiate intelligent, on-brand calls,' and its platform architecture (AOPs, integrations, guardrails) is shared across channels including chat, implying shared knowledge/actions. However, there's no explicit vendor statement confirming full parity of knowledge/actions between voice and chat, no technical detail on speech-to-speech quality, and no independent/hands-on corroboration of call handling. Missing for 10: explicit parity confirmation between voice and chat agent logic, technical/latency details of the voice pipeline, independent or customer testimonial evidence of voice call handling in production.",
    "evidenceIds": [
      "decagon-docs-14",
      "decagon-docs-23",
      "decagon-docs-13",
      "decagon-docs-16",
      "decagon-docs-25"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-agent-docs",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "A live probe confirms Fin publishes an llms.txt file at https://fin.ai/llms.txt (HTTP 200) explicitly framed to help LLMs understand its content structure, and Fin also offers an MCP server and a CLI agents can be pointed at for setup. Missing for 10: markdown-doc mirrors (docs-md probe 404) and a discoverable OpenAPI spec (all candidates 404), which would round out agent-friendly documentation.",
    "evidenceIds": [
      "intercom-fin-probe-1",
      "intercom-fin-docs-38",
      "intercom-fin-probe-4",
      "intercom-fin-probe-2",
      "intercom-fin-probe-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-ai-insights",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin's 'Insights', 'AI Recommendations', 'AI Topics', 'Trends', and CX Score features (docs-16, docs-47, docs-61) show it surfaces AI-generated analysis of conversation data, and 'Monitors'/'Custom AI Scorecards' extend this to proactive alerting. However this is framed around support-conversation analytics rather than a general 'insights from your data' experience, and there's no independent/hands-on validation of accuracy or usefulness of these AI-generated insights. missing for 10: independent corroboration of insight quality, broader data-source coverage beyond support conversations, concrete UI examples of AI-generated suggestions.",
    "evidenceIds": [
      "intercom-fin-docs-16",
      "intercom-fin-docs-47",
      "intercom-fin-docs-61",
      "intercom-fin-docs-66"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-autonomous-automation",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin's Procedures, Tasks, Workflows, and Proactive Outbound Suite are documented as letting Fin autonomously handle multi-step customer queries, business logic, and third-party system actions end-to-end without human intervention, and Fin resolves a majority of queries (76%) without agent involvement. However, evidence is entirely vendor marketing/help docs with no independent or hands-on confirmation of scheduling/background execution mechanics, and one community comment questions whether open-ended autonomous agent behavior is even desirable versus deterministic workflows. Missing for 10: independent verification of autonomous background execution, technical detail on triggers/scheduling for Procedures/Tasks, and hands-on confirmation that these run without ongoing human oversight.",
    "evidenceIds": [
      "intercom-fin-docs-23",
      "intercom-fin-docs-24",
      "intercom-fin-docs-26",
      "intercom-fin-docs-44",
      "intercom-fin-docs-15",
      "intercom-fin-docs-56",
      "intercom-fin-comm-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-builtin-assistant",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Fin is itself a built-in AI agent that customers and support teams delegate tasks to — it resolves ~76% of queries end-to-end, handles multi-step 'Procedures' with business logic, takes actions on external systems, and works across channels (chat, email, voice, Slack, WhatsApp). This is documented extensively in first-party docs/help articles covering delegation-style use (Fin Tasks, Fin Procedures, escalation rules, API/Agent API for programmatic delegation).\n\nmissing for 10: independent hands-on validation of delegation quality (one HN commenter says they've never seen Fin in the wild, and another questions whether autonomous agent reasoning is even the right approach vs deterministic workflows), and no live product screenshots/demo confirming smooth end-user delegation experience.",
    "evidenceIds": [
      "intercom-fin-docs-9",
      "intercom-fin-docs-11",
      "intercom-fin-docs-23",
      "intercom-fin-docs-26",
      "intercom-fin-docs-44",
      "intercom-fin-docs-56",
      "intercom-fin-docs-3",
      "intercom-fin-comm-1",
      "intercom-fin-comm-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-headless",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin exposes a programmatic Agent API and Node/TypeScript SDK (intercom-fin-docs-3, intercom-fin-gh-3) that allows calling Fin from external code/services, which supports headless/CI-style automation, and webhooks (intercom-fin-docs-6) enable event-driven automation without a UI. However, there is no documented CI-specific tooling (e.g., a CLI test runner, GitHub Action, or exit-code-based automation harness) and the orchestration endpoints needed for autonomous agent-style headless runs are only in a Preview API (intercom-fin-docs-4). missing for 10: dedicated CI/CD integration or GitHub Action, documented headless test/run mode with exit codes, independent confirmation of automation reliability outside preview APIs.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-docs-4",
      "intercom-fin-docs-6",
      "intercom-fin-gh-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-mcp-client",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin's own docs state it 'works through MCP or API Data Connectors for your business tools' (docs-40/54) and explicitly describe connecting the 'Zapier MCP connector to Fin' (intdir-3), confirming Fin can consume external MCP servers as tools. Missing for 10: independent/hands-on verification of MCP tool usage in production and a broader list of supported MCP servers beyond the Zapier example.",
    "evidenceIds": [
      "intercom-fin-docs-40",
      "intercom-fin-docs-54",
      "intercom-fin-intdir-3",
      "intercom-fin-docs-5",
      "intercom-fin-docs-32"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-mcp-server",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Intercom publishes official docs for an MCP server that lets AI agents securely access and interact with Intercom data, and this is corroborated by a dedicated probe confirming the documented endpoint plus Fin's own integrations page referencing MCP connectivity. missing for 10: independent/hands-on confirmation of the MCP server working in practice beyond vendor docs, and clearer detail on scope/auth setup.",
    "evidenceIds": [
      "intercom-fin-docs-5",
      "intercom-fin-docs-32",
      "intercom-fin-docs-40",
      "intercom-fin-docs-54",
      "intercom-fin-probe-4"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-nl-commands",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin's core function is responding to customer queries in natural language across channels, and admins can 'give Fin guidance' and train Procedures using natural-language instructions (docs-22, docs-44, docs-60); the Fin CLI also lets an AI coding agent configure Fin setup based on natural-language prompts (docs-38). However there's no clear evidence of a full natural-language command interface for operating/administering the product itself (e.g., configuring settings, running reports) beyond training/guidance content. Missing for 10: documented NL-driven admin console commands, evidence of broad NL task execution beyond content/guidance training, and independent verification of this capability in practice.",
    "evidenceIds": [
      "intercom-fin-docs-22",
      "intercom-fin-docs-44",
      "intercom-fin-docs-38",
      "intercom-fin-docs-60"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-official-cli",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows Fin has an API, Node SDK, and MCP server, but no official CLI tool is documented or referenced anywhere in the evidence pack.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-public-api",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin ships a documented Fin Agent API and an official TypeScript SDK (intercom-node) with concrete usage examples, plus a separate API platform for answer generation/knowledge retrieval, giving AI-native users a real programmatic path to drive Fin. However, key orchestration endpoints (Discover capabilities, Ask Fin, Run a procedure, Escalate) are only in Preview API version, and automated probes found no discoverable OpenAPI/swagger spec, suggesting the public API surface is not fully standardized/discoverable yet. Missing for 10: stable (non-preview) orchestration endpoints, a machine-readable OpenAPI spec, and independent hands-on developer corroboration beyond vendor docs/SDK repo.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-docs-4",
      "intercom-fin-docs-39",
      "intercom-fin-docs-53",
      "intercom-fin-gh-1",
      "intercom-fin-gh-2",
      "intercom-fin-gh-3",
      "intercom-fin-probe-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-scoped-keys",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "Evidence covers API access, MCP, webhooks, and integrations, but there is no mention of scoped or least-privilege API credential/token issuance for agents (e.g., granular permission scopes, per-agent API keys). Absence of evidence for an applicable capability yields 'none'. missing for 10: scoped/least-privilege credential issuance, API key/token permission granularity, agent-specific credential management docs.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-docs-5",
      "intercom-fin-docs-32"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-sdks",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Intercom publishes an official TypeScript/Node SDK (intercom-node) with convenient API access and iterator support, plus a documented Fin Agent API and API platform for programmatic use. However, only one official SDK language is evidenced, there's no public OpenAPI spec (probe returned 404s across all candidate paths), and no independent corroboration of SDK quality/adoption exists. Missing for 10: multi-language SDK coverage, publicly discoverable OpenAPI/schema, and third-party validation of SDK reliability.",
    "evidenceIds": [
      "intercom-fin-gh-1",
      "intercom-fin-gh-2",
      "intercom-fin-gh-3",
      "intercom-fin-docs-3",
      "intercom-fin-docs-39",
      "intercom-fin-probe-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "agentic-webhooks",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Intercom's developer docs explicitly document webhooks for subscribing to real-time event notifications (Contact created, Conversation received, Outbound Message receipt), directly matching the story. Missing for 10: independent/hands-on corroboration of webhook reliability and broader event coverage beyond the three examples cited.",
    "evidenceIds": [
      "intercom-fin-docs-6",
      "intercom-fin-docs-33"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "api-interactive-docs",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "While Fin has API docs (developers.intercom.com) and an SDK on GitHub, there is no evidence of an interactive API reference with runnable/try-it examples; a probe explicitly found no OpenAPI/Swagger spec published at any candidate path, and no docs mention a live API console.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-probe-3",
      "intercom-fin-gh-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "api-machine-spec",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Fin has documented REST APIs (Fin Agent API, Node/TS SDK) but no evidence of a downloadable machine-readable spec; a direct probe for OpenAPI/swagger files at common paths returned 404 for all candidates, and no docs page links such a spec.",
    "evidenceIds": [
      "intercom-fin-probe-3",
      "intercom-fin-docs-3",
      "intercom-fin-docs-4"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "api-sandbox",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Fin documents a testing/preview workflow ('How to preview and test Fin', testing across thousands of scenarios before going live) that implies an isolated test/preview mode, but no evidence explicitly describes a sandbox environment separate from production data or confirms production data isolation during testing. Missing for 10: explicit sandbox/staging environment documentation, confirmation that test scenarios don't touch live customer data, and any independent/hands-on corroboration of this isolation.",
    "evidenceIds": [
      "intercom-fin-docs-29",
      "intercom-fin-docs-45",
      "intercom-fin-docs-59"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "api-versioning-policy",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence shows Fin has a REST API with a 'Preview API version' for new endpoints (docs-4) and a general changes/changelog page (docs-1), but there is no documented versioning scheme or explicit deprecation policy (e.g., version sunset timelines, backward-compatibility guarantees) cited anywhere in the pack.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-docs-4",
      "intercom-fin-docs-1"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "automation-bulk-operations",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Fin exposes a programmatic API and TypeScript SDK with a pagination iterator for looping over items, plus a content-import endpoint, implying some capacity for scripted bulk actions, but there is no documented bulk-edit, bulk-resolve, or batch-action feature for AI-native users. missing for 10: explicit bulk operation/batch endpoint documentation, hands-on evidence of processing many items in one call, and any UI/CLI bulk-action support.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-gh-1",
      "intercom-fin-gh-2"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "automation-rules-engine",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin exposes several automation primitives that fit an event→action rule model: webhooks that fire on real-time Intercom events, configurable escalation guidance/rules, Fin Procedures/workflows, and Operator/Monitors for incident-triggered behavior. This is more than none, but the pack lacks a first-class 'if event X then action Y' rule-builder walkthrough or independent confirmation of how flexible/robust these triggers are. Missing for 10: a documented dedicated rules/automation builder UI, concrete example of a user-defined trigger-condition-action rule, and independent (non-vendor) evidence that automation rules work reliably in practice.",
    "evidenceIds": [
      "intercom-fin-docs-6",
      "intercom-fin-docs-25",
      "intercom-fin-docs-23",
      "intercom-fin-docs-24",
      "intercom-fin-docs-16"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "automation-scheduled-jobs",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Fin's evidence covers webhooks (event-driven, not scheduled), procedures, and a Proactive Outbound Suite, but nothing describes cron-like scheduling of recurring jobs or workflows on a time-based schedule. missing for 10: any documentation of scheduled/recurring workflow triggers, cron-style job scheduling, or recurring automation configuration.",
    "evidenceIds": [
      "intercom-fin-docs-6",
      "intercom-fin-docs-15",
      "intercom-fin-docs-44"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "automation-versioned-workflows",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence pack shows Fin Procedures/workflow management and testing before go-live ('test it across thousands of scenarios before anything goes live, roll out changes with control'), but there is no concrete documentation of version history, diffing/review of automation changes, or an explicit rollback mechanism for Procedures/Workflows/Guidance. Nothing describes a changelog, version comparison, or a 'revert to previous version' feature for automations.",
    "evidenceIds": [
      "intercom-fin-docs-23",
      "intercom-fin-docs-24",
      "intercom-fin-docs-45"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "brand-voice-tone",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs support tone/brand control (\"Customizable tone and answer length\", \"Train Fin on your knowledge, data, policies, and tone of voice\") and multi-language support is documented (\"Use Fin AI Agent in multiple languages\"), plus multi-channel consistency claims (email, Slack, live chat, voice) suggest tone carries across surfaces. However, there is no independent evidence or hands-on validation that tone/brand voice actually stays consistent across topics and languages in practice — only vendor marketing claims. Missing for 10: independent/third-party validation of tone consistency, concrete examples of cross-language tone fidelity, and evidence of consistency across many topics rather than just a training feature description.",
    "evidenceIds": [
      "intercom-fin-docs-10",
      "intercom-fin-docs-46",
      "intercom-fin-docs-60",
      "intercom-fin-docs-19",
      "intercom-fin-docs-65",
      "intercom-fin-docs-48"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "configurable-escalation-rules",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs explicitly cover configuring escalation/handoff behavior — 'Manage Fin AI Agent's escalation guidance and rules', 'Giving Fin Guidance', 'Hand-off with JavaScript', an orchestration 'Escalate to a human' API endpoint, and 'Transfers to agents directly in preferred Inbox' — showing topic/guidance-based and explicit-request handoff configuration exists. However, none of the evidence specifically documents sentiment- or customer-tier-based handoff triggers, nor is there independent/hands-on verification that the agent reliably obeys these rules in practice. Missing for 10: explicit sentiment/tier-based escalation config docs, and third-party evidence confirming reliable adherence to configured handoff rules.",
    "evidenceIds": [
      "intercom-fin-docs-25",
      "intercom-fin-docs-22",
      "intercom-fin-docs-31",
      "intercom-fin-docs-4",
      "intercom-fin-docs-12"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "content-gap-detection",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Fin's analytics stack (Insights, Monitors, AI Topics/Trends, AI Recommendations, Incident Detection) implies some ability to flag where Fin underperforms or topics trend poorly, and 'Giving Fin Guidance'/'Adding content to Fin' suggest content curation workflows, but no evidence explicitly describes detecting knowledge gaps or conflicting source content as a distinct surfaced capability. Missing for 10: explicit documentation of a 'knowledge gap' or 'content conflict' detection feature, and any hands-on/independent confirmation that such gaps are surfaced to support-ops leads.",
    "evidenceIds": [
      "intercom-fin-docs-16",
      "intercom-fin-docs-47",
      "intercom-fin-docs-61",
      "intercom-fin-docs-22",
      "intercom-fin-docs-28"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "context-rich-handoff",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs confirm Fin can transfer/escalate to human agents within the preferred inbox, has configurable escalation guidance/rules, and supports hand-off via JavaScript and an orchestration 'Escalate to a human' API — showing escalation handoff is a real, built feature. However, none of the evidence explicitly confirms that a generated summary or a structured collection of gathered details is passed along with the full conversation to the human agent, so the 'never repeats themselves' guarantee is not directly documented. Missing for 10: explicit documentation of automatic conversation summary generation at escalation, evidence that collected customer details are packaged and handed off, and independent/hands-on confirmation the handoff actually eliminates repetition.",
    "evidenceIds": [
      "intercom-fin-docs-4",
      "intercom-fin-docs-12",
      "intercom-fin-docs-25",
      "intercom-fin-docs-31"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "end-to-end-resolution",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin explicitly claims a resolution metric (\"Fin resolves 76% of customer queries on average, and handles multi-step queries end to end\") across multiple integration pages, plus dedicated testing/evaluation tooling (\"evaluate every live conversation\") and analytics (Insights, Monitors) framed around resolution outcomes rather than deflection. Community commentary raises philosophical skepticism about autonomous agents vs. deterministic workflows, but does not concretely contradict the stated resolution rate. Missing for 10: independent/third-party audit of the 76% figure, and clear definition distinguishing 'resolution' from deflection/bounce in the metric methodology.",
    "evidenceIds": [
      "intercom-fin-docs-56",
      "intercom-fin-docs-57",
      "intercom-fin-docs-58",
      "intercom-fin-docs-59",
      "intercom-fin-docs-45",
      "intercom-fin-docs-61",
      "intercom-fin-comm-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "grounded-cited-answers",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Fin's docs confirm it is trained/grounded on the customer's own content (help center, websites, docs, knowledge sources) via content import and knowledge retrieval APIs (e.g., docs-46/60 'train Fin on your knowledge', docs-28 'Adding content to Fin', docs-27 'Sync and manage websites', gh-1 content import API, docs-39 knowledge retrieval API), but no evidence explicitly shows that end-user answers surface or cite the specific source article they were drawn from. missing for 10: explicit documentation or UI evidence of in-answer source citation/attribution, independent/hands-on confirmation of citation behavior.",
    "evidenceIds": [
      "intercom-fin-docs-46",
      "intercom-fin-docs-60",
      "intercom-fin-docs-28",
      "intercom-fin-docs-27",
      "intercom-fin-gh-1",
      "intercom-fin-docs-39",
      "intercom-fin-docs-49"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "hallucination-guardrails",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs show Fin is 'trained on your knowledge, data, policies, and tone' and grounded in the knowledge base (docs-46/60/49), with configurable 'escalation guidance and rules' (docs-25) and pre-launch testing across scenarios (docs-45/59), which together imply guardrails against off-knowledge answers and hand-off rather than free invention. However there is no explicit documentation of a 'safe decline' behavior or hallucination-refusal mechanism, and no independent/hands-on evidence confirming Fin actually declines rather than guesses on off-knowledge questions. Missing for 10: explicit safe-decline/refusal documentation, third-party or hands-on verification that Fin avoids inventing policies/prices, and quantified guardrail testing results.",
    "evidenceIds": [
      "intercom-fin-docs-46",
      "intercom-fin-docs-60",
      "intercom-fin-docs-25",
      "intercom-fin-docs-45",
      "intercom-fin-docs-59",
      "intercom-fin-docs-49",
      "intercom-fin-docs-52",
      "intercom-fin-docs-67"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "helpdesk-platform-integrations",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Fin explicitly documents standalone deployment plus deep integrations with Zendesk, Salesforce, HubSpot, Freshdesk and 'any helpdesk', with two-way data sync (bring content/history over, surface data in inbox) and APIs/MCP/webhooks for programmatic ticket/context sync. Multiple dedicated integration pages (Zendesk, Salesforce, HubSpot, Freshdesk) and platform docs (Agent API, MCP, webhooks, data connectors) corroborate bidirectional context flow. missing for 10: independent/hands-on verification of the bidirectional sync working in production, and no direct evidence of Intercom-native standalone ticketing depth compared to third-party helpdesks.",
    "evidenceIds": [
      "intercom-fin-docs-7",
      "intercom-fin-docs-20",
      "intercom-fin-docs-21",
      "intercom-fin-docs-36",
      "intercom-fin-docs-37",
      "intercom-fin-docs-3",
      "intercom-fin-docs-5",
      "intercom-fin-docs-6",
      "intercom-fin-docs-40",
      "intercom-fin-intdir-8"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "knowledge-auto-sync",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Docs mention 'Sync and manage websites' and content import via API (createContentImportSource), implying re-syncing of website sources, but there is no explicit documentation of scheduled or change-triggered automatic re-syncing across all source types (docs, help centers, files) without manual re-upload. missing for 10: explicit scheduling/cron documentation for content refresh, confirmation of automatic change-detection re-sync for non-website sources (PDFs, articles), and independent/hands-on verification that sync happens without manual re-upload.",
    "evidenceIds": [
      "intercom-fin-docs-27",
      "intercom-fin-gh-1"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "knowledge-source-ingestion",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Fin explicitly supports ingesting websites/help center content (createContentImportSource, 'Sync and manage websites', 'Adding content to Fin'), integrates with helpdesks (Zendesk, Salesforce, HubSpot, Freshdesk) preserving existing content/history, and pulls in Confluence, Notion, Zendesk content into a unified knowledge source alongside data connectors for third-party tools. This covers help center, docs, past ticket/helpdesk content, and wiki tools (Confluence/Notion) without manual re-authoring. Missing for 10: no explicit mention of ingesting 'internal wikis' broadly beyond Confluence/Notion, and no independent/hands-on verification of ingestion fidelity or effort required.",
    "evidenceIds": [
      "intercom-fin-gh-1",
      "intercom-fin-docs-27",
      "intercom-fin-docs-28",
      "intercom-fin-intdir-6",
      "intercom-fin-docs-36",
      "intercom-fin-docs-20",
      "intercom-fin-docs-21",
      "intercom-fin-docs-43"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "live-api-actions",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin's Procedures feature explicitly supports multi-step business logic with third-party systems (refunds, order changes, etc. are common examples in docs like 'Takes action to update external systems' and Procedures training), and data connectors (Stripe, Shopify) plus MCP/API access enable real API calls. However, evidence does not detail per-action scoped authentication/authorization — no documentation on granular auth scoping per action, permission boundaries, or audit trails for individual API calls. missing for 10: explicit scoped-auth-per-action documentation, permission/authorization model for individual actions, and independent verification that real refunds/subscription updates are executed successfully in production.",
    "evidenceIds": [
      "intercom-fin-docs-11",
      "intercom-fin-docs-23",
      "intercom-fin-docs-44",
      "intercom-fin-docs-34",
      "intercom-fin-docs-39",
      "intercom-fin-docs-52"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "multi-turn-troubleshooting",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin's docs describe 'Procedures' that train it to handle multi-step queries with business logic across systems, and marketing claims it 'handles multi-step queries end to end,' supporting structured troubleshooting rather than single canned replies. However, no evidence explicitly documents Fin asking clarifying questions mid-conversation, and one HN commenter argues open-ended agentic reasoning is unnecessary versus deterministic workflows, adding some uncertainty. missing for 10: explicit documentation/example of clarifying-question behavior, independent hands-on proof of multi-step troubleshooting quality.",
    "evidenceIds": [
      "intercom-fin-docs-44",
      "intercom-fin-docs-56",
      "intercom-fin-docs-22",
      "intercom-fin-comm-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "multilingual-answers",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Intercom/Fin explicitly documents multi-language support ('Use Fin AI Agent in multiple languages') allowing Fin to answer in customers' languages even when the knowledge base content is authored in English, which directly matches the story. Missing for 10: independent/hands-on verification of translation quality or explicit detail on how English-only knowledge base content is translated/handled per language.",
    "evidenceIds": [
      "intercom-fin-docs-19"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "omnichannel-coverage",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin explicitly documents multi-channel support spanning live chat/messenger (in-app), email, Slack, WhatsApp, SMS, and voice, with dedicated channel pages for each (docs-9, docs-50, docs-51, docs-65, docs-48). Missing for 10: explicit named coverage of broader 'social' channels like Instagram/Facebook/X/Twitter — only 'and more' is implied, not documented by name, so the specific 'social' claim in the story is only partially evidenced.",
    "evidenceIds": [
      "intercom-fin-docs-9",
      "intercom-fin-docs-48",
      "intercom-fin-docs-50",
      "intercom-fin-docs-51",
      "intercom-fin-docs-65"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "ongoing-qa-review",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin ships explicit QA tooling: fin.ai/testing describes testing across thousands of scenarios before launch and evaluating every live conversation, fin.ai/analyze describes Insights, Monitors, and Recommendations continuously analyzing every conversation, and pricing docs list Custom AI Scorecards and Incident Detection alongside guidance and procedure-editing docs that let fixes feed back into the agent behavior. missing for 10: independent or hands-on verification of the scoring and review loop, and detail on how flagged failures are routed to specific fixes.",
    "evidenceIds": [
      "intercom-fin-docs-45",
      "intercom-fin-docs-59",
      "intercom-fin-docs-47",
      "intercom-fin-docs-61",
      "intercom-fin-docs-16",
      "intercom-fin-docs-22",
      "intercom-fin-docs-23"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "openness-api-parity",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Fin exposes an Agent API, Node SDK, MCP, webhooks, and data-connector APIs for programmatic interaction, but there is no evidence of full UI/API parity — no OpenAPI spec was discoverable (404s on all candidate paths) and no documentation claims that every configuration (e.g., inbox setup, workflows, Procedures, escalation rules) manageable in the UI is also manageable via API. Missing for 10: comprehensive API reference/OpenAPI spec, explicit parity claims, evidence that admin/config UI actions (not just conversational/data actions) are API-accessible.",
    "evidenceIds": [
      "intercom-fin-docs-3",
      "intercom-fin-docs-4",
      "intercom-fin-gh-3",
      "intercom-fin-probe-3",
      "intercom-fin-docs-39"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "openness-full-export",
    "verdict": "partial",
    "quality": 2,
    "confidence": "low",
    "rationale": "The only export-related evidence is a narrow feature to 'Export a saved View exactly as you see it,' not a comprehensive data export mechanism; migration language found in the evidence is framed only around bringing data INTO Intercom from other helpdesks, not exporting all data out in open formats. Missing for 10: bulk/full account data export, documented open file formats (CSV/JSON), a migration-out or offboarding process, and any independent confirmation of successful full data extraction.",
    "evidenceIds": [
      "intercom-fin-docs-2",
      "intercom-fin-docs-36",
      "intercom-fin-docs-55"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "openness-open-license",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Fin is a closed, commercial SaaS AI agent product; there is no indication it is or ever was open-source, and 'read the source under an open license' is not a fair axis for this kind of proprietary hosted service.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "openness-self-host",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Fin is a proprietary SaaS AI agent platform (Intercom); there is no evidence of, nor plausibility for, self-hosting the core product—self-hosting is a category mismatch for this hosted-service type of product.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "personalized-account-answers",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Docs show Fin connects to live business data via data connector templates (Stripe, Shopify, Statuspage), MCP/API data connectors, webhooks, and CRM integrations (Salesforce, HubSpot, Zendesk) that surface data like plan/order status directly in the inbox, plus Procedures that let Fin execute multi-step logic against third-party systems to update/retrieve real account data rather than just canned articles. Missing for 10: independent/hands-on verification that live data grounding actually improves accuracy in practice, and no case study quantifying resolution based specifically on live account/order data vs. help-article content.",
    "evidenceIds": [
      "intercom-fin-docs-34",
      "intercom-fin-docs-35",
      "intercom-fin-docs-11",
      "intercom-fin-intdir-4",
      "intercom-fin-docs-37",
      "intercom-fin-docs-44",
      "intercom-fin-docs-32",
      "intercom-fin-probe-4"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "pre-launch-simulation",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin has a dedicated testing capability documented at fin.ai/testing (\"test it across thousands of scenarios before anything goes live\") and a specific help article \"How to preview and test Fin,\" directly matching the support-ops need to validate the agent before it faces customers. Missing for 10: independent/hands-on corroboration of the testing workflow and explicit mention of importing historical ticket transcripts as test scenarios rather than only generic 'scenarios'.",
    "evidenceIds": [
      "intercom-fin-docs-29",
      "intercom-fin-docs-45",
      "intercom-fin-docs-59",
      "intercom-fin-docs-66"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "privacy-data-residency",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "The evidence pack contains extensive documentation on Fin's features, integrations, and compliance messaging (e.g., 'meet the world's leading compliance standards') but no mention of data residency, region selection, or storage location controls anywhere in the docs or community evidence.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "privacy-no-training",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence in the pack addresses data being used for AI model training or an opt-out/data-use control; general trust/compliance mentions (e.g., trust-reliability page) do not specify training data usage or opt-out mechanisms.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "privacy-retention-controls",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence in the pack addresses data retention policies, deletion controls, or user-facing controls over how long data is kept or when it is deleted; the docs cover integrations, channels, MCP access, and training but never retention/deletion settings.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "privacy-telemetry-optout",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence pack items mention telemetry opt-out, usage tracking controls, or privacy settings for Fin; only trust/compliance marketing language and unrelated docs are present.",
    "evidenceIds": []
  },
  {
    "productId": "intercom-fin",
    "storyId": "procedure-sop-builder",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Fin has a dedicated \"Building Fin Procedures\" feature explicitly for encoding multi-step SOPs with business logic and branching, plus \"Giving Fin Guidance\" and escalation-rule configuration to control deterministic branching to humans or other steps, and a testing suite to validate procedures before launch. This directly matches the story of encoding SOPs with deterministic branching for known issue types. Missing for 10: independent hands-on verification of branching logic depth/determinism and more detail on how conditional branches are authored (docs are descriptive marketing/help-article summaries rather than technical specs).",
    "evidenceIds": [
      "intercom-fin-docs-23",
      "intercom-fin-docs-44",
      "intercom-fin-docs-22",
      "intercom-fin-docs-25",
      "intercom-fin-docs-45",
      "intercom-fin-docs-66"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "resolution-analytics",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Fin's Insights/Monitors/AI Recommendations pages reference analyzing conversations and measuring CX, and CX Score/Trends are listed as features, implying some dashboard capability, but no evidence explicitly confirms a unified dashboard reporting resolution rate, CSAT, handoff rate, and cost per resolution together. missing for 10: explicit documentation or screenshots of a dashboard showing resolution rate, CSAT, handoff rate, and cost per resolution metrics; independent confirmation these specific KPIs are surfaced for exec reporting.",
    "evidenceIds": [
      "intercom-fin-docs-47",
      "intercom-fin-docs-61",
      "intercom-fin-docs-16",
      "intercom-fin-docs-56"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "restricted-topic-controls",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin has documented first-party features for exactly this: 'Giving Fin Guidance' and 'Manage Fin AI Agent's escalation guidance and rules' let ops teams define escalation/human-handoff rules, and the Agent API includes an 'Escalate to a human' capability. However this escalate-to-human orchestration endpoint is explicitly noted as Preview-only, and there is no independent or hands-on evidence confirming the agent reliably respects human-only topic boundaries without ever freelancing. Missing for 10: independent verification that escalation rules are strictly enforced (no freelancing), and GA (non-preview) status of escalate-to-human tooling.",
    "evidenceIds": [
      "intercom-fin-docs-22",
      "intercom-fin-docs-25",
      "intercom-fin-docs-4"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "supervised-draft-mode",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows Fin's guardrails center on escalation rules, hand-off to human agents, and pre-launch testing/preview (docs-25, docs-29, docs-31, docs-45), but nothing describes a supervised/draft-for-approval mode where Fin composes replies that a human must approve before they reach the customer on every interaction. Community commentary (comm-3) even argues for more deterministic, human-controlled workflows, but that's not evidence Fin ships this specific approval-gate mode.",
    "evidenceIds": [
      "intercom-fin-docs-25",
      "intercom-fin-docs-29",
      "intercom-fin-docs-31",
      "intercom-fin-docs-45",
      "intercom-fin-comm-3"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "topic-trend-insights",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Fin's pricing page lists analytics features like 'AI Topics, Trends... Monitors... Incident Detection' and 'Insights continuously analyzes every Fin and human conversation,' indicating topic clustering and emerging-issue detection capability, but there is no deeper documentation, screenshots, or independent corroboration of how this works in practice. Missing for 10: detailed product docs on topic clustering methodology, evidence of proactive alerting before ticket-volume spikes, and independent/hands-on validation of the insights/analytics feature.",
    "evidenceIds": [
      "intercom-fin-docs-16",
      "intercom-fin-docs-47",
      "intercom-fin-docs-61"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "transparent-per-resolution-pricing",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "A community source states Fin charges '$1 per successful customer session', which is consistent with outcome-based, per-resolution pricing, and a dedicated fin.ai/pricing page exists (cited repeatedly for feature lists). However, none of the evidence shows the actual published rate table, caps, or usage controls on that pricing page — the citations only reference features (integrations, channels, AI Agent capabilities), not price mechanics or caps. Missing for 10: first-party documentation of the per-resolution rate/caps/controls on fin.ai/pricing, and independent corroboration beyond a single HN comment.",
    "evidenceIds": [
      "intercom-fin-comm-2",
      "intercom-fin-docs-7"
    ]
  },
  {
    "productId": "intercom-fin",
    "storyId": "voice-phone-support",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Fin Voice is a documented dedicated product ('Deploy Fin Voice', 'Fin Voice 2 runs on Apex Flash... latency-sensitive tasks', 'grounded in your knowledge base, and trained to apply your specific policies on every call'), indicating speech-in/speech-out phone support using the same knowledge base as chat. Missing for 10: independent/hands-on validation of voice call quality or action-taking parity specifically for phone, and no detail on how actions (e.g., updating systems) work identically on voice vs chat.",
    "evidenceIds": [
      "intercom-fin-docs-18",
      "intercom-fin-docs-49",
      "intercom-fin-docs-64",
      "intercom-fin-docs-48"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-agent-docs",
    "verdict": "full",
    "quality": 9,
    "confidence": "high",
    "rationale": "Direct probe evidence confirms Lorikeet serves an llms.txt file at docs.lorikeetcx.ai/llms.txt returning HTTP 200 with structured agent-oriented reference links, and documentation is further organized around MCP/agent access. Missing for 10: no independent third-party report of an agent successfully consuming this llms.txt in practice.",
    "evidenceIds": [
      "lorikeet-probe-1",
      "lorikeet-probe-2",
      "lorikeet-docs-46"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-ai-insights",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Lorikeet's Coach and MCP-server capabilities generate AI-driven insights directly from customer data: ticket quality scoring across 100% of conversations, knowledge-base gap/quality audits, root-cause diagnosis of tickets, and analytics on resolution quality, satisfaction, and revenue impact, with Coach able to 'implement improvements... or make suggestions for you to action yourself.' This is all first-party documentation without independent hands-on corroboration. Missing for 10: independent/third-party validation of insight quality and real-world usage examples beyond vendor docs.",
    "evidenceIds": [
      "lorikeet-docs-9",
      "lorikeet-docs-16",
      "lorikeet-docs-15",
      "lorikeet-docs-23",
      "lorikeet-docs-48",
      "lorikeet-docs-33",
      "lorikeet-docs-30"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-autonomous-automation",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet's core product is an autonomous agent that resolves tickets end-to-end across channels and coordinates specialist agents for multi-step workflows, and Outbound campaigns run on scheduled cadences without human intervention, while Coach can autonomously implement improvements. However, the evidence centers on the vendor's own agent running in background rather than a user-configurable 'automation' builder with triggers/schedules exposed as a general-purpose feature. Missing for 10: explicit user-facing scheduling/trigger configuration UI or API for arbitrary automations, and independent/hands-on confirmation that these automations run unattended reliably.",
    "evidenceIds": [
      "lorikeet-docs-31",
      "lorikeet-docs-32",
      "lorikeet-docs-40",
      "lorikeet-docs-48",
      "lorikeet-docs-28"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-builtin-assistant",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet's 'Coach' is a built-in AI assistant accessible directly inside the Lorikeet platform (as well as via Slack/Claude/ChatGPT/MCP) that users can delegate tasks to — diagnosing tickets, auditing knowledge bases, building workflows via natural language, running simulations, and even implementing improvements automatically on the user's behalf. This is well documented across multiple first-party pages describing concrete delegated actions (e.g. doc-48 'Coach can implement improvements on your behalf').  missing for 10: independent/hands-on verification of Coach's assistant behavior, and clearer distinction between autonomous action vs. suggestion-only mode.",
    "evidenceIds": [
      "lorikeet-docs-9",
      "lorikeet-docs-10",
      "lorikeet-docs-15",
      "lorikeet-docs-16",
      "lorikeet-docs-17",
      "lorikeet-docs-33",
      "lorikeet-docs-48"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-headless",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Lorikeet is a customer-support AI concierge platform, not a dev-tool/CLI-type product meant to be run headlessly in CI pipelines; its MCP server is for interactive assistant use (Claude, ChatGPT), not CI automation. This axis is a category error for this product type.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-mcp-client",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Lorikeet's docs state the agent 'takes action through your APIs and MCP servers inside natural-language and deterministic workflows' (lorikeet-docs-28), indicating it can consume external MCP servers as tool sources, but the bulk of the MCP evidence pack actually describes the reverse direction — Lorikeet exposing its own MCP server for external clients like Claude/ChatGPT to connect to (lorikeet-docs-1, lorikeet-docs-46, lorikeet-probe-1/2). Missing for 10: dedicated documentation on how a user configures/adds third-party MCP servers into Lorikeet, a list of supported MCP integrations, and independent confirmation of this client-side tool-use capability.",
    "evidenceIds": [
      "lorikeet-docs-28",
      "lorikeet-docs-13"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-mcp-server",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet publishes an official MCP server (docs.lorikeetcx.ai/mcp/mcp-server) that lets Claude, ChatGPT, Claude Code, Codex, and MintMCP connect directly to a Lorikeet account, with documented capabilities like diagnosing tickets, auditing knowledge bases, building workflows, testing tools, and running simulations via MCP. This is confirmed live (HTTP 200) via probe evidence, not just marketing copy. Missing for 10: independent/hands-on third-party verification of the MCP server working end-to-end, and no community reports corroborating reliability.",
    "evidenceIds": [
      "lorikeet-docs-1",
      "lorikeet-docs-2",
      "lorikeet-docs-3",
      "lorikeet-docs-4",
      "lorikeet-docs-5",
      "lorikeet-docs-46",
      "lorikeet-probe-1",
      "lorikeet-probe-2"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-nl-commands",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Lorikeet ships an MCP server plus Claude/ChatGPT/Codex integrations that let users build workflows, diagnose tickets, audit knowledge bases, and run simulations using natural-language commands (e.g. 'Build workflows - create and iterate on workflows using natural language', example prompts like 'Test the get-order-status tool...'), and even exposes slash-command skills like /lorikeet:create-simulations. This is first-party documentation only, with no independent/hands-on corroboration of the NL command experience. Missing for 10: independent user reports or demos validating the natural-language MCP workflow in practice.",
    "evidenceIds": [
      "lorikeet-docs-1",
      "lorikeet-docs-4",
      "lorikeet-docs-17",
      "lorikeet-docs-19",
      "lorikeet-docs-24",
      "lorikeet-docs-46",
      "lorikeet-probe-2"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-official-cli",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Lorikeet is a customer-support AI platform; the story asks for an official CLI for AI-native workflows. Evidence shows an MCP server and Claude-code skills/slash commands but no standalone CLI tool is documented anywhere in the pack. missing for 10: any mention of a CLI binary/tool, installation instructions, or CLI command reference.",
    "evidenceIds": [
      "lorikeet-docs-1",
      "lorikeet-docs-24",
      "lorikeet-probe-2"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-public-api",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Lorikeet publishes a documented MCP server (docs.lorikeetcx.ai/mcp/mcp-server) that lets AI-native users drive the product directly from Claude, ChatGPT, Codex, and Claude Code — diagnosing tickets, auditing knowledge bases, building workflows, testing tools, and running simulations, all documented with concrete examples and even slash-command skills. This is a genuine documented programmatic interface built for AI agents, not just human UI docs. Missing for 10: no separate traditional REST/GraphQL API reference beyond MCP, and no independent third-party corroboration of the API's reliability.",
    "evidenceIds": [
      "lorikeet-docs-1",
      "lorikeet-docs-2",
      "lorikeet-docs-3",
      "lorikeet-docs-4",
      "lorikeet-docs-5",
      "lorikeet-docs-19",
      "lorikeet-docs-20",
      "lorikeet-docs-46",
      "lorikeet-docs-47",
      "lorikeet-probe-1",
      "lorikeet-probe-2"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-scoped-keys",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Lorikeet's guardrails page mentions 'workflow-scoped tool access' and 'server-side identity validation' enforced in code, suggesting some access scoping, but there is no explicit documentation of issuing or managing scoped/least-privilege API credentials or tokens for agents. Missing for 10: explicit credential/token issuance mechanism, documentation of API key scoping or permission granularity, and any user-facing controls for creating least-privilege credentials.",
    "evidenceIds": [
      "lorikeet-docs-34"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-sdks",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence only documents an MCP server and integrations/APIs for connecting tools, but there is no mention of an official SDK (e.g., Python/JS client library) for developers to build against. Missing for 10: any documented official SDK, its language support, or developer-facing library docs.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "agentic-webhooks",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "The evidence pack covers Lorikeet's MCP server, simulations, coach, and guardrails features, but contains no mention of webhooks or event subscription mechanisms for AI-native users to receive push notifications on events.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "api-interactive-docs",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Lorikeet is a customer-support AI agent platform, not a developer API/SDK product; there is no evidence of an interactive API reference or runnable-example explorer, and this axis doesn't fit its product category (its docs cover MCP server usage and skills, not an API playground).",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "api-machine-spec",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence of a downloadable OpenAPI spec or machine-readable API documentation; evidence only covers MCP server integration, workflows, and product features, not a formal API spec artifact.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "api-sandbox",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Evidence shows robust simulation/testing tooling (replay historical tickets, synthetic scenarios, guardrail adversarial tests) that approximates sandbox-style testing, but no explicit claim of an isolated sandbox environment distinct from production. missing for 10: explicit documentation of a dedicated sandbox/staging environment, confirmation that simulations do not touch or affect live production data/systems, and independent verification of this isolation.",
    "evidenceIds": [
      "lorikeet-docs-6",
      "lorikeet-docs-18",
      "lorikeet-docs-21",
      "lorikeet-docs-26",
      "lorikeet-docs-37",
      "lorikeet-docs-47"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "api-versioning-policy",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence pack item mentions API versioning, version numbers, or a deprecation policy for Lorikeet's APIs or MCP server; the docs discuss features and integrations but not API lifecycle/versioning commitments. Missing for 10: any documented API version scheme, changelog of breaking changes, or stated deprecation/support timeline.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "automation-bulk-operations",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet documents multiple bulk operations available to AI-native users via its MCP server and product surfaces: running simulations in bulk batches across hundreds of scenarios, auditing entire knowledge bases at scale, and reviewing 100% of conversations for quality (not manual spot checks). These are explicitly framed as batch/bulk actions accessible through natural-language or MCP-driven workflows. Missing for 10: independent/hands-on verification of bulk operation scale and performance, and no explicit example of bulk edits/updates to many tickets or records simultaneously (only bulk testing/auditing/review are documented).",
    "evidenceIds": [
      "lorikeet-docs-6",
      "lorikeet-docs-7",
      "lorikeet-docs-3",
      "lorikeet-docs-16",
      "lorikeet-docs-9",
      "lorikeet-docs-38",
      "lorikeet-docs-21",
      "lorikeet-docs-26",
      "lorikeet-docs-22"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "automation-rules-engine",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Lorikeet's workflows and guardrails encode conditional, event-triggered actions (e.g., prompt-injection detection triggers block/rewrite/escalate, bad QA scores trigger refunds, campaign cadences control scheduled outreach, escalation triggers hand off to humans), and workflows can be built/edited via natural language including deterministic steps. However this is more built-in platform logic than a general-purpose rule-definition interface for arbitrary custom events an AI-native user could freely wire up. Missing for 10: a documented general rules/automation engine or API letting users define arbitrary trigger-condition-action rules beyond the platform's fixed guardrail/workflow/outbound features, and independent confirmation of this working in practice.",
    "evidenceIds": [
      "lorikeet-docs-34",
      "lorikeet-docs-35",
      "lorikeet-docs-39",
      "lorikeet-docs-40",
      "lorikeet-docs-44",
      "lorikeet-docs-32",
      "lorikeet-docs-28"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "automation-scheduled-jobs",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "The only relevant evidence is outbound campaign cadences with 'scheduling windows' controlling when outreach happens, which implies some recurring/scheduled automation but is narrowly scoped to outbound messaging rather than general recurring jobs or workflow runs. Missing for 10: explicit cron-like or recurring workflow scheduling for MCP-driven tasks (simulations, audits, diagnostics), documentation of scheduling frequency/config options, and any independent confirmation of recurring job execution.",
    "evidenceIds": [
      "lorikeet-docs-40"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "automation-versioned-workflows",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Lorikeet's docs show workflow building/iteration via natural language and strong review tooling (simulations, batch comparisons showing how an edit changed outcomes across scenarios), which covers the 'review' part of the story. However, there is no explicit mention of a version history or a rollback/undo mechanism for automations. missing for 10: explicit versioning/change-history feature, explicit rollback/undo capability, evidence of restoring a prior workflow state.",
    "evidenceIds": [
      "lorikeet-docs-4",
      "lorikeet-docs-17",
      "lorikeet-docs-7",
      "lorikeet-docs-22",
      "lorikeet-docs-6"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "brand-voice-tone",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Lorikeet docs show the agent can be trained on 'business context, brand guidelines, help docs and standard operating procedures' (lorikeet-docs-25), and Coach's Ticket Quality Score reviews 100% of conversations against quality standards to catch drift (lorikeet-docs-9, lorikeet-docs-38), supporting brand-voice control and consistency monitoring. However, there is no explicit evidence of multi-language tone consistency or dedicated brand-voice/style configuration tooling beyond general training inputs. Missing for 10: explicit multilingual consistency support, dedicated tone/voice configuration UI, and independent evidence of voice consistency across topics/languages.",
    "evidenceIds": [
      "lorikeet-docs-25",
      "lorikeet-docs-9",
      "lorikeet-docs-38",
      "lorikeet-docs-27"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "configurable-escalation-rules",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Docs show human handoff when the AI can't resolve a ticket, escalation for regulatory/specialist cases with full context, guardrails that can 'escalate' on policy violations, and deployment scoped to trained topics — supporting topic-based and inability/explicit-need escalation. However there is no explicit mention of configuring handoff by sentiment or customer tier, nor independent verification that these rules are 'reliably obeyed'. Missing for 10: explicit sentiment-based trigger config, explicit customer-tier-based trigger config, and independent/hands-on evidence of reliability.",
    "evidenceIds": [
      "lorikeet-docs-14",
      "lorikeet-docs-27",
      "lorikeet-docs-35",
      "lorikeet-docs-44",
      "lorikeet-docs-12"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "content-gap-detection",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Lorikeet explicitly documents auditing knowledge bases at scale to find gaps and quality issues, and its simulation/replay tooling surfaces knowledge gaps by replaying historical tickets and synthetic scenarios to project resolution quality before deployment. Coach's Ticket Quality Score also reviews 100% of conversations to catch fumbled answers and feed fixes back. Missing for 10: independent/hands-on validation of the gap-detection accuracy and no explicit mention of detecting 'conflicting' content specifically (only gaps/outdated/quality issues).",
    "evidenceIds": [
      "lorikeet-docs-3",
      "lorikeet-docs-16",
      "lorikeet-docs-21",
      "lorikeet-docs-26",
      "lorikeet-docs-9",
      "lorikeet-docs-38",
      "lorikeet-docs-33"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "context-rich-handoff",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Lorikeet documents human handoff happening 'inside the same platform' when AI can't resolve, and specifically claims escalations carry 'full interaction history so your team picks up mid-conversation' (financial-services vertical) — directly supporting no-repeat handoff. However, evidence doesn't explicitly confirm a generated summary or structured 'collected details' package accompanying every escalation across all verticals, only interaction history. Missing for 10: explicit documentation of an auto-generated conversation summary at escalation, structured collected-details extraction, and independent/customer verification that customers never repeat themselves in practice.",
    "evidenceIds": [
      "lorikeet-docs-14",
      "lorikeet-docs-44",
      "lorikeet-docs-45"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "end-to-end-resolution",
    "verdict": "partial",
    "quality": 6,
    "confidence": "low",
    "rationale": "Lorikeet's docs repeatedly claim end-to-end resolution (chat, email, voice, SMS), with escalation to humans when it can't resolve, plus QA scoring (Ticket Quality Score) and analytics tracking 'resolution quality' as an outcome metric, and even a refund-on-bad-score mechanism tied to quality. However, all evidence is vendor-authored marketing/docs; there are no independent benchmarks, customer case studies, or hard resolution-rate numbers (e.g., % of conversations fully resolved) to substantiate the claims. Missing for 10: independent/third-party resolution-rate data, customer-reported metrics, and clear definition/measurement methodology distinguishing true resolution from deflection.",
    "evidenceIds": [
      "lorikeet-docs-31",
      "lorikeet-docs-49",
      "lorikeet-docs-14",
      "lorikeet-docs-23",
      "lorikeet-docs-38",
      "lorikeet-docs-39",
      "lorikeet-docs-9"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "grounded-cited-answers",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Lorikeet documents training its agent on 'business context, brand guidelines, help docs and standard operating procedures' (docs-25) and mentions 'grounding' as a guardrail check (docs-35), but there is no evidence that end-user-facing answers cite or display the specific article/source used to generate a response. Missing for 10: any documented citation/source-attribution UI or API in agent responses, and independent confirmation that answers reference specific knowledge-base articles.",
    "evidenceIds": [
      "lorikeet-docs-25",
      "lorikeet-docs-35",
      "lorikeet-docs-16"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "hallucination-guardrails",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Lorikeet documents explicit guardrails that enforce grounding and policy compliance in code rather than prompts, including response screening for grounding/policy with block/rewrite/escalate actions, and adversarial simulation testing for false authority claims and prompt injection to validate safe-decline behavior before deployment. This directly targets stopping invented policies/prices/promises via server-side enforcement rather than relying on model honesty. Missing for 10: independent/third-party verification or a concrete hands-on example showing an off-knowledge question actually triggering a safe decline rather than a hallucinated answer.",
    "evidenceIds": [
      "lorikeet-docs-34",
      "lorikeet-docs-35",
      "lorikeet-docs-36",
      "lorikeet-docs-37",
      "lorikeet-docs-45"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "helpdesk-platform-integrations",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Docs claim Lorikeet works 'alongside your existing tools' and connects 'in seconds to your ticketing system' via integrations, implying helpdesk co-existence, but there is no specific evidence of named Zendesk/Salesforce/Intercom connectors or two-way ticket/context sync — the pack focuses on MCP client integrations (Claude, ChatGPT) rather than helpdesk platforms. missing for 10: named Zendesk/Salesforce/Intercom integration docs, evidence of bidirectional ticket sync, standalone-mode confirmation, independent corroboration of integration reliability.",
    "evidenceIds": [
      "lorikeet-docs-11",
      "lorikeet-docs-13",
      "lorikeet-docs-14",
      "lorikeet-docs-28"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "knowledge-auto-sync",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows Lorikeet ingests knowledge bases and can audit them for gaps/outdated articles, and connects to ticketing/knowledge sources 'in seconds', but there is no mention of scheduled or change-triggered re-syncing of sources without manual re-upload. missing for 10: explicit scheduled/webhook-based re-sync mechanism, evidence of automatic detection of source updates, any documentation of sync cadence or on-change triggers.",
    "evidenceIds": [
      "lorikeet-docs-3",
      "lorikeet-docs-13",
      "lorikeet-docs-16",
      "lorikeet-docs-25"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "knowledge-source-ingestion",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet explicitly claims to 'seamlessly ingest data' from ticketing systems and knowledge bases (docs-13), to train the agent on 'business context, brand guidelines, help docs and standard operating procedures' (docs-25), to audit the entire knowledge base for gaps/outdated content (docs-3, docs-16), and to replay historical tickets to surface knowledge gaps (docs-21, docs-26) — covering help center, docs, and past tickets without manual re-authoring. Internal wikis are not explicitly named as a source type, and all evidence is first-party vendor documentation with no independent/hands-on corroboration of ingestion working end-to-end. Missing for 10: explicit mention of internal wiki ingestion, and independent verification of the ingestion pipeline's fidelity/accuracy.",
    "evidenceIds": [
      "lorikeet-docs-13",
      "lorikeet-docs-25",
      "lorikeet-docs-3",
      "lorikeet-docs-16",
      "lorikeet-docs-21",
      "lorikeet-docs-26"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "live-api-actions",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Docs explicitly state the agent 'takes action through your APIs and MCP servers' (docs-28), give concrete examples like refund workflows and order-status tool tests (docs-19, docs-24, docs-43), and describe workflow-scoped tool access, server-side identity validation, and hard execution caps enforced in code rather than prompts (docs-34). This directly matches the story of scoped, real-world API actions like refunds/order changes/subscriptions. Missing for 10: independent/hands-on verification of the auth-scoping mechanism, and more granular detail on how 'per action' scopes are configured/enforced beyond high-level guardrails language.",
    "evidenceIds": [
      "lorikeet-docs-28",
      "lorikeet-docs-34",
      "lorikeet-docs-19",
      "lorikeet-docs-24",
      "lorikeet-docs-43",
      "lorikeet-docs-13"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "multi-turn-troubleshooting",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Lorikeet's docs describe workflows that coordinate multi-step, multi-system troubleshooting (e.g. 'coordinates a team of specialist agents... to handle multi-party, multi-system workflows end-to-end' and 'diagnose tickets by tracing workflow execution') rather than single canned replies, and its agent resolves issues 'end-to-end' across channels. However, there is no explicit documentation showing the agent proactively asking clarifying questions mid-conversation or examples of dialogue-based troubleshooting turns. Missing for 10: explicit examples/transcripts of clarifying-question behavior, documentation describing conversational back-and-forth troubleshooting logic rather than just workflow/tool orchestration.",
    "evidenceIds": [
      "lorikeet-docs-32",
      "lorikeet-docs-31",
      "lorikeet-docs-15",
      "lorikeet-docs-17",
      "lorikeet-docs-28"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "multilingual-answers",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence in the pack addresses multilingual support or the ability to serve customers in languages other than English despite an English-only knowledge base; nothing about translation, language detection, or multilingual training is mentioned.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "omnichannel-coverage",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Lorikeet documents one agent resolving issues across chat, email, voice, SMS and WhatsApp (docs-31, docs-49), and mentions Slack as a place to interact with Coach (docs-10), but this is Coach access, not evidence that the customer-facing agent itself operates in Slack or social channels. missing for 10: explicit confirmation that the customer-facing support agent (not just Coach) handles Slack and social media channels, and independent/hands-on validation of omnichannel deployment.",
    "evidenceIds": [
      "lorikeet-docs-31",
      "lorikeet-docs-49",
      "lorikeet-docs-10",
      "lorikeet-docs-14"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "ongoing-qa-review",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Lorikeet's Coach product directly addresses this story: it reviews 100% of conversations against quality standards (not just samples), scores them via a Ticket Quality Score, flags failures (bad scores trigger refunds), and turns findings into fixes that feed back into the agent so it 'improves every week instead of drifting.' Simulations complement this with batch testing and adversarial scenario scoring tied to guardrails and analytics tracking. Missing for 10: independent/hands-on verification of the review loop in practice and more detail on how flagged failures are triaged/assigned to human reviewers.",
    "evidenceIds": [
      "lorikeet-docs-9",
      "lorikeet-docs-33",
      "lorikeet-docs-38",
      "lorikeet-docs-39",
      "lorikeet-docs-48",
      "lorikeet-docs-36",
      "lorikeet-docs-6",
      "lorikeet-docs-37"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "openness-api-parity",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Lorikeet's MCP server exposes a broad set of actions also available in the UI—diagnosing tickets, auditing knowledge bases, building/testing workflows, running simulations—letting AI-native users act via Claude/ChatGPT/Codex instead of the dashboard. However, there's no evidence of a general-purpose REST/GraphQL API or explicit claim of full UI/API parity (e.g., outbound campaign management, guardrail configuration, analytics dashboards aren't shown as MCP-accessible). Missing for 10: explicit full API parity statement, broader API coverage beyond MCP-listed actions, independent confirmation of completeness.",
    "evidenceIds": [
      "lorikeet-docs-1",
      "lorikeet-docs-2",
      "lorikeet-docs-3",
      "lorikeet-docs-4",
      "lorikeet-docs-5",
      "lorikeet-docs-20",
      "lorikeet-docs-47",
      "lorikeet-docs-28",
      "lorikeet-probe-2"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "openness-full-export",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence pack items mention data export, open data formats, or account/data portability for leaving the platform; the docs cover MCP integrations, workflows, simulations, and guardrails but nothing about exporting or migrating data out. Missing for 10: any documentation of export functionality, supported open formats (CSV/JSON), or data portability/exit process.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "openness-open-license",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Lorikeet is a closed commercial SaaS customer-support platform; no evidence of any open-source license or public source code repository is provided anywhere in the evidence pack.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "openness-self-host",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Lorikeet is a hosted SaaS customer-support AI platform with no evidence of any self-hosted/on-prem deployment option; self-hosting is not a fair axis for this category of cloud service, so this is a category mismatch rather than a missing capability.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "personalized-account-answers",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet's docs describe connecting to ticketing systems, knowledge bases, and internal tools/APIs to 'ingest data and take action for your customers' (lorikeet-docs-13, lorikeet-docs-28, lorikeet-docs-29), with concrete examples like testing a 'get-order-status' tool with a real order ID (lorikeet-docs-19) and financial-services use cases like disputes/loan inquiries requiring account-specific data (lorikeet-docs-43, lorikeet-docs-44). This shows the agent is designed to pull and act on live customer data rather than just static help content.  missing for 10: independent/hands-on verification that responses actually reflect real-time account state in production, and more detail on latency/freshness guarantees for live data lookups.",
    "evidenceIds": [
      "lorikeet-docs-13",
      "lorikeet-docs-19",
      "lorikeet-docs-28",
      "lorikeet-docs-29",
      "lorikeet-docs-43",
      "lorikeet-docs-44"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "pre-launch-simulation",
    "verdict": "full",
    "quality": 9,
    "confidence": "high",
    "rationale": "Lorikeet's Simulations product directly supports this story: it generates simulations from actual historical tickets and runs them in bulk batches before workflow changes go live, with side-by-side batch comparisons and per-conversation drill-downs, plus authored adversarial/guardrail scenarios to pre-test against tricky real-world behavior. Docs also describe replaying historical tickets and synthetic scenarios in bulk to project resolution quality before deploying on trained topics. Missing for 10: independent/hands-on corroboration beyond first-party docs.",
    "evidenceIds": [
      "lorikeet-docs-6",
      "lorikeet-docs-7",
      "lorikeet-docs-8",
      "lorikeet-docs-21",
      "lorikeet-docs-26",
      "lorikeet-docs-47"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "privacy-data-residency",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "The evidence covers compliance/trust items (Vanta certifications, zero-data-retention with model vendors) but no mention of data residency or region-selection options for storage. Missing for 10: any documentation of regional data storage choices, residency guarantees, or data localization controls.",
    "evidenceIds": []
  },
  {
    "productId": "lorikeet",
    "storyId": "privacy-no-training",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet explicitly states zero-data-retention agreements with all model vendors and no fine-tuning on customer data, directly addressing the story's request to prevent data from being used for AI training. This is backed by independently verified trust/compliance reports on their Vanta Trust Center. Missing for 10: independent hands-on verification or third-party audit confirmation of this specific claim, and no detail on user-level opt-out controls or granularity of enforcement.",
    "evidenceIds": [
      "lorikeet-docs-42",
      "lorikeet-docs-41"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "privacy-retention-controls",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Lorikeet mentions zero-data-retention agreements with model vendors and no fine-tuning on customer data, plus SOC2-style independently verified reports on a trust center, which touches data retention posture at the vendor-model level. However, there is no evidence of user-facing controls letting an AI-native user configure or request deletion/retention of their own conversation or account data within Lorikeet itself. Missing for 10: explicit customer-data deletion/export controls, retention period configuration, and user-initiated deletion workflows.",
    "evidenceIds": [
      "lorikeet-docs-41",
      "lorikeet-docs-42"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "privacy-telemetry-optout",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Lorikeet is a customer-support AI platform, and telemetry opt-out for the product itself is a fair privacy-posture question, but none of the evidence mentions any telemetry/usage-tracking opt-out mechanism for users of the product; it only discusses data retention with model vendors and consent handling for outbound customer messaging, which is unrelated to product telemetry opt-out.",
    "evidenceIds": [
      "lorikeet-docs-41",
      "lorikeet-docs-42"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "procedure-sop-builder",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Lorikeet explicitly supports training agents on SOPs (docs-25) and building workflows with 'natural-language and deterministic workflows' plus 'pockets of determinism for regulated steps' (docs-28, docs-32), with workflow-scoped tool access and hard execution caps enforced in code (docs-34), directly matching the deterministic-branching SOP story. Missing for 10: independent/hands-on validation of branching logic in practice and more detail on how branching conditions are authored beyond natural-language workflow builder claims.",
    "evidenceIds": [
      "lorikeet-docs-25",
      "lorikeet-docs-28",
      "lorikeet-docs-32",
      "lorikeet-docs-34",
      "lorikeet-docs-4",
      "lorikeet-docs-17"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "resolution-analytics",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Lorikeet claims analytics tracking 'resolution quality, customer satisfaction, revenue impact, and operational efficiency' and that 'every event lands in your analytics as a tracked outcome,' which loosely maps to resolution rate, CSAT, and cost metrics, but there is no explicit mention of a handoff-rate metric or a dedicated exec-facing dashboard view combining these four numbers. missing for 10: explicit handoff-rate metric, dedicated dashboard UI/screenshot evidence, cost-per-resolution specifics, and independent corroboration of dashboard usability.",
    "evidenceIds": [
      "lorikeet-docs-23",
      "lorikeet-docs-30",
      "lorikeet-docs-36"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "restricted-topic-controls",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Lorikeet documents topic-scoped deployment (\"deploy on the topics the agent is trained for and leave the rest with your team\"), workflow-scoped tool access with hard execution caps enforced in code, and escalation paths for regulated/specialist cases with full history handoff — all consistent with restricting the agent from acting on sensitive topics. However, there's no explicit documentation of a UI/config for support-ops to designate specific topics (e.g., legal threats, cancellations, security) as strictly human-only with enforced non-freelancing. Missing for 10: explicit topic-tagging/human-only designation feature, confirmation that agent cannot even attempt those topics (vs. escalating mid-conversation), and independent verification of this behavior in production.",
    "evidenceIds": [
      "lorikeet-docs-27",
      "lorikeet-docs-32",
      "lorikeet-docs-34",
      "lorikeet-docs-44",
      "lorikeet-docs-45"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "supervised-draft-mode",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Lorikeet's documented model is autonomous resolution with post-hoc QA (Coach reviews 100% of tickets after the fact) and escalation to humans only when the AI can't resolve a case, not a pre-send draft-for-approval workflow. No evidence describes a mode where every agent reply is queued for human sign-off before reaching a customer. Missing for 10: any mention of a draft/approve workflow, human-in-the-loop send gating, or a configurable 'supervised' launch mode.",
    "evidenceIds": [
      "lorikeet-docs-14",
      "lorikeet-docs-35",
      "lorikeet-docs-38",
      "lorikeet-docs-44"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "topic-trend-insights",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows general analytics and quality-review features (e.g., 'Track resolution quality, customer satisfaction... with industry-leading analytics', Coach reviewing 100% of conversations) but no mention of topic clustering of conversations or proactive surfacing of emerging product issues before ticket-volume spikes. missing for 10: topic clustering of conversations, trend/anomaly detection for emerging issues, ticket-volume spike prediction or alerting.",
    "evidenceIds": [
      "lorikeet-docs-23",
      "lorikeet-docs-30",
      "lorikeet-docs-9"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "transparent-per-resolution-pricing",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "The only evidence touching pricing-economics is the refund policy tying billing to Coach's quality score ('When Coach gives a conversation a bad score, we refund the AI portion of that interaction'), which shows some outcome-linkage in billing, but there is no published price list, per-resolution rate, cap structure, or self-serve pricing page. Missing for 10: a published price sheet or rate card, explicit per-resolution cost, spend/volume caps, and any evidence pricing is transparent versus custom enterprise quoting.",
    "evidenceIds": [
      "lorikeet-docs-39"
    ]
  },
  {
    "productId": "lorikeet",
    "storyId": "voice-phone-support",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Lorikeet's marketing claims 'one agent that resolves issues end-to-end across chat, email, voice and SMS' and lists phone/voice among supported channels, implying the same agent handles calls. However, there is no detail on speech-in/speech-out mechanics, telephony integration, or evidence that voice interactions carry the same knowledge/actions/guardrails as chat beyond a generic channel list. missing for 10: specifics on speech recognition/TTS, call-handling architecture, latency/quality benchmarks, and confirmation that voice shares the same knowledge base and action set as chat.",
    "evidenceIds": [
      "lorikeet-docs-31",
      "lorikeet-docs-49"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-agent-docs",
    "verdict": "full",
    "quality": 9,
    "confidence": "high",
    "rationale": "Direct probe evidence confirms Parahelp's docs serve a working llms.txt (HTTP 200) and per-page .md mirrors via Mintlify, making the docs agent-legible and directly consumable by an AI agent. Missing for 10: no independent third-party report of an agent actually consuming this llms.txt successfully in practice.",
    "evidenceIds": [
      "parahelp-probe-1",
      "parahelp-probe-rt-1"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-ai-insights",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Parahelp's Internal Agent generates insights from data: it analyzes historical tickets to auto-build a Customer Agent (docs-19) and runs automations that surface gaps in tools/support queues on their own (docs-17). This is AI-generated insight/suggestion behavior tied to the user's own data, but it's narrowly scoped to ticket/gap analysis rather than a general-purpose insights feature across all product data. Missing for 10: a dedicated insights/analytics dashboard surfacing trends or recommendations beyond gap detection and agent bootstrapping, and independent/hands-on confirmation of insight quality.",
    "evidenceIds": [
      "parahelp-docs-17",
      "parahelp-docs-19"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-autonomous-automation",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Docs explicitly describe automations as Internal Agent tasks that run on a schedule or trigger (e.g., weekday 9am, GitHub PR merge), configurable with auto-approve so they act without waiting for manual approval, enabling background autonomous operation. Additional support covers auto-approve per tool/task, API-driven triggering, and monitoring/reverting via releases. Missing for 10: independent third-party verification of long-running unattended automations in production and more detail on failure/alerting handling during autonomous runs.",
    "evidenceIds": [
      "parahelp-docs-5",
      "parahelp-docs-3",
      "parahelp-docs-8",
      "parahelp-docs-14",
      "parahelp-docs-4",
      "parahelp-docs-17"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-builtin-assistant",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Parahelp's Internal Agent is a built-in AI assistant you delegate configuration and operational tasks to via natural language (docs-1), which can act autonomously per auto-approve settings (docs-8, docs-14), run on schedules/triggers (docs-5), and proactively surface gaps (docs-17), with the agent also bootstrapping itself from historical tickets (docs-19). Missing for 10: independent/hands-on user reports of delegating tasks and outcomes, and more detail on the breadth/limits of what can be delegated beyond configuration and support workflows.",
    "evidenceIds": [
      "parahelp-docs-1",
      "parahelp-docs-19",
      "parahelp-docs-17",
      "parahelp-docs-5",
      "parahelp-docs-8",
      "parahelp-docs-14"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-headless",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Parahelp exposes a documented API for triggering agent actions programmatically, supports auto-approve mode to bypass manual gating, and automations can run on schedules or triggers like a GitHub PR merge — all consistent with headless/CI-style automation. However, there's no explicit CLI, CI-pipeline example, or SDK documentation, and no independent/hands-on confirmation of running it in a real CI environment. Missing for 10: dedicated CLI or CI-integration docs, and third-party confirmation of headless automation in production.",
    "evidenceIds": [
      "parahelp-docs-5",
      "parahelp-docs-8",
      "parahelp-docs-14",
      "parahelp-docs-13",
      "parahelp-docs-20",
      "parahelp-probe-rt-2"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-mcp-client",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Parahelp's tools connect via its own APIs and integrations (Stripe, internal endpoints, ticketing systems) but there is no mention of MCP server support or the ability to plug in MCP servers so the agent can use their tools. Missing for 10: any reference to MCP protocol, MCP server connection, or standardized tool-plugin interface.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-mcp-server",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Parahelp is a customer/internal support agent product, not itself an agent client seeking to connect to external MCP servers, nor is there evidence it exposes a first-party MCP server for other agents to connect to — its integration surface is API/webhooks and ticketing-system connections, not MCP. This is a category mismatch for the MCP-server axis.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-nl-commands",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Parahelp's core interaction model is describing desired flows, policies, and guardrails in natural language, which the Internal Agent then configures into tasks, automations, and guardrail rules (e.g. 'Require Slack approval before refund tool runs above $100'). This covers configuring agent behavior via NL, though it's scoped to support/ops workflows rather than general-purpose NL command execution. Missing for 10: independent/hands-on evidence of NL command accuracy, and clarity on whether all product actions (not just configuration) can be triggered via natural language versus UI/API.",
    "evidenceIds": [
      "parahelp-docs-1",
      "parahelp-docs-6",
      "parahelp-docs-7",
      "parahelp-docs-19"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-official-cli",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Parahelp is a customer-support/internal-agent SaaS platform, not a developer tool whose category typically ships a CLI; its interfaces are documented as API, dashboard, and ticketing integrations, not a command-line tool. No evidence suggests a CLI exists or is relevant to its workflow.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-public-api",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Docs explicitly describe driving both the Internal Agent and Customer Agent via API (ticket-free API intake, automation triggers, run-without-approval via API), and a runtime probe confirms a public API reference is reachable keylessly at app.parahelp.com/api/docs. However, standard OpenAPI/swagger spec endpoints all 404, so there's no machine-readable spec confirmed, only prose docs and a reference UI. Missing for 10: a confirmed OpenAPI/swagger spec, and independent hands-on confirmation of actually calling the API successfully.",
    "evidenceIds": [
      "parahelp-docs-13",
      "parahelp-docs-14",
      "parahelp-docs-20",
      "parahelp-docs-22",
      "parahelp-probe-rt-2",
      "parahelp-probe-2"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-scoped-keys",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Parahelp documents role-based access control, approval guardrails per-tool, and API access for triggering agents, but no evidence describes issuing scoped or least-privilege API credentials/keys specifically for an agent's tool access. Missing for 10: any documentation of API key scoping, credential minting, or permission-limited tokens issued to an agent.",
    "evidenceIds": [
      "parahelp-docs-10",
      "parahelp-docs-3",
      "parahelp-docs-8",
      "parahelp-docs-13"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-sdks",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Parahelp documents a customer-facing API and webhook/automation triggers, but no evidence anywhere mentions official client SDKs (Python, JS, etc.); the OpenAPI/swagger probe also 404s across all candidate paths, suggesting no machine-readable spec for SDK generation either.",
    "evidenceIds": [
      "parahelp-docs-13",
      "parahelp-docs-14",
      "parahelp-probe-2",
      "parahelp-probe-rt-2"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "agentic-webhooks",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence pack item mentions webhooks or event subscriptions; Parahelp's automation/trigger docs reference schedules and GitHub PR merges but nothing about outbound webhook subscriptions for third-party consumption. missing for 10: any mention of webhook subscription/registration, event types, or push-notification mechanism.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "api-interactive-docs",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "There is an API reference surface (app.parahelp.com/api/docs) and docs are agent-legible, but no evidence shows an interactive reference with runnable examples (e.g. try-it console, code samples execution); explicit openapi.json probes all 404'd, suggesting no standard interactive spec is exposed.",
    "evidenceIds": [
      "parahelp-probe-2",
      "parahelp-probe-rt-2"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "api-machine-spec",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Parahelp has a customer-facing API and a public API reference page (app.parahelp.com/api/docs), but explicit probes for a machine-readable OpenAPI/Swagger spec at all standard paths (openapi.json, swagger.json, etc.) returned 404, and no evidence shows a downloadable spec file.",
    "evidenceIds": [
      "parahelp-probe-2",
      "parahelp-probe-rt-2",
      "parahelp-docs-13"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "api-sandbox",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Docs describe simulations that replay historical tickets or mock scenarios against configurations, which functions as a sandbox for testing without touching production, and separately note that 'production tests verify your tool connections' implying a distinction between simulated and production environments. However, there is no explicit mention of a dedicated sandbox environment, isolated test data store, or guarantee that simulations cannot write to production systems/tools. missing for 10: explicit documentation of an isolated sandbox environment separate from production data/tools, confirmation that simulated runs cannot trigger real side-effects, and independent/hands-on verification of this isolation.",
    "evidenceIds": [
      "parahelp-docs-2",
      "parahelp-docs-4"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "api-versioning-policy",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "There is evidence of an API existing (customer requests via API, run automations via API) and public API docs, but no mention of API versioning or a documented deprecation policy anywhere in the evidence pack. Missing for 10: versioning scheme, deprecation policy, changelog entries about breaking changes or API version sunset.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "automation-bulk-operations",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Evidence shows batch-style processing (analyzing 500-1,000 historical tickets, simulations replaying many tickets, scheduled/triggered automations) which imply working across many items, but there's no explicit documentation of a user-facing bulk-action feature (e.g., mass-editing or resolving many tickets/items in one command).\nmissing for 10: explicit bulk-operation UI/API (e.g., batch resolve/tag/update across selected items), evidence of throughput/rate limits for bulk actions, independent confirmation of bulk workflows in production.",
    "evidenceIds": [
      "parahelp-docs-19",
      "parahelp-docs-2",
      "parahelp-docs-5",
      "parahelp-docs-17"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "automation-rules-engine",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Docs explicitly describe automations as Internal Agent tasks that run on schedules or triggers (e.g., GitHub PR merge), plus guardrail rules that trigger actions on events (Slack approval above $100, transfer enterprise tickets), matching the rule-based, event-triggered automation story. Missing for 10: independent/third-party corroboration of automation reliability and a broader list of supported trigger event types beyond the examples given.",
    "evidenceIds": [
      "parahelp-docs-5",
      "parahelp-docs-6",
      "parahelp-docs-7",
      "parahelp-docs-17",
      "parahelp-docs-3"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "automation-scheduled-jobs",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Parahelp's docs explicitly describe automations as Internal Agent tasks that run on a schedule (e.g. every weekday at 9am) or on triggers, which directly matches recurring scheduled workflows, and these can be run via API for custom automation depth. Missing for 10: no independent/hands-on corroboration of scheduling reliability, and no detail on schedule configuration flexibility (cron-like options, timezone handling, etc.).",
    "evidenceIds": [
      "parahelp-docs-5",
      "parahelp-docs-17",
      "parahelp-docs-14"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "automation-versioned-workflows",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Parahelp supports release publishing with one-click revert and knowledge version history with diffs (parahelp-docs-4, parahelp-docs-9), directly covering versioning and rollback of configurations/automations. However, there's no explicit review/approval workflow for the versioning itself (approvals exist for actions, not for config changes review), and no evidence of branching, staged review, or audit trail beyond 'who published what and when'. Missing for 10: dedicated review/approval step before publishing a release, granular version comparison across automations (not just knowledge), and independent/hands-on confirmation of rollback working in practice.",
    "evidenceIds": [
      "parahelp-docs-4",
      "parahelp-docs-9",
      "parahelp-docs-5"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "brand-voice-tone",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence covers configuration via natural language, knowledge base versioning, guardrails, and tool integrations, but nothing addresses tone/brand-voice control or consistency across topics and languages. Missing for 10: any documentation or claim about persona/tone settings, brand voice configuration, or multilingual consistency testing.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "configurable-escalation-rules",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs show configurable guardrails that trigger handoff/escalation by topic (e.g., transfer enterprise-account tickets) and by threshold (Slack approval on refunds >$100), plus auto-approve vs manual-approval controls and release tracking for changes — evidence of rule-based handoff configuration and enforcement. However, there is no explicit documentation of sentiment-based triggers, customer-tier-specific rules beyond the enterprise example, or handling of explicit 'transfer me to a human' requests, and no independent/hands-on verification that these rules are reliably obeyed in production. Missing for 10: sentiment-based escalation rules, explicit-request handoff configuration, customer-tier granularity beyond one example, and third-party evidence of reliability.",
    "evidenceIds": [
      "parahelp-docs-6",
      "parahelp-docs-7",
      "parahelp-docs-3",
      "parahelp-docs-4",
      "parahelp-docs-8"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "content-gap-detection",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs show explicit gap-surfacing automations that watch tools/support queues (parahelp-docs-17) plus knowledge version history with diffs and publisher attribution (parahelp-docs-9), which support detecting stale or conflicting knowledge. However, there's no evidence of conflict detection between multiple knowledge sources, no dashboard/reporting UI shown, and no independent/hands-on corroboration of gap surfacing actually catching missed or fumbled questions. missing for 10: evidence of explicit conflicting-content detection across knowledge sources, a surfaced-gaps UI/report example, and independent verification that this reduces missed/fumbled tickets.",
    "evidenceIds": [
      "parahelp-docs-17",
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "context-rich-handoff",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Docs show guardrail rules that transfer tickets 'with a summary note' to human teams (parahelp-docs-7) and ticketing-system integrations where the agent works as a team member (parahelp-docs-15), implying handoff context is passed, but there's no explicit description of a full escalation package (full transcript + generated summary + structured collected details) reaching the human agent. missing for 10: dedicated documentation of the escalation handoff artifact itself (transcript, structured details form, summary format), evidence that customers never have to repeat themselves, and any hands-on/independent confirmation of this workflow.",
    "evidenceIds": [
      "parahelp-docs-7",
      "parahelp-docs-15",
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "end-to-end-resolution",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Docs describe the Customer Agent handling tickets end-to-end with tools/guardrails and note integrations 'resolve tickets across email, live chat, and Slack,' implying resolution capability, but there is no quantified resolution rate, case study, or benchmark distinguishing true resolutions from deflections/bounces. Missing for 10: measured resolution-rate metrics, customer case studies with numbers, and independent verification of end-to-end resolution vs. escalation rates.",
    "evidenceIds": [
      "parahelp-docs-15",
      "parahelp-docs-16",
      "parahelp-docs-13"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "grounded-cited-answers",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "Evidence shows Parahelp maintains a versioned knowledge base (docs-9) that the Customer Agent draws on, but there is no evidence that individual answers to users cite or surface the specific article/source they were grounded in. missing for 10: any documentation of per-answer source citation or attribution UI, evidence of end-user-facing 'grounded in X article' display.",
    "evidenceIds": [
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "hallucination-guardrails",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Parahelp's guardrails docs focus on approval gates for actions (e.g., refund thresholds, enterprise transfers) and versioned knowledge bases, which constrain what the agent can execute and what knowledge it draws from — but there is no explicit documentation of a 'safe decline' behavior when a question falls outside known knowledge/policy. Missing for 10: explicit fallback/decline behavior for off-knowledge queries, evidence of hallucination prevention or confidence thresholds, and independent confirmation the agent refuses rather than guesses.",
    "evidenceIds": [
      "parahelp-docs-6",
      "parahelp-docs-7",
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "helpdesk-platform-integrations",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Docs explicitly confirm the Customer Agent works inside Zendesk, Intercom, Front, Plain, or Pylon as a 'team member,' with a standalone API mode for ticket-free workflows, and shared tool/knowledge connections across both agents. Ticket sync/context bidirectionality is implied by the 'team member' integration model and API mode, but no explicit two-way sync mechanics or Salesforce-specific mention are documented. Missing for 10: explicit mention of Salesforce support, and detailed description of two-way ticket/context syncing mechanics beyond high-level integration claims.",
    "evidenceIds": [
      "parahelp-docs-15",
      "parahelp-docs-16",
      "parahelp-docs-13",
      "parahelp-docs-18",
      "parahelp-docs-20"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "knowledge-auto-sync",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Parahelp documents a generic automations feature that can run on a schedule or trigger (parahelp-docs-5) and shows live knowledge with version history (parahelp-docs-9), but there is no explicit documentation that knowledge sources are automatically re-synced on a schedule or on change rather than manually updated. missing for 10: explicit knowledge-source auto-sync mechanism, evidence of scheduled/triggered re-ingestion tied specifically to knowledge base content, independent confirmation of this behavior.",
    "evidenceIds": [
      "parahelp-docs-5",
      "parahelp-docs-9",
      "parahelp-docs-19"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "knowledge-source-ingestion",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Parahelp's Internal Agent explicitly analyzes past resolved tickets and existing context to build the Customer Agent's knowledge/configuration automatically, and there's a live knowledge base with version history, but evidence only covers tickets and general 'existing context' rather than explicit ingestion of help center articles, docs, or internal wikis as named sources. missing for 10: explicit documentation of connectors/ingestion for help-center articles, internal wikis, and docs repositories as knowledge sources; independent evidence of successful re-ingestion at scale beyond ticket history.",
    "evidenceIds": [
      "parahelp-docs-19",
      "parahelp-docs-9",
      "parahelp-docs-1"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "live-api-actions",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Docs show the agent executes real API actions (Stripe or internal endpoints) like refunds, with per-tool guardrails (e.g., Slack approval above $100), tool-level auto-approve settings, and role-based access control across the product. This covers real-action execution with approval/authorization gating per action, though evidence doesn't detail fine-grained per-action auth scoping (e.g., API-key/permission scoping distinct from approval workflows) beyond general RBAC mentions. Missing for 10: explicit documentation of per-action credential/auth scoping mechanics (vs. approval gates), and independent/hands-on verification of action execution in production.",
    "evidenceIds": [
      "parahelp-docs-6",
      "parahelp-docs-3",
      "parahelp-docs-8",
      "parahelp-docs-12",
      "parahelp-docs-10",
      "parahelp-docs-14"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "multi-turn-troubleshooting",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence pack covers configuration, guardrails, approvals, integrations, and knowledge versioning, but nothing describes the Customer Agent's actual conversational behavior—no mention of it asking clarifying questions or performing multi-step troubleshooting dialogue rather than a single scripted reply.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "multilingual-answers",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence in the pack mentions multilingual support, language detection, or translation capabilities for the Customer Agent; documentation focuses on ticketing integrations, tool connections, and knowledge management without addressing cross-language customer support.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "omnichannel-coverage",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Parahelp's Customer Agent operates as a team member inside major ticketing/chat platforms (Intercom, Zendesk, Front, Plain, Pylon), and its Pylon integration explicitly covers email, live chat, and Slack tickets, showing some multi-channel reach. However, there is no evidence of native WhatsApp or social-media channel support, so the 'channels customers actually use' claim is only partially substantiated. Missing for 10: explicit WhatsApp support, explicit social media channel support, and independent confirmation beyond vendor docs.",
    "evidenceIds": [
      "parahelp-docs-15",
      "parahelp-docs-16"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "ongoing-qa-review",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Parahelp offers simulation-based testing before release (docs-2), release tracking/revert (docs-4), and automations that 'watch your tools and support queue and surface gaps' (docs-17), which map loosely to a QA/review loop feeding fixes back. However there is no evidence of a structured scored-sample QA program, explicit failure flagging on live conversations, or a dedicated review dashboard for support-ops to audit agent conversations. Missing for 10: scored sampling of live conversations, explicit failure-flag workflow, dedicated QA review UI/report distinct from gap-detection automations.",
    "evidenceIds": [
      "parahelp-docs-2",
      "parahelp-docs-4",
      "parahelp-docs-17",
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "openness-api-parity",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Parahelp documents API access for core operational tasks—receiving customer requests via API (docs-13, docs-20) and running the Internal Agent's automations via API (docs-14, docs-22)—and a live API reference exists (parahelp-probe-rt-2). However, no evidence shows that configuration authoring, simulation/testing, release management, or knowledge-diff review (all UI-centric per docs-1/2/4/9) can be done via API, and no discoverable OpenAPI spec was found (parahelp-probe-2 shows 404s across standard paths). Missing for 10: API parity for configuration/build workflows, testing/simulation via API, release/rollback via API, formal OpenAPI spec confirming full surface coverage.",
    "evidenceIds": [
      "parahelp-docs-13",
      "parahelp-docs-14",
      "parahelp-docs-20",
      "parahelp-docs-22",
      "parahelp-probe-2",
      "parahelp-probe-rt-2"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "openness-full-export",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence of any data export functionality, open-format export, or account-closure/data-portability capability in the docs or probes; the material covers agent configuration, tools, and integrations but nothing about exporting user/knowledge data or leaving the platform with data intact.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "openness-open-license",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Parahelp is a closed commercial SaaS platform for customer/internal support agents, not an open-source project; there is no evidence of any open-license source availability, and the product category (proprietary SaaS) doesn't typically ship source under an open license.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "openness-self-host",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Parahelp is a hosted SaaS customer-support/internal-agent platform with no evidence of any self-hostable deployment option; self-hosting is not a fair axis for this kind of managed product, and nothing in the evidence suggests it ships an on-prem or open-source core.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "personalized-account-answers",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Docs show the Customer Agent can connect tools to any API (Stripe, internal endpoints) and both agents share connections, which enables pulling live account/order data rather than just static knowledge articles, but no concrete example or case study shows it actually surfacing plan/order-status/account-history data in an answer. missing for 10: a documented example or customer case showing live account/order lookups in a real resolution, and specifics on how live data is merged into agent responses vs. knowledge base content.",
    "evidenceIds": [
      "parahelp-docs-12",
      "parahelp-docs-18",
      "parahelp-docs-21",
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "pre-launch-simulation",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Docs explicitly describe simulations that replay historical tickets or mock scenarios against configurations, plus production tests to verify tool connections, directly matching the story of pre-launch testing before facing real customers. Missing for 10: no independent/hands-on corroboration of simulation accuracy or third-party validation of the testing workflow.",
    "evidenceIds": [
      "parahelp-docs-2",
      "parahelp-docs-19"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "privacy-data-residency",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence of data residency or region selection options; only related mention is disabling personal data storage entirely, not choosing storage region. Missing for 10: any mention of region/residency controls, data center location choices, or geographic compliance options.",
    "evidenceIds": [
      "parahelp-docs-11"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "privacy-no-training",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Parahelp only offers a data-retention control (admins can disable personal data storage) but no explicit statement about excluding customer data from AI model training. missing for 10: explicit no-training-on-customer-data guarantee, opt-out toggle for model training, third-party/independent confirmation of training data handling.",
    "evidenceIds": [
      "parahelp-docs-11"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "privacy-retention-controls",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Parahelp's security docs state that workspace admins can disable all personal data storage, giving explicit control over retention, but there is no documented deletion mechanism (e.g., data export/erasure API, retention period settings) beyond this single toggle. missing for 10: explicit data deletion/export controls, granular retention period settings, independent confirmation of the disable-storage feature.",
    "evidenceIds": [
      "parahelp-docs-11",
      "parahelp-docs-10"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "privacy-telemetry-optout",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "Evidence shows a workspace-level admin control to disable personal data storage, but nothing about opting out of telemetry/usage tracking for the product itself as an AI-native/dev-tool user; missing for 10: any telemetry opt-out setting, docs on analytics collection, or user-level tracking controls.",
    "evidenceIds": [
      "parahelp-docs-11"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "procedure-sop-builder",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Parahelp lets ops encode conditional rules (e.g. 'if enterprise account, transfer with summary', 'if refund >$100, require Slack approval') and lets the Internal Agent configure task flows from natural-language descriptions of policies/tools, which supports rule-based branching for known issue types. However, there's no evidence of an explicit step-by-step SOP builder or visual decision-tree/flowchart for deterministic multi-step branching beyond simple guardrail conditions. Missing for 10: explicit multi-step SOP/workflow authoring UI, evidence of complex nested/deterministic branching logic beyond single-condition guardrails, and independent validation of branching behavior in production.",
    "evidenceIds": [
      "parahelp-docs-1",
      "parahelp-docs-6",
      "parahelp-docs-7",
      "parahelp-docs-9"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "resolution-analytics",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence of any dashboard or reporting feature covering resolution rate, CSAT, handoff rate, or cost per resolution; the pack only covers agent configuration, guardrails, integrations, and API access, none of which mention analytics/metrics reporting for support leaders.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "restricted-topic-controls",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Parahelp's guardrails let ops leads set conditional rules such as escalating enterprise-account tickets to a human team or requiring approval before certain tools run, and every agent action requires manual approval by default unless explicitly auto-approved — which supports building 'never freelance, hand off to human' rules for chosen topics. However, there's no dedicated 'mark topic as human-only' feature or explicit examples for legal threats/cancellations/security specifically, only generic escalation/approval guardrail examples. Missing for 10: explicit topic-blocklist UI, named legal/cancellation/security guardrail templates, and evidence the escalation is airtight (agent truly cannot act) rather than just gated by approval.",
    "evidenceIds": [
      "parahelp-docs-6",
      "parahelp-docs-7",
      "parahelp-docs-3",
      "parahelp-docs-8"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "supervised-draft-mode",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Docs explicitly state that by default every action requires manual approval, with per-tool/task auto-approve override, and cite concrete examples like Slack approval before refund actions run — directly matching a supervised human-approval-before-customer-impact mode. Releases page tracks and reverts config changes, adding auditability to this guardrail workflow. Missing for 10: independent/hands-on verification of the approval UI in practice and confirmation that 'draft for approval' specifically (vs. blocking tool actions) is the default behavior for all reply types.",
    "evidenceIds": [
      "parahelp-docs-3",
      "parahelp-docs-6",
      "parahelp-docs-8",
      "parahelp-docs-4"
    ]
  },
  {
    "productId": "parahelp",
    "storyId": "topic-trend-insights",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Parahelp's docs describe automations that 'watch your tools and your support queue and surface gaps on their own' (parahelp-docs-17), but this is about detecting configuration/knowledge gaps, not clustering conversations by topic or flagging emerging issues before ticket volume spikes. No evidence pack item describes topic clustering, trend detection, or predictive volume analytics.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "transparent-per-resolution-pricing",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence anywhere in the pack of published pricing, per-resolution billing, caps, or outcome-based pricing controls — evidence is entirely about product functionality (agents, guardrails, integrations), not commercial terms.",
    "evidenceIds": []
  },
  {
    "productId": "parahelp",
    "storyId": "voice-phone-support",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Parahelp's evidence describes a text-based Customer Agent that works via ticketing systems and API, with no mention of voice/telephony channels, speech-to-text, or text-to-speech capabilities. No citation addresses phone call handling.",
    "evidenceIds": []
  },
  {
    "productId": "pylon",
    "storyId": "agentic-agent-docs",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Probes confirm a live llms.txt index at docs.usepylon.com/llms.txt and working per-page .md mirrors (e.g. pylon-mcp.md), so an agent can be pointed at agent-oriented docs and consume them directly. missing for 10: independent third-party confirmation beyond the vendor's own probe, and some per-page md endpoints (e.g. pylon-docs.md) 404 showing coverage is inconsistent.",
    "evidenceIds": [
      "pylon-probe-1",
      "pylon-probe-rt-1",
      "pylon-probe-2"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-ai-insights",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon documents AI summarization of customer conversations synced to CRM accounts (Salesforce), AI-powered knowledge base features, and analytics dashboards, plus AI Support Agents that surface context and take action — all suggesting embedded AI insight generation. However, there's no dedicated 'insights' feature that proactively surfaces suggestions from the user's own data (e.g., trend detection, anomaly alerts) beyond conversation summarization and agent responses. Missing for 10: a first-party 'insights' or 'suggestions' dashboard feature explicitly generating recommendations from account/ticket data, and independent/hands-on validation of insight quality.",
    "evidenceIds": [
      "pylon-docs-33",
      "pylon-docs-29",
      "pylon-docs-23",
      "pylon-docs-13",
      "pylon-docs-19"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-autonomous-automation",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon documents triggers/automations to execute business processes, webhooks, API access, and AI agents that autonomously handle issues assigned to them with monitoring via issue logs, supporting background autonomous automation. However, evidence is mostly first-party docs with no independent/hands-on confirmation of unattended reliability or long-running autonomy at scale. Missing for 10: independent/hands-on verification of autonomous background execution, detail on failure handling/retries in triggers, and third-party confirmation of agent autonomy in production.",
    "evidenceIds": [
      "pylon-docs-30",
      "pylon-docs-13",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-32",
      "pylon-docs-31"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-builtin-assistant",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Pylon ships built-in AI Support Agents that users can build, assign to issues, give runbooks/skills, and monitor outcomes—clearly a built-in assistant users delegate tasks to (assign issues, define workflows, review AI-taken actions). Missing for 10: independent/third-party hands-on validation of assistant task quality and no user-facing UI screenshots/demos beyond docs.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-14",
      "pylon-docs-15",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-20"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-headless",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon exposes a REST API, webhooks, and triggers that allow programmatic, non-UI interaction with the platform, which supports automation-style usage outside the web UI (pylon-docs-31, pylon-docs-32, pylon-docs-30, pylon-probe-rt-2). However there is no evidence of a CLI, headless mode, or explicit CI/CD integration story — the product is a SaaS support platform accessed via API/OAuth, not something designed to be invoked in a CI pipeline. Missing for 10: explicit CLI/headless execution mode, CI/CD pipeline examples or docs, evidence of automated/scripted runs outside a live SaaS API context.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-docs-30",
      "pylon-probe-rt-2"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-mcp-client",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon's agent 'connectors' documentation states that if a third-party system provides an MCP server, agents can use it directly (falling back to a Custom API tool only when none exists), indicating Pylon can consume external MCP servers as tool sources for its AI agents. However, this is a single thin doc line with no setup walkthrough, supported-server list, or hands-on/runtime confirmation of actually connecting to a third-party MCP server. missing for 10: a dedicated 'add MCP server' configuration flow/UI, documentation of supported transports/auth for third-party servers, and independent or runtime evidence of a working third-party MCP connection.",
    "evidenceIds": [
      "pylon-docs-21"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-mcp-server",
    "verdict": "full",
    "quality": 9,
    "confidence": "high",
    "rationale": "Pylon is a customer-support platform (not itself an agent), and it ships a first-party MCP server documented at docs.usepylon.com/pylon-docs/integrations/pylon-mcp with OAuth 2.0 authentication, hosted at mcp.usepylon.com. Runtime probes independently confirm the server is live (proper 401/WWW-Authenticate challenge with oauth-protected-resource metadata), corroborating the docs beyond vendor claims. Missing for 10: no independent third-party hands-on report of an agent successfully completing a task via this MCP server.",
    "evidenceIds": [
      "pylon-docs-34",
      "pylon-probe-4",
      "pylon-probe-rt-1",
      "pylon-probe-rt-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-nl-commands",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon exposes a live first-party MCP server (confirmed by runtime probes) that lets AI tools authenticate via OAuth and read/update issues, accounts, and contacts, effectively letting an AI-native user drive Pylon via natural-language-mediated agent actions; runbooks and skills also let admins encode natural-language instructions for agent behavior. However, this is achieved through external AI-tool/agent integration rather than a native in-app NL command interface for end users. Missing for 10: evidence of a built-in chat/command-bar UI where a human directly issues natural-language commands to operate Pylon itself, and independent confirmation of real-world NL-driven task completion.",
    "evidenceIds": [
      "pylon-docs-34",
      "pylon-docs-15",
      "pylon-docs-20",
      "pylon-probe-rt-1",
      "pylon-probe-rt-3",
      "pylon-docs-16"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-official-cli",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence of an official Pylon CLI anywhere in docs, changelog, or probes; Pylon offers REST API, webhooks, and an MCP server but no CLI tool is mentioned. missing for 10: any mention of a CLI, CLI installation instructions, or CLI command reference.",
    "evidenceIds": []
  },
  {
    "productId": "pylon",
    "storyId": "agentic-public-api",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Pylon documents a public REST API (docs-31) and webhooks (docs-32), and runtime probes confirm the API is live and properly auth-gated (probe-rt-2). Docs also describe programmatic/agentic access patterns (Custom API tools, OAuth-based AI tool access via docs-34) reinforcing that the API is designed for agentic/programmatic drivers. missing for 10: no discoverable OpenAPI/swagger spec (probe-3 shows 404s across candidate paths) and no independent third-party corroboration of API robustness beyond Pylon's own docs.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-docs-34",
      "pylon-probe-rt-2",
      "pylon-probe-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-scoped-keys",
    "verdict": "partial",
    "quality": 4,
    "confidence": "medium",
    "rationale": "Pylon's MCP integration uses OAuth 2.0 (via AuthKit) and the REST API requires Bearer tokens, so credentials are authenticated and gated (pylon-docs-34, pylon-probe-rt-1/2/3), and access is limited to 'the data you can see in Pylon,' giving a coarse form of least-privilege tied to the issuing user's own permissions. However there is no evidence of a dedicated mechanism to mint distinct, granularly-scoped API keys or OAuth scopes specifically for an agent (e.g., read-only vs write, per-object scoping) separate from a full user account's access. Missing for 10: documented ability to configure fine-grained scopes/permissions per API credential, evidence of agent-specific token restriction UI, and independent confirmation of least-privilege enforcement beyond basic auth gating.",
    "evidenceIds": [
      "pylon-docs-34",
      "pylon-probe-rt-1",
      "pylon-probe-rt-2",
      "pylon-probe-rt-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-sdks",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon exposes a documented REST API, webhooks, and an OAuth-secured official MCP server (confirmed live via runtime probes), which together let AI-native developers build integrations and agents against Pylon data. However, there is no evidence of traditional language-specific client SDKs (e.g., Python/Node packages) or an OpenAPI spec (probe found 404s for openapi/swagger endpoints), so 'SDK' support is really API+MCP+webhooks rather than packaged SDKs. Missing for 10: dedicated official SDK packages, publicly discoverable OpenAPI/swagger spec, independent developer corroboration of SDK usage.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-docs-34",
      "pylon-probe-rt-2",
      "pylon-probe-rt-3",
      "pylon-probe-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "agentic-webhooks",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Pylon documents a dedicated webhooks feature for receiving events happening in Pylon, directly matching the story. missing for 10: no independent/hands-on corroboration of webhook delivery, no detail on event types/payload schema, or subscription management UI/API specifics.",
    "evidenceIds": [
      "pylon-docs-32",
      "pylon-docs-16"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "api-interactive-docs",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Pylon documents a REST API and webhooks (pylon-docs-31, pylon-docs-32) but there is no evidence of an interactive API reference with runnable/try-it examples — probes explicitly found no OpenAPI/Swagger spec at any candidate path (pylon-probe-3), only static markdown docs.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-probe-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "api-machine-spec",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Pylon has a REST API and docs, but explicit probes for OpenAPI/Swagger spec files at standard paths all returned 404, and no evidence anywhere shows a downloadable machine-readable spec (OpenAPI, JSON Schema, etc.) for its API.",
    "evidenceIds": [
      "pylon-probe-3",
      "pylon-docs-31",
      "pylon-probe-rt-2"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "api-sandbox",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Pylon's Support Agent docs mention a 'test' feature to simulate interactions with the AI without deploying it live (pylon-docs-17), which suggests some sandboxed testing before touching real customer issues, but there's no documented sandbox environment, staging API, or test dataset separate from production data. Missing for 10: explicit sandbox/staging environment, test API keys or test accounts isolated from production, and confirmation that simulated interactions don't touch real customer/production data.",
    "evidenceIds": [
      "pylon-docs-17",
      "pylon-docs-18"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "api-versioning-policy",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows Pylon has a REST API, webhooks, and MCP server, but no documentation of API versioning scheme or a deprecation policy anywhere in docs/changelog; openapi spec probes 404 and no version headers or changelog entries reference API deprecations. Missing for 10: versioning scheme (e.g., v1/v2 paths), documented deprecation timeline/policy, changelog entries about breaking changes.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-probe-3",
      "pylon-probe-rt-2"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "automation-bulk-operations",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon's API (pylon-docs-31), MCP server (pylon-docs-34, pylon-probe-rt-1/3), and triggers/automations (pylon-docs-30) give AI agents and scripts programmatic read/update access to issues, accounts, and contacts, which could be used to script bulk changes, but no evidence describes a dedicated bulk-edit/bulk-action UI or batch-endpoint for acting on many items in one call. Missing for 10: an explicit bulk-update/bulk-action API or UI feature, and any documentation or demo of multi-item batch operations rather than single-record CRUD via API/MCP.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-34",
      "pylon-docs-30",
      "pylon-probe-rt-1",
      "pylon-probe-rt-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "automation-rules-engine",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Pylon has a dedicated Triggers platform docs page describing setting up complex workflows/automations tied to business processes, plus webhooks for events and an API to programmatically act on data, enabling event-driven automated actions. This directly matches rule-based automation triggering actions on events, with runtime probes confirming the API/webhook infrastructure is live. Missing for 10: detailed trigger-rule syntax/examples, independent hands-on validation of complex trigger logic, and explicit AI-agent-specific trigger configuration beyond general workflow automation.",
    "evidenceIds": [
      "pylon-docs-30",
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-probe-rt-2"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "automation-scheduled-jobs",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Pylon's docs mention 'triggers' for building automations and workflows (pylon-docs-30) plus API/webhooks for programmatic action, but nothing in the evidence pack specifies time-based or recurring scheduling (e.g., cron-like jobs, scheduled runs) as opposed to event-triggered automations. No mention of a scheduler, recurring job configuration, or interval-based execution.",
    "evidenceIds": [
      "pylon-docs-30",
      "pylon-docs-31",
      "pylon-docs-32"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "automation-versioned-workflows",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Pylon has automations/triggers, agent runbooks, and issue logs for monitoring outcomes, but there is no evidence of version history, diffing, review workflows, or rollback capability for automations/triggers/skills/runbooks. Missing for 10: version history for triggers/runbooks, change review/approval process, rollback to a prior automation version.",
    "evidenceIds": [
      "pylon-docs-30",
      "pylon-docs-15",
      "pylon-docs-19"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "brand-voice-tone",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Pylon's AI Agent can be configured with natural-language runbooks and reusable 'skills' that dictate how it handles specific scenarios, which could be used to encode tone/brand-voice guidance, but the docs never explicitly mention persona/tone control or consistency across languages. missing for 10: explicit tone/brand-voice configuration settings, evidence of multilingual consistency, and any customer/independent proof the agent's voice stays consistent across topics.",
    "evidenceIds": [
      "pylon-docs-15",
      "pylon-docs-20",
      "pylon-docs-13"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "configurable-escalation-rules",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Pylon's Support Agent framework lets admins write natural-language runbooks for 'specific scenarios' and assign only certain issues to the agent, plus monitor outcomes via the issue log — this could be used to script handoff conditions, but there is no explicit documentation of built-in triggers for topic, sentiment, or customer-tier-based escalation, nor any evidence about reliability/guardrails ensuring the agent 'reliably obeys' handoff rules. Missing for 10: dedicated sentiment/tier/topic escalation configuration UI, explicit-request handoff trigger documentation, and evidence (docs or hands-on) of enforcement reliability.",
    "evidenceIds": [
      "pylon-docs-15",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-20",
      "pylon-docs-13"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "content-gap-detection",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Pylon documents training-data ingestion, a knowledge base, agent simulation/testing, and issue logs for reviewing agent outcomes, but nothing in the evidence describes a feature that proactively surfaces knowledge gaps or flags conflicting content driving agent misses. This is a fair axis for a support-agent platform (comparable products ship content-quality/coverage analytics), so absence of evidence yields 'none' rather than 'na'.",
    "evidenceIds": [
      "pylon-docs-19",
      "pylon-docs-17",
      "pylon-docs-22",
      "pylon-docs-23"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "context-rich-handoff",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Pylon's AI agent operates within the same ticket/issue thread that a human later takes over, and the issue log records the agent's steps and outcomes (pylon-docs-19), while the agent explicitly gathers internal context for the team (pylon-docs-13) and elsewhere Pylon auto-summarizes conversations for CRM sync (pylon-docs-33), suggesting summarization capability exists. However, there is no explicit documentation of an escalation-specific handoff package (conversation + AI summary + structured collected details) being automatically attached to a ticket when an agent escalates to a human. missing for 10: explicit escalation-handoff feature docs, evidence the summary/details are surfaced to the human agent at hand-off time, confirmation customer doesn't need to repeat themselves.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-19",
      "pylon-docs-33",
      "pylon-docs-18"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "end-to-end-resolution",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Pylon's own docs describe the Support Agent primarily in terms of 'deflecting' questions and taking actions, with an issue log to inspect outcomes/steps, but there is no first-party or independent metric quantifying full end-to-end resolutions distinct from deflections or bounces. missing for 10: quantified resolution-rate reporting, evidence distinguishing true resolutions from deflections, independent/customer validation of resolution outcomes.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-29"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "grounded-cited-answers",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon's Support Agents are documented to answer/deflect using a dedicated Knowledge Base and connected external knowledge sources (pylon-docs-13, pylon-docs-22, pylon-docs-23), which supports grounding in the user's own content, but there is no evidence that responses explicitly cite or display which specific article/source was used to generate an answer. missing for 10: explicit citation/source-attribution UI in agent answers, evidence of per-answer source linking, independent confirmation of grounding accuracy.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-22",
      "pylon-docs-23",
      "pylon-docs-20"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "hallucination-guardrails",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Pylon's Support Agent is grounded in a curated knowledge base and runbooks/skills that define scenario-specific instructions, which implicitly limits it to trained content rather than free invention, and 'deflect using content' suggests answers are content-bound. However, there is no explicit documentation of a safe-decline behavior for off-knowledge questions, confidence thresholds, or anti-hallucination guardrails around prices/policies. Missing for 10: explicit safe-decline/escalation mechanism for unknown questions, documentation of hallucination prevention, and evidence of guardrails specifically around prices/policies/promises.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-15",
      "pylon-docs-20",
      "pylon-docs-23"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "helpdesk-platform-integrations",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Evidence shows a one-way sync of AI-summarized conversations into Salesforce Accounts and general 'omnichannel' support language, but there is no evidence of the AI agent embedding inside Zendesk or Intercom, nor of two-way ticket/context syncing with any named helpdesk — Pylon's docs otherwise position it as the standalone system of record. Missing for 10: Zendesk/Intercom integration evidence, bidirectional ticket sync (not just one-way conversation summaries), and confirmation the agent itself operates 'inside' another helpdesk's UI/workflow.",
    "evidenceIds": [
      "pylon-docs-33",
      "pylon-docs-24",
      "pylon-docs-16"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "knowledge-auto-sync",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Pylon's knowledge base and training-data docs mention connecting external knowledge sources (pylon-docs-22, pylon-docs-23), and product-data sync can be disabled mid-run (pylon-docs-8), implying some sync mechanism exists, but there is no documentation of a re-sync schedule, change-detection triggers, or automatic re-indexing cadence for knowledge sources. missing for 10: explicit scheduling/frequency controls for knowledge source re-sync, change-detection or webhook-triggered re-sync of external knowledge, and any evidence contradicting or confirming this works hands-on.",
    "evidenceIds": [
      "pylon-docs-22",
      "pylon-docs-23",
      "pylon-docs-8"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "knowledge-source-ingestion",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Pylon documents a 'training-data' page for connecting external knowledge sources and a Knowledge Base feature for hosting support content and runbooks, suggesting some support for knowledge grounding, but there is no concrete evidence of automated ingestion of help center content, past tickets, or internal wikis without manual re-authoring. Missing for 10: explicit documentation of connectors for help center/docs platforms, wiki integrations (Confluence/Notion), automatic ticket-history ingestion, and any hands-on evidence confirming these sources sync without manual re-entry.",
    "evidenceIds": [
      "pylon-docs-22",
      "pylon-docs-23",
      "pylon-docs-13"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "live-api-actions",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon documents that AI agents can 'take action' and that admins can define Custom API tools so agents can hit internal or third-party API endpoints (e.g., a company's own refund/order APIs), and that Pylon's own API/MCP integration uses OAuth authentication for issues/accounts/contacts. However, there's no concrete example or documentation of granular, action-level scoped auth (e.g., distinct permission scopes for 'refund' vs 'subscription update') — the OAuth scoping described applies to Pylon's own data objects, not to arbitrary external business APIs the agent might call. Missing for 10: explicit examples of refunds/order/subscription actions being executed, and documentation of per-action auth scoping for custom API tools beyond general OAuth.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-16",
      "pylon-docs-21",
      "pylon-docs-34",
      "pylon-probe-rt-2",
      "pylon-probe-rt-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "multi-turn-troubleshooting",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon's Support Agent docs describe multi-step behavior—runbooks with natural-language instructions for specific scenarios, skills as reusable workflow instructions, gathering internal context, and taking action—plus an issue log showing 'outcomes and steps taken by the AI' (pylon-docs-13, pylon-docs-15, pylon-docs-20, pylon-docs-19). However, there is no explicit documentation that the agent proactively asks clarifying questions to the customer rather than issuing a single canned reply; the story's core 'asks clarifying questions' behavior is only implied, not confirmed. Missing for 10: explicit product documentation or example of the agent asking clarifying follow-up questions mid-conversation, and independent/hands-on evidence of multi-turn troubleshooting flows in practice.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-15",
      "pylon-docs-20",
      "pylon-docs-19",
      "pylon-docs-17"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "multilingual-answers",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence in the pack mentions multilingual support, language detection, translation of knowledge base content, or the AI agent responding in customer languages other than English; docs cover agent building, knowledge base, and integrations but never address language coverage.",
    "evidenceIds": []
  },
  {
    "productId": "pylon",
    "storyId": "omnichannel-coverage",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon clearly supports one-agent omnichannel routing across chat, email, and in-app (chat widget, Slack, email integrations, omnichannel-support doc), but the evidence pack never explicitly documents WhatsApp or social-media channel integrations, which the story specifically calls out. missing for 10: explicit WhatsApp integration docs, explicit social media (e.g. Twitter/Instagram) channel support, and confirmation these route into the same unified agent workflow as chat/email/Slack.",
    "evidenceIds": [
      "pylon-docs-24",
      "pylon-docs-25",
      "pylon-docs-26",
      "pylon-docs-27",
      "pylon-docs-13"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "ongoing-qa-review",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon documents a monitor-and-iterate loop: a 'test' simulation mode to probe agent behavior, an issue log to inspect outcomes/steps taken by the AI, and runbooks/skills that let ops feed corrections back into the agent's instructions. However, there's no evidence of structured QA scoring of samples or a formal flagged-failure workflow with review states — it's more free-form inspection than a scored QA/review pipeline. Missing for 10: sample scoring/rubrics, dedicated failure-flagging workflow with reviewer roles, and quantitative QA metrics tied back to agent training.",
    "evidenceIds": [
      "pylon-docs-17",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-15",
      "pylon-docs-20"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "openness-api-parity",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon documents a public REST API and webhooks for programmatically accessing/acting on issues, accounts, and contacts (pylon-docs-31, pylon-docs-32, pylon-docs-4), plus an MCP server that lets AI tools read/update the same core objects (pylon-docs-34, pylon-probe-rt-3). However, there's no evidence the API/MCP surface covers the full breadth of UI capabilities shown in the changelog (custom object forms, analytics dashboards, chat widget branding/customization, knowledge base authoring, triggers/automation config) — these appear UI-only in the evidence pack. Missing for 10: documented API/MCP parity for analytics dashboards, chat widget customization, knowledge base management, and workflow/trigger configuration.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-docs-4",
      "pylon-docs-34",
      "pylon-probe-rt-2",
      "pylon-probe-rt-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "openness-full-export",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows a REST API and webhooks for programmatic access to issues/accounts/contacts, but nothing documents a bulk data-export feature, open-format dump, or data-portability/account-closure workflow that would let a user extract all their data and leave.",
    "evidenceIds": [
      "pylon-docs-31",
      "pylon-docs-32",
      "pylon-probe-rt-2"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "openness-open-license",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Pylon is a closed-source SaaS product; no evidence of any open-source license or public source repository is present in the pack—only product docs, API/MCP references, and security/compliance pages. missing for 10: any open-source license, public source code repository, or licensing statement.",
    "evidenceIds": []
  },
  {
    "productId": "pylon",
    "storyId": "openness-self-host",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Pylon is a hosted SaaS customer-support/helpdesk platform with no evidence of any self-hosted/on-prem deployment option; all evidence points to a cloud-only product (api.usepylon.com, mcp.usepylon.com, Vanta-managed compliance). Self-hosting is not a plausible axis for this SaaS product's offering model.",
    "evidenceIds": [
      "pylon-docs-10",
      "pylon-probe-rt-2",
      "pylon-probe-rt-3"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "personalized-account-answers",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon documents product-data syncs, Account Fields, custom objects, and Connectors/Custom API tools/MCP that let AI Support Agents pull in live account/product data (e.g., pylon-docs-8, pylon-docs-21, pylon-docs-22, pylon-docs-33) rather than relying solely on the knowledge base, and Skills/Runbooks let agents act on that data during resolution (pylon-docs-13, pylon-docs-20). However there's no first-party or independent case study showing an actual generated answer citing live plan/order-status/account-history data in a real resolution. Missing for 10: concrete example/transcript of an AI answer referencing live order status or plan data, and independent verification that synced data is actually surfaced in agent responses rather than just accessible to admins.",
    "evidenceIds": [
      "pylon-docs-8",
      "pylon-docs-21",
      "pylon-docs-22",
      "pylon-docs-33",
      "pylon-docs-13",
      "pylon-docs-20",
      "pylon-docs-34"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "pre-launch-simulation",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon's docs explicitly describe a test/simulate feature ('Simulate an interaction with the AI and see how it handles different questions') plus an issue log to inspect outcomes, which supports pre-deployment validation and monitoring. However, there's no explicit mention of testing against historical ticket datasets in bulk, regression testing, or scoring accuracy across a corpus of past conversations. missing for 10: bulk/historical-ticket backtesting workflow, quantitative accuracy/regression metrics from test runs, independent/hands-on confirmation of the simulate feature's depth.",
    "evidenceIds": [
      "pylon-docs-17",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-15"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "privacy-data-residency",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence anywhere in the pack mentions data residency, region selection, or storage location options; only compliance/certification via Vanta is mentioned, which does not address data residency choice. Missing for 10: any mention of regional data centers, residency options, or ability to select storage location.",
    "evidenceIds": [
      "pylon-docs-10"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "privacy-no-training",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "Evidence pack mentions Pylon's compliance is managed via Vanta (SOC2-type certification) but there is no mention of any AI-training data opt-out, data usage policy for model training, or controls letting users prevent their data from being used to train AI models. Missing for 10: explicit AI-training opt-out/data-usage policy, contractual or product-level control preventing training use, and any documentation addressing this specific privacy concern.",
    "evidenceIds": [
      "pylon-docs-10"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "privacy-retention-controls",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "The evidence pack shows security/compliance certification via Vanta and general data handling but contains no documentation of data retention periods, deletion policies, or user/admin controls to delete or export data. Missing for 10: retention policy documentation, deletion/export controls, data lifecycle settings.",
    "evidenceIds": [
      "pylon-docs-10"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "privacy-telemetry-optout",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence Pylon offers any telemetry opt-out or usage-tracking controls; evidence only covers compliance certification (Vanta) and product features, not telemetry settings.",
    "evidenceIds": []
  },
  {
    "productId": "pylon",
    "storyId": "procedure-sop-builder",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Pylon's Runbooks feature lets ops leads write natural-language, step-by-step instructions for the AI Support Agent to follow for specific scenarios, and Skills provide reusable instructions for consistent handling of known workflows — directly matching the SOP-encoding need. However, the docs describe these as natural-language guidance rather than an explicit deterministic branching/decision-tree engine, and Triggers/automations are a separate rule system not tied to agent SOPs. Missing for 10: explicit conditional/branching logic constructs within runbooks, and evidence of deterministic (non-LLM-interpreted) path selection.",
    "evidenceIds": [
      "pylon-docs-15",
      "pylon-docs-20",
      "pylon-docs-14",
      "pylon-docs-30"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "resolution-analytics",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Pylon docs confirm built-in analytics dashboards with 'common support metrics' filterable by attributes, supporting general resolution/ticket reporting, but no evidence explicitly names CSAT, handoff rate, or cost-per-resolution as tracked metrics. missing for 10: explicit documentation of CSAT scoring, handoff-rate metric, and cost-per-resolution calculation in dashboards.",
    "evidenceIds": [
      "pylon-docs-29"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "restricted-topic-controls",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Pylon documents assignment-based control—agents only interact with issues explicitly assigned to them, giving admins 'full control of which customer issues your AI agent interacts with'—which could be used to withhold sensitive categories from the agent, and runbooks/skills let you write natural-language handling instructions. But there is no documented feature for marking specific topics (legal threats, cancellations, security) as categorically human-only or any guardrail that blocks the agent from acting on flagged topic types. Missing for 10: an explicit topic/category exclusion or block-list mechanism, enforcement guarantees against agent 'freelancing' on excluded topics, and independent verification that assignment-based control reliably prevents agent action on sensitive categories.",
    "evidenceIds": [
      "pylon-docs-18",
      "pylon-docs-15",
      "pylon-docs-19",
      "pylon-docs-20"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "supervised-draft-mode",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Pylon's docs describe agent deployment controls (assigning which issues the agent handles, testing/simulating responses, and an issue log for post-hoc review) but nowhere document a supervised 'draft-for-human-approval-before-send' mode; the language instead emphasizes agents 'taking action' autonomously once deployed.",
    "evidenceIds": [
      "pylon-docs-13",
      "pylon-docs-18",
      "pylon-docs-19",
      "pylon-docs-17"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "topic-trend-insights",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Evidence shows generic analytics dashboards with filterable metrics (pylon-docs-29) and AI summarization of Slack conversations synced to Salesforce (pylon-docs-33), but nothing describes topic clustering of conversations or proactive surfacing of emerging product issues before ticket volume spikes.",
    "evidenceIds": [
      "pylon-docs-29",
      "pylon-docs-33"
    ]
  },
  {
    "productId": "pylon",
    "storyId": "transparent-per-resolution-pricing",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "No evidence anywhere in the pack of outcome-based or per-resolution pricing, published pricing tiers, caps, or usage-based billing controls; Pylon's documentation covers product features, agents, and APIs but is silent on pricing model or billing transparency.",
    "evidenceIds": []
  },
  {
    "productId": "pylon",
    "storyId": "voice-phone-support",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Pylon's evidence covers omnichannel support (chat, email, Slack), AI support agents, knowledge bases, and MCP/API integrations, but nothing describes voice/phone call handling with speech-in/speech-out capability. The omnichannel doc (pylon-docs-24) is generic and doesn't mention telephony or voice at all.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "agentic-agent-docs",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Direct probes show no llms.txt at sierra.ai (404) and docs.sierra.ai/llms.txt merely resolves to the login SPA HTML shell rather than an actual plain-text agent-oriented index; the real docs are login-gated to contracted customers, so an agent cannot be pointed at a genuine llms.txt or open agent-oriented docs.",
    "evidenceIds": [
      "sierra-probe-1",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-ai-insights",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Sierra's Explorer and Insights products explicitly deliver AI-generated insights: natural-language querying across conversations, automatic weekly briefings on trends/issues with recommendations, and click-to-investigate drill-downs on report data, plus explainability of agent reasoning. This is a first-party documented feature set directly matching the story, with some independent corroboration of Sierra's data-driven production use (sierra-comm-2/3). Missing for 10: independent hands-on validation of the insights/briefing feature specifically (community evidence is about agent setup, not analytics quality), and no detail on data freshness/accuracy limits.",
    "evidenceIds": [
      "sierra-docs-10",
      "sierra-docs-11",
      "sierra-docs-12",
      "sierra-docs-13",
      "sierra-comm-2",
      "sierra-comm-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-autonomous-automation",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's agents are built to operate autonomously across channels (chat, phone, email, SMS) once deployed, and Explorer explicitly runs in the background to deliver automatic weekly briefings and trend detection 'without you having to ask,' which is genuine unprompted automation. However, there's no documentation of user-configurable scheduled tasks/triggers beyond the always-on conversational agent and the one named automatic-briefing feature. Missing for 10: explicit support for user-defined scheduled/triggered background jobs beyond conversation handling and briefings, and independent/hands-on confirmation of autonomous background execution reliability.",
    "evidenceIds": [
      "sierra-docs-11",
      "sierra-docs-3",
      "sierra-docs-14",
      "sierra-comm-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-builtin-assistant",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's Ghostwriter feature lets users delegate agent-building tasks to a built-in AI assistant via natural-language prompts (build/modify agents, generate journeys from SOPs/transcripts), fitting the 'delegate tasks to built-in AI assistant' story. However, this is vendor-only documentation with no independent/hands-on corroboration of Ghostwriter specifically; community evidence instead describes manual point-and-click agent setup, not AI-assistant delegation. Missing for 10: independent/hands-on validation of Ghostwriter's delegation capability, and detail on scope/limits of what can be delegated.",
    "evidenceIds": [
      "sierra-docs-8",
      "sierra-docs-9",
      "sierra-docs-6"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-headless",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Sierra's Agent SDK is described as code-based with dev workflows retained, and one blog post explicitly notes publishing to ChatGPT can be done 'via CI/CD', implying some automation/pipeline support. However there is no dedicated CLI, headless runtime docs, or CI-specific guidance, and docs/API references are login-gated (probe shows /llms.txt is a login SPA and no public OpenAPI spec), so full headless/CI operation is unconfirmed. Missing for 10: explicit CLI/headless execution docs, public API/OpenAPI spec, independent CI usage reports.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-17",
      "sierra-probe-1",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-mcp-client",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence anywhere in the pack mentions MCP servers or the ability to plug external MCP tool servers into Sierra's agents; Sierra's integration mentions are about internal APIs and custom systems, not MCP. missing for 10: any mention of MCP protocol support, MCP client configuration, or third-party tool server integration via MCP.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "agentic-mcp-server",
    "verdict": "na",
    "quality": 0,
    "confidence": "high",
    "rationale": "Sierra is itself an AI agent platform (the agent role), not a service being connected to by external agents via MCP; this axis is about serving as an MCP server for other agents, which is a category mismatch for a product that is the agent itself. No evidence shows Sierra exposing an official MCP server, and the story's axis is more appropriate for Sierra being a client integrating others' tools than serving as one.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "agentic-nl-commands",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Sierra provides explicit natural-language control surfaces: Ghostwriter lets users 'build or modify agents by describing how you want them to behave' with 'simple prompts' for workflows, integrations, guardrails, tone and style, and Explorer lets users 'ask any question about your customer experience in natural language.' This directly matches an AI-native user operating the product via NL commands rather than only clicking through UI. Missing for 10: independent hands-on corroboration of NL command reliability/scope beyond vendor docs, and no evidence of NL control over the entire platform (e.g., release governance, channel deployment) rather than just Ghostwriter/Explorer.",
    "evidenceIds": [
      "sierra-docs-8",
      "sierra-docs-9",
      "sierra-docs-10",
      "sierra-comm-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-official-cli",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence of an official Sierra CLI tool; the Agent SDK mentions a code-based development workflow but nothing describing a CLI, and probes for llms.txt/openapi return 404s with no CLI reference anywhere in the pack.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-probe-1",
      "sierra-probe-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-public-api",
    "verdict": "partial",
    "quality": 4,
    "confidence": "medium",
    "rationale": "Sierra markets an Agent SDK for building 'customer journeys as code' with API-call inspection, implying a programmatic interface exists, but there's no publicly discoverable OpenAPI spec, llms.txt, or open API reference — docs.sierra.ai is login-gated to contracted customers rather than a documented public API. Missing for 10: publicly accessible API reference/OpenAPI spec, evidence of self-serve API keys or open documentation, independent developer confirmation of using the API without a sales contract.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-2",
      "sierra-probe-1",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-scoped-keys",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence anywhere in the pack of scoped or least-privilege API credential/token issuance for agents; docs describe agent building, workflows, channels, and analytics but nothing about credential scoping, permissions, or API key management. Probes even show no public OpenAPI/API docs are accessible (sierra-probe-2, sierra-probe-rt-3), reinforcing the absence of evidence.",
    "evidenceIds": [
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-sdks",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Sierra advertises an official \"Agent SDK\" for writing customer journeys as code with logic traces and testing (sierra-docs-1/2/4), but runtime probes show no public OpenAPI/swagger spec and the SDK/docs reference is login-gated to contracted customers rather than openly available to any AI-native developer (sierra-probe-2, sierra-probe-rt-3). Missing for 10: publicly accessible API reference/OpenAPI spec, evidence of open sign-up or trial SDK access, and independent developer accounts of building against it outside a paid contract.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-2",
      "sierra-docs-4",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "agentic-webhooks",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence in the pack mentions webhooks or event subscription capabilities for Sierra; docs pages describe SDK, studio, insights, voice, and channels features but nothing about outbound event notifications. Probes further show no public API spec or open documentation to confirm such a mechanism exists.",
    "evidenceIds": [
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "api-interactive-docs",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Probes show no public OpenAPI/interactive reference (404s at openapi.json paths, llms.txt returns login SPA shell), and docs are login-gated to contracted customers rather than an open interactive API reference with runnable examples.",
    "evidenceIds": [
      "sierra-probe-1",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "api-machine-spec",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Direct probes show no OpenAPI/Swagger spec at common paths (all 404) and no machine-readable llms.txt index; docs are login-gated rather than publicly exposing a spec. missing for 10: publicly downloadable OpenAPI/Swagger file, any machine-readable API spec endpoint.",
    "evidenceIds": [
      "sierra-probe-1",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "api-sandbox",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Sierra's docs mention scenario testing and regression avoidance (sierra-docs-4) and 'Agent Checks and Simulations' for proactive problem catching (sierra-docs-18), implying some pre-production testing capability, but there is no explicit mention of a dedicated sandbox environment isolated from production data. Missing for 10: explicit sandbox/staging environment documentation, confirmation that test runs don't touch production data, and independent/hands-on verification of this separation.",
    "evidenceIds": [
      "sierra-docs-4",
      "sierra-docs-18"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "api-versioning-policy",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence of a public versioned API reference or a documented deprecation policy; probes show no OpenAPI spec and docs are login-gated, and llms.txt/openapi.json all 404 or resolve to a login shell rather than API docs.",
    "evidenceIds": [
      "sierra-probe-1",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "automation-bulk-operations",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence describes bulk operations across many items (e.g., batch editing knowledge entries, mass workflow updates, or bulk conversation actions). The docs describe individual agent building, knowledge editing, and analytics tools, but nothing about performing actions at scale across many items simultaneously.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "automation-rules-engine",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Sierra supports building workflows/journeys and guardrail-based release governance (e.g., agent checks, split traffic, merge approval) which implies some rule-based triggering, and Ghostwriter/Agent Studio let users define step-by-step logic, but there's no explicit documentation of an event-driven 'if X happens, trigger Y automatically' rules engine for AI-native users to configure independently. missing for 10: explicit event-trigger/rules-engine documentation, API/SDK examples of automated action-on-event configuration, independent verification of this specific automation capability.",
    "evidenceIds": [
      "sierra-docs-6",
      "sierra-docs-18",
      "sierra-docs-15"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "automation-scheduled-jobs",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Sierra's evidence covers agent building, workflows, channels, analytics, and release governance, but nothing describes scheduling recurring jobs/workflows (e.g., cron-like triggers or timed automation runs) for AI-native users. Missing for 10: any documentation of scheduled/recurring job execution, trigger-based automation, or timer-based workflow runs.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "automation-versioned-workflows",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Sierra's Agent SDK is described as 'journeys as code' with change tracking, and release-governance docs mention merge approval, agent checks/simulations, and split-traffic gradual rollouts — pointing to versioned, reviewable release workflows. However, no explicit rollback/revert mechanism is documented, and docs are login-gated so no independent confirmation exists. missing for 10: explicit rollback capability, independent/hands-on confirmation of version history and revert function.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-18",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "brand-voice-tone",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Sierra provides explicit tone/brand-voice controls (Ghostwriter prompts for tone and style, Voice Personas tuned across 59 languages) plus testing/simulation tools to verify agent behavior across scenarios and avoid regressions, which supports consistency claims. However, all evidence is vendor-authored with no independent or hands-on verification that voice/tone actually stays consistent across topics and languages in production. Missing for 10: independent case studies or benchmarks confirming cross-topic/cross-language tone consistency, and detail on how brand-voice guardrails are enforced at scale.",
    "evidenceIds": [
      "sierra-docs-8",
      "sierra-docs-16",
      "sierra-docs-4",
      "sierra-docs-18"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "configurable-escalation-rules",
    "verdict": "partial",
    "quality": 3,
    "confidence": "low",
    "rationale": "Sierra's docs mention configurable guardrails, workflows, and a 'Live Assist' human-handoff product, implying some escalation mechanism exists, but no evidence describes explicit configuration of handoff triggers by topic, sentiment, customer tier, or explicit customer request, nor proof of reliable adherence to such rules. Missing for 10: documentation of specific trigger types (topic/sentiment/tier/request), configuration UI/API for these rules, and evidence (first-party or independent) that the agent reliably obeys them in production.",
    "evidenceIds": [
      "sierra-docs-8",
      "sierra-docs-14",
      "sierra-docs-18",
      "sierra-docs-6"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "content-gap-detection",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Sierra's Explorer and Insights products surface emerging issues, trends, and the reasoning/knowledge sources behind agent answers, and Agent Studio lets teams view/manage knowledge content (FAQs, policies), which could help surface gaps — but no evidence explicitly describes detecting knowledge gaps or conflicting content as a distinct feature. Missing for 10: explicit conflicting-content/contradiction detection, explicit 'knowledge gap' flagging, and independent/hands-on confirmation that Explorer or Insights actually pinpoints such gaps rather than general conversation trends.",
    "evidenceIds": [
      "sierra-docs-7",
      "sierra-docs-10",
      "sierra-docs-11",
      "sierra-docs-12",
      "sierra-docs-13"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "context-rich-handoff",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's Live Assist product explicitly targets escalation handoff, claiming reps are guided through next steps with automatically captured and updated customer context and can pick up conversations 'no tab-switching, no referencing instructions, no lost context' — directly addressing the no-repeat-yourself goal. However, there's no explicit mention of a generated conversation summary artifact or hands-on/independent confirmation that reps actually receive full transcript + summary + collected details in practice. Missing for 10: explicit summary-generation evidence, independent/hands-on validation of the handoff experience, confirmation reps see full conversation history alongside context.",
    "evidenceIds": [
      "sierra-docs-14",
      "sierra-docs-15"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "end-to-end-resolution",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's marketing emphasizes agent-driven end-to-end workflow execution (journeys, integrations, payments, guardrails, testing) implying resolution rather than mere deflection, and community evidence corroborates real production use handling customer processes end-to-end for at least one customer. However, no quantitative resolution-rate metrics, benchmarks, or third-party validation of 'meaningful share resolved end-to-end' are provided — tau-bench is a research benchmark, not a customer outcome metric. Missing for 10: published resolution-rate statistics, case studies with concrete resolution percentages, independent audits distinguishing resolution from deflection.",
    "evidenceIds": [
      "sierra-docs-6",
      "sierra-docs-9",
      "sierra-docs-19",
      "sierra-docs-20",
      "sierra-comm-2",
      "sierra-comm-3",
      "sierra-probe-rt-1"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "grounded-cited-answers",
    "verdict": "full",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Sierra explicitly supports grounding agent answers in customer-owned knowledge (Help Center content, FAQs, policies) via Agent Studio, and Insights lets teams 'understand the reasoning behind every agent action or answer—view knowledge sources, systems accessed' which directly maps to showing which source an answer drew from. Missing for 10: independent/hands-on verification that end-user-facing answers visibly cite specific articles, and detail on citation UX rather than just admin-side reasoning traces.",
    "evidenceIds": [
      "sierra-docs-7",
      "sierra-docs-13",
      "sierra-docs-10"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "hallucination-guardrails",
    "verdict": "partial",
    "quality": 4,
    "confidence": "low",
    "rationale": "Sierra's docs mention grounding agents in knowledge/FAQs/policies (sierra-docs-7) and general 'guardrails' as a configurable behavior via Ghostwriter (sierra-docs-8), plus release-governance guardrails like Agent Checks/Simulations (sierra-docs-18) and visibility into reasoning/knowledge sources (sierra-docs-13). However, no evidence explicitly describes a 'safe decline' mechanism for off-knowledge questions or confirms the agent won't invent prices/policies rather than guess. Missing for 10: explicit documentation or hands-on proof of decline-on-unknown behavior, third-party validation that hallucinated policies/prices are prevented, and detail on the unexplored 'trust-and-reliability' product page.",
    "evidenceIds": [
      "sierra-docs-7",
      "sierra-docs-8",
      "sierra-docs-13",
      "sierra-docs-18"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "helpdesk-platform-integrations",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence pack citation mentions Zendesk, Salesforce, Intercom, or bidirectional ticket/context syncing with existing helpdesk platforms; only generic 'systems integrations' and 'internal APIs' are referenced without naming any helpdesk system or describing two-way ticket sync.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "knowledge-auto-sync",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "Evidence shows Sierra lets teams manually 'view, manage, and edit knowledge data' (sierra-docs-7) but nothing describes automatic re-syncing of knowledge sources on a schedule or on-change detection; docs are login-gated so no further detail is visible.",
    "evidenceIds": [
      "sierra-docs-7",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "knowledge-source-ingestion",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's docs explicitly claim agent-studio lets teams 'view, manage, and edit knowledge data such as Help Center content, FAQs, and policies' and Ghostwriter can ingest SOPs, transcripts, and audio interviews to build journeys, supporting knowledge ingestion beyond just help center docs. However, there's no explicit evidence of ingesting past tickets or internal wikis specifically, and docs/reference material is login-gated so independent verification of breadth of source-type ingestion is limited. Missing for 10: documented support for tickets/wiki ingestion specifically, and independent/hands-on confirmation of the ingestion workflow beyond marketing copy.",
    "evidenceIds": [
      "sierra-docs-7",
      "sierra-docs-9",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "live-api-actions",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Sierra's Agent SDK/Studio docs show agents can call customer APIs to execute actions (order changes, refunds implied by 'internal APIs' use), with guardrails, release governance, and human-in-the-loop approval, and a HN commenter confirms it wires directly into a customer's internal APIs. However there's no explicit documentation of per-action scoped auth/permissioning model for API calls. missing for 10: explicit scoped-auth/permission model per action, concrete refund/subscription action examples, independent confirmation of granular auth scoping.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-2",
      "sierra-docs-18",
      "sierra-comm-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "multi-turn-troubleshooting",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's docs describe building step-by-step, multi-step 'customer journeys' and workflows (sierra-docs-6, sierra-docs-9) rather than single canned replies, and community commentary confirms agents are configured to work through processes with internal APIs (sierra-comm-2). However, no evidence explicitly demonstrates the agent proactively asking clarifying questions or a hands-on troubleshooting transcript showing multi-turn dialogue in practice. Missing for 10: concrete transcript/demo of clarifying-question behavior, independent evaluation of troubleshooting depth, and explicit mention of clarification-seeking as a designed capability.",
    "evidenceIds": [
      "sierra-docs-6",
      "sierra-docs-9",
      "sierra-comm-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "multilingual-answers",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra documents multilingual voice capability explicitly ('Voice Personas... tuned across 59 languages') and general multi-channel deployment, implying broad language coverage for customer-facing conversations. However, there is no direct evidence about handling a knowledge base that exists only in English—no mention of automatic translation/grounding of English-only content into other languages, so the specific claim in the story (KB gap bridging) is unevidenced. missing for 10: explicit documentation on cross-language grounding from English-only knowledge base, independent verification of multilingual quality beyond voice.",
    "evidenceIds": [
      "sierra-docs-16",
      "sierra-docs-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "omnichannel-coverage",
    "verdict": "partial",
    "quality": 5,
    "confidence": "medium",
    "rationale": "Sierra's docs confirm one agent can be built once and deployed across chat, phone, email, SMS, and messaging channels (sierra-docs-3), with channel-specific tuning like Voice Personas (sierra-docs-16) and even ChatGPT publishing (sierra-docs-17), showing broad omnichannel intent. However, none of the evidence explicitly names Slack, WhatsApp, or social media as supported channels, leaving the specific channels support leaders care about unconfirmed. Missing for 10: explicit documentation or independent confirmation that Slack, WhatsApp, and social platforms are supported channels, plus hands-on evidence of a single agent operating consistently across these specific channels.",
    "evidenceIds": [
      "sierra-docs-3",
      "sierra-docs-16",
      "sierra-docs-17"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "ongoing-qa-review",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Sierra documents Agent Checks and Simulations that proactively catch problems, merge-approval workflows for human review, and split-traffic releases (sierra-docs-18), plus regression testing (sierra-docs-4) and Explorer/Insights tools that surface conversation trends, flag emerging issues, and explain agent reasoning (sierra-docs-10, -11, -13). Ghostwriter lets teams feed fixes back by updating workflows/guardrails via prompts (sierra-docs-8). Missing for 10: explicit description of scored/sampled QA reviews with quantitative scoring rubrics, and independent evidence of the review loop actually closing the gap between flagged failures and shipped fixes.",
    "evidenceIds": [
      "sierra-docs-4",
      "sierra-docs-18",
      "sierra-docs-10",
      "sierra-docs-11",
      "sierra-docs-13",
      "sierra-docs-8"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "openness-api-parity",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Sierra markets an Agent SDK for code-based building alongside a separate no-code Agent Studio, but there is no evidence of API/UI feature parity — no public OpenAPI/swagger spec was found (404s), and the SDK reference docs themselves are login-gated to contracted customers, meaning even documented API scope can't be verified as matching UI capabilities.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-5",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "openness-full-export",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "No evidence of any data export or portability feature for customer journeys, knowledge data, or conversation logs in open formats; probes even show docs/API surfaces are gated or unavailable (404s, login-gated /llms.txt). Missing for 10: any export functionality, data portability documentation, or open-format data dump capability.",
    "evidenceIds": [
      "sierra-probe-1",
      "sierra-probe-2",
      "sierra-probe-rt-3"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "openness-open-license",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Sierra is presented purely as a closed commercial SaaS platform (Agent SDK, Studio, Ghostwriter, Voice, etc.) with no evidence of an open-source license for the product itself; the only open artifact found is the unrelated tau-bench research benchmark repo, not Sierra's product source code, and docs/API surfaces are login-gated or 404. Missing for 10: any open-license repository for Sierra's actual product code, license file, or public source release.",
    "evidenceIds": [
      "sierra-probe-rt-1",
      "sierra-probe-rt-3",
      "sierra-probe-1",
      "sierra-probe-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "openness-self-host",
    "verdict": "none",
    "quality": 0,
    "confidence": "high",
    "rationale": "Sierra is offered as a hosted SaaS platform with no evidence of any self-hosted or on-premise deployment option; docs are login-gated to contracted customers rather than exposing an installable/self-hostable core product. missing for 10: any documentation of self-hosting, on-prem deployment, or open-source release of the core agent platform.",
    "evidenceIds": [
      "sierra-probe-rt-3",
      "sierra-probe-1",
      "sierra-probe-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "personalized-account-answers",
    "verdict": "partial",
    "quality": 7,
    "confidence": "medium",
    "rationale": "Sierra's docs and community evidence show the agent integrates with customers' internal APIs and systems (not just static knowledge), with Insights explicitly noting 'systems accessed' during agent actions, and Live Assist capturing live customer context. However, no evidence explicitly confirms real-time retrieval of specifics like plan, order status, or account history. missing for 10: explicit documented example of pulling live order/plan/account data, independent verification beyond one HN anecdote.",
    "evidenceIds": [
      "sierra-docs-2",
      "sierra-docs-13",
      "sierra-docs-14",
      "sierra-comm-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "pre-launch-simulation",
    "verdict": "full",
    "quality": 8,
    "confidence": "high",
    "rationale": "Sierra explicitly documents 'Agent Checks and Simulations' and 'verify your agent performs as expected across a wide range of scenarios and avoid regressions' (sierra-docs-4, sierra-docs-18), directly matching pre-release testing against scenarios/regressions. This is further corroborated by Sierra's public tau-bench research benchmark for evaluating conversational agents on simulated user interactions (sierra-probe-rt-1), showing real investment in simulation-based testing methodology. Missing for 10: no explicit mention of testing against historical/real ticket transcripts specifically (only 'scenarios' and simulations), and no independent hands-on customer account of the simulation workflow in practice.",
    "evidenceIds": [
      "sierra-docs-4",
      "sierra-docs-18",
      "sierra-probe-rt-1"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "privacy-data-residency",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence pack item addresses data residency, region selection, or storage location controls; the docs cover agent building, workflows, and channels but not data governance/residency options. missing for 10: any mention of data region selection, residency guarantees, or storage location controls.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "privacy-no-training",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence in the pack addresses data-training opt-out or AI model training policies; Sierra's docs focus on product features (agent building, analytics, channels) with no privacy/data-use policy statements provided.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "privacy-retention-controls",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence pack item addresses data retention controls, deletion policies, or user-facing data lifecycle management; docs cover agent building, analytics, and channels but never mention retention/deletion settings, and docs.sierra.ai is login-gated so no public verification exists.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "privacy-telemetry-optout",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence pack item discusses telemetry, usage tracking, opt-out controls, or privacy settings for AI-native/developer users; docs focus on product features and docs are gated behind login. Missing for 10: any documentation of telemetry collection practices, opt-out mechanism, or privacy controls.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "procedure-sop-builder",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Sierra's Agent SDK/Studio explicitly supports authoring step-by-step journeys ('workflows') as code or no-code, with branching logic ('sophisticated logic'), SOP ingestion via Ghostwriter, and simulation/testing to verify deterministic behavior across scenarios. This directly matches encoding SOPs with deterministic branching for known issue types.\n\nmissing for 10: independent/hands-on verification of deterministic branching behavior specifically (community evidence is generic, not focused on SOP branching), and docs are login-gated so full workflow-editor detail isn't independently viewable.",
    "evidenceIds": [
      "sierra-docs-1",
      "sierra-docs-6",
      "sierra-docs-9",
      "sierra-docs-4",
      "sierra-comm-2"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "resolution-analytics",
    "verdict": "partial",
    "quality": 5,
    "confidence": "low",
    "rationale": "Sierra's Insights and Explorer products provide dashboards/reports with drill-down and natural-language trend analysis (sierra-docs-10, sierra-docs-11, sierra-docs-12, sierra-docs-13), showing the platform surfaces analytics for support leaders, but no evidence specifies the exact metrics named in the story (resolution rate, CSAT, handoff rate, cost per resolution). Missing for 10: explicit documentation naming these specific KPIs, screenshots/examples of the actual dashboard metrics, and independent confirmation that leaders use it for exec-level reporting.",
    "evidenceIds": [
      "sierra-docs-10",
      "sierra-docs-11",
      "sierra-docs-12",
      "sierra-docs-13"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "restricted-topic-controls",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Sierra's guardrail evidence covers release governance (Agent Checks, Simulations, merge approval, split traffic) and general knowledge/workflow editing, but nothing describes topic-level human-only flags for categories like legal threats, cancellations, or security that the agent is barred from handling. Live-assist shows human+AI collaboration but not a mechanism to designate certain topics as strictly human-only with enforced escalation.",
    "evidenceIds": [
      "sierra-docs-18",
      "sierra-docs-14"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "supervised-draft-mode",
    "verdict": "none",
    "quality": 0,
    "confidence": "medium",
    "rationale": "Sierra's evidence covers release governance (merge-approval for deploying agent changes) and Live Assist (guiding human reps in real time), but neither describes a mode where the agent drafts individual customer replies that a human must approve before they are sent. No docs or community evidence mention message-level human-in-the-loop approval for live customer conversations.",
    "evidenceIds": [
      "sierra-docs-18",
      "sierra-docs-14",
      "sierra-docs-15"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "topic-trend-insights",
    "verdict": "partial",
    "quality": 6,
    "confidence": "medium",
    "rationale": "Sierra's Explorer product explicitly surfaces trends and emerging issues via natural-language querying of conversations and proactive weekly briefings, and Insights lets users drill into data points—directly matching the 'surface emerging issues before they spike' theme. However, evidence doesn't explicitly confirm topic clustering methodology, predictive spike detection, or independent/hands-on validation of this specific capability. Missing for 10: explicit description of topic-clustering mechanics, evidence of predictive alerting before ticket-volume spikes (vs. reactive weekly summaries), and independent/customer corroboration of this specific insights capability.",
    "evidenceIds": [
      "sierra-docs-10",
      "sierra-docs-11",
      "sierra-docs-12"
    ]
  },
  {
    "productId": "sierra",
    "storyId": "transparent-per-resolution-pricing",
    "verdict": "none",
    "quality": 0,
    "confidence": "low",
    "rationale": "No evidence pack item mentions pricing, resolution-based billing, caps, or published rates; docs and probes cover product features and access gating only. Sierra is widely known anecdotally for outcome-based pricing but nothing in this evidence pack substantiates published, capped, self-serve pricing terms.",
    "evidenceIds": []
  },
  {
    "productId": "sierra",
    "storyId": "voice-phone-support",
    "verdict": "full",
    "quality": 8,
    "confidence": "medium",
    "rationale": "Sierra explicitly markets voice as a first-class channel with the same agent logic/knowledge as chat ('Build once and deploy across any channel—chat, phone, email, SMS'), plus voice-specific features like Voice Personas across 59 languages, replacing IVR with empathetic voice agents, and phone payments without IVR handoff. Missing for 10: independent/hands-on verification that voice calls actually share identical knowledge/actions with chat in production (only vendor docs, no community confirmation of phone-specific parity).",
    "evidenceIds": [
      "sierra-docs-3",
      "sierra-docs-16",
      "sierra-docs-19",
      "sierra-docs-20",
      "sierra-docs-14"
    ]
  }
]
