Skip to content

Langfuse wins · 2319 (8 drawn)

Agenticness — how well agents can access and operate the productAgenticness

How well agents can access and operate the product

Agent access

  1. ai-native userPoint an agent at llms.txt or agent-oriented docs

    weight 2 · round drawn
    Langfusefullprobed9/10

    Langfuse serves both llms.txt and docs.md, confirmed by direct HTTP 200 probes, making its docs directly consumable by agents. missing for 10: independent third-party confirmation of agent usage of these endpoints in practice.

    • [probe] PROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](htt…
    • [probe] PROBE docs-md: HTTP 200 at https://langfuse.com/docs.md --- title: Overview seoTitle: Open Source AI Engineering Platform description: Langf…
    Cekurafullprobed9/10

    Cekura has a confirmed live llms.txt endpoint (HTTP 200) plus per-page .md documentation format and an index pointing agents to llms.txt, directly enabling agent-oriented doc consumption; this is corroborated by both docs and live probes. Missing for 10: independent third-party confirmation of an agent actually consuming llms.txt successfully in the wild.

    • [probe] PROBE llms.txt: HTTP 200 at https://docs.cekura.ai/llms.txt # Cekura > Cekura is the testing and observability platform for voice AI agents…
    • [probe] PROBE docs-md: HTTP 200 at https://docs.cekura.ai/documentation/introduction.md > ## Documentation Index > Fetch the complete documentation …
    • [claimed-docs] create a comprehensive plan to integrate cekura observability
  2. ai-native userRun the product headlessly / in CI for automation

    weight 2 · round to Cekura
    Langfusepartialclaimed6/10

    Langfuse supports headless/CI automation via SDKs (Python/JS), OTLP endpoint, Metrics/Query APIs, CI/CD experiment blocking on regressions, and self-hosted Docker deployment — all usable without a UI. However, there's no dedicated CI recipe/example workflow (e.g., GitHub Actions template) or hands-on community confirmation of running it fully headlessly in pipelines. Missing for 10: concrete CI pipeline examples/templates, independent/hands-on evidence of headless automation success, and confirmation that all core evaluation features work without any manual UI step.

    • [claimed-docs] Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
    • [claimed-docs] Block deploys on regressions | [CI/CD experiments](/docs/evaluation/experiments/experiments-ci-cd)
    • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
    • [claimed-docs] Langfuse can receive traces on the `/api/public/otel` (OTLP) endpoint.
    • [claimed-docs] Block deploys on regressions | CI/CD experiments
    • [claimed-docs] Block deploys on regressions CI/CD experiments
    • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
    • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
    • [claimed-docs] You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
    Cekurafullprobed8/10

    Cekura documents a CLI/SDK for programmatic and terminal-based agent/scenario/run management, a REST API with OpenAPI spec and API keys, cron job scheduling for automated evaluation workflows, and explicit GitHub Actions CI/CD integration for testing agents on every change—together these clearly support headless/CI automation. Missing for 10: independent hands-on confirmation of the GitHub Actions or cron setup actually working in a real pipeline beyond docs.

    • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
    • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
    • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
    • [claimed-docs] Learn how to set up GitHub Actions workflows to automatically test your agents with each change
    • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
    • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
    • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli
  3. ai-native userConnect an agent via an official MCP server

    weight 3 · round to Cekura
    Langfusefullprobed7/10

    Langfuse documents an official MCP server (docs-mcp) that lets AI coding agents (e.g., Cursor) integrate Langfuse tracing, and separately offers an MCP Server for connecting AI tools that cannot run shell commands, confirmed via a live probe of the docs page. missing for 10: independent/hands-on verification of the MCP server's reliability and no community discussion specifically validating this feature.

    • [claimed-docs] Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
    • [claimed-docs] Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
    • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
    • [claimed-docs] Connect an AI tool that cannot run shell commands MCP Server
    • [probe] official MCP server documented at https://langfuse.com/docs/docs-mcp
    Cekurafullprobed8/10

    Cekura is not itself an agent but a testing/observability platform, so the axis applies; it documents an official MCP server ("Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations") confirmed live via probe. Missing for 10: independent hands-on verification of the MCP server working end-to-end and more detail on supported tool/resource surface.

    • [claimed-docs] Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations.
    • [probe] official MCP server documented at https://docs.cekura.ai/mcp/overview
  4. ai-native userUse an official CLI

    weight 2 · round to Cekura
    Langfusefullclaimed7/10

    Docs explicitly advertise a CLI for working with Langfuse from a terminal or coding agent ('Work with Langfuse from a terminal or coding agent | CLI'), directly matching the story. Missing for 10: no independent/hands-on corroboration of CLI usage or details on install/command coverage beyond the single doc mention.

    • [claimed-docs] Work with Langfuse from a terminal or coding agent | CLI
    Cekurafullprobed8/10

    Cekura documents an official CLI ('Manage agents, scenarios, runs, and call data from your terminal') alongside an SDK, confirmed by both docs and probe evidence. Missing for 10: independent hands-on review of the CLI's usability/completeness beyond first-party docs.

    • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
    • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli
  5. ai-native userDrive the product through a documented public API

    weight 3 · round to Cekura
    Langfusepartialprobed6/10

    Langfuse documents a public API usable via SDKs (Python/JS), a Metrics API v2, blob export, and even a CLI and MCP server for driving it programmatically, and community evidence confirms API usage (though one user found the trace/generation retrieval API slow for bulk operations, bypassing it via direct DB queries). However, the probe found no discoverable OpenAPI/swagger spec at standard locations, which weakens the 'documented public API' claim despite docs referencing it. missing for 10: a verifiable OpenAPI/swagger spec link, deeper API reference docs beyond feature-page mentions, and stronger independent corroboration of full API coverage.

    • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
    • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
    • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
    • [claimed-docs] Work with Langfuse from a terminal or coding agent | CLI
    • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
    • [probe] PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/opena…
    • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
    Cekurafullprobed9/10

    Cekura publishes a documented public API with OpenAPI spec (probe confirms live openapi.json), API key auth, dedicated API reference endpoints (e.g. send-calls), plus SDKs (sync/async) and a CLI for programmatic/terminal control—clearly enabling an AI-native user to drive the product via documented API access. Missing for 10: independent third-party review/hands-on validation of the API beyond vendor docs and probes.

    • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
    • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
    • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
    • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
    • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
    • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli
  6. ai-native userIssue scoped/least-privilege API credentials for an agent

    weight 2 · round to Cekura
    Langfusenone0/10

    The evidence pack documents Langfuse's tracing, prompt management, evaluation, and API/export features, but contains no mention of API key scoping, role-based permissions, or least-privilege credential issuance for agents. Since Langfuse exposes public API keys and a self-hosted deployment model, scoped credentials would be a fair capability to expect, but none is evidenced here.

      Cekurapartialclaimed4/10

      Cekura documents API key creation restricted to project-level scope, showing some credential scoping capability, but there's no evidence of finer-grained least-privilege controls like read/write permission tiers, resource-specific scoping, or per-agent credential issuance. missing for 10: granular permission scopes (read-only vs write), per-agent or per-resource credential issuance, role-based access control documentation.

      • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
    • ai-native userBuild against official SDKs

      weight 2 · round drawn
      Langfusefullprobed8/10

      Langfuse documents native SDKs for Python/JS across observability, prompt management, and API/data platform querying (e.g., 'Query via SDKs'), plus 100+ framework integrations and OpenTelemetry support, giving AI-native developers first-class programmatic access. Community feedback corroborates strong DX/integration quality, though some users note prompt management SDK features lag competitors and bulk API retrieval performance issues. missing for 10: independent hands-on verification of SDK completeness/versioning, and no explicit public API reference (OpenAPI probe 404s).

      • [claimed-docs] Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
      • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
      • [claimed-docs] Langfuse Prompt Management adds no latency to your application. Prompts are cached client-side by the SDK, so retrieving them is as fast as …
      • [community] Langfuse is by far the best of the langs (-chain, -graph, -smith, -flow) in terms of UI/DX/integration/docs/quality.
      • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
      • [community] Langfuse seemed way better at tracing though [vs PromptLayer]... but Langfuse's prompt management was pretty behind in terms of features whe…
      • [probe] PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/opena…
      Cekurafullprobed8/10

      Cekura documents official sync/async SDKs and a CLI for programmatic access ('Manage agents, scenarios, runs, and call data from your terminal', 'sync and async clients for application code'), plus an OpenAPI spec, backed by probe confirmation of live docs and endpoints. missing for 10: independent hands-on developer reports specifically about SDK usage (only community praise for the platform generally, not SDK-specific).

      • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
      • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
      • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
      • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli
      • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
    • ai-native userSubscribe to events via webhooks

      weight 2 · round to Langfuse
      Langfusepartialclaimed5/10

      Langfuse's docs explicitly mention webhook notifications as part of its alerting feature ('Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold'), confirming webhook subscription support exists. However, this is only in the context of alert thresholds, not a general-purpose event subscription system, and there is no documentation of webhook payload schemas, event types, or configuration details, nor independent/community corroboration. Missing for 10: detailed webhook event catalog/API reference, broader event-subscription capabilities beyond alerting, and hands-on or community validation of webhook reliability.

      • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
      Cekuranone0/10

      Evidence shows Cekura receiving inbound webhooks (agent's provider posts call data to Cekura) but nothing about Cekura emitting outbound webhooks that a user could subscribe to for events like evaluation completion or drift detection. Missing for 10: documented outbound webhook/event subscription mechanism, webhook configuration UI/API, event types list.

      • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…

    Agentic features

    1. ai-native userGet AI-generated insights and suggestions from my data inside the product

      weight 2 · round to Cekura
      Langfusepartialclaimed4/10

      Langfuse provides LLM-as-a-Judge to automatically score/evaluate production traces, which is a form of AI-generated assessment of data, plus dashboards and alerts, but these are threshold/metric-based rather than proactive AI-generated insights or suggestions (e.g., natural-language recommendations, anomaly explanations). Missing for 10: evidence of narrative AI-generated insights/summaries, proactive suggestions for improving prompts/agents beyond LLM-as-a-Judge scoring, and any independent corroboration that these AI features are used for 'insights' rather than pure evaluation.

      • [claimed-docs] LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
      • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
      • [claimed-docs] In Langfuse, that score can be numeric, categorical, or boolean.
      • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
      Cekurapartialclaimed6/10

      Cekura provides LLM-judge evaluation metrics, live drift detection, and auto-generated fix suggestions ('flags issues → reproduces in simulation → suggests fixes automatically'), showing AI-generated insights derived from call/agent data. However, this is narrowly scoped to voice-agent testing/observability rather than broad, ad-hoc AI-generated insights across arbitrary product data, and there's no independent hands-on evidence of the 'suggests fixes automatically' claim beyond marketing copy. Missing for 10: independent/hands-on verification of the auto-fix-suggestion feature, broader insight/analytics generation beyond voice-agent metrics, and detail on how suggestions are surfaced to users.

      • [claimed-docs] Evaluate AI voice agent calls using natural language criteria with LLM-powered judging
      • [claimed-docs] Live drift detection across every call.
      • [claimed-docs] Cekura flags issues → reproduces in simulation → suggests fixes automatically.
      • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.
    2. ai-native userSet up automations that run autonomously in the background

      weight 2 · round to Cekura
      Langfusepartialclaimed4/10

      Langfuse supports some autonomous background processes within its observability/evaluation domain: automatic alerts when metrics cross thresholds, scheduled blob storage exports, and automatic LLM-as-a-Judge scoring of live production traces without manual intervention. However, it is not a general agentic automation/workflow engine — these are narrow, observability-specific triggers rather than a broad 'set up automations that run autonomously' capability. Missing for 10: a general-purpose scheduler/automation builder, evidence of arbitrary background task orchestration, and independent confirmation that alerts/scoring reliably run unattended at scale.

      • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
      • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
      • [claimed-docs] Set up alerts: get notified automatically when spend crosses a threshold.
      • [claimed-docs] Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
      • [claimed-docs] Export large volumes on a schedule | Blob Storage Export
      Cekurapartialclaimed6/10

      Cekura supports autonomous background automation via cron jobs for scheduled testing/evaluation runs, GitHub Actions CI/CD integration, and live drift detection across calls, plus automatic flagging/reproduction/fix-suggestion pipelines. However, evidence is entirely first-party docs describing scheduling/triggering mechanisms rather than a broad general-purpose 'automation builder' for arbitrary agentic workflows. missing for 10: independent/hands-on confirmation that cron/CI automations run reliably unattended, evidence of more flexible trigger types beyond cron/CI/webhooks, and detail on monitoring/error-handling for long-running autonomous automations.

      • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
      • [claimed-docs] Learn how to set up GitHub Actions workflows to automatically test your agents with each change
      • [claimed-docs] Live drift detection across every call.
      • [claimed-docs] Cekura flags issues → reproduces in simulation → suggests fixes automatically.
      • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
    3. ai-native userDelegate tasks to a built-in AI assistant inside the product

      weight 3 · round drawn
      Langfusenone0/10

      Langfuse's evidence covers observability, prompt management, evaluation, MCP server connectivity, and self-hosting, but nothing describes a built-in AI assistant within the product itself that users can delegate tasks to; the MCP/docs-mcp features are for external coding agents integrating with Langfuse, not an assistant embedded in the Langfuse UI.

      • [claimed-docs] Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
      • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
      • [probe] official MCP server documented at https://langfuse.com/docs/docs-mcp
      Cekuranone0/10

      Cekura's AI-assistant integrations (Skills, MCP, CLI) are designed so external AI assistants like Claude or Cursor can drive Cekura's testing/evaluation features — this is the reverse relationship of an AI-native user delegating tasks to a built-in assistant inside Cekura itself. The 'flags issues → reproduces → suggests fixes automatically' pipeline is an automated feature, not a delegatable in-product assistant, so there is no evidence of a built-in AI assistant you can converse with or delegate open-ended tasks to.

      • [claimed-docs] Cekura flags issues → reproduces in simulation → suggests fixes automatically.
      • [claimed-docs] Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations.
      • [claimed-docs] npx skills add cekura-ai/cekura-skills --all
    4. ai-native userOperate the product with natural-language commands

      weight 2 · round drawn
      Langfusepartialprobed6/10

      Langfuse documents an official MCP server that lets AI tools/agents interact with it without running shell commands, plus a CLI explicitly for 'terminal or coding agent' use and AI-coding-agent integration (Cursor) for auto-instrumenting tracing — all first-party evidence that AI-native users can drive Langfuse via natural-language-directed agent workflows. However this is exposed through MCP/CLI tooling rather than a native NL command interface, and there is no independent/hands-on evidence validating the experience of controlling Langfuse purely via natural language. Missing for 10: independent/community validation of NL-driven operation, and richer documentation of exactly which actions/commands are exposed via the MCP server or CLI.

      • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
      • [claimed-docs] Connect an AI tool that cannot run shell commands MCP Server
      • [claimed-docs] Work with Langfuse from a terminal or coding agent | CLI
      • [claimed-docs] Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
      • [claimed-docs] Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
      • [probe] official MCP server documented at https://langfuse.com/docs/docs-mcp
      Cekurapartialprobed6/10

      Cekura ships an MCP server and Skills package so an AI assistant can design, run, and improve voice-agent evaluations via natural-language commands, plus an LLM-judge metric for natural-language evaluation criteria; this supports agentic/NL operation for a subset of workflows. missing for 10: independent hands-on evidence of natural-language command execution beyond vendor docs, and confirmation that the full breadth of platform actions (not just evaluation/testing tasks) can be driven purely by NL commands.

      • [claimed-docs] npx skills add cekura-ai/cekura-skills --all
      • [claimed-docs] Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations.
      • [github] 13 Skills, 14 Commands in one plugin
      • [claimed-docs] Evaluate AI voice agent calls using natural language criteria with LLM-powered judging
      • [probe] official MCP server documented at https://docs.cekura.ai/mcp/overview

    Api quality

    1. ai-native userExplore an interactive API reference with runnable examples

      weight 2 · round to Cekura
      Langfusenone0/10

      The evidence pack shows Langfuse has an API (Metrics API v2, SDKs, Query via SDKs) but no mention of an interactive API reference page with runnable/try-it-out examples; a direct probe for OpenAPI/Swagger endpoints returned 404s, suggesting no such interactive reference was discoverable. Missing for 10: any documentation or screenshot of an interactive API explorer, runnable code snippets in an API reference UI, or a working OpenAPI/Swagger spec.

      • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
      • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
      • [probe] PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/opena…
      Cekurapartialprobed4/10

      Cekura publishes an OpenAPI spec and structured API-reference docs (e.g., send-calls, create-api-key endpoints) confirming a machine-readable interactive reference exists, but no evidence shows runnable 'try it' examples or a sandboxed console in the docs. Missing for 10: explicit runnable/try-it example UI, evidence of live request execution, independent confirmation of interactivity.

      • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
      • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
      • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
    2. ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

      weight 2 · round to Cekura
      Langfusenone0/10

      While Langfuse's docs reference an API, SDKs, and a Metrics API v2, a direct probe for a machine-readable spec (openapi.json, swagger.json, etc.) returned 404 on all candidate paths, and no evidence pack item links to a downloadable OpenAPI/Swagger file. No documentation citation confirms an accessible spec despite the API existing.

      • [probe] PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/opena…
      • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
      • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
      Cekurafullprobed9/10

      A direct probe confirms an OpenAPI spec is publicly downloadable at https://docs.cekura.ai/openapi.json, returning HTTP 200 with a valid "openapi" key, corroborated by API reference docs built from it. Missing for 10: no independent third-party confirmation of spec completeness/versioning beyond the probe.

      • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
      • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
      • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
    3. ai-native userTest against a sandbox environment without touching production data

      weight 1 · round to Cekura
      Langfusepartialclaimed4/10

      Langfuse supports self-hosting via Docker/local VM for testing and low-scale deployments, which effectively allows a separate sandbox environment distinct from Langfuse Cloud production, and datasets/experiments can be run without touching live production traces. However, there is no explicit documented 'sandbox mode' or staging environment feature, no test-data isolation guarantees, and no evidence of a first-party sandbox/test environment distinct from self-hosting workarounds. missing for 10: explicit sandbox/staging environment feature, documented test-data isolation from production, hands-on confirmation of safe sandbox testing workflow.

      • [claimed-docs] You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
      • [claimed-docs] Build a reusable set of test cases | Datasets
      • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
      Cekurapartialclaimed5/10

      Cekura explicitly supports testing voice agents via synthetic/simulated conversations 'before go-live' and running scenarios without external API keys, which implies a sandbox-like testing flow separate from live production calls. However, there is no explicit documentation of a dedicated sandbox environment, data isolation guarantees, or explicit statements that test runs never touch production data/logs. missing for 10: explicit sandbox/staging environment documentation, data isolation guarantees between test and production calls, independent confirmation that simulated runs are fully separated from production data.

      • [claimed-docs] Learn how to quickly create and test voice agents on Cekura without requiring any external API keys or integrations.
      • [claimed-docs] Run thousands of synthetic conversations before go-live.
      • [claimed-docs] Run the same scenarios across platforms and models. Pick the one that actually performs.
    4. ai-native userRely on versioned APIs with a documented deprecation policy

      weight 2 · round drawn
      Langfusenone0/10

      Evidence shows an API exists (e.g., 'Metrics API v2') but there is no documentation of a versioning scheme or deprecation policy; the OpenAPI spec probe even returned 404s across candidate paths, suggesting no discoverable API spec/versioning docs.

      • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
      • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
      • [probe] PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/opena…
      Cekuranone0/10

      There is an OpenAPI spec and API-key auth docs, but no evidence of API versioning scheme or a documented deprecation policy anywhere in the pack. missing for 10: versioning scheme documentation, deprecation policy, changelog/migration guides.

      • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
      • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.

    Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards

    Stories about alerting dashboards in this arena

    Monitoring

    1. ml engineerBuild custom dashboards over latency, error, cost, and eval-score metrics

      weight 2 · round to Langfuse
      Langfusefullclaimed8/10

      Docs explicitly state custom dashboards can be built over cost, latency, volume, and quality/eval-score metrics (docs-33, docs-38, docs-53), backed by a Metrics API v2 for querying aggregate cost, usage, latency, and score metrics (docs-15/28) and alerting on threshold crossings (docs-4, docs-53). Missing for 10: independent/hands-on confirmation of the dashboard-building UI experience and any evidence of error-rate-specific dashboard metrics (only cost/latency/volume/quality are explicitly named).

      • [claimed-docs] Create dashboards: monitor cost across models, tags, or users.
      • [claimed-docs] Analyze cost, latency, volume, and quality in custom dashboards
      • [claimed-docs] Set up alerts: get notified automatically when spend crosses a threshold.
      • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
      • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
      • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
      Cekurapartialclaimed6/10

      Cekura docs explicitly describe building custom dashboards with widgets to visualize call data, metrics, and metadata, which covers eval-score and general metric visualization for an ML engineer's needs. However, the docs don't explicitly enumerate latency, error rate, or cost as first-class dashboard metrics, and there's no independent/hands-on corroboration of dashboard flexibility. Missing for 10: explicit mention of latency/error/cost metric widgets, and third-party evidence of dashboard customization in practice.

      • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.
      • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
    2. developerSet alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email

      weight 2 · round to Langfuse
      Langfusepartialclaimed6/10

      Langfuse docs confirm alerting on cost spikes/thresholds and metric-based alerts delivered via Slack, GitHub Actions, or Webhooks, but there is no explicit mention of PagerDuty or email as notification channels, nor specific alerting tied to eval-score drops (only general metric thresholds). missing for 10: PagerDuty integration, email notification channel, explicit eval-score-drop alert examples, independent/hands-on confirmation of alerting reliability.

      • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
      • [claimed-docs] Set up alerts: get notified automatically when spend crosses a threshold.
      • [claimed-docs] Create dashboards: monitor cost across models, tags, or users.
      Cekuranone0/10

      Evidence covers dashboards, metrics, PII redaction, CI/CD, and cron jobs but no mention of alerting/notification integrations for error rates, cost spikes, or eval-score drops, nor any Slack/PagerDuty/email notification channels.

      Automation depth — how much of the product can run unattendedAutomation depth

      How much of the product can run unattended

      1. ai-native userPerform bulk operations across many items at once

        weight 2 · round to Cekura

        Langfuse docs describe multiple bulk-operation features — running experiments against datasets, exporting large trace volumes to blob storage on a schedule, querying aggregate metrics via API, and reviewing/annotating traces in bulk via annotation queues — which collectively support performing operations across many items. However, a hands-on community report states the trace/generation retrieval API was 'brutally slow for bulk scanning operations,' forcing the user to bypass the API and query the database directly, directly contradicting the claimed bulk-operation experience for API-based retrieval. missing for 10: independent verification that bulk export/experiments perform well at scale, resolution of the reported API slowness, and clearer first-party benchmarks for bulk throughput.

        • [claimed-docs] Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
        • [claimed-docs] Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
        • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
        • [claimed-docs] Review and rate traces manually | Annotation Queues, Scores via UI
        • [claimed-docs] Build a reusable set of test cases | Datasets
        • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
        Cekurapartialclaimed6/10

        Cekura's docs advertise running thousands of synthetic conversations/scenarios at once and provide a CLI/SDK to manage agents, scenarios, runs, and call data programmatically, which supports bulk test execution and automation via cron/CI. However, there's no explicit documentation of bulk edit/delete/update operations across many existing items (e.g., batch API endpoints for updating many scenarios or calls at once). Missing for 10: explicit bulk CRUD API/CLI commands operating on many items simultaneously, and any hands-on evidence of large-scale batch throughput.

        • [claimed-docs] Run thousands of synthetic conversations before go-live.
        • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
        • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
        • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
        • [claimed-docs] Learn how to set up GitHub Actions workflows to automatically test your agents with each change
      2. ai-native userDefine rules that trigger actions automatically on events

        weight 3 · round to Cekura
        Langfusepartialclaimed4/10

        Langfuse supports threshold-based alerts (Slack/GitHub Actions/Webhooks) and automated scoring of production traces via LLM-as-a-Judge, which are event-triggered automations, but there is no evidence of a general-purpose rule-engine where arbitrary user-defined conditions trigger arbitrary custom actions across the platform. missing for 10: user-defined conditional rule builder, broader action types beyond notifications/scoring (e.g. auto-remediation, workflow triggers), independent confirmation of alerting reliability.

        • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
        • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
        • [claimed-docs] Set up alerts: get notified automatically when spend crosses a threshold.
        • [claimed-docs] LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
        Cekurapartialclaimed6/10

        Cekura supports automation triggers like scheduled cron jobs for test workflows, GitHub Actions CI/CD hooks that fire tests on code changes, and an automatic flag→reproduce→suggest-fix pipeline triggered by detected issues, which are event-driven automations. However, there's no evidence of a general-purpose user-defined rules engine (e.g., 'if X metric drops, then do Y') beyond these fixed automation patterns. Missing for 10: a documented custom rule/condition builder for arbitrary events, and independent confirmation these automations behave as advertised.

        • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
        • [claimed-docs] Learn how to set up GitHub Actions workflows to automatically test your agents with each change
        • [claimed-docs] Cekura flags issues → reproduces in simulation → suggests fixes automatically.
        • [claimed-docs] Live drift detection across every call.
      3. ai-native userSchedule recurring jobs or workflows

        weight 2 · round to Cekura
        Langfusenone0/10

        Langfuse is an observability/evaluation platform for LLM apps; while it has scheduled exports and alerts, there is no evidence of user-defined recurring job/workflow scheduling (e.g., cron-like automation of arbitrary tasks) as an ai-native automation capability.

          Cekurafullclaimed7/10

          Cekura documents dedicated cron job support for scheduling automated testing/evaluation workflows (cekura-docs-13), plus GitHub Actions integration for triggering workflows on changes (cekura-docs-14), giving AI-native users concrete recurring automation options via CLI/SDK/API. Missing for 10: independent/hands-on confirmation that cron scheduling works reliably in production and more detail on scheduling granularity/limits.

          • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
          • [claimed-docs] Learn how to set up GitHub Actions workflows to automatically test your agents with each change
          • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
          • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
        • ai-native userVersion, review, and roll back my automations

          weight 1 · round to Langfuse
          Langfusepartialclaimed6/10

          Langfuse documents prompt version control and labels to manage deployments across environments, plus review mechanisms (annotation queues, experiments, LLM-as-judge) and linking prompts to trace performance — covering versioning and review well for prompt-based automations. However, there is no explicit 'rollback' feature described beyond implied re-labeling of a prior version, and no rollback capability for broader agent/workflow automations. missing for 10: explicit rollback/revert UI or API, evidence of rollback for full agent workflows (not just prompts), independent confirmation of rollback usage.

          • [claimed-docs] Use version control and labels to manage deployments across environments
          • [claimed-docs] Link prompts to traces to analyze performance by prompt version
          • [claimed-docs] When prompts live in Langfuse, non-technical team members update them directly in the UI while your application automatically fetches the la…
          • [claimed-docs] Review and rate traces manually | Annotation Queues, Scores via UI
          • [claimed-docs] Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments vi…
          • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
          Cekurapartialclaimed3/10

          Cekura supports comparing two versions of an agent side-by-side (A/B testing) and provides dashboards/metrics for reviewing agent performance, which touches 'version' and 'review' aspects, but there is no evidence of a rollback mechanism to revert an automation/agent to a prior version. Missing for 10: explicit version history/versioning system, rollback/undo capability, and any documented workflow for reverting agent configurations.

          • [claimed-docs] Compare two versions of your agent side-by-side to measure prompt, model, or configuration changes.
          • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.

        Cost monitoring — stories about cost monitoring in this arenaCost monitoring

        Stories about cost monitoring in this arena

        Cost tracking

        1. developerAttribute cost and usage to users, sessions, and features via custom metadata

          weight 2 · round to Langfuse
          Langfusefullclaimed7/10

          Langfuse docs describe cost/usage tracking with dashboards that break down spend 'across models, tags, or users' plus native session and user tracking, and a Metrics API to query aggregate cost/usage/latency by these dimensions. This directly supports attributing cost to users, sessions, and tag-based feature/metadata groupings. Missing for 10: explicit documentation of arbitrary custom metadata fields (beyond tags) as a first-class cost-attribution dimension, and independent/hands-on confirmation of this specific attribution workflow.

          • [claimed-docs] Create dashboards: monitor cost across models, tags, or users.
          • [claimed-docs] Support for tracking multi-turn conversations as sessions and user tracking
          • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
          • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
          • [claimed-docs] Analyze cost, latency, volume, and quality in custom dashboards
          • [claimed-docs] For most users, cost tracking works out of the box. Langfuse ships with prices for popular models from OpenAI, Anthropic, and Google
          Cekuranone0/10

          Cekura's docs show metadata being sent with calls and custom dashboards for visualizing metrics/metadata, but there is no evidence of cost tracking, cost attribution, or usage-cost breakdown by user/session/feature — the platform is framed purely around voice-agent evaluation and observability, not cost accounting.

          • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
          • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.
        2. developerSee cost and token usage per request, model, and time period in dashboards

          weight 3 · round to Langfuse
          Langfusefullcommunity9/10

          Langfuse docs explicitly describe out-of-the-box token/cost tracking, custom dashboards to monitor cost across models, tags, or users, and a Metrics API to query aggregate cost/usage/latency by time period, plus alerting on spend thresholds. Community feedback corroborates real-world use for tracking cost/latency at scale (e.g. $60k+/month LLM spend monitoring). Missing for 10: independent third-party verification of dashboard usability/accuracy beyond vendor docs and anecdotal community praise.

          • [claimed-docs] For most users, cost tracking works out of the box. Langfuse ships with prices for popular models from OpenAI, Anthropic, and Google
          • [claimed-docs] Create dashboards: monitor cost across models, tags, or users.
          • [claimed-docs] Langfuse ships with prices for popular models from OpenAI, Anthropic, and Google, and most integrations capture usage and cost automatically
          • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
          • [claimed-docs] Set up alerts: get notified automatically when spend crosses a threshold.
          • [claimed-docs] Analyze cost, latency, volume, and quality in custom dashboards
          • [community] Been using it. Happy customer. It gave me sanity into otherwise very complex LLM infrastructure. We spend 60k+ every month on LLM calls, so …
          Cekuranone0/10

          Cekura documents customizable dashboards for call data, metrics, and metadata (cekura-docs-11), but nothing in the evidence pack mentions cost or token usage tracking per request, model, or time period. As an observability platform for voice agents, this axis plausibly applies, but no evidence supports the specific cost/token dashboard capability.

          Data access export — stories about data access export in this arenaData access export

          Stories about data access export in this arena

          Data export

          1. developerBulk-export traces and datasets to blob storage or my data warehouse

            weight 2 · round to Langfuse
            Langfusefullcommunity8/10

            Langfuse explicitly documents a Blob Storage Export feature for exporting large volumes of traces/datasets on a schedule, plus Metrics API v2 and SDK-based querying for programmatic data access, supporting warehouse/blob-storage export workflows. Community evidence corroborates that self-hosted/open-source deployments allow direct DB access for bulk operations, though one user noted the bulk retrieval API can be slow. Missing for 10: independent hands-on verification specifically of blob storage export (vs. just docs) and more detail on supported destinations/formats.

            • [claimed-docs] Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
            • [claimed-docs] Export large volumes on a schedule | Blob Storage Export
            • [claimed-docs] Export large volumes on a schedule Blob Storage Export
            • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
            • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
            • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
            Cekuranone0/10

            Evidence shows Cekura has an API, CLI, SDK, and dashboards for accessing call data, but there is no mention of bulk-export functionality to blob storage (S3/GCS) or data warehouse integrations (Snowflake, BigQuery, etc). Missing for 10: bulk export feature, blob storage integration, data warehouse connectors/sinks, any documentation or API endpoint for exporting datasets in bulk.

            • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
            • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
            • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.

          Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets

          Measuring quality — datasets, eval runs, regression tracking

          Ai eval ops

          1. ai-native userHave an agent create a dataset, trigger an eval run programmatically, and read back the results

            weight 2 · round to Cekura
            Langfusepartialclaimed7/10

            Langfuse's docs describe the building blocks needed: Datasets for test cases, Experiments run via SDK to test versions programmatically, and Scores/Metrics accessible via API/SDK for reading back results (docs-12, docs-27, docs-42, docs-49), plus an MCP server for AI tools without shell access (docs-39/50). However, there's no explicit end-to-end example or documentation showing an agent specifically creating a dataset programmatically and then triggering+reading an eval run in one flow. missing for 10: an explicit walkthrough of an agent creating a dataset via SDK/API, triggering an eval, and reading results back; independent/hands-on confirmation of this specific workflow.

            • [claimed-docs] Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments vi…
            • [claimed-docs] Build a reusable set of test cases | Datasets
            • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
            • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
            • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
            • [claimed-docs] Connect an AI tool that cannot run shell commands MCP Server
            Cekurafullprobed8/10

            Cekura ships a CLI/SDK explicitly for managing agents, scenarios (datasets), runs, and call data from code, plus an official MCP server letting an AI assistant 'design, run, and improve' evaluations, backed by a documented OpenAPI spec and API-key auth for programmatic access — together covering dataset creation, triggering eval runs, and reading back results. missing for 10: independent/hands-on confirmation that an agent using the MCP/CLI actually completes the full create-dataset-to-trigger-run-to-read-results loop end-to-end.

            • [claimed-docs] Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations.
            • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
            • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
            • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
            • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
            • [probe] official MCP server documented at https://docs.cekura.ai/mcp/overview
            • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli

          Human review

          1. ml engineerRoute outputs to human annotation queues for review and labeling

            weight 2 · round to Langfuse
            Langfusefullclaimed8/10

            Langfuse docs explicitly describe Annotation Queues for manually reviewing and rating/labeling traces via the UI, tied into its broader evaluation/scoring system (numeric, categorical, boolean scores). Missing for 10: independent hands-on confirmation of annotation queue workflows and detail on routing/assignment logic for reviewers beyond the doc snippets.

            • [claimed-docs] Review and rate traces manually | Annotation Queues, Scores via UI
            • [claimed-docs] Review and rate traces manually | Annotation Queues
            • [claimed-docs] Review and rate traces manually Annotation Queues, Scores via UI
            • [claimed-docs] In Langfuse, that score can be numeric, categorical, or boolean.
            Cekuranone0/10

            Cekura's evidence pack covers voice-agent testing, evaluation metrics (LLM-judge, Python), dashboards, PII redaction, and observability, but there is no mention of routing outputs to human annotation queues, labeling workflows, or human-in-the-loop review pipelines. Absence of evidence for this applicable ML-eval capability yields none. Missing for 10: any mention of human annotation queues, labeling UI/workflow, or reviewer assignment features.

            Offline evals

            1. developerRun evals in CI and gate deployments on their results

              weight 2 · round to Langfuse
              Langfusefullclaimed7/10

              Langfuse explicitly documents a dedicated CI/CD experiments feature described as 'Block deploys on regressions' with a linked doc page (evaluation/experiments/experiments-ci-cd), directly matching the story of gating deployments on eval results, and this is paired with dataset-based experiments and code/LLM-as-judge evaluators for scoring runs. Missing for 10: independent/hands-on corroboration of the CI/CD gating workflow in practice and more detail on how the pass/fail threshold integrates with pipelines beyond the doc link.

              • [claimed-docs] Block deploys on regressions | [CI/CD experiments](/docs/evaluation/experiments/experiments-ci-cd)
              • [claimed-docs] Block deploys on regressions | CI/CD experiments
              • [claimed-docs] Block deploys on regressions CI/CD experiments
              • [claimed-docs] Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
              • [claimed-docs] Run deterministic checks | Code Evaluators
              • [claimed-docs] LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
              Cekurapartialclaimed6/10

              Cekura provides a dedicated GitHub Actions CI/CD guide plus CLI/SDK and cron job automation for testing agents on each change, indicating CI integration is a first-class capability. However, the evidence pack only shows the guide's title/description, not concrete details on exit-code/gating behavior for deployments, and there is no independent corroboration of this specific workflow. Missing for 10: explicit documentation of deployment-gating logic (pass/fail thresholds blocking merges/deploys), and independent/hands-on confirmation of the GitHub Actions integration.

              • [claimed-docs] Learn how to set up GitHub Actions workflows to automatically test your agents with each change
              • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
              • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
              • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
            2. ml engineerWrite custom code-based scorers and metrics for my evaluations

              weight 2 · round to Cekura
              Langfusefullclaimed7/10

              Langfuse's evaluation docs explicitly list 'Code Evaluators' for running deterministic/custom checks alongside LLM-as-a-Judge and API/SDK score submission, giving ML engineers a code-based path to define custom metrics and push scores via SDK. Missing for 10: deeper documentation/examples of the Code Evaluator API surface and independent hands-on confirmation of custom scorer authoring.

              • [claimed-docs] Run deterministic checks | Code Evaluators
              • [claimed-docs] Run deterministic checks Code Evaluators
              • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
              • [claimed-docs] In Langfuse, that score can be numeric, categorical, or boolean.
              • [claimed-docs] Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments vi…
              Cekurafullclaimed8/10

              Cekura explicitly documents Python metrics that let engineers 'write custom evaluation logic in Python to evaluate your AI agent's performance with complete control', alongside LLM-judge metrics as an alternative, giving ML engineers code-based scorer capability. missing for 10: no independent hands-on example of a custom Python metric in use, and no detail on SDK/CLI invocation of custom scorers beyond the docs page.

              • [claimed-docs] Write custom evaluation logic in Python to evaluate your AI agent's performance with complete control
              • [claimed-docs] Evaluate AI voice agent calls using natural language criteria with LLM-powered judging
              • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
            3. ml engineerCompare eval runs side by side to catch regressions between prompt or model versions

              weight 3 · round to Langfuse
              Langfusefullclaimed8/10

              Langfuse's Evaluation feature explicitly supports comparing prompt, model, or code changes side by side via Experiments (UI and SDK), and can block deploys on regressions via CI/CD experiments, directly matching the story of catching regressions between versions. Datasets, LLM-as-a-Judge, and code evaluators provide the scoring backbone, and prompt-version linking to traces supports drill-down analysis. Missing for 10: independent hands-on validation of the side-by-side comparison UI specifically, and no community evidence directly confirming regression-catching workflows in practice.

              • [claimed-docs] Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments vi…
              • [claimed-docs] Block deploys on regressions | [CI/CD experiments](/docs/evaluation/experiments/experiments-ci-cd)
              • [claimed-docs] Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
              • [claimed-docs] Build a reusable set of test cases | Datasets
              • [claimed-docs] Link prompts to traces to analyze performance by prompt version
              • [claimed-docs] LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
              • [claimed-docs] Run deterministic checks | Code Evaluators
              Cekurapartialclaimed6/10

              Cekura has an explicit A/B testing feature to compare two agent versions side-by-side to measure prompt, model, or configuration changes, plus dashboards for visualizing metrics/call data, which supports regression detection between versions. However, the story specifically frames this as an ML-engineer 'eval run' comparison workflow, and evidence is limited to a single docs page description without detail on run history diffing, metric-level regression flags, or independent hands-on confirmation. Missing for 10: detailed UI/workflow evidence of comparing full eval-run history (not just two agent versions), regression-flagging mechanics, and independent/community corroboration of the comparison feature itself.

              • [claimed-docs] Compare two versions of your agent side-by-side to measure prompt, model, or configuration changes.
              • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.
              • [claimed-docs] Run the same scenarios across platforms and models. Pick the one that actually performs.
            4. ml engineerScore outputs with configurable LLM-as-a-judge evaluators

              weight 3 · round drawn
              Langfusefullclaimed8/10

              Langfuse's docs explicitly document LLM-as-a-Judge evaluators that produce numeric, categorical, or boolean scores, can run automatically on live production traces, and can be configured against datasets/experiments; this is a first-party, dedicated feature rather than a workaround. Missing for 10: independent/hands-on confirmation of configuring custom judge prompts or judge model selection, and no community evidence specifically validating LLM-as-a-judge quality in practice.

              • [claimed-docs] LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
              • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
              • [claimed-docs] In Langfuse, that score can be numeric, categorical, or boolean.
              • [claimed-docs] Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments vi…
              • [claimed-docs] Build a reusable set of test cases | Datasets
              Cekurafullclaimed8/10

              Cekura documents an LLM-judge metric type that evaluates calls using natural language criteria, alongside custom Python metrics for full control, giving configurable LLM-as-a-judge evaluators for scoring agent outputs. This is corroborated by dashboards/A-B testing docs showing these metrics feed into broader evaluation workflows. Missing for 10: no independent hands-on benchmark of judge accuracy/configurability, and no detail on rubric customization depth or multi-model judge support.

              • [claimed-docs] Evaluate AI voice agent calls using natural language criteria with LLM-powered judging
              • [claimed-docs] Write custom evaluation logic in Python to evaluate your AI agent's performance with complete control
              • [claimed-docs] Compare two versions of your agent side-by-side to measure prompt, model, or configuration changes.
            5. ml engineerCurate datasets from production traces and run offline evaluations against them

              weight 3 · round to Langfuse
              Langfusefullclaimed8/10

              Langfuse docs explicitly support building reusable datasets ('Build a reusable set of test cases | Datasets') and running offline evaluations against them via Experiments (UI/SDK/CI-CD), plus LLM-as-a-Judge and code evaluators, and traces from production can feed datasets. This directly matches the story of curating datasets from production traces and running offline evals. missing for 10: explicit documented workflow step showing 'add trace to dataset' UI action, and independent/hands-on community corroboration specifically of the dataset-curation-from-traces workflow (community evidence covers observability/prompt mgmt broadly but not this specific feature).

              • [claimed-docs] Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
              • [claimed-docs] Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments vi…
              • [claimed-docs] Block deploys on regressions | [CI/CD experiments](/docs/evaluation/experiments/experiments-ci-cd)
              • [claimed-docs] Run Experiments against datasets to test new prompt versions directly within Langfuse
              • [claimed-docs] Run deterministic checks | Code Evaluators
              • [claimed-docs] Build a reusable set of test cases | Datasets
              • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
              • [claimed-docs] Run deterministic checks Code Evaluators
              Cekurapartialclaimed5/10

              Cekura stores production call data as CallLog objects via its observability API and lets teams run metric evaluations (Python/LLM-judge) and scenario-based tests, which supports building and evaluating datasets from real traces, but there is no explicit documentation of a 'dataset curation' workflow (filtering/tagging/versioning production traces into an eval set) or an offline batch-eval mode distinct from live/simulated runs. missing for 10: dedicated dataset-curation/versioning feature for production traces, explicit offline-evaluation-on-curated-dataset workflow, independent evidence of this specific use case.

              • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
              • [claimed-docs] Write custom evaluation logic in Python to evaluate your AI agent's performance with complete control
              • [claimed-docs] Evaluate AI voice agent calls using natural language criteria with LLM-powered judging
              • [claimed-docs] Learn how to set up and manage automated cron jobs for your AI agent testing and evaluation workflows.
              • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.

            Online evals

            1. ml engineerRun evaluators continuously on live production traffic, not just offline datasets

              weight 2 · round to Cekura
              Langfusefullclaimed7/10

              Langfuse docs explicitly describe automatically scoring live production traces via LLM-as-a-Judge and Scores API/SDK, plus code evaluators for deterministic checks and alerting when metrics cross thresholds — directly matching continuous evaluation on live traffic, not just offline datasets. missing for 10: independent/hands-on confirmation that continuous production scoring works reliably at scale, and more detail on how evaluators are scheduled/triggered continuously.

              • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
              • [claimed-docs] LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
              • [claimed-docs] Run deterministic checks | Code Evaluators
              • [claimed-docs] Run deterministic checks Code Evaluators
              • [claimed-docs] Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
              • [claimed-docs] In Langfuse, that score can be numeric, categorical, or boolean.
              Cekurafullprobed8/10

              Cekura's observability pipeline explicitly ingests live production calls via webhook/API, stores them as CallLogs, and 'schedules metric evaluation' automatically, with marketing copy specifically citing 'Live drift detection across every call' as a continuous production-monitoring feature distinct from offline simulation runs. missing for 10: independent/hands-on confirmation of evaluator behavior on live traffic at scale (only vendor docs/marketing cited).

              • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
              • [claimed-docs] Live drift detection across every call.
              • [claimed-docs] Automatically detect and remove sensitive information from transcripts and audio recordings
              • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
              • [probe] PROBE llms.txt: HTTP 200 at https://docs.cekura.ai/llms.txt # Cekura > Cekura is the testing and observability platform for voice AI agents…

            Openness — open source, data portability, and self-hosting storiesOpenness

            Open source, data portability, and self-hosting stories

            1. ai-native userDo everything through the API that I can do in the UI

              weight 2 · round to Cekura
              Langfusepartialprobed6/10

              Langfuse exposes broad API/SDK coverage for core workflows — tracing (OTLP endpoint, SDKs), Metrics API v2, Scores via API/SDK, prompt management and experiments via SDK, and blob storage export — indicating most UI actions have API equivalents. However, some features are documented as UI-first (annotation queues, LLM Playground, dashboard creation) without explicit API parity confirmation, and the community reports the retrieval API was 'brutally slow for bulk scanning', suggesting API usability gaps. missing for 10: explicit API endpoints/docs for annotation queues and playground actions, a published OpenAPI/swagger reference (probe found only 404s), and independent confirmation that all UI actions are fully API-mirrored.

              • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
              • [claimed-docs] Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK
              • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
              • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
              • [claimed-docs] Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
              • [claimed-docs] Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
              • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
              • [probe] PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/opena…
              Cekurapartialprobed7/10

              Cekura ships a full OpenAPI spec, CLI, and SDK that cover core workflows—managing agents, scenarios, runs, call data, sending calls, and metrics—giving strong programmatic parity with the UI (cekura-docs-4, cekura-docs-5, cekura-docs-6, cekura-probe-3, cekura-probe-5). However, some UI-specific features like custom dashboard building (cekura-docs-11) and cron/CI setup aren't clearly confirmed as fully API-driven rather than UI-configured. Missing for 10: explicit confirmation that dashboard creation and all UI-configurable settings are exposed via API/CLI, and independent verification of complete feature parity.

              • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
              • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
              • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
              • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
              • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
              • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli
              • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.
            2. ai-native userExport all of my data in open formats and leave

              weight 3 · round to Langfuse
              Langfusepartialcommunity6/10

              Langfuse supports data export via Blob Storage Export, Metrics API, SDK queries, and is fully open-source/self-hostable (Docker, Kubernetes, Terraform), giving strong open-format export and no lock-in. However, evidence doesn't detail specific export formats (e.g., JSON/CSV/Parquet) or confirm full-fidelity bulk export of all trace/prompt/eval data, and community feedback notes the trace retrieval API was slow for bulk scanning, requiring direct DB queries as a workaround. missing for 10: explicit documentation of exported data formats, confirmation of complete data portability across all entities (traces, prompts, scores, datasets), and independent verification of clean/reliable bulk export at scale.

              • [claimed-docs] Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
              • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
              • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
              • [claimed-docs] When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
              • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
              • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
              Cekuranone0/10

              Cekura offers CLI/SDK/API access to call data and metrics, but there is no evidence of a bulk data export feature in open/portable formats or any account-closure/data-portability guarantee for users wishing to leave the platform.

              • ai-native userRead the product's source under an open license

                weight 2 · round to Langfuse
                Langfusefullprobed8/10

                Langfuse explicitly documents that it is open-source and self-hostable, with source on GitHub, and community members confirm using 'Langfuse OSS' and querying the underlying DB directly since it's open source. Missing for 10: no explicit statement of the specific license type (e.g. MIT/Apache) in the evidence pack, and no independent audit confirming license terms.

                • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
                • [claimed-docs] When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
                • [probe] PROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](htt…
                • [probe] PROBE docs-md: HTTP 200 at https://langfuse.com/docs.md --- title: Overview seoTitle: Open Source AI Engineering Platform description: Langf…
                • [community] Been using Langfuse OSS for almost 15 months from the start. By far the best solution. No dark patterns found in other projects such as Port…
                • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
                Cekuranone0/10

                Cekura is a closed SaaS platform with a GitHub repo only for a skills/plugin package (cekura-skills), not the product's core source code, and no evidence of an open-source license for the platform itself. No license file or open-source repo for the core product is mentioned anywhere in the evidence.

                • ai-native userSelf-host the core product

                  weight 3 · round to Langfuse
                  Langfusefullprobed9/10

                  Langfuse is explicitly open source and documents multiple self-hosting paths (Docker Compose for local/testing, and production-scale Kubernetes/Helm, AWS/Azure/GCP Terraform), noting the self-hosted stack runs the same infrastructure as Langfuse Cloud, and this is corroborated by community users citing long-term OSS self-hosted usage. Missing for 10: independent hands-on verification of a full self-hosted deployment matching cloud parity, and no direct comment on self-hosting from third parties beyond general OSS praise.

                  • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
                  • [claimed-docs] When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
                  • [claimed-docs] Kubernetes (Helm) ... AWS (Terraform) ... Azure (Terraform) ... GCP (Terraform)
                  • [claimed-docs] Kubernetes (Helm) | AWS (Terraform) | Azure (Terraform) | GCP (Terraform)
                  • [claimed-docs] You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
                  • [claimed-docs] For production-scale deployments, we recommend one of the following options: Kubernetes (Helm), AWS (Terraform), Azure (Terraform), GCP (Ter…
                  • [community] Been using Langfuse OSS for almost 15 months from the start. By far the best solution. No dark patterns found in other projects such as Port…
                  • [probe] PROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](htt…
                  Cekuranone0/10

                  Cekura is presented as a hosted SaaS platform (API keys, cloud dashboards, webhooks) with no mention of a self-hosted or on-premises deployment option anywhere in the docs, CLI, SDK, or website copy. No evidence of Docker images, self-hosting instructions, or open-source core.

                  • [claimed-docs] Create a new API key. When using API key authentication, only project-level API keys can be created.
                  • [claimed-docs] create a comprehensive plan to integrate cekura observability
                  • [probe] PROBE llms.txt: HTTP 200 at https://docs.cekura.ai/llms.txt # Cekura > Cekura is the testing and observability platform for voice AI agents…

                Privacy posture — data-handling and privacy storiesPrivacy posture

                Data-handling and privacy stories

                1. ai-native userChoose where my data is stored (region/residency)

                  weight 2 · round to Langfuse

                  Langfuse's self-hosting docs (Docker, Kubernetes/Helm, AWS/Azure/GCP Terraform) let a user deploy the platform anywhere, which is the primary mechanism for choosing data region/residency. However, community evidence directly contradicts the vendor's EU/GDPR data-residency marketing for Langfuse Cloud, stating that after the ClickHouse-related change the Cloud offering is 'no longer GDPR compliant' since ClickHouse is US-headquartered — a concrete counter-example to the residency claim for the hosted product. Missing for 10: explicit region-selection settings within Langfuse Cloud itself, first-party documentation addressing the GDPR/residency gap, and independent confirmation that self-hosting fully resolves residency for non-technical users.

                  • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
                  • [claimed-docs] When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
                  • [claimed-docs] Kubernetes (Helm) ... AWS (Terraform) ... Azure (Terraform) ... GCP (Terraform)
                  • [claimed-docs] Kubernetes (Helm) | AWS (Terraform) | Azure (Terraform) | GCP (Terraform)
                  • [claimed-docs] You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
                  • [claimed-docs] For production-scale deployments, we recommend one of the following options: Kubernetes (Helm), AWS (Terraform), Azure (Terraform), GCP (Ter…
                  • [community] Very sad, for all their marketing around EU, GDPR, privacy... Since clickhouse is headquartered in the US that means the langfuse cloud is n…
                  Cekuranone0/10

                  No evidence anywhere in the pack mentions data residency, region selection, or geographic storage options for Cekura; PII redaction is mentioned but that's a different concern than data location choice.

                  • ai-native userPrevent my data from being used to train AI models

                    weight 3 · round to Langfuse
                    Langfusepartialclaimed3/10

                    Langfuse does not publish any explicit statement about not using customer data to train AI models, but its self-hosting option (Docker, Kubernetes, Terraform for AWS/Azure/GCP) means an org can keep all trace/prompt data entirely within its own infrastructure, which indirectly prevents that data from reaching Langfuse's own systems or being used for model training. Missing for 10: explicit data-usage/training policy, DPA or privacy documentation addressing model training, and independent confirmation of this stance for Langfuse Cloud users.

                    • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
                    • [claimed-docs] When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
                    • [claimed-docs] You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
                    • [claimed-docs] For production-scale deployments, we recommend one of the following options: Kubernetes (Helm), AWS (Terraform), Azure (Terraform), GCP (Ter…
                    Cekuranone0/10

                    Evidence shows PII redaction features for transcripts but nothing about opting out of AI model training on customer data, data-use policies, or training-data controls.

                    • ai-native userControl data retention and deletion

                      weight 2 · round to Langfuse
                      Langfusepartialcommunity3/10

                      Langfuse offers self-hosting (Docker, Kubernetes, Terraform) which gives users infrastructure-level control over their own data, implying they can manage retention/deletion themselves, but there is no explicit documentation of a retention policy setting, deletion API, or data lifecycle management feature in the evidence pack. Missing for 10: explicit retention/deletion API or UI controls, documented data retention policy, GDPR-specific deletion tooling, and any independent confirmation that these controls work as described.

                      • [claimed-docs] Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
                      • [claimed-docs] When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
                      • [claimed-docs] You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
                      • [claimed-docs] For production-scale deployments, we recommend one of the following options: Kubernetes (Helm), AWS (Terraform), Azure (Terraform), GCP (Ter…
                      • [community] Very sad, for all their marketing around EU, GDPR, privacy... Since clickhouse is headquartered in the US that means the langfuse cloud is n…
                      Cekuranone0/10

                      Cekura documents PII redaction to remove sensitive info from transcripts, but there is no evidence of user-facing data retention controls or deletion policies/APIs (e.g., data export, retention windows, right-to-delete). missing for 10: documented retention policy, deletion/export mechanism, user control over stored call data lifecycle.

                      • [claimed-docs] Automatically detect and remove sensitive information from transcripts and audio recordings
                    • ai-native userOpt out of telemetry and usage tracking

                      weight 2 · round drawn
                      Langfusenone0/10

                      The evidence pack covers Langfuse's LLM tracing/observability features and self-hosting deployment options, but contains no mention of Langfuse's own product telemetry (e.g., anonymized usage stats phoned home from self-hosted instances) or any documented opt-out mechanism (like a TELEMETRY_ENABLED flag). Absence of evidence for this applicable capability yields none.

                        Cekuranone0/10

                        No evidence of any telemetry opt-out or usage tracking controls; Cekura's docs focus on PII redaction for customer call data, not on the vendor's own product telemetry. Missing for 10: any mention of telemetry/analytics collection, opt-out mechanism, or privacy settings for usage data.

                        Prompt management — stories about prompt management in this arenaPrompt management

                        Stories about prompt management in this arena

                        Prompt workflow

                        1. developerIterate on prompts in a playground against real models and variables

                          weight 2 · round to Langfuse
                          Langfusefullcommunity7/10

                          Langfuse explicitly documents an LLM Playground for interactively testing prompts, plus experiment/versioning tooling that supports iteration against real models (docs-7, docs-20, docs-8, docs-21, docs-9). Missing for 10: detailed documentation on variable injection specifics in the playground and independent hands-on corroboration; one community comment even notes prompt management felt 'behind' competitors (langfuse-comm-11).

                          • [claimed-docs] Test prompts interactively in the [LLM Playground](/docs/prompt-management/features/playground)
                          • [claimed-docs] Test prompts interactively in the LLM Playground
                          • [claimed-docs] Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
                          • [claimed-docs] Run Experiments against datasets to test new prompt versions directly within Langfuse
                          • [claimed-docs] Link prompts to traces to analyze performance by prompt version
                          • [community] Langfuse seemed way better at tracing though [vs PromptLayer]... but Langfuse's prompt management was pretty behind in terms of features whe…
                          Cekuranone0/10

                          Cekura's docs describe testing/evaluation workflows (scenario runs, A/B testing agent versions, synthetic conversations) but there is no evidence of an interactive playground where a developer iterates on prompt text against real models with variables in real time — the closest items (ab-testing, testing-agents/overview) describe comparing pre-built agent versions, not a live prompt-editing sandbox.

                          • [claimed-docs] Compare two versions of your agent side-by-side to measure prompt, model, or configuration changes.
                          • [claimed-docs] Learn how to quickly create and test voice agents on Cekura without requiring any external API keys or integrations.
                          • [claimed-docs] Run the same scenarios across platforms and models. Pick the one that actually performs.

                        Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation

                        Instrumenting code and tracing requests end to end

                        Ai observability

                        1. ai-native userHave an agent query my traces, metrics, and eval results through an API or MCP server to debug my app

                          weight 3 · round drawn
                          Langfusepartialprobed6/10

                          Langfuse offers both API access (Metrics API v2, trace/score APIs, SDKs) and a documented MCP Server for connecting AI tools (langfuse-docs-15/28/49/39/50, probe-4), giving an agent programmatic access to traces, metrics, and evals. However, the MCP server's headline documented use case is auto-instrumenting code via Cursor, not querying data for debugging (langfuse-docs-19/36), and one hands-on report notes the trace retrieval API was too slow for bulk scanning, requiring a DB workaround (langfuse-comm-8). Missing for 10: explicit documentation/examples of an agent using MCP or API specifically to pull eval results and metrics for debugging workflows, and independent corroboration that this querying pattern works smoothly at scale.

                          • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
                          • [claimed-docs] Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
                          • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
                          • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
                          • [claimed-docs] Connect an AI tool that cannot run shell commands MCP Server
                          • [claimed-docs] Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
                          • [claimed-docs] Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
                          • [probe] official MCP server documented at https://langfuse.com/docs/docs-mcp
                          • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
                          Cekurapartialprobed6/10

                          Cekura has documented API access (OpenAPI spec, send-calls, SDK/CLI for agents/scenarios/runs/call data) and an official MCP server, but the MCP server's documented purpose is to 'design, run, and improve voice-agent evaluations' rather than explicitly querying traces/metrics/eval results for debugging — no direct evidence the MCP tools expose trace/metric/eval query endpoints for agent-driven debugging. missing for 10: explicit MCP tool list showing trace/metric/eval query capabilities, independent hands-on confirmation of an agent using MCP or API to debug via traces/metrics/evals.

                          • [claimed-docs] Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations.
                          • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
                          • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
                          • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
                          • [probe] PROBE openapi: HTTP 200 at https://docs.cekura.ai/openapi.json — contains "openapi" key
                          • [probe] official MCP server documented at https://docs.cekura.ai/mcp/overview
                          • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli

                        Data controls

                        1. developerMask or redact sensitive data before it is stored in traces

                          weight 2 · round to Cekura
                          Langfusenone0/10

                          No evidence in the pack mentions masking, redaction, or PII scrubbing before trace storage; the docs cover tracing, prompt management, evaluation, and deployment but not data masking capabilities.

                            Cekurafullclaimed8/10

                            Cekura has a dedicated PII redaction feature that automatically detects and removes sensitive information from transcripts and audio recordings before/at storage, directly addressing masking/redaction in traces. missing for 10: no independent/hands-on corroboration of redaction accuracy or configurability details beyond the doc summary.

                            • [claimed-docs] Automatically detect and remove sensitive information from transcripts and audio recordings
                            • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…

                          Sdk coverage

                          1. developerInstrument apps in both Python and JS/TS with officially supported SDKs

                            weight 2 · round to Langfuse
                            Langfusefullcommunity8/10

                            Langfuse docs explicitly state native SDKs for both Python and JS are available for capturing traces (langfuse-docs-5), plus SDK-based API querying for Python/JS (langfuse-docs-49) and general community validation of strong tracing/DX quality (langfuse-comm-4). Missing for 10: no direct evidence of separate SDK repos/versioning details or independent hands-on confirmation specifically of both SDKs' feature parity.

                            • [claimed-docs] Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
                            • [claimed-docs] Use the API from Python or JS/TS Query via SDKs
                            • [community] Langfuse is by far the best of the langs (-chain, -graph, -smith, -flow) in terms of UI/DX/integration/docs/quality.
                            Cekurapartialprobed3/10

                            Cekura documents a general SDK ('sync and async clients for application code') and a Python-specific custom metric guide, implying Python SDK support, but there is no explicit mention of an officially supported JS/TS SDK anywhere in the evidence pack. missing for 10: explicit JS/TS SDK documentation, parity of features across languages, and confirmation the generic SDK is multi-language rather than Python-only.

                            • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
                            • [claimed-docs] Write custom evaluation logic in Python to evaluate your AI agent's performance with complete control
                            • [probe] official CLI documented at https://docs.cekura.ai/cli-sdk/cli

                          Trace capture

                          1. developerTrace multi-step agent runs as nested spans grouped into sessions or threads

                            weight 2 · round to Langfuse
                            Langfusefullcommunity8/10

                            Langfuse docs describe traces capturing all LLM and non-LLM calls (implying nested spans), session/multi-turn conversation tracking, and agent graph visualization for complex agentic workflows, which directly matches the story. Community evidence corroborates real-world use for tracing complex LLM infrastructure and debugging (comm-7, comm-10), though one user noted retries complicate trace collapsibility (comm-9), a minor edge case rather than a core failure. Missing for 10: explicit documentation of 'thread' terminology and a deeper technical walkthrough of nested span hierarchy structure.

                            • [claimed-docs] Traces include all LLM and non-LLM calls, including retrieval, embedding, API calls, and more
                            • [claimed-docs] Support for tracking multi-turn conversations as sessions and user tracking
                            • [claimed-docs] LLM agents can be visualized as a graph to illustrate the flow of complex agentic workflows.
                            • [community] Been using it. Happy customer. It gave me sanity into otherwise very complex LLM infrastructure. We spend 60k+ every month on LLM calls, so …
                            • [community] One user described the difference to be that they use observability usually to know that nothing is going wrong whereas they use Langfuse ma…
                            • [community] retries for instructor like structured outputs mess up the traces, i wonder if they can be tracked and collapsible
                            Cekuranone0/10

                            Cekura's observability evidence centers on capturing call-level logs (CallLog with transcript, recording, metadata) and LiveKit SDK tracing, plus dashboards and PII redaction, but nothing in the evidence describes a hierarchical trace model with nested spans grouped into sessions or threads for multi-step agent runs. Missing for 10: explicit span/trace data model, parent-child span nesting, and session/thread grouping documentation.

                            • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
                            • [claimed-docs] Enhanced observability for LiveKit agents using the Cekura SDK
                            • [claimed-docs] Build custom dashboards with widgets to visualize and analyze your call data, metrics, and metadata.
                          2. developerInstrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK

                            weight 2 · round to Langfuse
                            Langfusefullcommunity7/10

                            Langfuse docs state native SDKs plus "100+ library/framework integrations," OpenTelemetry support, and LLM Gateway compatibility, and community reviews praise its tracing/integration quality relative to competitors. However, the pack never explicitly names LangChain, OpenAI SDK, or Vercel AI SDK integrations, only the general '100+' claim. Missing for 10: explicit citation naming LangChain, OpenAI SDK, and Vercel AI SDK integrations specifically, plus independent hands-on confirmation of setup ease for each.

                            • [claimed-docs] Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
                            • [claimed-docs] Langfuse can receive traces on the `/api/public/otel` (OTLP) endpoint.
                            • [claimed-docs] Langfuse can receive traces on the /api/public/otel (OTLP) endpoint.
                            • [community] Langfuse is by far the best of the langs (-chain, -graph, -smith, -flow) in terms of UI/DX/integration/docs/quality.
                            • [community] Langfuse seemed way better at tracing though [vs PromptLayer]... but Langfuse's prompt management was pretty behind in terms of features whe…
                            Cekurapartialclaimed3/10

                            Cekura documents a generic observability API/webhook (send-calls) and SDK/CLI that any agent could POST to, plus a named integration for LiveKit voice agents, but there is no documented integration for LangChain, the OpenAI SDK, or the Vercel AI SDK specifically as the story requests. missing for 10: named LangChain integration, named OpenAI SDK integration, named Vercel AI SDK integration.

                            • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
                            • [claimed-docs] Enhanced observability for LiveKit agents using the Cekura SDK
                            • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
                          3. developerCapture multimodal payloads (images, audio, files) inside my traces

                            weight 1 · round to Cekura
                            Langfusenone0/10

                            No evidence pack item mentions capturing images, audio, or file attachments in traces; the closest docs only reference generic LLM/non-LLM call tracing, not multimodal payload support.

                              Cekurapartialclaimed5/10

                              Cekura's observability pipeline explicitly captures audio (recording URL) and transcript/metadata as part of CallLog traces, and PII redaction applies to both transcripts and audio recordings, showing audio is a first-class multimodal payload. However, there is no evidence of support for images or generic file attachments within traces — the product is voice-agent focused and docs never mention image/file capture. Missing for 10: documented support for image payloads, documented support for arbitrary file attachments in traces.

                              • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
                              • [claimed-docs] Automatically detect and remove sensitive information from transcripts and audio recordings
                            • developerSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary format

                              weight 3 · round to Langfuse
                              Langfusefullclaimed8/10

                              Langfuse documents a native OTLP endpoint (/api/public/otel) for receiving traces alongside its SDKs, explicitly supporting OpenTelemetry as an alternative to proprietary SDK-based capture. missing for 10: no evidence of exporting/sending traces out via OTLP (only receiving), and no independent/hands-on corroboration of OTLP interoperability.

                              • [claimed-docs] Langfuse can receive traces on the `/api/public/otel` (OTLP) endpoint.
                              • [claimed-docs] Langfuse can receive traces on the /api/public/otel (OTLP) endpoint.
                              • [claimed-docs] Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
                              Cekuranone0/10

                              Cekura's observability ingestion uses a proprietary POST endpoint (transcript, recording URL, metadata) and its own API/CLI/SDK, with no mention of OpenTelemetry or OTLP support anywhere in the evidence pack.

                              • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
                              • [claimed-docs] Programmatic access to Cekura — sync and async clients for application code
                              • [claimed-docs] Manage agents, scenarios, runs, and call data from your terminal
                            • developerCapture traces of my LLM calls with inputs, outputs, latency, and token usage

                              weight 3 · round to Langfuse
                              Langfusefullcommunity9/10

                              Langfuse's core product is LLM tracing capturing inputs/outputs/latency/token usage via native SDKs, OTel, and 100+ integrations, with automatic cost tracking based on token usage, and community testimonials confirm real-world use for debugging cost/latency/quality at scale. Missing for 10: independent hands-on verification of trace fidelity (one community comment notes slow bulk retrieval API and issues with retries messing up traces, minor secondary caveats).

                              • [claimed-docs] Traces include all LLM and non-LLM calls, including retrieval, embedding, API calls, and more
                              • [claimed-docs] Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
                              • [claimed-docs] For most users, cost tracking works out of the box. Langfuse ships with prices for popular models from OpenAI, Anthropic, and Google
                              • [claimed-docs] Langfuse ships with prices for popular models from OpenAI, Anthropic, and Google, and most integrations capture usage and cost automatically
                              • [community] Been using it. Happy customer. It gave me sanity into otherwise very complex LLM infrastructure. We spend 60k+ every month on LLM calls, so …
                              • [community] One user described the difference to be that they use observability usually to know that nothing is going wrong whereas they use Langfuse ma…
                              • [community] The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But th…
                              Cekurapartialprobed4/10

                              Cekura's observability captures call-level data (transcript, recording URL, metadata) via its send-calls API and stores it as a CallLog, and offers 'enhanced observability' for LiveKit agents, but the docs never explicitly mention capturing per-LLM-call inputs/outputs, latency, or token usage — the focus is on transcript/audio-level call logging for voice-agent evaluation rather than granular LLM tracing. Missing for 10: explicit documentation of input/output payload capture per LLM call, latency metrics, and token usage tracking.

                              • [claimed-docs] Your agent (or its provider's post-call webhook) POSTs here with the transcript, recording URL, and metadata; Cekura stores it as a `CallLog…
                              • [claimed-docs] Enhanced observability for LiveKit agents using the Cekura SDK
                              • [probe] official MCP server documented at https://docs.cekura.ai/mcp/overview

                            Not comparable on these axes

                            1. ai-native userPlug MCP servers into this product so it can use their tools

                              weight 3 · not comparable
                              Langfusen/a

                              Langfuse is an observability/tracing and prompt-management platform, not an agent with a tool-use loop; the evidence shows Langfuse ships an MCP *server* so other coding agents (e.g. Cursor) can call Langfuse's own tools, which is the opposite role from the story's request that Langfuse itself consume external MCP servers' tools. This client-side MCP-consumption capability is a category mismatch for this kind of product.

                              • [claimed-docs] Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
                              • [claimed-docs] Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
                              • [claimed-docs] Connect an AI tool that cannot run shell commands | MCP Server
                              • [claimed-docs] Connect an AI tool that cannot run shell commands MCP Server
                              • [probe] official MCP server documented at https://langfuse.com/docs/docs-mcp
                              Cekuranone0/10

                              All Cekura MCP evidence describes Cekura exposing its own MCP server so external AI assistants can call Cekura's tools (docs-3, probe-4), not Cekura itself acting as an MCP client that consumes third-party MCP servers' tools. No evidence shows Cekura can plug in and use external MCP servers.

                              • [claimed-docs] Install Cekura Skills and MCP so your AI assistant can design, run, and improve voice-agent evaluations.
                              • [probe] official MCP server documented at https://docs.cekura.ai/mcp/overview
                            2. developerVersion prompts and deploy changes to production without shipping code

                              weight 3 · not comparable
                              Langfusefullcommunity8/10

                              Langfuse's docs directly support this story: prompts are managed centrally with version control and labels for deployments across environments (docs-10), non-technical users can update prompts in the UI without code changes while apps auto-fetch the latest version (docs-6/22), and SDK-side caching ensures no latency penalty (docs-23/41). There's one HN comment noting prompt management was 'behind' compared to a competitor a couple months prior (langfuse-comm-11), a mild caveat but not a concrete failure of the core capability. Missing for 10: independent hands-on verification of production deployment workflows and rollback mechanisms beyond docs claims.

                              • [claimed-docs] non-technical team members update them directly in the UI while your application automatically fetches the latest version
                              • [claimed-docs] Use version control and labels to manage deployments across environments
                              • [claimed-docs] When prompts live in Langfuse, non-technical team members update them directly in the UI while your application automatically fetches the la…
                              • [claimed-docs] Langfuse Prompt Management adds no latency to your application. Prompts are cached client-side by the SDK, so retrieving them is as fast as …
                              • [claimed-docs] Langfuse Prompt Management adds no latency to your application. Prompts are cached client-side by the SDK
                              • [community] Langfuse seemed way better at tracing though [vs PromptLayer]... but Langfuse's prompt management was pretty behind in terms of features whe…
                              Cekuran/a

                              Cekura is a testing/observability platform for voice AI agents (simulated calls, metrics, A/B comparison of agent versions), not a prompt-management or deployment system. Prompt versioning and production deployment without code changes belongs to the underlying agent-building platform, not to a third-party evaluation/observability tool like Cekura, making this axis a category error for this product type.

                              • [claimed-docs] Compare two versions of your agent side-by-side to measure prompt, model, or configuration changes.
                              • [probe] PROBE llms.txt: HTTP 200 at https://docs.cekura.ai/llms.txt # Cekura > Cekura is the testing and observability platform for voice AI agents…