Rank #2 of 8 in LLM Evals & Observability
Try itExperimental
See what an agent can do with Cekura before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$uvx --from 'cekura[cli]' cekura --helprecorded session — replayed, not liveVerified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboardsevidence →
Stories about alerting dashboards in this arena
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Cost monitoring — stories about cost monitoring in this arenaCost monitoringevidence →
Stories about cost monitoring in this arena
Data access export — stories about data access export in this arenaData access exportevidence →
Stories about data access export in this arena
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasetsevidence →
Measuring quality — datasets, eval runs, regression tracking
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Prompt management — stories about prompt management in this arenaPrompt managementevidence →
Stories about prompt management in this arena
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentationevidence →
Instrumenting code and tracing requests end to end
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Cekura stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Webhooks · Versioning policy · Full data export
Subscribe to events via webhooks
—0/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
~4/10
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
✓9/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
~5/10
Explore an interactive API reference with runnable examples
~4/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
—0/10
Get AI-generated insights and suggestions from my data inside the product
~6/10
Set up automations that run autonomously in the background
~6/10
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards
Stories about alerting dashboards in this arena
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Cost monitoring — stories about cost monitoring in this arenaCost monitoring
Stories about cost monitoring in this arena
Data access export — stories about data access export in this arenaData access export
Stories about data access export in this arena
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets
Measuring quality — datasets, eval runs, regression tracking
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Prompt management — stories about prompt management in this arenaPrompt management
Stories about prompt management in this arena
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation
Instrumenting code and tracing requests end to end
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
~6/10
Mask or redact sensitive data before it is stored in traces
✓8/10
Instrument apps in both Python and JS/TS with officially supported SDKs
~3/10
Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Cclaimed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 5/10 | Cclaimed | |
Score outputs with configurable LLM-as-a-judge evaluators C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 8/10 | Cclaimed | |
Compare eval runs side by side to catch regressions between prompt or model versions C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | partial | 6/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 6/10 | Cclaimed | |
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app C Ai observability | ai-native user | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | partial | 6/10 | Tprobed | |
Curate datasets from production traces and run offline evaluations against them C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | partial | 5/10 | Cclaimed | |
Capture traces of my LLM calls with inputs, outputs, latency, and token usage C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | partial | 4/10 | Tprobed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | none | 0/10 | ||
Version prompts and deploy changes to production without shipping code C Prompt workflow | developer | Prompt management — stories about prompt management in this arenaPrompt management | 3 | n/a | 0/10 | ||
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
See cost and token usage per request, model, and time period in dashboards C Cost tracking | developer | Cost monitoring — stories about cost monitoring in this arenaCost monitoring | 3 | none | untested | none yet | |
Have an agent create a dataset, trigger an eval run programmatically, and read back the results C Ai eval ops | ai-native user | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | full | 8/10 | Tprobed | |
Mask or redact sensitive data before it is stored in traces C Data controls | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | full | 8/10 | Cclaimed | |
Run evaluators continuously on live production traffic, not just offline datasets C Online evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | full | 8/10 | Tprobed | |
Write custom code-based scorers and metrics for my evaluations C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | full | 8/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 7/10 | Tprobed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | full | 7/10 | Cclaimed | |
Build custom dashboards over latency, error, cost, and eval-score metrics C Monitoring | ml engineer | Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Run evals in CI and gate deployments on their results C Offline evals | developer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 6/10 | Cclaimed | |
Instrument apps in both Python and JS/TS with officially supported SDKs G Sdk coverage | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | partial | 3/10 | Tprobed | |
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | partial | 3/10 | Cclaimed | |
Attribute cost and usage to users, sessions, and features via custom metadata C Cost tracking | developer | Cost monitoring — stories about cost monitoring in this arenaCost monitoring | 2 | none | 0/10 | ||
Bulk-export traces and datasets to blob storage or my data warehouse C Data export | developer | Data access export — stories about data access export in this arenaData access export | 2 | none | 0/10 | ||
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Iterate on prompts in a playground against real models and variables C Prompt workflow | developer | Prompt management — stories about prompt management in this arenaPrompt management | 2 | none | 0/10 | ||
Trace multi-step agent runs as nested spans grouped into sessions or threads C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Route outputs to human annotation queues for review and labeling C Human review | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | none | untested | none yet | |
Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email C Monitoring | developer | Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards | 2 | none | untested | none yet | |
Capture multimodal payloads (images, audio, files) inside my traces C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 1 | partial | 5/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 3/10 | Cclaimed |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 38 stories with headroom
What would move Cekura’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Cekura's AI-assistant integrations (Skills, MCP, CLI) are designed so external AI assistants like Claude or Cursor can drive Cekura's testing/evaluation features — this is the reverse relationship of an AI-native user delegating tasks to a built-in assistant inside Cekura itself.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
All Cekura MCP evidence describes Cekura exposing its own MCP server so external AI assistants can call Cekura's tools (docs-3, probe-4), not Cekura itself acting as an MCP client that consumes third-party MCP servers' tools.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
Cekura offers CLI/SDK/API access to call data and metrics, but there is no evidence of a bulk data export feature in open/portable formats or any account-closure/data-portability guarantee for users wishing to leave the platform.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Cekura is presented as a hosted SaaS platform (API keys, cloud dashboards, webhooks) with no mention of a self-hosted or on-premises deployment option anywhere in the docs, CLI, SDK, or website copy.
Tracing instrumentation — instrumenting code and tracing requests end to endSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary format
nonemoves PA Scoreimpact 30
Cekura's observability ingestion uses a proprietary POST endpoint (transcript, recording URL, metadata) and its own API/CLI/SDK, with no mention of OpenTelemetry or OTLP support anywhere in the evidence pack.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
Evidence shows PII redaction features for transcripts but nothing about opting out of AI model training on customer data, data-use policies, or training-data controls.
Cost monitoring — stories about cost monitoring in this arenaSee cost and token usage per request, model, and time period in dashboards
nonemoves PA Scoreimpact 30
Cekura documents customizable dashboards for call data, metrics, and metadata (cekura-docs-11), but nothing in the evidence pack mentions cost or token usage tracking per request, model, or time period.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
Missing: documented outbound webhook/event subscription mechanism, webhook configuration UI/API, event types list.
Showing the top 8 of 38 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 32 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Documentation docs23 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Build custom dashboards over latency, error, cost, and eval-score metrics
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Version, review, and roll back my automations
- Run evals in CI and gate deployments on their results
- Write custom code-based scorers and metrics for my evaluations
- Compare eval runs side by side to catch regressions between prompt or model versions
- Score outputs with configurable LLM-as-a-judge evaluators
- Curate datasets from production traces and run offline evaluations against them
- Run evaluators continuously on live production traffic, not just offline datasets
- Do everything through the API that I can do in the UI
- Mask or redact sensitive data before it is stored in traces
- Instrument apps in both Python and JS/TS with officially supported SDKs
- Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK
- Capture multimodal payloads (images, audio, files) inside my traces
- Capture traces of my LLM calls with inputs, outputs, latency, and token usage
API reference18 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Set up automations that run autonomously in the background
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Build custom dashboards over latency, error, cost, and eval-score metrics
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Run evals in CI and gate deployments on their results
- Curate datasets from production traces and run offline evaluations against them
- Run evaluators continuously on live production traffic, not just offline datasets
- Do everything through the API that I can do in the UI
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Mask or redact sensitive data before it is stored in traces
- Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK
- Capture multimodal payloads (images, audio, files) inside my traces
- Capture traces of my LLM calls with inputs, outputs, latency, and token usage
CLI SDK docs13 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Build against official SDKs
- Perform bulk operations across many items at once
- Schedule recurring jobs or workflows
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Run evals in CI and gate deployments on their results
- Write custom code-based scorers and metrics for my evaluations
- Do everything through the API that I can do in the UI
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Instrument apps in both Python and JS/TS with officially supported SDKs
- Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK
OpenAPI spec8 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Do everything through the API that I can do in the UI
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
cekura.ai7 stories
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Compare eval runs side by side to catch regressions between prompt or model versions
- Run evaluators continuously on live production traffic, not just offline datasets
MCP docs5 stories
- Connect an agent via an official MCP server
- Operate the product with natural-language commands
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Capture traces of my LLM calls with inputs, outputs, latency, and token usage
llms.txt2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$uvx --from 'cekura[cli]' cekura --helpreproduced$ uvx --from 'cekura[cli]' cekura --help ⠋ Resolving dependencies... ⠙ Resolving dependencies... ⠋ Resolving dependencies... ⠙ Resolving dependencies... ⠙ cekura==1.6.7 ⠙ cekura==1.6.7 ⠙ aiohttp==3.14.3 ⠙ httpx==0.28.1 ⠙ typer==0.27.2 ⠙ rich==15.0.0 ⠙ aiohappyeyeballs==2.7.1 ⠙ aiosignal==1.4.0 ⠙ attrs==26.1.0 ⠙ frozenlist==1.8.0 ⠙ multidict==6.8.0 ⠙ propcache==0.5.2 ⠙ yarl==1.24.5 ⠙ anyio==4.15.1 ⠙ certifi==2026.7.22 ⠙ httpcore==1.0.9 ⠙ idna==3.19 ⠙ shellingham==1.5.4 Usage: cekura [OPTIONS] COMMAND [ARGS]... Cekura CLI — testing and observability for voice AI agents. ╭─ Options ────────────────────────────────────────────────────────────────────╮ │ --install-completion Install completion for the current shell. │ │ --show-completion Show completion for the current shell, to copy │ │ it or customize the installation. │ │ --help Show this message and exit. │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭─ Commands ───────────────────────────────────────────────────────────────────╮ │ version Show the Cekura SDK version. │ │ auth Authentication management. │ │ agents Manage AI voice agents. │ │ scenarios Manage test scenarios (evaluators). │ │ metrics Manage evaluation metrics. │ │ run Run evaluations and view results. │ │ runs View and manage individual evaluation runs. │ │ results View and manage evaluation results. │ │ calls View and manage production calls. │ │ alerts Configure and review observability alerts. │ │ projects Manage projects. │ │ personalities Manage caller personalities. │ │ test-profiles Manage reusable test configuration profiles. │ │ cron Schedule recurring evaluation runs. │ │ config CLI configuration management. │ │ dashboards Manage analytics dashboards and widgets. │ │ metric-reviews Process metric review feedbacks (labs │ │ pipeline). │ │ critical-metric-scenarios Read and update critical-metric-scenarios. │ │ test-sets Create and manage test sets. │ │ predefined-metrics Browse the predefined-metrics catalog. │ │ scenario-improvement-sessions Manage scenario improvement sessions. │ │ insights Audit metric failure modes and generate │ │ scenarios from them. │ │ deep-research Run and review project-wide Deep Research │ │ audits. │ │ slack Inspect connected Slack workspaces. │ │ phone-numbers Phone number tooling. │ │ billing View billing info. │ │ organizations List the current user's organizations. │ │ api-[redacted]s Create Cekura API [redacted]s (requires bearer-[redacted] │ │ auth). │ ╰──────────────────────────────────────────────────────────────────────────────╯ Tip: Use --format json for machine-readable output.
$curl -s https://docs.cekura.ai/llms.txt | head -6reproduced$ curl -s https://docs.cekura.ai/llms.txt | head -6 # Cekura > Cekura is the testing and observability platform for voice AI agents. Run simulated conversations, evaluate performance with LLM-judge and code metrics, and monitor production calls. ## Getting Started - [Introduction](https://docs.cekura.ai/documentation/introduction.md): Testing for AI Voice Agents. Launch in minutes not weeks by ensuring your agents deliver a seamless experience in every conversational scenario
$curl -si -X POST https://api.cekura.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -si -X POST https://api.cekura.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401
date: Mon, 14 Sep 2026 23:35:43 GMT
content-type: application/json
content-length: 98
server: uvicorn
www-authenticate: Bearer resource_metadata="https://api.cekura.ai/mcp/.well-known/oauth-protected-resource", error="invalid_[redacted]"
{"error":"unauthorized","error_description":"Authenticate via OAuth (Bearer) or X-CEKURA-API-[redacted]"}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
5 of 17 testable claims verified · 1 contradicted → integrity 18/100
21 distinct capability claims found in Cekura’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
5
Verified
11
Unverified
1
Contradicted
16
Undersold
Verified (6)
“AI assistant can draft a full plan to integrate Cekura observability into your codebase”
Operate the product with natural-language commandspartialproof ↗
“Provides an official MCP server (plus skills) so an AI assistant can design, run, and improve voice-agent evaluations”
“CLI lets you manage agents, scenarios, runs, and call data from the terminal”
“Offers sync and async Python clients for programmatic access from application code”
“Detects behavioral drift live across every production call”
Run evaluators continuously on live production traffic, not just offline datasetsfullproof ↗
“Plugin bundles 13 skills and 14 commands an AI assistant can invoke directly”
Operate the product with natural-language commandspartialproof ↗
Unverified (12)
“Supports creating project-scoped API keys for authentication”
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
“Lets you write custom evaluation logic in Python for full control over scoring”
Write custom code-based scorers and metrics for my evaluationsfullproof ↗
“Evaluates voice agent calls against natural-language criteria using an LLM judge”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Compares two versions of an agent side-by-side to measure prompt/model/config changes”
Compare eval runs side by side to catch regressions between prompt or model versionspartialproof ↗
“Build custom dashboards with widgets over call data, metrics, and metadata”
Build custom dashboards over latency, error, cost, and eval-score metricspartialproof ↗
“Automatically detects and redacts sensitive info from transcripts and audio recordings”
Mask or redact sensitive data before it is stored in tracesfullproof ↗
“Set up automated cron jobs to run recurring agent testing and evaluation workflows”
“GitHub Actions integration automatically tests agents on every code change”
Run evals in CI and gate deployments on their resultspartialproof ↗
“Provides enhanced observability instrumentation for LiveKit agents via Cekura SDK”
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDKpartialproof ↗
“Lets you quickly create and test voice agents without needing external API keys or integrations”
Test against a sandbox environment without touching production datapartialproof ↗
“Runs the same test scenarios across different platforms and models to compare real performance”
Compare eval runs side by side to catch regressions between prompt or model versionspartialproof ↗
“Automatically flags issues, reproduces them in simulation, and suggests fixes”
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Contradicted (1)
“Accepts webhook POSTs of transcripts/recordings which are stored and auto-scheduled for metric evaluation”
Undersold (16)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Drive the product through a documented public APIfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)fullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Have an agent create a dataset, trigger an eval run programmatically, and read back the resultsfullproof ↗
Curate datasets from production traces and run offline evaluations against thempartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my apppartialproof ↗
Instrument apps in both Python and JS/TS with officially supported SDKspartialproof ↗
Capture multimodal payloads (images, audio, files) inside my tracespartialproof ↗
Capture traces of my LLM calls with inputs, outputs, latency, and token usagepartialproof ↗
Claims outside our story set (2)
Real capability claims found in Cekura’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Probes agents for jailbreaks, data leaks, and off-script behavior”
source ↗“Runs thousands of synthetic conversations to stress-test an agent before launch”
source ↗
Business model
Pay-as-you-go voice testing at $0.25/min and $0.05/monitored call with 300 free starter credits; $500/mo Startup plan covers ~2,000 test minutes, ~10k monitored calls and 10 seats; enterprise is custom (VPC/on-prem, SSO, BAA).
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
