Skip to content

Rank #2 of 8 in LLM Evals & Observability

Cekura logo

Cekura

YC F24

Cekura · commercial

713/yr

Access

Install

pippip install cekura
npxnpx skills add cekura-ai/cekura-skills --all

Compare head-to-head

Alternatives to Cekura

Try itExperimental

See what an agent can do with Cekura before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$uvx --from 'cekura[cli]' cekura --helprecorded session — replayed, not live
recorded 2026-09-14 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

45.7/100

Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboardsevidence →

Stories about alerting dashboards in this arena

18.0/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

42.3/100

Cost monitoring — stories about cost monitoring in this arenaCost monitoringevidence →

Stories about cost monitoring in this arena

0.0/100

Data access export — stories about data access export in this arenaData access exportevidence →

Stories about data access export in this arena

0.0/100

Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasetsevidence →

Measuring quality — datasets, eval runs, regression tracking

52.1/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

8.4/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Prompt management — stories about prompt management in this arenaPrompt managementevidence →

Stories about prompt management in this arena

0.0/100

Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentationevidence →

Instrumenting code and tracing requests end to end

24.6/100

Story verdicts — every judged story with its evidenceStory verdicts

?

Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10T

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3none0/10

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3none0/10

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10C

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10T

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10C

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10T

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10C

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial5/10C

Score outputs with configurable LLM-as-a-judge evaluators C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3full8/10C

Compare eval runs side by side to catch regressions between prompt or model versions C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3partial6/10C

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial6/10C

Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app C

Ai observability

ai-native userTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3partial6/10T

Curate datasets from production traces and run offline evaluations against them C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3partial5/10C

Capture traces of my LLM calls with inputs, outputs, latency, and token usage C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3partial4/10T

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3none0/10

Version prompts and deploy changes to production without shipping code C

Prompt workflow

developerPrompt management — stories about prompt management in this arenaPrompt management3n/a0/10

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3noneuntestednone yet

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

See cost and token usage per request, model, and time period in dashboards C

Cost tracking

developerCost monitoring — stories about cost monitoring in this arenaCost monitoring3noneuntestednone yet

Have an agent create a dataset, trigger an eval run programmatically, and read back the results C

Ai eval ops

ai-native userEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10T

Mask or redact sensitive data before it is stored in traces C

Data controls

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2full8/10C

Run evaluators continuously on live production traffic, not just offline datasets C

Online evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10T

Write custom code-based scorers and metrics for my evaluations C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial7/10T

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2full7/10C

Build custom dashboards over latency, error, cost, and eval-score metrics C

Monitoring

ml engineerAlerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards2partial6/10C

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial6/10C

Run evals in CI and gate deployments on their results C

Offline evals

developerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2partial6/10C

Instrument apps in both Python and JS/TS with officially supported SDKs G

Sdk coverage

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2partial3/10T

Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2partial3/10C

Attribute cost and usage to users, sessions, and features via custom metadata C

Cost tracking

developerCost monitoring — stories about cost monitoring in this arenaCost monitoring2none0/10

Bulk-export traces and datasets to blob storage or my data warehouse C

Data export

developerData access export — stories about data access export in this arenaData access export2none0/10

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2none0/10

Iterate on prompts in a playground against real models and variables C

Prompt workflow

developerPrompt management — stories about prompt management in this arenaPrompt management2none0/10

Trace multi-step agent runs as nested spans grouped into sessions or threads C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2noneuntestednone yet

Route outputs to human annotation queues for review and labeling C

Human review

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2noneuntestednone yet

Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email C

Monitoring

developerAlerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards2noneuntestednone yet

Capture multimodal payloads (images, audio, files) inside my traces C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation1partial5/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial3/10C

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 38 stories with headroom

What would move Cekura’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product

    nonemoves Built-in AIimpact 45

    Cekura's AI-assistant integrations (Skills, MCP, CLI) are designed so external AI assistants like Claude or Cursor can drive Cekura's testing/evaluation features — this is the reverse relationship of an AI-native user delegating tasks to a built-in assistant inside Cekura itself.

  2. Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools

    nonemoves agent-readyimpact 45

    All Cekura MCP evidence describes Cekura exposing its own MCP server so external AI assistants can call Cekura's tools (docs-3, probe-4), not Cekura itself acting as an MCP client that consumes third-party MCP servers' tools.

  3. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    Cekura offers CLI/SDK/API access to call data and metrics, but there is no evidence of a bulk data export feature in open/portable formats or any account-closure/data-portability guarantee for users wishing to leave the platform.

  4. Openness — open source, data portability, and self-hosting storiesSelf-host the core product

    nonemoves PA Scoreimpact 30

    Cekura is presented as a hosted SaaS platform (API keys, cloud dashboards, webhooks) with no mention of a self-hosted or on-premises deployment option anywhere in the docs, CLI, SDK, or website copy.

  5. Tracing instrumentation — instrumenting code and tracing requests end to endSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary format

    nonemoves PA Scoreimpact 30

    Cekura's observability ingestion uses a proprietary POST endpoint (transcript, recording URL, metadata) and its own API/CLI/SDK, with no mention of OpenTelemetry or OTLP support anywhere in the evidence pack.

  6. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    Evidence shows PII redaction features for transcripts but nothing about opting out of AI model training on customer data, data-use policies, or training-data controls.

  7. Cost monitoring — stories about cost monitoring in this arenaSee cost and token usage per request, model, and time period in dashboards

    nonemoves PA Scoreimpact 30

    Cekura documents customizable dashboards for call data, metrics, and metadata (cekura-docs-11), but nothing in the evidence pack mentions cost or token usage tracking per request, model, or time period.

  8. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    Missing: documented outbound webhook/event subscription mechanism, webhook configuration UI/API, event types list.

Showing the top 8 of 38 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 32 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$uvx --from 'cekura[cli]' cekura --helpreproduced
$ uvx --from 'cekura[cli]' cekura --help
⠋ Resolving dependencies...                                                     
⠙ Resolving dependencies...                                                     
⠋ Resolving dependencies...                                                     
⠙ Resolving dependencies...                                                     
⠙ cekura==1.6.7                                                                 
⠙ cekura==1.6.7                                                                 
⠙ aiohttp==3.14.3                                                               
⠙ httpx==0.28.1                                                                 
⠙ typer==0.27.2                                                                 
⠙ rich==15.0.0                                                                  
⠙ aiohappyeyeballs==2.7.1                                                       
⠙ aiosignal==1.4.0                                                              
⠙ attrs==26.1.0                                                                 
⠙ frozenlist==1.8.0                                                             
⠙ multidict==6.8.0                                                              
⠙ propcache==0.5.2                                                              
⠙ yarl==1.24.5                                                                  
⠙ anyio==4.15.1                                                                 
⠙ certifi==2026.7.22                                                            
⠙ httpcore==1.0.9                                                               
⠙ idna==3.19                                                                    
⠙ shellingham==1.5.4                                                            

 Usage: cekura [OPTIONS] COMMAND [ARGS]...

 Cekura CLI — testing and observability for voice AI agents.

╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --install-completion          Install completion for the current shell.      │
│ --show-completion             Show completion for the current shell, to copy │
│                               it or customize the installation.              │
│ --help                        Show this message and exit.                    │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────╮
│ version                        Show the Cekura SDK version.                  │
│ auth                           Authentication management.                    │
│ agents                         Manage AI voice agents.                       │
│ scenarios                      Manage test scenarios (evaluators).           │
│ metrics                        Manage evaluation metrics.                    │
│ run                            Run evaluations and view results.             │
│ runs                           View and manage individual evaluation runs.   │
│ results                        View and manage evaluation results.           │
│ calls                          View and manage production calls.             │
│ alerts                         Configure and review observability alerts.    │
│ projects                       Manage projects.                              │
│ personalities                  Manage caller personalities.                  │
│ test-profiles                  Manage reusable test configuration profiles.  │
│ cron                           Schedule recurring evaluation runs.           │
│ config                         CLI configuration management.                 │
│ dashboards                     Manage analytics dashboards and widgets.      │
│ metric-reviews                 Process metric review feedbacks (labs         │
│                                pipeline).                                    │
│ critical-metric-scenarios      Read and update critical-metric-scenarios.    │
│ test-sets                      Create and manage test sets.                  │
│ predefined-metrics             Browse the predefined-metrics catalog.        │
│ scenario-improvement-sessions  Manage scenario improvement sessions.         │
│ insights                       Audit metric failure modes and generate       │
│                                scenarios from them.                          │
│ deep-research                  Run and review project-wide Deep Research     │
│                                audits.                                       │
│ slack                          Inspect connected Slack workspaces.           │
│ phone-numbers                  Phone number tooling.                         │
│ billing                        View billing info.                            │
│ organizations                  List the current user's organizations.        │
│ api-[redacted]s                       Create Cekura API [redacted]s (requires bearer-[redacted] │
│                                auth).                                        │
╰──────────────────────────────────────────────────────────────────────────────╯

 Tip: Use --format json for machine-readable output.
proves: Use an official CLIrecorded 2026-09-14
$curl -s https://docs.cekura.ai/llms.txt | head -6reproduced
$ curl -s https://docs.cekura.ai/llms.txt | head -6
# Cekura

> Cekura is the testing and observability platform for voice AI agents. Run simulated conversations, evaluate performance with LLM-judge and code metrics, and monitor production calls.

## Getting Started
- [Introduction](https://docs.cekura.ai/documentation/introduction.md): Testing for AI Voice Agents. Launch in minutes not weeks by ensuring your agents deliver a seamless experience in every conversational scenario
$curl -si -X POST https://api.cekura.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced
$ curl -si -X POST https://api.cekura.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401

date: Mon, 14 Sep 2026 23:35:43 GMT

content-type: application/json

content-length: 98

server: uvicorn

www-authenticate: Bearer resource_metadata="https://api.cekura.ai/mcp/.well-known/oauth-protected-resource", error="invalid_[redacted]"

{"error":"unauthorized","error_description":"Authenticate via OAuth (Bearer) or X-CEKURA-API-[redacted]"}

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

5 of 17 testable claims verified · 1 contradictedintegrity 18/100

21 distinct capability claims found in Cekura’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

5

Verified

11

Unverified

1

Contradicted

16

Undersold

Verified (6)
Unverified (12)
Contradicted (1)
Undersold (16)
Claims outside our story set (2)

Real capability claims found in Cekura’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Probes agents for jailbreaks, data leaks, and off-script behavior

    source ↗
  • Runs thousands of synthetic conversations to stress-test an agent before launch

    source ↗
Suggest a story for these →

Business model

free-tierusage-basedsubscription-flatenterprise-custom

Pay-as-you-go voice testing at $0.25/min and $0.05/monitored call with 300 free starter credits; $500/mo Startup plan covers ~2,000 test minutes, ~10k monitored calls and 10 seats; enterprise is custom (VPC/on-prem, SSO, BAA).

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Scoretracked since Sep 14 '26 — no movement recorded yet
Agent-readytracked since Sep 14 '26 — no movement recorded yet

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data