LLM Evals & Observability arenaLLM Evals & Observability
Tracing, evaluation, and observability platforms for LLM apps and agents, judged on trace capture and OpenTelemetry support, offline and online evals, prompt management, cost visibility, self-hostability, and how well an agent can drive the platform itself through its API.
52 user stories · 416 judged cells · updated 2026-09-14 · Evidence as of 2026-09-14
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | free-tier vs Cekura ↗ | 51/100 | 67/100 | 3/100 | 28/100 | 40/100 | ★ 27▲ 9/yrnpm 1.3M/wk | 10/39 verified | 21/100 integrity | ||
| 2 | free-tier vs Braintrust ↗ | 58/100 | 24/100 | 37/100 | 8/100 | 42/100 | ★ 7▲ 13/yr | 15/32 verified | 18/100 integrity | ||
| 3 | free-tier vs Braintrust ↗ | 56/100 | 34/100 | 33/100 | 13/100 | 28/100 | ★ 1.1k▲ 345/yrnpm 348.4k/wk | 11/36 verified | 24/100 integrity | ||
| 4 | free-tier vs Braintrust ↗ | 43/100 | 24/100 | 37/100 | 30/100 | 29/100 | ★ 6.2k▲ 1.7k/yrnpm 3.6k/wk | 16/35 verified · 1 disputed | 0/100 integrity | ||
| 5 | free-tier vs Braintrust ↗ | 52/100 | 19/100 | 3/100 | 61/100 | 17/100 | ★ 34.6k▲ 10.4k/yrnpm 1.4M/wkpypi 5M/wk | 19/42 verified · 2 disputed | 54/100 integrity | ||
| 6 | free-tier vs Braintrust ↗ | 53/100 | 22/100 | 4/100 | 50/100 | 11/100 | ★ 11.5k▲ 3k/yrnpm 48.7k/wk | 6/37 verified | 0/100 integrity | ||
| 7 | free-tier vs Braintrust ↗ | 36/100 | 21/100 | 4/100 | 24/100 | 23/100 | ★ 1.1k▲ 320/yrnpm 4.8M/wk | 14/36 verified | 29/100 integrity | ||
| 8 | free-tier vs Braintrust ↗ | 36/100 | 17/100 | 3/100 | 10/100 | 16/100 | npm 1.1k/wk | 8/28 verified | 8/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Story matrix — every product × every judged storyStory matrix
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 9/10 | fullT 8/10 | partialT 6/10 | fullT 9/10 | fullT 9/10 | fullT 8/10 | fullT 8/10 | fullT 9/10 |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | partialC 6/10 | partialC 6/10 | fullT 9/10 | partialC 7/10 | partialC 6/10 | partialT 6/10 | partialC 6/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | n/a | n/a | none 0/10 | n/a | n/a | n/a | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | fullT 7/10 | partialT 4/10 | fullT 8/10 | fullT 8/10 | fullT 8/10 | none 0/10 | fullT 8/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | fullC 7/10 | none 0/10 | fullT 9/10 | partialC 6/10 | partialT 5/10 | none 0/10 | none 0/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | partialT 6/10 | partialT 6/10 | fullT 8/10 | fullT 7/10 | fullT 8/10 | fullT 8/10 | partialT 6/10 | fullT 9/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | fullT 8/10 | fullX 8/10 | fullT 8/10 | fullC 8/10 | fullT 8/10 | fullX 8/10 | partialT 6/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | partialC 5/10 | partialC 6/10 | none 0/10 | none 0/10 | partialC 4/10 | fullC 7/10 | none 0/10 | none 0/10 |
Agentic features
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | partialC 4/10 | partialX 5/10 | fullC 8/10 | partialC 5/10 | partialC 6/10 | partialC 4/10 | partialC 5/10 | partialC 6/10 |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | partialC 4/10 | partialC 6/10 | partialC 6/10 | none 0/10 | partialC 5/10 | n/a | partialC 3/10 | partialC 6/10 |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | none 0/10 | partialC 3/10 | fullC 7/10 | none 0/10 | n/a | n/a | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | partialT 6/10 | none 0/10 | fullT 8/10 | fullT 7/10 | partialT 6/10 | n/a | partialT 5/10 | partialT 6/10 |
Api quality
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialT 3/10 | partialT 5/10 | none 0/10 | partialT 4/10 |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | fullT 9/10 | fullT 9/10 | none 0/10 | fullT 9/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | partialC 4/10 | partialC 5/10 | partialC 4/10 | partialC 5/10 | partialC 3/10 | partialC 3/10 | partialC 4/10 | partialC 5/10 |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards
Monitoring
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Alerting dashboards — stories about alerting dashboards in this arenaBuild custom dashboards over latency, error, cost, and eval-score metrics | ml-engineer | fullC 8/10 | fullC 8/10 | partialC 5/10 | none 0/10 | partialC 6/10 | partialX 6/10 | none 0/10 | partialC 6/10 |
| Alerting dashboards — stories about alerting dashboards in this arenaSet alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email | developer | partialC 6/10 | fullC 8/10 | none 0/10 | none 0/10 | partialC 4/10 | partialC 4/10 | partialC 5/10 | none 0/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | disputedD 5/10 | partialC 6/10 | fullC 7/10 | partialC 5/10 | partialC 5/10 | partialC 4/10 | partialC 5/10 | partialC 6/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | partialC 4/10 | partialC 6/10 | partialC 5/10 | none 0/10 | partialC 5/10 | partialC 5/10 | partialC 3/10 | partialC 6/10 |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | none 0/10 | none 0/10 | partialC 5/10 | none 0/10 | n/a | n/a | none 0/10 | fullC 7/10 |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | partialC 6/10 | none 0/10 | partialC 5/10 | partialC 5/10 | partialC 3/10 | partialC 6/10 | n/a | partialC 3/10 |
Cost monitoring — stories about cost monitoring in this arenaCost monitoring
Cost tracking
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Cost monitoring — stories about cost monitoring in this arenaAttribute cost and usage to users, sessions, and features via custom metadata | developer | fullC 7/10 | fullC 7/10 | partialC 4/10 | partialC 4/10 | partialC 5/10 | partialC 4/10 | partialC 4/10 | none 0/10 |
| Cost monitoring — stories about cost monitoring in this arenaSee cost and token usage per request, model, and time period in dashboards | developer | fullX 9/10 | fullC 8/10 | partialC 5/10 | partialC 5/10 | fullC 8/10 | partialX 5/10 | none 0/10 | none 0/10 |
Data access export — stories about data access export in this arenaData access export
Data export
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Data access export — stories about data access export in this arenaBulk-export traces and datasets to blob storage or my data warehouse | developer | fullX 8/10 | partialC 3/10 | partialC 5/10 | none 0/10 | none 0/10 | partialX 4/10 | none 0/10 | none 0/10 |
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets
Ai eval ops
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Evals datasets — measuring quality — datasets, eval runs, regression trackingHave an agent create a dataset, trigger an eval run programmatically, and read back the results | ai-native | partialC 7/10 | fullT 7/10 | fullC 8/10 | partialC 6/10 | fullT 8/10 | none 0/10 | fullT 7/10 | fullT 8/10 |
Human review
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Evals datasets — measuring quality — datasets, eval runs, regression trackingRoute outputs to human annotation queues for review and labeling | ml-engineer | fullC 8/10 | fullC 8/10 | partialC 5/10 | partialC 6/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 |
Offline evals
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Evals datasets — measuring quality — datasets, eval runs, regression trackingRun evals in CI and gate deployments on their results | developer | fullC 7/10 | partialC 5/10 | fullC 8/10 | partialC 5/10 | partialC 4/10 | none 0/10 | none 0/10 | partialC 6/10 |
| Evals datasets — measuring quality — datasets, eval runs, regression trackingWrite custom code-based scorers and metrics for my evaluations | ml-engineer | fullC 7/10 | fullC 8/10 | fullC 8/10 | fullC 8/10 | fullC 9/10 | partialC 3/10 | partialC 6/10 | fullC 8/10 |
| Evals datasets — measuring quality — datasets, eval runs, regression trackingCompare eval runs side by side to catch regressions between prompt or model versions | ml-engineer | fullC 8/10 | fullC 8/10 | fullC 9/10 | fullC 8/10 | fullC 8/10 | partialC 3/10 | partialC 5/10 | partialC 6/10 |
| Evals datasets — measuring quality — datasets, eval runs, regression trackingScore outputs with configurable LLM-as-a-judge evaluators | ml-engineer | fullC 8/10 | fullC 8/10 | fullC 8/10 | fullC 9/10 | fullC 9/10 | partialC 4/10 | fullC 8/10 | fullC 8/10 |
| Evals datasets — measuring quality — datasets, eval runs, regression trackingCurate datasets from production traces and run offline evaluations against them | ml-engineer | fullC 8/10 | fullC 8/10 | fullC 9/10 | fullC 8/10 | fullC 8/10 | partialC 4/10 | fullC 8/10 | partialC 5/10 |
Online evals
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Evals datasets — measuring quality — datasets, eval runs, regression trackingRun evaluators continuously on live production traffic, not just offline datasets | ml-engineer | fullC 7/10 | fullC 8/10 | fullC 8/10 | partialC 5/10 | partialC 6/10 | partialC 5/10 | fullC 8/10 | fullT 8/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialT 6/10 | partialT 5/10 | partialT 7/10 | partialT 5/10 | partialT 6/10 | partialT 6/10 | partialT 5/10 | partialT 7/10 |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | partialX 6/10 | partialX 4/10 | partialC 5/10 | partialC 6/10 | partialC 3/10 | partialX 5/10 | none 0/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | fullT 8/10 | none 0/10 | none 0/10 | partialC 5/10 | none 0/10 | disputedD 5/10 | partialC 3/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | fullT 9/10 | partialX 6/10 | partialC 6/10 | fullC 9/10 | none 0/10 | partialX 6/10 | none 0/10 | none 0/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | disputedD 5/10 | partialC 4/10 | partialC 6/10 | partialC 6/10 | none 0/10 | partialC 4/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | partialC 3/10 | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 | n/a | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | partialX 3/10 | none 0/10 | partialC 4/10 | partialC 5/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | none 0/10 | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 |
Prompt management — stories about prompt management in this arenaPrompt management
Prompt workflow
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Prompt management — stories about prompt management in this arenaIterate on prompts in a playground against real models and variables | developer | fullX 7/10 | none 0/10 | fullC 8/10 | fullC 8/10 | fullC 8/10 | partialC 6/10 | partialC 5/10 | none 0/10 |
| Prompt management — stories about prompt management in this arenaVersion prompts and deploy changes to production without shipping code | developer | fullX 8/10 | partialC 4/10 | partialC 6/10 | fullC 8/10 | none 0/10 | fullC 8/10 | none 0/10 | n/a |
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation
Ai observability
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Tracing instrumentation — instrumenting code and tracing requests end to endHave an agent query my traces, metrics, and eval results through an API or MCP server to debug my app | ai-native | partialT 6/10 | partialT 6/10 | fullT 9/10 | fullT 8/10 | fullT 8/10 | partialT 5/10 | partialT 5/10 | partialT 6/10 |
Data controls
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Tracing instrumentation — instrumenting code and tracing requests end to endMask or redact sensitive data before it is stored in traces | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | fullC 8/10 |
Sdk coverage
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Tracing instrumentation — instrumenting code and tracing requests end to endInstrument apps in both Python and JS/TS with officially supported SDKs | developer | fullX 8/10 | fullX 8/10 | partialC 6/10 | partialC 6/10 | fullC 8/10 | partialX 6/10 | partialC 4/10 | partialT 3/10 |
Trace capture
| Story | Persona | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Tracing instrumentation — instrumenting code and tracing requests end to endTrace multi-step agent runs as nested spans grouped into sessions or threads | developer | fullX 8/10 | partialX 6/10 | partialC 6/10 | fullC 8/10 | fullC 9/10 | fullC 8/10 | fullC 8/10 | none 0/10 |
| Tracing instrumentation — instrumenting code and tracing requests end to endInstrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK | developer | fullX 7/10 | partialX 6/10 | partialX 4/10 | fullC 8/10 | partialC 5/10 | fullX 8/10 | none 0/10 | partialC 3/10 |
| Tracing instrumentation — instrumenting code and tracing requests end to endCapture multimodal payloads (images, audio, files) inside my traces | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 5/10 |
| Tracing instrumentation — instrumenting code and tracing requests end to endSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary format | developer | fullC 8/10 | fullC 7/10 | none 0/10 | fullC 9/10 | partialC 7/10 | none 0/10 | partialC 5/10 | none 0/10 |
| Tracing instrumentation — instrumenting code and tracing requests end to endCapture traces of my LLM calls with inputs, outputs, latency, and token usage | developer | fullX 9/10 | fullX 8/10 | fullC 8/10 | fullC 9/10 | fullC 9/10 | fullX 9/10 | fullC 7/10 | partialT 4/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.