Skip to content

Arena

LLM Evals & Observability arenaLLM Evals & Observability

Tracing, evaluation, and observability platforms for LLM apps and agents, judged on trace capture and OpenTelemetry support, offline and online evals, prompt management, cost visibility, self-hostability, and how well an agent can drive the platform itself through its API.

52 user stories · 416 judged cells · updated 2026-09-14 · Evidence as of 2026-09-14

Buyer checklist →Procurement report →

Leaderboard — every product ranked by evidenceLeaderboard

Best by user type — persona-weighted winnersBest by user type

Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.

Best for developer

Langfuse logo

Langfuse

69/100

Runner-up: Arize Phoenix logo Arize Phoenix (49/100)

14 developer stories scored

Best for ml-engineer

LangSmith logo

LangSmith

80/100

Runner-up: Langfuse logo Langfuse (78/100)

7 ml-engineer stories scored

Best for ai-native

Braintrust logo

Braintrust

41/100

Runner-up: Arize Phoenix logo Arize Phoenix (36/100)

31 ai-native stories scored

Story matrix — every product × every judged storyStory matrix

52/52 stories shown · legend

Agenticness — how well agents can access and operate the productAgenticness

Agent access

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docsai-native
fullT
9/10
fullT
8/10
partialT
6/10
fullT
9/10
fullT
9/10
fullT
8/10
fullT
8/10
fullT
9/10
Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automationai-native
partialC
6/10
partialC
6/10
fullT
9/10
partialC
7/10
partialC
6/10
partialT
6/10
partialC
6/10
fullT
8/10
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their toolsai-native
n/a
n/a
none
0/10
n/a
n/a
n/a
n/a
none
0/10
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP serverai-native
fullT
7/10
partialT
4/10
fullT
8/10
fullT
8/10
fullT
8/10
none
0/10
fullT
8/10
fullT
8/10
Agenticness — how well agents can access and operate the productUse an official CLIai-native
fullC
7/10
none
0/10
fullT
9/10
partialC
6/10
partialT
5/10
none
0/10
none
0/10
fullT
8/10
Agenticness — how well agents can access and operate the productDrive the product through a documented public APIai-native
partialT
6/10
partialT
6/10
fullT
8/10
fullT
7/10
fullT
8/10
fullT
8/10
partialT
6/10
fullT
9/10
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agentai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
partialC
4/10
Agenticness — how well agents can access and operate the productBuild against official SDKsai-native
fullT
8/10
fullX
8/10
fullT
8/10
fullC
8/10
fullT
8/10
fullX
8/10
partialT
6/10
fullT
8/10
Agenticness — how well agents can access and operate the productSubscribe to events via webhooksai-native
partialC
5/10
partialC
6/10
none
0/10
none
0/10
partialC
4/10
fullC
7/10
none
0/10
none
0/10

Agentic features

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the productai-native
partialC
4/10
partialX
5/10
fullC
8/10
partialC
5/10
partialC
6/10
partialC
4/10
partialC
5/10
partialC
6/10
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the backgroundai-native
partialC
4/10
partialC
6/10
partialC
6/10
none
0/10
partialC
5/10
n/a
partialC
3/10
partialC
6/10
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the productai-native
none
0/10
partialC
3/10
fullC
7/10
none
0/10
n/a
n/a
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productOperate the product with natural-language commandsai-native
partialT
6/10
none
0/10
fullT
8/10
fullT
7/10
partialT
6/10
n/a
partialT
5/10
partialT
6/10

Api quality

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examplesai-native
none
0/10
none
0/10
none
0/10
none
0/10
partialT
3/10
partialT
5/10
none
0/10
partialT
4/10
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)ai-native
none
0/10
none
0/10
none
0/10
none
0/10
fullT
9/10
fullT
9/10
none
0/10
fullT
9/10
Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production dataai-native
partialC
4/10
partialC
5/10
partialC
4/10
partialC
5/10
partialC
3/10
partialC
3/10
partialC
4/10
partialC
5/10
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policyai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards

Monitoring

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Alerting dashboards — stories about alerting dashboards in this arenaBuild custom dashboards over latency, error, cost, and eval-score metricsml-engineer
fullC
8/10
fullC
8/10
partialC
5/10
none
0/10
partialC
6/10
partialX
6/10
none
0/10
partialC
6/10
Alerting dashboards — stories about alerting dashboards in this arenaSet alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or emaildeveloper
partialC
6/10
fullC
8/10
none
0/10
none
0/10
partialC
4/10
partialC
4/10
partialC
5/10
none
0/10

Automation depth — how much of the product can run unattendedAutomation depth

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Automation depth — how much of the product can run unattendedPerform bulk operations across many items at onceai-native
disputedD
5/10
partialC
6/10
fullC
7/10
partialC
5/10
partialC
5/10
partialC
4/10
partialC
5/10
partialC
6/10
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on eventsai-native
partialC
4/10
partialC
6/10
partialC
5/10
none
0/10
partialC
5/10
partialC
5/10
partialC
3/10
partialC
6/10
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflowsai-native
none
0/10
none
0/10
partialC
5/10
none
0/10
n/a
n/a
none
0/10
fullC
7/10
Automation depth — how much of the product can run unattendedVersion, review, and roll back my automationsai-native
partialC
6/10
none
0/10
partialC
5/10
partialC
5/10
partialC
3/10
partialC
6/10
n/a
partialC
3/10

Cost monitoring — stories about cost monitoring in this arenaCost monitoring

Cost tracking

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Cost monitoring — stories about cost monitoring in this arenaAttribute cost and usage to users, sessions, and features via custom metadatadeveloper
fullC
7/10
fullC
7/10
partialC
4/10
partialC
4/10
partialC
5/10
partialC
4/10
partialC
4/10
none
0/10
Cost monitoring — stories about cost monitoring in this arenaSee cost and token usage per request, model, and time period in dashboardsdeveloper
fullX
9/10
fullC
8/10
partialC
5/10
partialC
5/10
fullC
8/10
partialX
5/10
none
0/10
none
0/10

Data access export — stories about data access export in this arenaData access export

Data export

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Data access export — stories about data access export in this arenaBulk-export traces and datasets to blob storage or my data warehousedeveloper
fullX
8/10
partialC
3/10
partialC
5/10
none
0/10
none
0/10
partialX
4/10
none
0/10
none
0/10

Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets

Ai eval ops

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Evals datasets — measuring quality — datasets, eval runs, regression trackingHave an agent create a dataset, trigger an eval run programmatically, and read back the resultsai-native
partialC
7/10
fullT
7/10
fullC
8/10
partialC
6/10
fullT
8/10
none
0/10
fullT
7/10
fullT
8/10

Human review

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Evals datasets — measuring quality — datasets, eval runs, regression trackingRoute outputs to human annotation queues for review and labelingml-engineer
fullC
8/10
fullC
8/10
partialC
5/10
partialC
6/10
none
0/10
none
0/10
partialC
4/10
none
0/10

Offline evals

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Evals datasets — measuring quality — datasets, eval runs, regression trackingRun evals in CI and gate deployments on their resultsdeveloper
fullC
7/10
partialC
5/10
fullC
8/10
partialC
5/10
partialC
4/10
none
0/10
none
0/10
partialC
6/10
Evals datasets — measuring quality — datasets, eval runs, regression trackingWrite custom code-based scorers and metrics for my evaluationsml-engineer
fullC
7/10
fullC
8/10
fullC
8/10
fullC
8/10
fullC
9/10
partialC
3/10
partialC
6/10
fullC
8/10
Evals datasets — measuring quality — datasets, eval runs, regression trackingCompare eval runs side by side to catch regressions between prompt or model versionsml-engineer
fullC
8/10
fullC
8/10
fullC
9/10
fullC
8/10
fullC
8/10
partialC
3/10
partialC
5/10
partialC
6/10
Evals datasets — measuring quality — datasets, eval runs, regression trackingScore outputs with configurable LLM-as-a-judge evaluatorsml-engineer
fullC
8/10
fullC
8/10
fullC
8/10
fullC
9/10
fullC
9/10
partialC
4/10
fullC
8/10
fullC
8/10
Evals datasets — measuring quality — datasets, eval runs, regression trackingCurate datasets from production traces and run offline evaluations against themml-engineer
fullC
8/10
fullC
8/10
fullC
9/10
fullC
8/10
fullC
8/10
partialC
4/10
fullC
8/10
partialC
5/10

Online evals

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Evals datasets — measuring quality — datasets, eval runs, regression trackingRun evaluators continuously on live production traffic, not just offline datasetsml-engineer
fullC
7/10
fullC
8/10
fullC
8/10
partialC
5/10
partialC
6/10
partialC
5/10
fullC
8/10
fullT
8/10

Openness — open source, data portability, and self-hosting storiesOpenness

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UIai-native
partialT
6/10
partialT
5/10
partialT
7/10
partialT
5/10
partialT
6/10
partialT
6/10
partialT
5/10
partialT
7/10
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leaveai-native
partialX
6/10
partialX
4/10
partialC
5/10
partialC
6/10
partialC
3/10
partialX
5/10
none
0/10
none
0/10
Openness — open source, data portability, and self-hosting storiesRead the product's source under an open licenseai-native
fullT
8/10
none
0/10
none
0/10
partialC
5/10
none
0/10
disputedD
5/10
partialC
3/10
none
0/10
Openness — open source, data portability, and self-hosting storiesSelf-host the core productai-native
fullT
9/10
partialX
6/10
partialC
6/10
fullC
9/10
none
0/10
partialX
6/10
none
0/10
none
0/10

Privacy posture — data-handling and privacy storiesPrivacy posture

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency)ai-native
disputedD
5/10
partialC
4/10
partialC
6/10
partialC
6/10
none
0/10
partialC
4/10
none
0/10
none
0/10
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI modelsai-native
partialC
3/10
none
0/10
none
0/10
partialC
6/10
none
0/10
n/a
none
0/10
none
0/10
Privacy posture — data-handling and privacy storiesControl data retention and deletionai-native
partialX
3/10
none
0/10
partialC
4/10
partialC
5/10
none
0/10
none
0/10
none
0/10
none
0/10
Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage trackingai-native
none
0/10
none
0/10
none
0/10
partialC
6/10
none
0/10
partialC
3/10
none
0/10
none
0/10

Prompt management — stories about prompt management in this arenaPrompt management

Prompt workflow

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Prompt management — stories about prompt management in this arenaIterate on prompts in a playground against real models and variablesdeveloper
fullX
7/10
none
0/10
fullC
8/10
fullC
8/10
fullC
8/10
partialC
6/10
partialC
5/10
none
0/10
Prompt management — stories about prompt management in this arenaVersion prompts and deploy changes to production without shipping codedeveloper
fullX
8/10
partialC
4/10
partialC
6/10
fullC
8/10
none
0/10
fullC
8/10
none
0/10
n/a

Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation

Ai observability

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Tracing instrumentation — instrumenting code and tracing requests end to endHave an agent query my traces, metrics, and eval results through an API or MCP server to debug my appai-native
partialT
6/10
partialT
6/10
fullT
9/10
fullT
8/10
fullT
8/10
partialT
5/10
partialT
5/10
partialT
6/10

Data controls

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Tracing instrumentation — instrumenting code and tracing requests end to endMask or redact sensitive data before it is stored in tracesdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
fullC
8/10

Sdk coverage

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Tracing instrumentation — instrumenting code and tracing requests end to endInstrument apps in both Python and JS/TS with officially supported SDKsdeveloper
fullX
8/10
fullX
8/10
partialC
6/10
partialC
6/10
fullC
8/10
partialX
6/10
partialC
4/10
partialT
3/10

Trace capture

StoryPersona
Langfuse logoLangfuse
LangSmith logoLangSmith
Braintrust logoBraintrust
Arize Phoenix logoArize Phoenix
W&B Weave logoW&B Weave
Helicone logoHelicone
Galileo logoGalileo
Cekura logoCekura
Tracing instrumentation — instrumenting code and tracing requests end to endTrace multi-step agent runs as nested spans grouped into sessions or threadsdeveloper
fullX
8/10
partialX
6/10
partialC
6/10
fullC
8/10
fullC
9/10
fullC
8/10
fullC
8/10
none
0/10
Tracing instrumentation — instrumenting code and tracing requests end to endInstrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDKdeveloper
fullX
7/10
partialX
6/10
partialX
4/10
fullC
8/10
partialC
5/10
fullX
8/10
none
0/10
partialC
3/10
Tracing instrumentation — instrumenting code and tracing requests end to endCapture multimodal payloads (images, audio, files) inside my tracesdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
partialC
5/10
Tracing instrumentation — instrumenting code and tracing requests end to endSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary formatdeveloper
fullC
8/10
fullC
7/10
none
0/10
fullC
9/10
partialC
7/10
none
0/10
partialC
5/10
none
0/10
Tracing instrumentation — instrumenting code and tracing requests end to endCapture traces of my LLM calls with inputs, outputs, latency, and token usagedeveloper
fullX
9/10
fullX
8/10
fullC
8/10
fullC
9/10
fullC
9/10
fullX
9/10
fullC
7/10
partialT
4/10
Verdict✓ fullclear evidence~ partialwith caveats! disputedevidence conflicts— noneno evidence foundn/aquestion doesn't apply to this kind of product
ProofT probedtested by usX communityusers back itC claimedvendor claim onlyD contradictedevidence disagrees⚿ auth-gatedprobe hit a live sign-in wall — verified reachable, untestable keylessly
quality 0–10 · PA Score /100 · A–D = evidence confidence · full guide

Adjacent arenas — categories often shopped togetherAdjacent arenas

Shopping this category often means shopping these too.