Rank #6 of 8 in LLM Evals & Observability
Install
Showcase


Verified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboardsevidence →
Stories about alerting dashboards in this arena
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Cost monitoring — stories about cost monitoring in this arenaCost monitoringevidence →
Stories about cost monitoring in this arena
Data access export — stories about data access export in this arenaData access exportevidence →
Stories about data access export in this arena
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasetsevidence →
Measuring quality — datasets, eval runs, regression tracking
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Prompt management — stories about prompt management in this arenaPrompt managementevidence →
Stories about prompt management in this arena
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentationevidence →
Instrumenting code and tracing requests end to end
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 8 free · 0 paid · 0 enterprise · 29 not stated in evidence
Follow the green: where the map greys out is where Arize Phoenix stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓7/10
unlocks → Webhooks · Scoped API keys · Machine-readable spec · Versioning policy
Subscribe to events via webhooks
—–
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
~5/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
✓7/10
unlocks → Autonomous automations
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
~5/10
Set up automations that run autonomously in the background
—–
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards
Stories about alerting dashboards in this arena
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Cost monitoring — stories about cost monitoring in this arenaCost monitoring
Stories about cost monitoring in this arena
Data access export — stories about data access export in this arenaData access export
Stories about data access export in this arena
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets
Measuring quality — datasets, eval runs, regression tracking
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Prompt management — stories about prompt management in this arenaPrompt management
Stories about prompt management in this arena
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation
Instrumenting code and tracing requests end to end
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
✓8/10
Mask or redact sensitive data before it is stored in traces
—–
Instrument apps in both Python and JS/TS with officially supported SDKs
~6/10
Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 7/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialfree | 7/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partialfree | 5/10 | Cclaimed | |
Capture traces of my LLM calls with inputs, outputs, latency, and token usage C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | full | 9/10 | Cclaimed | |
Score outputs with configurable LLM-as-a-judge evaluators C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 9/10 | Cclaimed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullfree | 9/10 | Cclaimed | |
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | full | 9/10 | Cclaimed | |
Compare eval runs side by side to catch regressions between prompt or model versions C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 8/10 | Cclaimed | |
Curate datasets from production traces and run offline evaluations against them C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 8/10 | Cclaimed | |
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app C Ai observability | ai-native user | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | full | 8/10 | Tprobed | |
Version prompts and deploy changes to production without shipping code C Prompt workflow | developer | Prompt management — stories about prompt management in this arenaPrompt management | 3 | full | 8/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partialfree | 6/10 | Cclaimed | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | partialfree | 6/10 | Cclaimed | |
See cost and token usage per request, model, and time period in dashboards C Cost tracking | developer | Cost monitoring — stories about cost monitoring in this arenaCost monitoring | 3 | partial | 5/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | none | untested | none yet | |
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | full | 8/10 | Cclaimed | |
Iterate on prompts in a playground against real models and variables C Prompt workflow | developer | Prompt management — stories about prompt management in this arenaPrompt management | 2 | full | 8/10 | Cclaimed | |
Trace multi-step agent runs as nested spans grouped into sessions or threads C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | full | 8/10 | Cclaimed | |
Write custom code-based scorers and metrics for my evaluations C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | full | 8/10 | Cclaimed | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialfree | 6/10 | Cclaimed | |
Have an agent create a dataset, trigger an eval run programmatically, and read back the results C Ai eval ops | ai-native user | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 6/10 | Cclaimed | |
Instrument apps in both Python and JS/TS with officially supported SDKs G Sdk coverage | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | partial | 6/10 | Cclaimed | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialfree | 6/10 | Cclaimed | |
Route outputs to human annotation queues for review and labeling C Human review | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 6/10 | Cclaimed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialfree | 5/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Tprobed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 5/10 | Cclaimed | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Cclaimed | |
Run evals in CI and gate deployments on their results C Offline evals | developer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 5/10 | Cclaimed | |
Run evaluators continuously on live production traffic, not just offline datasets C Online evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 5/10 | Cclaimed | |
Attribute cost and usage to users, sessions, and features via custom metadata C Cost tracking | developer | Cost monitoring — stories about cost monitoring in this arenaCost monitoring | 2 | partial | 4/10 | Cclaimed | |
Bulk-export traces and datasets to blob storage or my data warehouse C Data export | developer | Data access export — stories about data access export in this arenaData access export | 2 | none | 0/10 | ||
Build custom dashboards over latency, error, cost, and eval-score metrics C Monitoring | ml engineer | Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards | 2 | none | untested | none yet | |
Mask or redact sensitive data before it is stored in traces C Data controls | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email C Monitoring | developer | Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards | 2 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Cclaimed | |
Capture multimodal payloads (images, audio, files) inside my traces C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 1 | none | 0/10 |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 34 stories with headroom
What would move Arize Phoenix’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Phoenix is an observability/evaluation platform; the evidence describes tracing, evals, prompt management, datasets, and an MCP server that lets *external* agents (Claude Code, Cursor, etc.) operate on Phoenix data — not a built-in AI assistant living inside Phoenix that users delegate tasks to.
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events
nonemoves PA Scoreimpact 30
Phoenix's evidence covers tracing, evaluation, datasets, prompt management, and MCP integration, but nothing describes a rules/triggers engine that automatically fires actions on events (e.g., alerting, auto-remediation, webhooks on thresholds).
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background
nonemoves Built-in AIimpact 30
Phoenix's evidence covers tracing, evaluation, prompt management, and datasets, but nothing describes scheduled or autonomous background automations (e.g., recurring eval jobs, alerting rules, or triggers) that run without user initiation.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
No evidence in the pack describes scoped or least-privilege API key/credential issuance for agents; Phoenix's docs cover tracing, evaluation, prompt management, and an MCP endpoint, but nothing about credential scoping or access control granularity.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
No evidence anywhere in the pack of a webhook subscription mechanism; Phoenix's integration surface is OTLP tracing ingestion, an MCP server, and SDKs, but nothing about outbound event webhooks for subscribing to Phoenix events.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
Docs mention an 'sdk-api-reference' page listing decorators and SDK features, but there is no evidence of an interactive, runnable API reference (e.g., a Swagger/OpenAPI explorer or live code sandbox); a direct probe for OpenAPI/swagger specs returned 404 on all candidate paths, indicating no such interactive reference is discoverable.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
Missing: any OpenAPI/Swagger spec, documented REST API reference, or SDK-generated schema.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Evidence shows only generic container/image version pinning (e.g., 'version-8.0.0' Docker tags) but no documented API versioning scheme or deprecation policy for Phoenix's SDK/API; an OpenAPI probe also returned 404s, finding no formal API spec to review versioning against.
Showing the top 8 of 34 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map5 surfaces · 37 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs37 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Version, review, and roll back my automations
- Attribute cost and usage to users, sessions, and features via custom metadata
- See cost and token usage per request, model, and time period in dashboards
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Route outputs to human annotation queues for review and labeling
- Run evals in CI and gate deployments on their results
- Write custom code-based scorers and metrics for my evaluations
- Compare eval runs side by side to catch regressions between prompt or model versions
- Score outputs with configurable LLM-as-a-judge evaluators
- Curate datasets from production traces and run offline evaluations against them
- Run evaluators continuously on live production traffic, not just offline datasets
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- Choose where my data is stored (region/residency)
- Prevent my data from being used to train AI models
- Control data retention and deletion
- Opt out of telemetry and usage tracking
- Iterate on prompts in a playground against real models and variables
- Version prompts and deploy changes to production without shipping code
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Instrument apps in both Python and JS/TS with officially supported SDKs
- Trace multi-step agent runs as nested spans grouped into sessions or threads
- Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK
- Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format
- Capture traces of my LLM calls with inputs, outputs, latency, and token usage
GitHub README6 stories
- Build against official SDKs
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Compare eval runs side by side to catch regressions between prompt or model versions
- Curate datasets from production traces and run offline evaluations against them
- Export all of my data in open formats and leave
- Read the product's source under an open license
Phoenix docs4 stories
OpenAPI spec2 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
2 of 20 testable claims verified · 1 contradicted → integrity 0/100
39 distinct capability claims found in Arize Phoenix’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
2
Verified
17
Unverified
1
Contradicted
18
Undersold
Verified (3)
“Query and filter spans with powerful filtering capabilities”
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my appfullproof ↗
“MCP endpoint lets any MCP-compatible client search, query, and operate on projects, traces, datasets, experiments, prompts, and annotations”
“MCP endpoint lets any MCP-compatible client search, query, and operate on projects, traces, datasets, experiments, prompts, and annotations”
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my appfullproof ↗
Unverified (37)
“Send detailed trace logs from your app to see what happened during a run”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Score outputs with evaluation tests to catch failures and regressions”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Score outputs with evaluation tests to catch failures and regressions”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Iterate on prompts using real production examples”
Iterate on prompts in a playground against real models and variablesfullproof ↗
“Iterate on prompts using real production examples”
Version prompts and deploy changes to production without shipping codefullproof ↗
“Run experiments that compare changes across the same inputs”
Compare eval runs side by side to catch regressions between prompt or model versionsfullproof ↗
“Accepts traces over OTLP with auto-instrumentation for popular frameworks”
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary formatfullproof ↗
“Accepts traces over OTLP with auto-instrumentation for popular frameworks”
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDKfullproof ↗
“Attach human ground-truth annotations directly in the UI”
Route outputs to human annotation queues for review and labelingpartialproof ↗
“Version, store, and deploy prompts (prompt management)”
Version prompts and deploy changes to production without shipping codefullproof ↗
“Sync prompts across environments via SDK (prompts in code)”
Version prompts and deploy changes to production without shipping codefullproof ↗
“Group traces into datasets, rerun through app versions, and compare eval results”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Group traces into datasets, rerun through app versions, and compare eval results”
Compare eval runs side by side to catch regressions between prompt or model versionsfullproof ↗
“Coding agent can auto-set-up tracing via a CLI command”
“Group related traces into conversations and user sessions”
Trace multi-step agent runs as nested spans grouped into sessions or threadsfullproof ↗
“Supports deterministic code-based evaluators and LLM-as-a-judge evaluators”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Supports deterministic code-based evaluators and LLM-as-a-judge evaluators”
Write custom code-based scorers and metrics for my evaluationsfullproof ↗
“Configure evaluators in the UI and auto-score experiment results with no code”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Model-agnostic judge via adapters (OpenAI, LiteLLM, LangChain, AI SDK, etc.)”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“LLM evaluations return explanations by default for richer signal”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Prompt playground to test prompts/models/params while tracking via tracing and experiments”
Iterate on prompts in a playground against real models and variablesfullproof ↗
“Datasets can be collected from production, staging, evaluations, or manual entry”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Dataset evaluators act as automated test cases scoring experiment outputs, like a unit test suite”
Run evals in CI and gate deployments on their resultspartialproof ↗
“Dataset evaluators act as automated test cases scoring experiment outputs, like a unit test suite”
Write custom code-based scorers and metrics for my evaluationsfullproof ↗
“Free to self-host with no feature limits; data stays fully within your infra, air-gappable”
“Free to self-host with no feature limits; data stays fully within your infra, air-gappable”
Choose where my data is stored (region/residency)partialproof ↗
“Zero-config tracing via auto_instrument=True for supported AI libraries”
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDKfullproof ↗
“Manual instrumentation via tracing decorators like @tracer.chain and @tracer.tool”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Single-command server start via uvx arize-phoenix serve”
“Create versioned datasets of examples for experiments”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Modular Python SDK lets you install only the components you need”
Instrument apps in both Python and JS/TS with officially supported SDKspartialproof ↗
“Traces capture model calls, retrieval, tool use, and custom logic for debugging”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Score traces and spans with LLM-based evaluators, code-based checks, or human labels”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Score traces and spans with LLM-based evaluators, code-based checks, or human labels”
Write custom code-based scorers and metrics for my evaluationsfullproof ↗
“Score traces and spans with LLM-based evaluators, code-based checks, or human labels”
Route outputs to human annotation queues for review and labelingpartialproof ↗
“Breakdown of token usage per LLM call to find and optimize expensive invocations”
See cost and token usage per request, model, and time period in dashboardspartialproof ↗
“Built on OpenTelemetry and powered by OpenInference instrumentation”
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary formatfullproof ↗
Contradicted (1)
“Pin deployments to a specific Docker image version for production”
Rely on versioned APIs with a documented deprecation policynoneproof ↗
Undersold (18)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationpartialproof ↗
Drive the product through a documented public APIfullproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Attribute cost and usage to users, sessions, and features via custom metadatapartialproof ↗
Have an agent create a dataset, trigger an eval run programmatically, and read back the resultspartialproof ↗
Run evaluators continuously on live production traffic, not just offline datasetspartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Read the product's source under an open licensepartialproof ↗
Prevent my data from being used to train AI modelspartialproof ↗
Claims outside our story set (8)
Real capability claims found in Arize Phoenix’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Debug by replaying LLM calls with different inputs (span replay)”
source ↗“Identify and address slow invocations of LLMs, retrievers, and other components”
source ↗“Inspect retrieved documents from a Retriever call including score and order”
source ↗“View tool descriptions and function signatures available to the LLM”
source ↗“Organize traces into separate projects per application”
source ↗“Run thousands of evaluations without writing retry or concurrency logic”
source ↗“Use trace viewer to explore eval traces and identify systematic evaluator bias”
source ↗“Every evaluation run captures inputs, exact judge prompts, model reasoning, scores, and timing”
source ↗
Business model
Phoenix is fully open source (ELv2) and free to self-host with no feature gates; Phoenix Cloud has a free tier, and Arize's paid AX platform (usage/custom pricing) sits above it.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
