Skip to content

Rank #1 of 8 in LLM Evals & Observability

Braintrust logo

Braintrust Data, Inc. · commercial

279/yrnpm 1.3M/wk ±0npm/wk -364.2kpypi/wk -51.5k

Access

Install

pippip install braintrust openai autoevals
installercurl -fsSL https://bt.dev/cli/install.sh | sh

Vendor-official, but review any script before piping it to a shell.

Compare head-to-head

Alternatives to Braintrust

Showcase

Braintrust homepage screenshot
homepage · captured Sep 2026 · view live ↗
Braintrust docs screenshot
docs · captured Sep 2026 · view live ↗

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

45.9/100

Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboardsevidence →

Stories about alerting dashboards in this arena

15.0/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

40.0/100

Cost monitoring — stories about cost monitoring in this arenaCost monitoringevidence →

Stories about cost monitoring in this arena

27.6/100

Data access export — stories about data access export in this arenaData access exportevidence →

Stories about data access export in this arena

30.0/100

Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasetsevidence →

Measuring quality — datasets, eval runs, regression tracking

77.9/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

28.2/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

13.3/100

Prompt management — stories about prompt management in this arenaPrompt managementevidence →

Stories about prompt management in this arena

53.6/100

Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentationevidence →

Instrumenting code and tracing requests end to end

39.0/100

Story verdicts — every judged story with its evidenceStory verdicts

?

Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full7/10C

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3none0/10

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial±6/10T

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial±6/10C

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial4/10C

Compare eval runs side by side to catch regressions between prompt or model versions C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3full9/10C

Curate datasets from production traces and run offline evaluations against them C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3full9/10C

Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app C

Ai observability

ai-native userTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3full9/10T

Capture traces of my LLM calls with inputs, outputs, latency, and token usage C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3full8/10C

Score outputs with configurable LLM-as-a-judge evaluators C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3full8/10C

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partial6/10C

Version prompts and deploy changes to production without shipping code C

Prompt workflow

developerPrompt management — stories about prompt management in this arenaPrompt management3partial6/10C

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial5/10C

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partial5/10C

See cost and token usage per request, model, and time period in dashboards C

Cost tracking

developerCost monitoring — stories about cost monitoring in this arenaCost monitoring3partial5/10C

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3noneuntestednone yet

Have an agent create a dataset, trigger an eval run programmatically, and read back the results C

Ai eval ops

ai-native userEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10C

Iterate on prompts in a playground against real models and variables C

Prompt workflow

developerPrompt management — stories about prompt management in this arenaPrompt management2full8/10C

Run evals in CI and gate deployments on their results C

Offline evals

developerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10C

Run evaluators continuously on live production traffic, not just offline datasets C

Online evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10C

Write custom code-based scorers and metrics for my evaluations C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2full8/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial7/10T

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2full7/10C

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partial6/10C

Instrument apps in both Python and JS/TS with officially supported SDKs G

Sdk coverage

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2partial6/10C

Trace multi-step agent runs as nested spans grouped into sessions or threads C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2partial6/10C

Build custom dashboards over latency, error, cost, and eval-score metrics C

Monitoring

ml engineerAlerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards2partial5/10C

Bulk-export traces and datasets to blob storage or my data warehouse C

Data export

developerData access export — stories about data access export in this arenaData access export2partial5/10C

Route outputs to human annotation queues for review and labeling C

Human review

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2partial5/10C

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial5/10C

Attribute cost and usage to users, sessions, and features via custom metadata C

Cost tracking

developerCost monitoring — stories about cost monitoring in this arenaCost monitoring2partial4/10C

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partial4/10C

Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2partial4/10X

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2none0/10

Mask or redact sensitive data before it is stored in traces C

Data controls

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email C

Monitoring

developerAlerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards2noneuntestednone yet

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial5/10C

Capture multimodal payloads (images, audio, files) inside my traces C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation1none0/10

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 33 stories with headroom

What would move Braintrust’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools

    nonemoves agent-readyimpact 45

    All MCP evidence describes Braintrust exposing an MCP server that other clients (Claude Code, Cursor, Codex) connect to in order to use Braintrust's tools — the reverse of this story, which asks whether Braintrust can consume external MCP servers' tools.

  2. Tracing instrumentation — instrumenting code and tracing requests end to endSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary format

    nonemoves PA Scoreimpact 30

    Missing: any mention of OTLP endpoint, OpenTelemetry SDK compatibility, or OTel collector integration.

  3. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    Missing: any explicit no-training-on-customer-data policy, opt-out controls, or terms-of-service statement about AI training use.

  4. Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent

    nonemoves agent-readyimpact 30

    The evidence describes Braintrust's general API, CLI, and MCP integrations but contains no mention of scoped, role-based, or least-privilege API key/credential issuance for agents; the only security-related item is a breach report telling customers to rotate keys, which does not demonstrate a scoping/least-privilege capability.

  5. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    No evidence in the pack mentions webhooks or any event-subscription mechanism; Braintrust's documented interfaces are API, CLI, MCP server, and UI, none of which are shown to support webhook subscriptions.

  6. Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples

    nonemoves API qualityimpact 30

    Missing: evidence of an interactive API console, runnable code snippets embedded in the reference, or a machine-readable OpenAPI spec powering such interactivity.

  7. Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)

    nonemoves API qualityimpact 30

    Braintrust documents a REST API (api-reference) but explicit probes for OpenAPI/swagger specs at all standard paths returned 404, and no docs mention a downloadable machine-readable spec.

  8. Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy

    nonemoves API qualityimpact 30

    Missing: versioning scheme documentation, explicit deprecation policy, changelog/migration guides.

Showing the top 8 of 33 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 39 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

docs39 stories

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

7 of 24 testable claims verified · 1 contradictedintegrity 21/100

28 distinct capability claims found in Braintrust’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

7

Verified

16

Unverified

1

Contradicted

16

Undersold

Verified (11)
Unverified (24)
Contradicted (2)
Undersold (16)

Business model

free-tiersubscription-flatusage-basedenterprise-custom

Free tier with capped traces and scores; Pro is a flat monthly platform fee plus usage-based ingestion/processing overages; Enterprise (incl. hybrid self-hosting) is custom.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score32 (Sep 4 '26)38 (Sep 4 '26)
Agent-ready53 (Sep 4 '26)51 (Sep 4 '26)

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)