Rank #3 of 8 in LLM Evals & Observability
Install
pip install weave openaiShowcase


Verified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboardsevidence →
Stories about alerting dashboards in this arena
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Cost monitoring — stories about cost monitoring in this arenaCost monitoringevidence →
Stories about cost monitoring in this arena
Data access export — stories about data access export in this arenaData access exportevidence →
Stories about data access export in this arena
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasetsevidence →
Measuring quality — datasets, eval runs, regression tracking
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Prompt management — stories about prompt management in this arenaPrompt managementevidence →
Stories about prompt management in this arena
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentationevidence →
Instrumenting code and tracing requests end to end
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where W&B Weave stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Scoped API keys · Versioning policy
Subscribe to events via webhooks
~4/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
✓9/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
~3/10
Explore an interactive API reference with runnable examples
~3/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
n/an/a
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
~6/10
Set up automations that run autonomously in the background
~5/10
Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards
Stories about alerting dashboards in this arena
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Cost monitoring — stories about cost monitoring in this arenaCost monitoring
Stories about cost monitoring in this arena
Data access export — stories about data access export in this arenaData access export
Stories about data access export in this arena
Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets
Measuring quality — datasets, eval runs, regression tracking
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Prompt management — stories about prompt management in this arenaPrompt management
Stories about prompt management in this arena
Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation
Instrumenting code and tracing requests end to end
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
✓8/10
Mask or redact sensitive data before it is stored in traces
—0/10
Instrument apps in both Python and JS/TS with officially supported SDKs
✓8/10
Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a± | untested | none yet | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial± | 5/10 | Tprobed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Cclaimed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial± | 3/10 | Tprobed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 3/10 | Cclaimed | |
Capture traces of my LLM calls with inputs, outputs, latency, and token usage C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | full | 9/10 | Cclaimed | |
Score outputs with configurable LLM-as-a-judge evaluators C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 9/10 | Cclaimed | |
Compare eval runs side by side to catch regressions between prompt or model versions C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 8/10 | Cclaimed | |
Curate datasets from production traces and run offline evaluations against them C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 3 | full | 8/10 | Cclaimed | |
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app C Ai observability | ai-native user | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | full | 8/10 | Tprobed | |
See cost and token usage per request, model, and time period in dashboards C Cost tracking | developer | Cost monitoring — stories about cost monitoring in this arenaCost monitoring | 3 | full | 8/10 | Cclaimed | |
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 3 | partial | 7/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 5/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 3/10 | Cclaimed | |
Version prompts and deploy changes to production without shipping code C Prompt workflow | developer | Prompt management — stories about prompt management in this arenaPrompt management | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Trace multi-step agent runs as nested spans grouped into sessions or threads C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | full | 9/10 | Cclaimed | |
Write custom code-based scorers and metrics for my evaluations C Offline evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | full | 9/10 | Cclaimed | |
Have an agent create a dataset, trigger an eval run programmatically, and read back the results C Ai eval ops | ai-native user | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | full | 8/10 | Tprobed | |
Instrument apps in both Python and JS/TS with officially supported SDKs G Sdk coverage | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | full | 8/10 | Cclaimed | |
Iterate on prompts in a playground against real models and variables C Prompt workflow | developer | Prompt management — stories about prompt management in this arenaPrompt management | 2 | full | 8/10 | Cclaimed | |
Build custom dashboards over latency, error, cost, and eval-score metrics C Monitoring | ml engineer | Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards | 2 | partial | 6/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 6/10 | Tprobed | |
Run evaluators continuously on live production traffic, not just offline datasets C Online evals | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 6/10 | Cclaimed | |
Attribute cost and usage to users, sessions, and features via custom metadata C Cost tracking | developer | Cost monitoring — stories about cost monitoring in this arenaCost monitoring | 2 | partial | 5/10 | Cclaimed | |
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | partial | 5/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 5/10 | Cclaimed | |
Run evals in CI and gate deployments on their results C Offline evals | developer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | partial | 4/10 | Cclaimed | |
Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email C Monitoring | developer | Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards | 2 | partial | 4/10 | Cclaimed | |
Bulk-export traces and datasets to blob storage or my data warehouse C Data export | developer | Data access export — stories about data access export in this arenaData access export | 2 | none | 0/10 | ||
Mask or redact sensitive data before it is stored in traces C Data controls | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Route outputs to human annotation queues for review and labeling C Human review | ml engineer | Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | n/a | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 3/10 | Cclaimed | |
Capture multimodal payloads (images, audio, files) inside my traces C Trace capture | developer | Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 33 stories with headroom
What would move W&B Weave’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Weave is documented as a hosted SaaS platform (weave.init() connecting to W&B's cloud) with no evidence pack mentions of a self-hosted or on-prem deployment option for the core Weave product itself; only W&B Models/Platform is known to have enterprise self-hosting but that's not evidenced here for Weave specifically.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
Missing: any privacy policy statement, training opt-out mechanism, or data usage terms documentation.
Prompt management — stories about prompt management in this arenaVersion prompts and deploy changes to production without shipping code
nonemoves PA Scoreimpact 30
The evidence pack covers tracing, evaluation, cost tracking, and a Playground for prompt editing/model comparison, but nothing describes a prompt versioning/registry system or a mechanism to push prompt changes to production without redeploying code.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
Missing: any documentation of API key scoping, permission granularity, or credential management for agent access.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Missing: versioning scheme documentation, deprecation policy/notice process, changelog or migration guides for breaking changes.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
partialq3/10moves PA Scoreimpact 21
Missing: explicit bulk export/download feature, documented open-format export (e.g., JSON/OTLP dump of all traces/evals), and any guidance for full data portability or platform exit.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
partialq3/10moves API qualityimpact 21
Missing: confirmation of an interactive UI with 'try it now' runnable examples, evidence of live execution from the docs, and any independent confirmation of this feature.
Data access export — stories about data access export in this arenaBulk-export traces and datasets to blob storage or my data warehouse
nonemoves PA Scoreimpact 20
Evidence shows Weave has a Service API for programmatic access and OTel import for bringing trace data in, but nothing documents bulk export of traces/datasets to blob storage (S3/GCS) or a data warehouse (Snowflake/BigQuery), which is a reasonable ask for an observability/eval platform.
Showing the top 8 of 33 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 36 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Weave docs29 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Test against a sandbox environment without touching production data
- Build custom dashboards over latency, error, cost, and eval-score metrics
- Perform bulk operations across many items at once
- Version, review, and roll back my automations
- Attribute cost and usage to users, sessions, and features via custom metadata
- See cost and token usage per request, model, and time period in dashboards
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Run evals in CI and gate deployments on their results
- Write custom code-based scorers and metrics for my evaluations
- Compare eval runs side by side to catch regressions between prompt or model versions
- Score outputs with configurable LLM-as-a-judge evaluators
- Curate datasets from production traces and run offline evaluations against them
- Run evaluators continuously on live production traffic, not just offline datasets
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Iterate on prompts in a playground against real models and variables
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Instrument apps in both Python and JS/TS with officially supported SDKs
- Trace multi-step agent runs as nested spans grouped into sessions or threads
- Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK
- Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format
- Capture traces of my LLM calls with inputs, outputs, latency, and token usage
Site docs22 stories
- Connect an agent via an official MCP server
- Drive the product through a documented public API
- Subscribe to events via webhooks
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Build custom dashboards over latency, error, cost, and eval-score metrics
- Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Run evals in CI and gate deployments on their results
- Compare eval runs side by side to catch regressions between prompt or model versions
- Score outputs with configurable LLM-as-a-judge evaluators
- Curate datasets from production traces and run offline evaluations against them
- Run evaluators continuously on live production traffic, not just offline datasets
- Do everything through the API that I can do in the UI
- Iterate on prompts in a playground against real models and variables
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Trace multi-step agent runs as nested spans grouped into sessions or threads
GitHub README16 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Connect an agent via an official MCP server
- Drive the product through a documented public API
- Build against official SDKs
- Operate the product with natural-language commands
- Have an agent create a dataset, trigger an eval run programmatically, and read back the results
- Write custom code-based scorers and metrics for my evaluations
- Compare eval runs side by side to catch regressions between prompt or model versions
- Score outputs with configurable LLM-as-a-judge evaluators
- Curate datasets from production traces and run offline evaluations against them
- Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app
- Instrument apps in both Python and JS/TS with officially supported SDKs
- Trace multi-step agent runs as nested spans grouped into sessions or threads
- Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK
- Capture traces of my LLM calls with inputs, outputs, latency, and token usage
OpenAPI spec6 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
4 of 17 testable claims verified · 0 contradicted → integrity 24/100
33 distinct capability claims found in W&B Weave’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
4
Verified
13
Unverified
0
Contradicted
19
Undersold
Verified (4)
“Weave Service API exposes REST endpoints for programmatic access to Weave functionality”
Drive the product through a documented public APIfullproof ↗
“Via W&B skills and an MCP server, coding agents can read live production data, run evaluations, and execute automatic iteration loops”
“Via W&B skills and an MCP server, coding agents can read live production data, run evaluations, and execute automatic iteration loops”
Have an agent create a dataset, trigger an eval run programmatically, and read back the resultsfullproof ↗
“Via W&B skills and an MCP server, coding agents can read live production data, run evaluations, and execute automatic iteration loops”
Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my appfullproof ↗
Unverified (31)
“Manually instrument LLM calls and arbitrary functions to trace, version, and collect feedback”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Evaluate agent/app responses using LLM judges and custom scorers”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Trace and collect metrics from agents built with other SDKs via an OTel-compatible SDK”
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary formatpartialproof ↗
“Decorating/wrapping a function with weave.op automatically captures code, inputs, outputs, and execution metadata”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Threads group related traces into a session/conversation for analysis or scoring as a whole”
Trace multi-step agent runs as nested spans grouped into sessions or threadsfullproof ↗
“Evaluation object defines a dataset plus one or more scoring functions for structured test runs”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Custom Scorers let users encode their own evaluation criteria beyond built-in scorers”
Write custom code-based scorers and metrics for my evaluationsfullproof ↗
“Predefined scorers ship out of the box, including hallucination detection, summarization quality, and embedding similarity”
Score outputs with configurable LLM-as-a-judge evaluatorsfullproof ↗
“Playground supports prompt editing, message retrying, and model comparison for iterating on LLM apps”
Iterate on prompts in a playground against real models and variablesfullproof ↗
“Playground can test OpenAI-compatible API endpoints for custom model runtimes”
Iterate on prompts in a playground against real models and variablesfullproof ↗
“Saved models let users create reusable model presets for their workflow”
Iterate on prompts in a playground against real models and variablesfullproof ↗
“Automatic cost tracking captures token usage and applies built-in pricing for supported integrations with no extra code”
See cost and token usage per request, model, and time period in dashboardsfullproof ↗
“Custom cost tracking can be added via add_cost with llm_id and token cost fields for unsupported models”
Attribute cost and usage to users, sessions, and features via custom metadatapartialproof ↗
“Supports importing OpenTelemetry-compatible trace data through a dedicated endpoint”
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary formatpartialproof ↗
“Built-in and custom signals automatically capture and classify agent interactions for behavior monitoring”
Run evaluators continuously on live production traffic, not just offline datasetspartialproof ↗
“Alerts route signals through Slack notifications and trigger webhook automations”
Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or emailpartialproof ↗
“Sessions, turns, steps, tools, and sub-agents are first-class concepts for navigating agent execution”
Trace multi-step agent runs as nested spans grouped into sessions or threadsfullproof ↗
“Flexible imperative evaluation API with comparisons and visualizations to catch regressions before production”
Compare eval runs side by side to catch regressions between prompt or model versionsfullproof ↗
“New LLMs/custom models can be tested against production traces to assess performance for a use case”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Weave Guardrails provides pre-built safety/quality scorers including toxicity, bias, PII detection, and hallucinations”
Run evaluators continuously on live production traffic, not just offline datasetspartialproof ↗
“Evaluations can be aggregated into shareable leaderboards of best performers”
Compare eval runs side by side to catch regressions between prompt or model versionsfullproof ↗
“weave.op can trace any function, including calls to OpenAI, Anthropic, Google AI Studio, Hugging Face, and custom validation/transform functions”
Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDKpartialproof ↗
“Weave automatically records token usage and calculates cost per call, shown in trace tree and calls table UI”
See cost and token usage per request, model, and time period in dashboardsfullproof ↗
“Automatic cost tracking works when using weave.init() with supported LLM integrations like OpenAI, Anthropic, Cohere, or Mistral”
See cost and token usage per request, model, and time period in dashboardsfullproof ↗
“Custom cost tracking supports fine-tuned, self-hosted, or unintegrated model providers”
Attribute cost and usage to users, sessions, and features via custom metadatapartialproof ↗
“OTel integration lets users instrument with the OpenTelemetry standard while traces appear alongside other Weave data without replacing existing pipelines”
Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary formatpartialproof ↗
“Decorating functions generates a trace tree of inputs and outputs across all traced functions”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Logs and debugs language model inputs, outputs, and traces”
Capture traces of my LLM calls with inputs, outputs, latency, and token usagefullproof ↗
“Builds rigorous, apples-to-apples evaluations for language model use cases”
Curate datasets from production traces and run offline evaluations against themfullproof ↗
“Weave organizes traces into sessions and turns from the ground up, aiding multi-agent behavior analysis”
Trace multi-step agent runs as nested spans grouped into sessions or threadsfullproof ↗
“Evaluation guide covers measuring app performance against repeatable test cases, comparing changes over time, and identifying regressions”
Compare eval runs side by side to catch regressions between prompt or model versionsfullproof ↗
Undersold (19)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)fullproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Build custom dashboards over latency, error, cost, and eval-score metricspartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Run evals in CI and gate deployments on their resultspartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Instrument apps in both Python and JS/TS with officially supported SDKsfullproof ↗
Business model
Free personal tier with capped ingestion; Pro plans are per-seat with usage-based ingested-data overages; dedicated/on-prem enterprise deployments are custom-priced.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% · openapi.json 100% (30d, checked every 6h since Sep 8 '26)
