How Langfuse’s scores are calculated
The full audit trail, recomputed from the verdict data at build time through the same code that produced the leaderboard: verdict × quality × story weight per cell, cells sum to dimension scores, dimensions blend into the PA Score. Every number on the product page is reproducible from this page alone; for why the formula looks like this, see the methodology.
verdict factors: full ×1.0 · partial ×0.6 · disputed ×0.3 · none ×0.0 · n/a excluded from both sides · cell points = weight × quality × factor · cell max = weight × 10
PA Score34/100
Agent-ready 51.7 × 0.30 = 15.51
API quality 3.4 × 0.20 = 0.68
Openness 61.0 × 0.20 = 12.20
Built-in AI 18.7 × 0.15 = 2.80
Automation 17.3 × 0.15 = 2.60
(15.51 + 0.68 + 12.20 + 2.80 + 2.60) ÷ (0.30 + 0.20 + 0.20 + 0.15 + 0.15) = 33.79 ÷ 1.00 = 33.8
Scores are stored to 1 decimal; the product page’s pills round to whole numbers for display. Each dimension below shows the stories, verdicts, and cited evidence behind its number.
Agent-ready51.7/100×0.30 of the PA blend
Outside-in: can YOUR agent reach and drive this product — API, MCP, CLI, headless runs, agent docs.
Point an agent at llms.txt or agent-oriented docsweight 2
2 (weight) × 9 (quality) × 1.0 (full) = 18.0 of 20 max
- [probe] https://langfuse.com/llms.txt“PROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](https://github.com/langfuse/langfuse)) th”
- [probe] https://langfuse.com/docs.md“PROBE docs-md: HTTP 200 at https://langfuse.com/docs.md --- title: Overview seoTitle: Open Source AI Engineering Platform description: Langfuse is an open-source AI engineering”
Run the product headlessly / in CI for automationweight 2
2 (weight) × 6 (quality) × 0.6 (partial) = 7.2 of 20 max
- [claimed-docs] https://langfuse.com/docs/observability/overview“Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Block deploys on regressions | [CI/CD experiments](/docs/evaluation/experiments/experiments-ci-cd)”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [claimed-docs] https://langfuse.com/integrations/native/opentelemetry“Langfuse can receive traces on the `/api/public/otel` (OTLP) endpoint.”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Block deploys on regressions | CI/CD experiments”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Block deploys on regressions CI/CD experiments”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [claimed-docs] https://langfuse.com/self-hosting“Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.”
- [claimed-docs] https://langfuse.com/self-hosting“You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments”
Plug MCP servers into this product so it can use their toolsweight 3
n/a — not applicable to this product: excluded from numerator and denominator
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands | MCP Server”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands MCP Server”
- [probe] https://langfuse.com/docs/docs-mcp“official MCP server documented at https://langfuse.com/docs/docs-mcp”
Connect an agent via an official MCP serverweight 3
3 (weight) × 7 (quality) × 1.0 (full) = 21.0 of 30 max
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands | MCP Server”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands MCP Server”
- [probe] https://langfuse.com/docs/docs-mcp“official MCP server documented at https://langfuse.com/docs/docs-mcp”
Use an official CLIweight 2
2 (weight) × 7 (quality) × 1.0 (full) = 14.0 of 20 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Work with Langfuse from a terminal or coding agent | CLI”
Drive the product through a documented public APIweight 3
3 (weight) × 6 (quality) × 0.6 (partial) = 10.8 of 30 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Work with Langfuse from a terminal or coding agent | CLI”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands | MCP Server”
- [probe] https://langfuse.com/openapi.json“PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)”
- [community] https://news.ycombinator.com/item?id=42441258“The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.”
Issue scoped/least-privilege API credentials for an agentweight 2
2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max
no evidence cited — the verdict rests on absence of evidence, re-checked on refresh
Build against official SDKsweight 2
2 (weight) × 8 (quality) × 1.0 (full) = 16.0 of 20 max
- [claimed-docs] https://langfuse.com/docs/observability/overview“Capture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [claimed-docs] https://langfuse.com/docs/prompt-management/overview“Langfuse Prompt Management adds no latency to your application. Prompts are cached client-side by the SDK, so retrieving them is as fast as reading from memory.”
- [community] https://news.ycombinator.com/item?id=46650290“Langfuse is by far the best of the langs (-chain, -graph, -smith, -flow) in terms of UI/DX/integration/docs/quality.”
- [community] https://news.ycombinator.com/item?id=42441258“The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.”
- [community] https://hn.algolia.com/api/v1/items/46656552“Langfuse seemed way better at tracing though [vs PromptLayer]... but Langfuse's prompt management was pretty behind in terms of features when I looked a couple months back, so I went with PromptLayer for that.”
- [probe] https://langfuse.com/openapi.json“PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)”
Subscribe to events via webhooksweight 2
2 (weight) × 5 (quality) × 0.6 (partial) = 6.0 of 20 max
- [claimed-docs] https://langfuse.com/docs“Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts”
Agent-ready = 93.0 ÷ 180 × 100 = 51.7
API quality3.4/100×0.20 of the PA blend
The programmable surface once an agent is there — machine-readable spec, interactive docs, sandbox, versioning discipline.
Explore an interactive API reference with runnable examplesweight 2
2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [probe] https://langfuse.com/openapi.json“PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)”
Download a machine-readable API spec (OpenAPI or equivalent)weight 2
2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max
- [probe] https://langfuse.com/openapi.json“PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2”
Test against a sandbox environment without touching production dataweight 1
1 (weight) × 4 (quality) × 0.6 (partial) = 2.4 of 10 max
- [claimed-docs] https://langfuse.com/self-hosting“You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Build a reusable set of test cases | Datasets”
- [claimed-docs] https://langfuse.com/self-hosting“Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.”
Rely on versioned APIs with a documented deprecation policyweight 2
2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2”
- [probe] https://langfuse.com/openapi.json“PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)”
API quality = 2.4 ÷ 70 × 100 = 3.4
Openness61.0/100×0.20 of the PA blend
Can you leave, inspect, or self-host — data export, open source, portability.
Do everything through the API that I can do in the UIweight 2
2 (weight) × 6 (quality) × 0.6 (partial) = 7.2 of 20 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | Metrics API v2”
- [claimed-docs] https://langfuse.com/docs“Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)”
- [community] https://news.ycombinator.com/item?id=42441258“The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.”
- [probe] https://langfuse.com/openapi.json“PROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)”
Export all of my data in open formats and leaveweight 3
3 (weight) × 6 (quality) × 0.6 (partial) = 10.8 of 30 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [claimed-docs] https://langfuse.com/self-hosting“Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.”
- [claimed-docs] https://langfuse.com/self-hosting“When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Use the API from Python or JS/TS Query via SDKs”
- [community] https://news.ycombinator.com/item?id=42441258“The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.”
Read the product's source under an open licenseweight 2
2 (weight) × 8 (quality) × 1.0 (full) = 16.0 of 20 max
- [claimed-docs] https://langfuse.com/self-hosting“Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.”
- [claimed-docs] https://langfuse.com/self-hosting“When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.”
- [probe] https://langfuse.com/llms.txt“PROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](https://github.com/langfuse/langfuse)) th”
- [probe] https://langfuse.com/docs.md“PROBE docs-md: HTTP 200 at https://langfuse.com/docs.md --- title: Overview seoTitle: Open Source AI Engineering Platform description: Langfuse is an open-source AI engineering”
- [community] https://news.ycombinator.com/item?id=42441258“Been using Langfuse OSS for almost 15 months from the start. By far the best solution. No dark patterns found in other projects such as Portkey.”
- [community] https://news.ycombinator.com/item?id=42441258“The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.”
Self-host the core productweight 3
3 (weight) × 9 (quality) × 1.0 (full) = 27.0 of 30 max
- [claimed-docs] https://langfuse.com/self-hosting“Langfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.”
- [claimed-docs] https://langfuse.com/self-hosting“When self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.”
- [claimed-docs] https://langfuse.com/self-hosting“Kubernetes (Helm) ... AWS (Terraform) ... Azure (Terraform) ... GCP (Terraform)”
- [claimed-docs] https://langfuse.com/self-hosting“Kubernetes (Helm) | AWS (Terraform) | Azure (Terraform) | GCP (Terraform)”
- [claimed-docs] https://langfuse.com/self-hosting“You can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments”
- [claimed-docs] https://langfuse.com/self-hosting“For production-scale deployments, we recommend one of the following options: Kubernetes (Helm), AWS (Terraform), Azure (Terraform), GCP (Terraform)”
- [community] https://news.ycombinator.com/item?id=42441258“Been using Langfuse OSS for almost 15 months from the start. By far the best solution. No dark patterns found in other projects such as Portkey.”
- [probe] https://langfuse.com/llms.txt“PROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](https://github.com/langfuse/langfuse)) th”
Openness = 61.0 ÷ 100 × 100 = 61.0
Built-in AI18.7/100×0.15 of the PA blend
Inside-out: how agentic the product itself is for its users — built-in assistants, autonomous features.
Get AI-generated insights and suggestions from my data inside the productweight 2
2 (weight) × 4 (quality) × 0.6 (partial) = 4.8 of 20 max
- [claimed-docs] https://langfuse.com/docs/evaluation/evaluation-methods/llm-as-a-judge“LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK”
- [claimed-docs] https://langfuse.com/docs/evaluation/evaluation-methods/llm-as-a-judge“In Langfuse, that score can be numeric, categorical, or boolean.”
- [claimed-docs] https://langfuse.com/docs“Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts”
Set up automations that run autonomously in the backgroundweight 2
2 (weight) × 4 (quality) × 0.6 (partial) = 4.8 of 20 max
- [claimed-docs] https://langfuse.com/docs“Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK”
- [claimed-docs] https://langfuse.com/docs/observability/features/token-and-cost-tracking“Set up alerts: get notified automatically when spend crosses a threshold.”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Export large volumes on a schedule | Blob Storage Export”
Delegate tasks to a built-in AI assistant inside the productweight 3
3 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 30 max
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands | MCP Server”
- [probe] https://langfuse.com/docs/docs-mcp“official MCP server documented at https://langfuse.com/docs/docs-mcp”
Operate the product with natural-language commandsweight 2
2 (weight) × 6 (quality) × 0.6 (partial) = 7.2 of 20 max
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands | MCP Server”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Connect an AI tool that cannot run shell commands MCP Server”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Work with Langfuse from a terminal or coding agent | CLI”
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Core use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [claimed-docs] https://langfuse.com/docs/docs-mcp“Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase”
- [probe] https://langfuse.com/docs/docs-mcp“official MCP server documented at https://langfuse.com/docs/docs-mcp”
Built-in AI = 16.8 ÷ 90 × 100 = 18.7
Automation17.3/100×0.15 of the PA blend
Depth of automation primitives — rules, scheduling, bulk operations, webhooks.
Perform bulk operations across many items at onceweight 2
2 (weight) × 5 (quality) × 0.3 (disputed) = 3.0 of 20 max
- [claimed-docs] https://langfuse.com/docs“Run [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Export large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)”
- [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overview“Query aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Review and rate traces manually | Annotation Queues, Scores via UI”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Build a reusable set of test cases | Datasets”
- [community] https://news.ycombinator.com/item?id=42441258“The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.”
Define rules that trigger actions automatically on eventsweight 3
3 (weight) × 4 (quality) × 0.6 (partial) = 7.2 of 30 max
- [claimed-docs] https://langfuse.com/docs“Get notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK”
- [claimed-docs] https://langfuse.com/docs/observability/features/token-and-cost-tracking“Set up alerts: get notified automatically when spend crosses a threshold.”
- [claimed-docs] https://langfuse.com/docs/evaluation/evaluation-methods/llm-as-a-judge“LLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.”
Schedule recurring jobs or workflowsweight 2
2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max
no evidence cited — the verdict rests on absence of evidence, re-checked on refresh
Version, review, and roll back my automationsweight 1
1 (weight) × 6 (quality) × 0.6 (partial) = 3.6 of 10 max
- [claimed-docs] https://langfuse.com/docs/prompt-management/overview“Use version control and labels to manage deployments across environments”
- [claimed-docs] https://langfuse.com/docs/prompt-management/overview“Link prompts to traces to analyze performance by prompt version”
- [claimed-docs] https://langfuse.com/docs/prompt-management/overview“When prompts live in Langfuse, non-technical team members update them directly in the UI while your application automatically fetches the latest version.”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Review and rate traces manually | Annotation Queues, Scores via UI”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Compare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments via SDK]”
- [claimed-docs] https://langfuse.com/docs/evaluation/overview“Automatically score live production traces LLM-as-a-Judge, Scores via API/SDK”
Automation = 13.8 ÷ 80 × 100 = 17.3