Skip to content

How Langfuse’s scores are calculated

The full audit trail, recomputed from the verdict data at build time through the same code that produced the leaderboard: verdict × quality × story weight per cell, cells sum to dimension scores, dimensions blend into the PA Score. Every number on the product page is reproducible from this page alone; for why the formula looks like this, see the methodology.

verdict factors: full ×1.0 · partial ×0.6 · disputed ×0.3 · none ×0.0 · n/a excluded from both sides · cell points = weight × quality × factor · cell max = weight × 10

PA Score34/100

Agent-ready 51.7 × 0.30 = 15.51

API quality 3.4 × 0.20 = 0.68

Openness 61.0 × 0.20 = 12.20

Built-in AI 18.7 × 0.15 = 2.80

Automation 17.3 × 0.15 = 2.60

(15.51 + 0.68 + 12.20 + 2.80 + 2.60) ÷ (0.30 + 0.20 + 0.20 + 0.15 + 0.15) = 33.79 ÷ 1.00 = 33.8

Scores are stored to 1 decimal; the product page’s pills round to whole numbers for display. Each dimension below shows the stories, verdicts, and cited evidence behind its number.

Agent-ready51.7/100×0.30 of the PA blend

Outside-in: can YOUR agent reach and drive this product — API, MCP, CLI, headless runs, agent docs.

Point an agent at llms.txt or agent-oriented docsweight 2

2 (weight) × 9 (quality) × 1.0 (full) = 18.0 of 20 max

  • [probe] https://langfuse.com/llms.txtPROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](https://github.com/langfuse/langfuse)) th
  • [probe] https://langfuse.com/docs.mdPROBE docs-md: HTTP 200 at https://langfuse.com/docs.md --- title: Overview seoTitle: Open Source AI Engineering Platform description: Langfuse is an open-source AI engineering

Run the product headlessly / in CI for automationweight 2

2 (weight) × 6 (quality) × 0.6 (partial) = 7.2 of 20 max

  • [claimed-docs] https://langfuse.com/docs/observability/overviewCapture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewBlock deploys on regressions | [CI/CD experiments](/docs/evaluation/experiments/experiments-ci-cd)
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [claimed-docs] https://langfuse.com/integrations/native/opentelemetryLangfuse can receive traces on the `/api/public/otel` (OTLP) endpoint.
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewBlock deploys on regressions | CI/CD experiments
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewBlock deploys on regressions CI/CD experiments
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [claimed-docs] https://langfuse.com/self-hostingLangfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
  • [claimed-docs] https://langfuse.com/self-hostingYou can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments

Plug MCP servers into this product so it can use their toolsweight 3

n/a — not applicable to this product: excluded from numerator and denominator

  • [claimed-docs] https://langfuse.com/docs/docs-mcpCore use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [claimed-docs] https://langfuse.com/docs/docs-mcpUse Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands | MCP Server
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands MCP Server
  • [probe] https://langfuse.com/docs/docs-mcpofficial MCP server documented at https://langfuse.com/docs/docs-mcp

Connect an agent via an official MCP serverweight 3

3 (weight) × 7 (quality) × 1.0 (full) = 21.0 of 30 max

  • [claimed-docs] https://langfuse.com/docs/docs-mcpCore use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [claimed-docs] https://langfuse.com/docs/docs-mcpUse Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands | MCP Server
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands MCP Server
  • [probe] https://langfuse.com/docs/docs-mcpofficial MCP server documented at https://langfuse.com/docs/docs-mcp

Use an official CLIweight 2

2 (weight) × 7 (quality) × 1.0 (full) = 14.0 of 20 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewWork with Langfuse from a terminal or coding agent | CLI

Drive the product through a documented public APIweight 3

3 (weight) × 6 (quality) × 0.6 (partial) = 10.8 of 30 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewWork with Langfuse from a terminal or coding agent | CLI
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands | MCP Server
  • [probe] https://langfuse.com/openapi.jsonPROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)
  • [community] https://news.ycombinator.com/item?id=42441258The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.

Issue scoped/least-privilege API credentials for an agentweight 2

2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max

no evidence cited — the verdict rests on absence of evidence, re-checked on refresh

Build against official SDKsweight 2

2 (weight) × 8 (quality) × 1.0 (full) = 16.0 of 20 max

  • [claimed-docs] https://langfuse.com/docs/observability/overviewCapture traces via our native SDKs for Python/JS, 100+ library/framework integrations, OpenTelemetry, or via an LLM Gateway such as LiteLLM
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [claimed-docs] https://langfuse.com/docs/prompt-management/overviewLangfuse Prompt Management adds no latency to your application. Prompts are cached client-side by the SDK, so retrieving them is as fast as reading from memory.
  • [community] https://news.ycombinator.com/item?id=46650290Langfuse is by far the best of the langs (-chain, -graph, -smith, -flow) in terms of UI/DX/integration/docs/quality.
  • [community] https://news.ycombinator.com/item?id=42441258The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.
  • [community] https://hn.algolia.com/api/v1/items/46656552Langfuse seemed way better at tracing though [vs PromptLayer]... but Langfuse's prompt management was pretty behind in terms of features when I looked a couple months back, so I went with PromptLayer for that.
  • [probe] https://langfuse.com/openapi.jsonPROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)

Subscribe to events via webhooksweight 2

2 (weight) × 5 (quality) × 0.6 (partial) = 6.0 of 20 max

  • [claimed-docs] https://langfuse.com/docsGet notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts

Agent-ready = 93.0 ÷ 180 × 100 = 51.7

API quality3.4/100×0.20 of the PA blend

The programmable surface once an agent is there — machine-readable spec, interactive docs, sandbox, versioning discipline.

Explore an interactive API reference with runnable examplesweight 2

2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [probe] https://langfuse.com/openapi.jsonPROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)

Download a machine-readable API spec (OpenAPI or equivalent)weight 2

2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max

  • [probe] https://langfuse.com/openapi.jsonPROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | Metrics API v2

Test against a sandbox environment without touching production dataweight 1

1 (weight) × 4 (quality) × 0.6 (partial) = 2.4 of 10 max

  • [claimed-docs] https://langfuse.com/self-hostingYou can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewBuild a reusable set of test cases | Datasets
  • [claimed-docs] https://langfuse.com/self-hostingLangfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.

Rely on versioned APIs with a documented deprecation policyweight 2

2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
  • [probe] https://langfuse.com/openapi.jsonPROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)

API quality = 2.4 ÷ 70 × 100 = 3.4

Openness61.0/100×0.20 of the PA blend

Can you leave, inspect, or self-host — data export, open source, portability.

Do everything through the API that I can do in the UIweight 2

2 (weight) × 6 (quality) × 0.6 (partial) = 7.2 of 20 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewAutomatically score live production traces LLM-as-a-Judge, Scores via API/SDK
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | Metrics API v2
  • [claimed-docs] https://langfuse.com/docsRun [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewExport large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
  • [community] https://news.ycombinator.com/item?id=42441258The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.
  • [probe] https://langfuse.com/openapi.jsonPROBE openapi: all candidate paths 404 (https://langfuse.com/openapi.json, https://langfuse.com/swagger.json, https://langfuse.com/api/openapi.json, https://langfuse.com/.well-known/openapi.json)

Export all of my data in open formats and leaveweight 3

3 (weight) × 6 (quality) × 0.6 (partial) = 10.8 of 30 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewExport large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [claimed-docs] https://langfuse.com/self-hostingLangfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
  • [claimed-docs] https://langfuse.com/self-hostingWhen self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewUse the API from Python or JS/TS Query via SDKs
  • [community] https://news.ycombinator.com/item?id=42441258The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.

Read the product's source under an open licenseweight 2

2 (weight) × 8 (quality) × 1.0 (full) = 16.0 of 20 max

  • [claimed-docs] https://langfuse.com/self-hostingLangfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
  • [claimed-docs] https://langfuse.com/self-hostingWhen self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
  • [probe] https://langfuse.com/llms.txtPROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](https://github.com/langfuse/langfuse)) th
  • [probe] https://langfuse.com/docs.mdPROBE docs-md: HTTP 200 at https://langfuse.com/docs.md --- title: Overview seoTitle: Open Source AI Engineering Platform description: Langfuse is an open-source AI engineering
  • [community] https://news.ycombinator.com/item?id=42441258Been using Langfuse OSS for almost 15 months from the start. By far the best solution. No dark patterns found in other projects such as Portkey.
  • [community] https://news.ycombinator.com/item?id=42441258The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.

Self-host the core productweight 3

3 (weight) × 9 (quality) × 1.0 (full) = 27.0 of 30 max

  • [claimed-docs] https://langfuse.com/self-hostingLangfuse is open source and can be self-hosted using Docker. This section contains guides for different deployment scenarios.
  • [claimed-docs] https://langfuse.com/self-hostingWhen self-hosting Langfuse, you run the same infrastructure that powers Langfuse Cloud.
  • [claimed-docs] https://langfuse.com/self-hostingKubernetes (Helm) ... AWS (Terraform) ... Azure (Terraform) ... GCP (Terraform)
  • [claimed-docs] https://langfuse.com/self-hostingKubernetes (Helm) | AWS (Terraform) | Azure (Terraform) | GCP (Terraform)
  • [claimed-docs] https://langfuse.com/self-hostingYou can run Langfuse on a VM or locally using Docker Compose. This is recommended for testing and low-scale deployments
  • [claimed-docs] https://langfuse.com/self-hostingFor production-scale deployments, we recommend one of the following options: Kubernetes (Helm), AWS (Terraform), Azure (Terraform), GCP (Terraform)
  • [community] https://news.ycombinator.com/item?id=42441258Been using Langfuse OSS for almost 15 months from the start. By far the best solution. No dark patterns found in other projects such as Portkey.
  • [probe] https://langfuse.com/llms.txtPROBE llms.txt: HTTP 200 at https://langfuse.com/llms.txt # Langfuse > Langfuse is an **open-source AI engineering platform** ([GitHub](https://github.com/langfuse/langfuse)) th

Openness = 61.0 ÷ 100 × 100 = 61.0

Built-in AI18.7/100×0.15 of the PA blend

Inside-out: how agentic the product itself is for its users — built-in assistants, autonomous features.

Get AI-generated insights and suggestions from my data inside the productweight 2

2 (weight) × 4 (quality) × 0.6 (partial) = 4.8 of 20 max

  • [claimed-docs] https://langfuse.com/docs/evaluation/evaluation-methods/llm-as-a-judgeLLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewAutomatically score live production traces LLM-as-a-Judge, Scores via API/SDK
  • [claimed-docs] https://langfuse.com/docs/evaluation/evaluation-methods/llm-as-a-judgeIn Langfuse, that score can be numeric, categorical, or boolean.
  • [claimed-docs] https://langfuse.com/docsGet notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts

Set up automations that run autonomously in the backgroundweight 2

2 (weight) × 4 (quality) × 0.6 (partial) = 4.8 of 20 max

  • [claimed-docs] https://langfuse.com/docsGet notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewAutomatically score live production traces LLM-as-a-Judge, Scores via API/SDK
  • [claimed-docs] https://langfuse.com/docs/observability/features/token-and-cost-trackingSet up alerts: get notified automatically when spend crosses a threshold.
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewExport large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewExport large volumes on a schedule | Blob Storage Export

Delegate tasks to a built-in AI assistant inside the productweight 3

3 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 30 max

  • [claimed-docs] https://langfuse.com/docs/docs-mcpCore use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands | MCP Server
  • [probe] https://langfuse.com/docs/docs-mcpofficial MCP server documented at https://langfuse.com/docs/docs-mcp

Operate the product with natural-language commandsweight 2

2 (weight) × 6 (quality) × 0.6 (partial) = 7.2 of 20 max

  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands | MCP Server
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewConnect an AI tool that cannot run shell commands MCP Server
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewWork with Langfuse from a terminal or coding agent | CLI
  • [claimed-docs] https://langfuse.com/docs/docs-mcpCore use case: Use Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [claimed-docs] https://langfuse.com/docs/docs-mcpUse Cursor (or other AI Coding Agent) to automatically integrate Langfuse Tracing into your codebase
  • [probe] https://langfuse.com/docs/docs-mcpofficial MCP server documented at https://langfuse.com/docs/docs-mcp

Built-in AI = 16.8 ÷ 90 × 100 = 18.7

Automation17.3/100×0.15 of the PA blend

Depth of automation primitives — rules, scheduling, bulk operations, webhooks.

Perform bulk operations across many items at onceweight 2

2 (weight) × 5 (quality) × 0.3 (disputed) = 3.0 of 20 max

  • [claimed-docs] https://langfuse.com/docsRun [Experiments](/docs/evaluation/features/prompt-experiments) against datasets to test new prompt versions directly within Langfuse
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewExport large volumes on a schedule | [Blob Storage Export](/docs/api-and-data-platform/features/export-to-blob-storage)
  • [claimed-docs] https://langfuse.com/docs/api-and-data-platform/overviewQuery aggregate cost, usage, latency, volume, or score metrics | [Metrics API v2]
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewReview and rate traces manually | Annotation Queues, Scores via UI
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewBuild a reusable set of test cases | Datasets
  • [community] https://news.ycombinator.com/item?id=42441258The trace/generation retrieval API was brutally slow for bulk scanning operations, so I bypassed it and just queried the db directly. But that is the beauty of open source/self hosted code.

Define rules that trigger actions automatically on eventsweight 3

3 (weight) × 4 (quality) × 0.6 (partial) = 7.2 of 30 max

  • [claimed-docs] https://langfuse.com/docsGet notified over Slack, GitHub Actions, or Webhooks when a metric crosses a threshold with alerts
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewAutomatically score live production traces LLM-as-a-Judge, Scores via API/SDK
  • [claimed-docs] https://langfuse.com/docs/observability/features/token-and-cost-trackingSet up alerts: get notified automatically when spend crosses a threshold.
  • [claimed-docs] https://langfuse.com/docs/evaluation/evaluation-methods/llm-as-a-judgeLLM-as-a-Judge is an evaluation methodology where an LLM is used to assess the quality of outputs produced by another LLM application.

Schedule recurring jobs or workflowsweight 2

2 (weight) × 0 (quality) × 0.0 (none) = 0.0 of 20 max

no evidence cited — the verdict rests on absence of evidence, re-checked on refresh

Version, review, and roll back my automationsweight 1

1 (weight) × 6 (quality) × 0.6 (partial) = 3.6 of 10 max

  • [claimed-docs] https://langfuse.com/docs/prompt-management/overviewUse version control and labels to manage deployments across environments
  • [claimed-docs] https://langfuse.com/docs/prompt-management/overviewLink prompts to traces to analyze performance by prompt version
  • [claimed-docs] https://langfuse.com/docs/prompt-management/overviewWhen prompts live in Langfuse, non-technical team members update them directly in the UI while your application automatically fetches the latest version.
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewReview and rate traces manually | Annotation Queues, Scores via UI
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewCompare prompt, model, or code changes side by side | [Experiments via UI](/docs/evaluation/experiments/experiments-via-ui), [Experiments via SDK]
  • [claimed-docs] https://langfuse.com/docs/evaluation/overviewAutomatically score live production traces LLM-as-a-Judge, Scores via API/SDK

Automation = 13.8 ÷ 80 × 100 = 17.3