Skip to content

Rank #4 of 8 in LLM Evals & Observability

Helicone logo

Helicone

Open SourceYC W23

Helicone, Inc.

6.2k1.7k/yrnpm 3.6k/wk +25npm/wk +340

Showcase

Helicone homepage screenshot
homepage · captured Sep 2026 · view live ↗
Helicone docs screenshot
docs · captured Sep 2026 · view live ↗

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

39.9/100

Alerting dashboards — stories about alerting dashboards in this arenaAlerting dashboardsevidence →

Stories about alerting dashboards in this arena

30.0/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

29.0/100

Cost monitoring — stories about cost monitoring in this arenaCost monitoringevidence →

Stories about cost monitoring in this arena

27.6/100

Data access export — stories about data access export in this arenaData access exportevidence →

Stories about data access export in this arena

24.0/100

Evals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasetsevidence →

Measuring quality — datasets, eval runs, regression tracking

15.5/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

30.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

14.0/100

Prompt management — stories about prompt management in this arenaPrompt managementevidence →

Stories about prompt management in this arena

62.4/100

Tracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentationevidence →

Instrumenting code and tracing requests end to end

41.8/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 3 free · 0 paid · 0 enterprise · 31 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3noneuntestednone yet

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10X

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10T

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10T

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10C

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial3/10C

Capture traces of my LLM calls with inputs, outputs, latency, and token usage C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3fullfree9/10X

Version prompts and deploy changes to production without shipping code C

Prompt workflow

developerPrompt management — stories about prompt management in this arenaPrompt management3full8/10C

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partialfree6/10X

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial5/10C

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partial5/10X

Have an agent query my traces, metrics, and eval results through an API or MCP server to debug my app C

Ai observability

ai-native userTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3partial5/10T

See cost and token usage per request, model, and time period in dashboards C

Cost tracking

developerCost monitoring — stories about cost monitoring in this arenaCost monitoring3partialfree5/10X

Curate datasets from production traces and run offline evaluations against them C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3partial4/10C

Score outputs with configurable LLM-as-a-judge evaluators C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3partial4/10C

Compare eval runs side by side to catch regressions between prompt or model versions C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets3partial3/10C

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3n/auntestednone yet

Send and receive traces over OpenTelemetry (OTLP) instead of a proprietary format C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation3noneuntestednone yet

Instrument my app through existing integrations for frameworks like LangChain, the OpenAI SDK, or the Vercel AI SDK C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2full8/10X

Trace multi-step agent runs as nested spans grouped into sessions or threads C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2full8/10C

Build custom dashboards over latency, error, cost, and eval-score metrics C

Monitoring

ml engineerAlerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards2partial6/10X

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial6/10T

Instrument apps in both Python and JS/TS with officially supported SDKs G

Sdk coverage

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2partial6/10X

Iterate on prompts in a playground against real models and variables C

Prompt workflow

developerPrompt management — stories about prompt management in this arenaPrompt management2partial6/10C

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2disputed5/10D

Run evaluators continuously on live production traffic, not just offline datasets C

Online evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2partial5/10C

Attribute cost and usage to users, sessions, and features via custom metadata C

Cost tracking

developerCost monitoring — stories about cost monitoring in this arenaCost monitoring2partial4/10C

Bulk-export traces and datasets to blob storage or my data warehouse C

Data export

developerData access export — stories about data access export in this arenaData access export2partial4/10X

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partial4/10C

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial4/10C

Set alerts on error rates, cost spikes, or eval-score drops and get notified in Slack, PagerDuty, or email C

Monitoring

developerAlerting dashboards — stories about alerting dashboards in this arenaAlerting dashboards2partial4/10C

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partial3/10C

Write custom code-based scorers and metrics for my evaluations C

Offline evals

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2partial3/10C

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Have an agent create a dataset, trigger an eval run programmatically, and read back the results C

Ai eval ops

ai-native userEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2noneuntestednone yet

Mask or redact sensitive data before it is stored in traces C

Data controls

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation2noneuntestednone yet

Route outputs to human annotation queues for review and labeling C

Human review

ml engineerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2noneuntestednone yet

Run evals in CI and gate deployments on their results C

Offline evals

developerEvals datasets — measuring quality — datasets, eval runs, regression trackingEvals datasets2noneuntestednone yet

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2n/auntestednone yet

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial6/10C

Capture multimodal payloads (images, audio, files) inside my traces C

Trace capture

developerTracing instrumentation — instrumenting code and tracing requests end to endTracing instrumentation1noneuntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 36 stories with headroom

What would move Helicone’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server

    nonemoves agent-readyimpact 45

    Helicone is an LLM observability/gateway platform, and this axis (offering an official MCP server so agents can connect) plausibly applies since it has an ecosystem of integrations, docs, and APIs.

  2. Tracing instrumentation — instrumenting code and tracing requests end to endSend and receive traces over OpenTelemetry (OTLP) instead of a proprietary format

    nonemoves PA Scoreimpact 30

    No evidence pack items mention OpenTelemetry, OTLP, or any standard tracing protocol support; Helicone's docs describe proprietary logging via SDK integration, sessions, and REST API, not OTLP ingestion/export.

  3. Agenticness — how well agents can access and operate the productUse an official CLI

    nonemoves agent-readyimpact 30

    No evidence of an official Helicone CLI tool; integration is via SDKs, API keys, gateway, and REST/OpenAPI, but no CLI is mentioned anywhere in docs, GitHub, or community sources.

  4. Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent

    nonemoves agent-readyimpact 30

    No evidence Helicone supports issuing scoped or least-privilege API credentials/keys for agents; docs mention bringing your own provider keys or using Helicone's own key, but nothing about granular permission scoping.

  5. Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy

    nonemoves API qualityimpact 30

    An OpenAPI spec exists (helicone-probe-3) confirming a REST API, but there is no evidence of API versioning scheme (e.g., v1/v2 paths with migration guides) or a documented deprecation policy for endpoints/models; nothing in the docs pack addresses lifecycle or backward-compatibility commitments.

  6. Evals datasets — measuring quality — datasets, eval runs, regression trackingCompare eval runs side by side to catch regressions between prompt or model versions

    partialq3/10moves PA Scoreimpact 21

    Missing: a documented eval-run comparison UI, dataset-based batch evaluation runs, and any hands-on/community confirmation of side-by-side regression catching.

  7. Evals datasets — measuring quality — datasets, eval runs, regression trackingHave an agent create a dataset, trigger an eval run programmatically, and read back the results

    nonemoves PA Scoreimpact 20

    Evidence shows observability features (sessions, webhooks, HQL query, REST API for point queries) and a vague mention of 'real-time evaluation' scoring, but there is no documentation of a dataset-creation API, a way to programmatically trigger an eval run, or an API to read back eval results — the core building blocks of this story are absent from the evidence pack.

  8. Evals datasets — measuring quality — datasets, eval runs, regression trackingRun evals in CI and gate deployments on their results

    nonemoves PA Scoreimpact 20

    Helicone offers observability, webhooks, real-time scoring, caching, and prompt versioning, but there is no evidence of a CI-integrated eval runner, test suite, or deployment gating mechanism tied to eval results.

Showing the top 8 of 36 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 35 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

4 of 14 testable claims verified · 2 contradictedintegrity 0/100

26 distinct capability claims found in Helicone’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

4

Verified

8

Unverified

2

Contradicted

22

Undersold

Verified (6)
Unverified (10)
Contradicted (2)
Undersold (22)
Claims outside our story set (9)

Real capability claims found in Helicone’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Switch between 100+ models just by changing the model name

    source ↗
  • Requests sharing the same path are grouped as the same type of work over time

    source ↗
  • Caching stores complete responses on Cloudflare's edge network to cut redundant calls, latency, and cost

    source ↗
  • Create separate cache namespaces for different users or contexts

    source ↗
  • Supports point queries to fetch individual requests

    source ↗
  • Supports a query language (HQL) for querying data

    source ↗
  • Free tier includes 10,000 requests

    source ↗
  • AI Gateway provides access to 100+ models via one API key with intelligent routing and automatic fallbacks

    source ↗
  • Supports fine-tuning through partners OpenPipe or Autonomi

    source ↗
Suggest a story for these →

Business model

open-sourcefree-tiersubscription-per-seatusage-basedenterprise-custom

Apache-licensed and self-hostable free; Helicone Cloud has a free tier (10k requests/month), then per-seat Pro with usage-based request overages and custom enterprise plans.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score27 (Sep 4 '26)34 (Sep 4 '26)
Agent-ready38 (Sep 4 '26)43 (Sep 4 '26)

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime llms.txt 100% · openapi.json 100% (30d, checked every 6h since Sep 8 '26)