Rank #4 of 5 in Feature Flags & Experimentation
Install
Showcase


Try itExperimental
See what an agent can do with LaunchDarkly before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$ldcli --version # installed via `brew tap launchdarkly/homebrew-tap && brew install ldcli`recorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self hostevidence →
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentationevidence →
Stories about experimentation in this arena
Flag management — stories about flag management in this arenaFlag managementevidence →
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance auditevidence →
Stories about governance audit in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plansevidence →
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releasesevidence →
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk deliveryevidence →
Stories about sdk delivery in this arena
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where LaunchDarkly stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓7/10
unlocks → Machine-readable spec · Versioning policy · Full data export
Subscribe to events via webhooks
✓8/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
~6/10
Connect an agent via an official MCP server
✓9/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
✓7/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
✓7/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
—0/10
Set up automations that run autonomously in the background
~5/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self host
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentation
Stories about experimentation in this arena
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
~5/10
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
~6/10
Run A/B and multivariate experiments on flags and see which variation wins on my metrics
✓8/10
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
—0/10
Flag management — stories about flag management in this arenaFlag management
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance audit
Stories about governance audit in this arena
Restrict who can change which flags with roles, permissions, and scoped API tokens
✓8/10
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
✓7/10
Require approvals or change requests before production flag changes go live
✓8/10
Every flag change is recorded in an audit log — who changed what, when, and to which value
~6/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releases
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk delivery
Stories about sdk delivery in this arena
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
✓8/10
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
~6/10
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
✓8/10
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
✓8/10
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
✓7/10
Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 7/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | full | 7/10 | Cclaimed | |
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying C Rollouts | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 3 | full | 9/10 | Xcommunity | |
Target flags with attribute-based rules and reusable segments so the right users see the right variation C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 3 | full | 9/10 | Xcommunity | |
Require approvals or change requests before production flag changes go live C Approvals | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | full | 8/10 | Cclaimed | |
Run A/B and multivariate experiments on flags and see which variation wins on my metrics C Experiments | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 3 | full | 8/10 | Xcommunity | |
Create a feature flag and toggle it live in production within minutes of signing up C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 3 | full | 7/10 | Xcommunity | |
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed C Agent ops | ai agent | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | full | 7/10 | Tprobed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | full | 7/10 | Cclaimed | |
Every flag change is recorded in an audit log — who changed what, when, and to which value G Audit | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | partial | 6/10 | Xcommunity | |
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check C Evaluation | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 3 | partial | 6/10 | Xcommunity | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure C Self host | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | n/a | untested | none yet | |
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys C Environments | developer | Flag management — stories about flag management in this arenaFlag management | 2 | full | 9/10 | Xcommunity | |
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users C Edge | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | full | 8/10 | Xcommunity | |
Restrict who can change which flags with roles, permissions, and scoped API tokens C Access | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 2 | full | 8/10 | Cclaimed | |
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior C Sdks | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | full | 8/10 | Xcommunity | |
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills C Streaming | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | full | 7/10 | Xcommunity | |
Find stale flags and code references so temporary flags actually get removed from the codebase C Lifecycle | platform engineer | Flag management — stories about flag management in this arenaFlag management | 2 | partial | 6/10 | Xcommunity | |
Guard a rollout with metrics so a regression is detected and the release is rolled back automatically C Rollouts | platform engineer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 2 | partial | 6/10 | Cclaimed | |
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing G Pricing | product manager | Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans | 2 | partial | 6/10 | Xcommunity | |
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 2 | partial | 6/10 | Xcommunity | |
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results C Analysis | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | partial | 6/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Tprobed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 5/10 | Tprobed | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store C Metrics | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Schedule flag changes and releases to happen at a specific future time C Scheduling | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 1 | full | 9/10 | Cclaimed | |
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network C Proxy | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 1 | full | 8/10 | Xcommunity | |
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts) C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 1 | full | 8/10 | Xcommunity | |
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API C Standards | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 1 | full | 8/10 | Cclaimed | |
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions C Agent ops | ai agent | Experimentation — stories about experimentation in this arenaExperimentation | 1 | partial | 5/10 | Tprobed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Cclaimed |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 28 stories with headroom
What would move LaunchDarkly’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
LaunchDarkly's evidence shows AI-related feature-flag tooling (AI Runs, LLM evaluation flags) and an MCP server that lets external AI agents call into LaunchDarkly as a tool provider, but there is no evidence of a built-in AI assistant inside the product itself that a user can delegate tasks to.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
LaunchDarkly offers API access tokens and a CLI (ldcli) that could programmatically read flag/segment configurations, but no evidence pack item documents a dedicated bulk data-export feature, an open-format export tool, or a stated data-portability/exit policy for account data.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
LaunchDarkly is a hosted SaaS platform; the only self-hostable component is the Relay Proxy, a caching/streaming layer that still depends on the LaunchDarkly SaaS control plane for flag configuration, not the core product itself.
Deployment self host — stories about deployment self host in this arenaSelf-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure
nonemoves PA Scoreimpact 30
LaunchDarkly's evidence only describes an open-source Relay Proxy that caches/streams flag data from the LaunchDarkly cloud for resiliency and low latency, not a full self-hostable flag-management platform; the control plane, rules engine, and dashboard remain SaaS-hosted, and community comments explicitly flag this as a third-party critical-path dependency rather than a self-hosted deployment (launchdarkly-docs-13, launchdarkly-comm-2, launchdarkly-comm-6).
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs
nonemoves agent-readyimpact 30
Direct probes show llms.txt and docs.md both return 404, and no evidence of any agent-oriented documentation endpoint; while LaunchDarkly ships an MCP server and CLI, these do not satisfy the specific 'llms.txt or agent-oriented docs' story.
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product
nonemoves Built-in AIimpact 30
The evidence shows LaunchDarkly's AI-related features (AI Configs, 'AI Runs', LLM-as-judge evals) are about managing and evaluating AI/LLM application behavior via flags, not about the product itself surfacing AI-generated insights or suggestions from a user's flag/experiment/metrics data.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
The evidence pack shows only static docs pages and confirms no OpenAPI/swagger spec discoverable (probe-3 all 404s) and no llms.txt/docs.md; nothing indicates an interactive, runnable API reference/playground exists.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
The evidence pack shows explicit probes for an OpenAPI/swagger spec at LaunchDarkly's standard candidate paths all returning 404, and no other citation surfaces a downloadable machine-readable API spec (only human-readable API access token docs are referenced).
Showing the top 8 of 28 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map4 surfaces · 37 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs36 stories
- Run the product headlessly / in CI for automation
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Subscribe to events via webhooks
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Find stale flags and code references so temporary flags actually get removed from the codebase
- Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)
- Target flags with attribute-based rules and reusable segments so the right users see the right variation
- Restrict who can change which flags with roles, permissions, and scoped API tokens
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Require approvals or change requests before production flag changes go live
- Every flag change is recorded in an audit log — who changed what, when, and to which value
- Do everything through the API that I can do in the UI
- Guard a rollout with metrics so a regression is detected and the release is rolled back automatically
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Schedule flag changes and releases to happen at a specific future time
- Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
- Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Hacker News15 stories
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Find stale flags and code references so temporary flags actually get removed from the codebase
- Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)
- Target flags with attribute-based rules and reusable segments so the right users see the right variation
- Every flag change is recorded in an audit log — who changed what, when, and to which value
- See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Pricing docs7 stories
- Run the product headlessly / in CI for automation
- Build against official SDKs
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Create a feature flag and toggle it live in production within minutes of signing up
- See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
- Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$ldcli --version # installed via `brew tap launchdarkly/homebrew-tap && brew install ldcli`reproduced$ ldcli --version # installed via `brew tap launchdarkly/homebrew-tap && brew install ldcli` ldcli version 3.11.0
$curl -si -X POST https://mcp.launchdarkly.com/mcp/launchdarkly -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -si -X POST https://mcp.launchdarkly.com/mcp/launchdarkly -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401
date: Sat, 05 Sep 2026 05:00:59 GMT
content-type: application/json
content-length: 81
access-control-allow-credentials: true
access-control-allow-headers: Accept, Content-Type, Content-Length, Accept-Encoding, Authorization, User-Agent, Gram-Session, Gram-Project, Gram-[redacted], Gram-[redacted], idempotency-[redacted], Gram-Admin-Override, Gram-Chat-ID, Gram-Assistant-ID, Gram-Chat-Session, MCP-Protocol-Version, Mcp-Method, Mcp-Name, Mcp-Session-Id, X-Gram-Scope-Override, X-Gram-Source
access-control-allow-methods: POST, GET, OPTIONS, PUT, DELETE
access-control-allow-origin: https://app.getgram.ai
access-control-expose-headers: Accept, Content-Type, Content-Length, Accept-Encoding, x-trace-id, Gram-Session, Gram-Chat-ID, Gram-Chat-Session, Mcp-Session-Id
mcp-session-id: e95fa55c-7f4e-422f-9a5d-3efe7d34cffc
www-authenticate: Bearer resource_metadata="https://mcp.launchdarkly.com/.well-known/oauth-protected-resource/mcp/launchdarkly"
x-request-id: f8e7c64866e956b96616ddf413214425
x-trace-id: 0cee9fca19143252ada73ae6e3c75694
strict-transport-security: max-age=31536000; includeSubDomains
content-security-policy: base-uri 'none'; connect-src 'self' https://api.github.com https://metrics.speakeasy.com https://browser-intake-datadoghq.com https://*.usepylon.com wss://*.pusher.com https://*.pusher.com https://*.posthog.com https://chat.speakeasy.com https://chat.dev.speakeasy.com https://app.getgram.ai https://dev.getgram.ai https://cdn.prod.getgram.ai https://cdn.dev.getgram.ai https://app.cal.com https://storage.googleapis.com https://*.clairedefermat.com; default-src 'self'; font-src 'self' https://*.getgram.ai https://fonts.googleapis.com https://fonts.gstatic.com https://*.usepylon.com https://*.posthog.com; frame-ancestors 'self'; frame-src 'self' https://polar.sh https://*.polar.sh https://app.svix.com https://app.cal.com https://cal.com; img-src 'self' blob: https: android-webview-video-poster: data:; media-src 'self' https://*.posthog.com; object-src 'none'; script-src 'self' 'wasm-unsafe-eval' https://www.datadoghq-browser-agent.com https://metrics.speakeasy.com https://widget.usepylon.com https://*.posthog.com https://*.getgram.ai https://app.cal.com https://*.clairedefermat.com; style-src 'self' 'unsafe-inline' https://*.getgram.ai https://fonts.googleapis.com https://*.usepylon.com https://*.posthog.com; worker-src 'self' blob:; report-to browser-intake-datadoghq
permissions-policy: fullscreen=(), compute-pressure=(), camera=(), microphone=(self), geolocation=(), accelerometer=(), bluetooth=(), gyroscope=(), payment=(), usb=(), midi=(), magnetometer=()
referrer-policy: strict-origin-when-cross-origin
reporting-endpoints: browser-intake-datadoghq="https://browser-intake-datadoghq.com/api/v2/logs?dd-[redacted]
x-content-type-options: nosniff
x-frame-options: deny
x-xss-protection: 1; mode=block
{"error":{"code":-32001,"message":"unauthorized access"},"id":1,"jsonrpc":"2.0"}
$echo '<jsonrpc initialize>' | npx -y @launchdarkly/mcp-server start --transport stdioreproduced$ echo '<jsonrpc initialize>' | npx -y @launchdarkly/mcp-server start --transport stdio
\|/{"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"LaunchDarkly","version":"0.6.2"}},"jsonrpc":"2.0","id":1}
\
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
15 of 23 testable claims verified · 1 contradicted → integrity 57/100
33 distinct capability claims found in LaunchDarkly’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
15
Verified
7
Unverified
1
Contradicted
15
Undersold
Verified (21)
“Define attribute-based targeting rules and custom rules to control which users see which flag variation”
Target flags with attribute-based rules and reusable segments so the right users see the right variationfullproof ↗
“Target or exclude specific individual users (allowlist/beta testers) for a flag”
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)fullproof ↗
“Create progressive percentage-based rollouts that ramp exposure over time with consistent bucketing”
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeployingfullproof ↗
“Every flag change is recorded in a change history / audit log”
Every flag change is recorded in an audit log — who changed what, when, and to which valuepartialproof ↗
“Choose from different metric types when defining experiment success criteria”
Run A/B and multivariate experiments on flags and see which variation wins on my metricsfullproof ↗
“Manage separate environments (e.g. dev/staging/prod) with independent flag state”
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keysfullproof ↗
“Provides an official (self-hosted or hosted) MCP server so agents can connect and use LaunchDarkly's tools”
“Official CLI for managing LaunchDarkly resources from the command line”
“Relay Proxy is a small Go app you run on your own infrastructure to serve flags at the edge/within your network”
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my networkfullproof ↗
“Provides a Cloudflare Workers SDK for evaluating flags at the edge”
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to usersfullproof ↗
“Plans include unlimited feature flags regardless of tier”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Published usage limits include a monthly quota of session replays and error events”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Create a new feature flag directly in the product”
Create a feature flag and toggle it live in production within minutes of signing upfullproof ↗
“Contexts define the attributes (users, devices, orgs, etc.) used to evaluate and target flags”
Target flags with attribute-based rules and reusable segments so the right users see the right variationfullproof ↗
“Create reusable segments of users/contexts to target consistently across multiple flags”
Target flags with attribute-based rules and reusable segments so the right users see the right variationfullproof ↗
“Scans code for flag references so stale/temporary flags can be identified and cleaned up”
Find stale flags and code references so temporary flags actually get removed from the codebasepartialproof ↗
“Offers roughly 30 idiomatic SDKs spanning backend, web, and mobile platforms”
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behaviorfullproof ↗
“Convert existing targeting rules directly into reusable segments”
Target flags with attribute-based rules and reusable segments so the right users see the right variationfullproof ↗
“Run A/B tests and multivariate experiments on flags to see which variation performs best”
Run A/B and multivariate experiments on flags and see which variation wins on my metricsfullproof ↗
“No per-seat pricing — unlimited seats, invite the whole team, pricing based on usage”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Issue API access tokens for scoped, programmatic access to the platform”
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Unverified (8)
“Guard a rollout with metrics so regressions can be detected and rolled back automatically”
Guard a rollout with metrics so a regression is detected and the release is rolled back automaticallypartialproof ↗
“Require approval workflows before production flag changes go live”
Require approvals or change requests before production flag changes go livefullproof ↗
“Supports adaptive triggers that respond to conditions during a rollout”
Guard a rollout with metrics so a regression is detected and the release is rolled back automaticallypartialproof ↗
“Ships OpenFeature-compatible providers so flag code isn't locked to LaunchDarkly's SDK API”
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK APIfullproof ↗
“Experiment results are computed using a documented Bayesian statistics engine”
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment resultspartialproof ↗
“Schedule flag changes to take effect at a specific future time”
Schedule flag changes and releases to happen at a specific future timefullproof ↗
“Assign roles to team members to control who can change what”
Restrict who can change which flags with roles, permissions, and scoped API tokensfullproof ↗
“Subscribe to platform events via webhooks”
Contradicted (1)
“Offers a version/deployment for federal environments (data residency/compliance)”
Choose where my data is stored (region/residency)noneproof ↗
Undersold (15)
Run the product headlessly / in CI for automationfullproof ↗
Drive the product through a documented public APIfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Test against a sandbox environment without touching production datafullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventsfullproof ↗
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisionspartialproof ↗
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleanspartialproof ↗
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewedfullproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag checkpartialproof ↗
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually killsfullproof ↗
Claims outside our story set (3)
Real capability claims found in LaunchDarkly’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Preview/test the impact of targeting rule changes before applying them”
source ↗“Configure, evaluate, roll out, and observe AI models end-to-end within the platform”
source ↗“Generates an 'AI Run' record when tracking an LLM request or running an online LLM-as-a-judge eval”
source ↗
Business model
Developer free tier, then per-seat Foundation plans with usage-based add-ons (service connections, experimentation keys, data export) and custom Enterprise/Guardian tiers.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP 100% (30d, checked every 6h since Sep 8 '26)
