Rank #2 of 5 in Feature Flags & Experimentation
Showcase


Try itExperimental
See what an agent can do with GrowthBook before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$docker run -d mongo:6 && docker run -d -p 13100:3100 growthbook/growthbook && POST /auth/register && POST /organization && POST /feature {id: pa-probe-flag} && POST /sdk-connections && curl /api/features/<sdk-key>recorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self hostevidence →
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentationevidence →
Stories about experimentation in this arena
Flag management — stories about flag management in this arenaFlag managementevidence →
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance auditevidence →
Stories about governance audit in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plansevidence →
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releasesevidence →
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk deliveryevidence →
Stories about sdk delivery in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 11 free · 10 paid · 0 enterprise · 22 not stated in evidence
Follow the green: where the map greys out is where GrowthBook stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Webhooks · Machine-readable spec · Versioning policy · Official CLI · Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
Subscribe to events via webhooks
—–
Build against official SDKs
✓9/10
Issue scoped/least-privilege API credentials for an agent
~5/10
Connect an agent via an official MCP server
✓9/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
✓8/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
✓8/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
~4/10
Set up automations that run autonomously in the background
~4/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self host
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentation
Stories about experimentation in this arena
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
✓8/10
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
✓9/10
Run A/B and multivariate experiments on flags and see which variation wins on my metrics
✓8/10
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
✓8/10
Flag management — stories about flag management in this arenaFlag management
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance audit
Stories about governance audit in this arena
Restrict who can change which flags with roles, permissions, and scoped API tokens
~5/10
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
~6/10
Require approvals or change requests before production flag changes go live
✓8/10
Every flag change is recorded in an audit log — who changed what, when, and to which value
~5/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releases
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk delivery
Stories about sdk delivery in this arena
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
✓8/10
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
✓9/10
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
✓7/10
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
—0/10
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
~4/10
Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | fullfree | 9/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Xcommunity | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 8/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialpaid | 4/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | fullfree | 8/10 | Tprobed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullfree | 10/10 | Tprobed | |
Create a feature flag and toggle it live in production within minutes of signing up C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 3 | fullfree | 9/10 | Tprobed | |
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check C Evaluation | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 3 | full | 9/10 | Tprobed | |
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure C Self host | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 3 | fullfree | 9/10 | Tprobed | |
Require approvals or change requests before production flag changes go live C Approvals | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | fullpaid | 8/10 | Cclaimed | |
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying C Rollouts | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 3 | fullpaid | 8/10 | Tprobed | |
Run A/B and multivariate experiments on flags and see which variation wins on my metrics C Experiments | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 3 | fullfree | 8/10 | Xcommunity | |
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed C Agent ops | ai agent | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | partialpaid | 6/10 | Tprobed | |
Target flags with attribute-based rules and reusable segments so the right users see the right variation C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 3 | partial | 6/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 5/10 | Cclaimed | |
Every flag change is recorded in an audit log — who changed what, when, and to which value G Audit | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | partialpaid | 5/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partialfree | 5/10 | Xcommunity | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results C Analysis | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | full | 9/10 | Xcommunity | |
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store C Metrics | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | full | 8/10 | Xcommunity | |
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users C Edge | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | full | 8/10 | Cclaimed | |
Find stale flags and code references so temporary flags actually get removed from the codebase C Lifecycle | platform engineer | Flag management — stories about flag management in this arenaFlag management | 2 | full | 8/10 | Cclaimed | |
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys C Environments | developer | Flag management — stories about flag management in this arenaFlag management | 2 | fullpaid | 8/10 | Tprobed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 7/10 | Tprobed | |
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 2 | full | 7/10 | Tprobed | |
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior C Sdks | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | full | 7/10 | Xcommunity | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialfree | 6/10 | Tprobed | |
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing G Pricing | product manager | Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 5/10 | Tprobed | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partialfree | 5/10 | Xcommunity | |
Restrict who can change which flags with roles, permissions, and scoped API tokens C Access | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 2 | partialpaid | 5/10 | Cclaimed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialfree | 4/10 | Cclaimed | |
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills C Streaming | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | partial | 4/10 | Xcommunity | |
Guard a rollout with metrics so a regression is detected and the release is rolled back automatically C Rollouts | platform engineer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 2 | partialpaid | 4/10 | Cclaimed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network C Proxy | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 1 | full | 9/10 | Xcommunity | |
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions C Agent ops | ai agent | Experimentation — stories about experimentation in this arenaExperimentation | 1 | full | 8/10 | Tprobed | |
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts) C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 1 | full | 7/10 | Cclaimed | |
Schedule flag changes and releases to happen at a specific future time C Scheduling | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 1 | partialpaid | 6/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partialpaid | 6/10 | Xcommunity | |
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API C Standards | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 1 | none | 0/10 |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 28 stories with headroom
What would move GrowthBook’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
GrowthBook documents an MCP server that lets external AI tools (Cursor, VS Code, Claude) connect to and query GrowthBook — this is the opposite direction (GrowthBook as a data/tool source for external assistants), not a built-in AI assistant living inside the GrowthBook product itself.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence addresses whether GrowthBook uses customer data to train AI models or offers an opt-out; while self-hosting (growthbook-docs-10) keeps data on-prem, there is no explicit privacy statement about AI training data usage.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
The evidence pack documents SDKs, REST API, self-host Docker, and an official MCP server, but there is no mention anywhere of an official standalone CLI tool for GrowthBook.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
The evidence pack covers SDKs, REST API, MCP integration, self-hosting, and experimentation features but contains no mention of webhooks or event subscription mechanisms anywhere, so there is no evidence GrowthBook supports this capability.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
GrowthBook documents a REST API (growthbook-docs-12) but there is no evidence of an interactive, runnable API reference (e.g., Swagger/OpenAPI try-it console) — a direct probe for openapi/swagger specs returned 404 on all candidate paths, and no docs mention runnable code samples in an API explorer.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
While GrowthBook documents a full REST API (growthbook-docs-12) and an SDK Connection Endpoint (growthbook-docs-20), the evidence pack shows an explicit probe for OpenAPI/Swagger spec files at expected paths that all returned 404, and no other evidence surfaces a downloadable OpenAPI or machine-readable spec.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Missing: any mention of API version numbers, backward-compatibility guarantees, or a deprecation/sunset policy.
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows
nonemoves PA Scoreimpact 20
Evidence only shows one-time scheduling of feature-flag rule start/stop dates and 'ramp schedules' for gradual rollouts (growthbook-docs-22, growthbook-docs-37) — nothing resembling recurring cron-like jobs or workflow automation that an AI-native user could schedule to run repeatedly.
Showing the top 8 of 28 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map11 surfaces · 43 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
App docs20 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Get AI-generated insights and suggestions from my data inside the product
- Perform bulk operations across many items at once
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Restrict who can change which flags with roles, permissions, and scoped API tokens
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Guard a rollout with metrics so a regression is detected and the release is rolled back automatically
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Self host docs19 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Test against a sandbox environment without touching production data
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- Choose where my data is stored (region/residency)
- Control data retention and deletion
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Features docs19 stories
- Set up automations that run autonomously in the background
- Test against a sandbox environment without touching production data
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Find stale flags and code references so temporary flags actually get removed from the codebase
- Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)
- Target flags with attribute-based rules and reusable segments so the right users see the right variation
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Require approvals or change requests before production flag changes go live
- Choose where my data is stored (region/residency)
- Control data retention and deletion
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Schedule flag changes and releases to happen at a specific future time
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Hacker News14 stories
- Build against official SDKs
- Version, review, and roll back my automations
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure
- Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
- Create a feature flag and toggle it live in production within minutes of signing up
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Pricing docs13 stories
- Issue scoped/least-privilege API credentials for an agent
- Set up automations that run autonomously in the background
- Version, review, and roll back my automations
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Restrict who can change which flags with roles, permissions, and scoped API tokens
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Require approvals or change requests before production flag changes go live
- Every flag change is recorded in an audit log — who changed what, when, and to which value
- Export all of my data in open formats and leave
- See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing
- Guard a rollout with metrics so a regression is detected and the release is rolled back automatically
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Schedule flag changes and releases to happen at a specific future time
Integrations docs10 stories
- Point an agent at llms.txt or agent-oriented docs
- Connect an agent via an official MCP server
- Issue scoped/least-privilege API credentials for an agent
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Do everything through the API that I can do in the UI
GitHub README7 stories
- Build against official SDKs
- Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Read the product's source under an open license
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
Lib docs5 stories
- Build against official SDKs
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
Experiments docs4 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$docker run -d mongo:6 && docker run -d -p 13100:3100 growthbook/growthbook && POST /auth/register && POST /organization && POST /feature {id: pa-probe-flag} && POST /sdk-connections && curl /api/features/<sdk-key>reproduced$ docker run -d mongo:6 && docker run -d -p 13100:3100 growthbook/growthbook && POST /auth/register && POST /organization && POST /feature {id: pa-probe-flag} && POST /sdk-connections && curl /api/features/<sdk-[redacted]>
{
"status": 200,
"healthy": true
}
org: org_3owr3imtnx0gvv
{
"status": 200,
"feature": {
"defaultValue": "true",
"valueType": "boolean",
"owner": "u_3owr3imtnx0guc
sdk [redacted]: sdk-9bYqWRtJ6RqRGNxz
{
"status": 200,
"features": {
"pa-probe-flag": {
"defaultValue": true
}
},
"dateUpdated": "2026-09-05T05:01:09.550Z"
}
$echo '<jsonrpc initialize>' | npx -y @growthbook/mcpreproduced$ echo '<jsonrpc initialize>' | npx -y @growthbook/mcp
\|/-{"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"GrowthBook MCP Thin","version":"2.1.0","title":"GrowthBook MCP Thin","websiteUrl":"https://growthbook.io"},"instructions":"You are a helpful assistant that interacts with GrowthBook via a thin MCP server.\n\nTools:\n- growthbook_list_skills — discover available GrowthBook skills (name + description)\n- growthbook_read_skill — load a skill's full workflow and guardrails\n- growthbook_api_read — authenticated GET against the GrowthBook REST API\n- growthbook_api_write — authenticated POST/PUT/PATCH/DELETE against the GrowthBook REST API\n\nWorkflow:\n1. growthbook_list_skills (or growthbook_read_skill if you already know the skill name) to load competence.\n2. Follow the skill's steps. Skills show bash like `gb-call <METHOD> <PATH> [body]` —\n map GET to growthbook_api_read and POST/PUT/PATCH/DELETE to growthbook_api_write with the same path and optional JSON body string.\n3. Do not invent endpoints; prefer the paths listed in the skill.\n\nSkill content is the source of truth for GrowthBook task workflows and API footguns.\ngrowthbook_api_read / growthbook_api_write are dumb authenticated passthroughs — they do not validate payloads."},"jsonrpc":"2.0","id":1}
\
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
14 of 20 testable claims verified · 0 contradicted → integrity 70/100
31 distinct capability claims found in GrowthBook’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
14
Verified
6
Unverified
0
Contradicted
23
Undersold
Verified (21)
“Get feature flags and experiments running in your app with a few lines of code using a language-specific SDK”
Create a feature flag and toggle it live in production within minutes of signing upfullproof ↗
“Manage separate dev/staging/production environments with independently enabled/disabled flags”
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keysfullproof ↗
“Choose between a Bayesian or Frequentist statistics engine when viewing experiment results”
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment resultsfullproof ↗
“Define Fact Tables and metrics so results are computed directly from your own data warehouse”
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric storefullproof ↗
“Self-host the full platform (MongoDB, API, front-end) via Docker Compose on your own infrastructure”
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructurefullproof ↗
“Run a GrowthBook Proxy next to your app to cache SDK payloads, cut latency, and keep serving flags if the API is unreachable”
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my networkfullproof ↗
“Offers a full REST API for interacting with the application”
Drive the product through a documented public APIfullproof ↗
“Official MCP server connects AI tools like Cursor, VS Code, and Claude to GrowthBook”
“Ships 24 official SDKs covering languages/platforms including React, Python, Android, and iOS”
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behaviorfullproof ↗
“Stats engine supports CUPED, sequential testing, Bayesian, post-stratification, bandits, and SRM checks”
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment resultsfullproof ↗
“Warehouse-native: can query 11 data sources including BigQuery, Snowflake, and Databricks”
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric storefullproof ↗
“SDK hooks (e.g. useFeatureIsOn, useFeatureValue) serve boolean or typed dynamic config values, not just booleans”
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleansfullproof ↗
“Run server-side experiments via inline SDK calls with no third-party network requests”
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag checkfullproof ↗
“SDK Connection Endpoint provides readonly, scoped access to just the flag data needed for evaluation”
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
“MCP server loads GrowthBook Agent Skills (workflows/guardrails) and makes authenticated calls to the REST API on an agent's behalf”
“Rules overriding a flag's default value can be scoped, ramped, and scheduled for gradual rollout or experimentation”
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeployingfullproof ↗
“"Chance to Win" metric shows the probability a variation is better, highlighting clear winners above 95%”
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment resultsfullproof ↗
“Supports multiple experiment assignment methods: feature flags, URL redirects, the Visual Editor, or custom bucketing systems”
Run A/B and multivariate experiments on flags and see which variation wins on my metricsfullproof ↗
“Experiment rules randomly assign users to variations using a consistent hashing attribute”
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeployingfullproof ↗
“Quickstart via 'docker compose up -d' to run the entire app locally in one command”
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructurefullproof ↗
“On app load, the SDK checks whether the user should be in an experiment and serves the assigned variation client-side”
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag checkfullproof ↗
Unverified (9)
“Add attribute-based targeting conditions to a rule to control which value a user sees”
Target flags with attribute-based rules and reusable segments so the right users see the right variationpartialproof ↗
“Automatically detects unused/stale flags so they can be cleaned up from code”
Find stale flags and code references so temporary flags actually get removed from the codebasefullproof ↗
“Draft feature revisions, request review, resolve merge conflicts, and require approval before publishing changes”
Require approvals or change requests before production flag changes go livefullproof ↗
“Evaluate feature flags and Visual Editor experiments at the edge on Cloudflare Workers to reduce flicker”
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to usersfullproof ↗
“Starter tier adds visual editor, multi-arm bandits, safe rollouts, custom dashboards, advanced permissioning, and a power calculator”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Pro tier adds custom environments, ramp schedules, approval workflows, advanced access control, SSO/SCIM, and exportable audit logs”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Rules overriding a flag's default value can be scoped, ramped, and scheduled for gradual rollout or experimentation”
Schedule flag changes and releases to happen at a specific future timepartialproof ↗
“Plans include unlimited feature flags, unlimited experiments, and unlimited traffic”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Free plan supports up to 3 users, 1 project, and unlimited flags/experiments/traffic”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
Undersold (23)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Test against a sandbox environment without touching production datafullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisionsfullproof ↗
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)fullproof ↗
Restrict who can change which flags with roles, permissions, and scoped API tokenspartialproof ↗
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewedpartialproof ↗
Every flag change is recorded in an audit log — who changed what, when, and to which valuepartialproof ↗
Do everything through the API that I can do in the UIfullproof ↗
Export all of my data in open formats and leavepartialproof ↗
Read the product's source under an open licensepartialproof ↗
Choose where my data is stored (region/residency)partialproof ↗
Guard a rollout with metrics so a regression is detected and the release is rolled back automaticallypartialproof ↗
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually killspartialproof ↗
Claims outside our story set (2)
Real capability claims found in GrowthBook’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Includes a built-in product analytics suite to build and share dashboards”
source ↗“Actual targeting attribute values (user IDs, emails) are never sent to GrowthBook, only kept in memory locally in the SDK”
source ↗
Business model
MIT-licensed core you can self-host free with unlimited traffic; GrowthBook Cloud has a free starter tier, per-seat Pro pricing, and custom enterprise plans.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
