Rank #5 of 5 in Feature Flags & Experimentation
Showcase


Try itExperimental
See what an agent can do with Statsig before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$npx -y @statsig/siggy --versionrecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self hostevidence →
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentationevidence →
Stories about experimentation in this arena
Flag management — stories about flag management in this arenaFlag managementevidence →
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance auditevidence →
Stories about governance audit in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plansevidence →
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releasesevidence →
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk deliveryevidence →
Stories about sdk delivery in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 1 free · 0 paid · 0 enterprise · 31 not stated in evidence
Follow the green: where the map greys out is where Statsig stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Webhooks · Machine-readable spec · Versioning policy · API sandbox · Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users · Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
Subscribe to events via webhooks
—–
Build against official SDKs
✓7/10
Issue scoped/least-privilege API credentials for an agent
~3/10
Connect an agent via an official MCP server
✓9/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—0/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—–
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
~5/10
Set up automations that run autonomously in the background
~6/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self host
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentation
Stories about experimentation in this arena
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
✓8/10
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
~6/10
Run A/B and multivariate experiments on flags and see which variation wins on my metrics
✓9/10
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
✓7/10
Flag management — stories about flag management in this arenaFlag management
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance audit
Stories about governance audit in this arena
Restrict who can change which flags with roles, permissions, and scoped API tokens
~4/10
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
~4/10
Require approvals or change requests before production flag changes go live
—–
Every flag change is recorded in an audit log — who changed what, when, and to which value
—–
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releases
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk delivery
Stories about sdk delivery in this arena
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
—0/10
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
~5/10
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
~5/10
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
—–
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
~7/10
Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed⚿ | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ⚿ | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed⚿ | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed⚿ | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 3/10 | Tprobed⚿ | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | 0/10 | ||
Run A/B and multivariate experiments on flags and see which variation wins on my metrics C Experiments | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 3 | full | 9/10 | Xcommunity | |
Create a feature flag and toggle it live in production within minutes of signing up C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 3 | fullfree | 8/10 | Xcommunity | |
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying C Rollouts | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 3 | full | 8/10 | Cclaimed | |
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check C Evaluation | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 3 | partial | 5/10 | Cclaimed | |
Target flags with attribute-based rules and reusable segments so the right users see the right variation C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 3 | partial | 5/10 | Cclaimed | |
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed C Agent ops | ai agent | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | partial | 4/10 | Tprobed⚿ | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 4/10 | Cclaimed | |
Every flag change is recorded in an audit log — who changed what, when, and to which value G Audit | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | n/a | untested | none yet | |
Require approvals or change requests before production flag changes go live C Approvals | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure C Self host | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 3 | none | untested | none yet | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 8/10 | Tprobed | |
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store C Metrics | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | full | 7/10 | Cclaimed | |
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills C Streaming | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | partial | 7/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Tprobed | |
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 2 | partial | 6/10 | Cclaimed | |
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results C Analysis | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | partial | 6/10 | Xcommunity | |
Guard a rollout with metrics so a regression is detected and the release is rolled back automatically C Rollouts | platform engineer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 2 | partial | 5/10 | Cclaimed | |
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior C Sdks | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | partial | 5/10 | Cclaimed | |
Restrict who can change which flags with roles, permissions, and scoped API tokens C Access | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 2 | partial | 4/10 | Cclaimed | |
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing G Pricing | product manager | Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans | 2 | partial | 3/10 | Cclaimed | |
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users C Edge | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Find stale flags and code references so temporary flags actually get removed from the codebase C Lifecycle | platform engineer | Flag management — stories about flag management in this arenaFlag management | 2 | none | untested | none yet | |
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys C Environments | developer | Flag management — stories about flag management in this arenaFlag management | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions C Agent ops | ai agent | Experimentation — stories about experimentation in this arenaExperimentation | 1 | full | 8/10 | Tprobed⚿ | |
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts) C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 1 | full | 8/10 | Cclaimed | |
Schedule flag changes and releases to happen at a specific future time C Scheduling | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 1 | partial | 6/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 4/10 | Cclaimed | |
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network C Proxy | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 1 | none | untested | none yet | |
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API C Standards | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 39 stories with headroom
What would move Statsig’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Evidence covers Statsig's MCP server, which lets external AI tools (Cursor, Claude Code, Codex) query Statsig data — the opposite direction of an in-product assistant that users delegate tasks to.
Governance audit — stories about governance audit in this arenaRequire approvals or change requests before production flag changes go live
nonemoves PA Scoreimpact 30
No evidence describes an approval workflow, change request, or review gate before flag changes go live in production; docs mention scheduled rollouts, overrides, and audit-adjacent features like exposure logging but not a governance approval mechanism.
Governance audit — stories about governance audit in this arenaEvery flag change is recorded in an audit log — who changed what, when, and to which value
nonemoves PA Scoreimpact 30
The evidence pack covers feature gates, experiments, SSO/SCIM, CLI, Console API, and MCP integrations, but nowhere mentions an audit log or change history capturing who changed a flag, when, or to what value.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
No evidence anywhere in the pack mentions self-hosting or an on-prem/open-source deployment of the core Statsig platform; all documentation describes a hosted SaaS console, hosted MCP/API endpoints, and cloud-based warehouse-native analysis.
Deployment self host — stories about deployment self host in this arenaSelf-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure
nonemoves PA Scoreimpact 30
No evidence in the pack mentions an open-source self-hosted distribution of the full Statsig platform; all material describes the hosted SaaS console, SDKs, CLI, and MCP integrations that connect to Statsig's cloud API.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
The evidence pack contains no mention of webhooks or event subscription mechanisms anywhere in Statsig's docs, community posts, or probes — only feature gates, experiments, CLI, MCP, and access management are covered.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
Statsig documents an HTTP API and Console API but there is no evidence of an interactive API reference with runnable/try-it examples; explicit probes for OpenAPI/Swagger endpoints returned 404s, indicating no such interactive explorer exists.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
Statsig documents a Console API and HTTP API but a direct probe for OpenAPI/swagger specs at common paths (openapi.json, swagger.json, api/openapi.json, .well-known/openapi.json) all returned 404, and no docs reference a downloadable machine-readable spec.
Showing the top 8 of 39 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 32 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
API reference32 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)
- Target flags with attribute-based rules and reusable segments so the right users see the right variation
- Restrict who can change which flags with roles, permissions, and scoped API tokens
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing
- Guard a rollout with metrics so a regression is detected and the release is rolled back automatically
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Schedule flag changes and releases to happen at a specific future time
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
- Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Statsigcli docs8 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Set up automations that run autonomously in the background
- Perform bulk operations across many items at once
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Do everything through the API that I can do in the UI
Integrations docs7 stories
- Point an agent at llms.txt or agent-oriented docs
- Connect an agent via an official MCP server
- Issue scoped/least-privilege API credentials for an agent
- Get AI-generated insights and suggestions from my data inside the product
- Operate the product with natural-language commands
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
Hacker News3 stories
OpenAPI spec2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$npx -y @statsig/siggy --versionreproduced$ npx -y @statsig/siggy --version \|/-0.1.0 \
$curl -s -X POST https://docs.statsig.com/api/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -s -X POST https://docs.statsig.com/api/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
event: message
data: {"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"statsig-docs","version":"1.1.0"}},"jsonrpc":"2.0","id":1}
$curl -si -X POST https://api.statsig.com/v1/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -si -X POST https://api.statsig.com/v1/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401
access-control-allow-origin: *
content-type: application/json; charset=utf-8
www-authenticate: Bearer realm="statsig", resource="https://api.statsig.com/v1/mcp", resource_metadata="https://api.statsig.com/.well-known/oauth-protected-resource/v1/mcp"
statsig-final-byte-size: 138
content-length: 138
vary: Accept-Encoding
date: Sat, 05 Sep 2026 05:00:59 GMT
referrer-policy: strict-origin-when-cross-origin
referrer-policy: strict-origin-when-cross-origin
x-content-type-options: nosniff
x-content-type-options: nosniff
content-security-policy: frame-ancestors *.statsig.com
set-cookie: GCLB="c091aaf8ec1589e1"; Max-Age=1; Path=/; HttpOnly
via: 1.1 google
alt-svc: h3=":443"; ma=2592000,h3-29=":443"; ma=2592000
{"statusCode":401,"message":"This endpoint only accepts an active CONSOLE [redacted], but an invalid [redacted] was sent. [redacted]: ","error":"Unauthorized"}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
9 of 19 testable claims verified · 0 contradicted → integrity 47/100
24 distinct capability claims found in Statsig’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
9
Verified
10
Unverified
0
Contradicted
13
Undersold
Verified (10)
“Feature gates (flags) let you toggle product behavior live without redeploying code”
Create a feature flag and toggle it live in production within minutes of signing upfullproof ↗
“Experimentation runs randomized controlled A/B or A/B/n trials to measure impact on key metrics”
Run A/B and multivariate experiments on flags and see which variation wins on my metricsfullproof ↗
“Scorecard lets you enter a hypothesis and pick primary/secondary metrics for an experiment”
Run A/B and multivariate experiments on flags and see which variation wins on my metricsfullproof ↗
“A CLI manages feature gates, experiments, and dynamic configs from the command line”
“An HTTP API lets you retrieve gate/experiment/config values and log events without an SDK”
Drive the product through a documented public APIfullproof ↗
“A Console API provides full CRUD access equivalent to everything doable in the web console UI”
Do everything through the API that I can do in the UIfullproof ↗
“An MCP server integrates Statsig data and actions into AI tools like Codex, Cursor, and Claude Code”
“A Docs MCP server lets AI clients read public Statsig documentation directly via an API endpoint”
Point an agent at llms.txt or agent-oriented docsfullproof ↗
“Offers advanced statistical methods including CUPED, Stratified Sampling, and Switchback Tests”
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment resultspartialproof ↗
“The CLI can be integrated into CI/CD pipelines to automate flag and experiment management”
Run the product headlessly / in CI for automationfullproof ↗
Unverified (11)
“An emergency kill switch can immediately disable a code branch in production”
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually killspartialproof ↗
“Feature gates can be configured as scheduled rollouts to gradually deploy a feature over time”
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeployingfullproof ↗
“Feature gates can be configured as scheduled rollouts to gradually deploy a feature over time”
Schedule flag changes and releases to happen at a specific future timepartialproof ↗
“On-Device Eval SDKs cache experiment/gate definitions on the client for faster local evaluation”
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag checkpartialproof ↗
“Supports SSO and SCIM together for enterprise identity management”
Restrict who can change which flags with roles, permissions, and scoped API tokenspartialproof ↗
“Warehouse Native runs experimentation analysis directly inside your own data warehouse”
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric storefullproof ↗
“Free tier includes feature flags plus 2 million metered events”
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committingpartialproof ↗
“Infra Analytics ingests OpenTelemetry metrics/traces and sets alerts to catch regressions alongside product outcomes”
Guard a rollout with metrics so a regression is detected and the release is rolled back automaticallypartialproof ↗
“Tutorial shows targeting a gate to mobile platforms and internal testers, then verifying it live via the JS SDK”
Target flags with attribute-based rules and reusable segments so the right users see the right variationpartialproof ↗
“Feature gate overrides let specific individual users bypass the gate”
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)fullproof ↗
“Supports auto-adding same-domain users and SSO to simplify inviting teams to a project”
Restrict who can change which flags with roles, permissions, and scoped API tokenspartialproof ↗
Undersold (13)
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisionsfullproof ↗
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleanspartialproof ↗
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewedpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behaviorpartialproof ↗
Claims outside our story set (4)
Real capability claims found in Statsig’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Autotune automatically shifts traffic toward the winning variant to maximize a target metric”
source ↗“Evaluation is deterministic — same user and experiment state always yield the same result across platforms”
source ↗“Console lets you view feature gate exposures to see who encountered a live gate”
source ↗“A Layer lets you manage multiple experiments and feature flags together as a group”
source ↗
Business model
Generous free tier (unlimited flags and seats, event-volume caps), then usage-based pricing on analytics events and session replays with custom enterprise contracts.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP 100% · llms.txt 100% (30d, checked every 6h since Sep 8 '26)
