Rank #1 of 5 in Feature Flags & Experimentation
Install
Showcase


Try itExperimental
See what an agent can do with Flagsmith before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; the live MCP handshake runs real requests from our edge, right now — including, where the server allows it, one real read-only tool call (bring your own key for auth-gated servers); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$docker run -d postgres:16-alpine && docker run -d -p 18000:8000 flagsmith/flagsmith && POST /api/v1/auth/users/ && POST /organisations/ && POST /projects/ && POST /environments/ && POST /features/ {pa_probe_flag} && curl /api/v1/flags/ -H "X-Environment-Key: <key>"recorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self hostevidence →
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentationevidence →
Stories about experimentation in this arena
Flag management — stories about flag management in this arenaFlag managementevidence →
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance auditevidence →
Stories about governance audit in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plansevidence →
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releasesevidence →
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk deliveryevidence →
Stories about sdk delivery in this arena
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Flagsmith stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Machine-readable spec · Versioning policy · My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check · Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API · Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
Subscribe to events via webhooks
~5/10
Build against official SDKs
~6/10
Issue scoped/least-privilege API credentials for an agent
~5/10
Connect an agent via an official MCP server
✓9/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
~7/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
~5/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
~7/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
—0/10
Set up automations that run autonomously in the background
~7/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment self host — stories about deployment self host in this arenaDeployment self host
Stories about deployment self host in this arena
Experimentation — stories about experimentation in this arenaExperimentation
Stories about experimentation in this arena
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
~5/10
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
~6/10
Run A/B and multivariate experiments on flags and see which variation wins on my metrics
✓7/10
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
—0/10
Flag management — stories about flag management in this arenaFlag management
Stories about flag management in this arena
Governance audit — stories about governance audit in this arenaGovernance audit
Stories about governance audit in this arena
Restrict who can change which flags with roles, permissions, and scoped API tokens
✓9/10
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
✓7/10
Require approvals or change requests before production flag changes go live
✓9/10
Every flag change is recorded in an audit log — who changed what, when, and to which value
✓9/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rollouts releases — stories about rollouts releases in this arenaRollouts releases
Stories about rollouts releases in this arena
Sdk delivery — stories about sdk delivery in this arenaSdk delivery
Stories about sdk delivery in this arena
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
~5/10
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
!3/10
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
~6/10
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API
—–
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
—0/10
Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 7/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 7/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 7/10 | Tprobed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | full | 10/10 | Tprobed | |
Create a feature flag and toggle it live in production within minutes of signing up C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 3 | full | 9/10 | Tprobed | |
Every flag change is recorded in an audit log — who changed what, when, and to which value G Audit | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | full | 9/10 | Cclaimed | |
Require approvals or change requests before production flag changes go live C Approvals | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | full | 9/10 | Cclaimed | |
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying C Rollouts | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 3 | full | 9/10 | Tprobed | |
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure C Self host | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 3 | full | 9/10 | Tprobed | |
Target flags with attribute-based rules and reusable segments so the right users see the right variation C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 3 | full | 9/10 | Tprobed | |
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed C Agent ops | ai agent | Governance audit — stories about governance audit in this arenaGovernance audit | 3 | full | 7/10 | Tprobed | |
Run A/B and multivariate experiments on flags and see which variation wins on my metrics C Experiments | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 3 | full | 7/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 6/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 6/10 | Tprobed | |
My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check C Evaluation | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 3 | disputed | 3/10 | Dcontradicted | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | n/a | untested | none yet | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 9/10 | Tprobed | |
Restrict who can change which flags with roles, permissions, and scoped API tokens C Access | platform engineer | Governance audit — stories about governance audit in this arenaGovernance audit | 2 | full | 9/10 | Cclaimed | |
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans C Flags | developer | Flag management — stories about flag management in this arenaFlag management | 2 | full | 9/10 | Tprobed | |
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys C Environments | developer | Flag management — stories about flag management in this arenaFlag management | 2 | full | 8/10 | Tprobed | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 7/10 | Cclaimed | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 6/10 | Tprobed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Tprobed | |
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results C Analysis | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | partial | 6/10 | Cclaimed | |
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior C Sdks | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | partial | 6/10 | Xcommunity | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 5/10 | Cclaimed | |
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users C Edge | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | partial | 5/10 | Xcommunity | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 4/10 | Cclaimed | |
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store C Metrics | product manager | Experimentation — stories about experimentation in this arenaExperimentation | 2 | none | 0/10 | ||
Flag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills C Streaming | developer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 2 | none | 0/10 | ||
Guard a rollout with metrics so a regression is detected and the release is rolled back automatically C Rollouts | platform engineer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 2 | none | 0/10 | ||
See published pricing and understand what drives cost (seats, MAUs, events, requests) before committing G Pricing | product manager | Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans | 2 | none | 0/10 | ||
Find stale flags and code references so temporary flags actually get removed from the codebase C Lifecycle | platform engineer | Flag management — stories about flag management in this arenaFlag management | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | n/a | untested | none yet | |
Schedule flag changes and releases to happen at a specific future time C Scheduling | developer | Rollouts releases — stories about rollouts releases in this arenaRollouts releases | 1 | full | 9/10 | Cclaimed | |
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts) C Targeting | developer | Flag management — stories about flag management in this arenaFlag management | 1 | full | 9/10 | Cclaimed | |
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network C Proxy | platform engineer | Deployment self host — stories about deployment self host in this arenaDeployment self host | 1 | full | 8/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | full | 7/10 | Cclaimed | |
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions C Agent ops | ai agent | Experimentation — stories about experimentation in this arenaExperimentation | 1 | partial | 5/10 | Tprobed | |
Use the vendor through OpenFeature providers so my flag code isn't locked to one vendor's SDK API C Standards | platform engineer | Sdk delivery — stories about sdk delivery in this arenaSdk delivery | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 28 stories with headroom
What would move Flagsmith’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Flagsmith exposes an MCP server so external AI agents can call its Admin API, and calls its automated release pipeline an 'intelligent release assistant,' but neither is a built-in AI assistant inside the Flagsmith product that a user can converse with or delegate tasks to — the MCP server is a server-side integration point for external agents, not a first-party in-app assistant.
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product
nonemoves Built-in AIimpact 30
Flagsmith provides an MCP server for external AI agents to call its Admin API (docs-17/29) and a 'release pipeline' described as an 'intelligent release assistant' (docs-37), but this is rule-based automation and API access, not evidence of the product itself generating AI insights or suggestions from data (e.g., anomaly detection, usage analysis, recommended flags/segments).
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
The evidence pack explicitly shows probes for an OpenAPI/Swagger spec and llms.txt returning 404s, and no citation describes an interactive, runnable API reference (e.g., Swagger UI, Postman collection, or live code playground).
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
Flagsmith documents a full Admin API and CLI/MCP integrations, but explicit probes for a machine-readable spec (openapi.json, swagger.json, .well-known/openapi.json) all returned 404, and no evidence pack item points to a downloadable OpenAPI/Swagger spec.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Missing: any statement of API versioning strategy, deprecation timelines, or migration guides for breaking changes.
Rollouts releases — stories about rollouts releases in this arenaGuard a rollout with metrics so a regression is detected and the release is rolled back automatically
nonemoves PA Scoreimpact 20
Flagsmith documents percentage rollouts, scheduled flags, and Release Pipelines with 'triggers and actions,' plus an Experimentation module with Bayesian stats — but none of the evidence describes automatic metric-based regression detection that triggers an automatic rollback of a release.
Experimentation — stories about experimentation in this arenaDefine experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric store
nonemoves PA Scoreimpact 20
Flagsmith's experimentation feature explicitly funnels events into its own 'managed data warehouse' and evaluates them via a built-in Bayesian engine, the opposite of letting a PM define metrics from their own warehouse tables or ingested events — there is no evidence of BYO-warehouse or custom event-source metric definition.
Sdk delivery — stories about sdk delivery in this arenaFlag changes propagate to connected SDKs in seconds via streaming or fast polling — a kill switch actually kills
nonemoves PA Scoreimpact 20
Missing: documented real-time/streaming SDK update mechanism, polling interval specs, and evidence of kill-switch propagation latency.
Showing the top 8 of 28 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map13 surfaces · 40 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Deployment self hosting docs17 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Test against a sandbox environment without touching production data
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructure
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Target flags with attribute-based rules and reusable segments so the right users see the right variation
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- Choose where my data is stored (region/residency)
- Control data retention and deletion
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
Integrating with flagsmith docs16 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Restrict who can change which flags with roles, permissions, and scoped API tokens
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Control data retention and deletion
Managing flags docs12 stories
- Set up automations that run autonomously in the background
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Version, review, and roll back my automations
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Create a feature flag and toggle it live in production within minutes of signing up
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
- Target flags with attribute-based rules and reusable segments so the right users see the right variation
- Require approvals or change requests before production flag changes go live
- Control data retention and deletion
- Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeploying
- Schedule flag changes and releases to happen at a specific future time
Administration and security docs11 stories
- Issue scoped/least-privilege API credentials for an agent
- Subscribe to events via webhooks
- Test against a sandbox environment without touching production data
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keys
- Restrict who can change which flags with roles, permissions, and scoped API tokens
- Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewed
- Require approvals or change requests before production flag changes go live
- Every flag change is recorded in an audit log — who changed what, when, and to which value
- Control data retention and deletion
GitHub README6 stories
- Build against official SDKs
- Test against a sandbox environment without touching production data
- Version, review, and roll back my automations
- Create a feature flag and toggle it live in production within minutes of signing up
- Read the product's source under an open license
- Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behavior
Experimentation docs4 stories
- Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisions
- Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment results
- Run A/B and multivariate experiments on flags and see which variation wins on my metrics
- Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleans
Performance docs4 stories
- Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my network
- Choose where my data is stored (region/residency)
- Evaluate flags at the edge (CDN workers or an edge/relay layer) close to users
- My server SDKs evaluate flags locally from a cached ruleset — microsecond decisions with no network call per flag check
OpenAPI spec3 stories
Hacker News3 stories
docs.flagsmith.com3 stories
llms.txt2 stories
Flagsmith concepts docs2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$docker run -d postgres:16-alpine && docker run -d -p 18000:8000 flagsmith/flagsmith && POST /api/v1/auth/users/ && POST /organisations/ && POST /projects/ && POST /environments/ && POST /features/ {pa_probe_flag} && curl /api/v1/flags/ -H "X-Environment-Key: <key>"reproduced$ docker run -d postgres:16-alpine && docker run -d -p 18000:8000 flagsmith/flagsmith && POST /api/v1/auth/users/ && POST /organisations/ && POST /projects/ && POST /environments/ && POST /features/ {pa_probe_flag} && curl /api/v1/flags/ -H "X-Environment-[redacted]: <[redacted]>"
HTTP/1.1 200 OK
env [redacted]: Gj2V3EBcYavE98vaWQHVxx
{"id":1,"name":"pa_probe_flag","type":"STANDARD","default_enabled":true,"initial_value":null,"created_date":"2026-09-05T
[{"id":1,"feature":{"id":1,"name":"pa_probe_flag","created_date":"2026-09-05T05:01:29.235712Z","description":null,"initial_value":null,"default_enabled":true,"type":"STANDARD"},"feature_state_value":null,"environment":1,"identity":null,"feature_segment":null,"enabled":true}]
$curl -si -X POST https://mcp.flagsmith.com -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -si -X POST https://mcp.flagsmith.com -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401
date: Sat, 05 Sep 2026 05:01:10 GMT
content-type: application/json
content-length: 301
server: uvicorn
www-authenticate: Bearer error="invalid_[redacted]", error_description="Authentication failed. The provided bearer [redacted] is invalid, expired, or no longer recognized by the server. To resolve: clear authentication [redacted]s in your MCP client and reconnect. Your client should automatically re-register and obtain new [redacted]s.", resource_metadata="https://mcp.flagsmith.com/.well-known/oauth-protected-resource"
{"error": "invalid_[redacted]", "error_description": "Authentication failed. The provided bearer [redacted] is invalid, expired, or no longer recognized by the server. To resolve: clear authentication [redacted]s in your MCP client and reconnect. Your client should automatically re-register and obtain new [redacted]s."}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
12 of 24 testable claims verified · 1 contradicted → integrity 42/100
32 distinct capability claims found in Flagsmith’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
12
Verified
11
Unverified
1
Contradicted
16
Undersold
Verified (14)
“Toggle features on/off remotely without deploying new code”
Create a feature flag and toggle it live in production within minutes of signing upfullproof ↗
“Wrap a section of code in a flag and toggle it per environment, user, or segment”
Manage separate environments (dev/staging/prod) with independent flag states and scoped SDK keysfullproof ↗
“Target features to users based on specific traits like geography, subscription level or device type”
Target flags with attribute-based rules and reusable segments so the right users see the right variationfullproof ↗
“Gradually enable a feature for a percentage of users, globally or within a segment”
Roll a flag out progressively by percentage with consistent bucketing, ramping from 1% to 100% without redeployingfullproof ↗
“Multivariate flags define multiple variants with percentage weightings (A/B/n)”
Serve multivariate flags and dynamic configuration values (strings, numbers, JSON), not just booleansfullproof ↗
“Self-host the complete Flagsmith platform on your own infrastructure, including via Docker”
Self-host the full flag platform from an open-source distribution, keeping evaluation data on my infrastructurefullproof ↗
“Self-host the complete Flagsmith platform on your own infrastructure, including via Docker”
“Flagsmith MCP Server exposes the Admin API to AI assistants and agents via Model Context Protocol”
“Official CLI manages flags, segments, features, projects and environments, and evaluates flags like an SDK”
“Manage feature flags and remote config across web, mobile, and server-side apps”
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behaviorpartialproof ↗
“Offers 15+ language SDKs and framework integrations such as React and Next.js”
Use official SDKs across my whole stack — backend, web, and mobile — with consistent flag behaviorpartialproof ↗
“Anything doable in the dashboard can also be done via the Admin API”
Do everything through the API that I can do in the UIfullproof ↗
“Define segments of users via trait-based rules and override flags for whole segments”
Target flags with attribute-based rules and reusable segments so the right users see the right variationfullproof ↗
“Admin API lets you programmatically create, update, and delete projects, environments, flags, segments and users”
Drive the product through a documented public APIfullproof ↗
Unverified (13)
“Override feature flags for individual users (identities) for QA, support or personalization”
Target or exclude specific individual users for a flag (allowlists, beta testers, internal accounts)fullproof ↗
“Scheduled flags queue flag changes to apply automatically at a specified future time”
Schedule flag changes and releases to happen at a specific future timefullproof ↗
“Release pipelines define stages flags progress through automatically, with triggers/actions controlling rollout to environments and audiences”
Define rules that trigger actions automatically on eventspartialproof ↗
“Run end-to-end A/B and multivariate experiments on the platform”
Run A/B and multivariate experiments on flags and see which variation wins on my metricsfullproof ↗
“Analyze experiment results with a built-in Bayesian statistics engine”
Trust a documented statistics engine (Bayesian or frequentist, with variance-reduction options) behind experiment resultspartialproof ↗
“Every action in the admin app is tracked and logged for auditing”
Every flag change is recorded in an audit log — who changed what, when, and to which valuefullproof ↗
“Stream audit logs to your own infrastructure via Audit Log Webhooks”
“Feature Change Requests give a create/approve/publish workflow like GitHub Pull Requests”
Require approvals or change requests before production flag changes go livefullproof ↗
“Role-based access control (RBAC) offers fine-grained permission management of resources”
Restrict who can change which flags with roles, permissions, and scoped API tokensfullproof ↗
“Run a self-hosted Edge Proxy so the Flagsmith Engine runs close to your servers”
Run a relay/edge proxy so flags stay served when the vendor is unreachable and SDK traffic stays inside my networkfullproof ↗
“Restrict which users can modify production environments”
Restrict who can change which flags with roles, permissions, and scoped API tokensfullproof ↗
“Release Pipelines can automatically generate change requests when promoting flags to production”
“Release Pipelines act as an automated release assistant for the entire deployment process”
Define rules that trigger actions automatically on eventspartialproof ↗
Contradicted (1)
“Collect application events into a managed data warehouse to power experiment metrics”
Define experiment metrics from my own data — warehouse tables or ingested events — instead of a black-box metric storenoneproof ↗
Undersold (16)
Point an agent at llms.txt or agent-oriented docspartialproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Read experiment configurations and results programmatically to summarize outcomes and recommend ship/rollback decisionspartialproof ↗
Create and toggle flags through documented APIs, CLIs, or MCP — and the platform can force my changes through approval workflows instead of letting me write to production unreviewedfullproof ↗
Export all of my data in open formats and leavepartialproof ↗
Choose where my data is stored (region/residency)partialproof ↗
Evaluate flags at the edge (CDN workers or an edge/relay layer) close to userspartialproof ↗
Claims outside our story set (5)
Real capability claims found in Flagsmith’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“SAML/SSO authentication support for login”
source ↗“Two-Factor Authentication (2FA) available for account security”
source ↗“Mark flags as protected to prevent accidental deletion”
source ↗“Mark a flag as server-side only so it can't be accessed by client-side SDKs”
source ↗“Environment Key is a public, non-secret key safe to expose in client-side code”
source ↗
Business model
BSD-3 open-source core you can self-host free; SaaS has a free tier then request-volume-based Start-Up pricing and custom Enterprise (SaaS, private cloud, or on-prem).
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP 100% (30d, checked every 6h since Sep 8 '26)
