Rank #1 of 7 in Vector Databases & Memory Stores
Showcase


Try itExperimental
See what an agent can do with Chroma before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$chroma --helprecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Data lifecycle — stories about data lifecycle in this arenaData lifecycleevidence →
Stories about data lifecycle in this arena
Deployment modes — stories about deployment modes in this arenaDeployment modesevidence →
Stories about deployment modes in this arena
Embeddings pipeline — stories about embeddings pipeline in this arenaEmbeddings pipelineevidence →
Stories about embeddings pipeline in this arena
Filtering metadata — stories about filtering metadata in this arenaFiltering metadataevidence →
Stories about filtering metadata in this arena
Multi tenancy scale — stories about multi tenancy scale in this arenaMulti tenancy scaleevidence →
Stories about multi tenancy scale in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Performance latency — stories about performance latency in this arenaPerformance latencyevidence →
Stories about performance latency in this arena
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plansevidence →
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Sdk integrations — stories about sdk integrations in this arenaSdk integrationsevidence →
Stories about sdk integrations in this arena
Search quality hybrid — stories about search quality hybrid in this arenaSearch quality hybridevidence →
Stories about search quality hybrid in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 6 free · 5 paid · 0 enterprise · 22 not stated in evidence
Follow the green: where the map greys out is where Chroma stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Webhooks · Scoped API keys · Versioning policy
Subscribe to events via webhooks
—–
Build against official SDKs
~7/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
✓9/10
Rely on versioned APIs with a documented deprecation policy
—–
Test against a sandbox environment without touching production data
~6/10
Explore an interactive API reference with runnable examples
~4/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓8/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—–
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
—–
Set up automations that run autonomously in the background
n/an/a
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Data lifecycle — stories about data lifecycle in this arenaData lifecycle
Stories about data lifecycle in this arena
Deployment modes — stories about deployment modes in this arenaDeployment modes
Stories about deployment modes in this arena
Embeddings pipeline — stories about embeddings pipeline in this arenaEmbeddings pipeline
Stories about embeddings pipeline in this arena
Filtering metadata — stories about filtering metadata in this arenaFiltering metadata
Stories about filtering metadata in this arena
Multi tenancy scale — stories about multi tenancy scale in this arenaMulti tenancy scale
Stories about multi tenancy scale in this arena
Scaling
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Performance latency — stories about performance latency in this arenaPerformance latency
Stories about performance latency in this arena
Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans
Plan structure and value — what each tier costs and what it unlocks
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Sdk integrations — stories about sdk integrations in this arenaSdk integrations
Stories about sdk integrations in this arena
Search quality hybrid — stories about search quality hybrid in this arenaSearch quality hybrid
Stories about search quality hybrid in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 7/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialfree | 6/10 | Tprobed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partialpaid | 6/10 | Xcommunity | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullfree | 9/10 | Tprobed | |
Have the database generate embeddings at ingest and query time using built-in or configured model providers, instead of running a separate embedding pipeline C Embeddings | ml-engineer | Embeddings pipeline — stories about embeddings pipeline in this arenaEmbeddings pipeline | 3 | full | 8/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partialfree | 5/10 | Xcommunity | |
Isolate many tenants cheaply using namespaces, partitions, or per-tenant collections with documented limits C Tenancy | platform-engineer | Multi tenancy scale — stories about multi tenancy scale in this arenaMulti tenancy scale | 3 | partialpaid | 5/10 | Xcommunity | |
Run approximate nearest-neighbor similarity search over embeddings with configurable distance metrics C Core search | developer | Search quality hybrid — stories about search quality hybrid in this arenaSearch quality hybrid | 3 | partial | 5/10 | Cclaimed | |
Combine dense vector search with keyword or sparse (BM25-style) signals in one hybrid query with fusion ranking C Hybrid | developer | Search quality hybrid — stories about search quality hybrid in this arenaSearch quality hybrid | 3 | partial | 4/10 | Cclaimed | |
Filter vector search by structured metadata conditions without wrecking recall or latency C Filtering | developer | Filtering metadata — stories about filtering metadata in this arenaFiltering metadata | 3 | partial | 4/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | n/a | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | fullfree | 9/10 | Xcommunity | |
Run the database embedded in-process or as a lightweight local instance for development and small workloads C Local dev | developer | Deployment modes — stories about deployment modes in this arenaDeployment modes | 2 | fullfree | 8/10 | Tprobed | |
Pay serverless usage-based pricing with transparent per-unit costs instead of provisioning fixed clusters G Pricing | developer | Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans | 2 | fullpaid | 7/10 | Xcommunity | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Plug the database into RAG and agent frameworks (LangChain, LlamaIndex, etc.) through maintained first-class integrations C Integrations | ml-engineer | Sdk integrations — stories about sdk integrations in this arenaSdk integrations | 2 | partial | 6/10 | Tprobed | |
Run keyword/full-text search over documents inside the database without bolting on a separate search engine C Hybrid | developer | Search quality hybrid — stories about search quality hybrid in this arenaSearch quality hybrid | 2 | partial | 6/10 | Xcommunity | |
Scale beyond one node with sharding or distributed deployment C Scaling | platform-engineer | Multi tenancy scale — stories about multi tenancy scale in this arenaMulti tenancy scale | 2 | partial | 6/10 | Xcommunity | |
Use a fully managed cloud version of the database with programmatic provisioning C Managed cloud | developer | Deployment modes — stories about deployment modes in this arenaDeployment modes | 2 | partialpaid | 6/10 | Xcommunity | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Tprobed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 4/10 | Cclaimed | |
Back up collections with snapshots and restore them C Backup | platform-engineer | Data lifecycle — stories about data lifecycle in this arenaData lifecycle | 2 | partialpaid | 3/10 | Cclaimed | |
Build against official SDKs in the major languages (Python, TypeScript, Go, Java) G Sdks | developer | Sdk integrations — stories about sdk integrations in this arenaSdk integrations | 2 | partial | 3/10 | Cclaimed | |
Express rich filter conditions (ranges, geo, nested boolean logic, array membership) in queries C Filtering | developer | Filtering metadata — stories about filtering metadata in this arenaFiltering metadata | 2 | partial | 3/10 | Cclaimed | |
Replicate data across nodes or zones for high availability with a documented consistency model C Scaling | platform-engineer | Multi tenancy scale — stories about multi tenancy scale in this arenaMulti tenancy scale | 2 | partial | 3/10 | Xcommunity | |
Enable vector quantization or compression to cut memory and storage cost with a documented accuracy trade-off C Index tuning | ml-engineer | Performance latency — stories about performance latency in this arenaPerformance latency | 2 | none | 0/10 | ||
Rerank search results with built-in or first-party-integrated reranking models C Reranking | ml-engineer | Search quality hybrid — stories about search quality hybrid in this arenaSearch quality hybrid | 2 | none | 0/10 | ||
Bulk-import and bulk-export vectors plus metadata in documented formats C Portability | developer | Data lifecycle — stories about data lifecycle in this arenaData lifecycle | 2 | none | untested | none yet | |
Enforce granular access control (API keys, roles, per-collection permissions) on database operations C Tenancy | platform-engineer | Multi tenancy scale — stories about multi tenancy scale in this arenaMulti tenancy scale | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | n/a | untested | none yet | |
See published benchmarks or measured latency/recall numbers backing the database's performance claims C Benchmarks | platform-engineer | Performance latency — stories about performance latency in this arenaPerformance latency | 2 | none | untested | none yet | |
Tune index parameters (HNSW graph settings, index types) to trade recall against latency and memory C Index tuning | ml-engineer | Performance latency — stories about performance latency in this arenaPerformance latency | 2 | none | untested | none yet | |
Upsert and delete records continuously and have changes reflected in search results quickly, with documented freshness/consistency behavior C Freshness | developer | Data lifecycle — stories about data lifecycle in this arenaData lifecycle | 2 | none | untested | none yet | |
Prototype on a meaningful free tier before paying anything G Pricing | developer | Pricing plans — plan structure and value — what each tier costs and what it unlocksPricing plans | 1 | fullfree | 8/10 | Xcommunity | |
Deploy to production on Kubernetes with an official Helm chart or operator C Self managed | platform-engineer | Deployment modes — stories about deployment modes in this arenaDeployment modes | 1 | none | 0/10 | ||
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 38 stories with headroom
What would move Chroma’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
Chroma is a vector database; the evidence pack contains no statement about data-training opt-out policies, data usage terms, or privacy commitments regarding whether user data is used to train AI models.
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product
nonemoves Built-in AIimpact 30
Chroma is positioned as a vector/embedding database and retrieval backend (storage, indexing, querying, MCP-based agent access) rather than a product that itself surfaces AI-generated insights or suggestions to the user; none of the evidence describes built-in analytics, summarization, or recommendation features inside Chroma's own interface.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
No evidence of scoped or least-privilege API credential issuance for agents; the evidence pack covers embeddings, querying, CLI, MCP server, and pricing but never mentions API key scoping, RBAC, or per-agent credential management in Chroma Cloud or self-hosted deployments.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
No evidence in the pack mentions webhooks or event subscription mechanisms for Chroma; the product is a vector database and its documented integrations (MCP, CLI, APIs) do not include any webhook/event-push feature.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Missing: any mention of API versioning, changelog, or deprecation/backward-compatibility policy.
Multi tenancy scale — stories about multi tenancy scale in this arenaEnforce granular access control (API keys, roles, per-collection permissions) on database operations
nonemoves PA Scoreimpact 20
No evidence in the pack mentions API keys, roles, RBAC, or per-collection permission scoping; docs cover embedding, querying, CLI, MCP, forking, and cloud pricing but nothing about access control mechanisms.
Data lifecycle — stories about data lifecycle in this arenaBulk-import and bulk-export vectors plus metadata in documented formats
nonemoves PA Scoreimpact 20
No evidence pack item documents a bulk-import or bulk-export feature, file format spec, or CLI/API command for moving vectors+metadata in/out of Chroma; forking (chroma-docs-8/18/26) is copy-on-write cloning, not data export/import.
Showing the top 8 of 38 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map9 surfaces · 33 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs20 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Build against official SDKs
- Operate the product with natural-language commands
- Explore an interactive API reference with runnable examples
- Perform bulk operations across many items at once
- Run the database embedded in-process or as a lightweight local instance for development and small workloads
- Use a fully managed cloud version of the database with programmatic provisioning
- Have the database generate embeddings at ingest and query time using built-in or configured model providers, instead of running a separate embedding pipeline
- Filter vector search by structured metadata conditions without wrecking recall or latency
- Express rich filter conditions (ranges, geo, nested boolean logic, array membership) in queries
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Self-host the core product
- Build against official SDKs in the major languages (Python, TypeScript, Go, Java)
- Run approximate nearest-neighbor similarity search over embeddings with configurable distance metrics
- Run keyword/full-text search over documents inside the database without bolting on a separate search engine
- Combine dense vector search with keyword or sparse (BM25-style) signals in one hybrid query with fusion ranking
docs.trychroma.com16 stories
- Run the product headlessly / in CI for automation
- Build against official SDKs
- Test against a sandbox environment without touching production data
- Run the database embedded in-process or as a lightweight local instance for development and small workloads
- Use a fully managed cloud version of the database with programmatic provisioning
- Have the database generate embeddings at ingest and query time using built-in or configured model providers, instead of running a separate embedding pipeline
- Scale beyond one node with sharding or distributed deployment
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- Prototype on a meaningful free tier before paying anything
- Pay serverless usage-based pricing with transparent per-unit costs instead of provisioning fixed clusters
- Choose where my data is stored (region/residency)
- Control data retention and deletion
- Run approximate nearest-neighbor similarity search over embeddings with configurable distance metrics
Hacker News12 stories
- Test against a sandbox environment without touching production data
- Run the database embedded in-process or as a lightweight local instance for development and small workloads
- Use a fully managed cloud version of the database with programmatic provisioning
- Scale beyond one node with sharding or distributed deployment
- Replicate data across nodes or zones for high availability with a documented consistency model
- Isolate many tenants cheaply using namespaces, partitions, or per-tenant collections with documented limits
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- Prototype on a meaningful free tier before paying anything
- Pay serverless usage-based pricing with transparent per-unit costs instead of provisioning fixed clusters
- Run keyword/full-text search over documents inside the database without bolting on a separate search engine
trychroma.com10 stories
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Back up collections with snapshots and restore them
- Scale beyond one node with sharding or distributed deployment
- Replicate data across nodes or zones for high availability with a documented consistency model
- Isolate many tenants cheaply using namespaces, partitions, or per-tenant collections with documented limits
- Choose where my data is stored (region/residency)
- Control data retention and deletion
- Run keyword/full-text search over documents inside the database without bolting on a separate search engine
- Combine dense vector search with keyword or sparse (BM25-style) signals in one hybrid query with fusion ranking
Cloud docs9 stories
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Back up collections with snapshots and restore them
- Scale beyond one node with sharding or distributed deployment
- Isolate many tenants cheaply using namespaces, partitions, or per-tenant collections with documented limits
- Do everything through the API that I can do in the UI
- Pay serverless usage-based pricing with transparent per-unit costs instead of provisioning fixed clusters
- Choose where my data is stored (region/residency)
- Control data retention and deletion
OpenAPI spec5 stories
llms.txt4 stories
GitHub README4 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$chroma --helpreproduced$ chroma --help A CLI for Chroma Usage: chroma <COMMAND> Commands: browse Browse Chroma collections copy Copy collection between local and Chroma Cloud db Manage Chroma Cloud databases docs Open Chroma online documentation install Install sample applications login Log in to Chroma Cloud profile Manage Chroma Cloud profiles run Start a local Chroma server support Open the Chroma Discord update Check for Chroma CLI updates vacuum Vacuum a local Chroma persistent directory help Print this message or the help of the given subcommand(s) Options: -h, --help Print help -V, --version Print version
$chroma run --path /tmp/pa-chroma-probe --port 8765 & curl -X POST localhost:8765/api/v2/tenants/default_tenant/databases/default_database/collections -d '{"name":"pa_probe"}' && curl localhost:8765/api/v2/tenants/default_tenant/databases/default_database/collectionsreproduced$ chroma run --path /tmp/pa-chroma-probe --port 8765 & curl -X POST localhost:8765/api/v2/tenants/default_tenant/databases/default_database/collections -d '{"name":"pa_probe"}' && curl localhost:8765/api/v2/tenants/default_tenant/databases/default_database/collections
{"nanosecond heartbeat":1788560721960487000}
{"id":"cb4da00a-c4cb-4e9d-9781-9b5621b9aaec","name":"pa_probe","configuration_json":{"hnsw":{"space":"l2","ef_construction":100,"ef_search":100,"max_neighbors":16,"resize_factor":1.2,"sync_threshold":1000},"spann":null,"embedding_function":
[{"id":"cb4da00a-c4cb-4e9d-9781-9b5621b9aaec","name":"pa_probe","configuration_json":{"hnsw":{"space":"l2","ef_construction":100,"ef_search":100,"max_neighbors":16,"resize_factor":1.2,"sync_threshold":1000},"spann":null,"embedding_function"
$echo '<jsonrpc initialize>' | uvx chroma-mcp --client-type ephemeralreproduced$ echo '<jsonrpc initialize>' | uvx chroma-mcp --client-type ephemeral
⠋ Resolving dependencies...
⠙ Resolving dependencies...
⠋ Resolving dependencies...
⠙ Resolving dependencies...
⠙ chroma-mcp==0.2.6
⠙ mcp==1.6.0
⠙ mcp==1.6.0
⠙ chromadb==1.5.9
⠙ cohere==7.1.1
⠙ httpx==0.28.1
⠙ openai==3.8.0
⠙ pillow==12.3.0
⠙ python-dotenv==1.2.3
⠙ typing-extensions==4.16.0
⠙ voyageai==0.5.0
⠙ anyio==4.15.0
⠙ typing-extensions==4.16.0
⠙ httpx-sse==0.4.3
⠙ pydantic-settings==2.15.0
⠙ pydantic==2.13.5
⠙ pydantic-core==2.46.5
⠙ sse-starlette==3.4.10
/Users/judegomila/.cache/uv/archive-v0/JD11Ic7cM6QK165B/lib/python3.13/site-packages/pydantic_settings/sources/utils.py:47: IncompleteFieldDefinitionWarning: Field 'lifespan' has an incomplete definition: its annotation contains an unresolved forward reference, so settings sources may fail to correctly resolve its value. Call `model_rebuild()` on the model where the field is defined, once all the referenced types are defined.
warnings.warn(
Successfully initialized Chroma client
Starting MCP server
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2024-11-05","capabilities":{"experimental":{},"prompts":{"listChanged":false},"resources":{"subscribe":false,"listChanged":false},"tools":{"listChanged":false}},"serverInfo":{"name":"chroma","version":"1.6.0"}}}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
10 of 16 testable claims verified · 0 contradicted → integrity 63/100
20 distinct capability claims found in Chroma’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
10
Verified
6
Unverified
0
Contradicted
17
Undersold
Verified (11)
“Supports full-text search via $contains/$not_contains and regex pattern matching via $regex/$not_regex”
Run keyword/full-text search over documents inside the database without bolting on a separate search enginepartialproof ↗
“Official CLI runs a local server, installs sample apps, browses collections, and manages Chroma Cloud DBs”
“Open source under Apache 2.0 license”
“Can run locally, self-host, or use a managed serverless Chroma Cloud offering”
“Can run locally, self-host, or use a managed serverless Chroma Cloud offering”
Use a fully managed cloud version of the database with programmatic provisioningpartialproof ↗
“Provides ready-made agent setup prompts so tools like Claude Code, Cursor, or Codex can configure Chroma automatically”
Point an agent at llms.txt or agent-oriented docsfullproof ↗
“MCP server lets Claude directly use Chroma's search capabilities via a standardized protocol”
“MCP integration enables persistent memory across conversations”
“First-class integration with LangChain for building RAG applications”
Plug the database into RAG and agent frameworks (LangChain, LlamaIndex, etc.) through maintained first-class integrationspartialproof ↗
“BYOC in your VPC with multi-cloud/multi-region replication for a resilient, scalable search system”
Replicate data across nodes or zones for high availability with a documented consistency modelpartialproof ↗
“You can create a database and try it out in under 30 seconds with $5 of free credits”
Prototype on a meaningful free tier before paying anythingfullproof ↗
Unverified (7)
“Automatically embeds and indexes stored text without a separate embedding pipeline”
Have the database generate embeddings at ingest and query time using built-in or configured model providers, instead of running a separate embedding pipelinefullproof ↗
“Lets you swap in different embedding models (OpenAI, Cohere, Hugging Face, sentence-transformers, etc.)”
Have the database generate embeddings at ingest and query time using built-in or configured model providers, instead of running a separate embedding pipelinefullproof ↗
“Query a collection with text queries and get back the n most similar results”
Run approximate nearest-neighbor similarity search over embeddings with configurable distance metricspartialproof ↗
“get/query support combining document content search (where_document) with metadata filters”
Filter vector search by structured metadata conditions without wrecking recall or latencypartialproof ↗
“Offers sparse vector / lexical search using BM25 and SPLADE alongside dense vector search”
Combine dense vector search with keyword or sparse (BM25-style) signals in one hybrid query with fusion rankingpartialproof ↗
“Offers BYOC (bring your own cloud) for single-tenant deployments”
Choose where my data is stored (region/residency)partialproof ↗
“Point-in-time recovery is supported as part of the resilient BYOC/Cloud architecture”
Back up collections with snapshots and restore thempartialproof ↗
Undersold (17)
Run the product headlessly / in CI for automationpartialproof ↗
Drive the product through a documented public APIfullproof ↗
Operate the product with natural-language commandspartialproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)fullproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Run the database embedded in-process or as a lightweight local instance for development and small workloadsfullproof ↗
Express rich filter conditions (ranges, geo, nested boolean logic, array membership) in queriespartialproof ↗
Scale beyond one node with sharding or distributed deploymentpartialproof ↗
Isolate many tenants cheaply using namespaces, partitions, or per-tenant collections with documented limitspartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Pay serverless usage-based pricing with transparent per-unit costs instead of provisioning fixed clustersfullproof ↗
Build against official SDKs in the major languages (Python, TypeScript, Go, Java)partialproof ↗
Claims outside our story set (3)
Real capability claims found in Chroma’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Supports collection forking for dataset versioning, A/B testing, and roll-outs”
source ↗“Forks are copy-on-write, only charging for incremental storage written after the fork”
source ↗“Can index and search images, audio, and other modalities alongside text”
source ↗
Pricing signals
- $2.5per GiB writtenpay-as-you-goCost per logical GiB written via add/update/upsertsource ↗as of 2026-09-07
- $0.33per GB-monthpay-as-you-goStorage cost per GiB per month, prorated hourlysource ↗as of 2026-09-07
- $0.0075per TB queriedpay-as-you-goRead cost per TiB of data queriedsource ↗as of 2026-09-07
Extracted verbatim from the vendor’s own pricing page — hover a figure for the exact quote.
Business model
Apache-2.0 open-source embedding database that runs embedded or as a server for free; Chroma Cloud charges usage-based pricing for writes, queries, and storage with $5 starting credit.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% · openapi.json 100% (30d, checked every 6h since Sep 8 '26)
