Rank #1 of 4 in Data Warehouses & Lakehouses
Install
Showcase


Try itExperimental
See what an agent can do with Databricks before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$databricks aitools install --path /tmp/pa-dbx-skills # vendor-shipped agent skills, then list + print one SKILL.md frontmatterrecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agent analytics — stories about agent analytics in this arenaAgent analyticsevidence →
Stories about agent analytics in this arena
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Cost economics — stories about cost economics in this arenaCost economicsevidence →
Stories about cost economics in this arena
Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrationsevidence →
The surrounding ecosystem — integrations, marketplaces, community packages
Governance access — stories about governance access in this arenaGovernance accessevidence →
Stories about governance access in this arena
Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelinesevidence →
Stories about ingestion pipelines in this arena
Notebooks workspace — stories about notebooks workspace in this arenaNotebooks workspaceevidence →
Stories about notebooks workspace in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Semantic layer — stories about semantic layer in this arenaSemantic layerevidence →
Stories about semantic layer in this arena
Sharing marketplace — stories about sharing marketplace in this arenaSharing marketplaceevidence →
Stories about sharing marketplace in this arena
Sql analytics — stories about sql analytics in this arenaSql analyticsevidence →
Stories about sql analytics in this arena
Streaming realtime — stories about streaming realtime in this arenaStreaming realtimeevidence →
Stories about streaming realtime in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 2 free · 1 paid · 0 enterprise · 40 not stated in evidence
Follow the green: where the map greys out is where Databricks stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agent analytics — stories about agent analytics in this arenaAgent analytics
Stories about agent analytics in this arena
Agent ops
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Webhooks · Machine-readable spec · API sandbox · API/UI parity
Subscribe to events via webhooks
—–
Build against official SDKs
✓9/10
Issue scoped/least-privilege API credentials for an agent
~4/10
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
~6/10
Test against a sandbox environment without touching production data
—–
Explore an interactive API reference with runnable examples
~4/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓8/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓8/10
Operate the product with natural-language commands
✓7/10
Plug MCP servers into this product so it can use their tools
✓8/10
Get AI-generated insights and suggestions from my data inside the product
✓8/10
Set up automations that run autonomously in the background
✓7/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Cost economics — stories about cost economics in this arenaCost economics
Stories about cost economics in this arena
Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations
The surrounding ecosystem — integrations, marketplaces, community packages
Standard drivers (JDBC/ODBC) and documented BI-tool integrations connect my dashboards without custom glue
✓7/10
I get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production compute
~5/10
Dbt is a first-class citizen — a documented adapter or native dbt project support with vendor docs to match
✓8/10
Governance access — stories about governance access in this arenaGovernance access
Stories about governance access in this arena
Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines
Stories about ingestion pipelines in this arena
Notebooks workspace — stories about notebooks workspace in this arenaNotebooks workspace
Stories about notebooks workspace in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Semantic layer — stories about semantic layer in this arenaSemantic layer
Stories about semantic layer in this arena
Sharing marketplace — stories about sharing marketplace in this arenaSharing marketplace
Stories about sharing marketplace in this arena
Sql analytics — stories about sql analytics in this arenaSql analytics
Stories about sql analytics in this arena
Query open table formats and files in object storage — Iceberg, Delta, Parquet — without first loading them into proprietary storage
~6/10
Inspect query profiles and execution plans to find why a query is slow or expensive
✓8/10
Time-travel — query data as of a past point and restore dropped or corrupted tables from history
✓9/10
I get a full analytical SQL surface — window functions, CTEs, semi-structured JSON, arrays, and rich date/time types — without bolt-on extensions
~5/10
Streaming realtime — stories about streaming realtime in this arenaStreaming realtime
Stories about streaming realtime in this arena
Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Cclaimed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | untested | none yet | |
My agent can run governed SQL end to end — authenticate, discover schemas, query, and read results back through a CLI or API with no dashboard in the loop C Agent ops | ai-native user | Agent analytics — stories about agent analytics in this arenaAgent analytics | 3 | full | 8/10 | Tprobed | |
Bulk-load CSV, JSON, and Parquet from cloud object storage with a single documented command C Loading | data-engineer | Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines | 3 | partial | 6/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 6/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 6/10 | Xcommunity | |
I get a full analytical SQL surface — window functions, CTEs, semi-structured JSON, arrays, and rich date/time types — without bolt-on extensions C Sql | analyst | Sql analytics — stories about sql analytics in this arenaSql analytics | 3 | partial | 5/10 | Cclaimed | |
The pricing model is documented clearly enough that I can estimate a monthly bill for my workload before committing G Pricing | platform-engineer | Cost economics — stories about cost economics in this arenaCost economics | 3 | partialpaid | 5/10 | Xcommunity | |
Access control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safely C Access | platform-engineer | Governance access — stories about governance access in this arenaGovernance access | 3 | disputed | 4/10 | Dcontradicted | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Share live datasets with another account or organization without copying data or building an export pipeline C Sharing | data-engineer | Sharing marketplace — stories about sharing marketplace in this arenaSharing marketplace | 2 | full | 9/10 | Cclaimed | |
Time-travel — query data as of a past point and restore dropped or corrupted tables from history C Recovery | data-engineer | Sql analytics — stories about sql analytics in this arenaSql analytics | 2 | full | 9/10 | Cclaimed | |
A built-in AI assistant writes, fixes, and explains SQL against my schemas from natural language, inside the product C Agent ops | ai-native user | Agent analytics — stories about agent analytics in this arenaAgent analytics | 2 | full | 8/10 | Cclaimed | |
A managed service continuously ingests new files or events as they arrive, without me running my own pipeline infrastructure C Loading | data-engineer | Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines | 2 | full | 8/10 | Cclaimed | |
Dbt is a first-class citizen — a documented adapter or native dbt project support with vendor docs to match C Transformation | data-engineer | Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations | 2 | full | 8/10 | Cclaimed | |
Define a governed semantic model — metrics, dimensions, and joins declared once — that queries and AI tools answer against consistently C Semantics | analyst | Semantic layer — stories about semantic layer in this arenaSemantic layer | 2 | full | 8/10 | Cclaimed | |
First-party notebooks let me mix SQL and Python against warehouse data, with results and charts inline C Notebooks | analyst | Notebooks workspace — stories about notebooks workspace in this arenaNotebooks workspace | 2 | full | 8/10 | Cclaimed | |
I get audit logs of who ran what and column-level lineage of where data came from G Governance | platform-engineer | Governance access — stories about governance access in this arenaGovernance access | 2 | full | 8/10 | Cclaimed | |
Inspect query profiles and execution plans to find why a query is slow or expensive C Performance | data-engineer | Sql analytics — stories about sql analytics in this arenaSql analytics | 2 | full | 8/10 | Cclaimed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | full | 8/10 | Tprobed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | full | 7/10 | Tprobed | |
Streaming writes land queryable within seconds through a documented streaming ingestion API C Streaming | data-engineer | Streaming realtime — stories about streaming realtime in this arenaStreaming realtime | 2 | full | 7/10 | Cclaimed | |
First-party and partner connectors cover my sources — SaaS apps, databases, and ETL/ELT tools — with documented setup C Connectors | data-engineer | Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines | 2 | partial | 6/10 | Cclaimed | |
Query open table formats and files in object storage — Iceberg, Delta, Parquet — without first loading them into proprietary storage C Lakehouse | data-engineer | Sql analytics — stories about sql analytics in this arenaSql analytics | 2 | partial | 6/10 | Cclaimed | |
Budgets, resource monitors, or auto-suspend stop a runaway query or idle compute from burning money overnight G Pricing | platform-engineer | Cost economics — stories about cost economics in this arenaCost economics | 2 | partial | 5/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | disputed | 5/10 | Dcontradicted | |
I get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production compute C Dev loop | data-engineer | Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations | 2 | partialfree | 5/10 | Xcommunity | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | n/a | untested | none yet | |
Run continuous or incremental transformations — streams, tasks, declarative pipelines, or continuous queries — inside the platform C Streaming | data-engineer | Streaming realtime — stories about streaming realtime in this arenaStreaming realtime | 1 | full | 9/10 | Cclaimed | |
A marketplace of third-party datasets lets me enrich my own data directly inside the platform C Sharing | analyst | Sharing marketplace — stories about sharing marketplace in this arenaSharing marketplace | 1 | full | 8/10 | Cclaimed | |
Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joins C Agent ops | ai-native user | Agent analytics — stories about agent analytics in this arenaAgent analytics | 1 | full | 8/10 | Cclaimed | |
Compliance attestations (SOC 2, HIPAA, PCI) are documented so security review does not stall the rollout C Governance | platform-engineer | Governance access — stories about governance access in this arenaGovernance access | 1 | full | 8/10 | Cclaimed | |
Evaluate with a free tier or trial — real queries on real data without a credit card or a sales call G Trial | analyst | Cost economics — stories about cost economics in this arenaCost economics | 1 | fullfree | 7/10 | Cclaimed | |
Standard drivers (JDBC/ODBC) and documented BI-tool integrations connect my dashboards without custom glue C Bi | analyst | Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations | 1 | full | 7/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 5/10 | Cclaimed |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 21 stories with headroom
What would move Databricks’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Databricks is offered exclusively as a managed cloud service (AWS/Azure/GCP workspaces, pay-as-you-go billing) with no documented option to self-host the core platform on your own infrastructure; the closest analog (Free Edition) is still a hosted SaaS trial, not a self-hostable deployment.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
Missing: an explicit AI-training opt-out/data-use policy, documentation of contractual or technical controls preventing model training on customer data, and any independent confirmation of such a guarantee.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
No evidence in the pack describes any webhook subscription mechanism (e.g., event notifications pushed to external endpoints); Databricks docs reference REST APIs, SQL alerts, and MCP tools but nothing about webhooks for event subscription.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
Databricks publishes an extensive human-readable REST API reference (databricks-docs-65, databricks-supp-versioned-apis, databricks-supp-api-reference-examples) but there is no evidence of a downloadable machine-readable spec — the probe explicitly found openapi.json/swagger.json/api/openapi.json/.well-known/openapi.json all return 404 (databricks-probe-2), and the API reference page confirms 'No live try-it runner is documented.' This is an applicable axis for a platform with a large REST API surface, so absent evidence of an OpenAPI/Swagger artifact this is 'none' rather than 'na'.
Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency)
nonemoves PA Scoreimpact 20
No evidence in the pack addresses regional deployment, data residency controls, or workspace region selection for Databricks; the compliance mention only covers SOC 2 attestations, not data-location choice.
Privacy posture — data-handling and privacy storiesControl data retention and deletion
nonemoves PA Scoreimpact 20
Missing: explicit retention configuration, right-to-delete/erasure mechanisms, data lifecycle/expiry policy documentation, and any independent verification of deletion behavior.
Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking
nonemoves PA Scoreimpact 20
Missing: any mention of telemetry collection, opt-out settings, or usage-tracking controls.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
partialq4/10moves agent-readyimpact 18
Missing: explicit credential/token issuance workflow scoped to an agent identity, documentation of OAuth/service-principal scoping granularity, and any hands-on or independent confirmation that credentials can be narrowly scoped per-agent.
Showing the top 8 of 21 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 45 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Aws docs45 stories
- My agent can run governed SQL end to end — authenticate, discover schemas, query, and read results back through a CLI or API with no dashboard in the loop
- A built-in AI assistant writes, fixes, and explains SQL against my schemas from natural language, inside the product
- Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joins
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Plug MCP servers into this product so it can use their tools
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Explore an interactive API reference with runnable examples
- Rely on versioned APIs with a documented deprecation policy
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Version, review, and roll back my automations
- The pricing model is documented clearly enough that I can estimate a monthly bill for my workload before committing
- Budgets, resource monitors, or auto-suspend stop a runaway query or idle compute from burning money overnight
- Evaluate with a free tier or trial — real queries on real data without a credit card or a sales call
- Standard drivers (JDBC/ODBC) and documented BI-tool integrations connect my dashboards without custom glue
- I get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production compute
- Dbt is a first-class citizen — a documented adapter or native dbt project support with vendor docs to match
- Access control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safely
- I get audit logs of who ran what and column-level lineage of where data came from
- Compliance attestations (SOC 2, HIPAA, PCI) are documented so security review does not stall the rollout
- First-party and partner connectors cover my sources — SaaS apps, databases, and ETL/ELT tools — with documented setup
- Bulk-load CSV, JSON, and Parquet from cloud object storage with a single documented command
- A managed service continuously ingests new files or events as they arrive, without me running my own pipeline infrastructure
- First-party notebooks let me mix SQL and Python against warehouse data, with results and charts inline
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Define a governed semantic model — metrics, dimensions, and joins declared once — that queries and AI tools answer against consistently
- A marketplace of third-party datasets lets me enrich my own data directly inside the platform
- Share live datasets with another account or organization without copying data or building an export pipeline
- Query open table formats and files in object storage — Iceberg, Delta, Parquet — without first loading them into proprietary storage
- Inspect query profiles and execution plans to find why a query is slow or expensive
- Time-travel — query data as of a past point and restore dropped or corrupted tables from history
- I get a full analytical SQL surface — window functions, CTEs, semi-structured JSON, arrays, and rich date/time types — without bolt-on extensions
- Run continuous or incremental transformations — streams, tasks, declarative pipelines, or continuous queries — inside the platform
- Streaming writes land queryable within seconds through a documented streaming ingestion API
API reference9 stories
- My agent can run governed SQL end to end — authenticate, discover schemas, query, and read results back through a CLI or API with no dashboard in the loop
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Rely on versioned APIs with a documented deprecation policy
- Perform bulk operations across many items at once
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
Hacker News5 stories
- The pricing model is documented clearly enough that I can estimate a monthly bill for my workload before committing
- I get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production compute
- Access control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safely
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
databricks.com4 stories
- A built-in AI assistant writes, fixes, and explains SQL against my schemas from natural language, inside the product
- Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joins
- Get AI-generated insights and suggestions from my data inside the product
- Operate the product with natural-language commands
Product docs1 story
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$databricks aitools install --path /tmp/pa-dbx-skills # vendor-shipped agent skills, then list + print one SKILL.md frontmatterreproduced$ databricks aitools install --path /tmp/pa-dbx-skills # vendor-shipped agent skills, then list + print one SKILL.md frontmatter Using skills version 0.2.10 Wrote 29 skills to /tmp/pa-dbx-skills --- installed skills --- /tmp/pa-dbx-skills/databricks-agent-bricks/SKILL.md /tmp/pa-dbx-skills/databricks-ai-functions/SKILL.md /tmp/pa-dbx-skills/databricks-aibi-dashboards/SKILL.md /tmp/pa-dbx-skills/databricks-app-design/SKILL.md /tmp/pa-dbx-skills/databricks-apps-python/SKILL.md /tmp/pa-dbx-skills/databricks-apps/SKILL.md /tmp/pa-dbx-skills/databricks-core/SKILL.md /tmp/pa-dbx-skills/databricks-dabs/SKILL.md --- frontmatter of databricks-docs --- --- name: databricks-docs description: "Databricks documentation reference via llms.txt index. Use when other skills do not cover a topic, looking up unfamiliar Databricks features, or needing authoritative docs on APIs, configurations, or platform capabilities."
$databricks --version # installed via `brew tap databricks/tap && brew install databricks`reproduced$ databricks --version # installed via `brew tap databricks/tap && brew install databricks` Databricks CLI v1.15.0
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
7 of 32 testable claims verified · 1 contradicted → integrity 16/100
47 distinct capability claims found in Databricks’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
7
Verified
24
Unverified
1
Contradicted
12
Undersold
Verified (10)
“Official CLI lets you interact with the platform from a local terminal or automation scripts”
“Host a custom MCP server as a Databricks app to expose your own tools”
“External clients like Claude, Cursor, and MCP Inspector can connect to MCP servers hosted on Databricks”
“Agents can be given governed access to third-party SaaS tools such as Slack, GitHub, Google Drive/Calendar, and Gmail”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Create AI agent tools from Unity Catalog functions, including third-party integrations and code-interpreter tools”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Agent code can discover, authenticate, and call managed/custom MCP servers, then be deployed on Databricks Apps”
Plug MCP servers into this product so it can use their toolsfullproof ↗
“Unity Catalog objects can be managed via Catalog Explorer, SQL, the CLI, and REST APIs”
Drive the product through a documented public APIfullproof ↗
“Pay-as-you-go pricing with no upfront cost, billed per-second for products used”
The pricing model is documented clearly enough that I can estimate a monthly bill for my workload before committingpartialproof ↗
“Databricks Connect lets developers write and debug code on Databricks compute from any IDE's native tools”
I get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production computepartialproof ↗
“Databricks SDK for Python lets you automate Databricks operations and speed up development”
Unverified (32)
“Build ETL pipelines for data orchestration using Lakeflow pipelines and Auto Loader”
Run continuous or incremental transformations — streams, tasks, declarative pipelines, or continuous queries — inside the platformfullproof ↗
“Structured Streaming lets you write streaming computations the same way as batch jobs”
Run continuous or incremental transformations — streams, tasks, declarative pipelines, or continuous queries — inside the platformfullproof ↗
“Auto Loader incrementally and efficiently processes new files as they arrive in cloud storage”
A managed service continuously ingests new files or events as they arrive, without me running my own pipeline infrastructurefullproof ↗
“Genie Code is a built-in AI coding and data assistant for developers in the workspace”
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
“Genie Code generates and runs code, builds pipelines/dashboards, debugs errors, and works with Unity Catalog tables and lineage”
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
“Genie Code can run as an autonomous agent that plans, executes, fixes errors and asks approval before using tools”
Set up automations that run autonomously in the backgroundfullproof ↗
“Write and run SQL with integrated AI assistance, inline code comments, and version history”
A built-in AI assistant writes, fixes, and explains SQL against my schemas from natural language, inside the productfullproof ↗
“Automatic insights and recommendations are surfaced when queries run inefficiently”
Get AI-generated insights and suggestions from my data inside the productfullproof ↗
“A semantic layer lets you define a metric once and let users group it by any available field”
Define a governed semantic model — metrics, dimensions, and joins declared once — that queries and AI tools answer against consistentlyfullproof ↗
“Metric views can be queried from external BI tools like Power BI, Tableau, and Sigma”
Standard drivers (JDBC/ODBC) and documented BI-tool integrations connect my dashboards without custom gluefullproof ↗
“Metric views can be queried from external BI tools like Power BI, Tableau, and Sigma”
Define a governed semantic model — metrics, dimensions, and joins declared once — that queries and AI tools answer against consistentlyfullproof ↗
“Ask data questions in natural language and get answers grounded in your organization's data”
Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joinsfullproof ↗
“Genie One interface lets business users view dashboards, ask NL questions, and discover shared assets”
Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joinsfullproof ↗
“AI/BI supports natural-language dashboard creation and deep conversational analytics from the start”
Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joinsfullproof ↗
“dbt Core support: install, connect, write dbt code locally and run from the CLI”
Dbt is a first-class citizen — a documented adapter or native dbt project support with vendor docs to matchfullproof ↗
“Notebooks provide real-time multi-language coauthoring, automatic versioning, and built-in visualizations”
First-party notebooks let me mix SQL and Python against warehouse data, with results and charts inlinefullproof ↗
“Notebooks can query sample data in Unity Catalog and visualize results inline”
First-party notebooks let me mix SQL and Python against warehouse data, with results and charts inlinefullproof ↗
“OpenSharing lets you share data and AI assets with users outside your organization, Databricks or not”
Share live datasets with another account or organization without copying data or building an export pipelinefullproof ↗
“Open Marketplace listings include datasets, AI models, notebooks, apps, and MCP servers”
A marketplace of third-party datasets lets me enrich my own data directly inside the platformfullproof ↗
“Discover and interact with shared assets across all workspaces from a single entry point”
A marketplace of third-party datasets lets me enrich my own data directly inside the platformfullproof ↗
“Vendors can publish offerings that customers discover, evaluate, and connect to directly from their workspace”
A marketplace of third-party datasets lets me enrich my own data directly inside the platformfullproof ↗
“Unity Catalog automatically enforces access control, tracks lineage, and logs activity for auditing across data and AI use”
I get audit logs of who ran what and column-level lineage of where data came fromfullproof ↗
“Standard connectors can be chosen by data source and pipeline customization level”
First-party and partner connectors cover my sources — SaaS apps, databases, and ETL/ELT tools — with documented setuppartialproof ↗
“Budgets can track account-wide spend or be filtered to specific teams, projects, or workspaces”
Budgets, resource monitors, or auto-suspend stop a runaway query or idle compute from burning money overnightpartialproof ↗
“Free trial account lets you start using Databricks without commitment”
Evaluate with a free tier or trial — real queries on real data without a credit card or a sales callfullproof ↗
“Runs directly on the data lake, supports ANSI SQL with Delta Lake extensions, enabling warehouses without moving data”
Query open table formats and files in object storage — Iceberg, Delta, Parquet — without first loading them into proprietary storagepartialproof ↗
“Runs directly on the data lake, supports ANSI SQL with Delta Lake extensions, enabling warehouses without moving data”
I get a full analytical SQL surface — window functions, CTEs, semi-structured JSON, arrays, and rich date/time types — without bolt-on extensionspartialproof ↗
“Real-time workloads can be processed with end-to-end latency as low as five milliseconds”
Streaming writes land queryable within seconds through a documented streaming ingestion APIfullproof ↗
“Time travel lets you audit operations, roll back a table, or query it at a past point in time”
Time-travel — query data as of a past point and restore dropped or corrupted tables from historyfullproof ↗
“Query execution plans can be inspected to identify bottlenecks and optimization opportunities”
Inspect query profiles and execution plans to find why a query is slow or expensivefullproof ↗
“Alerts can monitor query results, evaluate conditions, and deliver notifications automatically”
Define rules that trigger actions automatically on eventspartialproof ↗
“SOC 2 Type II report and broader compliance certification portfolio are documented in the Trust Center”
Compliance attestations (SOC 2, HIPAA, PCI) are documented so security review does not stall the rolloutfullproof ↗
Contradicted (2)
“Unity Catalog automatically enforces access control, tracks lineage, and logs activity for auditing across data and AI use”
Access control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safelydisputedproof ↗
“Create a table and grant privileges using the Unity Catalog governance model”
Access control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safelydisputedproof ↗
Undersold (12)
My agent can run governed SQL end to end — authenticate, discover schemas, query, and read results back through a CLI or API with no dashboard in the loopfullproof ↗
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Rely on versioned APIs with a documented deprecation policypartialproof ↗
Perform bulk operations across many items at oncefullproof ↗
Bulk-load CSV, JSON, and Parquet from cloud object storage with a single documented commandpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Claims outside our story set (6)
Real capability claims found in Databricks’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Create interactive AI/BI dashboards with AI-assisted authoring to share across the organization”
source ↗“Genie One mobile apps (iOS/Android) let users ask data questions, view dashboards, and run Databricks Apps”
source ↗“AI Playground lets you query and compare LLMs side-by-side, prototype a tool-calling agent, and export it to code”
source ↗“MLflow tracks training runs, hyperparameters and metrics, with Hyperopt for automated hyperparameter tuning”
source ↗“Lakebase provides serverless Postgres for applications that need to scale”
source ↗“Framework to build production-ready AI agents grounded in your own data”
source ↗
Business model
Consumption pricing in DBUs (Databricks Units) per-second, with rates varying by workload (SQL, jobs, DLT, serving) and cloud; a Free Edition and 14-day trial exist, enterprise commitments discount usage.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
