Skip to content

Rank #1 of 4 in Data Warehouses & Lakehouses

Databricks logo

Databricks, Inc. · commercial

pypi 18.9M/wkpypi/wk -668.7k

Access

Install

brewbrew tap databricks/tap && brew install databricks
curlcurl -fsSL https://raw.githubusercontent.com/databricks/setup-cli/main/install.sh | sh

Vendor-official, but review any script before piping it to a shell.

Compare head-to-head

Alternatives to Databricks

Showcase

Databricks homepage screenshot
homepage · captured Sep 2026 · view live ↗
Databricks docs screenshot
docs · captured Sep 2026 · view live ↗

Try itExperimental

See what an agent can do with Databricks before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$databricks aitools install --path /tmp/pa-dbx-skills # vendor-shipped agent skills, then list + print one SKILL.md frontmatterrecorded session — replayed, not live
recorded 2026-09-07 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agent analytics — stories about agent analytics in this arenaAgent analyticsevidence →

Stories about agent analytics in this arena

80.0/100

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

61.6/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

54.8/100

Cost economics — stories about cost economics in this arenaCost economicsevidence →

Stories about cost economics in this arena

36.7/100

Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrationsevidence →

The surrounding ecosystem — integrations, marketplaces, community packages

58.0/100

Governance access — stories about governance access in this arenaGovernance accessevidence →

Stories about governance access in this arena

46.0/100

Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelinesevidence →

Stories about ingestion pipelines in this arena

48.6/100

Notebooks workspace — stories about notebooks workspace in this arenaNotebooks workspaceevidence →

Stories about notebooks workspace in this arena

80.0/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

17.3/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Semantic layer — stories about semantic layer in this arenaSemantic layerevidence →

Stories about semantic layer in this arena

80.0/100

Sharing marketplace — stories about sharing marketplace in this arenaSharing marketplaceevidence →

Stories about sharing marketplace in this arena

86.7/100

Sql analytics — stories about sql analytics in this arenaSql analyticsevidence →

Stories about sql analytics in this arena

55.8/100

Streaming realtime — stories about streaming realtime in this arenaStreaming realtimeevidence →

Stories about streaming realtime in this arena

76.7/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 2 free · 1 paid · 0 enterprise · 40 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10T

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10C

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10C

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10T

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10C

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1noneuntestednone yet

My agent can run governed SQL end to end — authenticate, discover schemas, query, and read results back through a CLI or API with no dashboard in the loop C

Agent ops

ai-native userAgent analytics — stories about agent analytics in this arenaAgent analytics3full8/10T

Bulk-load CSV, JSON, and Parquet from cloud object storage with a single documented command C

Loading

data-engineerIngestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines3partial6/10C

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial6/10C

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partial6/10X

I get a full analytical SQL surface — window functions, CTEs, semi-structured JSON, arrays, and rich date/time types — without bolt-on extensions C

Sql

analystSql analytics — stories about sql analytics in this arenaSql analytics3partial5/10C

The pricing model is documented clearly enough that I can estimate a monthly bill for my workload before committing G

Pricing

platform-engineerCost economics — stories about cost economics in this arenaCost economics3partialpaid5/10X

Access control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safely C

Access

platform-engineerGovernance access — stories about governance access in this arenaGovernance access3disputed4/10D

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

Share live datasets with another account or organization without copying data or building an export pipeline C

Sharing

data-engineerSharing marketplace — stories about sharing marketplace in this arenaSharing marketplace2full9/10C

Time-travel — query data as of a past point and restore dropped or corrupted tables from history C

Recovery

data-engineerSql analytics — stories about sql analytics in this arenaSql analytics2full9/10C

A built-in AI assistant writes, fixes, and explains SQL against my schemas from natural language, inside the product C

Agent ops

ai-native userAgent analytics — stories about agent analytics in this arenaAgent analytics2full8/10C

A managed service continuously ingests new files or events as they arrive, without me running my own pipeline infrastructure C

Loading

data-engineerIngestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines2full8/10C

Dbt is a first-class citizen — a documented adapter or native dbt project support with vendor docs to match C

Transformation

data-engineerEcosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations2full8/10C

Define a governed semantic model — metrics, dimensions, and joins declared once — that queries and AI tools answer against consistently C

Semantics

analystSemantic layer — stories about semantic layer in this arenaSemantic layer2full8/10C

First-party notebooks let me mix SQL and Python against warehouse data, with results and charts inline C

Notebooks

analystNotebooks workspace — stories about notebooks workspace in this arenaNotebooks workspace2full8/10C

I get audit logs of who ran what and column-level lineage of where data came from G

Governance

platform-engineerGovernance access — stories about governance access in this arenaGovernance access2full8/10C

Inspect query profiles and execution plans to find why a query is slow or expensive C

Performance

data-engineerSql analytics — stories about sql analytics in this arenaSql analytics2full8/10C

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2full8/10T

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2full7/10T

Streaming writes land queryable within seconds through a documented streaming ingestion API C

Streaming

data-engineerStreaming realtime — stories about streaming realtime in this arenaStreaming realtime2full7/10C

First-party and partner connectors cover my sources — SaaS apps, databases, and ETL/ELT tools — with documented setup C

Connectors

data-engineerIngestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines2partial6/10C

Query open table formats and files in object storage — Iceberg, Delta, Parquet — without first loading them into proprietary storage C

Lakehouse

data-engineerSql analytics — stories about sql analytics in this arenaSql analytics2partial6/10C

Budgets, resource monitors, or auto-suspend stop a runaway query or idle compute from burning money overnight G

Pricing

platform-engineerCost economics — stories about cost economics in this arenaCost economics2partial5/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2disputed5/10D

I get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production compute C

Dev loop

data-engineerEcosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations2partialfree5/10X

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2n/auntestednone yet

Run continuous or incremental transformations — streams, tasks, declarative pipelines, or continuous queries — inside the platform C

Streaming

data-engineerStreaming realtime — stories about streaming realtime in this arenaStreaming realtime1full9/10C

A marketplace of third-party datasets lets me enrich my own data directly inside the platform C

Sharing

analystSharing marketplace — stories about sharing marketplace in this arenaSharing marketplace1full8/10C

Business users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joins C

Agent ops

ai-native userAgent analytics — stories about agent analytics in this arenaAgent analytics1full8/10C

Compliance attestations (SOC 2, HIPAA, PCI) are documented so security review does not stall the rollout C

Governance

platform-engineerGovernance access — stories about governance access in this arenaGovernance access1full8/10C

Evaluate with a free tier or trial — real queries on real data without a credit card or a sales call G

Trial

analystCost economics — stories about cost economics in this arenaCost economics1fullfree7/10C

Standard drivers (JDBC/ODBC) and documented BI-tool integrations connect my dashboards without custom glue C

Bi

analystEcosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations1full7/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial5/10C

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 21 stories with headroom

What would move Databricks’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Openness — open source, data portability, and self-hosting storiesSelf-host the core product

    nonemoves PA Scoreimpact 30

    Databricks is offered exclusively as a managed cloud service (AWS/Azure/GCP workspaces, pay-as-you-go billing) with no documented option to self-host the core platform on your own infrastructure; the closest analog (Free Edition) is still a hosted SaaS trial, not a self-hostable deployment.

  2. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    Missing: an explicit AI-training opt-out/data-use policy, documentation of contractual or technical controls preventing model training on customer data, and any independent confirmation of such a guarantee.

  3. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    No evidence in the pack describes any webhook subscription mechanism (e.g., event notifications pushed to external endpoints); Databricks docs reference REST APIs, SQL alerts, and MCP tools but nothing about webhooks for event subscription.

  4. Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)

    nonemoves API qualityimpact 30

    Databricks publishes an extensive human-readable REST API reference (databricks-docs-65, databricks-supp-versioned-apis, databricks-supp-api-reference-examples) but there is no evidence of a downloadable machine-readable spec — the probe explicitly found openapi.json/swagger.json/api/openapi.json/.well-known/openapi.json all return 404 (databricks-probe-2), and the API reference page confirms 'No live try-it runner is documented.' This is an applicable axis for a platform with a large REST API surface, so absent evidence of an OpenAPI/Swagger artifact this is 'none' rather than 'na'.

  5. Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency)

    nonemoves PA Scoreimpact 20

    No evidence in the pack addresses regional deployment, data residency controls, or workspace region selection for Databricks; the compliance mention only covers SOC 2 attestations, not data-location choice.

  6. Privacy posture — data-handling and privacy storiesControl data retention and deletion

    nonemoves PA Scoreimpact 20

    Missing: explicit retention configuration, right-to-delete/erasure mechanisms, data lifecycle/expiry policy documentation, and any independent verification of deletion behavior.

  7. Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking

    nonemoves PA Scoreimpact 20

    Missing: any mention of telemetry collection, opt-out settings, or usage-tracking controls.

  8. Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent

    partialq4/10moves agent-readyimpact 18

    Missing: explicit credential/token issuance workflow scoped to an agent identity, documentation of OAuth/service-principal scoping granularity, and any hands-on or independent confirmation that credentials can be narrowly scoped per-agent.

Showing the top 8 of 21 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 45 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Aws docs45 stories

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$databricks aitools install --path /tmp/pa-dbx-skills # vendor-shipped agent skills, then list + print one SKILL.md frontmatterreproduced
$ databricks aitools install --path /tmp/pa-dbx-skills  # vendor-shipped agent skills, then list + print one SKILL.md frontmatter
Using skills version 0.2.10
Wrote 29 skills to /tmp/pa-dbx-skills
--- installed skills ---
/tmp/pa-dbx-skills/databricks-agent-bricks/SKILL.md
/tmp/pa-dbx-skills/databricks-ai-functions/SKILL.md
/tmp/pa-dbx-skills/databricks-aibi-dashboards/SKILL.md
/tmp/pa-dbx-skills/databricks-app-design/SKILL.md
/tmp/pa-dbx-skills/databricks-apps-python/SKILL.md
/tmp/pa-dbx-skills/databricks-apps/SKILL.md
/tmp/pa-dbx-skills/databricks-core/SKILL.md
/tmp/pa-dbx-skills/databricks-dabs/SKILL.md
--- frontmatter of databricks-docs ---
---
name: databricks-docs
description: "Databricks documentation reference via llms.txt index. Use when other skills do not cover a topic, looking up unfamiliar Databricks features, or needing authoritative docs on APIs, configurations, or platform capabilities."
$databricks --version # installed via `brew tap databricks/tap && brew install databricks`reproduced
$ databricks --version  # installed via `brew tap databricks/tap && brew install databricks`
Databricks CLI v1.15.0
proves: Use an official CLIrecorded 2026-09-07

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

7 of 32 testable claims verified · 1 contradictedintegrity 16/100

47 distinct capability claims found in Databricks’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

7

Verified

24

Unverified

1

Contradicted

12

Undersold

Verified (10)
Unverified (32)
Contradicted (2)
Undersold (12)
Claims outside our story set (6)

Real capability claims found in Databricks’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Create interactive AI/BI dashboards with AI-assisted authoring to share across the organization

    source ↗
  • Genie One mobile apps (iOS/Android) let users ask data questions, view dashboards, and run Databricks Apps

    source ↗
  • AI Playground lets you query and compare LLMs side-by-side, prototype a tool-calling agent, and export it to code

    source ↗
  • MLflow tracks training runs, hyperparameters and metrics, with Hyperopt for automated hyperparameter tuning

    source ↗
  • Lakebase provides serverless Postgres for applications that need to scale

    source ↗
  • Framework to build production-ready AI agents grounded in your own data

    source ↗
Suggest a story for these →

Business model

usage-basedfree-tierenterprise-custom

Consumption pricing in DBUs (Databricks Units) per-second, with rates varying by workload (SQL, jobs, DLT, serving) and cloud; a Free Edition and 14-day trial exist, enterprise commitments discount usage.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score39 (Sep 7 '26)48 (Sep 10 '26)
Agent-ready55 (Sep 7 '26)70 (Sep 10 '26)

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)