Data Warehouses & Lakehouses Arena
Databricks vs BigQuery
Databricks
Databricks, Inc.
Databricks wins · 23–13 (16 drawn)
Agent analytics — stories about agent analytics in this arenaAgent analytics
Stories about agent analytics in this arena
Agent ops
ai-native userMy agent can run governed SQL end to end — authenticate, discover schemas, query, and read results back through a CLI or API with no dashboard in the loop
weight 3 · round drawnDatabricks documents a full non-UI path: CLI/API authentication (databricks-docs-2/48/64), schema and object discovery governed by Unity Catalog accessible via CLI/SQL/REST (databricks-docs-56, 19, 38, 70, 75), SQL execution via the Databricks SQL API/CLI (databricks-supp-api-versioning, databricks-docs-32), and reading results back programmatically (workspace REST API reference, databricks-docs-65), with runtime evidence that the CLI actually installs and exposes a raw `databricks api` wrapper keylessly (databricks-probe-rt-1). Missing for 10: an end-to-end documented/hands-on example chaining auth→discover→query→result specifically for an agent (no single walkthrough), and no independent corroboration of the full flow outside vendor docs.
- [claimed-docs] “The Databricks CLI (command-line interface) allows you to interact with the Databricks platform from your local terminal or automation scrip…”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Data and AI assets such as tables, views, volumes, functions, models, and services (model services and MCP services) follow a three-level na…”
- [claimed-docs] “It runs directly on your data lake, supports ANSI SQL with Delta Lake extensions, and provides the tools to build highly performant, cost-ef…”
- [claimed-docs] “Docs, "Update to the latest Databricks SQL API version": "The legacy API is deprecated and support will end soon. Use this page to migrate y…”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
BigQuery ships a documented, non-interactive bq CLI and a machine-readable REST API (verified live: keyless discovery doc exposes datasets/tables/routines/jobs resources) alongside IAM-based authentication/authorization, letting an agent authenticate, discover schemas, run SQL, and pull results with zero dashboard interaction. Google also ships an official MCP Toolbox with a prebuilt BigQuery toolset for agent access, confirmed running via npx (credential-gated as expected). Missing for 10: a first-party end-to-end hands-on trace showing an agent chaining auth to schema discovery to query to result-read purely via CLI/API without any human/dashboard step, and independent non-Google confirmation of this exact workflow.
- [claimed-docs] “vai aprender a usar o `bq`, a ferramenta de interface de linha de comandos (CLI) baseada em Python para o BigQuery, para criar um conjunto d…”
- [claimed-docs] “learn how to use bq, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query tables”
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
- [claimed-docs] “BigQuery: Roles and permissions that apply to BigQuery resources such as datasets, tables, views, and routines.”
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
- [probe] “PROBE runtime (recorded 2026-09-06): Google's official MCP Toolbox for Databases ran from npm — `npx -y @toolbox-sdk/server --version` print…”
- [probe] “official MCP server documented at https://github.com/googleapis/mcp-toolbox”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
ai-native userA built-in AI assistant writes, fixes, and explains SQL against my schemas from natural language, inside the product
weight 2 · round drawnDatabricks documents a built-in AI assistant (Genie/Genie Code) that writes, debugs, and explains code and queries directly inside notebooks/SQL editor against Unity Catalog tables, columns, and lineage, plus a separate Genie conversational analytics surface for asking natural-language data questions grounded in org data. This is native, in-product functionality (not a bolt-on), covering SQL generation, fixing errors, and natural-language explanation/answers over schemas. Missing for 10: independent/hands-on third-party verification of SQL-writing accuracy and no explicit example transcript showing it explaining SQL syntax step-by-step.
- [claimed-docs] “Genie Code is the AI coding and data assistant for developers and technical practitioners in the Databricks workspace.”
- [claimed-docs] “Write and run SQL queries with integrated AI assistance, code comments, and version history.”
- [claimed-docs] “Ask data questions in natural language and get answers grounded in your organization's data.”
- [claimed-docs] “It generates and runs code, builds pipelines and AI/BI dashboards, debugs errors, and works directly with Unity Catalog tables, columns, and…”
- [claimed-docs] “Chat with Genie Code, get inline suggestions, and run agentic tasks in your workspace.”
- [claimed-docs] “Run Genie Code as an autonomous agent that plans, runs code, fixes errors, and asks for approval before it uses tools.”
- [claimed-docs] “It generates and runs code, builds pipelines and AI/BI dashboards, debugs errors, and works directly with Unity Catalog tables, columns, and…”
- [claimed-docs] “From natural language dashboard creation to deep conversational analytics with Genie, this is BI built on AI from the start.”
Gemini in BigQuery provides an in-product natural-language assistant that generates/suggests SQL and Python code and explains existing queries, plus conversational analytics for chatting with data (docs-3,4,5,39,66,21,13). This matches the story of a built-in assistant writing, fixing, and explaining SQL from natural language directly inside the product. missing for 10: independent/hands-on user reports specifically validating the assistant's SQL-writing/fixing/explaining quality (only vendor docs cited).
- [claimed-docs] “BigQuery의 Gemini를 사용하여 SQL 또는 Python에서 코드를 생성하거나 제안하고 기존 SQL 쿼리를 설명할 수 있습니다.”
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “BigQuery 데이터 캔버스로 데이터 탐색, 변환, 쿼리, 시각화”
- [claimed-docs] “BigQuery의 Gemini로 자연어를 사용하여 테이블 애셋을 찾고 조인하고 쿼리하고 결과를 시각화하며 전체 프로세스에서 다른 사용자와 원활하게 공동작업할 수 있습니다.”
- [claimed-docs] “您还可以使用自然语言查询来开始数据分析。如需了解如何生成、补全和总结代码”
- [claimed-docs] “Conversational analytics now supports questions about market basket analysis.”
ai-native userBusiness users can ask questions in natural language and get governed, semantically-grounded answers rather than hallucinated joins
weight 1 · round to DatabricksGenie provides natural language Q&A grounded in organizational data (docs-8, docs-22, docs-57), backed by Unity Catalog governance enforcing access control and lineage (docs-19, docs-38) and metric views providing a semantic layer so metrics are defined once and computed consistently rather than via ad-hoc joins (docs-7, docs-16, docs-42, docs-72). Business-user-focused Genie One interface is explicitly documented for non-technical users (docs-57). missing for 10: independent/hands-on evidence of Genie's accuracy avoiding hallucinated joins in practice, and no community corroboration of semantic grounding quality specifically (community evidence pack is generic platform commentary, not about Genie/semantic layer).
- [claimed-docs] “Ask data questions in natural language and get answers grounded in your organization's data.”
- [claimed-docs] “From natural language dashboard creation to deep conversational analytics with Genie, this is BI built on AI from the start.”
- [claimed-docs] “Navigate the Genie One interface designed for business users. View dashboards, ask natural language data questions, and discover assets shar…”
- [claimed-docs] “you define the metric once, for example _sum of revenue divided by distinct customer count_, and users can group by any available field.”
- [claimed-docs] “you define the metric once, for example sum of revenue divided by distinct customer count, and users can group by any available field”
- [claimed-docs] “Define business metrics with consistent calculations using a semantic layer. Reuse metrics across queries and dashboards.”
- [claimed-docs] “you define the metric once, for example sum of revenue divided by distinct customer count, and users can group by any available field. The q…”
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
BigQuery's Gemini-powered conversational analytics and data canvas explicitly let business users ask natural-language questions to find, join, and query tables, with collaboration and visualization built in (docs-3,4,5,39,66,13/21), and IAM-based access control (docs-31,65) provides a governance layer over which data can be queried. However, there's no first-party evidence of a dedicated semantic/business-glossary layer preventing hallucinated joins, and no independent or hands-on corroboration that these NL answers are reliably 'governed' or free of hallucination. Missing for 10: evidence of an explicit semantic modeling/metadata layer enforcing join correctness, and independent validation of answer accuracy/hallucination rates.
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “BigQuery 데이터 캔버스로 데이터 탐색, 변환, 쿼리, 시각화”
- [claimed-docs] “BigQuery의 Gemini를 사용하여 SQL 또는 Python에서 코드를 생성하거나 제안하고 기존 SQL 쿼리를 설명할 수 있습니다.”
- [claimed-docs] “BigQuery의 Gemini로 자연어를 사용하여 테이블 애셋을 찾고 조인하고 쿼리하고 결과를 시각화하며 전체 프로세스에서 다른 사용자와 원활하게 공동작업할 수 있습니다.”
- [claimed-docs] “您还可以使用自然语言查询来开始数据分析。如需了解如何生成、补全和总结代码”
- [claimed-docs] “Conversational analytics now supports questions about market basket analysis.”
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
- [claimed-docs] “BigQuery: Roles and permissions that apply to BigQuery resources such as datasets, tables, views, and routines.”
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to DatabricksDatabricks hosts an llms.txt file (probe confirms HTTP 200 with a structured documentation index) and also ships agent-oriented skill docs (SKILL.md files installable via `databricks aitools install`) that agents can be pointed at, going beyond a bare llms.txt. missing for 10: no independent third-party report of an agent successfully consuming llms.txt or the skills in practice, and no evidence of an /llms-full.txt or deeper machine-readable agent doc index.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.databricks.com/llms.txt # Databricks Documentation > Comprehensive documentation for the Databricks…”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI SHIPS vendor agent skills — `databricks aitools install --path /tmp/pa-dbx-skills` k…”
- [claimed-docs] “Genie Code is the AI coding and data assistant for developers and technical practitioners in the Databricks workspace.”
ai-native userRun the product headlessly / in CI for automation
weight 2 · round drawnDatabricks provides a documented, keylessly-installable CLI supporting scripting/automation of workspace, jobs, and Lakeflow pipelines (databricks-docs-2, databricks-docs-48, databricks-probe-rt-1), plus REST APIs and Python SDK explicitly for automating operations (databricks-docs-49, databricks-docs-65), enabling headless/CI use for orchestrating pipelines, jobs, and ML training. Jobs/pipelines can be triggered/orchestrated without UI interaction, matching CI automation needs. Missing for 10: no explicit CI/CD pipeline integration guide (e.g., GitHub Actions) or independent case study confirming CI usage beyond docs and CLI probes.
- [claimed-docs] “The Databricks CLI (command-line interface) allows you to interact with the Databricks platform from your local terminal or automation scrip…”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
- [claimed-docs] “Develop and deploy your first ETL (extract, transform, and load) pipeline for data orchestration with Apache Spark™.”
BigQuery ships a documented Python-based CLI (`bq`) for scriptable dataset/table/query operations, a fully machine-readable REST API (verified live via discovery document), ODBC/JDBC drivers, and scheduled/batch load jobs — all standard building blocks for headless CI automation. missing for 10: explicit CI/CD pipeline examples (e.g. GitHub Actions/Cloud Build recipes) and independent hands-on reports of using bq/REST in automated pipelines.
- [claimed-docs] “vai aprender a usar o `bq`, a ferramenta de interface de linha de comandos (CLI) baseada em Python para o BigQuery, para criar um conjunto d…”
- [claimed-docs] “learn how to use bq, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query tables”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “The Simba Open Database Connectivity (ODBC) and Java Database Connectivity (JDBC) drivers for BigQuery connect your applications to BigQuery…”
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round to DatabricksDatabricks documents connecting agents to external, managed, and custom MCP servers (docs-45, docs-50, docs-71) and using their tools/third-party SaaS integrations (docs-31, docs-43), plus a Marketplace listing MCP servers for discovery (docs-34/69). This directly matches the story of plugging in MCP servers to gain tool access. Missing for 10: independent/hands-on community confirmation of the MCP-client integration working end-to-end (only vendor docs, no third-party validation).
- [claimed-docs] “Wire up Claude, Cursor, MCP Inspector, and other external clients to MCP servers hosted on Databricks.”
- [claimed-docs] “Discover, authenticate to, and call managed, MCP Service, and custom MCP servers from your agent code, then deploy the agent on Databricks A…”
- [claimed-docs] “Give your agent governed access to third-party and SaaS tools such as Slack, GitHub, Google Drive, Google Calendar, and Gmail — through buil…”
- [claimed-docs] “Create AI agent tools using Unity Catalog functions, including third-party integrations and code interpreter tools.”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers.”
- [probe] “official MCP server documented at https://docs.databricks.com/aws/en/agents/mcp-tools”
BigQuerynone0/10The evidence only shows Google's MCP Toolbox exposing BigQuery *as* an MCP server/tool provider for external AI agents (bigquery-probe-3, bigquery-probe-rt-2) — the opposite direction from this story, which asks whether a user can plug external MCP servers *into* BigQuery so BigQuery itself can consume their tools (e.g., within Gemini in BigQuery's conversational analytics). No evidence shows BigQuery or its Gemini features acting as an MCP client that ingests external tool servers.
- [probe] “official MCP server documented at https://github.com/googleapis/mcp-toolbox”
- [probe] “PROBE runtime (recorded 2026-09-06): Google's official MCP Toolbox for Databases ran from npm — `npx -y @toolbox-sdk/server --version` print…”
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “BigQuery의 Gemini로 자연어를 사용하여 테이블 애셋을 찾고 조인하고 쿼리하고 결과를 시각화하며 전체 프로세스에서 다른 사용자와 원활하게 공동작업할 수 있습니다.”
ai-native userConnect an agent via an official MCP server
weight 3 · round to DatabricksDatabricks documents official MCP server support extensively: it hosts managed/custom MCP servers as Databricks Apps, exposes Unity Catalog functions as MCP tools, provides built-in system.ai MCP Services for third-party SaaS tools, and explicitly supports wiring external clients (Claude, Cursor, MCP Inspector) to MCP servers hosted on Databricks; Marketplace also lists MCP servers as discoverable assets. Missing for 10: independent/hands-on third-party verification of the MCP connection flow (only first-party docs available).
- [claimed-docs] “Host a custom MCP server as a Databricks app to expose your own tools.”
- [claimed-docs] “Give your agent governed access to third-party and SaaS tools such as Slack, GitHub, Google Drive, Google Calendar, and Gmail”
- [claimed-docs] “Create AI agent tools using Unity Catalog functions, including third-party integrations and code interpreter tools.”
- [claimed-docs] “Wire up Claude, Cursor, MCP Inspector, and other external clients to MCP servers hosted on Databricks.”
- [claimed-docs] “Discover, authenticate to, and call managed, MCP Service, and custom MCP servers from your agent code, then deploy the agent on Databricks A…”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers. This gives customers a single catalog for f…”
- [claimed-docs] “Give your agent governed access to third-party and SaaS tools such as Slack, GitHub, Google Drive, Google Calendar, and Gmail — through buil…”
- [probe] “official MCP server documented at https://docs.databricks.com/aws/en/agents/mcp-tools”
Google's official MCP Toolbox for Databases (googleapis/mcp-toolbox) provides a documented, runnable MCP server with a prebuilt BigQuery toolset, confirmed to run via npx and require ADC credentials plus a project ID. However, it is a general 'Toolbox for Databases' rather than a BigQuery-specific first-party product page, and the runtime probe shows the handshake is credential-gated with no independent hands-on confirmation of full agent connectivity in production use. Missing for 10: independent/third-party corroboration of agent use, BigQuery-specific (not generic toolbox) branding, and evidence of production-scale reliability or broader ecosystem adoption.
ai-native userUse an official CLI
weight 2 · round to DatabricksDatabricks ships an official CLI (docs, migration guides, REST wrapper) confirmed by hands-on runtime probe (v1.15.0), and it is explicitly AI-native: the CLI ships an `aitools install` subcommand that installs vendor agent skills (SKILL.md files) for Claude Code, Codex, Cursor, Copilot, etc. missing for 10: independent/community corroboration beyond the vendor docs and runtime probe.
- [claimed-docs] “The Databricks CLI (command-line interface) allows you to interact with the Databricks platform from your local terminal or automation scrip…”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “To migrate from Databricks CLI version 0.18 or below to Databricks CLI version 0.205 or above, see Databricks CLI migration.”
- [probe] “official CLI documented at https://docs.databricks.com/aws/en/dev-tools/cli/”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI SHIPS vendor agent skills — `databricks aitools install --path /tmp/pa-dbx-skills` k…”
BigQuery ships 'bq', an official Python-based CLI for creating datasets, loading data, and querying tables, well documented across multiple doc pages and confirmed as a distinct probe finding. This is a general-purpose CLI rather than one purpose-built for AI-native agentic workflows (e.g., no agent-specific CLI features), so missing for 10: evidence of AI-agent-specific CLI features or tooling, independent hands-on community validation of CLI usage in agentic contexts.
- [claimed-docs] “vai aprender a usar o `bq`, a ferramenta de interface de linha de comandos (CLI) baseada em Python para o BigQuery, para criar um conjunto d…”
- [claimed-docs] “learn how to use bq, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query tables”
- [claimed-docs] “you learn how to use `bq`, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query …”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
ai-native userDrive the product through a documented public API
weight 3 · round to DatabricksDatabricks publishes a documented REST API reference (workspace-level APIs, versioned Jobs/SQL APIs with request/response examples), plus CLI, SDKs (Python), and Databricks Connect that all wrap this public API for automation and agentic use, and this is corroborated by a runtime probe confirming a working CLI with a raw `databricks api` REST wrapper. Missing for 10: a live interactive API try-it console or independent third-party API-quality corroboration beyond vendor docs.
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [claimed-docs] “API reference hub lists versioned REST APIs side by side, e.g. "Jobs v2.0 API — REST API reference for version 2.0 version of the Jobs REST …”
- [claimed-docs] “API reference (docs.databricks.com/api): "This reference describes the types, paths, and any request payload or query parameters, for each s…”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [claimed-docs] “Just like a JDBC driver, the Databricks Connect library can be embedded in any application to interact with Databricks.”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
- [probe] “official CLI documented at https://docs.databricks.com/aws/en/dev-tools/cli/”
BigQuery exposes a documented, machine-readable REST API (v2 discovery doc confirmed live with resources for datasets/jobs/models/tables), plus gRPC Storage Write API, client SDKs, ODBC/JDBC drivers, and a Python-based bq CLI, all publicly documented and independently verified via runtime probe. missing for 10: no llms.txt or standalone OpenAPI spec file was found (404s), and the MCP server integration requires credential-gated setup rather than being fully turnkey.
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
- [claimed-docs] “The Storage Write API (gRPC) has lower pricing and more robust features, including exactly-once delivery semantics.”
- [claimed-docs] “we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has lower pricing and …”
- [claimed-docs] “The Simba Open Database Connectivity (ODBC) and Java Database Connectivity (JDBC) drivers for BigQuery connect your applications to BigQuery…”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
- [probe] “PROBE llms.txt: HTTP 404 at https://cloud.google.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://cloud.google.com/openapi.json, https://cloud.google.com/swagger.json, https://cloud.google.c…”
- [probe] “PROBE runtime (recorded 2026-09-06): Google's official MCP Toolbox for Databases ran from npm — `npx -y @toolbox-sdk/server --version` print…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round to BigQueryDatabricks documents Unity Catalog governance that enforces access control on data/AI assets and lets agents get 'governed access' to third-party tools via Unity Catalog functions and MCP services (docs-19, 38, 70, 71, 43), implying a permissions model that could scope agent access. However, there is no explicit documentation of an API/credential-issuance mechanism (e.g., scoped service-principal tokens or OAuth scopes specifically for agents) described as 'least-privilege API credentials for an agent.' Missing for 10: explicit credential/token issuance workflow scoped to an agent identity, documentation of OAuth/service-principal scoping granularity, and any hands-on or independent confirmation that credentials can be narrowly scoped per-agent.
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Give your agent governed access to third-party and SaaS tools such as Slack, GitHub, Google Drive, Google Calendar, and Gmail — through buil…”
- [claimed-docs] “Create AI agent tools using Unity Catalog functions, including third-party integrations and code interpreter tools.”
- [claimed-docs] “Data and AI assets such as tables, views, volumes, functions, models, and services (model services and MCP services) follow a three-level na…”
BigQuery documents predefined IAM roles/permissions scoped to specific resources (datasets, tables, views, routines) and custom quotas, which are the building blocks for issuing least-privilege access to any caller including an agent, but there is no explicit documentation of a workflow for minting scoped/short-lived API credentials specifically for an AI agent (e.g., service account impersonation, workload identity federation, or OAuth scope restriction guidance tied to agentic use). Missing for 10: explicit agent-credential-issuance guidance, short-lived/ephemeral token support, workload identity federation docs, and IAM Conditions examples for fine-grained scoping.
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
- [claimed-docs] “BigQuery: Roles and permissions that apply to BigQuery resources such as datasets, tables, views, and routines.”
- [claimed-docs] “複数の BigQuery プロジェクトとユーザーが存在している場合は、カスタム割り当てを要求することで費用を管理できます。この割り当てでは、1 日に処理されるデータ量の上限を指定します。”
ai-native userBuild against official SDKs
weight 2 · round to DatabricksDatabricks publishes official SDKs (e.g., Databricks SDK for Python), a Databricks Connect library embeddable in any IDE, official CLI, and REST API reference with SDK-based examples, all first-party documented and independently verified via CLI runtime probes. This directly supports AI-native developers building against official SDKs/tools, further reinforced by agent-skill installs and MCP tool integration docs. missing for 10: no independent third-party benchmark of SDK reliability/coverage across languages beyond Python.
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [claimed-docs] “Databricks Connect enables developers to develop and debug their code on Databricks compute using any IDE's native running and debugging fun…”
- [claimed-docs] “Just like a JDBC driver, the Databricks Connect library can be embedded in any application to interact with Databricks.”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI SHIPS vendor agent skills — `databricks aitools install --path /tmp/pa-dbx-skills` k…”
- [claimed-docs] “API reference hub lists versioned REST APIs side by side, e.g. "Jobs v2.0 API — REST API reference for version 2.0 version of the Jobs REST …”
BigQuery documents multiple official SDK/API surfaces for building integrations: the Storage Write API (gRPC) client, a Rust SDK now in Preview, ODBC/JDBC drivers, the Python-based bq CLI, and a machine-readable REST v2 discovery document verified reachable keylessly at runtime. This gives an AI-native builder concrete, documented, and runtime-verified official surfaces to build against. Missing for 10: explicit listing/mention of mainstream language client libraries (Python, Java, Node.js, Go, .NET) by name and any independent hands-on developer report confirming SDK ergonomics beyond docs.
- [claimed-docs] “The Rust SDK for BigQuery is now in Preview.”
- [claimed-docs] “The Rust SDK for BigQuery is now in Preview”
- [claimed-docs] “we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has lower pricing and …”
- [claimed-docs] “The Simba Open Database Connectivity (ODBC) and Java Database Connectivity (JDBC) drivers for BigQuery connect your applications to BigQuery…”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
ai-native userSubscribe to events via webhooks
weight 2 · round drawnDatabricksnone0/10No evidence in the pack describes any webhook subscription mechanism (e.g., event notifications pushed to external endpoints); Databricks docs reference REST APIs, SQL alerts, and MCP tools but nothing about webhooks for event subscription.
BigQuerynone0/10The evidence pack shows BigQuery streaming ingestion via Pub/Sub subscriptions and continuous queries for real-time analysis of incoming data, but nothing documents an outbound webhook mechanism for subscribing external clients to BigQuery events (e.g., job completion, table changes) via HTTP callbacks. Pub/Sub subscriptions described here are for loading data in, not for AI-native agents subscribing to events out.
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub. Pub/Sub can handle high throughput of data loads into BigQuery…”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub.”
- [claimed-docs] “Pub/Sub can handle high throughput of data loads into BigQuery. It supports real-time data streaming, loading data as it's generated.”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round drawnDatabricks ships multiple first-party AI-insight surfaces: Genie for natural-language Q&A grounded in org data, AI/BI dashboards with AI-assisted authoring, automatic query performance insights/recommendations, and a semantic layer (metric views) that standardizes metrics for consistent AI-driven analysis. This directly satisfies the story of getting AI-generated insights/suggestions inside the product. Missing for 10: independent/hands-on user corroboration specifically validating insight quality (community evidence is generic platform sentiment, not about Genie/AI-BI insight accuracy), and no benchmark of suggestion usefulness.
- [claimed-docs] “Ask data questions in natural language and get answers grounded in your organization's data.”
- [claimed-docs] “Create interactive AI/BI dashboards with AI-assisted authoring to share insights across your organization.”
- [claimed-docs] “Get automatic insights and recommendations when queries run inefficiently.”
- [claimed-docs] “From natural language dashboard creation to deep conversational analytics with Genie, this is BI built on AI from the start.”
- [claimed-docs] “Define business metrics with consistent calculations using a semantic layer. Reuse metrics across queries and dashboards.”
- [claimed-docs] “you define the metric once, for example sum of revenue divided by distinct customer count, and users can group by any available field. The q…”
BigQuery's Gemini integration provides conversational analytics (natural language Q&A over data), data canvas for exploration, SQL/Python code generation and query explanation, plus AI functions for summarization/sentiment/enrichment and BigQuery ML predictive insights — directly matching the story of in-product AI-generated insights and suggestions. missing for 10: independent/hands-on community validation of Gemini-in-BigQuery's actual output quality (only first-party docs cited) and more detail on proactive 'suggestions' beyond query generation.
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “BigQuery 데이터 캔버스로 데이터 탐색, 변환, 쿼리, 시각화”
- [claimed-docs] “BigQuery의 Gemini를 사용하여 SQL 또는 Python에서 코드를 생성하거나 제안하고 기존 SQL 쿼리를 설명할 수 있습니다.”
- [claimed-docs] “Conversational analytics now supports questions about market basket analysis.”
- [claimed-docs] “Train, evaluate, and deploy predictive analytics models directly within BigQuery using SQL.”
- [claimed-docs] “Build sophisticated context-retrieval and RAG applications with embeddings and vector, text, or hybrid search to find information based on m…”
- [claimed-docs] “BigQuery의 Gemini로 자연어를 사용하여 테이블 애셋을 찾고 조인하고 쿼리하고 결과를 시각화하며 전체 프로세스에서 다른 사용자와 원활하게 공동작업할 수 있습니다.”
- [claimed-docs] “Use generative AI in your workflows with AI functions for text summarization, sentiment analysis, and data enrichment.”
- [claimed-docs] “您还可以使用自然语言查询来开始数据分析。如需了解如何生成、补全和总结代码”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to DatabricksDatabricks supports autonomous background automation via Lakeflow pipelines/jobs scheduling, Genie Code running as an autonomous agent that plans, runs code, fixes errors and asks approval before tool use, and CLI/SDK scripting for automation, plus MCP-based agent tooling for background agent workflows. missing for 10: independent/hands-on verification of fully unattended (no-human-in-loop) agent runs, and clearer documentation of scheduling/triggers specifically for autonomous agent tasks rather than just pipelines.
- [claimed-docs] “Run Genie Code as an autonomous agent that plans, runs code, fixes errors, and asks for approval before it uses tools.”
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
- [claimed-docs] “The Databricks CLI (command-line interface) allows you to interact with the Databricks platform from your local terminal or automation scrip…”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [claimed-docs] “Create AI agent tools using Unity Catalog functions, including third-party integrations and code interpreter tools.”
- [claimed-docs] “Discover, authenticate to, and call managed, MCP Service, and custom MCP servers from your agent code, then deploy the agent on Databricks A…”
BigQuery offers background-automation primitives — continuous queries that run indefinitely to process streaming data in real time, scheduled load jobs, and Dataform for scheduling and orchestrating multi-step transformation workflows — which can run autonomously without a user actively invoking them. However, none of this is framed or documented as 'AI-native' agentic automation (e.g., AI-triggered actions, agent orchestration, or autonomous decision loops using BigQuery's AI/ML functions), and there is no independent/community corroboration of these automation features being used this way. Missing for 10: AI-agent-specific automation/orchestration evidence, independent hands-on validation of continuous queries/Dataform running unattended long-term, and any framing tying automation to autonomous AI workflows rather than plain data-pipeline scheduling.
- [claimed-docs] “BigQuery の継続的クエリは、継続的に実行される SQL ステートメントです。継続的クエリを使用すると、BigQuery で受信データをリアルタイムで分析できます。”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “Dataform is a service for data analysts to develop, test, control versions, and schedule complex workflows for data transformation in BigQue…”
- [claimed-docs] “Dataform lets you manage data transformation in the Extraction, Loading, and Transformation (ELT) process for data integration.”
- [claimed-docs] “View a visualization of the dependency tree of your workflow.”
- [claimed-docs] “Collaborate with team members on workflow development through Git.”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to DatabricksDatabricks ships Genie Code, a built-in AI assistant that generates/runs code, builds pipelines and dashboards, debugs errors, and can run as an autonomous agent that plans, executes code, fixes errors and asks approval before tool use — directly matching 'delegate tasks to a built-in AI assistant.' It also supports parallel chats and can be given governed tool access (Slack, GitHub, etc.) via MCP for agentic task execution. missing for 10: independent/hands-on user validation of agentic delegation quality is absent from community evidence.
- [claimed-docs] “It generates and runs code, builds pipelines and AI/BI dashboards, debugs errors, and works directly with Unity Catalog tables, columns, and…”
- [claimed-docs] “Chat with Genie Code, get inline suggestions, and run agentic tasks in your workspace.”
- [claimed-docs] “Run Genie Code as an autonomous agent that plans, runs code, fixes errors, and asks for approval before it uses tools.”
- [claimed-docs] “It generates and runs code, builds pipelines and AI/BI dashboards, debugs errors, and works directly with Unity Catalog tables, columns, and…”
- [claimed-docs] “Start work directly from Genie Code rather than from an asset like a notebook, run multiple chats in parallel, and persona”
- [claimed-docs] “Give your agent governed access to third-party and SaaS tools such as Slack, GitHub, Google Drive, Google Calendar, and Gmail”
BigQuery ships Gemini in BigQuery, a built-in AI assistant offering conversational analytics (natural-language Q&A over data), a data canvas for exploring/transforming/querying/visualizing data, and SQL/Python code generation, explanation, and suggestion — directly matching the story of delegating tasks to an in-product AI assistant. Release notes also show ongoing expansion (e.g., market basket analysis support in conversational analytics). missing for 10: independent/hands-on user reports validating Gemini in BigQuery's assistant behavior (only first-party docs cited), and no evidence of broader agentic task delegation (e.g., multi-step autonomous execution) beyond query/code assistance.
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “BigQuery 데이터 캔버스로 데이터 탐색, 변환, 쿼리, 시각화”
- [claimed-docs] “BigQuery의 Gemini를 사용하여 SQL 또는 Python에서 코드를 생성하거나 제안하고 기존 SQL 쿼리를 설명할 수 있습니다.”
- [claimed-docs] “Conversational analytics now supports questions about market basket analysis.”
- [claimed-docs] “BigQuery의 Gemini로 자연어를 사용하여 테이블 애셋을 찾고 조인하고 쿼리하고 결과를 시각화하며 전체 프로세스에서 다른 사용자와 원활하게 공동작업할 수 있습니다.”
- [claimed-docs] “您还可以使用自然语言查询来开始数据分析。如需了解如何生成、补全和总结代码”
ai-native userOperate the product with natural-language commands
weight 2 · round to DatabricksDatabricks documents multiple natural-language interaction surfaces: Genie for asking data questions in plain English (docs-8), Genie Code as an autonomous coding/data agent that plans, runs code, fixes errors, and builds pipelines/dashboards from NL prompts (docs-15, docs-28, docs-29, docs-30), and NL-driven dashboard authoring (docs-22). This covers the core of 'operate via natural language' across both business and technical personas. Missing for 10: independent/hands-on validation of NL command reliability and no evidence of NL control over broader ops (e.g., cluster/job management) beyond data/coding tasks.
- [claimed-docs] “Ask data questions in natural language and get answers grounded in your organization's data.”
- [claimed-docs] “It generates and runs code, builds pipelines and AI/BI dashboards, debugs errors, and works directly with Unity Catalog tables, columns, and…”
- [claimed-docs] “From natural language dashboard creation to deep conversational analytics with Genie, this is BI built on AI from the start.”
- [claimed-docs] “Chat with Genie Code, get inline suggestions, and run agentic tasks in your workspace.”
- [claimed-docs] “Run Genie Code as an autonomous agent that plans, runs code, fixes errors, and asks for approval before it uses tools.”
- [claimed-docs] “It generates and runs code, builds pipelines and AI/BI dashboards, debugs errors, and works directly with Unity Catalog tables, columns, and…”
BigQuery's Gemini/Conversational Analytics features let users query, join, and visualize data using natural language, and generate/explain SQL or Python code from NL prompts, which directly supports natural-language operation of the product. However, this is all first-party doc-based evidence with no independent/hands-on corroboration, and NL support seems scoped to querying/exploration rather than full operational control (e.g., admin, pipeline management). Missing for 10: independent verification of NL feature reliability, evidence of NL commands controlling broader BigQuery operations beyond analytics/query generation.
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “BigQuery 데이터 캔버스로 데이터 탐색, 변환, 쿼리, 시각화”
- [claimed-docs] “BigQuery의 Gemini를 사용하여 SQL 또는 Python에서 코드를 생성하거나 제안하고 기존 SQL 쿼리를 설명할 수 있습니다.”
- [claimed-docs] “Conversational analytics now supports questions about market basket analysis.”
- [claimed-docs] “BigQuery의 Gemini로 자연어를 사용하여 테이블 애셋을 찾고 조인하고 쿼리하고 결과를 시각화하며 전체 프로세스에서 다른 사용자와 원활하게 공동작업할 수 있습니다.”
- [claimed-docs] “您还可以使用自然语言查询来开始数据分析。如需了解如何生成、补全和总结代码”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round to DatabricksDatabricks publishes a detailed REST API reference with request/response payload examples and code samples via CLI/Terraform/SDKs (databricks-docs-65, databricks-supp-api-reference-examples), but the same doc explicitly states 'No live try-it runner is documented,' and probes found no OpenAPI/Swagger endpoint to power an interactive explorer (databricks-probe-2). This gives static, copyable examples rather than a true in-browser runnable API reference. Missing for 10: an in-page 'try it now' execution console, OpenAPI-based interactive explorer, and confirmation of live request execution against a user's workspace from the docs site.
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [claimed-docs] “API reference (docs.databricks.com/api): "This reference describes the types, paths, and any request payload or query parameters, for each s…”
- [claimed-docs] “API reference hub lists versioned REST APIs side by side, e.g. "Jobs v2.0 API — REST API reference for version 2.0 version of the Jobs REST …”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.databricks.com/openapi.json, https://docs.databricks.com/swagger.json, https://docs.dat…”
BigQuerynone0/10The evidence shows only a machine-readable REST discovery document (bigquery-probe-rt-1) and standard docs pages, with explicit probe failures for llms.txt and openapi.json (bigquery-probe-1, bigquery-probe-2). There is no evidence of an interactive, runnable-example API reference (e.g., a try-it console or embedded runnable code snippets) for AI-native exploration.
- [probe] “PROBE llms.txt: HTTP 404 at https://cloud.google.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://cloud.google.com/openapi.json, https://cloud.google.com/swagger.json, https://cloud.google.c…”
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round to BigQueryDatabricksnone0/10Databricks publishes an extensive human-readable REST API reference (databricks-docs-65, databricks-supp-versioned-apis, databricks-supp-api-reference-examples) but there is no evidence of a downloadable machine-readable spec — the probe explicitly found openapi.json/swagger.json/api/openapi.json/.well-known/openapi.json all return 404 (databricks-probe-2), and the API reference page confirms 'No live try-it runner is documented.' This is an applicable axis for a platform with a large REST API surface, so absent evidence of an OpenAPI/Swagger artifact this is 'none' rather than 'na'.
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.databricks.com/openapi.json, https://docs.databricks.com/swagger.json, https://docs.dat…”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [claimed-docs] “API reference (docs.databricks.com/api): "This reference describes the types, paths, and any request payload or query parameters, for each s…”
BigQuery does not host a standard OpenAPI/swagger.json file (probe confirms openapi.json/swagger.json paths 404), but it exposes Google's own machine-readable API Discovery Document (bigquery:v2) at the REST reference which was verified to download keylessly and parse cleanly into resources/methods — functionally equivalent to an OpenAPI spec though in a different format. Missing for 10: a canonical OpenAPI/Swagger file, first-party statement equating the Discovery doc to OpenAPI, and independent community confirmation of AI-native tooling consuming it.
- [probe] “PROBE openapi: all candidate paths 404 (https://cloud.google.com/openapi.json, https://cloud.google.com/swagger.json, https://cloud.google.c…”
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
ai-native userTest against a sandbox environment without touching production data
weight 1 · round to BigQueryDatabricksnone0/10Evidence shows Unity Catalog governance, Free Edition for learning, and an AI Playground for prototyping agents, but nothing documents a dedicated sandbox/test environment that isolates AI-native testing from production data (e.g., dev/test catalog cloning, data masking for test runs, or a documented sandbox mode). Missing for 10: explicit sandbox/test-environment feature, data isolation guarantees for testing, and any customer/community confirmation of safe non-production testing workflows.
BigQuery Sandbox is explicitly documented as a way to explore and test BigQuery capabilities without a billing account or credit card, functioning as an isolated environment separate from a paid production setup. This directly matches the story of testing without touching production data, though evidence is purely first-party docs. Missing for 10: independent/hands-on confirmation of sandbox isolation guarantees, and specifics on how sandbox environments map to keeping AI-agent test workloads fully separate from production datasets.
- [claimed-docs] “Mit der BigQuery-Sandbox können Sie BigQuery nutzen, ohne eine Kreditkarte anzugeben oder ein Rechnungskonto für Ihr Projekt zu erstellen.”
- [claimed-docs] “The BigQuery sandbox lets you explore limited BigQuery capabilities at no cost to confirm whether BigQuery fits your needs.”
- [claimed-docs] “The BigQuery sandbox lets you experience BigQuery without providing a credit card or creating a billing account for your project.”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round to DatabricksDatabricks documents versioned REST APIs (e.g., Jobs v2.0) with explicit recommendations to use the latest version, and shows a concrete deprecation/migration case (legacy SQL API path deprecated in favor of api/2.0/sql/queries with a migration guide). This shows real-world versioning and deprecation practice, but there's no single consolidated deprecation policy document, no stated support timelines/sunset dates, and no explicit versioning scheme description (e.g., semver, LTS windows) across the whole API surface. missing for 10: a unified deprecation-policy page with timelines/sunset commitments, explicit API versioning scheme documentation, and independent confirmation that deprecations are reliably telegraphed in advance across all APIs.
- [claimed-docs] “Docs, "Update to the latest Databricks SQL API version": "The legacy API is deprecated and support will end soon. Use this page to migrate y…”
- [claimed-docs] “API reference hub lists versioned REST APIs side by side, e.g. "Jobs v2.0 API — REST API reference for version 2.0 version of the Jobs REST …”
- [claimed-docs] “API reference (docs.databricks.com/api): "This reference describes the types, paths, and any request payload or query parameters, for each s…”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
BigQuery exposes a versioned, machine-readable REST API (v2 discovery document) confirmed via runtime probe, showing a genuine versioned API surface, but there is no evidence in the pack of a documented deprecation policy, versioning cadence, or sunset guarantees for that API. missing for 10: explicit deprecation/sunset policy docs, versioning changelog commitments, and independent confirmation of long-term API stability guarantees.
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
- [claimed-docs] “The Rust SDK for BigQuery is now in Preview.”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round drawnDatabricks exposes bulk/automation capabilities via CLI, SDKs, REST APIs, and Lakeflow pipelines that let AI-native users script operations across many jobs, tables, or workspace objects at once (databricks-docs-2, databricks-docs-48, databricks-docs-49, databricks-docs-56, databricks-docs-65), and Unity Catalog governs bulk operations across catalogs/schemas with consistent access control. Runtime probes confirm the CLI actually works keylessly and covers workspace/compute/jobs/pipelines plus a raw REST wrapper (databricks-probe-rt-1). missing for 10: explicit documentation or example of a single bulk-operation command (e.g., batch update/delete across many items in one call) rather than iterating via scripts, and independent/hands-on evidence of bulk-operation reliability at scale.
- [claimed-docs] “The Databricks CLI (command-line interface) allows you to interact with the Databricks platform from your local terminal or automation scrip…”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
BigQuery natively supports bulk operations via DML for batch insert/update/delete, batch load jobs with scheduling, and the high-throughput Storage Write API for large-scale bulk ingestion, plus the bq CLI for scripting these operations across many items/tables. Community evidence corroborates BigQuery's ability to process petabyte-scale data efficiently. Missing for 10: explicit evidence of AI-agent-driven orchestration of bulk operations (e.g., an agent issuing many bulk actions via MCP/API in one flow) and independent benchmarks specifically on DML/bulk-write scale rather than just query scale.
- [claimed-docs] “Data Manipulation Language (DML) statements enable you to update, insert, and delete data from your BigQuery tables.”
- [claimed-docs] “we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has lower pricing and …”
- [claimed-docs] “For new projects, we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has …”
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “learn how to use bq, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query tables”
- [community] “I've worked with much larger datasets on BQ (petabyte scale) and managed to not spend more than $1000 in an hour; BQ tells you how much data…”
- [claimed-docs] “the distributed and scalable analysis engine of BigQuery allows querying terabytes in seconds and petabytes in minutes”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to DatabricksDatabricks supports rule-like automation via SQL Alerts (monitor query results, evaluate conditions, deliver notifications automatically) and event-driven Lakeflow/Auto Loader pipelines that trigger on new file arrival, plus job triggers and the CLI/SDK for scripting automated actions. However, there is no dedicated general-purpose 'if event then action' rules engine documented beyond these specific mechanisms (alerts, streaming triggers, job schedules). Missing for 10: a unified event-rule/trigger API spanning arbitrary events, independent hands-on validation of alert/trigger reliability, and more detail on custom action types beyond notifications.
- [claimed-docs] “Monitor query results, evaluate conditions, and deliver notifications automatically.”
- [claimed-docs] “Incrementally and efficiently process new data files as they arrive in cloud storage.”
- [claimed-docs] “Process data for real-time workloads with end-to-end latency as low as five milliseconds.”
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
BigQuery's continuous queries feature runs SQL statements continuously to analyze incoming data in real time, which is the closest capability to event-triggered automation, but there is no evidence of a general rules/condition-action engine that fires arbitrary actions (notifications, function calls, workflows) on events—continuous queries are limited to SQL-based streaming analysis. Missing for 10: explicit rule definition UI/API, broader action types (not just SQL output), and evidence of integration with external triggers/alerts as an automation mechanism.
- [claimed-docs] “BigQuery の継続的クエリは、継続的に実行される SQL ステートメントです。継続的クエリを使用すると、BigQuery で受信データをリアルタイムで分析できます。”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to DatabricksDatabricks Lakeflow Jobs (referenced via docs-1, docs-40, docs-4/45/50 MCP/agent orchestration context, and CLI docs-2/48/rt-1) natively supports scheduling recurring jobs/workflows, and the CLI/SDK (docs-49) allow programmatic/automated job management for AI-native workflows. Missing for 10: no explicit docs excerpt detailing cron/schedule syntax or trigger configuration, and no independent/hands-on confirmation of scheduling reliability beyond vendor docs.
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
- [claimed-docs] “Develop and deploy your first ETL (extract, transform, and load) pipeline for data orchestration with Apache Spark™.”
- [claimed-docs] “The Databricks CLI (command-line interface) allows you to interact with the Databricks platform from your local terminal or automation scrip…”
- [claimed-docs] “allows you to interact with the Databricks platform from your local terminal or automation scripts”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
BigQuery documents native support for recurring automation: scheduled load jobs ('schedule one-time or batch data transfers at regular intervals') and Dataform for developing, testing, versioning, and scheduling complex data-transformation workflows, plus continuous queries for always-on real-time processing. These are first-party, concrete capabilities directly matching the story. Missing for 10: explicit mention of the dedicated BigQuery Scheduled Queries feature by name, and independent/hands-on confirmation of scheduling reliability in production.
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “Dataform is a service for data analysts to develop, test, control versions, and schedule complex workflows for data transformation in BigQue…”
- [claimed-docs] “Dataform lets you manage data transformation in the Extraction, Loading, and Transformation (ELT) process for data integration.”
- [claimed-docs] “BigQuery の継続的クエリは、継続的に実行される SQL ステートメントです。継続的クエリを使用すると、BigQuery で受信データをリアルタイムで分析できます。”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
- [claimed-docs] “View a visualization of the dependency tree of your workflow.”
- [claimed-docs] “Collaborate with team members on workflow development through Git.”
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnDatabricks documents automatic versioning for notebooks, Delta table time-travel/rollback, and Unity Catalog lineage/audit logging, which together cover generic version/review/rollback capabilities, but none of this evidence explicitly ties into versioning or rolling back Lakeflow pipelines/Jobs (the actual 'automations') rather than notebooks/tables. Missing for 10: explicit job/pipeline version history and rollback workflow, CI/CD or git-based pipeline versioning evidence, and any hands-on confirmation of rolling back an automation run.
- [claimed-docs] “Databricks notebooks provide real-time coauthoring in multiple languages, automatic versioning, and built-in data visualizations”
- [claimed-docs] “Databricks notebooks provide real-time coauthoring in multiple languages, automatic versioning, and built-in data visualizations for develop…”
- [claimed-docs] “Use history information to audit operations, roll back a table, or query a table at a specific point in time using time travel.”
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
Dataform (BigQuery's workflow/transformation service) explicitly supports version control and Git-based collaboration for data pipelines, plus Git folders for scheduled pipelines, which cover versioning and team review to a degree. However, there's no explicit documentation of a rollback mechanism or formal review/approval workflow for automations beyond generic Git usage. Missing for 10: explicit rollback/revert documentation, structured review/approval process (e.g., PR-based approvals) for automations, and independent/hands-on evidence confirming these workflows in practice.
- [claimed-docs] “Dataform is a service for data analysts to develop, test, control versions, and schedule complex workflows for data transformation in BigQue…”
- [claimed-docs] “View a visualization of the dependency tree of your workflow.”
- [claimed-docs] “Collaborate with team members on workflow development through Git.”
- [claimed-docs] “You can now create, store, and manage pipelines in Git folders.”
- [claimed-docs] “Dataform lets you manage data transformation in the Extraction, Loading, and Transformation (ELT) process for data integration.”
Cost economics — stories about cost economics in this arenaCost economics
Stories about cost economics in this arena
Pricing
platform-engineerThe pricing model is documented clearly enough that I can estimate a monthly bill for my workload before committing
weight 3 · round to DatabricksDatabricks documents a pay-as-you-go, per-second-billing model (docs-21) and provides account/workspace-level budget tracking tools (docs-58, docs-61) that help monitor spend, but there is no evidence of a pricing calculator, DBU rate tables, or concrete per-workload cost examples that would let a platform engineer pre-estimate a monthly bill before committing. Community comments (databricks-comm-2, databricks-comm-4) describing the platform as 'crazy expensive' with 'surprise gotchas' around serverless pricing further suggest cost estimation in practice is harder than the high-level docs imply, though this falls short of a concrete documented failure of the pricing docs themselves. Missing for 10: a pricing/cost calculator, itemized DBU/unit rate tables, example workload cost breakdowns, and independent confirmation that pre-commitment estimates match actual bills.
- [claimed-docs] “Databricks offers you a pay-as-you-go approach with no up-front costs. Only pay for the products you use at per second granularity.”
- [claimed-docs] “You can set up budgets to either track account-wide spending, or apply filters to track the spending of specific teams, projects, or workspa…”
- [claimed-docs] “Budgets enable you to monitor usage across your account. You can set up budgets to either track account-wide spending, or apply filters to t…”
- [community] “Coming from hadoop, databricks is utopia. It's stable, fast, scales really well if you have massive datasets. The biggest gripe I have is ho…”
- [community] “They push Serverless so hard but there are SO MANY limitations and surprise gotchas. It's driving me absolutely insane.”
BigQuerydisputedcontradicted5/10Google documents pricing clearly (on-demand per-TiB, first 1TiB free, slot-based flat-rate reservations, custom quotas to cap spend) and community reports confirm a dry-run/bytes-scanned estimator exists that lets users predict costs before running queries. However, a concrete hands-on report describes a user being billed $14,000 with zero warning, and community critique that BigQuery 'hides query cost behind abstracted TBs scanned or slots' making bill estimation harder than it should be — a direct contradiction to the claim that pricing is 'documented clearly enough to estimate a bill' in practice. missing for 10: independent case studies showing accurate monthly forecasting at scale, clearer treatment of storage/streaming/BI Engine cost interactions, and resolution of the surprise-billing complaint.
- [claimed-docs] “On-demand pricing (per TiB). With this pricing model, you are charged for the number of bytes processed by each query. The first 1 TiB of qu…”
- [claimed-docs] “BigQuery offers a choice of two compute pricing models for running queries”
- [claimed-docs] “You can also reserve compute capacity ahead of time in the form of slots, which represent virtual CPUs.”
- [claimed-docs] “複数の BigQuery プロジェクトとユーザーが存在している場合は、カスタム割り当てを要求することで費用を管理できます。この割り当てでは、1 日に処理されるデータ量の上限を指定します。”
- [community] “User ran a script on BigQuery for HTTP Archive data and was billed $14,000 with zero warning; complained about lack of customer support and …”
- [community] “BQ hides query cost behind abstracted 'TBs scanned' or 'slots' mechanism; if GCP returned query cost directly in API/console it would be muc…”
- [community] “BigQuery provides a dry run option to estimate bytes/costs before running a query, and shows bytes-to-be-scanned in small text before you hi…”
- [community] “I've worked with much larger datasets on BQ (petabyte scale) and managed to not spend more than $1000 in an hour; BQ tells you how much data…”
- [community] “BigQuery announced pricing changes: annual flat rate going from 2.3c to 4.8c per slot hour, and on-demand pricing increasing 25% (from $5/TB…”
- [community] “BigQuery has on-demand pricing metered by data read, plus reserved slot pricing metered by time; reserved slots offer a considerable discoun…”
platform-engineerBudgets, resource monitors, or auto-suspend stop a runaway query or idle compute from burning money overnight
weight 2 · round to DatabricksDatabricks docs show account-level budgets for tracking spend by team/project/workspace and query-level alerting when conditions are met (databricks-docs-58, -61, -26), which partially covers cost monitoring, but the evidence never documents an automatic enforcement mechanism (e.g., warehouse auto-stop for idle compute or query kill/timeout) that actually halts a runaway query or idle cluster overnight — budgets are described as tracking/notifying, not stopping. missing for 10: explicit auto-suspend/auto-stop behavior for idle compute, automatic termination of runaway queries, and independent confirmation that budgets can enforce hard spend caps rather than just alert.
- [claimed-docs] “You can set up budgets to either track account-wide spending, or apply filters to track the spending of specific teams, projects, or workspa…”
- [claimed-docs] “Budgets enable you to monitor usage across your account. You can set up budgets to either track account-wide spending, or apply filters to t…”
- [claimed-docs] “Monitor query results, evaluate conditions, and deliver notifications automatically.”
BigQuerydisputedcontradicted4/10BigQuery docs describe custom quotas that cap daily bytes processed per project/user as a cost-control mechanism, and community discussion confirms dry-run/bytes-scanned previews exist to estimate cost before running a query — but a documented real-world case shows a runaway query still resulted in a surprise $14,000 bill with 'zero warning' and Google refusing to waive it, meaning no automatic budget/resource-monitor/auto-suspend actually intervened. Missing for 10: evidence of real-time Cloud Billing budget alerts wired to auto-suspend BigQuery usage, a per-query 'maximum bytes billed' kill-switch, or any documented fix/response after the community-reported billing incident.
- [claimed-docs] “複数の BigQuery プロジェクトとユーザーが存在している場合は、カスタム割り当てを要求することで費用を管理できます。この割り当てでは、1 日に処理されるデータ量の上限を指定します。”
- [community] “User ran a script on BigQuery for HTTP Archive data and was billed $14,000 with zero warning; complained about lack of customer support and …”
- [community] “BigQuery provides a dry run option to estimate bytes/costs before running a query, and shows bytes-to-be-scanned in small text before you hi…”
- [community] “I've worked with much larger datasets on BQ (petabyte scale) and managed to not spend more than $1000 in an hour; BQ tells you how much data…”
Trial
analystEvaluate with a free tier or trial — real queries on real data without a credit card or a sales call
weight 1 · round to BigQueryDatabricks offers a genuinely free, self-serve 'Free Edition' (replacing Community Edition) plus a free trial, explicitly for exploring datasets and running real queries/ML without needing a workspace purchase or sales call (docs-66, docs-73, docs-74, docs-39). This directly matches the analyst story of self-serve evaluation on real data. Missing for 10: explicit confirmation that no credit card is required to sign up, and independent/hands-on user corroboration of the free-tier experience (community evidence only discusses paid usage/cost complaints, not the free tier).
- [claimed-docs] “Free Edition gives you an easy-to-use Databricks workspace where you can explore datasets, build and share projects, and work with AI and ma…”
- [claimed-docs] “Databricks Free Edition is a no-cost version of Databricks designed for students, educators, hobbyists, and anyone interested in learning or…”
- [claimed-docs] “Free Edition replaced the legacy Databricks Community Edition, which was retired in 2025. If you previously used Community Edition, sign up …”
- [claimed-docs] “Start your journey with Databricks by signing up for a free trial account.”
BigQuery Sandbox explicitly lets analysts explore BigQuery capabilities and run real queries without providing a credit card or creating a billing account, and on-demand pricing includes 1 TiB of free query processing per month — directly matching the story of a no-card, no-sales-call evaluation path. Missing for 10: independent/hands-on user reports confirming the sandbox experience in practice (only first-party docs are cited, and community evidence focuses on billing surprises for paid usage rather than the sandbox itself).
- [claimed-docs] “Mit der BigQuery-Sandbox können Sie BigQuery nutzen, ohne eine Kreditkarte anzugeben oder ein Rechnungskonto für Ihr Projekt zu erstellen.”
- [claimed-docs] “The BigQuery sandbox lets you explore limited BigQuery capabilities at no cost to confirm whether BigQuery fits your needs.”
- [claimed-docs] “The BigQuery sandbox lets you experience BigQuery without providing a credit card or creating a billing account for your project.”
- [claimed-docs] “On-demand pricing (per TiB). With this pricing model, you are charged for the number of bytes processed by each query. The first 1 TiB of qu…”
Ecosystem integrations — the surrounding ecosystem — integrations, marketplaces, community packagesEcosystem integrations
The surrounding ecosystem — integrations, marketplaces, community packages
Bi
analystStandard drivers (JDBC/ODBC) and documented BI-tool integrations connect my dashboards without custom glue
weight 1 · round to BigQueryDocs show that external BI tools (Power BI, Tableau, Sigma) can query Databricks metric views directly, and a JDBC-style connection library (Databricks Connect) is documented, indicating standard driver-based BI connectivity without custom glue. However, no dedicated JDBC/ODBC driver certification page or explicit ODBC integration guide is cited in the pack. missing for 10: a direct JDBC/ODBC driver download/certification doc, independent BI-tool hands-on validation.
- [claimed-docs] “Query metric views from Power BI, Tableau, Sigma, and other external BI tools.”
- [claimed-docs] “Just like a JDBC driver, the Databricks Connect library can be embedded in any application to interact with Databricks.”
- [claimed-docs] “It runs directly on your data lake, supports ANSI SQL with Delta Lake extensions, and provides the tools to build highly performant, cost-ef…”
- [claimed-docs] “Define business metrics with consistent calculations using a semantic layer. Reuse metrics across queries and dashboards.”
BigQuery documents official Simba JDBC and ODBC drivers explicitly for connecting BI tools and applications to BigQuery, plus native integrations (bq CLI, Sheets, Analytics Hub/sharing) that support dashboard connectivity without custom glue. Missing for 10: no independent hands-on report confirming smooth BI-tool dashboard connection experience beyond vendor docs.
- [claimed-docs] “The Simba Open Database Connectivity (ODBC) and Java Database Connectivity (JDBC) drivers for BigQuery connect your applications to BigQuery…”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
Dev loop
data-engineerI get a fast local or free dev loop — a local engine, emulator, or sandbox — to develop transformations before touching production compute
weight 2 · round drawnDatabricks offers a free-tier workspace (Free Edition) and local-IDE tooling (Databricks Connect, dbt Core, CLI/SDK) that let a data-engineer author and test transformations from their laptop, but Databricks Connect still executes against real remote (Databricks) compute rather than a true local engine/emulator, and Free Edition is a hosted cloud workspace, not an offline sandbox. Community commentary even notes some engineers prefer spinning up their own local notebook/storage instead, underscoring the lack of a genuine local execution engine. missing for 10: a true local/offline execution engine or emulator that fully mimics production compute without any cloud dependency, and evidence of hands-on validation of that local dev loop.
- [claimed-docs] “Databricks Connect enables developers to develop and debug their code on Databricks compute using any IDE's native running and debugging fun…”
- [claimed-docs] “Just like a JDBC driver, the Databricks Connect library can be embedded in any application to interact with Databricks.”
- [claimed-docs] “Free Edition gives you an easy-to-use Databricks workspace where you can explore datasets, build and share projects, and work with AI and ma…”
- [claimed-docs] “Databricks Free Edition is a no-cost version of Databricks designed for students, educators, hobbyists, and anyone interested in learning or…”
- [claimed-docs] “dbt Core enables you to write dbt code in the IDE of your choice on your local development machine and then run dbt from the command line.”
- [claimed-docs] “dbt (data build tool) is a development environment for transforming data by writing select statements.”
- [community] “Not only is Databricks a deplorable company when it comes to HR, but their product is terrible. I really don't get what it's all about. Much…”
BigQuery offers a free 'Sandbox' mode that lets users try queries and transformations without a credit card or billing account, and on-demand pricing gives 1 TiB/month free — both lowering the barrier to a free dev loop. However, there is no evidence of a local engine or offline emulator (unlike e.g. Firestore's emulator); the Sandbox is still a hosted, quota-limited slice of the real cloud service, not a local dev-loop tool. Missing for 10: a documented local/offline emulator, evidence of local development without any cloud dependency, and any hands-on account confirming the sandbox is sufficient for full transformation development before touching production.
- [claimed-docs] “Mit der BigQuery-Sandbox können Sie BigQuery nutzen, ohne eine Kreditkarte anzugeben oder ein Rechnungskonto für Ihr Projekt zu erstellen.”
- [claimed-docs] “The BigQuery sandbox lets you explore limited BigQuery capabilities at no cost to confirm whether BigQuery fits your needs.”
- [claimed-docs] “The BigQuery sandbox lets you experience BigQuery without providing a credit card or creating a billing account for your project.”
- [claimed-docs] “On-demand pricing (per TiB). With this pricing model, you are charged for the number of bytes processed by each query. The first 1 TiB of qu…”
Transformation
data-engineerDbt is a first-class citizen — a documented adapter or native dbt project support with vendor docs to match
weight 2 · round to DatabricksDatabricks has a dedicated vendor docs page for dbt covering installation, connection, and dbt Core usage (databricks-docs-10, -14, -35), making dbt a documented first-class integration path for data engineers. missing for 10: no evidence of a native dbt-databricks adapter maintenance page, independent community corroboration of dbt workflow quality, or dbt Cloud-specific integration details.
- [claimed-docs] “This page explains what dbt is, how to install dbt Core, and how to connect.”
- [claimed-docs] “dbt Core enables you to write dbt code in the IDE of your choice on your local development machine and then run dbt from the command line.”
- [claimed-docs] “dbt (data build tool) is a development environment for transforming data by writing select statements.”
BigQuerynone0/10The evidence pack contains no vendor documentation from Google about a dbt adapter or native dbt project support for BigQuery; the only related item is a community comment casually noting BigQuery's pipe syntax is 'great for dbt macros,' which is a third-party observation, not vendor-documented adapter support. Absence of evidence for this applicable ecosystem-integration axis yields 'none.'
- [community] “Review after a week of using BigQuery's new SQL pipe syntax: much more productive for data exploration/cleaning, unifies WHERE/HAVING/QUALIF…”
Governance access — stories about governance access in this arenaGovernance access
Stories about governance access in this arena
Access
platform-engineerAccess control reaches tables, columns, and rows — roles plus masking policies — so one warehouse can serve many teams safely
weight 3 · round to BigQueryDatabricksdisputedcontradicted4/10Unity Catalog docs describe enforcing access control on tables and objects with grants (databricks-docs-19/38/46/56/70), which supports role-based table-level governance for many teams, but the evidence pack contains no first-party documentation of column-level masking policies or row-level filters. A community report explicitly contradicts the row/column-masking claim, stating Databricks security 'seems lacking - just table level, only in SQL and Spark, none in R' compared to competitors offering table/column/row-level security and dynamic masking (databricks-comm-9). Missing for 10: first-party docs on column masking policies, row-level security/filters, and independent confirmation these work as claimed to resolve the community-reported gap.
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Create a table and grant privileges in Databricks using the Unity Catalog data governance model.”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [community] “Snowflake has much more advanced data security - table, column, row level security, dynamic data masking, and zero-copy cloning. Databricks …”
IAM predefined roles for BigQuery resources (datasets, tables, views, routines) are documented, and the REST API surface exposes a rowAccessPolicies resource confirming row-level security exists, but the evidence pack contains no documentation of column-level access controls or masking policies (e.g., taxonomy-based dynamic data masking) needed to fully satisfy the story. missing for 10: column-level security/masking policy documentation, examples of combining row+column policies with IAM roles for multi-tenant governance, independent/hands-on corroboration of masking behavior.
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
- [claimed-docs] “BigQuery: Roles and permissions that apply to BigQuery resources such as datasets, tables, views, and routines.”
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
Governance
platform-engineerI get audit logs of who ran what and column-level lineage of where data came from
weight 2 · round to DatabricksUnity Catalog docs explicitly state it enforces access control, tracks lineage of data and AI assets, and logs activity for auditing automatically across the workspace, and Delta history supports auditing operations. missing for 10: no independent/hands-on corroboration of column-level lineage granularity or audit log query examples, and no explicit mention of 'who ran what' query-level attribution beyond general activity logging.
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Use history information to audit operations, roll back a table, or query a table at a specific point in time using time travel.”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
BigQuerynone0/10The evidence pack only covers IAM roles/permissions (bigquery-docs-31, bigquery-docs-65) and compliance certifications, with no mention of Cloud Audit Logs, job history, or column-level lineage tracking (e.g., Data Catalog/Dataplex lineage) for BigQuery. Governance/audit-and-lineage is a plausible and expected axis for a data warehouse, but no evidence in this pack substantiates it.
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
- [claimed-docs] “BigQuery: Roles and permissions that apply to BigQuery resources such as datasets, tables, views, and routines.”
platform-engineerCompliance attestations (SOC 2, HIPAA, PCI) are documented so security review does not stall the rollout
weight 1 · round drawnDatabricks' Trust Center compliance page explicitly documents SOC 2 Type II reports and a broader certification portfolio, directly addressing platform-engineer compliance-review needs, and this is reinforced by Unity Catalog's built-in access control, lineage, and audit logging documentation which supports such attestations operationally. Missing for 10: explicit mention of HIPAA and PCI attestations/BAA details in the evidence pack, and no independent/third-party audit confirmation beyond the vendor's own trust page.
- [claimed-docs] “Databricks Trust Center compliance page ("Ensuring Security, Privacy, & Compliance") documents attestations including a SOC 2 Type II report…”
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
Google Cloud's compliance resource center explicitly documents BigQuery-covered attestations including SOC 1/2/3, HIPAA support, ISO 27001-family certs, and downloadable audit reports, directly satisfying a platform-engineer's need for security-review documentation. Missing for 10: explicit PCI DSS attestation mention (only general 'sector/regional programs' referenced) and no independent third-party corroboration of the compliance claims.
- [claimed-docs] “Google Cloud compliance resource center lists BigQuery-covered attestations and certifications including ISO 9001:2015, ISO 22301:2019, ISO …”
Ingestion pipelines — stories about ingestion pipelines in this arenaIngestion pipelines
Stories about ingestion pipelines in this arena
Connectors
data-engineerFirst-party and partner connectors cover my sources — SaaS apps, databases, and ETL/ELT tools — with documented setup
weight 2 · round to DatabricksDatabricks documents an ingestion hub for choosing 'standard connectors by data source' (Lakeflow Connect/Auto Loader) and a documented partner integration for dbt, plus a Marketplace listing partner data/tool assets — showing first-party and partner connector coverage with setup docs. However, the evidence pack lacks an explicit named catalog of SaaS-app/database connectors (e.g., Salesforce, Workday, specific DB connectors) with individual setup guides beyond dbt and generic Auto Loader ingestion. Missing for 10: an exhaustive/documented list of named SaaS and database connectors with per-connector setup instructions, and independent confirmation that connector coverage meets diverse source needs.
- [claimed-docs] “Use it to choose a standard connector by data source and level of pipeline customization.”
- [claimed-docs] “This page explains what dbt is, how to install dbt Core, and how to connect.”
- [claimed-docs] “dbt Core enables you to write dbt code in the IDE of your choice on your local development machine and then run dbt from the command line.”
- [claimed-docs] “dbt (data build tool) is a development environment for transforming data by writing select statements.”
- [claimed-docs] “Incrementally and efficiently process new data files as they arrive in cloud storage.”
- [claimed-docs] “publish offerings that Databricks customers can discover, evaluate, and connect with directly from their workspace”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers. This gives customers a single catalog for f…”
BigQuery's docs show database replication ('replicating data from databases to BigQuery in near real time'), scheduled load jobs, Pub/Sub streaming ingestion, Dataform for ELT transformation workflows, and ODBC/JDBC drivers for tool connectivity — a real but partial connector story. However, there's no evidence of BigQuery Data Transfer Service or a documented catalog of first-party/partner SaaS-app connectors (e.g. Google Ads, Analytics, Salesforce, Fivetran/Stitch-style partner ecosystem), and one community report notes CloudSQL-to-BigQuery ingestion historically required a third-party tool rather than a native connector. Missing for 10: explicit SaaS-app connector documentation, a partner-connector directory/marketplace, and independent confirmation that documented setup for diverse source types (beyond DBs/Pub/Sub) is straightforward.
- [claimed-docs] “This method enables replicating data from databases to BigQuery in near real time.”
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub. Pub/Sub can handle high throughput of data loads into BigQuery…”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub.”
- [claimed-docs] “Pub/Sub can handle high throughput of data loads into BigQuery. It supports real-time data streaming, loading data as it's generated.”
- [claimed-docs] “Dataform is a service for data analysts to develop, test, control versions, and schedule complex workflows for data transformation in BigQue…”
- [claimed-docs] “Dataform lets you manage data transformation in the Extraction, Loading, and Transformation (ELT) process for data integration.”
- [claimed-docs] “The Simba Open Database Connectivity (ODBC) and Java Database Connectivity (JDBC) drivers for BigQuery connect your applications to BigQuery…”
- [community] “Last time I checked, it was still hard to get a Google CloudSQL DB into BigQuery, so it's surprising they did the Sheets integration first; …”
Loading
data-engineerBulk-load CSV, JSON, and Parquet from cloud object storage with a single documented command
weight 3 · round to BigQueryDatabricks documents ingestion tooling (a connector-selection page for cloud object storage sources) and Auto Loader for incrementally/bulk processing new files as they arrive in cloud storage, which covers CSV/JSON/Parquet ingestion, but no evidence names a single specific command (e.g., COPY INTO) or shows a concrete one-line bulk-load example spanning all three formats. Missing for 10: explicit single documented command syntax for bulk-loading CSV/JSON/Parquet, and independent/hands-on confirmation of ease-of-use.
- [claimed-docs] “Use it to choose a standard connector by data source and level of pipeline customization.”
- [claimed-docs] “Incrementally and efficiently process new data files as they arrive in cloud storage.”
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
BigQuery's documented `bq load` command and load jobs support single-command bulk loading of CSV, JSON, and Parquet directly from Cloud Storage, with scheduling and batch options documented (bigquery-docs-41, bigquery-docs-12/49/58). This is a well-established, heavily documented core capability. Missing for 10: independent hands-on confirmation specifically of the load command (community evidence covers pricing/perf but not load-command usage) and explicit mention of all three formats in one citation.
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “vai aprender a usar o `bq`, a ferramenta de interface de linha de comandos (CLI) baseada em Python para o BigQuery, para criar um conjunto d…”
- [claimed-docs] “learn how to use bq, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query tables”
- [claimed-docs] “you learn how to use `bq`, the Python-based command-line interface (CLI) tool for BigQuery to create a dataset, load sample data, and query …”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
data-engineerA managed service continuously ingests new files or events as they arrive, without me running my own pipeline infrastructure
weight 2 · round drawnDatabricks provides Auto Loader/Lakeflow pipelines to incrementally and efficiently ingest new files as they arrive in cloud storage, plus Structured Streaming for continuous event processing with low latency and exactly-once guarantees on Delta Lake, all as managed serverless/DLT-style pipelines rather than self-managed infra, and a broader ingestion connector catalog for various sources. Missing for 10: independent/hands-on evidence of production reliability at scale for continuous ingestion, and more detail on serverless auto-scaling/operational overhead reduction claims beyond docs.
- [claimed-docs] “Incrementally and efficiently process new data files as they arrive in cloud storage.”
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
- [claimed-docs] “Process data for real-time workloads with end-to-end latency as low as five milliseconds.”
- [claimed-docs] “Structured Streaming lets you express computation on streaming data in the same way you express a batch computation on static data.”
- [claimed-docs] “Use Delta Lake tables as streaming sources and sinks with exactly-once processing guarantees.”
- [claimed-docs] “Use it to choose a standard connector by data source and level of pipeline customization.”
BigQuery provides multiple fully-managed ingestion paths for continuously arriving data: the Storage Write API with exactly-once semantics for streaming writes, a native Pub/Sub-to-BigQuery subscription for high-throughput event streaming, Continuous Queries for real-time analysis of incoming data, and scheduled or near-real-time database replication load jobs, all without the user provisioning or operating pipeline infrastructure. Docs explicitly market BigQuery as having no infrastructure to set up or manage. Missing for 10: independent or hands-on evidence validating continuous streaming ingestion at scale, and more detail on CDC or file-arrival-triggered ingestion beyond Pub/Sub and replication mentions.
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub. Pub/Sub can handle high throughput of data loads into BigQuery…”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub.”
- [claimed-docs] “Pub/Sub can handle high throughput of data loads into BigQuery. It supports real-time data streaming, loading data as it's generated.”
- [claimed-docs] “we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has lower pricing and …”
- [claimed-docs] “For new projects, we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has …”
- [claimed-docs] “BigQuery の継続的クエリは、継続的に実行される SQL ステートメントです。継続的クエリを使用すると、BigQuery で受信データをリアルタイムで分析できます。”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
- [claimed-docs] “This method enables replicating data from databases to BigQuery in near real time.”
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “there's no infrastructure to set up or manage, letting you focus on finding meaningful insights using GoogleSQL or Python”
Notebooks workspace — stories about notebooks workspace in this arenaNotebooks workspace
Stories about notebooks workspace in this arena
Notebooks
analystFirst-party notebooks let me mix SQL and Python against warehouse data, with results and charts inline
weight 2 · round drawnDatabricks notebooks docs explicitly describe multi-language coauthoring, automatic versioning, and built-in visualizations, and the getting-started guide shows querying Unity Catalog data and visualizing results inline in a notebook — directly matching the analyst story of mixing SQL/Python with inline charts. Missing for 10: explicit documentation of magic-command language switching within a single notebook cell (%sql/%python) and independent/hands-on corroboration beyond first-party docs.
- [claimed-docs] “Databricks notebooks provide real-time coauthoring in multiple languages, automatic versioning, and built-in data visualizations”
- [claimed-docs] “Databricks notebooks provide real-time coauthoring in multiple languages, automatic versioning, and built-in data visualizations for develop…”
- [claimed-docs] “Use a Databricks notebook to query sample data stored in Unity Catalog and then visualize the query results in the notebook.”
BigQuery's notebooks feature explicitly combines SQL queries, Python code, rich text, and inline visualizations, with Colab Enterprise integration enabling end-to-end data science/ML workflows in one interface. This directly matches the analyst story of mixing SQL/Python with inline results and charts. Missing for 10: independent/hands-on corroboration of the notebook experience beyond first-party docs, and detail on how seamlessly SQL and Python cells interoperate in practice.
- [claimed-docs] “notebooks let you combine SQL queries with Python code, rich text, and visualizations to tell a comprehensive story with your data.”
- [claimed-docs] “Colab Enterprise notebooks in BigQuery let you perform end-to-end data science and machine learning workflows within a single, integrated in…”
- [claimed-docs] “End-to-end ML workflows: build, evaluate, and deploy a BigQuery ML model within a single notebook interface.”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to BigQueryDatabricksdisputedcontradicted5/10Databricks documents an extensive, versioned REST API surface (workspace-level API reference, SDKs, Terraform provider, CLI that itself is a thin wrapper around the REST API — `databricks api`) and states that Unity Catalog objects can be managed via 'Catalog Explorer, SQL, the Databricks CLI, and REST APIs,' suggesting broad UI/API parity. However, a hands-on community report specifically contradicts full parity, stating 'no good way to get usage info programmatically' and calling out other CLI/SQL config gaps (e.g., can't set Spark config easily), showing some UI-visible functionality isn't cleanly exposed via API. Missing for 10: a documented comprehensive parity guarantee, independent confirmation that admin/UI-only features (budgets, workspace settings) are fully scriptable, and resolution of the community-reported programmatic usage-info gap.
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
- [claimed-docs] “you learn how to automate Databricks operations and accelerate development with the Databricks SDK for Python”
- [probe] “PROBE runtime (recorded 2026-09-06): the Databricks CLI installed via the vendor's brew tap and printed `Databricks CLI v1.15.0` keylessly; …”
- [community] “No persist() so can't cache dataframes; no good way to get usage info programmatically; can't set Spark config easily (had to hack S3A crede…”
BigQuery exposes a comprehensive, self-describing REST API (datasets/jobs/models/tables/routines) plus the bq CLI and client libraries, and SQL/DML/ML training operations that appear in the UI are all reachable via API or SQL statements — core data and query operations have strong API parity. However, several newer UI-centric features (Gemini-powered conversational analytics, BigQuery Data Canvas, notebook-based workflows) are documented as console experiences without clear evidence of an equivalent programmatic/API path for all their functionality. Missing for 10: explicit documentation that conversational analytics/Data Canvas/notebook features are fully accessible via API rather than console-only, and independent hands-on confirmation of full UI/API parity.
- [probe] “PROBE runtime (recorded 2026-09-06): the BigQuery v2 REST discovery document downloaded keylessly from bigquery.googleapis.com and parsed cl…”
- [probe] “official CLI documented at https://cloud.google.com/bigquery/docs/bq-command-line-tool”
- [claimed-docs] “vai aprender a usar o `bq`, a ferramenta de interface de linha de comandos (CLI) baseada em Python para o BigQuery, para criar um conjunto d…”
- [claimed-docs] “Data Manipulation Language (DML) statements enable you to update, insert, and delete data from your BigQuery tables.”
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
- [claimed-docs] “BigQuery 데이터 캔버스로 데이터 탐색, 변환, 쿼리, 시각화”
- [claimed-docs] “대화형 분석을 사용하면 자연어로 데이터와 대화할 수 있습니다.”
- [claimed-docs] “notebooks let you combine SQL queries with Python code, rich text, and visualizations to tell a comprehensive story with your data.”
ai-native userExport all of my data in open formats and leave
weight 3 · round to DatabricksDatabricks stores data in the open Delta Lake/Parquet format and offers Delta Sharing (OpenSharing) to share or export data outside the organization regardless of platform, plus CLI/SDK/REST API access for programmatic extraction of data and metadata. However, there's no single documented 'export everything and leave' workflow or bulk account-export tool, and community commentary (e.g., migrating workloads to Postgres) suggests migration is done piecemeal rather than via a turnkey export feature. missing for 10: a dedicated full-account/bulk data export or migration tool, and independent hands-on confirmation of frictionless full data egress.
- [claimed-docs] “OpenSharing is the secure data sharing platform in Databricks that lets you share data and AI assets with users outside your organization, r…”
- [claimed-docs] “The Open Marketplace, which does not require access to a Databricks workspace.”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
- [claimed-docs] “This reference contains information about the Databricks workspace-level application programming interfaces (APIs).”
- [community] “We've been moving our workflows out of Databricks to PostgreSQL to save a ton.”
BigQuery documents Iceberg managed tables that store data in customer-owned buckets in the open Apache Iceberg format for interoperability with third-party/open-source engines, and states general support for open table formats like Iceberg, Delta, and Hudi, addressing data portability/lock-in concerns. However, evidence does not show a dedicated bulk-export mechanism (e.g. bq extract to CSV/Avro/Parquet/JSON) for standard native BigQuery tables, nor confirmation that ALL data (not just tables opted into Iceberg format) can be freely exported without proprietary lock-in. Missing for 10: documented export/extract tooling for native tables in open formats, evidence of full-dataset export workflows, and independent confirmation that migrating away is friction-free.
- [claimed-docs] “Iceberg managed tables offer the same fully managed experience as standard BigQuery tables, but store data in customer-owned storage buckets…”
- [claimed-docs] “Iceberg managed tables support the open Iceberg table format for better interoperability with open-source and third-party compute engines”
- [claimed-docs] “Iceberg managed tables support the open Iceberg table format for better interoperability with open-source and third-party compute engines on…”
- [claimed-docs] “BigQuery provides a uniform way to work with both structured and unstructured data and supports open table formats like Apache Iceberg, Delt…”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnDatabricksnone0/10No evidence in the pack addresses regional deployment, data residency controls, or workspace region selection for Databricks; the compliance mention only covers SOC 2 attestations, not data-location choice.
BigQuerynone0/10The evidence pack contains no documentation or mention of choosing a dataset/table region, multi-region options, or data residency controls for BigQuery; only general compliance certifications and unrelated feature descriptions are provided. Missing for 10: explicit docs on selecting dataset location/region, data residency guarantees, or region-locking configuration.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnDatabricksnone0/10The evidence pack covers Databricks' data platform, governance, agent, and dev-tool features but contains no mention of any control to opt out of, or prevent, customer data being used to train Databricks' (or third-party) AI models — no data-use policy, model-training opt-out setting, or contractual guarantee is cited. missing for 10: an explicit AI-training opt-out/data-use policy, documentation of contractual or technical controls preventing model training on customer data, and any independent confirmation of such a guarantee.
BigQuerynone0/10The evidence pack covers BigQuery's features (SQL, Gemini in BigQuery, pricing, sharing, compliance certifications) but contains no documentation of a specific data-usage policy or customer control for opting out of having data used to train Google's AI models, despite BigQuery now integrating Gemini AI features. Missing for 10: explicit AI-training data-usage policy, opt-out/opt-in controls, and any documentation addressing whether customer data feeds model training.
ai-native userControl data retention and deletion
weight 2 · round to BigQueryDatabricksnone0/10Evidence covers Unity Catalog governance (access control, lineage, audit logging) and Delta Lake time-travel/versioning, but nothing documents user-controllable data retention periods, deletion/erasure APIs, or lifecycle policies for AI-native privacy control. Missing for 10: explicit retention configuration, right-to-delete/erasure mechanisms, data lifecycle/expiry policy documentation, and any independent verification of deletion behavior.
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Use history information to audit operations, roll back a table, or query a table at a specific point in time using time travel.”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
BigQuery's DML supports DELETE operations on tables (bigquery-docs-2) and Iceberg managed tables reference 'time travel for historical data access' (bigquery-docs-44), giving users mechanisms to remove or manage historical data, and IAM/access-control docs show governance controls exist (bigquery-docs-31). However, there is no direct documentation of retention policy configuration (e.g., table/partition expiration settings) or a dedicated data-deletion/GDPR compliance workflow in the evidence pack. Missing for 10: explicit table/dataset expiration policy docs, a dedicated data retention & deletion API/console feature, and independent confirmation that deletion requests are honored end-to-end.
- [claimed-docs] “Data Manipulation Language (DML) statements enable you to update, insert, and delete data from your BigQuery tables.”
- [claimed-docs] “Time travel for historical data access in BigQuery.”
- [claimed-docs] “In diesem Dokument finden Sie eine Liste der vordefinierten IAM-Rollen (Identity and Access Management) und Berechtigungen für BigQuery.”
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnDatabricksnone0/10No evidence pack item addresses telemetry/usage-tracking opt-out settings for Databricks; documentation covers compliance/SOC2 but not a specific telemetry toggle. missing for 10: any mention of telemetry collection, opt-out settings, or usage-tracking controls.
BigQuerynone0/10The evidence pack contains no mention of BigQuery collecting telemetry/usage data from client tools (bq CLI, client libraries, notebooks) nor any documented opt-out mechanism, despite this being a plausible axis for a cloud data platform with CLI/SDK tooling. missing for 10: any documentation of telemetry collection practices, an opt-out flag/setting (e.g., in bq CLI or client SDKs), or privacy-posture statements addressing usage tracking.
Semantic layer — stories about semantic layer in this arenaSemantic layer
Stories about semantic layer in this arena
Semantics
analystDefine a governed semantic model — metrics, dimensions, and joins declared once — that queries and AI tools answer against consistently
weight 2 · round to DatabricksDatabricks Unity Catalog Metric Views let analysts define metrics, dimensions, and joins once ('you define the metric once... users can group by any available field. The query engine generates the correct computation') and this governed semantic layer is queryable consistently from SQL, AI/BI dashboards, Genie natural-language Q&A, and external BI tools like Power BI, Tableau, and Sigma. missing for 10: independent/hands-on corroboration of consistency across tools, and detail on join declaration beyond metric definition.
- [claimed-docs] “you define the metric once, for example _sum of revenue divided by distinct customer count_, and users can group by any available field.”
- [claimed-docs] “you define the metric once, for example sum of revenue divided by distinct customer count, and users can group by any available field”
- [claimed-docs] “you define the metric once, for example sum of revenue divided by distinct customer count, and users can group by any available field. The q…”
- [claimed-docs] “Query metric views from Power BI, Tableau, Sigma, and other external BI tools.”
- [claimed-docs] “Define business metrics with consistent calculations using a semantic layer. Reuse metrics across queries and dashboards.”
- [claimed-docs] “Ask data questions in natural language and get answers grounded in your organization's data.”
BigQuerynone0/10The evidence pack shows BigQuery's SQL engine, Dataform (ELT transformation), notebooks, and Gemini natural-language querying, but no dedicated semantic-layer feature where metrics, dimensions, and joins are declared once and consistently reused across queries and AI tools (e.g., no LookML/metrics-layer equivalent is documented here). Dataform manages transformation pipelines, not a governed semantic/metrics model, so the specific capability described in the story is unevidenced.
- [claimed-docs] “Dataform is a service for data analysts to develop, test, control versions, and schedule complex workflows for data transformation in BigQue…”
- [claimed-docs] “Dataform lets you manage data transformation in the Extraction, Loading, and Transformation (ELT) process for data integration.”
- [claimed-docs] “Query statements... are the primary method to analyze data in BigQuery.”
Sharing marketplace — stories about sharing marketplace in this arenaSharing marketplace
Stories about sharing marketplace in this arena
Sharing
analystA marketplace of third-party datasets lets me enrich my own data directly inside the platform
weight 1 · round drawnDatabricks Marketplace is documented as an in-platform catalog where customers can discover, evaluate, and connect directly to third-party datasets (plus AI models, notebooks, apps, MCP servers) from within their workspace, and an Open Marketplace variant even works without a workspace. This directly matches the analyst story of enriching own data with third-party datasets inside the platform. Missing for 10: no independent/hands-on evidence of the enrichment workflow in practice, and no detail on how a discovered dataset is joined/merged with an analyst's own tables.
- [claimed-docs] “The Open Marketplace, which does not require access to a Databricks workspace.”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers.”
- [claimed-docs] “publish offerings that Databricks customers can discover, evaluate, and connect with directly from their workspace”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers. This gives customers a single catalog for f…”
BigQuery sharing (formerly Analytics Hub) is explicitly documented as a data exchange platform for discovering, sharing and accessing third-party/Google curated datasets and combining them with internal data, directly inside the platform. Multiple docs confirm this cross-org discovery and combination workflow. Missing for 10: independent hands-on evidence/user reports specifically validating the marketplace discovery/enrichment workflow (community evidence only covers other BigQuery aspects like pricing/performance, not the marketplace).
- [claimed-docs] “La condivisione di BigQuery (precedentemente Analytics Hub) è una piattaforma di scambio di dati che consente di condividere, scoprire e acc…”
- [claimed-docs] “Puoi utilizzare BigQuery sharing per scoprire set di dati curati di terze parti e di Google e combinarli con i tuoi dati interni”
- [claimed-docs] “BigQuery sharing (formerly Analytics Hub) is a data exchange platform that lets you securely share, discover, and access data across organiz…”
- [claimed-docs] “La condivisione di BigQuery... è una piattaforma di scambio di dati che consente di condividere, scoprire e accedere in modo sicuro ai dati …”
data-engineerShare live datasets with another account or organization without copying data or building an export pipeline
weight 2 · round drawnDelta Sharing (OpenSharing) is documented as a secure data sharing platform for sharing data and AI assets outside the organization without requiring the recipient to be on Databricks, explicitly avoiding data copying/export pipelines, and the Marketplace/Open Marketplace extends this to discovery and listing of shared datasets across organizations. missing for 10: independent hands-on validation of cross-account sharing beyond vendor docs.
- [claimed-docs] “OpenSharing is the secure data sharing platform in Databricks that lets you share data and AI assets with users outside your organization, r…”
- [claimed-docs] “The Open Marketplace, which does not require access to a Databricks workspace.”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers.”
- [claimed-docs] “publish offerings that Databricks customers can discover, evaluate, and connect with directly from their workspace”
- [claimed-docs] “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers. This gives customers a single catalog for f…”
BigQuery sharing (formerly Analytics Hub) is explicitly documented as a data exchange platform enabling secure sharing, discovery, and access to data across organizational boundaries without replicating data—exactly matching the story of sharing live datasets with another account/org without copying or building an export pipeline. Missing for 10: independent/hands-on corroboration beyond first-party docs.
- [claimed-docs] “La condivisione di BigQuery (precedentemente Analytics Hub) è una piattaforma di scambio di dati che consente di condividere, scoprire e acc…”
- [claimed-docs] “La condivisione di BigQuery... è una piattaforma di scambio di dati che consente di condividere, scoprire e accedere in modo sicuro ai dati …”
- [claimed-docs] “Puoi utilizzare BigQuery sharing per scoprire set di dati curati di terze parti e di Google e combinarli con i tuoi dati interni”
- [claimed-docs] “BigQuery sharing (formerly Analytics Hub) is a data exchange platform that lets you securely share, discover, and access data across organiz…”
Sql analytics — stories about sql analytics in this arenaSql analytics
Stories about sql analytics in this arena
Lakehouse
data-engineerQuery open table formats and files in object storage — Iceberg, Delta, Parquet — without first loading them into proprietary storage
weight 2 · round to BigQueryDocs confirm Databricks SQL runs directly on the data lake with ANSI SQL and Delta Lake extensions 'without moving your data' (databricks-docs-32), and Unity Catalog governs tables/external locations directly in object storage (databricks-docs-19, databricks-docs-38, databricks-docs-46). However, the evidence pack never explicitly mentions querying Iceberg tables or plain Parquet files in-place (e.g., via UniForm or external tables) — missing for 10: explicit Iceberg format support, Parquet file querying without ingestion, and independent/hands-on confirmation of in-place multi-format querying.
- [claimed-docs] “It runs directly on your data lake, supports ANSI SQL with Delta Lake extensions, and provides the tools to build highly performant, cost-ef…”
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
- [claimed-docs] “Unity Catalog operates beneath every data and AI interaction in your workspaces automatically: enforcing access control when you query a tab…”
- [claimed-docs] “Create a table and grant privileges in Databricks using the Unity Catalog data governance model.”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
BigQuery explicitly supports querying open table formats (Iceberg, Delta, Hudi) and files in object storage without ingesting them into proprietary storage — Iceberg managed tables store data in customer-owned buckets with schema evolution and time travel, and BigQuery natively supports open table formats per its intro docs. Missing for 10: independent hands-on validation of querying Delta/Parquet files directly (only Iceberg is detailed with specifics), and no community corroboration of external-table performance/limitations for these formats.
- [claimed-docs] “BigQuery provides a uniform way to work with both structured and unstructured data and supports open table formats like Apache Iceberg, Delt…”
- [claimed-docs] “Iceberg managed tables offer the same fully managed experience as standard BigQuery tables, but store data in customer-owned storage buckets…”
- [claimed-docs] “Iceberg managed tables support the open Iceberg table format for better interoperability with open-source and third-party compute engines”
- [claimed-docs] “Iceberg managed tables support the open Iceberg table format for better interoperability with open-source and third-party compute engines on…”
- [claimed-docs] “_Schema evolution_, which lets you add, drop, and rename columns to suit your needs.”
- [claimed-docs] “Time travel for historical data access in BigQuery.”
Performance
data-engineerInspect query profiles and execution plans to find why a query is slow or expensive
weight 2 · round to BigQueryDocs explicitly describe inspecting execution plans to find bottlenecks/optimization opportunities and automatic insights/recommendations for inefficient queries, directly matching the story. This is first-party documentation without independent hands-on corroboration or deeper detail on cost/spill/skew diagnostics. Missing for 10: independent/community validation of query profile usability, and more detail on cost breakdown metrics beyond the brief doc snippets.
- [claimed-docs] “Get automatic insights and recommendations when queries run inefficiently.”
- [claimed-docs] “Inspect the execution plan for a query to identify bottlenecks and optimization opportunities.”
- [claimed-docs] “It runs directly on your data lake, supports ANSI SQL with Delta Lake extensions, and provides the tools to build highly performant, cost-ef…”
BigQuery documents dedicated query plan/execution diagnostics via EXPLAIN-like output and timing information, letting engineers inspect stages, slot usage, and bottlenecks; the pricing model also exposes bytes-scanned/dry-run cost estimation used to diagnose expensive queries. Missing for 10: independent hands-on walkthroughs of using the query plan UI to debug a real slow query (only docs mention the feature, no community corroboration of using execution plans specifically).
- [claimed-docs] “BigQuery includes diagnostic query plan and timing information. This is similar to the information provided by statements such as EXPLAIN in…”
- [claimed-docs] “BigQuery includes diagnostic query plan and timing information. This is similar to the information provided by statements such as `EXPLAIN` …”
- [community] “BigQuery provides a dry run option to estimate bytes/costs before running a query, and shows bytes-to-be-scanned in small text before you hi…”
- [community] “I've worked with much larger datasets on BQ (petabyte scale) and managed to not spend more than $1000 in an hour; BQ tells you how much data…”
- [claimed-docs] “You can also reserve compute capacity ahead of time in the form of slots, which represent virtual CPUs.”
Recovery
data-engineerTime-travel — query data as of a past point and restore dropped or corrupted tables from history
weight 2 · round to DatabricksDatabricks documents Delta Lake history/time travel explicitly: querying tables as of a past version/timestamp and restoring/rolling back dropped or corrupted tables using history information, directly matching the story. This is a core, well-documented Delta Lake feature integrated into Databricks SQL/Unity Catalog. Missing for 10: no independent/hands-on corroboration of restore-after-drop specifically, beyond first-party docs.
- [claimed-docs] “Use history information to audit operations, roll back a table, or query a table at a specific point in time using time travel.”
Only one brief doc mention confirms BigQuery supports time travel for historical data access, but the evidence pack lacks detail on the retention window, querying historical snapshots via SQL syntax, or the fail-safe/restore-dropped-table mechanism, and has no independent/hands-on corroboration. missing for 10: documentation on the time-travel query syntax (FOR SYSTEM_TIME AS OF), the default/configurable retention window, explicit restore-dropped-table (fail-safe) workflow, and any community validation of these mechanisms.
- [claimed-docs] “Time travel for historical data access in BigQuery.”
Sql
analystI get a full analytical SQL surface — window functions, CTEs, semi-structured JSON, arrays, and rich date/time types — without bolt-on extensions
weight 3 · round to BigQueryDatabricks SQL is documented as ANSI-SQL compliant with Delta Lake extensions (databricks-docs-32) and Unity Catalog governs SQL objects (databricks-docs-56), implying a broad native SQL surface without bolt-on tools, but the evidence pack never explicitly calls out window functions, CTEs, JSON/semi-structured handling, array types, or date/time type richness. missing for 10: explicit doc citations for window functions, CTE support, JSON/semi-structured query functions, array manipulation functions, and native date/time types.
- [claimed-docs] “It runs directly on your data lake, supports ANSI SQL with Delta Lake extensions, and provides the tools to build highly performant, cost-ef…”
- [claimed-docs] “You work with the objects Unity Catalog governs through Catalog Explorer, SQL, the Databricks CLI, and REST APIs.”
- [claimed-docs] “enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used, logging activity for audit…”
BigQuery's GoogleSQL is documented as ANSI-compliant with support for DML, query statements, and a full analytical surface including window functions, CTEs, arrays/structs, JSON handling, and date/time types (implied by GoogleSQL SQL reference and confirmed by community praise of pipe syntax and table-valued functions). Independent hands-on testimony (bigquery-comm-15, bigquery-comm-16) confirms real analyst usage of advanced SQL constructs without extensions. Missing for 10: explicit doc citations enumerating window functions, JSON/array functions, and DATETIME/TIMESTAMP types individually rather than inferred from general SQL intro docs.
- [claimed-docs] “Query statements... are the primary method to analyze data in BigQuery.”
- [claimed-docs] “GoogleSQL is an ANSI-compliant Structured Query Language (SQL) that includes the following types of supported statements”
- [claimed-docs] “Data Manipulation Language (DML) statements enable you to update, insert, and delete data from your BigQuery tables.”
- [community] “Review after a week of using BigQuery's new SQL pipe syntax: much more productive for data exploration/cleaning, unifies WHERE/HAVING/QUALIF…”
- [community] “BigQuery has table-valued functions already, which can be used with pipes with a CALL clause.”
Streaming realtime — stories about streaming realtime in this arenaStreaming realtime
Stories about streaming realtime in this arena
Streaming
data-engineerRun continuous or incremental transformations — streams, tasks, declarative pipelines, or continuous queries — inside the platform
weight 1 · round to DatabricksDatabricks documents Structured Streaming for continuous/incremental processing with exactly-once guarantees and low-latency options, Lakeflow declarative pipelines with Auto Loader for incremental ingestion/ETL, and Databricks Workflows/Jobs orchestration referenced via CLI/tasks. This directly covers streams, declarative pipelines, and continuous/incremental transformations inside the platform. Missing for 10: no independent/hands-on benchmark validating claimed low-latency figures, and no explicit mention of scheduled/triggered task orchestration UI beyond CLI/Lakeflow docs.
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
- [claimed-docs] “Incrementally and efficiently process new data files as they arrive in cloud storage.”
- [claimed-docs] “Process data for real-time workloads with end-to-end latency as low as five milliseconds.”
- [claimed-docs] “Structured Streaming lets you express computation on streaming data in the same way you express a batch computation on static data.”
- [claimed-docs] “Use Delta Lake tables as streaming sources and sinks with exactly-once processing guarantees.”
- [claimed-docs] “Develop and deploy your first ETL (extract, transform, and load) pipeline for data orchestration with Apache Spark™.”
BigQuery provides continuous queries (SQL statements running continuously to analyze streaming data in real time) as a first-class feature, plus Dataform for declarative, version-controlled, scheduled transformation pipelines with dependency graphs, and streaming ingestion via Storage Write API/Pub/Sub subscriptions for near-real-time pipelines. Together these cover streams, tasks/scheduled pipelines, declarative pipelines, and continuous queries as described in the story. Missing for 10: independent/hands-on validation of continuous queries at scale and clearer documentation of latency/cost tradeoffs in production use.
- [claimed-docs] “BigQuery の継続的クエリは、継続的に実行される SQL ステートメントです。継続的クエリを使用すると、BigQuery で受信データをリアルタイムで分析できます。”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
- [claimed-docs] “Dataform is a service for data analysts to develop, test, control versions, and schedule complex workflows for data transformation in BigQue…”
- [claimed-docs] “View a visualization of the dependency tree of your workflow.”
- [claimed-docs] “Collaborate with team members on workflow development through Git.”
- [claimed-docs] “Dataform lets you manage data transformation in the Extraction, Loading, and Transformation (ELT) process for data integration.”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub. Pub/Sub can handle high throughput of data loads into BigQuery…”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub.”
- [claimed-docs] “Pub/Sub can handle high throughput of data loads into BigQuery. It supports real-time data streaming, loading data as it's generated.”
- [claimed-docs] “This method enables replicating data from databases to BigQuery in near real time.”
- [claimed-docs] “you can schedule load jobs. You can schedule one-time or batch data transfers at regular intervals”
- [claimed-docs] “The Storage Write API (gRPC) has lower pricing and more robust features, including exactly-once delivery semantics.”
- [claimed-docs] “For new projects, we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has …”
data-engineerStreaming writes land queryable within seconds through a documented streaming ingestion API
weight 2 · round to BigQueryDatabricks documents Structured Streaming as a first-class streaming API that treats streaming like batch, supports Delta Lake tables as streaming sources/sinks with exactly-once guarantees, and Auto Loader for incremental ingestion from cloud storage, with end-to-end latency claimed as low as 5ms — directly supporting near-real-time queryable ingestion. missing for 10: independent/hands-on benchmark confirming 'queryable within seconds' end-to-end, and no community corroboration specific to streaming latency claims.
- [claimed-docs] “Structured Streaming lets you express computation on streaming data in the same way you express a batch computation on static data.”
- [claimed-docs] “Incrementally and efficiently process new data files as they arrive in cloud storage.”
- [claimed-docs] “Use Delta Lake tables as streaming sources and sinks with exactly-once processing guarantees.”
- [claimed-docs] “Process data for real-time workloads with end-to-end latency as low as five milliseconds.”
- [claimed-docs] “Create and deploy an ETL (extract, transform, and load) pipeline for data orchestration using Lakeflow pipelines and Auto Loader.”
BigQuery's Storage Write API (gRPC) is documented as a real-time streaming ingestion API with exactly-once delivery semantics, and continuous queries/Pub/Sub subscriptions confirm data becomes queryable in near real time. missing for 10: no independent/hands-on benchmark confirming actual seconds-level latency from write to queryability, and no explicit SLA/numeric latency figure in docs.
- [claimed-docs] “The Storage Write API (gRPC) has lower pricing and more robust features, including exactly-once delivery semantics.”
- [claimed-docs] “we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has lower pricing and …”
- [claimed-docs] “For new projects, we recommend using the Storage Write API (gRPC) instead of the Storage Write API (REST). The Storage Write API (gRPC) has …”
- [claimed-docs] “BigQuery continuous queries are SQL statements that run continuously. Continuous queries let you analyze incoming data in BigQuery in real t…”
- [claimed-docs] “BigQuery の継続的クエリは、継続的に実行される SQL ステートメントです。継続的クエリを使用すると、BigQuery で受信データをリアルタイムで分析できます。”
- [claimed-docs] “Pub/Sub can handle high throughput of data loads into BigQuery. It supports real-time data streaming, loading data as it's generated.”
- [claimed-docs] “To stream data into BigQuery, you can use a BigQuery subscription in Pub/Sub. Pub/Sub can handle high throughput of data loads into BigQuery…”
Not comparable on these axes
ai-native userRead the product's source under an open license
weight 2 · not comparableDatabricksn/aDatabricks is a proprietary commercial data/AI platform; there is no evidence of its core source code being available under an open license, and 'read the product's source' is not a fair expectation for this category of SaaS platform (unlike an open-source library or framework). This axis does not apply.
BigQueryn/aBigQuery is a proprietary, closed-source managed cloud data warehouse service; there is no source code released under an open license for users to inspect. This is a category error — BigQuery is a hosted SaaS product, not open-source software, so the axis of 'reading source under an open license' does not apply to the product itself.
ai-native userSelf-host the core product
weight 3 · not comparableDatabricksnone0/10Databricks is offered exclusively as a managed cloud service (AWS/Azure/GCP workspaces, pay-as-you-go billing) with no documented option to self-host the core platform on your own infrastructure; the closest analog (Free Edition) is still a hosted SaaS trial, not a self-hostable deployment. Community evidence even references how self-hosting Spark used to be painful specifically because Databricks replaced that with a hosted service, reinforcing that the core product is not self-hostable.
- [claimed-docs] “Databricks offers you a pay-as-you-go approach with no up-front costs. Only pay for the products you use at per second granularity.”
- [claimed-docs] “Databricks Free Edition is a no-cost version of Databricks designed for students, educators, hobbyists, and anyone interested in learning or…”
- [community] “They had an excellent Spark-as-a-Service product, at a time when you'd have better luck finding a leprechaun than a reliable self-hosted Spa…”