Install
Showcase


Try itExperimental
See what an agent can do with Airbyte before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -si https://api.airbyte.com/v1/connections | head -4recorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Ai pipelines — stories about ai pipelines in this arenaAi pipelinesevidence →
Stories about ai pipelines in this arena
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Code first portability — stories about code first portability in this arenaCode first portabilityevidence →
Stories about code first portability in this arena
Connectors catalog — stories about connectors catalog in this arenaConnectors catalogevidence →
Stories about connectors catalog in this arena
Observability reliability — stories about observability reliability in this arenaObservability reliabilityevidence →
Stories about observability reliability in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Orchestration scheduling — stories about orchestration scheduling in this arenaOrchestration schedulingevidence →
Stories about orchestration scheduling in this arena
Pricing cost — stories about pricing cost in this arenaPricing costevidence →
Stories about pricing cost in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Reverse etl activation — stories about reverse etl activation in this arenaReverse etl activationevidence →
Stories about reverse etl activation in this arena
Schema evolution — stories about schema evolution in this arenaSchema evolutionevidence →
Stories about schema evolution in this arena
Sync replication — stories about sync replication in this arenaSync replicationevidence →
Stories about sync replication in this arena
Transformations dbt — stories about transformations dbt in this arenaTransformations dbtevidence →
Stories about transformations dbt in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 4 free · 0 paid · 0 enterprise · 36 not stated in evidence
Follow the green: where the map greys out is where Airbyte stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Versioning policy · The catalog tells me each connector's maturity, support level, and maintainer before I depend on it · I see end-to-end lineage of my datasets — which sources, steps, and transformations produced each table
Subscribe to events via webhooks
~5/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
~4/10
Connect an agent via an official MCP server
✓9/10
Download a machine-readable API spec (OpenAPI or equivalent)
~4/10
unlocks → Interactive API docs
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
~5/10
Explore an interactive API reference with runnable examples
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
~5/10
unlocks → MCP client · AI insights
Operate the product with natural-language commands
~7/10
Plug MCP servers into this product so it can use their tools
—0/10
Get AI-generated insights and suggestions from my data inside the product
—0/10
Set up automations that run autonomously in the background
✓7/10
Ai pipelines — stories about ai pipelines in this arenaAi pipelines
Stories about ai pipelines in this arena
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Code first portability — stories about code first portability in this arenaCode first portability
Stories about code first portability in this arena
Connectors catalog — stories about connectors catalog in this arenaConnectors catalog
Stories about connectors catalog in this arena
Observability reliability — stories about observability reliability in this arenaObservability reliability
Stories about observability reliability in this arena
Tell how fresh each destination table is and get warned when a pipeline misses its expected cadence
~5/10
I see run status, logs, and row counts per sync, and failures alert me in Slack, email, or a webhook
~6/10
Transient failures retry automatically and interrupted syncs resume from checkpoints instead of restarting
—0/10
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Orchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling
Stories about orchestration scheduling in this arena
I run and test a pipeline locally against a lightweight destination before it touches production
~6/10
I see end-to-end lineage of my datasets — which sources, steps, and transformations produced each table
—0/10
I define dependencies between pipeline steps and datasets, and the platform orchestrates runs in the right order
~4/10
Pricing cost — stories about pricing cost in this arenaPricing cost
Stories about pricing cost in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Reverse etl activation — stories about reverse etl activation in this arenaReverse etl activation
Stories about reverse etl activation in this arena
Schema evolution — stories about schema evolution in this arenaSchema evolution
Stories about schema evolution in this arena
Sync replication — stories about sync replication in this arenaSync replication
Stories about sync replication in this arena
Backfill history or resync a single table without rebuilding the whole pipeline
~6/10
I replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the source
✓8/10
Syncs move only new and changed records — cursor and state management handled for me, not full reloads
~7/10
I control sync frequency per pipeline — from sub-hour schedules to cron expressions and manual triggers
✓8/10
Transformations dbt — stories about transformations dbt in this arenaTransformations dbt
Stories about transformations dbt in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed⚿ | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 5/10 | Cclaimed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed⚿ | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed⚿ | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 7/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Xcommunity | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed⚿ | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed⚿ | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ⚿ | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | partial | 5/10 | Tprobed | |
I pick from a broad catalog of maintained connectors for the SaaS APIs, databases, and files my company actually uses C Catalog | data engineer | Connectors catalog — stories about connectors catalog in this arenaConnectors catalog | 3 | full | 9/10 | Xcommunity | |
A coding agent can scaffold, configure, and run a complete pipeline headlessly through the CLI or API C Ai build | ai-native user | Ai pipelines — stories about ai pipelines in this arenaAi pipelines | 3 | full | 8/10 | Tprobed⚿ | |
I replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the source C Cdc | data engineer | Sync replication — stories about sync replication in this arenaSync replication | 3 | full | 8/10 | Xcommunity | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullfree | 8/10 | Xcommunity | |
Syncs move only new and changed records — cursor and state management handled for me, not full reloads C Incremental | data engineer | Sync replication — stories about sync replication in this arenaSync replication | 3 | partial | 7/10 | Xcommunity | |
Upstream schema changes are detected and propagated by a policy I choose, instead of silently breaking loads C Evolution | data engineer | Schema evolution — stories about schema evolution in this arenaSchema evolution | 3 | full | 7/10 | Cclaimed | |
I see run status, logs, and row counts per sync, and failures alert me in Slack, email, or a webhook C Monitoring | data engineer | Observability reliability — stories about observability reliability in this arenaObservability reliability | 3 | partial | 6/10 | Cclaimed | |
AI drafts a working connector from API documentation — auth, pagination, streams — that I review and ship C Ai build | ai-native user | Ai pipelines — stories about ai pipelines in this arenaAi pipelines | 3 | partial | 5/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partialfree | 5/10 | Xcommunity | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | 0/10 | ||
I build a custom connector for a long-tail API with a supported framework or low-code builder, not a fork C Custom connectors | data engineer | Connectors catalog — stories about connectors catalog in this arenaConnectors catalog | 2 | full | 8/10 | Cclaimed | |
I control sync frequency per pipeline — from sub-hour schedules to cron expressions and manual triggers C Scheduling | data engineer | Sync replication — stories about sync replication in this arenaSync replication | 2 | full | 8/10 | Cclaimed | |
Dbt transformations run against freshly loaded data as part of the pipeline, not on a blind timer C Dbt | analytics engineer | Transformations dbt — stories about transformations dbt in this arenaTransformations dbt | 2 | full | 7/10 | Cclaimed | |
I load to the major warehouses and lakes — Snowflake, BigQuery, Databricks, Postgres, object storage — without changing pipelines C Destinations | data engineer | Code first portability — stories about code first portability in this arenaCode first portability | 2 | full | 7/10 | Cclaimed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | full | 7/10 | Tprobed⚿ | |
Backfill history or resync a single table without rebuilding the whole pipeline C Backfill | data engineer | Sync replication — stories about sync replication in this arenaSync replication | 2 | partial | 6/10 | Xcommunity | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 6/10 | Tprobed⚿ | |
I run and test a pipeline locally against a lightweight destination before it touches production C Dev loop | data engineer | Orchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling | 2 | partialfree | 6/10 | Tprobed | |
My pipelines are plain code and config in my own repository — versioned, reviewed, and portable like any software C Code first | data engineer | Code first portability — stories about code first portability in this arenaCode first portability | 2 | partial | 6/10 | Tprobed⚿ | |
An agent can check sync status, diagnose a failed run, and re-trigger it through an API or MCP server C Ai operate | ai-native user | Ai pipelines — stories about ai pipelines in this arenaAi pipelines | 2 | partial | 5/10 | Tprobed⚿ | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 5/10 | Tprobed⚿ | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | disputed | 5/10 | Dcontradicted | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialfree | 4/10 | Cclaimed | |
I define dependencies between pipeline steps and datasets, and the platform orchestrates runs in the right order C Orchestration | data engineer | Orchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling | 2 | partial | 4/10 | Cclaimed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
I see end-to-end lineage of my datasets — which sources, steps, and transformations produced each table C Lineage | data engineer | Orchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling | 2 | none | 0/10 | ||
I sync modeled warehouse data back into SaaS tools (CRM, ads, support) to activate it where teams work C Reverse etl | analytics engineer | Reverse etl activation — stories about reverse etl activation in this arenaReverse etl activation | 2 | none | 0/10 | ||
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
The pricing model is published and predictable — I can estimate what a new source costs before connecting it G Pricing | data platform lead | Pricing cost — stories about pricing cost in this arenaPricing cost | 2 | none | 0/10 | ||
Transient failures retry automatically and interrupted syncs resume from checkpoints instead of restarting C Recovery | data engineer | Observability reliability — stories about observability reliability in this arenaObservability reliability | 2 | none | 0/10 | ||
Loaded data lands as typed, deduplicated destination tables ready to query, not raw JSON blobs C Normalization | analytics engineer | Schema evolution — stories about schema evolution in this arenaSchema evolution | 1 | full | 8/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 6/10 | Cclaimed | |
Pipelines load into vector stores and LLM-ready formats so my agents can retrieve what was synced C Ai destinations | ai-native user | Ai pipelines — stories about ai pipelines in this arenaAi pipelines | 1 | partial | 5/10 | Tprobed | |
Tell how fresh each destination table is and get warned when a pipeline misses its expected cadence C Freshness | analytics engineer | Observability reliability — stories about observability reliability in this arenaObservability reliability | 1 | partial | 5/10 | Cclaimed | |
The catalog tells me each connector's maturity, support level, and maintainer before I depend on it C Catalog | data engineer | Connectors catalog — stories about connectors catalog in this arenaConnectors catalog | 1 | none | 0/10 |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 35 stories with headroom
What would move Airbyte’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
Missing: any evidence of Airbyte consuming external MCP servers, an MCP-client configuration surface, or AI Assistant tool-use extended via third-party MCP servers.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
Missing: any documented AI-training data policy, an opt-out toggle/setting, or a privacy statement addressing model-training use of customer data.
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product
nonemoves Built-in AIimpact 30
Airbyte's AI features (AI Assistant, PyAirbyte, MCP servers) all help configure connectors, move data, or let external agents query data — none of the evidence shows Airbyte itself generating insights, summaries, or suggestions about the user's data inside its own UI.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
Airbyte publishes API documentation (airbyte-docs-3/17/18) but there is no evidence of an interactive reference with runnable/try-it examples; a probe for a standard OpenAPI/Swagger spec at expected paths returned all 404s (airbyte-probe-2), and no docs mention live sandbox or code-execution widgets.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
The evidence shows Airbyte has a documented API (airbyte-docs-3, airbyte-docs-17, airbyte-docs-18) and even a live auth-gated endpoint (airbyte-probe-rt-3), but nothing in the pack describes API versioning conventions or a documented deprecation policy for breaking changes.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
partialq5/10moves Built-in AIimpact 22.5
Missing: evidence of a general-purpose in-app assistant beyond Connector Builder, hands-on/independent validation of the AI Assistant's real-world reliability.
Orchestration scheduling — stories about orchestration scheduling in this arenaI see end-to-end lineage of my datasets — which sources, steps, and transformations produced each table
nonemoves PA Scoreimpact 20
Evidence shows sync scheduling, notifications, dbt integration, and a Connection Timeline of sync events, but nothing describing an actual lineage graph tracing which sources/steps/transformations produced a given table — no lineage UI, OpenLineage/dbt lineage integration, or column-level lineage is documented.
Pricing cost — stories about pricing cost in this arenaThe pricing model is published and predictable — I can estimate what a new source costs before connecting it
nonemoves PA Scoreimpact 20
The evidence pack shows only a generic pricing page listing feature tiers (Multiple Workspaces, SSO, RBAC) with no per-connector or per-source cost breakdown, and no documentation letting a buyer estimate cost before connecting a new source.
Showing the top 8 of 35 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 41 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Platform docs34 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Subscribe to events via webhooks
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Download a machine-readable API spec (OpenAPI or equivalent)
- Test against a sandbox environment without touching production data
- AI drafts a working connector from API documentation — auth, pagination, streams — that I review and ship
- An agent can check sync status, diagnose a failed run, and re-trigger it through an API or MCP server
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Version, review, and roll back my automations
- I load to the major warehouses and lakes — Snowflake, BigQuery, Databricks, Postgres, object storage — without changing pipelines
- I pick from a broad catalog of maintained connectors for the SaaS APIs, databases, and files my company actually uses
- I build a custom connector for a long-tail API with a supported framework or low-code builder, not a fork
- Tell how fresh each destination table is and get warned when a pipeline misses its expected cadence
- I see run status, logs, and row counts per sync, and failures alert me in Slack, email, or a webhook
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- I run and test a pipeline locally against a lightweight destination before it touches production
- I define dependencies between pipeline steps and datasets, and the platform orchestrates runs in the right order
- Choose where my data is stored (region/residency)
- Upstream schema changes are detected and propagated by a policy I choose, instead of silently breaking loads
- Loaded data lands as typed, deduplicated destination tables ready to query, not raw JSON blobs
- Backfill history or resync a single table without rebuilding the whole pipeline
- I replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the source
- Syncs move only new and changed records — cursor and state management handled for me, not full reloads
- I control sync frequency per pipeline — from sub-hour schedules to cron expressions and manual triggers
- Dbt transformations run against freshly loaded data as part of the pipeline, not on a blind timer
Developers docs21 stories
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Set up automations that run autonomously in the background
- Download a machine-readable API spec (OpenAPI or equivalent)
- Test against a sandbox environment without touching production data
- A coding agent can scaffold, configure, and run a complete pipeline headlessly through the CLI or API
- Pipelines load into vector stores and LLM-ready formats so my agents can retrieve what was synced
- An agent can check sync status, diagnose a failed run, and re-trigger it through an API or MCP server
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Version, review, and roll back my automations
- My pipelines are plain code and config in my own repository — versioned, reviewed, and portable like any software
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- I run and test a pipeline locally against a lightweight destination before it touches production
- I define dependencies between pipeline steps and datasets, and the platform orchestrates runs in the right order
- I control sync frequency per pipeline — from sub-hour schedules to cron expressions and manual triggers
Hacker News10 stories
- Set up automations that run autonomously in the background
- Pipelines load into vector stores and LLM-ready formats so my agents can retrieve what was synced
- I pick from a broad catalog of maintained connectors for the SaaS APIs, databases, and files my company actually uses
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- I run and test a pipeline locally against a lightweight destination before it touches production
- Backfill history or resync a single table without rebuilding the whole pipeline
- I replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the source
- Syncs move only new and changed records — cursor and state management handled for me, not full reloads
Community docs7 stories
- Point an agent at llms.txt or agent-oriented docs
- Connect an agent via an official MCP server
- Use an official CLI
- Operate the product with natural-language commands
- A coding agent can scaffold, configure, and run a complete pipeline headlessly through the CLI or API
- Pipelines load into vector stores and LLM-ready formats so my agents can retrieve what was synced
- An agent can check sync status, diagnose a failed run, and re-trigger it through an API or MCP server
GitHub README4 stories
- I load to the major warehouses and lakes — Snowflake, BigQuery, Databricks, Postgres, object storage — without changing pipelines
- I pick from a broad catalog of maintained connectors for the SaaS APIs, databases, and files my company actually uses
- Export all of my data in open formats and leave
- Read the product's source under an open license
OpenAPI spec3 stories
Pricing docs2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -si https://api.airbyte.com/v1/connections | head -4reproduced$ curl -si https://api.airbyte.com/v1/connections | head -4 HTTP/2 401 www-authenticate: Bearer resource_metadata="http://api.airbyte.com/.well-known/oauth-protected-resource/api/public/v1/connections" content-type: application/json date: Tue, 8 Sep 2026 20:41:00 GMT
$uv pip install airbyte && printf '<jsonrpc initialize>' | airbyte-mcp # first-party stdio MCP bundled with PyAirbytereproduced$ uv pip install airbyte && printf '<jsonrpc initialize>' | airbyte-mcp # first-party stdio MCP bundled with PyAirbyte
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-06-18","capabilities":{"experimental":{},"logging":{},"prompts":{"listChanged":false},"resources":{"subscribe":false,"listChanged":false},"tools":{"listChanged":true},"extensions":{"io.modelcontextprotocol/ui":{}}},"serverInfo":{"name":"airbyte-mcp","version":"3.4.7"},"instructions":"PyAirbyte connector management and data integration serve
$mktemp -d && uv venv && uv pip install airbyte && python -c "import airbyte; print('PA_PROBE_OK pyairbyte', version('airbyte'))"reproduced$ mktemp -d && uv venv && uv pip install airbyte && python -c "import airbyte; print('PA_PROBE_OK pyairbyte', version('airbyte'))"
PA_PROBE_OK pyairbyte 0.59.0
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
9 of 16 testable claims verified · 0 contradicted → integrity 56/100
21 distinct capability claims found in Airbyte’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
9
Verified
7
Unverified
0
Contradicted
24
Undersold
Verified (11)
“PyAirbyte lets Python and AI developers use Airbyte programmatically as a library”
“Airbyte configuration can be managed as code via an official Terraform provider”
My pipelines are plain code and config in my own repository — versioned, reviewed, and portable like any softwarepartialproof ↗
“A documented API lets you control Airbyte programmatically, e.g. from orchestrators like Airflow”
Drive the product through a documented public APIfullproof ↗
“MCP-capable agents (Claude, Cursor, VS Code, ChatGPT, Codex) can connect to your data through Airbyte Agents”
“A quickstart lets you deploy a local instance of open-source Airbyte Core in minutes”
“abctl CLI makes it easy to run Airbyte anywhere Docker is available”
“Airbyte can consume database log files to capture INSERT/UPDATE/DELETE changes for replication (CDC)”
I replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the sourcefullproof ↗
“Programmatic access to Airbyte generally requires an access token”
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
“Incremental sync mode pulls only data changed since the previous sync”
Syncs move only new and changed records — cursor and state management handled for me, not full reloadspartialproof ↗
“An installation guide covers deploying Airbyte into any Kubernetes cluster”
“Airbyte provides a catalog of 600+ maintained connectors for APIs, databases, warehouses, and lakes”
I pick from a broad catalog of maintained connectors for the SaaS APIs, databases, and files my company actually usesfullproof ↗
Unverified (8)
“A no-code Connector Builder in the UI lets you develop custom connectors without leaving Airbyte”
I build a custom connector for a long-tail API with a supported framework or low-code builder, not a forkfullproof ↗
“An AI Assistant can auto-prefill fields in the Connector Builder to speed up connector setup”
AI drafts a working connector from API documentation — auth, pagination, streams — that I review and shippartialproof ↗
“Each connection offers a choice of options for how/when a sync is triggered to run”
I control sync frequency per pipeline — from sub-hour schedules to cron expressions and manual triggersfullproof ↗
“Airbyte can send sync notifications to an email address or a webhook”
I see run status, logs, and row counts per sync, and failures alert me in Slack, email, or a webhookpartialproof ↗
“A Connection Timeline shows historical events and activity for a connection”
I see run status, logs, and row counts per sync, and failures alert me in Slack, email, or a webhookpartialproof ↗
“dbt Cloud integration runs dbt transformations automatically right after Airbyte Cloud syncs complete”
Dbt transformations run against freshly loaded data as part of the pipeline, not on a blind timerfullproof ↗
“Per-connection settings let you choose how Airbyte handles source schema changes”
Upstream schema changes are detected and propagated by a policy I choose, instead of silently breaking loadsfullproof ↗
“Data lands with one-to-one stream-to-table mapping, avoiding sub-tables in the warehouse”
Loaded data lands as typed, deduplicated destination tables ready to query, not raw JSON blobsfullproof ↗
Undersold (24)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Set up automations that run autonomously in the backgroundfullproof ↗
Delegate tasks to a built-in AI assistant inside the productpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)partialproof ↗
Test against a sandbox environment without touching production datapartialproof ↗
A coding agent can scaffold, configure, and run a complete pipeline headlessly through the CLI or APIfullproof ↗
Pipelines load into vector stores and LLM-ready formats so my agents can retrieve what was syncedpartialproof ↗
An agent can check sync status, diagnose a failed run, and re-trigger it through an API or MCP serverpartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
I load to the major warehouses and lakes — Snowflake, BigQuery, Databricks, Postgres, object storage — without changing pipelinesfullproof ↗
Tell how fresh each destination table is and get warned when a pipeline misses its expected cadencepartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
I run and test a pipeline locally against a lightweight destination before it touches productionpartialproof ↗
I define dependencies between pipeline steps and datasets, and the platform orchestrates runs in the right orderpartialproof ↗
Choose where my data is stored (region/residency)partialproof ↗
Backfill history or resync a single table without rebuilding the whole pipelinepartialproof ↗
Claims outside our story set (2)
Real capability claims found in Airbyte’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Sync modes define the combined behavior of how a source is read and a destination is written”
source ↗“Airbyte supports multiple workspaces with SSO and RBAC”
source ↗
Business model
Open-source core (ELv2/MIT mix) is free to self-host; Airbyte Cloud is usage-based (capacity/volume-priced syncs) with a trial, and Teams/Enterprise (incl. self-managed) are custom.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt up (tracking since Sep 10 '26)
