Skip to content

Rank #1 of 5 in Data Pipelines & ELT

dlt logo

dlt

Open Source

dltHub (Scale Vector GmbH)

5.8k1.3k/yrpypi 1.1M/wkpypi/wk ±0

Access

Install

pippip install dlt
uvxuvx dlthub-start@latest

Compare head-to-head

Alternatives to dlt

Showcase

dlt homepage screenshot
homepage · captured Sep 2026 · view live ↗
dlt docs screenshot
docs · captured Sep 2026 · view live ↗

Try itExperimental

See what an agent can do with dlt before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$uvx dlt --versionrecorded session — replayed, not live
recorded 2026-09-08 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

43.7/100

Ai pipelines — stories about ai pipelines in this arenaAi pipelinesevidence →

Stories about ai pipelines in this arena

65.3/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

37.8/100

Code first portability — stories about code first portability in this arenaCode first portabilityevidence →

Stories about code first portability in this arena

80.0/100

Connectors catalog — stories about connectors catalog in this arenaConnectors catalogevidence →

Stories about connectors catalog in this arena

69.7/100

Observability reliability — stories about observability reliability in this arenaObservability reliabilityevidence →

Stories about observability reliability in this arena

21.0/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

78.0/100

Orchestration scheduling — stories about orchestration scheduling in this arenaOrchestration schedulingevidence →

Stories about orchestration scheduling in this arena

46.0/100

Pricing cost — stories about pricing cost in this arenaPricing costevidence →

Stories about pricing cost in this arena

18.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

10.7/100

Reverse etl activation — stories about reverse etl activation in this arenaReverse etl activationevidence →

Stories about reverse etl activation in this arena

18.0/100

Schema evolution — stories about schema evolution in this arenaSchema evolutionevidence →

Stories about schema evolution in this arena

75.0/100

Sync replication — stories about sync replication in this arenaSync replicationevidence →

Stories about sync replication in this arena

44.6/100

Transformations dbt — stories about transformations dbt in this arenaTransformations dbtevidence →

Stories about transformations dbt in this arena

70.0/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 5 free · 0 paid · 0 enterprise · 38 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10T

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3none0/10

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/a0/10

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial7/10T

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full7/10C

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial3/10T

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1partial6/10T

A coding agent can scaffold, configure, and run a complete pipeline headlessly through the CLI or API C

Ai build

ai-native userAi pipelines — stories about ai pipelines in this arenaAi pipelines3full9/10T

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3fullfree9/10T

Syncs move only new and changed records — cursor and state management handled for me, not full reloads C

Incremental

data engineerSync replication — stories about sync replication in this arenaSync replication3full9/10C

I pick from a broad catalog of maintained connectors for the SaaS APIs, databases, and files my company actually uses C

Catalog

data engineerConnectors catalog — stories about connectors catalog in this arenaConnectors catalog3full8/10T

AI drafts a working connector from API documentation — auth, pagination, streams — that I review and ship C

Ai build

ai-native userAi pipelines — stories about ai pipelines in this arenaAi pipelines3full7/10T

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3fullfree7/10C

Upstream schema changes are detected and propagated by a policy I choose, instead of silently breaking loads C

Evolution

data engineerSchema evolution — stories about schema evolution in this arenaSchema evolution3full7/10X

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial4/10C

I see run status, logs, and row counts per sync, and failures alert me in Slack, email, or a webhook C

Monitoring

data engineerObservability reliability — stories about observability reliability in this arenaObservability reliability3partial3/10C

I replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the source C

Cdc

data engineerSync replication — stories about sync replication in this arenaSync replication3none0/10

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

I run and test a pipeline locally against a lightweight destination before it touches production C

Dev loop

data engineerOrchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling2fullfree9/10T

My pipelines are plain code and config in my own repository — versioned, reviewed, and portable like any software C

Code first

data engineerCode first portability — stories about code first portability in this arenaCode first portability2full9/10T

I build a custom connector for a long-tail API with a supported framework or low-code builder, not a fork C

Custom connectors

data engineerConnectors catalog — stories about connectors catalog in this arenaConnectors catalog2full8/10T

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2fullfree8/10T

Backfill history or resync a single table without rebuilding the whole pipeline C

Backfill

data engineerSync replication — stories about sync replication in this arenaSync replication2full7/10C

Dbt transformations run against freshly loaded data as part of the pipeline, not on a blind timer C

Dbt

analytics engineerTransformations dbt — stories about transformations dbt in this arenaTransformations dbt2full7/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2full7/10T

I load to the major warehouses and lakes — Snowflake, BigQuery, Databricks, Postgres, object storage — without changing pipelines C

Destinations

data engineerCode first portability — stories about code first portability in this arenaCode first portability2full7/10T

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2full7/10T

An agent can check sync status, diagnose a failed run, and re-trigger it through an API or MCP server C

Ai operate

ai-native userAi pipelines — stories about ai pipelines in this arenaAi pipelines2partial6/10T

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial6/10C

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partial5/10C

I define dependencies between pipeline steps and datasets, and the platform orchestrates runs in the right order C

Orchestration

data engineerOrchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling2partial4/10C

I see end-to-end lineage of my datasets — which sources, steps, and transformations produced each table C

Lineage

data engineerOrchestration scheduling — stories about orchestration scheduling in this arenaOrchestration scheduling2partial4/10C

Transient failures retry automatically and interrupted syncs resume from checkpoints instead of restarting C

Recovery

data engineerObservability reliability — stories about observability reliability in this arenaObservability reliability2partial4/10C

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partial3/10C

I control sync frequency per pipeline — from sub-hour schedules to cron expressions and manual triggers C

Scheduling

data engineerSync replication — stories about sync replication in this arenaSync replication2partial3/10C

I sync modeled warehouse data back into SaaS tools (CRM, ads, support) to activate it where teams work C

Reverse etl

analytics engineerReverse etl activation — stories about reverse etl activation in this arenaReverse etl activation2partial3/10C

The pricing model is published and predictable — I can estimate what a new source costs before connecting it G

Pricing

data platform leadPricing cost — stories about pricing cost in this arenaPricing cost2partialfree3/10C

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Loaded data lands as typed, deduplicated destination tables ready to query, not raw JSON blobs C

Normalization

analytics engineerSchema evolution — stories about schema evolution in this arenaSchema evolution1full9/10T

Pipelines load into vector stores and LLM-ready formats so my agents can retrieve what was synced C

Ai destinations

ai-native userAi pipelines — stories about ai pipelines in this arenaAi pipelines1partial6/10C

Tell how fresh each destination table is and get warned when a pipeline misses its expected cadence C

Freshness

analytics engineerObservability reliability — stories about observability reliability in this arenaObservability reliability1partial4/10C

The catalog tells me each connector's maturity, support level, and maintainer before I depend on it C

Catalog

data engineerConnectors catalog — stories about connectors catalog in this arenaConnectors catalog1partial3/10C

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1partial3/10C

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 28 stories with headroom

What would move dlt’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product

    nonemoves Built-in AIimpact 45

    dlt provides tooling (AI Harness, MCP server, context files) that lets *external* coding agents like Claude Code, Cursor, or Codex learn to build dlt pipelines — but there is no evidence of a built-in AI assistant embedded inside dlt itself that a user delegates tasks to.

  2. Sync replication — stories about sync replication in this arenaI replicate databases with log-based CDC (binlog/WAL) so I capture updates and deletes without hammering the source

    nonemoves PA Scoreimpact 30

    dlt's SQL database source documentation only covers SQLAlchemy-based batch extraction with incremental cursor fields and merge/upsert loading (dlt-docs-18, dlt-docs-24, dlt-docs-27), with no mention of binlog/WAL-based log CDC, Debezium integration, or any low-impact replication mechanism for capturing deletes without polling.

  3. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    No evidence in the pack addresses any policy, setting, or guarantee about user data being excluded from AI model training — despite dlt/dltHub featuring AI agent integrations (dlt-docs-13, dlt-docs-14, dlt-docs-17) that could plausibly raise this question, there is no documented opt-out or training-data policy.

  4. Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product

    nonemoves Built-in AIimpact 30

    dlt's AI-related features (AI Harness, MCP server, coding-agent skills) are documented as helping build/deploy/operate data pipelines, not as generating insights or suggestions from the data content itself once loaded.

  5. Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent

    nonemoves agent-readyimpact 30

    dlt manages secrets/config via .dlt/secrets.toml for pipeline credentials, and it has AI-harness/MCP integrations for coding agents, but there is no evidence of a feature to issue scoped or least-privilege API credentials specifically for an agent's use — this remains a plausible ask for a platform coordinating agent-driven pipeline access, but it is unaddressed in the evidence.

  6. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    The evidence pack describes dlt as a data extraction/loading library and dltHub as a pipeline deployment/monitoring platform, but nowhere mentions webhook-based event subscriptions, notifications, or an event system for external consumers.

  7. Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples

    nonemoves API qualityimpact 30

    dlt's docs provide static code tutorials with runnable snippets (e.g.

  8. Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy

    nonemoves API qualityimpact 30

    The evidence pack shows dlt has version numbers (e.g., dlt 1.30.0) and extensive feature docs, but nothing documents an explicit API versioning scheme or deprecation policy for AI-native consumers to rely on.

Showing the top 8 of 28 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map8 surfaces · 43 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

docs40 stories

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$uvx dlt --versionreproduced
$ uvx dlt --version
dlt 1.30.0
proves: Use an official CLIrecorded 2026-09-08
$mktemp -d && uvx --with 'dlt[duckdb]' dlt init chess duckdb && ls && head chess_pipeline.py # real pipeline scaffold, no credentialsreproduced
$ mktemp -d && uvx --with 'dlt[duckdb]' dlt init chess duckdb && ls && head chess_pipeline.py  # real pipeline scaffold, no credentials
Creating and configuring a new pipeline with the verified source chess (A source loading player profiles and games from chess.com api)

Verified source chess was added to your project!
* See the usage examples and code snippets to copy from chess_pipeline.py
* Add credentials for duckdb and other secrets to ./.dlt/secrets.toml
* requirements.txt was created. Install it with:
pip3 install -r requirements.txt
* Read https://dlthub.com/docs/walkthroughs/create-a-pipeline for more information
--- files ---
.
..
.dlt
.gitignore
chess
chess_pipeline.py
requirements.txt
import dlt
from chess import source
$printf '<jsonrpc initialize>' | uv run --with 'dlt-mcp[duckdb]' dlt-mcp # official dltHub MCP, stdio, keylessreproduced
$ printf '<jsonrpc initialize>' | uv run --with 'dlt-mcp[duckdb]' dlt-mcp  # official dltHub MCP, stdio, [redacted]less
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-06-18","capabilities":{"logging":{},"prompts":{"listChanged":false},"resources":{"subscribe":false,"listChanged":false},"tools":{"listChanged":true}},"serverInfo":{"name":"dlt MCP","version":"4.0.3"},"instructions":"Helps you build with the dlt Python library."}}

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

13 of 21 testable claims verified · 0 contradictedintegrity 62/100

29 distinct capability claims found in dlt’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

13

Verified

8

Unverified

0

Contradicted

22

Undersold

Verified (23)
Unverified (9)
Undersold (22)
Claims outside our story set (4)

Real capability claims found in dlt’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Infers schemas and data types, normalizes data, and handles nested structures automatically

    source ↗
  • Lets you query and transform loaded data in Python/SQL and inspect/visualize it in Marimo notebooks

    source ↗
  • Query loaded datasets with dataframe expressions, Ibis, or SQL and read results as records, Pandas, or Arrow

    source ↗
  • Migration path to dltHub is included

    source ↗
Suggest a story for these →

Business model

open-sourcefree-tierusage-basedenterprise-custom

The dlt Python library is Apache-2.0 open source and free; dltHub (managed platform, AI Harness, Playground workspace) is commercially licensed with usage-based plans and custom enterprise terms.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Scoretracked since Sep 8 '26 — no movement recorded yet
Agent-readytracked since Sep 8 '26 — no movement recorded yet

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime llms.txt up (tracking since Sep 10 '26)