Dosu LogoDosu Logo
Ask
Join our Discord
elviskahoroPublic
elviskahoro
Documentselviskahoro
advanced-course
advanced-course
Type
External
Status
Published
Created
Mar 3, 2026
Updated
Mar 3, 2026
Source
docs/website/docs/tutorial/advanced-course.md

dlt Advanced Course#

In this course, you'll go far beyond the basics. You’ll build production-grade data pipelines with custom implementations, advanced patterns, and performance optimizations.

Lessons#

Lesson 1: Custom Sources – REST APIs & RESTClient Open in molab Open In Colab GitHub badge#

Learn how to build flexible REST API connectors from scratch using @dlt.resource and the powerful RESTClient.

Lesson 2: Custom Sources – SQL Databases Open in molab Open In Colab GitHub badge#

Connect to any SQL-compatible database, reflect table schemas, write query adapters, and selectively ingest data using sql_database.

Lesson 3: Custom Sources – Filesystems & Cloud Storage Open in molab Open In Colab GitHub badge#

Build sources that read from local or remote files (S3, GCS, Azure).

Lesson 4: Custom Destinations – Reverse ETL Open in molab Open In Colab GitHub badge#

Use @dlt.destination to send data back to APIs like Notion, Slack, or Airtable. Learn batching, retries, and idempotent patterns.

Lesson 5: Transforming Data Before & After Load Open in molab Open In Colab GitHub badge#

Learn when and how to apply add_map, add_filter, @dlt.transformer, or even post-load transformations via SQL or Ibis. Control exactly how your data looks.

Lesson 6: Write Disposition Strategies & Advanced Tricks Open in molab Open In Colab GitHub badge#

Understand how to use replace and merge, and combine them with schema hints and incremental loading.

Lesson 7: Data Contracts Open in molab Open In Colab GitHub badge#

Define expectations on schema, enforce data types and behaviors, and lock down your schema evolution. Ensure reliable downstream use of your data.

Lesson 8: Logging & Tracing Open in molab Open In Colab GitHub badge#

Track every step of your pipeline: from extraction to load. Use logs, traces, and metadata to debug and analyze performance.

Lesson 9: Performance Optimization Open in molab Open In Colab GitHub badge#

Handle large datasets, tune buffer sizes, parallelize resource extraction, optimize memory usage, and reduce pipeline runtime.

Homework & Certification#

You’ve finished the dlt Advanced Course — well done! Test your skills with the Advanced Certification Homework.

Documents
1.12.1
1.13-1.14
1.15
1.16
1.17
1.18
1.19
1.21.2
AGENTS
CLAUDE
CONTRIBUTING
README
README
README
README
README
_book-onboarding-call
_source-info-header
add-incremental-configuration
add-map
add_credentials
adjust-a-schema
advanced
advanced
advanced
advanced
advanced-course
advanced-state
airtable
alerting
amazon_kinesis
arrow-pandas
asana
athena
basic
bigquery
build-a-pipeline-tutorial
chess
clickhouse
command-line-interface
command-line-interface
community-destinations
complex_types
configuration
create-a-pipeline
create-new-destination
csv
currency_conversion_data_enrichment
cursor
data-quality
data-quality-dashboard
data-quality-lifecycle
database-connector-app
databricks
dataset
datasets
dbt
dbt-transformations
dbt_cloud
delta
delta
delta-iceberg
deploy-with-airflow-composer
deploy-with-dagster
deploy-with-github-actions
deploy-with-google-cloud-functions
deploy-with-google-cloud-run
deploy-with-kestra
deploy-with-modal
deploy-with-orchestra
deploy-with-prefect
destination
destination
destination-tables
dispatch-to-multiple-tables
dremio
duckdb
ducklake
education
encryption
fabric
facebook_ads
filesystem
filesystem
frequently-asked-questions
freshdesk
full-loading
fundamentals-course
general_usage
github
glossary
google_ads
google_analytics
google_sheets
how-dlt-works
hubspot
ibis-backend
iceberg
iceberg
iceberg
inbox
incremental-loading
index
index
index
index
index
index
index
index
index
index
init
insert-format
installation
installation
intro
jira
jsonl
kafka
lag
lancedb
llm-native-workflow
load-data-from-an-api
marimo
matomo
merge-loading
mongodb
monitoring
motherduck
ms-sql
mssql
mux
naming-convention
notion
openapi-generator
overview
overview
parquet
performance
personio
pg_replication
pipedrive
pipeline
postgres
profiles-dlthub
pseudonymizing_columns
python
qdrant
redshift
removing_columns
renaming_columns
resource
rest-api
run-a-pipeline
run-in-snowflake
running
runtime-tutorial
salesforce
schema
schema-contracts
schema-evolution
scrapy
setup
setup
share-a-dataset
shopify
slack
snowflake
snowflake_plus
source
sql
sql-client
sql-database
sqlalchemy
staging
state
strapi
stripe
synapse
telemetry
tracing
tracing
troubleshooting
troubleshooting
url-parser-data-enrichment
usage
user_agent_device_data_enrichment
vaults
view-dlt-schema
weaviate
workable
working_with_schemas
zendesk
zendesk-weaviate