Public catalog
Browse tools
11 shown. Search ranks against problem language; browsing defaults to stronger public signals.
Airbyte
replicate data between SaaS APIs, databases, files, and warehouses
Verified 2026-08-10
DataHub
build a searchable data catalog with dataset ownership and end to end lineage
Verified 2026-08-10
dbt Core
manage warehouse SQL transformations as versioned tested models
Verified 2026-08-10
dlt (data load tool)
Hand-written Python scripts pulling from REST APIs keep breaking on pagination, typing, and upstream schema changes before loading into a warehouse
Verified 2026-08-10
DVC
version datasets and model artifacts alongside Git commits
Verified 2026-08-10
Great Expectations
validate data schemas, nulls, ranges, uniqueness, and business rules
Verified 2026-08-10
Kestra
Scheduled scripts and data pipelines are scattered across servers as cron jobs with no central visibility, retries, or backfill when one fails
Verified 2026-08-10
Marquez
collect OpenLineage events and visualize how jobs produce and consume datasets
Verified 2026-08-10
OpenMetadata
self host a data catalog with lineage governance quality tests and discovery
Verified 2026-08-10
Prefect
orchestrate Python data workflows with retries, schedules, caching, and run state
Verified 2026-08-10
Soda Core
write readable data quality checks against databases and warehouses
Verified 2026-08-10