MerchantryTidbits

Public catalog

Browse tools

11 shown. Search ranks against problem language; browsing defaults to stronger public signals.

data-wranglingclilocal

Airbyte

replicate data between SaaS APIs, databases, files, and warehouses

Verified 2026-08-10

data-engineeringservicefree

DataHub

build a searchable data catalog with dataset ownership and end to end lineage

Verified 2026-08-10

data-wranglingclifree

dbt Core

manage warehouse SQL transformations as versioned tested models

Verified 2026-08-10

data-engineeringlibraryfree

dlt (data load tool)

Hand-written Python scripts pulling from REST APIs keep breaking on pagination, typing, and upstream schema changes before loading into a warehouse

Verified 2026-08-10

ml-operationsclifree

DVC

version datasets and model artifacts alongside Git commits

Verified 2026-08-10

data-wranglinglibraryfree

Great Expectations

validate data schemas, nulls, ranges, uniqueness, and business rules

Verified 2026-08-10

workflow-automationappfree

Kestra

Scheduled scripts and data pipelines are scattered across servers as cron jobs with no central visibility, retries, or backfill when one fails

Verified 2026-08-10

data-engineeringservicefree

Marquez

collect OpenLineage events and visualize how jobs produce and consume datasets

Verified 2026-08-10

data-engineeringservicefree

OpenMetadata

self host a data catalog with lineage governance quality tests and discovery

Verified 2026-08-10

data-wranglinglibraryfree

Prefect

orchestrate Python data workflows with retries, schedules, caching, and run state

Verified 2026-08-10

data-wranglingclifree

Soda Core

write readable data quality checks against databases and warehouses

Verified 2026-08-10