Make AI-ready data trustworthy.
Build governed pipelines that turn messy operational data into reliable analytics, retrieval, training, and evaluation inputs.

- In planning
- No launch date is promised
- Remote-first
- Proposed delivery format
- Practical proof
- Capstone-led curriculum design
- Interest only
- Not an internship application
Why this belongs on the roadmap.
The World Economic Forum places big-data and data-warehousing roles among the fastest-growing job categories. AI systems deepen the need for tested schemas, lineage, quality checks, and datasets that teams can actually trust.
Research sources support the direction, not a launch date or outcome. This page does not promise a job, salary, client, income, or final curriculum.
The work this curriculum would cover.
Each planned module is designed to produce evidence, moving from foundations to work a real operator can inspect.
- 01Model operational, analytical, retrieval, and evaluation data
- 02Ingest batch and event data with reproducible pipelines
- 03Transform data with tests, contracts, and version control
- 04Track quality, lineage, freshness, privacy, and cost
- 05Prepare documents and metadata for retrieval systems
- 06Design reliable datasets for evaluation and model monitoring
The proof this track would demand.
A production-style data platform for a real business use case: source ingestion, transformations, tests, lineage, quality dashboard, retrieval-ready outputs, an evaluation dataset, and an operator runbook.

The working stack.
Tools can change before launch. The proposed workflow and quality bar are the durable part.
- Python
- SQL
- dbt
- DuckDB
- Airbyte
- Dagster
- Postgres
- Supabase
Register interest in this track.
Join the track-specific waitlist. This is separate from the internship intake waitlist and does not submit an application.