01
Data Engineering
Pipelines, warehouses, and lakehouses on Snowflake, BigQuery, Databricks, and AWS. Airflow and dbt with tests, CI, and lineage, so a broken source fails loudly at 3am instead of quietly in next quarter's board deck. We also do the unglamorous half: migrating off legacy ETL and cutting warehouse spend.
- Batch and streaming pipelines with Airflow, Spark, Kafka, and AWS Lambda, Step Functions, Glue, and EMR
- Warehouse and lakehouse design on Snowflake, PostgreSQL, Databricks, and S3 with Iceberg tables
- dbt and Snowpark for transformations, with tests, documentation, lineage, and CI on every pull request
- Migrations we have run more than once: Hadoop to AWS, legacy ETL and SSIS to Matillion and Snowflake
- Database development where it still matters: Oracle PL/SQL, PL/pgSQL, T-SQL, query optimisation
- Cost work: Glue to Lambda where the workload is small, warehouse sizing, partitioning, file formats