trajectory
py-etl-schema-01core-12pythonDifficulty tier 3/5

Load a vendor export whose schema changed, without losing rows

etlschema-evolutiondata-qualitycsv

Task parameters

Reference steps
12
Step ceiling
40
Runs
15
Solved
14 of 15
Models
5

What is broken, and what fixed means

A CSV to warehouse loader falls over on the vendor's May export. Upstream renamed region_code to market, split customer_name into two columns, added an optional discount_cents, and switched the placed date from YYYY-MM-DD to DD/MM/YYYY, all without telling anyone. The loader raises on the first row. The tempting repairs are worse than the crash: a blanket except around the row loop silently drops rows, and mapping the new columns straight through changes the output field set that the warehouse and everything downstream of it are built on. The contract is written down in CONTRACT.md, including four reject reasons of which the loader has only ever implemented one, and April's export has to keep loading because the warehouse reloads history.

Results by model

One group per model. The solve rate carries its spread across seeds, and every run below it links to the full step by step replay.

stub:hasty

66.7%+/- 47.1% over 3 seeds

3 runs, seeds 0, 1, 2

stub:methodical

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:reckless

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:sloppy

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2

stub:thrasher

100.0%+/- 0.0% over 3 seeds

3 runs, seeds 0, 1, 2