The upstream change nobody mentioned
By Dylan Wolpe
The short version
- The dangerous data failures are the ones that keep passing validation: a unit change, a rounding change, a category renamed, a timezone silently shifted.
- Schema validation checks that data has the shape you declared. It cannot check that the shape still means what it meant last month.
- Structural comparison against the recent past catches this class, because a change in how data is generated shows up as a change in how compressible it is, even when every field remains individually valid.
- It signals that something changed, not what, so it is an alarm that sends a human to look, not an explanation.
The data failures that hurt are not the ones that break the pipeline. A broken pipeline pages someone at 3am and is fixed by breakfast. The expensive ones keep running perfectly, pass every check, and produce numbers that are quietly wrong for a quarter.
What silent drift looks like
Every one of these has happened to somebody, and none of them errors:
- A supplier switches a measurement from one unit to another.
- Timestamps start arriving in local time instead of UTC.
- A currency field changes from cents to units.
- A category is renamed, so the old value drops to zero and a new one appears.
- Rounding changes from two decimals to none.
- A null starts being sent as an empty string, or as zero.
In each case the data is structurally valid. Types are right, ranges are plausible, nothing is missing. Your validation passes because your validation was built to check exactly these things, and they are all still true.
Validation asks whether the data is the shape you declared. It cannot ask whether the shape still means what it meant last month.
Why nobody tells you
Not malice. The supplier made a reasonable internal change and did not know you existed downstream of three intermediate systems. Or they announced it in a release note nobody in your organisation reads. Or the change was made by a team who did not know the field was exported at all.
Data contracts help where you can impose them. Most organisations consume feeds from parties they cannot impose anything on, regulators, exchanges, partners, government sources, and for those the change simply arrives.
Comparison catches what validation cannot
The way to notice is to stop comparing data against a specification and start comparing it against itself, over time.
A change in how data is produced almost always changes its structural regularity, the redundancy, the repetition, the relationships between columns, even when every field remains individually plausible. That is measurable, as set out in compression as measurement: how compactly a batch can be described is a reading, and a reading that moves is a change.
| Change | Passes schema validation? | Changes structure? |
|---|---|---|
| Unit switched | Yes | Yes |
| Rounding reduced | Yes | Yes |
| Category renamed | Yes | Yes |
| Timezone shifted | Yes | Yes |
| Nulls became zeros | Yes | Yes |
Rounding is the clearest illustration. Values with two decimal places have different regularity from whole numbers, measurably so, immediately, from the first batch. No range check will ever notice. A structural comparison notices on day one.
The cheap version
You do not need anything sophisticated to start. Take a fixed-size sample of each feed daily, compress it with the same settings, and record the ratio. Plot it. Alarm on a departure from the recent band.
That is a few lines of code and a metric, and it catches an entire class of failure that no amount of schema validation reaches, because it is asking a question validation cannot express. The organisations that get burned by silent drift are rarely the ones without validation. They are the ones for whom validation was the only question being asked.
Questions people ask about this
Why doesn't schema validation catch upstream changes?
Because validation tests structure, not meaning. If a supplier starts sending a temperature in Fahrenheit instead of Celsius, the field is still a number in range and every check passes. The data is valid and wrong, and validation has no concept of the second thing.
How can you detect a change in data you were not told about?
By comparing today's data structurally against the recent past rather than against a specification. A change in how data is produced almost always changes its statistical and structural regularity, which is measurable even when every individual value is plausible.
What does this actually alert on?
That the structure of a feed has shifted relative to its own baseline. It does not identify the cause. The correct response is a person spending twenty minutes comparing samples, which is a far better outcome than discovering it in a quarterly report.