Lab
Experiments and research
Two interactive pieces on the same theme: before trusting a result, check what the data is really saying.
From my MSc thesis
Same data, opposite conclusion
Simpson's Paradox as a check on the quality of integrated data
- Combined trend
- −1.9 pts per UGX 1M
- Masaka branch
- +3.4 pts per UGX 1M
- Kampala branch
- +3.0 pts per UGX 1M
Combined, each extra UGX 1M borrowed goes with 1.9 points lower on-time repayment. Bigger loans look riskier.
Synthetic data: 72 members, loan size in UGX millions against the share of instalments paid on time.
The idea
When you integrate data from several sources, the combined dataset can tell a story that none of its parts supports. In the demo above, merging two branches' loan books makes bigger loans look riskier. Split the data by branch again and the relationship reverses in both. Nothing is wrong with either branch's data; the combined view mixes a small-loan branch with a large-loan one, and that difference in mix drives the trend.
This reversal is Simpson's Paradox. My MSc thesis at the University of Trento (2018) proposed using it as a quality check after data integration: if a trend in the integrated data reverses inside its subgroups, the integration may be hiding a confounding variable, and conclusions drawn from the combined data can't be trusted as they stand.
How the check works
- Pick the outcome you care about (here, on-time repayment) and the features that might explain it (here, loan size).
- Fit a simple linear model of the outcome on each feature across the whole integrated dataset.
- Split the data by each candidate confounder (here, branch) and fit the same model within every subgroup.
- Flag pairs where the direction of the relationship flips and the subgroup result is statistically significant (p ≤ 0.05). Count them as a data-quality signal.
The thesis tested the check on an engineered synthetic dataset and then applied it to a real integrated public dataset, The Washington Post's Fatal Force database. It is a proposed check rather than a validated metric: it was evaluated on one synthetic set and one case study.
Why it matters in practice
The same trap appears whenever data from several branches, products or institutions is combined, which is most of the reporting I build. Before a combined dashboard drives a decision, it is worth checking that its headline trend survives being split along the lines the data was merged on.
Sensor data · 2024 – 2025, revisited 2026
When the data, not the model, is the story
Auditing a Kalman filter for cement kiln bearing temperatures
- Raw sensor error (RMSE)
- 1.74 °C
- Kalman estimate error
- 0.56 °C
Filtered error 0.56 °C against 1.74 °C for the raw sensor, 68% lower. 2 reading(s) rejected as spikes.
Synthetic data. On the real kiln data the lesson was different: no filter can recover what the sensors never measured.
The project
From November 2024 to September 2025 I worked as the data scientist on a research project using temperature data from the roller bearings of a cement kiln at a plant in Uganda. Twelve bearings across three kiln tyres are monitored by control-room sensors, and a separate set of reference readings was recorded for comparison. The goal was to estimate the true bearing temperature from the noisy sensors with a Kalman filter.
What I found when I re-checked it
My first evaluation reported a large improvement. Revisiting it in 2026, I found a data-leakage bug: the filter had been updated with the same reference readings it was then scored against, so the result measured the bug, not the filter. With the inputs in the right order, the original filter was only about 2% better than the raw sensors.
So I re-ran the problem properly:
- The filter fuses the control-room sensor with occasional reference spot checks, estimating both the temperature and each sensor's bias.
- Its noise settings were fixed in advance, not tuned on the test data.
- It was scored only on readings it never saw.
- It had to beat simple baselines: the raw sensor, the sensor with a constant offset correction, and simply carrying forward the last spot check.
| Scenario | Raw sensor | Offset-corrected | Last spot check | Kalman filter |
|---|---|---|---|---|
| Calibrate once, at the start of the night | 4.67 °C | 4.41 °C | 3.56 °C | 3.89 °C |
| Spot checks every two hours | 5.11 °C | 3.90 °C | 3.24 °C | 3.58 °C |
Average error (RMSE) against the reference readings, on held-out timestamps. Lower is better.
The real finding
Simply carrying forward the last manual reading beat every method that used the control-room sensors, the Kalman filter included. On that night the sensors added no information beyond the spot checks.
The reason showed up in the data: between 03:00 and 03:30, the reference readings jumped on all twelve bearings at once, by up to 15 °C, while the control-room sensors did not move. A simultaneous jump on every bearing points to a change in how the reference readings were taken, rather than real heating. The right next step is to check the reading procedure and the plant log around that time before investing in more filtering.
It is a small study: one night, 13 readings per bearing. But the lesson carries: evaluate on data the model hasn't seen, compare against simple baselines, and look hard at the data before blaming or crediting the model.