Barlow Kasule

Lab

Experiments and research

Two interactive pieces on the same theme: before trusting a result, check what the data is really saying.

From my MSc thesis

Same data, opposite conclusion

Simpson's Paradox as a check on the quality of integrated data

Loan size against on-time repaymentScatter plot of 72 synthetic SACCO members from two branches, with fitted trend lines.50%60%70%80%90%100%0246810Loan size (UGX millions)Instalments paid on time
Combined trend
−1.9 pts per UGX 1M
Masaka branch
+3.4 pts per UGX 1M
Kampala branch
+3.0 pts per UGX 1M

Combined, each extra UGX 1M borrowed goes with 1.9 points lower on-time repayment. Bigger loans look riskier.

Synthetic data: 72 members, loan size in UGX millions against the share of instalments paid on time.

The idea

When you integrate data from several sources, the combined dataset can tell a story that none of its parts supports. In the demo above, merging two branches' loan books makes bigger loans look riskier. Split the data by branch again and the relationship reverses in both. Nothing is wrong with either branch's data; the combined view mixes a small-loan branch with a large-loan one, and that difference in mix drives the trend.

This reversal is Simpson's Paradox. My MSc thesis at the University of Trento (2018) proposed using it as a quality check after data integration: if a trend in the integrated data reverses inside its subgroups, the integration may be hiding a confounding variable, and conclusions drawn from the combined data can't be trusted as they stand.

How the check works

  1. Pick the outcome you care about (here, on-time repayment) and the features that might explain it (here, loan size).
  2. Fit a simple linear model of the outcome on each feature across the whole integrated dataset.
  3. Split the data by each candidate confounder (here, branch) and fit the same model within every subgroup.
  4. Flag pairs where the direction of the relationship flips and the subgroup result is statistically significant (p ≤ 0.05). Count them as a data-quality signal.

The thesis tested the check on an engineered synthetic dataset and then applied it to a real integrated public dataset, The Washington Post's Fatal Force database. It is a proposed check rather than a validated metric: it was evaluated on one synthetic set and one case study.

Why it matters in practice

The same trap appears whenever data from several branches, products or institutions is combined, which is most of the reporting I build. Before a combined dashboard drives a decision, it is worth checking that its headline trend survives being split along the lines the data was merged on.

Read the thesis (PDF, 330 KB)

Sensor data · 2024 – 2025, revisited 2026

When the data, not the model, is the story

Auditing a Kalman filter for cement kiln bearing temperatures

Bearing temperature: truth, sensor readings and Kalman estimateOne hour of synthetic readings every 30 seconds, with the filter's estimate drawn over the raw sensor.60°70°80°90°0102030405060Minutes
Raw sensor error (RMSE)
1.74 °C
Kalman estimate error
0.56 °C

Filtered error 0.56 °C against 1.74 °C for the raw sensor, 68% lower. 2 reading(s) rejected as spikes.

Synthetic data. On the real kiln data the lesson was different: no filter can recover what the sensors never measured.

The project

From November 2024 to September 2025 I worked as the data scientist on a research project using temperature data from the roller bearings of a cement kiln at a plant in Uganda. Twelve bearings across three kiln tyres are monitored by control-room sensors, and a separate set of reference readings was recorded for comparison. The goal was to estimate the true bearing temperature from the noisy sensors with a Kalman filter.

What I found when I re-checked it

My first evaluation reported a large improvement. Revisiting it in 2026, I found a data-leakage bug: the filter had been updated with the same reference readings it was then scored against, so the result measured the bug, not the filter. With the inputs in the right order, the original filter was only about 2% better than the raw sensors.

So I re-ran the problem properly:

  • The filter fuses the control-room sensor with occasional reference spot checks, estimating both the temperature and each sensor's bias.
  • Its noise settings were fixed in advance, not tuned on the test data.
  • It was scored only on readings it never saw.
  • It had to beat simple baselines: the raw sensor, the sensor with a constant offset correction, and simply carrying forward the last spot check.
ScenarioRaw sensorOffset-correctedLast spot checkKalman filter
Calibrate once, at the start of the night4.67 °C4.41 °C3.56 °C3.89 °C
Spot checks every two hours5.11 °C3.90 °C3.24 °C3.58 °C

Average error (RMSE) against the reference readings, on held-out timestamps. Lower is better.

The real finding

Simply carrying forward the last manual reading beat every method that used the control-room sensors, the Kalman filter included. On that night the sensors added no information beyond the spot checks.

The reason showed up in the data: between 03:00 and 03:30, the reference readings jumped on all twelve bearings at once, by up to 15 °C, while the control-room sensors did not move. A simultaneous jump on every bearing points to a change in how the reference readings were taken, rather than real heating. The right next step is to check the reading procedure and the plant log around that time before investing in more filtering.

It is a small study: one night, 13 readings per bearing. But the lesson carries: evaluate on data the model hasn't seen, compare against simple baselines, and look hard at the data before blaming or crediting the model.