[{"data":1,"prerenderedAt":235},["ShallowReactive",2],{"site":3,"lab":21},{"id":4,"availability":5,"cvEnabled":6,"cvPath":7,"email":8,"extension":9,"github":10,"hmtUrl":11,"intro":12,"linkedin":13,"location":14,"meta":15,"name":16,"positioning":17,"roleLabel":18,"stem":19,"__hash__":20},"site\u002Fdata\u002Fsite.yml","Based in Kampala. Available for projects in Uganda and remotely.",false,"\u002Fcv\u002Fbarlow-kasule-cv.pdf","Barlow.kasule@outlook.com","yml","https:\u002F\u002Fgithub.com\u002Fkasbaros","","I lead the data and AI function at FutureLink Technologies, a Kampala fintech, and I have built web systems since 2018 for fintechs, national research and statistics bodies, schools and private companies. You get one engineer who can design, build and run both the product and the data behind it.","https:\u002F\u002Fwww.linkedin.com\u002Fin\u002Fkasulebarlow1989\u002F","Kampala, Uganda",{},"Barlow Kasule","I build the whole system, from the web app and mobile-money integration to the data pipelines and ML models behind it.","Data & software consultant","data\u002Fsite","wyvEemA03cYJSTnIor0Et6BCchMTuA22lW5uyID8iJ4",[22,99],{"id":23,"title":24,"body":25,"description":88,"extension":89,"kicker":90,"meta":91,"navigation":92,"order":93,"path":94,"seo":95,"stem":96,"subtitle":97,"__hash__":98},"lab\u002Flab\u002Fsimpsons-paradox.md","Same data, opposite conclusion",{"type":26,"value":27,"toc":82},"minimark",[28,33,37,40,44,60,68,72,75],[29,30,32],"h2",{"id":31},"the-idea","The idea",[34,35,36],"p",{},"When you integrate data from several sources, the combined dataset can tell a story that none of its parts supports. In the demo above, merging two branches' loan books makes bigger loans look riskier. Split the data by branch again and the relationship reverses in both. Nothing is wrong with either branch's data; the combined view mixes a small-loan branch with a large-loan one, and that difference in mix drives the trend.",[34,38,39],{},"This reversal is Simpson's Paradox. My MSc thesis at the University of Trento (2018) proposed using it as a quality check after data integration: if a trend in the integrated data reverses inside its subgroups, the integration may be hiding a confounding variable, and conclusions drawn from the combined data can't be trusted as they stand.",[29,41,43],{"id":42},"how-the-check-works","How the check works",[45,46,47,51,54,57],"ol",{},[48,49,50],"li",{},"Pick the outcome you care about (here, on-time repayment) and the features that might explain it (here, loan size).",[48,52,53],{},"Fit a simple linear model of the outcome on each feature across the whole integrated dataset.",[48,55,56],{},"Split the data by each candidate confounder (here, branch) and fit the same model within every subgroup.",[48,58,59],{},"Flag pairs where the direction of the relationship flips and the subgroup result is statistically significant (p ≤ 0.05). Count them as a data-quality signal.",[34,61,62,63,67],{},"The thesis tested the check on an engineered synthetic dataset and then applied it to a real integrated public dataset, The Washington Post's ",[64,65,66],"em",{},"Fatal Force"," database. It is a proposed check rather than a validated metric: it was evaluated on one synthetic set and one case study.",[29,69,71],{"id":70},"why-it-matters-in-practice","Why it matters in practice",[34,73,74],{},"The same trap appears whenever data from several branches, products or institutions is combined, which is most of the reporting I build. Before a combined dashboard drives a decision, it is worth checking that its headline trend survives being split along the lines the data was merged on.",[34,76,77],{},[78,79,81],"a",{"href":80},"\u002Fthesis\u002Fbarlow-kasule-msc-thesis-2018.pdf","Read the thesis (PDF, 330 KB)",{"title":11,"searchDepth":83,"depth":83,"links":84},2,[85,86,87],{"id":31,"depth":83,"text":32},{"id":42,"depth":83,"text":43},{"id":70,"depth":83,"text":71},"An interactive demo of Simpson's Paradox on synthetic SACCO loan data, and the MSc thesis that proposed using the paradox to check integrated data.","md","From my MSc thesis",{},true,1,"\u002Flab\u002Fsimpsons-paradox",{"title":24,"description":88},"lab\u002Fsimpsons-paradox","Simpson's Paradox as a check on the quality of integrated data","3ejqioEjyJ20iWh0VzpVMzjRysU6Gt7i58x--hwAtgw",{"id":100,"title":101,"body":102,"description":227,"extension":89,"kicker":228,"meta":229,"navigation":92,"order":83,"path":230,"seo":231,"stem":232,"subtitle":233,"__hash__":234},"lab\u002Flab\u002Fkiln-kalman.md","When the data, not the model, is the story",{"type":26,"value":103,"toc":222},[104,108,111,115,118,121,136,204,209,213,216,219],[29,105,107],{"id":106},"the-project","The project",[34,109,110],{},"From November 2024 to September 2025 I worked as the data scientist on a research project using temperature data from the roller bearings of a cement kiln at a plant in Uganda. Twelve bearings across three kiln tyres are monitored by control-room sensors, and a separate set of reference readings was recorded for comparison. The goal was to estimate the true bearing temperature from the noisy sensors with a Kalman filter.",[29,112,114],{"id":113},"what-i-found-when-i-re-checked-it","What I found when I re-checked it",[34,116,117],{},"My first evaluation reported a large improvement. Revisiting it in 2026, I found a data-leakage bug: the filter had been updated with the same reference readings it was then scored against, so the result measured the bug, not the filter. With the inputs in the right order, the original filter was only about 2% better than the raw sensors.",[34,119,120],{},"So I re-ran the problem properly:",[122,123,124,127,130,133],"ul",{},[48,125,126],{},"The filter fuses the control-room sensor with occasional reference spot checks, estimating both the temperature and each sensor's bias.",[48,128,129],{},"Its noise settings were fixed in advance, not tuned on the test data.",[48,131,132],{},"It was scored only on readings it never saw.",[48,134,135],{},"It had to beat simple baselines: the raw sensor, the sensor with a constant offset correction, and simply carrying forward the last spot check.",[137,138,139,161],"table",{},[140,141,142],"thead",{},[143,144,145,149,152,155,158],"tr",{},[146,147,148],"th",{},"Scenario",[146,150,151],{},"Raw sensor",[146,153,154],{},"Offset-corrected",[146,156,157],{},"Last spot check",[146,159,160],{},"Kalman filter",[162,163,164,185],"tbody",{},[143,165,166,170,173,176,182],{},[167,168,169],"td",{},"Calibrate once, at the start of the night",[167,171,172],{},"4.67 °C",[167,174,175],{},"4.41 °C",[167,177,178],{},[179,180,181],"strong",{},"3.56 °C",[167,183,184],{},"3.89 °C",[143,186,187,190,193,196,201],{},[167,188,189],{},"Spot checks every two hours",[167,191,192],{},"5.11 °C",[167,194,195],{},"3.90 °C",[167,197,198],{},[179,199,200],{},"3.24 °C",[167,202,203],{},"3.58 °C",[34,205,206],{},[64,207,208],{},"Average error (RMSE) against the reference readings, on held-out timestamps. Lower is better.",[29,210,212],{"id":211},"the-real-finding","The real finding",[34,214,215],{},"Simply carrying forward the last manual reading beat every method that used the control-room sensors, the Kalman filter included. On that night the sensors added no information beyond the spot checks.",[34,217,218],{},"The reason showed up in the data: between 03:00 and 03:30, the reference readings jumped on all twelve bearings at once, by up to 15 °C, while the control-room sensors did not move. A simultaneous jump on every bearing points to a change in how the reference readings were taken, rather than real heating. The right next step is to check the reading procedure and the plant log around that time before investing in more filtering.",[34,220,221],{},"It is a small study: one night, 13 readings per bearing. But the lesson carries: evaluate on data the model hasn't seen, compare against simple baselines, and look hard at the data before blaming or crediting the model.",{"title":11,"searchDepth":83,"depth":83,"links":223},[224,225,226],{"id":106,"depth":83,"text":107},{"id":113,"depth":83,"text":114},{"id":211,"depth":83,"text":212},"Re-evaluating a Kalman filter for kiln bearing sensors with held-out data and simple baselines, and finding that the sensors, not the filter, were the problem.","Sensor data · 2024 – 2025, revisited 2026",{},"\u002Flab\u002Fkiln-kalman",{"title":101,"description":227},"lab\u002Fkiln-kalman","Auditing a Kalman filter for cement kiln bearing temperatures","4dDlhhwmnKrC9F4HTay9xCsnEP7uPjjUR2jEcGfiIqc",1790854436823]