NARO Big Data Platform
Laying the groundwork for a national agricultural research data platform across 16 institutes, from field requirements to architecture roadmap and budget
- Role
- Consultant, Data Science & Engineering · led the data analysis workstream
- When
- 2022 – 2025
- Status
- Two rounds of requirements and design; architecture decision pending (on-premise, in-country proposed); paused after donor funding was cut in 2025.
- Stack
- Hadoop
- Spark
- Hive
- Data-lake design
- Survey design & analysis
- Capacity forecasting
- Cost-benefit analysis
Context
The National Agricultural Research Organisation (NARO) coordinates agricultural research across its Secretariat and 16 public research institutes, including Namulonge and Kawanda. Research data, from field trials to institutional databases, lived on personal laptops and external drives, with no shared storage or consistent backup. NARO set out to build a Big Data Centre and analytics platform to serve all 16 institutes.
Role
Consultant, Data Science & Engineering (part-time). I led the data analysis workstream within a 10-person multidisciplinary team of M&E specialists, statisticians and agricultural specialists, and co-authored the original concept note.
What I did
- 2022Round oneConcept note, benchmarking visits, needs assessment across the institutes, analytical report and budget.
- Early 2025Round twoRedone with new funding: more institutions, more detail, better evidence. Cut short when donor funding was withdrawn.
- NextArchitecture decisionOn-premise, in-country hosting proposed; decision pending.
- Requirements from the field. Fieldwork tours and stakeholder workshops across the NARO institutions, working with technicians, researchers, non-technical staff, scientists and leadership to gather requirements and document use cases and process flows.
- Evidence from the surveys. I analysed the needs-assessment surveys (in round one, 62 staff on their data practices and 19 institute ICT infrastructure audits) and wrote the analytical report the budget and requirements were built on.
- Requirements and specifications. Business requirements documents and system specifications for a unified data architecture.
- Architecture roadmap. The target ecosystem is Apache Hadoop, Spark and Hive, ingesting structured and unstructured data (field research datasets, institutional databases), sized for a projected 20TB+ across the Secretariat and the 16 institutes.
- Cost-benefit and deployment. Feasibility studies and cost-benefit analyses of technology stacks and of cloud against on-premise deployment, with forecasts for infrastructure, licensing and maintenance. The 2022 concept note recommended a managed-cloud data centre built around a data lake as a fast start. The fuller 2025 evidence pointed to on-premise, in-country hosting.
- Proposals that secured approval. Project proposals covering requirements, technical approach, resource estimates, timelines and budgets that won stakeholder approval for the initiative.
What the field evidence showed
- had no one responsible for data storage in their unit
- 48 of 62
- backed up research data to an institute server
- 0 of 53
- generated more than 100 GB of data a year each
- 11 of 60
- of those who knew cloud storage agreed with using it
- 36 of 45
- About a third relied on consumer cloud drives for backups.
- Infrastructure gaps included missing or makeshift server rooms, no system administrators at some institutes, low-end workstations and weak connectivity at remote stations.
- Some staff feared losing control of their data once it was centralised, so governance and access rules had to be part of the design from the start.
Outcome
Two rounds of requirements and design are complete, and the evidence moved the deployment recommendation from a managed cloud to on-premise, in-country hosting. The architecture decision is pending; the project paused after donor funding was cut in 2025.