Over the last couple of years, a lot of our work at datocracy has been around a fairly simple question: what makes data and AI capacity building actually stick?
We have delivered responsible data and AI training to learners across different countries and institutional settings, and the more contexts we work in, the less convinced I am that the answer is simply better content. A course can be technically strong and still not work for the people taking it. It can be well received and still have very little effect on what happens afterwards.
That was the question I took into a session at the Global Data Festival in Nairobi earlier this month, on the last morning of the conference at 8AM. The room had Kenyan ministry officials, engineers, health practitioners, educators, funders, and people working on open data governance. I wasn't sure whether one framework could hold a group that varied.
It did. But the more useful outcome was what the discussion brought out about where capacity-building programmes tend to go wrong.
Three things remained with me.
One of the easiest traps in capacity building is designing for an imagined learner.
We tend to build a curriculum first and then think about who might use it. That often means making assumptions about what people already know, what language they work in, what their institutions look like, and what they actually need the training to help them do.
Our work with Qhala during Africa AI Literacy Week in 2025 was a useful example of what happens when you reverse that order. For our training on How to Manage Data and AI Responsibly, the underlying content came from datocracy's broader course on responsible data and AI use, which by then had reached more than 600 learners across 80+ countries. For Qhala, we adapted the course for Africa AI Literacy Week rather than delivering the full programme as designed. We offered participants two focused sessions: one on the core principles of responsible data and AI use, and another on applying generative AI to research and writing. We also designed the sessions for bilingual delivery in English and French from the outset, rather than developing the material in one language and translating it afterwards. These choices were driven by the audience and what we knew they were likely to need, rather than by trying to fit them into a fixed course structure.
None of these are particularly radical decisions. But they come from asking more specific questions about the learner. What do they already know? What are they trying to do? What language do they need? What is useful to them now rather than simply useful to know?
That specificity can feel like a constraint when we are thinking about scale. In practice, I think it is what makes adaptation possible. When we know why something worked for one group, we have something concrete to carry into the next context. But when we design for everyone, we are often designing for an average learner who doesn't actually exist.
AI and data cannot be taught as separate conversations.
The second point came up when we started talking about responsible AI. It is easy to treat AI and data as two separate areas of learning: first teach people how to work with data, then introduce AI, and somewhere along the way add a session on ethics and governance.
But AI is built on data, so the two cannot really be separated. The data an AI system is trained or built on affects what it produces, who it represents, what it misses and where it can fail. The decisions around that data matter too: where it came from, who collected it, who is represented, who is missing, and who has the authority to use it.
This is why the point at which governance is introduced in a curriculum matters. If participants spend four modules learning how to work with a dataset and encounter questions about its quality, provenance or representation only in the fifth, we have already taught them to see those questions as separate from the technical work. A participant who has spent four modules learning to build with a dataset is unlikely to start questioning the dataset in the fifth.
For people working in public institutions and civil society organisations, this distinction matters even more. They may be using AI in research, health, education, public services or policy, where the consequences of poor data or a poorly understood system do not stop with the person using the tool. Responsible AI therefore cannot be something added to the curriculum after the technical content. The questions about what the system is built on, who it serves and who might be harmed by it need to be part of how people learn to work with AI from the beginning.
The real test is what happens afterwards.
It is relatively easy to measure what happens during a programme. How many people attended? How many completed it? Did they find it useful? Much harder is knowing whether anything changed once they went back to their organisations and had to use what they had learned without the facilitator in the room.
The design therefore needs to account for what happens next. Is there someone in the organisation who can take the work forward? Do participants have an opportunity to apply what they learned to a real problem? Is there something they can return to when they get stuck? Are there peers they can learn from after the formal programme ends?
These do not necessarily require a large post-training programme. But they do require thinking beyond the training itself. The conversations happening across the wider festival made this feel particularly relevant. I attended discussions about AI sovereignty across the African continent that looked at who controls the compute, where the data sits, who owns the models, and what happens when critical systems depend on infrastructure somewhere else.
Across all of these, one concern kept coming up: capability built on infrastructure you do not control can ultimately be withdrawn by a decision made somewhere else. That is obviously an infrastructure and sovereignty problem. But it also changes what we should mean by capacity.
If the systems people are being trained to use and govern can change, disappear or become inaccessible, then lasting capacity cannot just mean knowing how to use today's tools. It has to mean building enough understanding and institutional capability to keep making decisions when the tools, infrastructure or circumstances change. With four years left on the SDG clock, the gap between what data systems can do and who has the capacity to govern them is not closing as quickly as the technology is moving.
The workshop didn't give me a new framework so much as reinforce something I have been seeing through our own work at datocracy: capacity building works better when we stop treating training as the end product. The content matters. But so does who it is designed for, how governance is built into it, and what remains once the facilitator is gone. Those are design decisions, and unlike the pace of the technology itself, they are decisions we can still change.
That was what I wanted to test during the session, and the discussion gave me more confidence that these choices can hold across very different contexts. I left Nairobi more energised than I arrived, and grateful to the Global Partnership for Sustainable Development Data for creating the space for the conversation.