Part of the Next Chapter of Evaluation Leadership Series
Why outcomes expertise increasingly needs to be paired with delivery fluency, AI literacy, governance, and cross-functional leadership
Across international development, senior evaluation roles are beginning to ask for a different combination of capabilities. One recent position description made that shift especially visible.
Much of the profile would be familiar to experienced leaders in evaluation, development effectiveness, and organizational learning. It called for deep knowledge of outcomes, performance systems, institutional reform, target management, implementation monitoring, and government delivery. Many professionals in our field could read those qualifications and recognize work they have been doing for years.
Then the description went further. The successful candidate would also need to understand AI-enabled analytics, large language models, predictive tools, and real-time intelligence. This person would need to work comfortably with senior government leaders, multilateral institutions, implementation teams, researchers, and data scientists.
The role treated outcomes, delivery, technology, learning, institutional reform, and leadership as parts of the same job.
One position alone does not establish a field-wide trend. But senior roles often reveal the problems institutions are preparing to solve and the capabilities they believe they will need next. In this case, the signal was hard to miss.
Outcomes expertise still matters deeply. Increasingly, it may not be enough on its own.
Evidence Is Moving Closer to Delivery
Evaluation professionals have long helped institutions answer essential questions: what changed, for whom, why, and what should be learned?
Those questions remain central to responsible practice. They bring discipline to claims of success and help institutions distinguish meaningful progress from activity.
Organizations are now asking evidence to do something more: inform decisions while implementation is still underway, so leaders can catch a stalling rollout or a failing assumption while there is still time to respond.
None of this is new to anyone who has worked closely with program teams. Political conditions shift. Partners change. Staffing gaps appear. Resources tighten. Communities respond in ways that challenge the original design.
What has changed is the expectation that evidence systems should help institutions navigate these conditions as part of delivery, not only explain them afterward.
That changes what senior evaluation leaders are expected to contribute.
The Role Is Becoming More Integrated
Evaluation functions still carry real responsibility for measurement, accountability, and learning, and some questions genuinely need time. Long-term outcomes cannot be rushed, and credible conclusions about impact, sustainability, and contribution still require careful design.
The trouble starts when evaluation sits too far from where decisions actually get made.
A strong report arrives after a program has already changed course. Monitoring data gets collected but rarely reaches the meetings where decisions happen. A dashboard shows performance slipping without telling anyone why or what to do about it.
The problem in these cases is rarely the evidence itself. It is the system around it: evidence is produced in one part of the organization, interpreted in another, and expected to shape decisions somewhere else entirely.
Adaptive management, developmental evaluation, rapid-cycle learning, and Collaborating, Learning, and Adapting (CLA) have all grown out of this same gap. Each tries to close the distance between what organizations are learning and what they are doing.
Senior evaluation leaders are increasingly the ones asked to close it.
AI Fluency Is Entering the Leadership Brief
Development institutions hold decades of evaluations, project documents, monitoring reports, and learning briefs. Much of that knowledge exists but stays out of reach in practice. A team may know that a useful lesson is buried somewhere in the archive and still spend weeks looking for it.
AI-supported systems can help teams search that material, identify recurring themes, and retrieve relevant evidence far faster than before. Generative AI can support early synthesis as well, provided that sources remain visible and a person still checks the work.
That value comes with new questions.
A polished summary can still be incomplete. A pattern can be easy to detect and hard to interpret. A predictive model can look precise while drawing on weak or poorly governed data.
Faster analysis does not remove the need for judgment. It raises the stakes for leaders who can ask how an output was produced, what it rests on, and what might be missing.
The Leadership Profile Is Expanding
Methodological expertise remains the anchor. Senior evaluation leaders still need a strong grounding in research design, theories of change, qualitative and quantitative evidence, causal reasoning, and the limits of inference.
Around that foundation, several capabilities are becoming harder to do without.
Leaders need enough AI and data fluency to question an output rather than simply accept it. They should understand how it was produced, what informed it, where bias might enter, and what needs expert review before it shapes a decision.
Most evaluation leaders will not need to build models or write code. They do need enough understanding to guide responsible use and work well with the specialists who build the tools.
They also need a stronger grasp of implementation and delivery. It is difficult to design a useful evidence system without knowing how decisions actually get made, where work stalls, and what pressures program teams are under.
Relational skill matters just as much because evidence does not move through an institution on its own. Someone has to interpret it, debate what it means, and decide what happens next.
Governance now belongs on this list too. Privacy, bias, consent, data sovereignty, transparency, and human accountability cannot sit entirely with the technical team.
The emerging evaluation leader works across these boundaries. They connect program teams with technical specialists, translate evidence into terms executives can act on, and help institutions decide not only what can be measured, but what should matter.
Evaluation Leaders Are Well Positioned for This Moment
This shift can feel unsettling for people who have spent years building deep technical expertise. But evaluation leaders are not starting from behind.
Evaluation leaders already understand that more information does not automatically create more insight. They know that methods shape findings, missing information matters, and patterns only become meaningful when interpreted in context.
An indicator can be technically sound and still miss what a community actually values. Evidence is also shaped by power, incentives, language, and organizational culture long before anyone begins analyzing it.
These are exactly the instincts responsible AI use requires.
The opportunity is to carry evaluation’s strongest traditions into this wider arena. That means keeping rigor and independent judgment intact while building technological fluency and working more closely with delivery teams.
A Field in Motion
The profession has evolved before, expanding from measurement into learning, participation, systems thinking, and adaptive management. Each shift has asked practitioners to add new capabilities without giving up the discipline that makes the field credible in the first place.
AI and real-time decision support are part of the next evolution. They are not the whole story.
The larger shift is toward evidence that does more than explain what already happened. It is increasingly expected to help institutions steer while there is still time to change course.
Outcomes expertise remains essential to that work. What has changed is the environment in which it now has to operate.
For evaluation leaders, the real question is no longer only whether we understand outcomes and impact. It is whether we can connect that understanding to delivery, technology, learning, governance, and better decisions.
That is what the next chapter of evaluation leadership will ask of us.
Next in the series: The Evaluation Leader as Integrator will explore what it takes to connect program teams, technical specialists, senior decision-makers, and communities without losing rigor, context, or accountability.



