Author: Illuminate Team

  • What Today’s Senior Evaluation Roles Reveal About the Field’s Expanding Competencies

    What Today’s Senior Evaluation Roles Reveal About the Field’s Expanding Competencies

    Part of the Next Chapter of Evaluation Leadership Series

    Why outcomes expertise increasingly needs to be paired with delivery fluency, AI literacy, governance, and cross-functional leadership

    Across international development, senior evaluation roles are beginning to ask for a different combination of capabilities. One recent position description made that shift especially visible.

    Much of the profile would be familiar to experienced leaders in evaluation, development effectiveness, and organizational learning. It called for deep knowledge of outcomes, performance systems, institutional reform, target management, implementation monitoring, and government delivery. Many professionals in our field could read those qualifications and recognize work they have been doing for years.

    Then the description went further. The successful candidate would also need to understand AI-enabled analytics, large language models, predictive tools, and real-time intelligence. This person would need to work comfortably with senior government leaders, multilateral institutions, implementation teams, researchers, and data scientists.

    The role treated outcomes, delivery, technology, learning, institutional reform, and leadership as parts of the same job.

    One position alone does not establish a field-wide trend. But senior roles often reveal the problems institutions are preparing to solve and the capabilities they believe they will need next. In this case, the signal was hard to miss.

    Outcomes expertise still matters deeply. Increasingly, it may not be enough on its own.

    Evidence Is Moving Closer to Delivery

    Evaluation professionals have long helped institutions answer essential questions: what changed, for whom, why, and what should be learned?

    Those questions remain central to responsible practice. They bring discipline to claims of success and help institutions distinguish meaningful progress from activity.

    None of this is new to anyone who has worked closely with program teams. Political conditions shift. Partners change. Staffing gaps appear. Resources tighten. Communities respond in ways that challenge the original design.

    What has changed is the expectation that evidence systems should help institutions navigate these conditions as part of delivery, not only explain them afterward.

    That changes what senior evaluation leaders are expected to contribute.

    The Role Is Becoming More Integrated

    Evaluation functions still carry real responsibility for measurement, accountability, and learning, and some questions genuinely need time. Long-term outcomes cannot be rushed, and credible conclusions about impact, sustainability, and contribution still require careful design.

    The trouble starts when evaluation sits too far from where decisions actually get made.

    A strong report arrives after a program has already changed course. Monitoring data gets collected but rarely reaches the meetings where decisions happen. A dashboard shows performance slipping without telling anyone why or what to do about it.

    The problem in these cases is rarely the evidence itself. It is the system around it: evidence is produced in one part of the organization, interpreted in another, and expected to shape decisions somewhere else entirely.

    Senior evaluation leaders are increasingly the ones asked to close it.

    AI Fluency Is Entering the Leadership Brief

    Development institutions hold decades of evaluations, project documents, monitoring reports, and learning briefs. Much of that knowledge exists but stays out of reach in practice. A team may know that a useful lesson is buried somewhere in the archive and still spend weeks looking for it.

    AI-supported systems can help teams search that material, identify recurring themes, and retrieve relevant evidence far faster than before. Generative AI can support early synthesis as well, provided that sources remain visible and a person still checks the work.

    That value comes with new questions.

    A polished summary can still be incomplete. A pattern can be easy to detect and hard to interpret. A predictive model can look precise while drawing on weak or poorly governed data.

    Faster analysis does not remove the need for judgment. It raises the stakes for leaders who can ask how an output was produced, what it rests on, and what might be missing.

    The Leadership Profile Is Expanding

    Methodological expertise remains the anchor. Senior evaluation leaders still need a strong grounding in research design, theories of change, qualitative and quantitative evidence, causal reasoning, and the limits of inference.

    Around that foundation, several capabilities are becoming harder to do without.

    Leaders need enough AI and data fluency to question an output rather than simply accept it. They should understand how it was produced, what informed it, where bias might enter, and what needs expert review before it shapes a decision.

    Most evaluation leaders will not need to build models or write code. They do need enough understanding to guide responsible use and work well with the specialists who build the tools.

    They also need a stronger grasp of implementation and delivery. It is difficult to design a useful evidence system without knowing how decisions actually get made, where work stalls, and what pressures program teams are under.

    Relational skill matters just as much because evidence does not move through an institution on its own. Someone has to interpret it, debate what it means, and decide what happens next.

    Governance now belongs on this list too. Privacy, bias, consent, data sovereignty, transparency, and human accountability cannot sit entirely with the technical team.

    The emerging evaluation leader works across these boundaries. They connect program teams with technical specialists, translate evidence into terms executives can act on, and help institutions decide not only what can be measured, but what should matter.

    Evaluation Leaders Are Well Positioned for This Moment

    This shift can feel unsettling for people who have spent years building deep technical expertise. But evaluation leaders are not starting from behind.

    Evaluation leaders already understand that more information does not automatically create more insight. They know that methods shape findings, missing information matters, and patterns only become meaningful when interpreted in context.

    An indicator can be technically sound and still miss what a community actually values. Evidence is also shaped by power, incentives, language, and organizational culture long before anyone begins analyzing it.

    These are exactly the instincts responsible AI use requires.

    The opportunity is to carry evaluation’s strongest traditions into this wider arena. That means keeping rigor and independent judgment intact while building technological fluency and working more closely with delivery teams.

    A Field in Motion

    The profession has evolved before, expanding from measurement into learning, participation, systems thinking, and adaptive management. Each shift has asked practitioners to add new capabilities without giving up the discipline that makes the field credible in the first place.

    AI and real-time decision support are part of the next evolution. They are not the whole story.

    The larger shift is toward evidence that does more than explain what already happened. It is increasingly expected to help institutions steer while there is still time to change course.

    Outcomes expertise remains essential to that work. What has changed is the environment in which it now has to operate.

    For evaluation leaders, the real question is no longer only whether we understand outcomes and impact. It is whether we can connect that understanding to delivery, technology, learning, governance, and better decisions.

    That is what the next chapter of evaluation leadership will ask of us.


    Next in the series: The Evaluation Leader as Integrator will explore what it takes to connect program teams, technical specialists, senior decision-makers, and communities without losing rigor, context, or accountability.

  • The Hardest Part of Improvement

    The Hardest Part of Improvement

    There is no shortage of ideas for organizational improvement.

    Strategic planning processes identify priorities. Employee surveys surface recurring concerns. Customers explain where services fall short. Audits uncover risks. After-action reviews capture hard-won lessons while they are still fresh. Board retreats end with pages of flip charts covered in commitments and next steps. Evaluations add their own layer of recommendations to the pile. By the time those conversations are over, most teams have a fairly clear picture of what could be better.

    The harder question is what happens next.

    Why do some recommendations become everyday practice while others quietly disappear?

    Over the years, we have noticed that organizations are often better at identifying opportunities for improvement than they are at embedding those improvements into the way they actually work.

    That isn’t because people don’t care. Quite the opposite. Most teams genuinely want to improve. They spend time discussing findings, prioritizing recommendations, and developing action plans. There is often real energy in the room, and people leave believing the conversation mattered.

    Then everyone returns to their day jobs.

    Projects continue moving. New requests arrive. Vacancies need to be filled. Budgets shift. Leadership priorities change. The urgency of today’s work gradually overtakes the importance of yesterday’s reflection.

    Six months later, some recommendations have taken hold. Others remain exactly where they started: in a report, a meeting summary, or a strategic plan that everyone remembers but few have revisited.

    And people begin to realize that making improvement part of everyday work is harder than expected.

    Improvement Is a Different Kind of Work

    Perhaps one of the assumptions organizations make, without realizing it, is that once people understand what needs to change, change naturally follows.

    Experience suggests otherwise.

    Knowing what needs to change and making that change stick are two very different kinds of work. The first depends on insight. The second depends on systems.

    The first recommendations to move forward are often the easiest ones. They show that the process mattered and that the organization is taking action.

    The more difficult recommendations, however, are necessarily more difficult to address. They may require coordination across teams, changes to existing processes, additional resources, or sustained leadership attention. Those are much easier to postpone, even when everyone agrees they are important.

    And so, six months later, organizations are often in an uncomfortable position, having acted on some of what they learned, but definitely not all of it.

    It’s not that leaders are ignoring what they need to do.

    It is simply that harder changes are hard to implement, especially in light of all the other things that need to happen.

    The Real Work

    Organizations may not need more recommendations as much as they need stronger ways of carrying good recommendations forward.

    Knowing what needs to change is an important milestone. Designing an organization that can consistently act on what it learns is something else entirely.

    That may be the hardest part of improvement.

    Questions for Reflection

    Think back to a time that your organization was successful in driving improvements to a program or process. What did that look like? What enabled success? Does your organization know how to implement change successfully and on a regular basis?

    If so, then there is probably some combination of the following in place at the systems level:

    1. A Clear Review Process: A consistent mechanism for evaluating progress and identifying course corrections.
    2. Resource Realism: Structural clarity on the exact time, funding, and headcount required to implement changes.
    3. Executive Sponsorship: Sustained leadership buy-in and a willingness to fiercely protect the resources committed to the effort.
    4. Explicit Accountability: Defined roles, clear ownership, and transparent responsibilities for execution.
    5. Continuous Learning Loops: Simple, repeatable checkpoints to reflect on what is working and what isn’t.
    6. Cultural Reinforcement: Deliberate celebration of success and structural reinforcement of existing strengths.

    How does this compare with how your organization is structured today?

    If you could use a partner to help put these systems and habits in place, let’s connect. Whether you need a strategic diagnostic of your current operational processes or hands-on facilitation to design stronger learning loops, we are here to help you turn insight into sustained impact.

  • When the Evidence Is Right but the Moment Has Passed

    When the Evidence Is Right but the Moment Has Passed

    Artificial intelligence (AI) is changing evaluation in many ways. One of its most promising contributions may be helping evidence reach decision-makers while there is still time to act.

    Most evaluators have experienced some version of this.

    A program commissions an impact evaluation. It’s designed with real care: a defensible counterfactual, a sound methodology, the kind of work that holds up to scrutiny and deserves to. The team is proud of it, and they should be.

    Two years later, it arrives. Rigorous. Publishable. Genuinely good.

    By then, the program has already moved on. It pivoted, or scaled, or wrapped. The funding cycle turned over. The people who first asked the questions are working on something new.

    The evidence is everything we hoped it would be.

    The moment it was built to inform has quietly closed.

    That’s not a failure of rigor. The rigor was never the problem. What slipped was the timing, the distance between when the evidence became available and when decision-makers needed it. Evaluation often takes exactly as long as it should. Decisions, however, rarely wait.

    That gap exists in more of our work than we often recognize.

    What Well-Timed Evidence Has in Common

    Rather than asking why evaluations land late, a more useful question is this:

    When has evidence landed right on time, and what can we learn from those moments?

    Most of us can think of examples. They rarely come from the headline evaluation. More often, they come from something smaller and better timed: a midline finding that arrives while there is still room to adjust, a field team noticing a pattern that monitoring data confirms a week later, or a rapid evidence synthesis that helps shape a new initiative before major investments have already been made.

    Those moments are worth studying because they remind us what good evidence really offers. The power isn’t rigor alone. It’s rigor that arrives in time to matter.

    The encouraging part is that organizations already create these moments. Strong monitoring systems, developmental evaluation, rapid feedback loops, and reflective practice all help generate timely insight. Too often, however, we treat those moments as fortunate exceptions rather than something we can intentionally design for.

    Looking closely at these examples, three characteristics consistently appear. The evidence reaches decision-makers while there is still time to act. Someone already knows which decision the evidence is meant to inform. And the work of turning raw information into useful insight doesn’t sit untouched in a queue for months.

    Two Kinds of Time in Evaluation

    One way to think about this challenge is to distinguish between two kinds of time.

    The first is causal time: the time it takes outcomes to unfold and evidence to emerge. Some questions simply require patience. No amount of technology can tell us today what will only become visible a year from now.

    The second is effort time: the time spent cleaning datasets, coding interviews, synthesizing literature, drafting reports, preparing visualizations, and moving work through multiple rounds of review. That work is essential, but much of it reflects the effort required to process information rather than the passage of time itself.

    We can’t compress causal time. Outcomes unfold when they unfold. What we can change is the amount of effort it takes to move from raw information to usable insight. That is where the opportunity lies.

    Consider a training program that collects participant feedback after every session. If the team reviews those responses six months later, several cohorts receive essentially the same experience. Facilitators may never discover where participants struggled until long after the course has ended.

    If those same responses are reviewed and synthesized within days, facilitators can strengthen the very next session. They can clarify confusing concepts, adjust activities, and respond while the program is still unfolding.

    The evidence itself is exactly the same. The only difference is that it reaches people while they still have the ability to act on it.

    Where AI Can Help

    AI’s greatest contribution may be in its potential to reduce effort time.

    It can help organize information, synthesize existing evidence, support qualitative coding, identify themes across large volumes of text, and draft initial summaries or reports for human review.

    None of those tasks eliminate the need for evaluators. They simply reduce the time between collecting information and putting it in a form that people can use.

    Evaluation has never been only about collecting evidence. Its purpose has always been to help people make better decisions. When effort time becomes the bottleneck, reducing that effort can make evaluation more useful without making it any less rigorous.

    This is also where some caution is important. AI doesn’t create rigor, and it can’t accelerate outcomes that naturally take years to emerge. Asked to answer genuinely causal questions, it often produces answers that are fast, confident, and wrong.

    The real opportunity is much simpler. Reduce the effort that keeps useful evidence from reaching decision-makers while they still have time to act. Protect the rigor. Shorten the delay.

    Building an Evidence System That Works at Two Rhythms

    Rather than thinking about AI as an alternative to traditional evaluation, it may be more useful to think about the different roles evidence plays throughout the life of a program.

    Some questions benefit from fast feedback. Monitoring data, participant feedback, rapid evidence syntheses, qualitative insights, and lightweight analyses help teams learn while implementation is still underway. They support course corrections, surface emerging issues, and help people respond while there is still time to act.

    Other questions deserve a slower pace. Understanding whether an intervention contributed to meaningful outcomes often requires stronger designs, deeper analysis, and more time. Those questions should not be rushed simply because faster tools are available.

    Both kinds of evidence matter. They serve different purposes, and together they create a stronger evidence system.

    In many ways, evaluation has always worked like this. Monitoring systems, developmental evaluation, rapid feedback, and adaptive management all recognize that some information needs to reach decision-makers quickly, while other questions require more rigorous investigation before conclusions can be drawn.

    AI doesn’t change that balance. If anything, it reinforces it. Its greatest value may lie in helping organizations move more efficiently from raw information to usable insight, making timely learning easier while preserving space for the deeper work that evaluation has always required.

    That is why the most interesting conversation is not whether AI will replace evaluators or automate evaluation. Those questions tend to generate more headlines than insight. A more useful conversation asks where AI can reduce effort without replacing judgment. Where can it free evaluators from repetitive processing tasks so they have more time for interpretation, stakeholder engagement, systems thinking, ethical reflection, and the work that depends on experience rather than automation?

    That feels like a much more productive direction for the field.

    Designing for Better Decisions

    Organizations don’t need to redesign their entire monitoring and evaluation function to begin moving in this direction.

    A better place to start is by looking at moments when evidence genuinely influenced a decision.

    What made those moments possible? Which parts of the process required careful evaluation? Where did the work simply get bogged down in moving information from raw data to usable insight?

    Reflecting on those questions often reveals opportunities that have been there all along.

    The goal isn’t to ask where AI should replace existing practice. A more useful question is where it can remove friction. Perhaps information routinely sits untouched for weeks before anyone reviews it. Perhaps evaluators spend hours on repetitive tasks that add little analytical value. Or perhaps the real delay occurs because teams are still processing information long after a decision needs to be made.

    Those are practical problems with practical solutions. Addressing them doesn’t change the purpose of evaluation. It simply helps evidence reach the people who need it while there is still time to use it.

    The Opportunity Ahead

    Evaluation has spent decades refining the methods we use to produce credible evidence. That work remains essential. AI doesn’t change the need for thoughtful design, rigorous methods, careful interpretation, or meaningful engagement with stakeholders.

    What may be changing is something else entirely.

    Our ability to reduce the time between collecting evidence and using it.

    That may be one of the next frontiers for evaluation. The challenge isn’t choosing between rigor and timeliness. It’s designing evidence systems that support both.

    At Illuminate, we believe the strongest evidence systems do more than produce credible findings. They help organizations learn, adapt, and improve while implementation is still underway. Sometimes that means strengthening monitoring systems. Sometimes it means improving evaluation design or streamlining workflows. Increasingly, it may mean using AI thoughtfully to reduce effort while preserving the rigor and professional judgment that high-quality evaluation demands.

    Because evidence creates its greatest value when people can still act on it.

    Continue the Conversation

    Organizations across sectors are asking similar questions: How can we make better use of evidence? How can we shorten the time between learning and decision-making? And where can AI genuinely add value without compromising quality?

    These are questions we explore every day through our consulting, facilitation, and professional development programs.

    If your organization is thinking about the future of monitoring and evaluation, we’d welcome the opportunity to continue the conversation.