Without Criteria, Evaluation Is Just Opinion

The six OECD-DAC criteria are a shared vocabulary, not a checklist. Used as a menu, they discipline judgement. Used as a checklist, they produce thin findings or unread reports.

Without criteria, evaluation has no ground to stand on. To say a programme is “successful” is meaningless unless we have agreed in advance what success means. Criteria are the explicit standards against which judgement is rendered. They make judgement defensible — not because they remove the judgement, but because they expose its basis and allow others to disagree honestly.

The most widely used framework in international development is the OECD-DAC’s, originally from 1991 and significantly revised in 2019. The revision retained the original five criteria, sharpened their definitions, and added a sixth: coherence. The framework is not perfect and was never meant to be a checklist. Its virtue is that it offers a shared vocabulary across donors, implementers, and evaluators in dozens of countries.

The six, in working language

Relevance asks whether the intervention responds to the needs, policies, and priorities of the people and contexts it serves — and whether it continues to as those needs change. Coherence, the newest, asks how the intervention fits with everything else: other efforts of the same actor, and the work of other actors in the same space. Effectiveness asks whether it achieved its objectives, including how results varied across groups. Efficiency asks whether it delivered results in an economic and timely way. Impact asks about higher-level effects, positive and negative, intended and unintended — with credible reasoning about contribution rather than proof of attribution. Sustainability asks whether the net benefits will continue after the intervention ends.

The DAC framework is a menu, not a checklist. A common failure is treating it as a checklist and forcing every evaluation to answer six criteria with equal depth.

Proportionality, not completeness

The result of the checklist mentality is either thin, generic findings under each heading or an enormous, unread report. Proportionality means asking which criteria matter most for the decision this evaluation will inform. For an early-stage evaluation of a new approach, relevance and coherence usually matter more than impact and sustainability — the impact has not had time to materialise. For an end-of-cycle evaluation of a mature programme, the balance flips.

I ask evaluation teams to assign rough weights to the criteria during inception, and to defend those weights in the inception report. Sometimes the weights are equal, but more often two or three criteria are dominant and the others are addressed lightly. This conversation, held at the start, prevents a great deal of pain at the end. It is far easier to agree that impact will be addressed lightly because the programme is young than to explain, in the final report, why the impact section is thin.

When the framework is not enough

The DAC criteria are useful but not sufficient. Several traditions have argued, with reason, that they are too instrumental, too donor-centric, and insufficient for questions about equity, power, rights, and justice. The Equitable Evaluation Framework asks whose needs and whose definitions of merit are centred, and whether the methods themselves reproduce the inequities the programme is trying to address. Indigenous evaluation traditions ask whether the very paradigm of judging merit and worth from outside is appropriate for community-led work.

These critiques do not invalidate the framework. They invite us to add criteria when the situation demands them — equity of process, equity of results, alignment with rights frameworks, accountability to affected populations, contribution to local capacity, environmental footprint.

FROM THE FIELD

A multi-purpose cash assistance programme in a protracted displacement context — Case G — was evaluated using the DAC criteria, but the team, in consultation with affected populations and a protection working group, added two further criteria: dignity (did the modality protect the dignity of recipients, including those with disabilities and female-headed households) and accountability to affected people (were feedback mechanisms accessible, used, and responded to). These additions changed both the data collection and the findings. The cash modality scored well on effectiveness and efficiency but raised serious dignity concerns at one distribution site. Without the added criteria, that finding would not have surfaced.

That is the right relationship between an evaluator and a framework. The framework should serve the evaluation, not the other way around. Use the six DAC criteria as a starting point and a shared language; weight them according to the decision at hand; and add the criteria your purpose and your users demand. What you must not do — because it is the failure that produces both thin findings and inconvenient blind spots — is treat the menu as a checklist to be ticked with equal, mechanical depth.

Take It to Your Practice

For your next evaluation, weight the six DAC criteria from one to ten according to the decision they must inform, defend the weights in the inception report, and ask one more question: is there a criterion — dignity, equity, accountability — that the standard six would miss entirely?

← Back to Insights