n.V1-9.03 | The Price Nobody Calculates
01| The Center for Global Development's review of how evidence gets used in aid produced one finding that deserves to be better known than it is: only about one in five impact evaluations considers what a programme achieved per dollar spent.
02| Read that carefully. It is not a statement about how much aid goes unevaluated. It is about the evaluations that do happen — the serious ones, with budgets, designs and published results. Four out of five establish whether something worked without establishing what it cost to work.
03| Every one of those evaluations sat inside a project with a budget. The cost was known. It was simply never brought into the same document as the result.
04| That is the failure worth explaining, and it is not secrecy and not incompetence. Three separate things are going on, and they compound.
The units genuinely do not match
05| Start with the part that is nobody's fault. Comparing the cost of two interventions requires a shared definition of what is being bought, and there is none. One project reports households reached, another people reached, another communities engaged. One reports the capital cost of building a latrine; another reports the annual cost of keeping sanitation running, which includes emptying it. One figure is a one-off, the other recurs for twenty years. Neither is wrong. They answer different questions, and no convention says which question a project must answer.
06| This is a real obstacle. It is also convenient, because it makes the absence of comparison look like a measurement problem rather than a choice. The other two reasons are less comfortable.
The person who spent the money is not the person who delivered
07| In most large development work a funder finances, an implementing organisation delivers, a government or utility operates, and a household receives. Costs are recorded at the top of that chain and outcomes appear at the bottom, with two or three organisations in between, each with its own reporting cycle and accounts.
08| Working out what a result cost therefore means reassembling information held by parties with no obligation to each other and no common format. Nobody is responsible for the division, because the division belongs to no single organisation. It is the kind of task everyone agrees should be done and nobody's job description contains.
Comparison produces losers
09| Then the plainest reason. A cost-effectiveness comparison, unlike an evaluation, ranks. It produces a better and a worse, and both of them are somebody's programme.
10| An evaluation showing a project achieved its objectives is good news for everyone involved. An analysis showing the same objective was achieved elsewhere for a third of the money is bad news for a specific department, a specific contractor and a specific career. The first study gets commissioned readily. The second requires somebody senior to want it.
What happened when someone did want it
11| USAID commissioned exactly that, and the results are the most useful thing in this chapter. The method is cash benchmarking: test a programme against simply handing the intended beneficiaries the same money. Cash is a floor — its administrative costs are minimal and its household effects are well documented — so the question becomes whether a programme earns its premium over the simplest possible alternative.
12| Craig McIntosh and Andrew Zeitlin, working with USAID and GiveDirectly, tested Rwanda's Gikuriro child nutrition and sanitation programme, delivered by Catholic Relief Services, against cost-equivalent cash in the same population at the same time.
13| State the results precisely, because they get oversimplified in both directions. At a cost of $142 per household, Gikuriro produced gains in savings — a domain one of its components directly targeted — and no improvement in consumption, dietary diversity, wealth, child anthropometrics or anaemia. The cost-equivalent cash arm, at $124, produced significantly greater consumption than the programme, and also moved no child outcomes. A larger transfer costing $517 substantially improved consumption and investment and modestly improved dietary diversity and child growth.
14| Three conclusions in descending order of confidence. At equivalent cost, cash beat a well-designed integrated programme on consumption and matched it on child outcomes, which is to say both produced nothing there. Programme design does deliver what it targets — the savings component worked. And amount mattered more than modality: the transfer that changed things was four times the size of the benchmark.
15| This is what a real comparison looks like, and it is chastening for everybody. The in-kind programme did not beat cash. Cash did not move child health either. What moved outcomes was more money. A parallel study of a Rwandan youth workforce readiness programme found the same shape of result — no outperformance of cost-equivalent cash on most primary outcomes.
16| What is more telling than either result is the surrounding circumstance. The comparison had to be built from scratch as a research project, because nothing in the existing data could answer it. Reviewing the effort afterwards, CGD described cost-effectiveness benchmarking at the agency as relatively new and niche, dependent on having a champion inside the building. An exception that requires a champion is not a system. It is a person.
Cash as the unit of account
17| The most serious attempt to make interventions comparable is to price everything in cash. GiveWell does this explicitly, expressing cost-effectiveness as multiples of unconditional transfers: a programme rated ten times cash is estimated to deliver ten times the benefit per dollar. The bar has moved over time — around eight times for top charities in 2024, around six now — which itself shows the bar tracks how much money is available as much as what programmes achieve.
18| What it solves is real: a common denominator, every claim on one scale, the opportunity cost of any grant made explicit. That is more than the sector had before.
19| What it does not solve matters more here. The multiple depends entirely on which outcomes are counted and how they are valued against each other — a life saved against a year of schooling against a rise in consumption — and those are moral choices, stated openly and contested by others. And it systematically favours interventions with short causal chains and measurable endpoints. Nothing whose benefit arrives in twenty years, or accrues to a community rather than a person, scores well, because the evidence needed to score it does not exist.
20| That point should be conceded rather than argued around. A method that ranks by measurable individual outcomes will always rank collective, slow, institutional work poorly — not because it is worse, but because it is harder to evidence. The correct response is to say so, not to invent a rival ratio that happens to rank it highly.
The objection from the other direction
21| Natsios's argument applies here and cuts hard. A price list rewards whatever divides cleanly. Latrines divide cleanly. Building a functioning water utility does not. A regime ranking interventions by cost per countable unit will favour the countable and may starve the slow institutional work that determines whether the countable things still exist in ten years.
22| So the field holds two credible criticisms of itself pointing in opposite directions. One says an industry spending enormous sums almost never asks what a result cost, that the comparison is feasible, and that its absence protects incumbents. The evidence for the absence is solid: one in five, from the field's own review. The other says the measurement apparatus is already the problem and more ranking will deepen the distortion. The evidence for that is an argument by a former agency head.
23| Both cannot be straightforwardly right, and the field avoids the collision by treating cost-effectiveness as a technical matter for specialists rather than the contested question it is. The honest position is that nobody knows what a costing regime does to the work over a decade, because no institution has run one long enough for anyone to find out.



