n.V1-7.22 | COST-EFFECTIVENESS SCORING
01| The other method, and the one the sector actually uses. Choose a common unit of human benefit, estimate how much of it each intervention buys per dollar, rank, and allocate. It is the intellectual foundation of evidence-based philanthropy and of every league table this book part opened by dismissing. It deserves a chapter because it is a genuine attempt at the hardest problem in the field, and because the best available test of it is discouraging in a specific and instructive way.
02| The test is this. Two well-resourced, careful, publicly transparent organisations evaluated the same psychotherapy programme against cash transfers, using the same metric — wellbeing-adjusted life years per dollar — drawing on the same evidence base.
03| The Happier Lives Institute estimated cash transfers at 8 WELLBYs per $1,000 and psychotherapy at 77, making psychotherapy roughly ten times more cost-effective. GiveWell, assessing the same programme, estimated about 17 per $1,000 — roughly 2.3 times cash. HLI later revised to about 62, roughly 7.5 times cash.
04| The divergence does not come from different data. It comes from judgement calls about spillover effects onto household members, adjustments for social desirability bias in self-reported wellbeing, corrections for publication bias, and assumptions about how much effects decay when a programme moves from trial to scale. Each choice is defensible. Together they span a factor of four to ten.
05| GiveWell's own stated range for the programme runs from roughly 5 to 80 percent of the cost-effectiveness of its marginal dollar — a sixteen-fold interval inside a single analysis by a single organisation that agrees with itself.
06| The implication is not that these organisations are careless. They are the most careful people doing this work, they published their disagreement, and the exercise is more honest than almost anything else in the sector. The implication is about what the output is. If purpose-built cost-effectiveness analysis, using one explicit metric on one intervention with one evidence base, produces that spread, then a scoring framework with author-assigned weights across many dimensions is not measurement. It is a structured statement of the authors' priors.
07| Publishing priors in structured form is legitimate and useful. What is not legitimate is the next step: using the resulting ranking as evidence that resources should move. The conclusion was installed in the weights before any evidence was consulted, so citing the ranking to justify reallocation is circular. This applies to every capability-scoring instrument, every impact matrix and every framework in which an author decides both what counts and how much it counts for.
08| There is a test that would tell you how much precision such an instrument actually has, and it is standard in every field that takes rating seriously: give the same interventions to several independent raters using the same instrument and measure how far apart they land. As far as this review found, no anti-poverty capability-scoring instrument has been through it. Until one is, the scores have no established precision and should not be reported to two significant figures.
09| So the field is left where the previous chapter put it. There is no agreed scale. The attempt to build one produces defensible estimates that differ by an order of magnitude. And the working substitute — benchmark against cash at equal cost — is a practical answer to an unsolved theoretical problem, which is the standard any new approach will be held to.
10| One last thing that the scoring literature obscures and that this book part should end on. Nearly every framework treats the choice as an allocation problem: given the money, what should it buy. For the countries where poverty is worst that framing is wrong before it starts. Low-income countries spend 0.8 percent of GDP on social protection, their coverage has not moved since 2017, and the additional cost of a basic floor is $308.5 billion a year against total global aid of $212.1 billion. There is no ranking, however well constructed, that reallocates a sum which does not exist.



