n.V1-9.10 | Can Impact Be Bought?
01| The idea is elegant enough that it keeps being reinvented. Stop paying organisations for activities and start paying them for results. Define the outcome, agree a price, verify independently, pay on delivery. Everything wrong with input-based funding — the disbursement pressure, the unfalsifiable reporting, the absence of any test — dissolves, because nobody gets paid unless something measurably happened.
02| This has not been a thought experiment for some time. It has been built, funded and run for fifteen years, and the record is specific enough to learn from.
The mechanism works. It has not spread.
03| The first social impact bond launched at Peterborough prison in England in 2010. Private investors funded work to reduce reoffending among short-sentenced prisoners; government would repay them, with a return, only if reoffending fell by an agreed margin. The evaluation found reoffending among the target group fell by around nine percent through 2015, against a target of seven and a half. The mechanism paid out. It did what it said it would do.
04| The counter-case is more instructive. The Rikers Island bond, launched in New York in 2012 to reduce juvenile recidivism, did not hit its targets. Investors lost their principal and the programme ended early. That is usually described as a failure, and for the investors it was — but it is also the design working exactly as intended. The public paid nothing for an intervention that did not deliver, and the loss fell on the party that had chosen to take the risk. Under conventional funding, the same programme would have been paid for in full and its disappointing results filed in an evaluation nobody read.
05| So the model does what its designers claimed. The interesting question is why, fifteen years on, there are only around 300 such projects across some 35 countries — a rounding error against global development and social spending.
06| The field's own answer is transaction costs. Structuring one of these arrangements requires lawyers, intermediaries, an outcome definition everyone accepts, a baseline, an evaluator, and a payment schedule, all negotiated in advance between parties with different interests. The Government Outcomes Lab at Oxford, now the main repository of evidence on these instruments, tracks continuing attempts to reduce that overhead through standardisation — outcome rate cards, where a commissioner publishes in advance what it will pay for a given result, so each deal need not be invented from scratch.
07| And there is a more awkward finding underneath. Rigorous evidence is still lacking on whether outcomes-based financing actually produces better results than conventional financing of the same work. The instrument built to force proof of effectiveness has not itself been proven more effective than what it replaced. That is not an argument against it. It is a striking gap after fifteen years.
The harder problem is who says what happened
08| Suppose the transaction costs were solved. A market in impact still requires what markets in anything require: a shared measure. Buyers must be able to compare what they are buying. That means somebody must rate impact, and different somebodies must arrive at similar answers.
09| There is a large natural experiment on whether that happens, and it is not encouraging. It comes from outside the sector, which is why it is worth more than anything inside it.
10| Environmental, social and governance ratings are the biggest existing attempt to make non-financial performance comparable. Trillions of dollars are allocated with reference to them. Florian Berg, Julian Kölbel and Roberto Rigobon examined the ratings of six major agencies — KLD, Sustainalytics, Moody's ESG, S&P Global, Refinitiv and MSCI — in work published in the Review of Finance under the title "Aggregate Confusion."
11| The average pairwise correlation between agencies ranged from 38 to 71 percent. The raters rarely reach opposite conclusions about a firm, but the dispersion is wide enough that, in the authors' terms, it is difficult to distinguish ESG leaders from average performers. Which agency you consult determines what you conclude.
12| Their decomposition of where the divergence comes from is the part that should worry anyone designing an impact market. Measurement accounts for 56 percent of it — the raters disagree about the facts, not merely about how to weight them. Scope accounts for 38 percent: they are not measuring the same set of things in the first place. Weighting, the factor most people assume is the problem, accounts for only 6 percent. And they detect a rater effect: an agency's overall view of a firm bleeds into how it scores that firm on specific individual categories.
13| Set that against the benchmark. Credit ratings from the major agencies correlate at close to 0.99 — which is why a bond market functions. ESG ratings, assessing large listed companies that publish audited accounts and file mandatory disclosures under regulatory supervision, do not come close.
14| Now consider what verifying impact in a poor neighbourhood involves: no audited disclosure, no regulator, no standard reporting, contested outcome definitions, and an attribution problem that listed-company reporting does not have. If the comparatively easy case has not converged, there is no reason to expect the harder one to.
What that means for a market
15| A market needs buyers, sellers and a price. The price requires comparability, and comparability requires verification that different verifiers agree on. Where verification diverges, the buyer chooses the verifier — and choosing the verifier who returns the flattering answer is precisely the behaviour that gets called impact washing.
16| This is not an accusation of bad faith against anyone in particular. It is what happens structurally when ratings disagree and the rated party is paying. The ESG experience suggests the pressure does not need dishonest actors to produce distorted results. A rater effect and inconsistent scope will do it.
Where this leaves the idea
17| The evidence supports a narrower claim than either the enthusiasts or the critics usually make.
18| Paying for verified outcomes works, in the specific sense that it has been done, has paid out when results were achieved, and has protected public money when they were not. That is a real demonstration and should not be dismissed.
19| It has not scaled, in fifteen years, because each transaction is expensive to construct and because standardising the outcome — the thing that would make it cheap — is the hardest part rather than an administrative detail.
20| And the verification layer a genuine market would require has not converged in the one domain where it has been attempted at scale, with far better underlying data than social outcomes will ever have.
21| So the honest statement is not that impact cannot be bought. It is that after fifteen years of serious attempts nobody has made impact comparable enough to be traded — and the field's largest adjacent experiment suggests the obstacle is measurement itself, not effort.
22| Which closes the loop this book part opened. The first chapter said the sector has no test that closes. Every attempt described since has been an attempt to build one: cost-effectiveness comparison, verified outcome payment, digital eligibility, coordination compacts, community voice. Each was built. Each was partly right. None of them produced an institution that can be told, on evidence its own participants accept, that it has failed.



