Whether it was worth doing is the least evaluated question
The UK government reviewed evaluation across the 227 largest projects it runs, together worth £834 billion. A third have a plan capable of finding out whether they worked. Of the three kinds of evaluation, the one asking whether the benefits justified the cost is the weakest, and the reason is that it has to be designed before the project starts.
A business case is a prediction. It says that if this money is spent, these things will follow.
Predictions can be scored. You wait, you look, you write down what happened. It is the cheapest form of learning available to any organisation, and on capital programmes it is almost never done.
That is usually asserted. It has now been measured, by a government, on its own portfolio.
What the review found
The Evaluation Task Force, a joint HM Treasury and Cabinet Office unit, published a review in April 2025 of evaluation across the Government Major Projects Portfolio. The 2023-24 portfolio comprises 227 projects with a total cost of £834 billion, measured as whole life cost.
Of those 227, a third, 34 per cent, were assessed as having evaluation plans the review judged robust, which is its term for a design capable of answering the questions it sets. Those projects represent £378 billion. The remaining two thirds, representing £456 billion, did not.
Within that two thirds, the split is worth keeping. Forty-five per cent had an evaluation of some kind that did not meet the quality criteria, some of them early in development and likely to improve. The other fifty-five per cent provided no evidence of any evaluation plan at all.
This is a considerable improvement on where it started. A review by the Prime Minister's Implementation Unit in 2019 found that only 8 per cent of spend on the portfolio at that time had high quality impact evaluation plans.
The part that is not what you would guess
The obvious reading is that construction is the problem. It is not.
Broken down by category, infrastructure and construction projects are the best evaluated of the four. Eighty-four per cent have an evaluation plan of some kind, and 68 per cent have one that meets the standard, which is twice the portfolio average.
Information and communications technology sits at 23 per cent. Government transformation and service delivery, 27. Military capability, 2 per cent, for reasons the review discusses and which are genuinely harder.
So a publication about capital assets has to be honest here: on this evidence, the discipline it writes about is comparatively good at asking whether the thing worked. The weakness is elsewhere in the portfolio, and the useful question is what construction is doing right that the others are not.
Where it does fail, and it is the part that matters
There are three kinds of evaluation. Process evaluation asks how the thing was delivered. Impact evaluation asks whether it changed the outcomes it was meant to change. Value for money evaluation asks whether the benefits were worth what they cost.
Coverage is broadly similar across the three. Fifty-four per cent of projects have an impact evaluation plan, 55 per cent a process evaluation plan, 46 per cent a value for money plan.
Quality is not. Thirty-one per cent meet the standard on impact, 28 per cent on process, and 17 per cent on value for money.
Value for money is the weakest on both measures. It is the least planned and the least well planned, and it is the only one of the three that answers the question the business case actually asked.
Why it has to be designed at inception
The review is specific about the mechanism, and it is the same shape as several other failures at this stage.
Evaluation, it says, was often not planned early enough for teams to identify suitable data sources, set up the data sharing permissions needed to access them, or collect new data.
That is the whole problem in one sentence. An evaluation is a comparison, and a comparison needs a before. If nobody recorded the before, the evaluation cannot be built afterwards at any price, because the year it would have described has passed.
It is the same constraint that makes a condition baseline urgent on a new asset, and the same one that makes information requirements worthless if issued after design has started. The cost of the decision is trivial at inception and the option disappears quietly.
Why the business case is the largest single cause of overrun
, read: The largest contributor to optimism bias is the document that starts the projectThe loop this closes, and why optimism persists
Put the two pieces of evidence side by side and something follows that neither states on its own.
The recommended correction for systematic over-optimism is to build estimates from the outcomes of comparable completed projects rather than from a bottom-up view of the one in front of you. That method requires a stock of recorded outcomes.
If two thirds of major projects have no adequate plan to establish their outcomes, and value for money is the least evaluated of all, then the stock of outcomes is not being replenished. The correction for optimism depends on data that the failure to evaluate prevents anyone from collecting.
That is not a coincidence of two separate problems. It is one loop, and it runs in the wrong direction: optimism produces projects, projects are not evaluated, the absence of outcomes leaves optimism unchallenged, and the next business case is written with exactly as much evidence as the last one.
What the review says gets in the way
Three barriers, and the second is the one worth quoting.
Operationally, evaluation is planned too late to secure the data, and some teams lack people with the skills to plan it.
Culturally, some projects reported resistance among senior staff and decision makers, with evaluation perceived as a luxury or a lower priority against immediate delivery pressure.
On resources, some projects simply had no staff, funding or expertise allocated to it.
None of those is about methodology. All three are about what an organisation treats as optional when it is busy, which is a governance question rather than an analytical one.
The response is instructive. HM Treasury has strengthened the emphasis on evaluation in the Treasury Approval Process, so that having an appropriate evaluation plan becomes a requirement for spending approval, and the review records that action as complete in April 2024. The instrument chosen was not guidance. It was the money.
Reading this from here
There is no equivalent published review for this region, and it would be a mistake to assume the numbers transfer. Portfolio composition, governance and reporting obligations all differ.
What transfers is the mechanism and the test, and both are cheap to apply to a single programme.
Ask, for one sanctioned programme, what evidence will exist in five years to establish whether it delivered what the business case promised. Then ask who is collecting it, from when, and where it will be held once the delivery organisation has dissolved.
Where the answer is that somebody will look into it afterwards, the honest position is that the business case was never a prediction. It was a permission, and permissions are not scored.
The framework for this stage sets out how to examine whether an outcome exists separately from the solution, which is the precondition for evaluating anything at all.
Sources. Evaluation Task Force, Government Major Projects Evaluation Review, April 2025, for the portfolio size and cost, the 34 per cent robust figure and the £378 billion and £456 billion splits, the breakdown by Infrastructure and Projects Authority category, the coverage and robustness figures by evaluation type, the three categories of barrier, and the Treasury Approval Process action recorded as complete in April 2024. The 2019 figure of 8 per cent is from a review by the Prime Minister's Implementation Unit, cited in that document. Portfolio composition is from the Infrastructure and Projects Authority Annual Report 2023-24. HM Treasury, Supplementary Green Book Guidance, Optimism Bias, for the adjustment ranges and the contributory factors. ISO 21502:2020, guidance on project management, for benefit management as a practice.
Related reading
Doing the project right, and doing the right project
Almost every instrument the industry owns answers the first question. Estimating, cost control, scheduling, gateways, technical assurance, internal audit, all of them examine execution. Whether this should have been the project at all is settled once, early, by fewer people, and becomes progressively harder to ask.
ReadThe largest contributor to optimism bias is the document that starts the project
HM Treasury publishes uplifts by project type, from 24 per cent on a standard building to 200 on equipment and software. The instruction is to start at the upper bound and come down only against verified evidence. The single largest contributory factor it names is not the contractor, the ground or the weather.
ReadA gate that has never been failed is not a control
Assurance before anything is built costs less and changes more than assurance anywhere else in the lifecycle. It is also the assurance almost nobody buys, and the reason is not that owners are careless. It is that inception produces very little of the sort of evidence an assurance function knows how to examine.
Read