After multi-criteria decision analysis produces its ranked list, the step most teams skip is testing the assumptions behind each criterion weight. The scores look precise. The judgments underneath them are often unverified.

In 2002, the UK Department of Health launched the National Program for IT, the largest civilian technology program ever attempted in Britain. Procurement teams scored suppliers across weighted criteria covering technical capability, delivery confidence and commercial terms. By 2011, the program had consumed an estimated £12.7 billion of public money and was formally dismantled. The weighted scores had looked rigorous. The assumptions behind the weights had never been tested.

Multi-criteria decision analysis is a structured method for scoring and ranking options against weighted criteria, producing a numerical comparison intended to guide selection.

NPfIT and the satisficing assumption nobody checked

The National Program for IT, known as NPfIT, aimed to give every NHS patient in England a unified electronic care record. The procurement process was a textbook application of multi-criteria decision analysis. Evaluation panels assigned weights to technical criteria, commercial criteria, and delivery track records. Suppliers including Accenture, Fujitsu, BT and CSC submitted bids that were scored against those weights. The numbers produced rankings. Contracts worth billions were awarded on the strength of them.

The scoring grids assumed two things that no weighted criterion could verify. First, that NHS trusts across England would adopt centralised systems chosen for them by a national program. Second, that clinical workflows in a district general hospital in Devon could be unified with those in a London teaching hospital. These were not technical risks to be mitigated. They were structural premises on which every criterion weight depended.

What to do after multi-criteria decision analysis: weighted scores rank options without testing the assumptions behind each criterion's weight and each option's performance
Weighted scores rank the options. Whether the criteria describe conditions that will actually hold is a separate question.Click to expand

The National Audit Office found that the program had not secured adequate local buy-in before contracts were let. Trusts resisted systems they had not chosen and could not adapt. Clinicians refused to use software that did not match their workflows. Fujitsu's contract for the Southern cluster was terminated in 2008. Accenture walked away from its £2 billion agreement in 2006, paying a £63 million exit fee rather than continuing. The House of Commons Health Committee concluded that the program had been imposed from above with insufficient engagement of the clinicians who would use the systems.

The multi-criteria evaluation had done its job in the narrow sense. Suppliers were compared. Scores were calculated. Rankings were produced. But the criteria assumed a world in which local adoption was a given, and the weights reflected priorities set by a central program team that had never verified whether those priorities matched conditions on the ground. The scoring grid could not surface what it was not asked to examine. When adoption failed, no amount of recalculating the weights would have changed the outcome, because the problem sat underneath the weights entirely.

NPfIT did not fail because the wrong supplier scored highest. It failed because the criteria described a health system that did not exist.

What multi-criteria decision analysis gets right, and where it stops

Multi-criteria decision analysis exists to solve a real problem. When a decision involves multiple objectives that pull in different directions, intuition is unreliable. MCDA provides a structured way to make trade-offs explicit. Belton and Stewart (2002) describe it as a framework for organising and synthesising information in a way that makes the value judgements transparent and open to challenge. The DCLG manual on multi-criteria analysis positions it as a practical alternative to cost-benefit analysis when not all impacts can be monetised.

These are genuine contributions. A decision matrix that forces a team to state what it values and how much is doing more useful work than a meeting where the loudest voice wins. The framework makes it harder to smuggle unstated preferences into a group decision. It creates a record that can be revisited. When teams use it alongside sensitivity analysis, they can see which weights drive the ranking and where small changes in weighting flip the result.

8.4/10
Weighted score for top-ranked supplier
Assumes: the criteria reflect conditions the supplier will actually face, not conditions the evaluation team hoped for
35%
Weight assigned to technical capability
Assumes: technical capability is the binding constraint, rather than organisational adoption or workflow fit
#1
Rank position after scoring
Assumes: the gap between ranked options is meaningful, not an artefact of how scores were anchored

The method stops where its design stops. MCDA takes the criteria as given and scores options against them. It does not ask whether the criteria themselves rest on assumptions that could be wrong. A weight of 35% on technical capability is a value judgement, and the framework makes that judgement visible. But the premise that technical capability is the right criterion at all, that it describes a condition which will actually determine success, is not something the scoring grid examines. That premise is an assumption, and it sits outside the framework's scope.

This is not a flaw in MCDA any more than a thermometer is flawed for not measuring humidity. The problem arises when teams treat the output of the framework as the end of analysis rather than an input to a decision that still needs its premises checked. The distinction between quantitative and qualitative dimensions of a decision matters here: the numbers feel final in a way that the unstated conditions behind them do not.

Write down the assumption your top-ranked option's highest criterion weight depends on and ask whether it was tested or adopted in the scoring session. Start the Walk →

The checkpoint between analysis and action

The step most teams skip is short. Before the ranked list becomes a commitment, name the assumptions that the criteria and weights depend on. Not the scores. The assumptions underneath the scores.

For NPfIT, two assumptions would have surfaced in minutes: that trusts would accept centrally chosen systems, and that clinical workflows could be standardised nationally. Neither had been verified. Both were treated as background conditions rather than testable propositions. A checkpoint that forced them into the open would not have required a new study. It would have required someone to write them down and ask whether they had evidence or only belief.

The five-step method that Roger Estall and I set out in Deciding provides a practical structure for this. Frame the decision. Set out its tentative elements. Surface the assumptions. Decide what level of certainty is sufficient. Then implement with monitoring attached. The third step is where MCDA's output meets its unexamined premises. Each assumption behind a criterion weight gets stated plainly and classified: verified, or believed? If believed, testable before commitment, or needing a monitor with a trigger?

Supplier A ranked first: weighted score 8.4/10
Two paths forward
Standard path
Award the contract and begin implementation planning based on the ranking.
Assumes: the criteria weights reflect conditions that will hold through delivery, not just conditions assumed at evaluation time.
With assumption testing
Before awarding, list the conditions each criterion weight depends on and verify which have evidence and which are still belief.
Tests: whether the criteria describe the world the decision will operate in, not the world the evaluation team assumed.

This is not a second round of analysis. It is a checkpoint that takes the MCDA output seriously enough to ask what it rests on. Teams that skip it treat the ranking as a conclusion. Teams that take it treat the ranking as a starting point for the question that matters: do the conditions behind these numbers actually hold?

The pattern recurs across decision-making frameworks that produce scored outputs. A technical feasibility study can recommend go without testing the staffing and schedule assumptions behind the recommendation. A trade-off analysis can compare alternatives without examining whether the dimensions it compared will remain stable. A trade study can select a design without asking whether the requirements it weighed will still hold. The scored output earns a confidence it has not tested, and the decision inherits the confidence without inheriting the conditions.

You could rank every option by weighted criteria and still leave the assumption behind each weight untested.

Work through your decision

No sign-up. Just pick your decision and start.


Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.