Three people in a conference room. A product manager has drawn a weighted decision matrix on the whiteboard: five criteria across the top, four vendor options down the side. The scores are filled in. The weights are blank. Engineering says reliability should carry 0.40. Sales says time-to-market deserves 0.35. The product manager suggests splitting everything equally at 0.20 and moving on.

Nobody objects. A shared non-decision is easier to defend than admitting the room has not yet started the argument that matters.

I have watched this scene in engineering firms and procurement panels for decades. The matrix was supposed to produce a winner. Instead it stalled at the one point that requires actual judgement: the weights. If you are the person in that room, charged with presenting a recommendation by Friday, the instinct is to search for a better weighting method or a template that removes the subjectivity. That instinct is wrong. The problem is the conversation the team has not had, not the method.

The search for "weighted decision matrix" returns ten pages that explain the arithmetic and promise objectivity. They describe the tool well enough. None addresses the part that stalls the meeting: what to do when the team cannot agree on the weights. Everyone in that room already knows the weights carry the decision. The question nobody answers is how to resolve the argument. Four decades of multiattribute decision research address it. Those findings almost never reach the person at the whiteboard.

A weighted decision matrix is a scoring model that compares options by rating them against criteria and multiplying each score by a weight that represents the criterion's relative importance.

Why most weighted decision matrix arguments are about the wrong question

The room assumes the fight is about the number: is reliability 0.40 or 0.25? In practice, a substantial portion of what looks like disagreement about weights is procedural noise, not genuine value difference.

In 1982, Schoemaker and Waid gave MBA students the same career-choice decision and asked them to weight the criteria using three different elicitation methods: ratio weighting, swing weighting, and indifference-curve comparison. The subjects, the decision and the criteria were identical in every condition, but the resulting weights were statistically different across methods. Subjects could not explain why they produced different numbers depending on how the question was framed. The methods are supposed to be theoretically equivalent. In practice, the procedure changes the answer.

Weber and Borcherding reviewed the experimental literature and identified specific biases that make weight elicitation unreliable independent of the person doing the weighting. Splitting bias means that breaking one criterion into subcriteria inflates its total weight: a team that splits "quality" into reliability and defect rate gives quality more aggregate weight than a team that keeps it as a single line. Anchoring means the first weight the team assigns becomes the reference point against which every subsequent weight is adjusted. Both effects operate independently of what the team actually believes.

Two members of the same team can genuinely agree on what matters and still produce different weights because they framed the criteria differently or started from different reference points. There is a simpler question worth asking before the number: are we disagreeing about what matters, or about how the question was asked?

Most pages about weighted decision matrices treat choosing criteria weights as a two-line instruction: "assign weights based on importance, ensuring they sum to 1.0." That instruction assumes the team already agrees on importance and that the weighting procedure is neutral. Neither holds in a real meeting. RICE prioritization builds the same assumption into its confidence score, a percentage picked by feel with no external standard to check it against.

Diagram showing how the same person produces different criteria weights using ratio weighting, swing weighting, and indifference curves, based on Schoemaker and Waid 1982
The method changes the weights, even when the person and the decision do not.
Click to expand

When the rank is enough

If the team agrees on which criterion matters most and which matters least, the cardinal weight may not matter. Barron and Barrett simulated thousands of multiattribute decisions and tested how well approximate weights derived from ordinal rankings performed against weights from formal utility assessment. Rank-order centroid weights, calculated from nothing more than "criterion A matters more than B, B more than C," correctly identified the best alternative 75 to over 90 per cent of the time, depending on the number of attributes and alternatives. The more alternatives on the table, the better ordinal weights performed.

Stillwell, Seaver and Edwards pushed the finding further. Even equal weights, where every criterion receives the same importance, produced decisions within a few percentage points of optimal. Teams tend to fight hardest about weights when the decision is close. That is exactly the condition under which weight precision matters least, because any plausible weighting scheme produces a near-tie.

The practical test takes five minutes. Ask the room: do we agree on the ordinal ranking of criteria? Does quality matter more than cost, and cost more than schedule? If the team can agree on that ordering, run the weighted decision matrix with rank-derived weights and again with equal weights. If both produce the same winner, the cardinal argument was consuming meeting time without changing the decision. If they produce different winners, you have learned something more useful: you now know exactly which criterion's position drives the result, and that is where the real conversation belongs.

This is where a prioritisation matrix earns or loses its credibility. A ranking that holds across any plausible weighting confirms a stable ordering. The team can present that result with confidence because it does not depend on anyone's contested number. A ranking that flips depending on the weights has not failed. It has done something more useful: it has shown the team that the decision depends on a judgement nobody has stated clearly enough to contest.

For the product manager who needs a recommendation by Friday, this is the most useful step. Settle the rank and test whether it changes. Most of the time, it will not.

Settle the ordinal ranking of your criteria and test whether the cardinal weights change the winner before Friday's meeting. Start the Walk →

When a weighted decision matrix cannot settle the argument

Sometimes the team cannot agree on the ordinal ranking itself. Engineering places reliability above time-to-market. Sales places time-to-market above reliability. When the scoring criteria are genuinely contested, no procedural adjustment will reconcile these positions because the disagreement is about what the organisation values, not about how a questionnaire was phrased.

At this point, the weighted decision matrix has done the most useful thing it can do. It has surfaced a value conflict that was sitting underneath the spreadsheet the whole time. The temptation is to split the difference or let seniority decide. Both bury the conflict without resolving it, and the buried conflict will resurface the moment the decision produces a result that one faction predicted and the other denied.

The better move is to treat each proposed weight as an assumption and test it the way you would test any other assumption supporting a decision. "Quality should be 0.40" is not a measurement. It is the claim that quality matters roughly twice as much as any single other criterion. State the claim plainly and then ask: what would make that wrong? If a competitor launches first, does the reliability advantage hold? The Recognise-assumptions step in the Universal Decision-Making Method exists for exactly this reason: to move the conversation from numbers the team cannot agree on to claims the team can examine.

Roger Estall and I wrote in Deciding that influence and confidence cannot be accurately measured and, therefore, cannot be combined in a mathematically valid way. That observation was about formal risk analysis, but it applies with equal force to criteria weighting. The matrix output is a prompt for judgement, not the judgement itself. When the team treats it as a verdict, every unresolved weight disagreement becomes a permanent fault line in the decision.

An alternative to cardinal weighting exists. Instead of asking "how important is this criterion on a scale of 1 to 10?", ask two questions: how significant is this assumption to the desired outcome, and how volatile is the context it depends on? The significance question exposes which criteria actually drive the decision. The volatility question exposes which criteria need monitoring after the decision is made. Together they do what a contested weight argument cannot: they give the team something to investigate rather than something to keep debating.

The Universal Decision-Making Method calls this step Recognise assumptions, and it belongs before the spreadsheet produces a number, not after.

The weighted decision matrix nobody could inspect

In 1993, Minister Ros Kelly distributed approximately A$30 million in community sports and recreation grants through a programme that published merit criteria in its guidelines. Applications were scored. The Auditor-General later found that the actual allocation did not follow the stated criteria. Grants went disproportionately to marginal and government-held electorates. Kelly maintained she had tracked allocations on a whiteboard. The whiteboard was erased. She resigned from the ministry in February 1994.

The published criteria gave the programme an appearance of structured, merit-based assessment. Applicants believed their projects would be judged against those criteria. The real deciding factors were never disclosed, and because they were never disclosed, they could not be challenged until after the money had been spent and the Auditor-General had examined the pattern.

Most weighted matrices are built in good faith by teams who genuinely want a structured way to compare options. What the Kelly case exposes is how far a matrix can drift from its stated purpose when the weights sit outside the conversation. Six months from now, the team needs to be able to explain why each criterion mattered and what each weight assumed. A weighted decision matrix is exactly as defensible as the argument behind its weights. The team that stalls at the weighting step and compromises on equal weights to keep the meeting moving has buried the argument. The team that settles the ordinal ranking first and tests whether cardinal precision changes the winner has done something a spreadsheet cannot do on its own. It has made the actual decision visible enough to defend.

You could assign the weights and never discover which criterion carried the decision.

Work through your decision

No sign-up. Just pick your decision and start.


Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.