A Pugh matrix example shows you how to score design alternatives against a baseline. What it cannot show is why the team that filled it in still disagrees, because the format's ordinal scoring convention renders a contested row indistinguishable from genuine agreement. The matrix surfaces the split; it was never built to resolve it.

A plus means the alternative is better than the baseline on that criterion. A minus means it is worse. A zero means both sides agree the two are equivalent. It also means neither side could persuade the other and the facilitator moved on. In a Pugh matrix, those last two situations produce the same mark in the same cell.

I have looked at completed sheets where every column had a tidy sum at the bottom and a circled winner at the top, and the team sitting across the table from me did not believe the result. Not because anyone cheated the scoring, and not because the criteria were wrong, but because several of the zeros in the grid were not agreement. They were unresolved arguments the scoring convention had absorbed and concealed.

A Pugh matrix is a concept-selection tool that scores design alternatives as better, worse, or equal to a chosen baseline, using ordinal comparisons rather than numerical weights.

How a Pugh matrix makes disagreement invisible

Stuart Pugh designed the convention deliberately: no numerical weights, no fine-grained scales, just +, -, and 0 to keep the conversation on direction rather than magnitude. The cost of that simplicity is that the finished sheet cannot distinguish between two kinds of zero. Nicholson and Collopy (2019), in the Proceedings of the 22nd International Conference on Engineering Design, documented this precisely: the cell that means "we agree these are equal" looks identical to the cell that means "we could not agree at all".

Pugh matrix scored grid showing two identical zeros, one genuine agreement, one concealed 3-3 disagreement
The same zero hides two different conversations. The matrix cannot tell you which one your team just had.
Click to expand

Consider a team of six engineers comparing three pump designs against a baseline. The scored grid looks decisive.

Datum
(Pump A)
Pump BPump CPump D
Initial cost+0
Reliability0++
Energy efficiency+0
Serviceability0 †+
Ease of installation+0
Sum of +312
Sum of −022
Sum of 0221

† Three engineers scored +, three scored -. The convention recorded 0.

Pump B leads with three plusses and no minuses. The team circles it and moves on. But on serviceability, three engineers scored Pump B as better than the datum and three scored it as worse. The convention recorded a zero, and the matrix shows parity on that row.

The real situation is a room split evenly on a criterion that may be decisive, with neither side having stated what it believes about maintenance access or spare-part supply. That is where the trade-off lives, and the grid has just erased it.

Nicholson and Collopy also found that the arbitrary choice of datum alters the winning concept in roughly one third of cases. The baseline is chosen before scoring begins; it is rarely defended; and it silently steers the result.

Nobody who sells the exercise mentions that part. The facilitator who booked the room for two hours gets to close the meeting with a circled column. The consultant who designed the scorecard gets paid for a clear recommendation. Neither has an incentive to pause on a zero and ask which kind it is. In nearly fifty years of working with decision tools, I have never seen anyone stop the session to make that distinction.

Why the format cannot be fixed by better facilitation

The instinct after a contested result is to blame the way the team ran the session. Better facilitation, clearer criteria, a second round with a new datum. I have watched teams try all three, and in every instance the structural problem remained because it sits inside the ordinal format itself, not in the room.

Hazelrigg (1996) applied Kenneth Arrow's impossibility theorem directly to engineering design methods that aggregate several criteria into one ranking. I had reached the same conclusion from practice before I encountered the proof: any method that combines multiple independent criteria or judgements into a single ranking cannot, in general, be both internally consistent and a valid representation of the group's actual preferences. The conditions under which it can are the exception, not the rule, in real design work.

Scott and Antonsson at Caltech (1999) sharpened the boundary. The vulnerability is not inherent to all scoring methods; it is a specific consequence of discarding the size of a preference and keeping only its direction. A Pugh matrix does exactly this: the + says "better" without saying "by how much" or "with what confidence". That voluntary reduction to ordinal data is the point at which Arrow's constraints bite.

That is an uncomfortable fact for anyone told to run the exercise and bring back a winner. The format the team is using cannot, by mathematical proof, guarantee that its ranking reflects what the group actually values. In the Universal Decision-Making Method, this is the reason the ranking is treated as a starting point rather than a verdict: the matrix's job is to surface where the judgements split, not to settle the question by arithmetic.

Take each contested zero in the grid and write what the plus advocate believes will be true and what the minus advocate believes. Start the Walk →

What the contested result is telling you

A contested Pugh matrix result is evidence that the team has genuine differences of judgement the scoring convention was never built to resolve. Franceschini and Maisano at Politecnico di Torino (2019) measured how much real designers' individual rankings agree, then tested how much the collective ranking changes depending on which aggregation model is applied to the same honest rankings. Their finding was consistent with the literature: different, individually defensible aggregation rules applied to the same data can produce different winners. Which option comes out on top depends partly on which tallying convention the team happens to use, not only on what the designers believe.

That is not a failure of the exercise; it is the most important output the exercise produced. I have watched competent teams reach this point and then spend two hours adjusting scores or switching the datum, trying to make the matrix converge, when the productive move was to stop scoring and start talking about what they actually believed. In procurement and pharmaceutical development, I have seen the same deadlock. The subject matter changes; the pattern does not.

The disagreement names the confrontation the team has been avoiding. Every contested criterion represents something one faction values more than the other does. Until the team says out loud what it is prepared to give up in order to gain what it prefers, no amount of rescoring will move the decision forward.

Each contested zero stands for an unstated assumption: one engineer's belief about production volume, or material behaviour, or customer acceptance, that the engineer across the table does not share. Naming that assumption means going back to each contested row, asking what each side believes will be true about the future, and recording those beliefs explicitly. The question is not "which score is correct"; the question is "what would have to be true for each score to hold".

After the Pugh matrix: three questions before the team commits

Once the Pugh matrix has identified where the room disagrees, three questions move the team from a scored grid toward a decision they can defend.

The first is whether the datum was chosen or merely inherited. For a prioritisation matrix of any kind, the baseline silently anchors every comparison; Nicholson and Collopy's one-in-three finding means the team cannot know whether its winner would survive a different baseline unless it checks. In three decades of facilitating design reviews, I have never seen the datum put to a vote. It arrives as an assumption and stays as a fixture.

The second is to surface the assumptions each contested cell contains. For each row where the team could not agree, write down what the + advocate believes will be true and what the - advocate believes, because those are assumptions. Each one has an influence on the desired outcome, and the team will have different levels of confidence in each.

In the pump example, the serviceability row split 3-3. When I asked each side to state what it believed, two assumptions emerged.

AssumptionAdvocatesInfluenceConfidence
Plant stocks Pump B spares locally; downtime under 4 hours3 engineers (+)HighLow
Nearest depot 400 km away; failure means a 2-day wait3 engineers (−)HighLow

Both beliefs were about the same future: how quickly a failed pump gets back online. Both had high influence on the outcome and neither side had verified its claim. One phone call to the supplier would have resolved it. That is what a contested zero looks like when you open it up: not a values disagreement but an untested factual claim each side had assumed the other already verified.

The method asks the group to judge the significance of each assumption by combining its influence and the team's confidence, rather than by assigning an arbitrary point score. That is the step where the contested weight stops being a political problem and becomes a testable claim.

The third is to ask whether the team has sufficient certainty to act. This is not a vote and it is not a feeling; it is a judgement about whether the remaining assumptions are understood well enough, and whether the monitoring arrangements would catch the ones that may change.

If the answer is no, the team does not need more scoring; it needs more information or a modified option that would bring the decision within reach. If the answer is yes, the final act is to design the monitoring before implementation begins: name the assumptions that carried the decision, assign an owner, and agree what would trigger a reassessment.

Roger Estall and I set out this sequence in Deciding because we had seen the same pattern across engineering and procurement for decades: a tool produces a number, the number becomes the decision, and the assumptions that carried the number are never examined.

A contested result, read properly, has done exactly what a comparison tool should do. It has shown the team where its judgement splits. The step every textbook Pugh matrix example omits is the one that turns that split into a decision the team can defend.

You could rescore the matrix with a new datum and still leave the contested zero unexamined.

Work through your decision

No sign-up. Just pick your decision and start.


Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.