A head of strategy at a mid-sized engineering firm handed me a brief her board had set. They wanted a Kaplan Norton balanced scorecard deployed by mid-year, across all four perspectives, because the CEO had heard Robert Kaplan speak at a governance conference the month before. The brief ran three pages describing the framework and zero pages on the question that should have come first: whether the strategy they intended to measure was built on assumptions anyone had tested.

That is how most scorecard projects begin. Somebody encounters the model, recognises it as serious work, and commissions the rollout before the decision about what to measure has been made. The strategy document was approved months ago; the scorecard is asked to grade execution of a plan that arrived with its assumptions unexamined.

In the engineering firm's case, the board had set a margin growth objective resting on a pricing assumption nobody had revisited since the previous year's budget round. Four perspectives' worth of measures would have monitored that assumption with precision, and none of them would have told the board it was wrong.

A Kaplan Norton balanced scorecard is a performance measurement framework that groups an organisation's metrics around four perspectives: financial, customer, internal process, and learning and growth. It connects operational measures to strategic objectives to show whether a plan is being executed.

What Kaplan and Norton actually built

In early 1992, Robert Kaplan and David Norton published "The Balanced Scorecard: Measures that Drive Performance" in the Harvard Business Review. The article drew on a one-year research project with twelve companies sponsored by the Nolan Norton Institute, KPMG's research arm at the time. Their argument was narrow and well-made: financial metrics alone give a misleading picture of organisational health because they report the past and say nothing about the operational drivers of future performance.

The four perspectives were the answer. Financial results sit beside customer outcomes, internal process efficiency, and learning and growth capacity, so the board sees execution from multiple angles rather than waiting for the quarterly numbers to confirm what went right or wrong months earlier.

Over the following decade, Kaplan extended the model into strategy maps, tracing causal chains from learning and growth through processes to customer outcomes to financials. Strategy maps organised the logic within a given strategy but did not test whether the assumptions at the base of that chain were sound. The Balanced Scorecard Hall of Fame inducted over 200 organisations from nearly 40 countries, confirming the framework does exactly what it was designed to do: track execution with discipline and breadth.

What is easy to miss in the original research and Kaplan's own 2010 retrospective is what the model explicitly excludes. Kaplan drew the boundary himself: 'strategy precedes stakeholders.' The scorecard implements and communicates a strategy; it does not validate one.

The opening line of the 1992 paper, 'What you measure is what you get,' presupposes that the strategy feeding the scorecard is sound. That competence is precisely what makes the framework dangerous when the strategy rests on untested assumptions: the board receives a disciplined, multi-perspective report confirming that execution is on track, and the dashboard itself becomes the reason nobody asks whether the plan deserves the confidence.

The person who saw this earliest was Art Schneiderman, who built the first balanced scorecard at Analog Devices in 1987 and later published a critique identifying three structural reasons implementations fail. His first cause was the sharpest: non-financial variables are incorrectly identified as primary drivers of stakeholder satisfaction. In plainer language, the wrong things end up on the scorecard because nobody validated the assumptions upstream.

Two columns comparing what a balanced scorecard tracks versus what it never tests
The scorecard tracks execution across four perspectives. Whether the strategy deserves the measurement is a question it was never built to raise.

Why balanced scorecard adoption peaked and fell

Bain & Company's Management Tools survey tracked balanced scorecard adoption from 1996 onwards. Usage rose from 39 per cent of surveyed firms in 1996 to a peak of 66 per cent among 1,221 respondents in 2007. By 2018 it had fallen to 29 per cent. In the 2023 survey, Bain dropped the balanced scorecard from its 25-tool list entirely; it no longer met the relevance threshold for inclusion.

That arc does not mean measurement stopped mattering. It means organisations discovered that broader measurement did not, on its own, produce better strategic outcomes. The Kaplan Norton balanced scorecard answered 'are we executing?' well enough to spread to two thirds of surveyed companies.

The question it never asked, 'should we be executing this?', is the one that explains the decline. Kaplan and Norton delivered exactly what they designed; the disappointment came from expecting the tool to do something it was never built for.

I have seen this pattern in organisations with every perspective covered. The dashboard stayed green because each measure reported faithfully against objectives set when the market looked different. A standard balanced scorecard reports from inside the frame that a prior decision established, and it has no mechanism for testing whether that frame still fits.

Rank the assumptions behind your strategy by influence and confidence and see which ones the scorecard was never built to test. Start the Walk →

Three questions before you build a Kaplan Norton balanced scorecard

Before any organisation commits to a scorecard rollout, I want three things answered.

First: for each proposed measure, which assumption does it test? Most scorecards are populated by asking 'what should we track?' The better question is 'what do we believe about this strategy, and which beliefs carry the most weight?'

In the Universal Decision-Making Method, the step is to rank every assumption by how much influence it has on the outcome and how much confidence exists in it. A measure earns its place only if it tracks an assumption that is both influential and uncertain. Everything else is decoration that will cost time to maintain and attention to read.

A statutory public safety organisation I worked with had been allocating 0.03 per cent of its budget to the one function that directly prevented loss of life. When someone finally asked which assumption about public safety carried the most weight, spending shifted to 0.5 per cent, and deaths attributable to that function fell 60 per cent.

No scorecard would have surfaced this, because the assumption was upstream of every measure the framework could hold. The reallocation was small and the result was large, because someone stopped measuring everything and started measuring the thing the purpose actually depended on.

Second: for each measure, what would the organisation do if this number moved? A metric with no pre-agreed response threshold is not a control; it is a number someone looks at, notes, and moves past.

Roger Estall and I built monitoring into the secondary elements of every decision in Deciding because the moment of deciding is when the people in the room have the clearest awareness of which numbers matter and what each movement would mean. Build the scorecard after the fact and that awareness is gone; the measures end up being chosen by whoever fills in the template, not by whoever made the decision.

Third: how often does each assumption need checking? Some assumptions can move in a week; others barely shift in a year. A quarterly review cadence applied uniformly to both is tidy and mostly useless.

In the engineering firm I opened with, the pricing assumption at the heart of their margin target was tied to raw material costs that shifted monthly. A quarterly scorecard review would have reported the assumption as stable for two full cycles after it had already broken. The scorecard should carry a different review frequency for each line, matched to the volatility of the underlying assumption rather than to the committee calendar.

When the balanced scorecard earns its keep

The Kaplan Norton balanced scorecard earns its place only after the upstream work has been done. If the assumptions behind the strategy have been named, ranked by influence and confidence, and assigned to someone who holds both the authority to act and a pre-agreed threshold that triggers action, then the four perspectives are an effective way to arrange the measures so the board can read them in one sitting.

Skip that step and the scorecard does active harm: a disciplined, multi-perspective report tells the board that governance is working, when all it is actually confirming is that execution of an untested plan proceeds on schedule.

In my own work, I ran a maturity evaluation with a client over several years using a six-principle scoring model. The client's score moved from 13.3 to 31.8 as the soundness of their decision-making process improved. That score measured the quality of the process producing decisions, not just the outputs.

A balanced scorecard built on untested assumptions cannot do this; it reports whether execution hits targets, but it has no way to assess whether the process generating those targets was itself rigorous. Most scorecards report on outputs; the few that work report on whether the apparatus generating those outputs is itself sound.

Three measures are usually enough. For any significant decision, I want to know whether the controls in place have been tested rather than merely asserted, what the organisation loses in money and time if the key assumption breaks next quarter, and whether the current level of exposure sits within what the decision-maker originally accepted.

Each of those exists because a specific decision depends on it; the rest of what a typical scorecard carries exists because a template had rows and nobody felt authorised to delete them.

The scorecards that survive are the ones built backward: from tested assumptions to measures, not from a template to a dashboard. Rank what the strategy assumes to be true by influence and confidence, and the measures that belong on a balanced scorecard fall out of that exercise; much of what a standard template would have included does not survive the cut. A worked example shows the difference between a scorecard that monitors assumptions and one that merely reports activity.

You could deploy the Kaplan Norton balanced scorecard across all four perspectives and still measure the execution of a plan nobody tested.

Work through your decision

No sign-up. Just pick your decision and start.


Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.