I have spent decades watching teams argue about whether to trust the screen or the person standing at the machine. The dashboard wins because it looks objective, or the operator wins because the room defers to experience. Neither outcome produces a decision anyone can defend after the fact, because both sides skip the step that would settle it: stating what each signal actually claims and testing the assumption behind it.
The data-driven decision making literature treats this as a question of culture: are you data-driven or are you not? That framing is useless to someone facing a specific production call on Tuesday. The real question is whether the dashboard's input covers the situation the operator is describing.
Dashboard vs gut feeling decision making is a conflict between a measured indicator and an operator's contrary observation, resolved by testing the assumption each one depends on.
Dashboard vs Gut Feeling Decision Making Starts With Two Claims
Take a director who must approve a twelve-week production run. Her Power BI page shows a 96 percent fill rate and steady margin. The plant manager says a supplier changed its resin and first-shift rejects have doubled. The dashboard supports proceeding; the plant manager wants a pause.
Before anyone ranks these two positions, record them side by side. The dashboard claim: recent production data supports a full run. The operator claim: the resin change has raised rejects enough to threaten quality over twelve weeks. Recording both claims does something that an argument about trust cannot do. It turns the dispute into a question of evidence.
I would then ask what decision rule produces the dashboard's recommendation. A 96 percent fill rate supports a full run only if it predicts the quality result after the resin change. If the metric records orders shipped before the changed lot entered production, it may be accurate and still irrelevant to the decision now pending. The distinction matters when considering whether data should inform or decide.
A 1998 study by Whitecotton, Sanders, and Norris tested this directly. In controlled prediction tasks, a mechanical aid beat human judgement when both had the same information. The combination performed best when the person held a relevant cue that the aid did not contain. The published record supports giving the dashboard more weight when an operator relies on stale experience or overfits a memorable incident. It also supports the reverse: when the operator holds a cue the model cannot see, the model alone is worse than the combination.
The study sets a burden of proof for both sides. The operator must name the cue. The dashboard must be valid for the current setting. The same observation should not get two votes simply because one version is described as experience.
Take the dashboard reading your team trusts most and write the assumption it depends on, then name the operator cue that could change the call. Start the Walk →
Put the Operator's Cue Into Words
An operator's objection becomes useful when it identifies a specific change and a test that could reduce its weight. For the resin issue, the claim must identify the changed supplier lot. A comparable sample must measure rejects against the prior lot. The check moves discussion from rank or confidence to a physical observation that can be verified.
The Epic Sepsis Model shows what happens when a dashboard's authority is assumed rather than tested. Epic Systems deployed a sepsis prediction score across hundreds of hospitals. When researchers at Michigan Medicine tested it externally across 38,455 hospitalisations, the model missed 1,709 sepsis cases at the selected threshold and detected 33 percent of the total. Clinicians at the bedside, who could see the patient's condition changing, were working alongside a score that had never been validated in their facility. The JAMA Internal Medicine validation found that a score adopted widely needed local validation before it could direct clinical action in a new setting.
The lesson generalises. A model built on one population, or one production run, or one resin batch, is a claim about a specific set of conditions. Use it outside those conditions and its output becomes a number with a history but no current authority. A forecast calibrated on earlier production may fail after a supplier change, even when every number on the screen is current.
Record each override beside the dashboard output, then review whether the operator's cue improved the eventual result. Repeated overrides that add no predictive value should lose weight over time. This is how a team prevents a vivid incident from becoming permanent authority over every future call.
Set a Threshold for Changing the Call
A threshold names the evidence that would change the commitment. It also assigns authority to act when the evidence arrives. The threshold should reflect the consequence of being wrong and how easily the decision can be reversed.
BP's Texas City refinery illustrates what happens when the threshold measures the wrong outcome entirely. Before the March 2005 explosion that killed 15 workers and injured 180, executives relied on personal-injury frequency rates as their principal safety measure. The numbers looked good. Personal injuries were declining. The US Chemical Safety and Hazard Investigation Board later concluded in its final investigation report that personal-injury rates were not reliable indicators of major-accident risk. The board also found poor analysis of investigation information and a corporate emphasis on personal safety statistics that diverted attention from process safety.
Texas City was not a dashboard-vs-operator conflict in the narrow sense. It was a metric-governance failure. But its lesson applies here directly: a measurement needs authority over the outcome the decision must protect. A personal-injury rate answered a question smaller than the one the refinery needed answered. The metric that mattered, process-safety incidents, was not on the dashboard at all.
For the resin decision, the director could authorise a three-day trial instead of the full twelve-week run. The plant manager takes a 50-unit sample at the end of each shift. A reject rate above 2 percent stops the trial and returns approval to the director. That threshold is testable within three days, reversible before the commitment grows, and grounded in the specific claim the operator made.
Monitor the Smaller Commitment
Put the monitoring fields in the decision record before approval. Record the signal and its threshold, then assign an owner with the authority to act when the threshold is crossed. The daily sample is the review point because it tests the resin claim directly. A routine dashboard refresh merely supplies another reading of the metric the operator has already challenged.
If the daily sample produces results below the threshold for all three days, the trial converts to a full run and the dashboard resumes its authority as the primary signal. If the threshold is crossed, the trial stops and the director reassesses with the new evidence. Either way, the record shows what happened and why.
The Universal Decision-Making Method treats monitoring as part of the decision because the evidence supporting a commitment can change after the commitment is made. The monitoring design should match the speed at which the critical assumption can move. For a resin change, that speed is measured in shifts, not quarters.
The Decider owns a defensible decision when the record shows what would change the call and who has the authority to act.
You could read the dashboard and still leave the operator's cue unrecorded beside the metric it challenges.
Work through your decisionNo sign-up. Just pick your decision and start.
Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.