Zillow had plenty of data when it wrote down roughly $304 million in housing inventory in November 2021 and then shut its home-buying operation. If that is your picture of big data decision making, the problem is plain enough. Volume made the call look industrial. It did not make the assumptions about price movement, timing and resale sound.
I have sat through too many post-mortems where someone points at the dashboard as if that settles the matter. It never does. Data can tell you a great deal about what has happened and something about what may happen next, but it cannot carry the burden of judgement for you. The room still has to decide what matters, which assumption matters most, and what signal would prove the basis of the decision has started to rot.
Big data decision making is the use of large and varied datasets to test the assumption carrying a decision.
Big Data Decision Making Starts with the Wrong Question
The common question is, "What more can we measure?" The useful question is, "What has to be true for this decision to work?" Answering the second is also how you work out whether the numbers should decide or only inform the call in front of you. Those are not the same question, and the difference is where most teams lose their footing. If you start with the pipeline and the dashboard, you finish with a report. If you start with the decision, you at least have a chance of testing the right thing.
Google Flu Trends is still one of the cleanest warnings I know. The system had scale coming out of its ears, yet it badly over-predicted flu levels because the model chased patterns that did not hold up. The authors pointed out that Google had effectively fitted about 50 million search terms against 1,152 data points. It shows a discipline failure, not a data shortage.
The same trick turned up in Amazon's abandoned recruiting tool. Reuters reported that the system had been trained on ten years of resumes, most of them from men, and it learned exactly the lesson it had been given to learn. It treated the past as a permission slip for the future. More historical data simply made the bias look better supported. That is why I treat Amazon's recruiting failure as a decision failure before it is a model failure. The broader case against treating evidence as a verdict runs through data-driven decision making as a discipline. Data earns its place only when it tests the live assumption under the decision.
Big Data Decision Making Breaks After Approval

Most big-data systems do not create their worst trouble at the moment of approval. They create it afterwards, when everyone assumes the dashboard will do the watching. I have chaired reviews where the display stayed green while the instrument had drifted, the software had changed or the person reading the output no longer understood the proxy being used. A clean screen can be as superstitious as a clean audit report.
That is why the phrase "we are data-driven" often tells me very little. It may mean the organisation has a large estate of reports and model outputs. It may also mean nobody feels personally responsible for reopening the decision when the basis changes. Oracle found in 2023 that many business leaders were slowed or stopped by the sheer volume of data rather than improved by it. Dashboard owners, platform vendors and executives who prefer an instrument panel to an argument all benefit from that glut. The apparatus is bigger, the thinking is thinner.
Roger Estall and I built the Universal Decision-Making Method around a plain fact: decisions live on after the meeting. If the decision depends on an assumption about demand or supplier reliability, then someone must monitor the signal that would show that assumption has stopped holding. Without an owner for that signal, the organisation is just producing reports. Big data decision making fails at exactly this point: volume after approval without a named person watching the proxy that matters.
I see this same pattern in teams suffering data paralysis. They think the cure is another feed or a fresher dashboard. Usually the real problem is that they never wrote down what event would cause them to revisit the call in the first place. A monitoring design beats another pile of data almost every time.
Check your model-backed decision through the five steps and record the assumption your dashboard treats as settled fact. Start the Walk →
Big Data Decision Making Still Rests on Assumptions
Zillow is a useful case because nobody can claim it lacked volume. The company had transaction and market data in abundance, yet in late 2021 it disclosed a roughly $304 million write-down and wound down Zillow Offers. A forecasting premise about home prices and resale speed could not bear the volatility it was asked to bear, and no amount of additional data would have rescued that premise. Zillow's own results release makes the scale of the assumption failure plain.
I take the same view of IBM Watson for Oncology. The system was sold with the glamour that usually comes with giant data stories, yet internal material reported by STAT described unsafe or incorrect treatment recommendations and narrow training foundations. The system was dressing weak judgement in expensive machinery. If the training base is thin or synthetic, scale in the wrapper does not rescue what the model actually knows. The STAT report on Watson for Oncology makes grim reading.
This is where most writing on big data decision making goes soft. It says the answer is better analytics and better governance, then quietly leaves the Decider out of it. Those are comforting answers because they upgrade the apparatus without forcing anyone to own the call. A person still has to say which assumptions are carrying the result and what would count as disconfirming evidence. The question is not "How much data have we got?" It is "Which assumption is still unresolved, and what evidence would reduce it enough to act?" Sometimes that means gathering more information. Just as often it means narrowing the commitment or rejecting the option altogether. Roger Estall and I wrote Deciding around that distinction, and it has not changed.
What I Record Before I Trust a Model
I want the model's proxy named in plain English, not buried in a technical specification. I want the assumption that proxy rests on visible enough for an outsider to challenge. In a big-data context that means asking what the training base actually represents, how often the underlying variable is recalibrated, and who has the authority to override the output when the proxy drifts. Once those questions are on the table, the room stops admiring the machinery and starts discussing what would have to change in the world for the model to mislead.
After that, two things need settling. First, what is enough information to make a decision now, given the data already available? Second, what drift in the proxy or the input would trigger a reopening? Those are different questions, and most teams muddle them. If your committee neglects the second, read monitoring a decision. Both problems are more common than any shortage of data.
I do not object to big data. I object to the theatre that grows around it. Platform vendors sell certainty and data custodians sell the comfort of a complex apparatus that nobody wants to question. I have watched that performance for years in other forms, from audit reports to forecasting models. The machinery gets more impressive. The willingness to own the call does not.
When a team tells me it needs more data, I ask which assumption the current model is protecting and who gets cover from delaying that admission. Big data decision making becomes useful at the moment someone names the proxy, tests the assumption, and accepts that the model can be overridden. Until then it is theatre with better production values.
You could scale the data pipeline and still carry the same untested premise.
Work through your decisionNo sign-up. Just pick your decision and start.
Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.