When an operational failure becomes a business crisis, the damage almost never comes from the original breakdown but from the response. Technical people solve the technical problem while the organisation stays silent; someone else fills the silence; and from that point the organisation is reacting to a crisis it did not define. That is the threshold.

Here is the test: can you explain, with evidence, why the situation is still contained? If you cannot, the failure has crossed the line.

I have investigated enough of these over nearly fifty years to see the pattern clearly. The original problem is usually well within the organisation's technical competence. In almost every case I have reviewed, the engineers or the operations team had identified the fix before the crisis was declared.

The consequences, however, had already moved into territory the technical team was never equipped to manage. Journalists were already calling, and the board wanted to know why nobody had told them.

Those questions belong in a different room with different authority, and most organisations discover this only after the wrong people have been answering them for too long.

A business crisis is an operational failure whose consequences have moved beyond what normal operations can contain, forcing decisions the existing structure was not designed for.

How a Single File Update Became a $5.4 Billion Business Crisis

On 19 July 2024, a faulty configuration update to CrowdStrike's Falcon Sensor caused 8.5 million Windows devices to crash simultaneously, an operational failure so narrow that a single file was the entire cause. The fix at the code level took hours; CrowdStrike's architecture demanded kernel-level access to the operating system, which meant each affected machine required someone to physically boot it into Safe Mode, delete the file, and restart. Delta Air Lines alone cancelled over 1,200 flights and filed a $500 million lawsuit. Fortune 500 companies reported $5.4 billion in direct losses.

I have sat with boards that approved single-vendor concentration without ever questioning what happens when the vendor fails in a way the architecture cannot absorb. In CrowdStrike's case, the decision to centralise endpoint security carried an assumption nobody had named: that the vendor's update process would never fail catastrophically.

The vendor who sold the contract had no incentive to name this assumption. Nobody selling a security product volunteers the conditions under which it will fail. The procurement team who simplified the estate had no reason to test it.

When the assumption broke, the operational failure was already a business crisis before any affected organisation knew it had happened. That is the mechanism worth examining.

Threshold test for operational failure versus business crisis: two common excuses above the line, the evidence-based containment question below
The threshold test: if the answer requires speculation, the crisis has already begun.
Click to expand

CrowdStrike is the extreme case because the architecture made the threshold invisible; containment was structurally impossible once the update shipped. Most operational failures do not cross the line that way. They cross it through the response, which means the threshold is often within the organisation's control, provided someone is prepared to lead when the facts are still changing.

Name the assumption your containment story depends on and test whether the evidence supports it before the board meeting asks you to. Start the Walk →

When the Response Turns an Operational Failure Into a Crisis

In February 2018, KFC UK switched its logistics from six regional warehouses operated by Bidvest to a single DHL depot in Rugby. Within days, 604 of 870 restaurants had closed. A chicken restaurant with no chicken is an operational failure with obvious potential to become something worse; KFC contained it by acknowledging immediately that the assumption behind the switch had failed and reversing the decision. Within weeks, 97% of stores were back open. Same-store sales dipped 2% and operating profit took a 5% hit for the quarter.

I have reviewed hundreds of incident responses over the years, and the ones that work share a single feature: someone in the room recognised early that the situation had moved beyond normal operational channels and had the authority to say so. KFC did not succeed because it had a superior crisis management strategy; it succeeded because the reversal came before the consequences could outrun the response. The company acknowledged the failure publicly (including an advertisement that rearranged the letters on an empty bucket to read "FCK"), but the communication worked because the operational fix preceded it.

Toyota's trajectory over the preceding decade was the opposite. Between 2004 and 2009, the company received complaints about unintended acceleration across multiple vehicle models and largely dismissed them as driver error. A limited 2007 recall addressed floor-mat interference in select vehicles. In August 2009, a fatal crash killed four people, and Toyota's crisis began in earnest; by early 2010, approximately nine million vehicles were recalled, executives were testifying before the United States Congress, and the company eventually paid over $1 billion in settlements.

The operational failure (a defect affecting the accelerator mechanism) had existed for five years before the crisis arrived. I have watched this pattern in organisations far smaller than Toyota: complaints accumulate and the people closest to the product insist their analysis is complete. Nobody upstream wants to hear otherwise. Everyone in the chain has a reason to keep calling it operational, and nobody has an incentive to call it a crisis.

Toyota dismissed complaints for years, then offered explanations that contradicted each other in rapid succession. Floor mats, then sticky pedals, then electronic systems.

The narrative vacuum was filled by Congressional investigators and journalists who had no reason to give Toyota the benefit of the doubt. The threshold was crossed not when the defect was discovered but when the public lost confidence that Toyota understood its own product. That is what five years of silence purchases.

The Assumption That Was Never Written Down

Roger Estall and I kept encountering the same structural failure across industries while writing Deciding: decisions resting on assumptions nobody had articulated, let alone monitored. The conventional taxi industry illustrates the slow-burn version of this threshold crossing. For decades, operators assumed that regulation protected their market position and that customers would tolerate poor reliability and no visibility of when a car would arrive. Those were not risks on any register. They were structural assumptions about how the industry worked, and nobody wrote them down because the people who profited from the arrangement had no reason to question it.

The threshold crossed not when Uber arrived but years earlier, in the gap between "customers complain about this" and "those complaints describe a service so poor that any competitor who solves them will win." The smartphone put real-time tracking and cashless payment in every urban commuter's pocket, and Uber assembled those capabilities into a service that made the taxi industry's assumptions visibly false. By the time operators recognised the threat, the operational failure had already become a business crisis; the context had changed beyond recovery.

The Universal Decision-Making Method would have required the taxi industry to do something none of its beneficiaries wanted: name the assumptions the business depended on and design monitoring that would tell them when those assumptions no longer held. The taxi industry did none of this. Neither, in a different way, did the organisations that gave CrowdStrike kernel-level access without ever testing the assumption that sat behind it.

The Threshold Test

If something broke and the CEO meeting is tomorrow morning, finish this sentence:

"This remains operational because _____ still holds."

If you can fill in the blank with evidence, your problem is operational. If you find yourself reaching for "these things usually resolve themselves" or "we have a team working on it," you have answered a different question: whether you are doing something about the failure, not whether the failure is contained. That is the distinction most organisations miss.

In my experience, the delay in naming a crisis is rarely accidental. The operations team wants to fix it without outside interference. The executives want to avoid telling the board. Both have perfectly good reasons to keep calling the problem "operational," and neither reason has anything to do with whether the situation is actually contained.

The question, in the method's terms, is not whether this is definitely a crisis (that standard demands certainty nobody has 48 hours into a failure) but whether you have sufficient certainty that it is not. If the answer requires speculation rather than evidence, the situation has likely moved beyond operational containment.

CrowdStrike's customers assumed update validation would hold. Toyota assumed its engineering analysis was complete. The taxi industry assumed regulation would protect its market. In each case, the organisation's ability to contain the failure depended on an assumption that had never been tested against reality. KFC contained its failure precisely because it identified the failed assumption and reversed the decision before the consequences could escalate beyond the operational room.

Most leadership under pressure fails not because the leader lacks composure but because the assumptions behind the original decision were never visible when the pressure arrived. If you can name those assumptions now and demonstrate that they still hold, your problem is operational and you can say so with confidence when the CEO asks. If you cannot, you are in a situation that needs executive decisions, and the sooner you name it, the sooner the right people can start working on it.

You could keep calling it operational until someone outside the building names it for you.

Work through your decision

No sign-up. Just pick your decision and start.


Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.