After risk evaluation, most teams move straight to selecting treatments. The step they skip: testing whether the criteria they just evaluated against still reflect the conditions they are actually facing.
Risk evaluation compares the results of risk analysis against established criteria to decide whether a risk needs treatment.
The standard next step after risk evaluation
ISO 31000:2018 positions risk evaluation as the final stage of risk assessment, sitting between analysis and treatment. The evaluation takes analyzed risks and compares them against predetermined criteria. The output is a verdict: accept the risk, treat it, undertake further analysis, maintain existing controls, or reconsider objectives entirely.
In most organizations, this comparison produces a rating. High, medium, low. Red, amber, green. Acceptable or unacceptable. The rating then drives the treatment response, and the treatment response is assumed to be proportionate because the rating said so.

The practical mechanics vary. Regulated sectors such as nuclear, aviation, and pharmaceuticals require committee sign-off before treatment plans are approved. Project-based work often collapses evaluation and treatment into a single session where risks are rated and mitigations assigned in the same afternoon. In either case, the transition follows the same logic: if the evaluation says "unacceptable," something must change; if it says "acceptable," the organization proceeds.
Teams that follow structured risk assessment processes will document the verdicts in a risk register, assign owners, and schedule review cycles. Less formal operations skip the paperwork but follow the same pattern. The evaluation verdict sets the agenda. Treatment planning responds to it. And the criteria behind the verdict are rarely revisited once they have been set.
This is the standard choreography. Aven (2016) describes it in a review of risk management foundations. Every corporate risk framework that traces its lineage to ISO 31000 follows the same sequence: evaluation yields a verdict, the verdict yields treatment, treatment yields monitoring. Forward, always forward.
What that step adds
This sequence is not pointless. Evaluation against criteria gives decision-makers a consistent threshold for action. Without it, every risk conversation becomes an ad hoc negotiation about what counts and how much is too much. Criteria provide a shared language, and shared language prevents the loudest voice in the room from setting the tolerance by default.
Structured evaluation also creates a record. When an organization documents that a risk was rated "medium" against specific criteria, it leaves a trail that future teams can revisit. This is genuinely useful for building risk culture across divisions that would otherwise operate on incompatible assumptions about what counts as significant.
| What risk evaluation produced | What it assumed | Gap to test |
|---|---|---|
| Risk rated "high": treatment required | Likelihood and consequence scales reflect current operating conditions | Whether the scales were calibrated to conditions that still apply |
| Risk rated "acceptable": no action needed | Nothing material has changed since the criteria were set | Whether regulatory, market, or operational context has shifted since the criteria were drafted |
| Treatment priority ranking across the register | Criteria weightings correctly express organizational priorities | Whether those priorities have been revisited since the original weighting exercise |
The comparison step also forces prioritization. When there are more risks than resources, evaluation criteria create a ranking that, however imperfect, prevents organizations from treating every identified risk with equal urgency. Sorting precedes spending, and sorting against explicit criteria is better than sorting by recency or seniority.
None of this is trivial. The problem is not that risk evaluation lacks value. The problem is that teams treat the evaluation verdict as a conclusion rather than as an interim finding that depends entirely on the quality of the criteria it was measured against. Risk evaluation compresses a complex judgment into a manageable form. The compression itself is valuable because it allows organizations to compare unlike risks, allocate finite resources, and communicate priorities across reporting lines. The vulnerability emerges not from the compression, but from the failure to test what got compressed out.
Rewrite the evaluation criterion your highest-rated risk depends on as a claim and test it before the treatment plan commits resources to a verdict. Start the Walk →
Where the standard playbook breaks down
The evaluation verdict is only as reliable as the criteria it was measured against. And criteria are constructed from assumptions: assumptions about what has happened before, what is likely to happen next, and what the organization can tolerate if it does.
When those assumptions are wrong, the evaluation produces a confident answer to the wrong question.
Fukushima Daiichi is the defining modern example. Before the 2011 disaster, Tokyo Electric Power Company (TEPCO) maintained a risk evaluation framework for its nuclear plants that included seismic and tsunami hazards. The criteria for tsunami risk were derived from historical records going back roughly 400 years. On those records, the maximum credible tsunami height at the Daiichi site was set at 5.7 metres. The seawall was built to that height. The evaluation said: acceptable.
Geological evidence told a different story. Deposits from the 869 Jogan tsunami, documented in peer-reviewed research, indicated that waves had reached far inland along the same coastline. In 2008, TEPCO's own internal study calculated that a Jogan-scale event could produce waves of 15.7 metres at the plant. Management acknowledged the study and decided to continue investigating rather than act. The risk evaluation criteria remained unchanged. The verdict remained "acceptable." The tsunami that struck on 11 March 2011 measured approximately 14 metres. Synolakis and Kanoglu (2015) later concluded that the accident was preventable.
The IAEA investigation found that TEPCO's hazard assessments "did not adequately account for" low-probability, high-consequence events. The criteria assumed that historical records were a sufficient proxy for geological possibility. That assumption was never formally tested against the available evidence. The evaluation process worked exactly as designed. The criteria it relied on were the point of failure.
This is not a nuclear-specific problem. Any organization that evaluates risk against static criteria inherits the same vulnerability. A risk matrix rates likelihood and consequence against thresholds someone set at some point. A risk appetite statement codifies tolerance levels that may have been drafted under conditions that no longer apply. The evaluation step does not test these foundations. It uses them.
The step to take first
Before selecting treatments, interrogate the assumptions embedded in the evaluation outputs. Not the outputs themselves, but the criteria and conditions those outputs depend on.
The five-step method (Frame, Tentative Elements, Assumptions, Sufficient Certainty, Implement and Monitor) treats this interrogation as structural, not optional. The third step exists precisely because analytical frameworks produce verdicts that look definitive while resting on conditions that no one has named.
An evaluation verdict of "acceptable" does not mean the risk is acceptable. It means the risk is acceptable if the criteria are right. The productive question becomes: what would have to be true for the criteria to be wrong?
Start with the highest-rated risks, not because they are necessarily the most important, but because their treatment plans will absorb the most resources. If the rating rests on assumptions that have not been verified, the organization may be committing budget and attention to the wrong priorities. Then examine the risks rated "acceptable," particularly those near the threshold. A risk sitting just below the treatment line may cross it if one criterion shifts.
This does not require restarting the assessment process. It requires asking, for each evaluation output, what the criteria assumed, whether those assumptions still hold, and what monitoring would detect if they shifted. Most organizations already monitor outcomes. Few monitor the validity of the criteria they used to set the thresholds. The goal is not more analysis. The goal is sufficient certainty that the analysis already done can be relied on.
Risk evaluation without assumption testing is a grading exercise performed against an unexamined rubric. The rubric may be sound. But the only way to know is to name what it assumed and decide whether those assumptions still hold.
You could accept the evaluation verdict and commit resources without testing whether the criteria it relied on still hold.
Work through your decisionNo sign-up. Just pick your decision and start.
Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.