After a systematic review, most teams move straight to a certainty rating and a recommendation. The step they skip is testing whether the studies the review could see were the studies that mattered.

A systematic review identifies, appraises, and synthesises every eligible study on a defined question using pre-specified, reproducible methods, often pooling the results in a meta-analysis.

The standard next step after a systematic review

The route from review to action is well mapped. The Cochrane Handbook for Systematic Reviews of Interventions sets out how reviewers extract data, assess risk of bias, and summarise findings. The output is usually a pooled effect estimate, a forest plot, a heterogeneity statistic, and a summary of findings table. The next job is to decide how much confidence that estimate deserves.

Most guideline bodies use GRADE for this. Guyatt et al. (2008) describe the approach: evidence from randomised trials starts at high certainty and is rated down for risk of bias, inconsistency, indirectness, imprecision, and publication bias. Observational evidence starts low and can be rated up. The result is one of four grades: high, moderate, low, or very low.

What to do after a systematic review: test whether the visible evidence represents all the evidence before acting on the certainty grade
The review pools what was published and grades what it pooled. What never reached the record stays outside the grade.Click to expand

The graded evidence then moves into an evidence-to-decision framework. Alonso-Coello et al. (2016) set out the GRADE structure: the size of benefits and harms, certainty of evidence, patient values, resource use, equity, acceptability, and feasibility. A guideline panel works through each criterion and issues a recommendation, strong or conditional, for or against.

Implementation follows. The recommendation becomes a clinical guideline, a formulary listing, a procurement contract, or a public health policy. Outside medicine, the same sequence runs in policy units and strategy teams that commission evidence reviews before committing budget, the practice described in evidence-based decision making and central to data-driven decision making more broadly. By the time a recommendation reaches the person who approves the spend, the review has been compressed into a single grade and a single sentence.

What that step adds

The certainty rating adds discipline. GRADE forces a panel to state why it trusts or distrusts the evidence, domain by domain, instead of leaning on the reputation of a journal or the size of an effect. Two panels reading the same review can see exactly where they disagree and why.

The evidence-to-decision framework adds what the review cannot supply, which makes it a working answer to whether data should decide or just inform. A review answers whether an intervention works. It does not answer whether it is worth the money, whether the people affected want it, or whether a health system can deliver it. Separating "does it work" from "should this organisation do it" is the move most decision processes never make. Board papers routinely blur the two, which is why the question of which numbers matter before the call so often goes unasked.

The process also leaves a trail. When a recommendation is challenged years later, the panel can show which studies it relied on, how it rated them, and which trade-offs it accepted. That transparency is rare in organisational decision-making and worth copying outside clinical settings.

The limit is structural. Publication bias is one of the five GRADE downgrade domains, but it is usually judged from funnel plots and statistical tests run on the published record itself. Selective outcome reporting inside a published trial is harder still to detect. A certainty rating can only grade the evidence that was put in front of it.

Rewrite the pooled estimate as a claim about your own population and test it before the recommendation commits budget to someone else's trials. Start the Walk →

Where the standard playbook breaks down

In the years around the 2009 H1N1 influenza pandemic, governments stockpiled oseltamivir, sold by Roche as Tamiflu. The case for stockpiling rested on more than shorter illness. It rested on the claim that the drug reduced serious complications such as pneumonia and hospital admission. A key source for that claim was Kaiser et al. (2003), a Roche-supported pooled analysis of ten treatment trials, most of which had never been published in full. The UK alone spent more than £400 million on the drug, according to the House of Commons Public Accounts Committee.

In 2009, after a Japanese paediatrician, Keiji Hayashi, pointed out that the complications finding leaned on unpublished trials, the Cochrane team updating its review tried to verify the data. It could not. The reviewers asked for the full clinical study reports, the regulatory documents behind each trial. Obtaining the complete set took until 2013.

What a review producesWhat it assumedGap to test
Pooled reduction in complications across treatment trialsThe trials in the pool were all the trials run, and their summaries matched the underlying dataObtain the clinical study reports, or confirm the unpublished trials say what the summaries say
No significant heterogeneity across trialsAgreement between trials means the effect is real, not that the trials share a sponsor and a blind spotCheck whether every trial shared a funder, a design choice, or an outcome definition
Benefit estimated from seasonal influenza trialsThe effect transfers to a pandemic strain and to the high-risk patients the stockpile is meant to protectTest whether the trial population resembles the population the purchase targets
Complications counted as reported pneumoniaUnverified diagnoses measure the outcome a stockpile buyer cares aboutConfirm the measured outcome is the one the decision turns on: admissions avoided, deaths prevented

Working from those reports, Jefferson et al. (2014) found that oseltamivir shortened the time to first symptom relief in adults by about 16.8 hours, from roughly seven days to 6.3. It did not significantly reduce hospital admission in adults. The apparent reduction in pneumonia rested on unverified diagnoses and was not statistically significant in the five trials that used a more detailed diagnostic form. The drug increased nausea and vomiting.

The finding remains contested. A 2015 analysis in The Lancet, pooling individual patient data from the Roche trials, reported fewer lower respiratory tract complications among treated adults. Reasonable reviewers still disagree about the size of the benefit. That dispute is not the lesson. The US Food and Drug Administration, which had seen the same trials, had already required the drug's label to state that a reduction in complications had not been shown. The stockpile decisions were taken before the buyers had checked whether the published record matched the underlying trial data.

No step in the standard playbook was skipped. Reviews were done, advisory bodies weighed the evidence, governments bought. The failure sat in an assumption none of those steps was designed to test: that the visible evidence represented all the evidence. The same shape recurs across data-driven decision making examples where the numbers were real and the call was still wrong. The assumption that the published trials were representative carried hundreds of millions of pounds, and the buyers never examined it before the money was committed.

The step to take first

The review is not the problem, and neither is GRADE. The problem is treating the certainty grade as the end of scrutiny. Before a review's conclusion drives a commitment, the assumptions underneath it need naming, and the decision-maker needs to know which of them the commitment actually depends on.

Assumptions behind a systematic review's conclusion
The search strategy was pre-specified and reproducible
Each included trial was assessed for risk of bias
The published trials represent all the trials that were run
Journal summaries match the underlying study reports
The trial population resembles the population this decision affects
The measured outcome is the outcome this decision turns on

The review informs the decision. It does not make it. The five steps of the Sufficient Certainty method put that distinction into practice. Frame the decision the review is meant to serve: a stockpile, a formulary listing, a program rollout. Separate the tentative elements, the quantities, timing, and options still open. Surface the assumptions, including the ones the review cannot see from inside its own methods. Decide what level of certainty is sufficient for this commitment, which for an irreversible purchase in the hundreds of millions sits higher than for a reversible pilot. Then implement with monitoring built in, including a trigger to revisit the decision if the underlying data surface.

An assumption that shapes the commitment and cannot be checked from the published record is the one to test before acting. Sometimes the test is a request for unpublished data. Sometimes it is a smaller first purchase with a review date. Sometimes it is simply writing the assumption down so that the next person to read the grade knows what it rests on.

A high certainty grade answers how far the visible evidence can be trusted. It does not answer what the evidence left out.

You could close this tab and carry that decision into another week.

Work through your decision

No sign-up. Just pick your decision and start.


Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.