An ai decision making tool will rank your options and deliver a recommendation. ChatGPT and Claude will do it in seconds. None of them will name the assumption behind that recommendation or assign anyone to watch for the moment it breaks. The output is preparation. It is not the decision.
I have watched this gap persist for fifty years under different labels. The label used to be "risk management": risk registers and the consultants who sold them. The apparatus kept growing and decisions kept breaking on the same failure, because the assumption that mattered was never surfaced and never owned.
AI has inherited that problem and accelerated it. I built the Walk to force the step that every tool, old and new, still skips: surface the assumption behind the preferred option and assign someone to watch for the change that would reopen the call.
An AI decision making tool is software that uses machine learning or large language models to process information, score options, or generate recommendations as inputs to a human decision.
What AI decision tools cover
These tools are genuinely useful at a specific job: they process volume that no person can hold in working memory, ranking thousands of data points by weighted criteria and modelling permutations that would take a team weeks to run by hand. That is real value, and I have no interest in dismissing it.
AI earns its place in the preparation before the decision, not in the decision itself. It can organise information and present the options in a form the room can work with. What it cannot do is judge whether the assumptions underneath the preferred option are sound for this situation, in this context.
Ask ChatGPT to evaluate three warehouse sites for a distribution centre. It will score each location on a dozen criteria and recommend Site B with a confident summary. The output assumes last year's freight costs predict next year's, that the planning approval will hold, and that the labour market in that region will staff the facility. None of those assumptions appear anywhere in the response. The tool processed the data you gave it and returned a ranked list. The ranked list is preparation. The decision has not happened yet.
That judgment is human work regardless of how much data the platform has processed. The question is not whether data should decide or just inform; it is whether anyone in the room has tested what the data assumes.
When an AI decision making tool is enough
If the choice is bounded and the downside is containable, an AI tool is probably sufficient on its own. Sorting job applicants by published requirements or optimising a delivery route against known constraints: these are legitimate uses where the room has enough information to make a decision and the cost of a wrong call is recoverable.
The line moves when the commitment becomes material. A plant closure or a market entry rests on assumptions about future conditions that the data cannot settle, because the demand forecast is still a guess and the competitive response is unknown. At that point the ai decision making tool has finished its useful work.
I do not say this because I distrust the technology. I say it because I have watched rooms full of competent people defer to a system output rather than state the assumption they were actually worried about.
That pattern predates AI by decades. Matrices and scoring models produced the same deference long before any algorithm was involved. AI simply makes the deference feel more scientific, which makes it harder for anyone in the room to challenge. Challenging a machine output feels less legitimate than challenging a colleague's opinion, even though the machine output rests on assumptions that are far less visible than the colleague's reasoning.
Name the assumption underneath the AI recommendation your team is about to accept, and ask whether anyone in the room has tested it. Start the Walk →
What AI decision tools skip
In July 1988, the USS Vincennes was operating in the Persian Gulf when its AEGIS combat system identified Iran Air Flight 655 as a civilian airliner climbing on a standard commercial route with a civilian transponder code. The crew rejected the system's correct identification and fired. 290 people died.
A Georgetown CSET report on automation bias examined this and comparable incidents and found that the failure works in both directions: people trust the system when they should challenge it, and override the system when they should trust it. Neither path tests the assumption that actually matters.
The crew had an unstated assumption: this contact is hostile because we are in a combat situation. The system had a different conclusion based on transponder data and flight path. Nobody in that chain asked the question the Universal Decision-Making Method puts at the centre of every decision: what are we assuming here, and how confident are we in that assumption?
The system could not ask it. The crew did not ask it. That is not a technology problem. It is a decision problem.
The same pattern runs through modern AI tools. Paste a vendor shortlist into Claude and ask it to recommend one. You will get a thorough comparison, clean formatting, and a confident recommendation. What you will not get is the sentence: this recommendation assumes your largest customer's volume stays flat. That is the assumption the procurement team was actually worried about, and no prompt will produce it, because the tool does not know it exists.
When ProPublica investigated the COMPAS recidivism algorithm in Broward County, the judges treated a proprietary recidivism score as neutral input while the assumption inside it, that historical arrest patterns predict individual behaviour, stayed buried. The mechanism is the same whether the tool is a recidivism score, a weighted spreadsheet, or a ChatGPT output: it looks authoritative, and the confirmation bias of a room that has already formed a preference does the rest.
Amazon spent four years building an AI hiring system trained on ten years of resumes before discovering that it penalised the word "women's" and downgraded graduates of women's colleges. The training data reflected the existing gender composition of the workforce. MIT Technology Review reported that the team could not make the system gender-neutral and disbanded.
Four years of engineering to reproduce, at scale, the one pattern the project was supposed to eliminate. The assumption that past hiring patterns reflect merit rather than structural bias was so basic it was invisible. The ai decision making tool could not surface it because the tool was built on top of it.
The vendors selling these platforms have a commercial interest in the recommendation being the product. The moment a buyer realises that the recommendation is not the decision, that the real work starts where the platform stops, the value claim shrinks.
I watched the same incentive structure in risk management consulting for decades: the apparatus grew because the consultants who built it also sold it, and their clients' interests became merged with or supplanted by the consultants' own. AI decision tools have inherited that structure. The platform processes your data and delivers a ranked output that looks like it has done the work. The assumption it did not test is invisible by design, because surfacing it would reveal that the expensive output is still an input.
What the Walk produces
The Walk exists because this gap keeps producing the same failure across every domain. It does not compete with AI on data processing. It forces the step AI cannot perform: before the room accepts or overrides the system output, it structures a conversation that asks what is the assumption here, and what evidence would change your mind.
The output is a Decision Record. It is short because every field earns its place: the decision and its Purpose at the top, the dated Context that tells the reader how fresh the evidence still is, then the critical assumptions named and ranked by influence and confidence. The last section names the monitoring signal: the threshold or event that would force the room to reopen the call, and the name of the person responsible for watching it.
Take the procurement team from the Claude example above. The Decision Record names the assumption: projected volume from Customer X holds within 15% of current run rate. The monitoring signal: the account manager reports a scope reduction or a competitive review. The person watching: the CFO. That is what no ai decision making tool delivers. Not a recommendation. A record of what was assumed and what would trigger a change of course.
Research from Harvard Business School found that AI does not compensate for weak judgment. It multiplies whatever judgment the room already has. People with sound business judgment make better decisions with AI. People with weak judgment do not improve by adding a platform.
The Walk addresses that finding directly: it does not assume the room has good judgment. It gives the room a structured method that forces the judgment to be exercised and owned. The Decision Record captures the result so that a later review can distinguish what was known from what was assumed.
The gap between data and judgment was already wide before AI arrived. That is why Roger and I wrote Deciding. AI has made data faster and more abundant, but it has not made judgment easier. The volume of information reaching a decision table has grown while the discipline of testing what the information assumes has not kept pace.
When the room cannot name the assumption the AI output rests on, nobody in it has decided anything. The room has ratified.
The question no AI decision making tool asks
Every ai decision making tool on the market stops at the same point: it delivers a ranked output and leaves the room to commit. The step between the output and the commitment is where decisions actually happen, and that step requires one question: what are we assuming about the world that must remain true for this choice to work?
That is the question evidence-based decision making is supposed to answer. It is the question every decision making tool on page one avoids.
The Walk forces that question into every decision it touches, because in my experience no room asks it without a structure that requires it. The AI structures the conversation; the Decider answers the question and the Decision Record captures it so that monitoring has something to watch. That is the step the platforms skip, and it is the step that determines whether the decision holds or breaks.
You could trust the next AI recommendation and still carry the assumption it never tested.
Work through your decisionNo sign-up. Just pick your decision and start.
Grant Purdy is the co-author, with Roger Estall, of Deciding (2020), and the architect of the Universal Decision-Making Method.