Post-mortems are useful. Pre-mortems are cheaper. A post-mortem tells you why a project failed after the money is gone. A pre-mortem tells you why a project might fail while you can still change course. The technique is straightforward: assume the project has already failed, then work backward to identify what caused the failure. Applied to AI projects, this surfaces risks that optimistic planning systematically overlooks.
AI projects have a unique risk profile. The core assumption (that the model can perform the task well enough) is often untested until late in the project. Traditional project risks like schedule slippage and scope creep apply, but the dominant risk in AI projects is technical feasibility at production quality, and that risk is usually addressed last.
When to Run a Pre-Mortem
Run the pre-mortem after the initial scoping but before committing resources. The team has defined what they want to build and roughly how. They have not yet committed to a timeline or allocated a budget. This is the window where the pre-mortem has maximum impact: early enough to change the plan, late enough that the plan is specific enough to critique.
Do not run a pre-mortem on day one. The team needs enough context to generate specific failure scenarios, not generic ones. “The model might not work” is not a useful failure scenario. “The model cannot distinguish between two entity types that share surface-level features, causing extraction errors in 15 percent of cases” is useful because it points to a specific mitigation.
The Pre-Mortem Process
Step 1: Set the Scene (15 minutes)
The facilitator describes the project as if it has already failed. Write a brief failure narrative: “It is six months from now. The project is cancelled. The model never reached production quality. Users rejected the outputs. The budget was consumed with no business value delivered.” Read it aloud. Let the room sit with it for a moment.
The failure narrative should be specific to the project. Generic narratives produce generic risks. If the project is a document classification system, the narrative should say “The classifier achieved 70 percent accuracy, well below the 90 percent threshold required for production. The training data was insufficient for the edge cases that dominate real documents.” Specificity drives useful responses.
Step 2: Silent Brainstorm (10 minutes)
Each participant writes down reasons for the failure independently. No discussion. No ranking. Just write. This produces more diverse risks than group discussion, where early speakers anchor the conversation and quieter team members self-censor.
Ask participants to generate at least three failure reasons. Categories to prompt if the room is stuck: data quality, model capability, integration complexity, organisational readiness, user adoption, and cost overruns. Each category tends to produce distinct failure modes.
Step 3: Share and Cluster (20 minutes)
Go around the room. Each person reads one reason. Continue until all reasons are on the board. Cluster similar reasons. Do not debate whether a reason is likely yet. The goal is coverage, not consensus.
Watch for the pattern where every reason clusters into two or three categories. This means the team has blind spots. Prompt for risks outside those categories. Ask: “What would our competitors say is the reason this failed?” Outsiders see different risks than insiders.
Step 4: Rank and Mitigate (30 minutes)
Vote on the top three to five risks. For each top risk, define a specific mitigation and an owner. The mitigation must be actionable: not “improve data quality” but “run a data quality audit on the training set by week 2 and establish a minimum quality threshold of 85 percent completeness.”
Assign a trigger condition for each risk. “If model accuracy is below 80 percent after the second training run, we escalate to the steering committee and reassess scope.” Triggers turn risks from abstract concerns into decision points.
AI-Specific Failure Modes to Check
Every AI pre-mortem should explicitly address these failure modes, even if they seem obvious:
The training data does not match production data. The model performs well on the training distribution and poorly on the real inputs. This is the most common AI project failure mode and the most preventable. Mitigation: validate the training data against a sample of production data before training begins.
**The evaluation metric does not match business value. The model optimises for accuracy but the business cares about precision on a specific class. Or the business cares about latency and the model is too slow. Mitigation: define the business metric first, then map it to a model metric. If the mapping is unclear, that is a risk to flag.
The integration is harder than the model. The model works in a notebook but the production integration requires data pipelines, caching, error handling, and monitoring that were not in scope. Mitigation: estimate integration effort separately from model development effort. If integration is more than 50 percent of total effort, it should be in the project plan.
Users do not trust the output. The model produces good results but users do not believe them. They second-guess, override, or stop using the system. Mitigation: design the user experience for calibrated trust. Show confidence scores, explain reasoning, and provide easy override mechanisms from the start.
The cost exceeds the value. The model works but costs more to run than the problem it solves. Mitigation: estimate production inference costs during scoping, not after the model is built.
The Output
The pre-mortem produces a risk register with three to five prioritised risks, each with a mitigation, an owner, and a trigger condition. This register becomes part of the project plan. Review it at each milestone. Update it as new risks emerge or existing risks are mitigated.
The risk register is not a document that gets filed. It is a living artifact that drives decisions. When a trigger condition fires, the team acts on the predefined mitigation. This is the difference between a pre-mortem that changes outcomes and a pre-mortem that generates a document.
Next Step
Schedule a 90-minute pre-mortem for your current AI project. Send the failure narrative to participants 24 hours before the session. Run the process as described. Leave the session with a risk register that has owners and trigger conditions. If you cannot identify a facilitator who is not the project lead, use an external facilitator. The project lead is too close to the work to challenge assumptions effectively.