Abstract
A premortem asks a planning team to assume that a plan has already failed and to explain why. The technique rests on prospective hindsight: imagining an outcome as certain elicits more, and more specific, causal explanations than asking what might go wrong. Language models change the economics of this exercise. They can generate diverse, detailed failure narratives for any plan at negligible cost, which makes it possible to run a premortem many times, from many perspectives, before a decision is made.
That capability creates a problem as large as the opportunity. A model's fluency at explaining how a plan could fail is not evidence that it can predict how plans do fail. A plausible narrative is not a validated forecast, and we identify no current practice of AI-assisted premortems that distinguishes the two.
This paper proposes Failure-First Forecasting (FFF), a protocol that converts prospective hindsight into a falsifiable decision procedure. A plan is frozen at an information cutoff; a baseline outcome forecast is recorded; the plan is assumed to have failed and causal failure pathways are generated across multiple isolated runs; each pathway is characterized by probability, consequence, detectability, interventionability, timing, and observable precursors; interventions are proposed that target specific causal links; and the modified plan is subjected to the premortem again. This forecast–intervene–reforecast loop, which the paper names Iterative Premortem Hardening, terminates when the marginal value of another cycle falls below its cost or when remaining exposure is better addressed by structural resilience than by further mechanism-specific prediction.
The central methodological contribution is an evaluation design that separates three questions — whether AI-generated failure pathways correspond to mechanisms that subsequently occur (predictive validity), whether ex-ante interventions target links that later matter (diagnostic validity), and whether applying those interventions improves realized outcomes (intervention efficacy) — and states which evidence can answer each. Historical backtesting answers the first and partially informs the second; the third requires simulation, controlled intervention, or prospective causal evidence. Because modern language models may have been trained on the outcomes of the historical decisions used to test them, temporal isolation — preventing foreknowledge contamination — is treated as a first-order requirement rather than a caveat.
The paper closes with two secondary constructs. Residual epistemic exposure names the failure-relevant uncertainty that remains after the accessible failure space has been explored, with cross-run divergence and related signals proposed as indicators rather than measures. Structural resilience names the class of plan properties — staging, reversion, optionality, containment, redundancy, sensing — that limit damage without requiring the failure mechanism to have been predicted. The paper hypothesizes that as residual epistemic exposure rises, rational mitigation shifts from mechanism-specific intervention toward mechanism-agnostic resilience. No empirical result is reported; every quantitative claim is labeled as proposed, hypothesized, or established by prior literature.
Introduction
Plans fail through assumptions that were never tested. A launch assumes demand that does not arrive; a migration assumes a dependency that turns out to be undocumented; a policy assumes compliance that the incentives do not support. In retrospect the violated assumption is usually visible, and often it was visible in prospect to someone who was not asked. The discipline of anticipating failure exists to ask.
The premortem is the most direct instrument for doing so. Klein (2007) formalized it as a planning exercise: before commitment, the team is told the plan has failed, and each member writes down why. Its psychological basis is prospective hindsight — Mitchell, Russo, and Pennington (1989) found that treating an outcome as certain rather than merely possible produced more numerous and more specific causal explanations. (The figure of roughly thirty percent more reasons that circulates in secondary accounts, including Klein's, is attributed to this work; this paper does not restate it as a primary result without verification against the original.) Subsequent quantitative work by Peabody (2017) tested the technique in three experiments with differing results: in a field study with Army cadet teams, premortem groups showed fewer fouls and less fixation with no increase in planning or execution time; in a laboratory comparison against a worst-case-scenario method on engineering plans, no statistically significant difference in the number of reasons or solutions emerged, although the two methods produced different category distributions; and in a third experiment on complex, unfamiliar plans, premortem participants generated significantly more reasons than worst-case-scenario participants. The technique is supported, but not uniformly, and the conditions under which it outperforms alternatives are not settled.
Language models make the exercise cheap to repeat. A single plan can be subjected to dozens of isolated premortem runs, from different analytical stances, at a cost that would have been prohibitive for a human team. Early comparative work (Veinott and Lehman, 2024) has begun to examine how model-generated failure reasons compare with those produced by people. The opportunity is real: breadth of enumeration is exactly the dimension on which human premortems are limited, and it is exactly the dimension on which models excel.
The opportunity carries a specific hazard. Language models are trained to produce plausible text, and a plausible failure story is easy to produce for any plan. The question that matters for a decision is not whether a failure narrative is coherent but whether it identifies a mechanism that will actually operate. These are different properties. The literature on language-model forecasting (Halawi et al., 2024; Schoenegger et al., 2024) shows that models can approach the accuracy of human forecasters on well-defined binary questions under carefully controlled conditions — and also that those conditions are difficult to satisfy, that calibration is uneven, and that any evaluation on historical events must confront the possibility that the model has already seen the answer. None of this transfers automatically to the open-ended, mechanism-level prediction that a premortem demands.
The gap this paper addresses is therefore methodological. In the literature reviewed here we identify no established protocol for asking whether an AI-assisted premortem forecasts anything, as opposed to narrates; no agreed way to score a predicted causal pathway against a realized one, to measure whether a proposed warning signal would have provided lead time, or to test whether an intervention proposed before execution targeted the mechanism that later caused the failure; and no standard control, in the premortem literature specifically, against the most damaging confound in retrospective evaluation of language models — that the model may recognize the historical case and retrieve its outcome. Recent work in forecasting evaluation has begun to quantify that confound directly (Section 8).
The paper makes four contributions, each labeled by type.
PROPOSED. Failure-First Forecasting, a seven-stage protocol that turns prospective hindsight into a frozen, scorable forecast with attached interventions.
PROPOSED. Iterative Premortem Hardening, the forecast–intervene–reforecast loop, with an explicit stopping condition and an explicit account of the failure modes that interventions themselves introduce.
PROPOSED. A backtesting methodology and metric set for evaluating premortem forecasts against realized outcomes at the level of causal mechanism, precursor lead time, and intervention alignment, with temporal isolation as a mandatory control.
HYPOTHESIS. That residual epistemic exposure — uncertainty remaining after the accessible failure space has been explored — should govern the balance between mechanism-specific remediation and mechanism-agnostic structural resilience, and that proposed indicators of that exposure predict subsequent discovery of unanticipated consequential failures.
The paper does not claim that AI premortems outperform human ones, that the residual can be measured, or that the protocol has been validated. It claims that these questions can now be asked precisely, and it specifies how.
Background and Related Work
2.1 Premortem and prospective hindsight
Mitchell, Russo, and Pennington (1989) established the core effect: explaining an event as though it had already occurred, rather than as a possibility, produces more numerous and more concrete causal accounts. Klein (2007) operationalized this into the premortem as a project-planning ritual, arguing that it legitimizes dissent and surfaces concerns that hierarchy would otherwise suppress. Peabody (2017), working with Veinott, tested the technique quantitatively across three experiments. Experiment 1 (field, Army cadet teams) found fewer fouls and less fixation without added time. Experiment 2 (laboratory, engineering plans) found no statistically significant difference in reasons or solutions generated between the premortem and a worst-case-scenario method, though category distributions differed. Experiment 3 (complex, unfamiliar plans; individuals and groups) found significantly more reasons generated with the premortem. The fixation and overconfidence effects are treated here as ESTABLISHED; the claim that premortems generate more reasons than worst-case prompting is SUPPORTED under some conditions and not others, and is not stated unconditionally in this paper.
Two limits of the technique matter for what follows. First, a premortem produces reasons, not forecasts; the literature measures the number and specificity of reasons, not whether they came true. Second, the exercise has documented side effects. Gandhi, Duke, and Schweitzer report, across four studies (three with N between 615 and 937 and a fourth pilot with N = 85; total N = 2,345), that while premortems reduce overconfidence, they magnify self-serving attribution: participants disproportionately blamed prospective failures on factors outside their control — weather, luck — rather than on their own effort or skill, and, in the authors' words, "no differences emerged in taking action to improve results." The work is a conference poster and should be weighted accordingly. Its relevance here is precise: a premortem whose pathways cluster on uncontrollable causes produces analysis that does not change behavior. That is not because an uncontrollable cause implies nothing can be done — weather is uncontrollable, but exposure to weather can be reduced, buffered, or insured. It is because attribution to the cause diverts attention from the links in the pathway where an action at t₀ could intervene. This motivates the interventionability dimension introduced in Section 3, which is a property of the pathway, not of its initiating cause.
2.2 Failure mode and effects analysis
FMEA, formalized in reliability engineering through MIL-STD-1629A and its industrial descendants, characterizes each failure mode by severity, occurrence, and detection, and combines these into a risk priority number. Detectability is therefore a long-established axis of failure analysis, not an innovation of this paper. What this paper does with detectability is different in one respect: rather than folding it into a single priority score, it uses it to partition failure modes into those suitable for monitoring and mechanism-specific remediation and those that, being undetectable in prospect, require structural resilience instead. That partition is PROPOSED; the axis is not, and Section 6.4 specifies that low detectability is first a prompt to ask whether observability can be engineered, and only then a reason to reach for structural resilience. FMEA's documented weaknesses also inform the measurement discipline of Section 3: Bowles (2003) showed that the risk priority number — the product of ordinal severity, occurrence, and detection ratings — is not a valid ordering of risk, since the same RPN arises from very different rating combinations and the scales lack interval properties; Liu, Liu, and Liu (2013) survey the extensive literature of proposed replacements. This paper does not multiply ordinal ratings.
2.3 Resilience, real options, and robust decision-making
Several literatures address how to act well when the failure mechanism is unknown. Real options theory (Dixit and Pindyck, 1994; Trigeorgis, 1996) values the ability to defer, stage, or abandon a commitment as an asset in its own right. Resilience engineering (Hollnagel, Woods, and Leveson, 2006) studies how systems absorb disturbances they were not designed for. Robust decision-making (Lempert, Popper, and Bankes, 2003) seeks strategies that perform acceptably across many futures rather than optimally in one. Staged commitment, blast-radius containment, and graceful degradation are standard practice in software and infrastructure engineering. This paper draws its structural-resilience taxonomy from these antecedents and claims no invention of its members. Its contribution is the proposed coupling of that taxonomy to an indicator of residual epistemic exposure — a rule for when to reach for these instruments rather than a catalogue of them.
2.4 AI-assisted planning and premortems
Veinott and Lehman (2024) compared failure reasons generated by student participants and by GPT-4 for the same planning scenario, examining similarity and difference as a basis for human–AI teaming in decision-making. Related work has assessed language models against human practitioners on project risk identification. This body of work establishes that models can produce premortem-style content and that its character differs from human output; it does not yet establish whether that content forecasts realized failures. The present paper is positioned as the methodological successor to that question.
2.5 Language-model forecasting and calibration
Halawi et al. (2024, NeurIPS) built a retrieval-augmented forecasting system and reported accuracy approaching that of aggregate human forecasters on binary questions from prediction platforms, with attention to retrieval date cutoffs. Schoenegger et al. (2024, Science Advances) found that an ensemble of twelve language models matched the accuracy of a human forecasting tournament crowd. Both lines of work also document limits: uneven calibration, sensitivity to question framing, and dependence on the information available to the model at forecast time. Critics have noted that some claims of superhuman AI forecasting do not survive scrutiny of their leakage controls. The relevant lesson for this paper is twofold: ensembles of multiple runs, particularly across model families, improve forecast quality, and evaluation is only as good as its information boundary.
2.6 Temporal leakage in retrospective evaluation
Any evaluation that asks a model to forecast a historical event faces the possibility that the model's training data contains the outcome. The contamination literature in natural-language evaluation documents how benchmark answers leak into training corpora and inflate measured performance. For forecasting the problem is sharper: the model need not have memorized the answer verbatim; it need only recognize the case well enough to retrieve the direction of the outcome. Section 8 treats this as the primary validity threat to any historical backtest of AI premortems.
Failure-First Forecasting
3.1 The formal object
Let π denote a plan, defined by its intended actions, its desired outcome, an explicit success criterion, a failure criterion, a time horizon [t₀, t₁], and a stated set of assumptions A = {a₁, …, an} that must hold for the plan to succeed.
A failure pathway h is a causal account of how π fails. It is not a risk label. It is a structured object with the following components:
- an initiating condition,
- the assumption in A that is violated,
- a causal sequence from initiation to consequence,
- intermediate effects observable along that sequence,
- the terminal failure,
- an estimated time to manifestation.
A premortem produces a set of pathways H = {h₁, …, hm}. The protocol's outputs are this set, a characterization of each member, a set of interventions, and a revised plan, all frozen at t₀ so that they can later be scored.
3.2 Measurement discipline
Every quantity attached to a pathway is one of two kinds, and the paper does not mix them.
Cardinal quantities are used only where the research design can produce them. A pathway probability is defined as
where I₀ is the information available at t₀ and π is the plan as frozen. The event forecast is the terminal failure occurring by way of that causal pathway — not the initiating condition alone, and not partial activation of the chain that is interrupted before terminal failure. A model may emit such a number, but a model-generated probability is not a calibrated probability until calibration has been measured against outcomes (Section 9). The two are labeled differently throughout. Consequence Li is defined in a common unit — money, time, a stated utility scale — when expected-loss arithmetic is performed, and only then.
Ordinal quantities — high, medium, low — are used where cardinal estimates are not supported. Ordinal labels are never multiplied, summed, or differentiated. Where a decision must be made on ordinal inputs, the protocol uses explicit ordinal rules: dominance (a pathway rated high on both consequence and probability is prioritized over one rated high on only one), lexicographic ordering with a stated priority of dimensions, or threshold rules (any pathway rated high-consequence and low-detectability is first routed to a detectability-engineering assessment, and to structural resilience if observability cannot be created economically — Section 6.4). The choice of rule is recorded as part of the frozen protocol.
The FMEA practice of multiplying severity, occurrence, and detection ratings into a single number is a known violation of this discipline, and the paper does not reproduce it.
3.3 The seven stages
F1 — Freeze. Define the plan, the desired outcome, the success and failure criteria, the time horizon, the explicit assumptions, and the constraints and resources. Establish the information cutoff t₀ and freeze the permitted source material. Nothing generated after this point may draw on information dated after t₀. Where a model is used, freeze its execution metadata as well: provider, exact model and version identifier where available, execution date and time, system instructions, prompt template, sampling settings, run count, retrieval state, evidence corpus and version, and tool availability. This stage exists to make the subsequent forecast legitimately ex ante and reproducible; without it, nothing downstream can be scored.
F2 — Baseline forecast. Before prospective hindsight is induced, record an outcome distribution for the plan: the probability of success, partial success, and failure as defined in F1, with any confidence or calibration metadata the design supports. This baseline is the reference against which the effect of the premortem itself is measured. If the premortem changes the forecast, the change is a datum; if the plan is later modified, the baseline is what the reforecast is compared to.
F3 — Prospective-hindsight induction. Assume the plan has failed at t₁. Do not ask what could go wrong. Ask what did go wrong, and reconstruct the mechanism. Each generated pathway must carry the full structure of Section 3.1: initiating condition, violated assumption, causal sequence, intermediate effects, terminal failure, timing. A failure reason without a causal sequence is not admitted as a pathway.
F4 — Ensemble failure generation. Run F3 multiple times in isolated contexts — no run is exposed to another run's output — and where justified from differing analytical perspectives: operational, financial, adversarial, regulatory, technical. Distinguish a within-model ensemble (repeated isolated runs of one model) from a cross-model ensemble (different model families or providers). Repeated runs of the same model are not statistically independent samples of the failure space; they share training data, architecture, and therefore blind spots, and their agreement is evidence about that model's accessible space rather than about the space itself. A cross-model ensemble reduces but does not remove this correlation. The protocol records which kind of ensemble was used and does not describe same-model runs as independent without that qualification. Cluster semantically equivalent pathways so that the same mechanism described twice is counted once. Preserve disagreement rather than resolving it. Record, for each run, which pathways were newly discovered and which had already appeared, so that the rate of novel discovery across runs is observable. Model agreement is not treated as truth; consensus and dissent are both evidence to be evaluated at F5, and neither carries authority.
F5 — Failure characterization. For each clustered pathway, record:
- probability, labeled as model-generated or empirically calibrated;
- consequence, in a stated unit or on a stated ordinal scale;
- detectability — whether the pathway's intermediate effects would be observable before the terminal failure;
- interventionability — whether an action available at t₀ can alter, interrupt, buffer, contain, or otherwise change at least one consequential causal link before terminal failure; cause controllability (whether the initiating condition itself is within any party's control) may be recorded as descriptive metadata but is not the operative variable;
- expected time to manifestation;
- observable precursors, stated concretely enough that their occurrence could later be adjudicated;
- at least one candidate intervention, or an explicit statement that none is available.
Interventionability is the operative variable because of the attribution finding in Section 2.1, and because of a distinction that finding makes easy to miss. The controllability of a cause and the interventionability of a pathway are different properties. A hurricane is not controllable; a plan's exposure to a hurricane — its timing, its geographic concentration, its insurance, its fallback — is interventionable. A pathway that begins with an uncontrollable cause may still contain several links where an action at t₀ changes the outcome, and a premortem that stops at the cause has stopped one step too early. Recording interventionability per pathway makes visible whether the analysis has found leverage, and F6 acts only on pathways where it has.
F6 — Remediation. For each pathway rated interventionable, specify the intervention as a modification to the plan that alters, interrupts, buffers, or contains a specific causal link. Classify each intervention as mechanism-specific (it addresses one pathway) or mechanism-agnostic (it improves structural resilience against a class of pathways, including unenumerated ones — Section 6). For each intervention record its cost, implementation time, expected effect on the targeted link, the new dependencies it introduces, and its plausible second-order effects. The last two fields are mandatory because interventions are themselves plan changes and therefore themselves sources of failure.
F7 — Reforecast. Apply F3–F5 to the modified plan. The output is a comparison: which pathways were removed or weakened, which are unchanged, and — critically — which are new, introduced by the interventions of F6. Record the revised outcome forecast beside the F2 baseline. Record the change in the residual-exposure indicators of Section 5. Recommend either another iteration or a decision-gate evaluation.
3.4 Tripwires
A precursor recorded at F5 becomes a tripwire when it is committed to as a monitoring condition: a concretely stated observation that, if it occurs after execution begins, indicates that a specific pathway is in progress. Tripwires are the mechanism by which a premortem continues to act after the decision. They are also scorable: after t₁, one can ask whether the tripwire fired, whether it fired before the terminal failure, and by how much — the warning lead time of Section 9.
Iterative Premortem Hardening
4.1 The loop
A single premortem produces a set of pathways and a modified plan. There is no reason to assume the modified plan is safe; its interventions were designed against the pathways that were found, and they have introduced their own. Iterative Premortem Hardening is the discipline of attacking the modified plan again:
Each cycle produces a new pathway set Hk. The interesting quantities are the differences between successive sets: pathways eliminated, pathways persisting, and pathways newly introduced by the previous cycle's interventions. Pathway count alone cannot establish hardening. A cycle should be evaluated by the consequence, probability, detectability, and interventionability of the pathways removed, retained, and introduced, not by their raw number — a cycle that removes one catastrophic pathway and introduces three minor, detectable ones may have hardened the plan considerably, and the reverse is equally possible.
4.2 Interventions introduce failure
This is the loop's signature observation and the reason it is iterative rather than one-shot. An intervention is a change to the plan, and every change carries assumptions. Adding a fallback system assumes the fallback works and is maintained; staging a rollout assumes the stages are independent; adding a monitoring tripwire assumes someone will act when it fires. The second premortem exists to test those assumptions. A remediation that is never itself subjected to prospective hindsight is an untested assumption wearing the costume of a solution.
4.3 Stopping condition
The loop terminates under either of two conditions:
- the expected marginal value of another failure-discovery and remediation cycle falls below the cost of conducting it, or
- the exposure that remains is better addressed through structural resilience than through further mechanism-specific prediction — the condition developed in Sections 5 and 6.
Low novelty alone is not sufficient evidence of coverage. A cycle that finds few new pathways may have exhausted the failure space, or may have exhausted only the space accessible to a generator whose blind spots are shared across runs; convergence within a shared generator speaks only to that generator's accessible space. Iteration may therefore stop when one or more of the following applies: validated marginal consequential recall has flattened; additional cycles across multiple generator families yield negligible new high-consequence pathways; the predeclared analysis budget is exhausted; or the remaining exposure is better managed through structural resilience, as indicated by the residual-exposure indicators of Section 5.
4.4 What a lower forecast does not prove
A reforecast that reports a lower model-estimated probability of failure for Plan₁ than for Plan₀ is a claim by the same generator that produced both. It does not establish that Plan₁ is objectively safer. The value of the reforecast — like the value of the original forecast — is an empirical question that the backtesting methodology of Section 7 is designed to answer. Until calibration has been measured, a reduced model-estimated failure probability is a proposed improvement, not a demonstrated one.
Residual Epistemic Exposure
5.1 The limit of enumeration
No premortem enumerates every way a plan can fail. The pathways generated are a sample of the failure space, bounded by the experience, imagination, and information of whoever — or whatever — generated them. The uncertainty that remains outside the enumerated set is not zero, and a decision procedure that acts as though it were is acting on a false premise.
This paper names that remainder residual epistemic exposure, Uπ: failure-relevant uncertainty remaining after the accessible failure space has been explored. Residual epistemic exposure is not directly observed; it is a latent construct inferred from observable indicators and validated only insofar as those indicators predict subsequently discovered consequential failures. It is deliberately not called expected residual loss, because no validated estimator of expected loss from unenumerated failures exists, and calling it one would manufacture a precision the concept cannot support.
5.2 Three classes of unenumerated failure
The remainder is not homogeneous. Distinguish:
- HE — failure mechanisms explicitly enumerated by the premortem.
- HC — failure classes that are known to exist, whose specific manifestation in this plan is not enumerated. An undocumented dependency, an untested edge case, an omitted workflow: the analyst knows that dependencies can be undocumented and workflows can be omitted. What is unknown is which one.
- HU — mechanisms genuinely outside the current conceptual model, which the analyst would not recognize as a failure class at all.
This distinction matters because the three classes respond differently to effort. HE is addressed by mechanism-specific remediation. HC is reduced by more enumeration effort, more perspectives, and better information — a more thorough dependency audit converts members of HC into HE. HU is not reduced by enumeration, because the limitation is conceptual rather than diligent.
The paper does not classify an omitted workflow or an undocumented dependency as a true unknown-unknown. Those belong to HC. Genuine HU membership is rarer than the term's popularity suggests, and most of what a post-mortem calls "unforeseeable" was, in the taxonomy above, an HC failure that was not enumerated because the enumeration stopped.
5.3 Indicators
Uπ cannot be observed directly. The following are PROPOSED as indicators, not measures:
- Cross-run divergence. Disagreement among isolated F4 runs about which pathways exist, reported separately for within-model and cross-model ensembles. High divergence suggests the accessible space is under-sampled; different runs are still discovering different regions.
- Rate of novel discovery. The number and consequence-weight of pathways first appearing in run k. A declining rate suggests approaching exhaustion of the space accessible to the generators used — which is not the same as the failure space; a persistently high rate suggests the space is large relative to the sampling.
- Plan novelty. The absence of precedent weakens enumeration-from-experience.
- Component coupling. Tightly coupled systems produce interaction failures that are properties of the whole rather than of any part.
- Environmental adversariality. An adversary generates failure modes selected for not appearing in the defender's enumeration.
- Dependency opacity. The extent to which the plan relies on components whose internal behavior is not visible to the planner.
- Data scarcity and model disagreement. Where the forecast itself rests on thin evidence or on generators that disagree about outcome probabilities.
Each of these is a hypothesis about what correlates with subsequent discovery of unanticipated consequential failure. None has been validated. Section 9 specifies the test: whether these indicators, recorded at t₀, predict the later appearance of consequential failures that were absent from HE.
5.4 A caution about low divergence
Low cross-run divergence does not establish that the failure space is covered. Runs that share a generator, a prompt, or a corpus share blind spots, and their agreement may reflect a common limitation rather than completeness. Convergence is evidence about the accessible space for this setup; it is silent about HU. The paper does not claim otherwise.
Structural Resilience Under Unenumerated Failure
6.1 Two kinds of mitigation
Interventions fall into two classes. Mechanism-specific interventions break a link in an enumerated pathway; they work exactly to the extent that the pathway was correctly identified. Mechanism-agnostic interventions change the plan's structure so that damage is limited or adaptation remains possible across a range of mechanisms; they can work without the specific mechanism having been predicted. Mechanism-agnostic does not mean universally effective across all failure mechanisms; it means less dependent on precise ex-ante identification of one specific failure pathway. A blast-radius limit bounds many failure modes and none perfectly; a rollback capacity is worthless against a failure that corrupts the rollback target.
Structural resilience is the second class. Its members, drawn from the antecedent literatures of Section 2.3, include:
- Staged commitment — execute in increments, each observable before the next commits.
- Reversion capacity — the ability to undo, not merely to stop.
- Optionality — preservation of the ability to change course; deferred commitment.
- Blast-radius containment — a bound on the damage any single failure can cause.
- Redundancy — duplicate capacity that tolerates the loss of one component.
- Diversification — non-correlated exposure, so that one mechanism does not fail everything.
- Modularity — isolation of components so that failure does not propagate.
- Slack and reserves — capacity, time, or liquidity held back to absorb the unexpected.
- Graceful degradation — reduced function under failure rather than total loss.
- Sensing and instrumentation — observability that shortens time-to-detection.
- Fail-safe defaults — the state the system enters when control is lost.
- Exit rights — contractual or structural ability to leave.
- Adaptive policies — decision rules that update on observation rather than fixed commitments.
Reversibility is one instrument in this class. It is not the class, and it is not the only defense against unenumerated failure; containment, redundancy, and sensing each operate without any reversal occurring.
6.2 The central proposition
HYPOTHESIS. As residual epistemic exposure increases, rational mitigation effort should shift from mechanism-specific interventions toward mechanism-agnostic structural resilience.
The reasoning is direct. Mechanism-specific remediation has value proportional to the probability that the enumerated pathway is the one that will operate. As Uπ grows — as the indicators of Section 5.3 suggest that consequential failures lie outside HE — that probability falls, and effort spent patching enumerated pathways buys less. Structural resilience is less dependent on precise ex-ante identification of the failure mechanism, so its value is discounted less by enumeration incompleteness. The crossover is not derived here as a theorem, because Uπ has no validated estimator; it is stated as a hypothesis with a proposed test.
6.3 A candidate formalization
The following expression is a structural illustration of the proposition, not a model to be estimated. Let R denote a level of structural-resilience investment and xi ∈ {0, 1} the decision to apply mechanism-specific intervention i with effectiveness ρi ∈ [0, 1] on enumerated pathway i. Expected total decision cost under the plan may be decomposed as:
with 0 ≤ α(R) ≤ 1 and dα/dR < 0, where Ũπ stands for the expected loss attributable to unenumerated failure and C(R) is the cost of the resilience investment itself — which is why the whole is labeled total decision cost rather than failure loss. The expression is illustrative, non-estimative, and non-computational. It makes explicit that mechanism-specific interventions act on enumerated pathways, while structural-resilience investment is intended to attenuate exposure not captured by those pathways. No crossover point is derived, because neither residual exposure nor the relevant response functions are presently estimable. In practice the paper relies on the ordinal rule: high residual-exposure indicators route effort toward resilience.
6.4 Detectability routing
Detectability, imported from FMEA, does one further job, and the routing it drives has two steps rather than one. An enumerated pathway with low prospective detectability — no intermediate effect currently observable before the terminal failure — cannot be monitored as the plan stands, and a tripwire cannot be set for it. The first question is therefore whether detectability can be engineered upward: instrumentation, reconciliation, audit, a parallel run, a canary. Sensing is itself a resilience instrument (Section 6.1), and it is frequently the cheapest one, because it converts a pathway from unmonitorable to monitorable without requiring that its probability or consequence be changed at all. Only if useful detection cannot be created economically, or cannot be created early enough to permit a response, does the pathway route to the second step: reduce its probability or consequence where possible, and structurally contain it where not. Low detectability is a prompt to build observability first and to fall back on containment second. That two-step routing is the PROPOSED use of an ESTABLISHED axis.
Backtesting Methodology
7.1 Three questions, not one
The empirical question divides into three, and the division matters because different evidence answers each.
- RQ1 — Predictive validity. Does the protocol identify causal failure mechanisms that subsequently materialize?
- RQ2 — Diagnostic validity. Do the interventions proposed ex ante target causal links that subsequently prove to matter?
- RQ3 — Intervention efficacy. Does applying the proposed intervention improve realized outcomes?
Ordinary historical backtesting can answer RQ1 and can partially inform RQ2, because the historical record contains what happened and, sometimes, why. It cannot establish RQ3, because the revised plan was never executed: the counterfactual outcome under intervention does not exist in the record. Any claim that the protocol improves plans — as opposed to predicts failures — requires simulation in which the revised plan can be run, controlled intervention in which some plans are revised and others are not, or prospective causal evidence. This paper does not let a favorable RQ1 result stand in for RQ3.
7.1a Which evidence answers which question
| Evidence source | Answers | Does not answer |
|---|---|---|
| Historical backtest (frozen forecast vs. recorded outcome) | RQ1 predictive validity | RQ3 |
| Historical causal record (adjudicated mechanism vs. proposed intervention) | RQ2 diagnostic / remediation alignment, partially | RQ3 |
| Simulator or randomized execution of revised vs. unrevised plans | RQ3 intervention efficacy | External validity beyond the simulator |
| Prospective frozen forecast (outcome postdates execution) | RQ1 without contamination; RQ2 | RQ3 unless paired with controlled intervention |
No single design answers all three.
7.2 Comparison conditions
- A. Baseline plan assessment without premortem.
- B. Conventional prospective risk analysis (a risk register or FMEA-style enumeration without prospective-hindsight induction).
- C. Human premortem.
- D. Single-run AI premortem.
- E. Ensemble, multi-run AI premortem (F4).
- F. Full protocol: ensemble AI premortem with remediation and adversarial reforecast (F1–F7).
- G (optional). Human–AI collaborative premortem.
Comparing D against E isolates the value of ensemble generation. Comparing E against F isolates the value of the intervention–reforecast loop. Comparing C against D and E addresses whether model breadth exceeds human breadth on this task — a question the paper does not prejudge.
Comparisons across conditions are interpretable only when enumeration budgets are normalized. A condition that emits more pathways will achieve higher recall by volume alone; an unrestricted recall comparison between a human premortem that produced eight pathways and an ensemble that produced sixty measures the count, not the quality. The protocol therefore fixes a budget across conditions using at least one of: an equal number K of admitted pathways (the top K by the condition's own prioritization), equal analysis time, equal inference or token budget, or equal monetary cost. Results are reported at fixed K and, where possible, as curves across K, so that the marginal consequential recall of each additional pathway is visible. Unrestricted comparisons across conditions with different enumeration volumes are not reported as evidence.
7.3 Historical backtest protocol
For each historical decision in the evaluation set:
- Reconstruct the decision at t₀: the plan, its stated goals, its assumptions, and the information legitimately available.
- Provide the analyst — human or model — only information dated at or before t₀, under the isolation controls of Section 8.
- Freeze the baseline forecast, the predicted pathways, the proposed tripwires, and the proposed interventions.
- Reveal observed events from t₀ to t₁.
- Construct a ground truth: the realized outcome and, where the record supports it, the realized failure mechanism(s) and their consequence.
- Score the frozen predictions against the ground truth using the metrics of Section 9.
Step 5 is the hard step. A realized failure often has contested causes, and the historical record may not adjudicate among them. Ground-truth construction must be done by adjudicators blinded to the predictions and to experimental condition, with inter-adjudicator agreement reported and genuine causal ambiguity preserved rather than forced to a single answer. Step 6 uses the predeclared causal-similarity rubric of Section 9, not a holistic judgment. For every forecast, historical or prospective, the execution metadata frozen at F1 is recorded alongside the predictions, so that the forecast can be reproduced and its provenance audited.
7.4 Prospective protocol
Where feasible, the strongest design is prospective: preregister the frozen forecast, pathways, tripwires, and interventions for a live decision whose outcome has not yet occurred, and score after t₁. This eliminates foreknowledge contamination, at the cost of time. The condition that makes a study prospective is that the outcome postdates the actual, timestamped execution of the frozen forecast — not merely a model's nominal knowledge cutoff, which is neither reliably known nor reliably respected (Section 8). A research program should include both: historical backtests for volume, prospective cases for validity.
Temporal Isolation
8.1 The threat
A historical backtest of a language model is invalid if the model can retrieve, recall, infer, or recognize events that occurred after the intended cutoff. The threat is not limited to verbatim memorization. A model that recognizes a company, a product, a policy, or a distinctive situation may retrieve the direction of its outcome without recalling any specific document. A premortem that "predicts" a well-documented historical failure may be reporting, not forecasting.
An instruction to the model to pretend the current date is earlier is not an isolation procedure. It changes what the model is asked to say, not what it knows. This is now directly measured. Liu et al. (2026), introducing the ExAnte benchmark at EACL, find that language models frequently violate temporal constraints across tasks and that, even under explicit temporal cutoffs, they often rely on internalized post-cutoff knowledge. The failure is not confined to the model's parameters. El Lahib et al. (2026), auditing search-engine date filters for retrospective forecasting at ACL, found that at least one retrieved page contained major post-cutoff leakage for 71% of questions on Google and 81% on DuckDuckGo, that the answer was directly revealed for 41% and 55% respectively, and that forecasting models given these leaky documents showed artificially inflated accuracy (Brier 0.10 against 0.24 with properly filtered documents). The mechanisms — updated articles, related-content modules, unreliable metadata, absence-based signals — are ordinary features of the web, not edge cases. Ordinary date-filtered retrieval therefore cannot support credible retrospective evaluation, and a frozen, provenance-controlled corpus is not a refinement of date filtering but a replacement for it.
8.2 Controls
The following are treated as mandatory for any historical evaluation reported under this protocol:
- Frozen pre-t₀ evidence corpus. The model's inputs are a fixed document set with dates verified at or before t₀.
- No unrestricted retrieval. Web access, if any, is disabled or restricted to the frozen corpus during forecast generation.
- Source provenance. Every input document is logged with its date.
- Entity anonymization. Distinctive names, product identifiers, and unique historical markers are removed or replaced where feasible.
- Date normalization. Dates are shifted or abstracted where doing so does not destroy the decision's structure.
- Case obscurity. Preference for historical cases that are unlikely to be well represented in training data.
- Recognition probes. Before the forecast, test whether the model can identify the anonymized case; discard cases it recognizes.
- Prospective cases. For the strongest validation, use decisions whose outcomes postdate the model's training.
8.3 Terminology
The paper uses temporal leakage, foreknowledge contamination, and post-cutoff contamination for this threat. Results reported without the controls above should be labeled as contamination-uncontrolled and treated as upper bounds on forecast validity, not as estimates of it.
Evaluation Metrics
Each metric below is defined so that a frozen premortem can be scored against a realized outcome. Together they distinguish the properties that a plausible narrative and a validated forecast do not share.
- Failure-mode recall. The fraction of realized consequential failure mechanisms that appeared in the frozen pathway set HE.
- Failure-mode recall@K. The same, computed over only the top K pathways admitted under the condition's own prioritization (recall@K, self-ranked), with K fixed across conditions. This is the primary recall metric for cross-condition comparison; unrestricted recall is reported only within a condition. Because self-ranked recall@K conflates two abilities — generation quality, whether the system produces the consequential pathway at all, and ranking quality, whether it places that pathway within the top K — a secondary comparison re-ranks each condition's full pathway set by a common blinded adjudication procedure before truncating at K. A condition that generates the right pathway but ranks it at K+5 has a ranking problem, not a generation problem, and the two comparisons together show which.
- Failure-mode precision, and precision@K. The fraction of predicted pathways (or of the top K) that materialized within the evaluation horizon. Low precision is not by itself damning — a premortem is supposed to over-generate — but it must be reported, because a generator that predicts everything achieves recall trivially.
- Critical-failure recall, and critical-failure recall@K. Recall weighted by realized consequence, so that missing the failure that mattered is penalized more than missing a minor one; the @K form is the cross-condition version.
- Marginal consequential recall. The consequence-weighted recall added by the (K+1)th pathway, reported as a curve across K. This shows where additional enumeration stops paying.
- Probabilistic calibration. Where cardinal pathway probabilities were recorded: Brier score, log score where appropriate, reliability curves, and expected calibration error with binning disclosed. This is the metric that separates model-generated probability from calibrated probability.
- Causal similarity. The degree to which a predicted pathway matches the realized mechanism, scored against a predeclared rubric rather than holistically. Six components are scored separately: initiating-condition match; violated-assumption match; mechanism-class match; causal-link or sequence match; terminal-failure match; and precursor match. Each is scored by at least two adjudicators blinded to experimental condition and to the source of the prediction, with inter-rater agreement reported per component. Automated semantic or model-based judging may be used as a secondary analysis and for triage, never as the sole ground-truth mechanism. This is the metric that distinguishes predicting the outcome label from predicting how it came about, and the per-component scores show which part of the mechanism the forecast got right.
- Warning lead time. For each realized failure with a corresponding predicted tripwire: the interval between the tripwire's first observable occurrence and the terminal failure. Zero or negative lead time means the tripwire was useless even though the pathway was correct.
- Interventionability. Whether an action available at t₀ could reasonably have altered, interrupted, buffered, or contained a causal link that was later observed to operate. This scores the interventionability judgments of F5 against the record, and it is a property of the pathway, not of whether its initiating cause was controllable.
- Remediation alignment. Whether the intervention proposed ex ante targeted a mechanism that subsequently contributed materially to the observed failure. This is the metric for the F6–F7 loop: did the plan change for the right reason. It informs RQ2 only; it says nothing about RQ3, because it does not observe the outcome under the intervention.
- Enumeration convergence. The marginal number and consequence-weight of new pathways discovered across multiple isolated runs, as a function of run count, reported separately for within-model and cross-model ensembles.
- Residual-uncertainty validity. Whether the indicators of Section 5.3, recorded at t₀, predict the later discovery of consequential failure mechanisms absent from HE. This is the test of the paper's secondary hypothesis.
Worked Example
The following runs a recognizable decision through the complete protocol. It is an illustration of procedure, not evidence of efficacy; no probability stated here is calibrated, and every number is a placeholder for what the protocol would record.
F1 — Freeze. Plan₀: cut a live production system over to a rebuilt replacement in a single scheduled switch on a fixed date. Desired outcome: the replacement serves all traffic with no loss of function or data. Success: thirty days of operation with no severity-one incident and reconciled data. Failure: rollback, data loss, or a severity-one incident attributable to the cutover. Horizon: t₀ = decision date, t₁ = t₀ + 30 days. Explicit assumptions: the replacement handles production load; migrated data is faithful; all downstream consumers tolerate the new behavior; a failed cutover is detected quickly; the old system remains available for fallback. Information frozen at t₀.
F2 — Baseline forecast. Before prospective hindsight: model-generated P(success) recorded, with the label "uncalibrated." Whatever the number, it is the reference point.
F3 — Induction. It is t₁ and the cutover failed. Reconstruct why.
F4 — Ensemble. Multiple isolated runs from technical, operational, and financial stances. After clustering, the pathway set includes: (h₁) load failure — replacement degrades under peak traffic, initiating condition is untested peak load, violated assumption is load capacity, terminal failure is outage within hours; (h₂) migration corruption — a type mismatch silently alters a subset of records, violated assumption is data fidelity, intermediate effects are reconciliation discrepancies, terminal failure is discovered weeks later; (h₃) undocumented consumer — a downstream system nobody listed breaks on the new interface; (h₄) slow degradation — performance decays over days rather than failing at once; (h₅) calendar-triggered workflow — a quarter-end process never exercised in testing fails when the calendar reaches it; (h₆) fallback unavailable — the old system was partially decommissioned in preparation, and the assumed rollback is not possible. Runs diverged most on h₃ and h₅ — the mechanisms that depend on knowledge no single perspective holds — which is recorded as a residual-exposure signal.
F5 — Characterization (ordinal, with rules stated):
- h₁: probability medium, consequence high, detectability high (an outage announces itself), interventionability high, timing hours, precursor: latency rising under load test, intervention: load test and capacity.
- h₂: probability medium, consequence high, detectability low as the plan stands, interventionability high, timing weeks, precursor: reconciliation deltas — but only if reconciliation is run, which the plan does not do. Detectability can be engineered: a parallel run with reconciliation makes the corruption observable. Intervention: build that observability.
- h₃: probability unknown, consequence high, detectability medium, interventionability medium (one cannot test a consumer one does not know about, but one can inventory and canary), timing hours to days. Intervention: dependency audit, traffic-mirroring canary.
- h₄: probability low-medium, consequence medium, detectability low as the plan stands but readily engineered, interventionability medium. Intervention: performance instrumentation with thresholds.
- h₅: probability unknown, consequence high, detectability very low before the calendar date and not economically engineerable — one cannot instrument a workflow one has not identified. Cause controllability is low (the calendar is not in anyone's control); pathway interventionability is nonetheless present, because exposure to the untested workflow can be contained by preserving reversion until the calendar has exercised it. This is an HC member — the class is known, the instance is not.
- h₆: probability medium, consequence very high, detectability high (one can simply check), interventionability high. Intervention: verify and preserve the fallback.
Rule applied (Section 6.4): any high-consequence, low-detectability pathway is first assessed for engineered observability. h₂ and h₄ can be made observable and are routed to sensing. h₅ cannot, and is routed to containment and reversion.
F6 — Remediation. Mechanism-specific: load test (h₁); dependency audit and canary (h₃); performance thresholds (h₄); verify fallback (h₆). Mechanism-agnostic, structural: run the two systems in parallel and reconcile output before trusting the replacement — this is the engineered detectability for h₂ (sensing) and also bounds damage (containment); stage the migration by record type rather than all at once (staged commitment, blast-radius); preserve the old system intact through a full business cycle including quarter-end, so that h₅ is covered not by predicting the workflow but by keeping the ability to revert while the calendar exercises it (reversion capacity, optionality). Each intervention's new dependencies are recorded: parallel running assumes the reconciliation itself is correct; staging assumes record types are independent; preserving the old system assumes budget and attention to maintain it.
F7 — Reforecast. Plan₁: staged, parallel-run, fallback-preserved migration. Second premortem. Removed or weakened: h₁ (tested), h₂ (now detectable), h₆ (verified). Unchanged: h₄ partially. Newly introduced by the interventions: (h₇) reconciliation false-negative — the reconciliation logic misses a class of discrepancy, so the parallel run certifies corrupted data; (h₈) stage-boundary inconsistency — a record whose type straddles two stages is migrated twice or not at all; (h₉) dual-write drift — during parallel running, the two systems diverge under a race condition and the "correct" version is ambiguous. The loop has done its job: the interventions introduced three pathways that did not exist in Plan₀, and h₇ in particular attacks the very control that was supposed to fix h₂. Residual-exposure indicators: cross-run divergence fell for the technical pathways and remained high for the workflow-inventory pathways, consistent with HC exposure concentrated in what nobody holds a complete map of.
Gate. Plan₁ carries less enumerated exposure than Plan₀ as rated by the same generator, at the cost of delay, maintaining two systems, and the reconciliation dependency. The decision criterion of Section 6 weighs that cost against the expected value of the migration and the exposure removed. Under the illustrative ordinal decision rule used in this example — and noting that the reforecast's newly introduced pathways are each rated remediable and detectable — the gate would advance Plan₁ to execution with h₇–h₉ under active monitoring rather than trigger another full iteration. This demonstrates the gate mechanics; it does not establish that PROCEED is the objectively correct real-world decision, since the probabilities are uncalibrated, residual exposure is unmeasured, and intervention efficacy is unverified. What the metrics could later score is predictive and diagnostic validity (RQ1, RQ2): did any of h₁–h₉ materialize, did the tripwires lead the failure, and did the parallel run surface what it was built to surface. Whether Plan₁ produced a better outcome than Plan₀ would have is an intervention-efficacy question (RQ3) that the record of a single execution cannot answer.
What the unaided premortem would have produced is the list h₁–h₆ and, characteristically, a cutover on the scheduled date with a longer runbook. What the protocol produced is a different plan, a record of why it is different, and a set of frozen predictions that can be scored.
Limitations
- Hallucination and narrative plausibility bias. Models generate fluent, coherent failure stories regardless of whether the mechanism is real. Every metric in Section 9 exists because of this; none of them removes it at generation time.
- Probability miscalibration. Model-generated probabilities are not calibrated probabilities. The protocol labels them separately and does not treat a reforecast's lower number as evidence of safety.
- Temporal leakage. Historical backtests of language models are vulnerable to foreknowledge contamination. The controls of Section 8 reduce but do not eliminate the risk; only prospective evaluation eliminates it.
- Ground-truth ambiguity. Realized failures often have contested causes. Causal-similarity scoring depends on adjudication that may itself be uncertain, and inter-adjudicator disagreement must be reported.
- Unobservable counterfactuals. When an intervention is applied and the plan succeeds, one cannot observe whether the targeted pathway would have operated. Remediation alignment can be scored only when a pathway materialized despite or alongside intervention, which biases the sample.
- Correlated ensemble outputs. This is a first-order limitation, not a footnote. Repeated runs of one model are not independent samples; they share training data and architecture, and therefore blind spots. Within-model convergence measures the accessible space for that generator, not the failure space. Cross-model ensembles reduce the correlation without removing it, since model families share pretraining sources. The protocol reports which ensemble type was used and does not describe same-model runs as independent.
- Attribution bias. The premortem frame may magnify attribution to uncontrollable causes. The interventionability dimension makes visible whether the analysis found leverage in the pathway despite the cause; it does not prevent the bias.
- Analysis cost and paralysis. The protocol has cost. On low-stakes, well-precedented, easily reversible decisions, the correct amount of it is very little, and a discipline applied indiscriminately is soon ignored.
- Domain dependence. Both forecast validity and the value of resilience instruments vary by domain. Results in one domain do not transfer without evidence.
Research Agenda
The empirical successor to this paper should include:
- Retrospective datasets. A curated set of historical decisions with reconstructable t₀ information, documented outcomes, and adjudicated failure mechanisms, selected for low probability of appearing in model training data and processed under the isolation controls of Section 8.
- Simulated environments. Domains with a faithful simulator — where a plan can be executed many times against a model of the environment — so that forecast validity can be measured at volume without contamination. Financial strategy backtesting is one candidate; operations and logistics simulation are others.
- Prospective preregistered studies. Live decisions in cooperating organizations, with forecasts, pathways, tripwires, and interventions preregistered and scored after the horizon.
- Human–AI comparison. Conditions C through G of Section 7.2 on a common case set, scored on the same metrics at fixed enumeration budget.
- Ensemble structure. A three-arm comparison — single model, single run; same model, multiple isolated runs; cross-model ensemble — on the same cases, measuring recall@K, enumeration convergence, and the extent to which within-model agreement overstates coverage relative to cross-model agreement.
- Intervention efficacy (RQ3). Simulated or controlled execution of revised versus unrevised plans, since no backtest can supply it.
- Calibration studies. Whether model-generated pathway probabilities can be calibrated post hoc, and whether calibration transfers across domains.
- Residual-exposure validation. Whether the indicators of Section 5.3 predict subsequent discovery of unanticipated consequential failure — the test on which the paper's secondary hypothesis stands or falls.
- Intervention-induced failure. The rate at which F6 interventions introduce new consequential pathways, and whether F7 catches them.
Conclusion
The premortem is a good idea with an unmeasured output. It produces reasons a plan might fail, and the literature that supports it measures how many reasons and how specific — not how many came true. Language models make reasons abundant. That is the opportunity, and it is also the hazard: abundance of plausible failure narrative is not foresight, and a decision process that cannot tell the difference has acquired eloquence, not information.
Failure-First Forecasting is a proposal for telling the difference. Freeze the plan and the information. Record the forecast. Induce prospective hindsight and demand causal mechanism, not labels. Generate across multiple isolated runs, ideally across model families, and preserve the disagreement. Characterize each pathway by whether it can be detected and whether any link in it can be changed — not by whether its cause was in anyone's control. Intervene against specific causal links. Then attack the modified plan again, because interventions are assumptions too. Stop when another cycle is not worth its cost, or when what remains is better survived than predicted. And score all of it against what actually happened, with the model's foreknowledge held out.
The value of AI premortem analysis should be judged not by how convincing its failure stories sound, but by whether those stories identify consequential failure mechanisms before they occur and support interventions that improve the decision. This paper has not shown that they do. It has specified how to find out.
References
- Bowles, J. B. (2003). An assessment of RPN prioritization in a failure modes effects and criticality analysis. In Proceedings of the Annual Reliability and Maintainability Symposium (pp. 380–386). Tampa, FL: IEEE.
- Dixit, A. K., & Pindyck, R. S. (1994). Investment under uncertainty. Princeton, NJ: Princeton University Press.
- El Lahib, A., Xia, Y.-J., Li, Z., Wang, Y., & Pi, X. (2026). Temporal leakage in search-engine date-filtered web retrieval: A retrospective forecasting case study. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 787–795). Association for Computational Linguistics.
- Gandhi, L., Duke, A., & Schweitzer, M. (2021). Debiasing premortems: Self-serving attribution [Poster]. Annual Conference of the Society for Judgment and Decision Making.
- Halawi, D., Zhang, F., Yueh-Han, C., & Steinhardt, J. (2024). Approaching human-level forecasting with language models. In Advances in Neural Information Processing Systems 37. https://doi.org/10.52202/079017-1598
- Hollnagel, E., Woods, D. D., & Leveson, N. (Eds.). (2006). Resilience engineering: Concepts and precepts. Aldershot, UK: Ashgate.
- Klein, G. (2007). Performing a project premortem. Harvard Business Review, 85(9), 18–19.
- Lempert, R. J., Popper, S. W., & Bankes, S. C. (2003). Shaping the next one hundred years: New methods for quantitative, long-term policy analysis. Santa Monica, CA: RAND.
- Liu, H.-C., Liu, L., & Liu, N. (2013). Risk evaluation approaches in failure mode and effects analysis: A literature review. Expert Systems with Applications, 40(2), 828–838. https://doi.org/10.1016/j.eswa.2012.08.010
- Liu, Y., Wei, X., Shi, L., Li, X., Zhang, B., Dhillon, P., & Mei, Q. (2026). ExAnte: A benchmark for ex-ante inference in large language models. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1551–1571). Rabat, Morocco: Association for Computational Linguistics.
- Mitchell, D. J., Russo, J. E., & Pennington, N. (1989). Back to the future: Temporal perspective in the explanation of events. Journal of Behavioral Decision Making, 2(1), 25–38.
- Peabody, M. (2017). Improving planning: Quantitative evaluation of the premortem technique in field and laboratory settings [Master's thesis, Michigan Technological University].
- Schoenegger, P., Tuminauskaite, I., Park, P. S., Bastos, R. V. S., & Tetlock, P. E. (2024). Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy. Science Advances, 10(45), eadp1528. https://doi.org/10.1126/sciadv.adp1528
- Trigeorgis, L. (1996). Real options: Managerial flexibility and strategy in resource allocation. Cambridge, MA: MIT Press.
- U.S. Department of Defense. (1980). MIL-STD-1629A: Procedures for performing a failure mode, effects and criticality analysis. Washington, DC.
- Veinott, E., & Lehman, B. (2024). Adaptive planning: Comparing human and AI responses in premortem planning. In HCI International 2024 – Late Breaking Papers (pp. 256–268). Cham: Springer. https://doi.org/10.1007/978-3-031-76827-9_15