This report contains a mathematical analysis of human extinction risk, perturbing variables in an expected value calculation to explore the relative worth of risk mitigation in different scenarios. It is a generalization of existing, cited models of risk by Ord, Adamczewski, and Thorstad (OAT)[1][2][3] which relaxes several assumptions about risk to draw sharper conclusions about near-term payoffs of mitigating actions. The main result of this report is using its generalized model to analyze the value of mitigations in scenarios like Great Filtering and Time of Perils. In this review, I highlight several model-theoretic assumptions and discuss their consequences, which frame limitations of the model which I think are worth surfacing.
Assumptions in the report:
Finding more accurate and more useful models of risk from nuclear weapons or emerging technologies is worthwhile. It is also a necessary activity, because no one model is capable of capturing all aspects of risk, which is inherently uncertain and often unable to be precisely defined outside of limited contexts like reliability analysis in systems engineering (cf. “Reliability, Verification, and Analysis in Engineering Design” by Wasserman, 2003)[4] or actuarial analysis in mathematical finance (cf. “Nonlife Actuarial Models” by Tse, 2023)[5]. It is notable that these two example domains share the same methodology, i.e. statistics, as a risk theory, and that statistics is not likely to be useful in the same way, for predicting risk from nuclear conflict or emerging technologies, certainly because we have little observed data and less idea of what distributions are appropriate for the task.
So, I find this report’s exploration of technical approaches to be reasonable and well-motivated. However, I think it is important to analyze this model’s assumptions as carefully as its implications, to better understand what insight we are gaining from the formalism.
1. Like OAT, this model assumes that after each period \(t\) of time, a catastrophe occurs with probability \(p_t\), reducing the value of the world to 0. An important assumption here is that the \(p_t\) are all fixed and independent of each other: the probability of survival until time \(t\) is the product of probabilities of surviving all previous periods. In this model of risk, interventions act “product-wise,” with a reduction of risk of the form \((1-f)r_t\) for \(0<f<1\) and for times \(t\) within the mitigation timeline. The author assumes that \(0<1-f<1\), presumably because it is irrational for a society to decide on a course of mitigation that increases extinction risk. The author also offers that the window of risk reduction rarely exceeds 50 years, which I find to be unsupported; this assumption clearly demarcates an interval of time, after which a mitigation does not have an effect, positive or negative.
One important dynamic ruled out by this product-wise mitigation assumption is the possibility that actions at time \(t < s\) will unintentionally increase the extinction risk \(r_s\). While the framework formally permits risk-increasing actions (cf. p. 11), all of the worked analyses treat the risk path as fixed, deterministic, and exogenous, and model interventions as certain-sign fractional reductions over a known window. Essentially, by ruling out this dynamic, we have narrowed our scope of actions to those which we are completely certain will not increase risks at later periods of time. While this can still be a valuable exercise in resource allocation, I find that it misses a complex problem: how can we decide between courses of action, all of which are designed to mitigate short-term risk, when there is deep uncertainty about externalities? While current methods for answering this problem, like robust decisionmaking and decisionmaking under deep uncertainty, are far from predictive, they are intended precisely to address this methodological gap, ideally in a practical way.
The model therefore does not address the decision problem practitioners actually face, which is choosing among actions whose long-run effects on risk are uncertain in sign and magnitude, including well-intentioned mitigations that may raise later risk. Extending the framework to uncertain or endogenous risk paths would substantially increase its policy relevance.
2. Another assumption, which is explicitly outlined, is that the value of survival is monotone increasing in each of the five analyzed scenarios, although the general framework admits arbitrary value sequences (cf. section 2.3, functions 10-11). The greatest value of persistence across a few scenarios is derived from cubic growth, although this is not a special feature of cubic growth - if the report considered quartic growth, then presumably that would lead to even greater returns. In the end, I am not sure I understand the justification of cubic growth, especially since there is no ceiling from the formalism to how quickly value can grow.
3. A last assumption to highlight is about convergence. On page 21, the author states that a key issue is whether the expected value of the world converges, because this makes it meaningful to talk about long-term value of risk mitigation. I disagree with this statement because it is unlikely that human existence could persist infinitely long, making it unnecessary to consider infinite series convergence (cf. Sandberg and Manheim, 2021)[6]. We can argue to surface this disagreement from the model’s framework. Let’s assume, to eventually argue that finite timelines are all that is needed for the report’s analyses, that no matter what, there is always at least a baseline risk of extinction, \(0<c<1\), which could be incredibly small relative to the actual values \(r_t\) in the model. This is consistent with the author’s discussion of decaying risk, where \(c=r_\infty\). This baseline risk implies that the probability of surviving up to time \(T\) is bounded above by \((1-c)^T\), which decays exponentially in \(T\). This identifies an “approximate support” of finite time where mitigations are valuable - past a certain point, it is simply unlikely that humanity has survived. Technically, this decay in survival probability could be counteracted by an exponentially increasing value function - but this is too optimistic, and not a scenario that the report considers. So, using the model’s assumptions, we derive that we might be better off looking at long but finite lengths of time and not assume that the value and probability coefficients behave in any particular way to guarantee convergence.
As an aside about the baseline risk of extinction, we could also point out that this assumption excludes eventual-safety and decaying-to-zero scenarios. In those cases, assuming an infinite timeline, an almost-surely finite lifetime does not guarantee a finite expected-value series, and a finite numerical cutoff can conceal a material or divergent tail. If the report takes the position that an infinite timeline is necessary for its analyses, it should also analyze or at least discuss other scenarios that assume infinite timelines.
As a last note about convergence, I recognize that this point about convergence is meant to resolve a philosophical problem with OAT’s original model, that the expected value may not be well-defined with infinite time. If this is an important topic to the reader, I think the arguments on pages 21-22 could assume slightly less of the coefficients, using lim infs and lim sups for the ratio test.
Implications for model applications
From my view, these are strong assumptions on risk dynamics, behavior of mitigations, value of persistence, and timescales. While I fully believe in the utility of abstract modeling to inform decisions, I worry that in this case the assumptions are strong enough, in some sense, to simply be asserting the conclusions that are reached in scenarios. To make this precise: the report assumes that mitigations are deterministic and reduce the probability of extinction in a “product-wise” fashion, which has the effect of increasing any expected value of survival. Then, the analyses of the report assume that the value coefficient is monotone increasing, and so the effect on the expected value is predictable and, in some growth regimes, notably significant. So, while the report is careful to not draw explicit conclusions from these comparisons, noting a lack of confidence in any particular estimate of the value of risk mitigation efforts or any one scenario, a reasonable read of these calculations would be that when value grows sufficiently quickly, then in scenarios like Time of Perils or Two Great Filters, the value of persistence is considerable.
To contrast, let us weakly reject each of these assumptions in either Time of Perils or Two Great Filters and instead say:
1. Extinction risk at time \(t\) heavily depends on both the set of all actions prior to \(t\) and stochastic variables whose probability distributions we cannot easily characterize. Mitigations may reduce this risk in the short term but have strong and possibly unpredictable interactions with risks in the long term.
2. The value of surviving a period of time fluctuates and is not characterized by e.g. cubic growth.
3. There is a finite timescale that is relevant for the expected value calculation for survival.
Then, we may reach different conclusions from the report about the value of mitigation of risk and persistence. For example, if we are not confident that our mitigations will not make things worse, we may not be justified in attempting mitigation at all, because the uncertainty may make an expected value determination impossible. And, in this far dimmer view of the future, our decision-space would likely resemble hedging and optimizing over the time we have left.