For most people, AI-induced extinction is a topic reserved for late at night after a few too many drinks with friends. But when Jacob Coxon resigned from Anthropic earlier this month, the argument that AI might kill us all escaped its enclosure.
Pollster David Shor quickly picked up on the change. “AI salience has increased dramatically in the past week,” he wrote on X, “increasing as much in the last week as the previous year combined.” Now, “64% of voters think it’s either very or somewhat likely that AI could pose a threat to humanity’s survival.”
It’s important to be precise about our terms because extinction and catastrophe aren’t the same. I am worried that increasingly capable AI systems pose catastrophic risks that harm or even kill people. We need to be prepared. But none of this means that AI is going to kill everybody. Indeed, an unfortunate side effect of all this extinction talk is that the real risks of AI aren’t properly understood. And when you actually dive into the literature, it isn’t as bleak as you might think.
Some probability theory of p(doom)
Over the years, I’ve written about the idea of p(doom), a shorthand for AI extinction risk, in various places including here, here, and here. But my most comprehensive dive into the subject is in “Don’t Just Tell Me Your p(doom), Tell Me Your Conditionals” from last year. I wrote it after being frustrated while reading If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All by Eliezer Yudkowsky and Nate Soares. As I explained,
In practice, many p(doom) estimates tend to be unconditional, that is, doom happens regardless of what else might occur. To me, as someone who was trained in Bayesian econometrics, most p(doom) conversations feel like they are the first 5 minutes of an introductory course where you also learn conditionals, and start exploring if probability A increases, decreases, or perhaps follows a bathtub over time, depending on the probability of B.
Looking at it now, however, I would be much more direct about my point. In practice, many p(doom) estimates are presented as a single probability without any details on how it changes with technology or policy. In that piece, I went on to describe an interaction between Eliezer Yudkowsky and podcaster Dwarkesh Patel that, I think, “perfectly illustrates this fundamental misunderstanding of how probability should work in AI safety discussions.” Dwarkesh sees AI risk in a many things must go wrong model, while Yudkowsky frames it as a many things must go right model.
Just for my own sake, I thought it might be helpful to formally lay this out.
In Dwarkesh’s model, extinction is conjunctive. For p(doom) to happen, a series of things must all go wrong. As a result, uncertainty compounds. Let’s imagine that p(doom) requires ten distinct things to happen. If each event is 80 percent likely and conditional on the previous ones, then we can formulate it as,
Which gives us,
Basically, it takes on a Fermi-style formulation.
Yudkowsky’s model, however, assumes that survival is conjunctive in that a series of things must go right. Doom can arrive through any one of a potentially enormous number of routes. Formally, we can model it as humanity needing to get ten things right, with each succeeding at an 80 percent probability, such that,
The same compounding logic runs in the opposite direction. After all, it is the previous number subtracted from 1:
But you also can get a really high p(doom) if you think that there are many independent or partially independent routes, which would give us:
The probability of doom can therefore be very high even though Yudkowsky has little confidence about which particular route will occur. If there are twenty roughly 20-percent-probability ways for things to go irretrievably wrong, independence would give:
All of this is perfectly coherent. But in a policy setting, none of it is particularly helpful because extinction estimates hide the assumptions. We shouldn’t be asking, what’s your p(doom)? Instead, we should be asking, which policies will reduce p(doom)?
It is also worth pointing out that AI experts seem to be resistant to persuasion. In a 2023 paper by the Forecasting Research Institute, titled “Forecasting Existential Risks: Evidence from a Long-Run Forecasting Tournament,” AI experts were pitted against superforecasters in the Existential Risk Persuasion Tournament (XPT). Superforecasters are an important group because they are consistently found to be more accurate than the general public and experts in predicting the likelihoods of future events. Naturally then, the purpose of this game was “to produce high-quality forecasts of the risks facing humanity over the next century by incentivizing thoughtful forecasts, explanations, persuasion, and updating from 169 forecasters over a multi-stage tournament.”
The results were striking. The median AI expert put the probability of extinction by 2100 at around 3 percent, while the number among superforecasters was an order of magnitude lower, at 0.38 percent. Notably, there wasn’t convergence among the participants:
We document large-scale disagreement and minimal convergence of beliefs over the course of the XPT, with the largest disagreement about risks from artificial intelligence. The most pressing practical question for future work is: why were superforecasters so unmoved by experts’ much higher estimates of AI extinction risk, and why were experts so unmoved by the superforecasters’ lower estimates? The most puzzling scientific question is: why did rational forecasters, incentivized by the XPT to persuade each other, not converge after months of debate and the exchange of millions of words and thousands of forecasts?
Months of argument failed to move either side at all. And for what it’s worth, I’d put my lot in with the superforecasters, who are better at predicting events. To say it again, p(doom) tells us little about the probability of extinction, but a lot about the assumptions and the models that people bring to a conversation.
p(doom) v p(catastrophe)
It is fashionable to say we are all going to die from AI. But what would it take to actually accomplish that?
The best report, really the only report, I’ve seen that tries to wrestle with the pathways to extinction is Michael J. D. Vermeer, Emily Lathrop, and Alvin Moon’s “On the Extinction Risk from Artificial Intelligence.” Instead of beginning with abstract alignment arguments, these RAND researchers set out to falsify a null hypothesis, that “there is no describable scenario in which AI is conclusively an extinction threat to humanity.” They then used the scientific literature and RAND subject-matter experts to see whether that hypothesis can be rejected for nuclear weapons, engineered pathogens, and malicious geoengineering.
As the researchers made clear, all three pathways could produce global catastrophes where countless people lose their lives. Still, actual human extinction would be extraordinarily difficult. What would be needed is an actor deliberately seeking extinction, overcoming substantial physical constraints, and preventing humans from responding during what would often be a lengthy unfolding process.
An unfortunate side effect of all this extinction talk is that the policy discussions lose focus on the real problems, the biggest of which is biosecurity risk.
There is no doubt that LLMs can inform users about harmful materials. Back in 2023, MIT students reported that, within an hour conversation, ChatGPT had suggested “four potential pandemic pathogens, explained how they can be generated from synthetic DNA using reverse genetics, supplied the names of DNA synthesis companies unlikely to screen orders, identified detailed protocols and how to troubleshoot them, and recommended that anyone lacking the skills to perform reverse genetics engage a core facility or contract research organization.” But more than likely, the AI served up those responses because it is trained on the internet, which includes all of those details.
Still, knowledge of pathogens is just a small part of the overall production process. There is an important difference between informational uplift and capability uplift, between the information that LLMs tell an average user and whether AI meaningfully lowers the barrier to making something dangerous. A laptop can tell you how something works, but it cannot build you a containment lab, successfully order specialized precursors, and magically grant you the practical know-how and equipment needed to make it all work.
The first major look into this, a study out of RAND in 2024 led by Christopher Mouton, Caleb Lucas, and Ella Guest, found no statistically significant difference in the viability of biological attacks proposed by teams granted LLM and Internet access compared to those who just had access to the Internet. In “Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology,” Hong et al. (2026) ran a similar evaluation and “observed no significant difference in the primary endpoint of workflow completion (5.2% LLM vs. 6.6% Internet; P = 0.759), nor in the success rate of individual tasks.”
OpenAI’s evaluation of GPT-4 found that it “provides at most a mild uplift in biological threat creation accuracy.” AI didn’t make people better at thinking of new techniques, nor did it speed them up. Most importantly, participants didn’t find the tasks any less difficult. Peppin et al. (2025) came to a similar conclusion. This study conducted by 13 researchers at Cohere Labs, Stanford University, MIT, and other institutions collected all of the research on this problem and concluded that current large language models and AI-enabled biological tools “do not pose an immediate risk.” However, the researchers were clear that “more work is needed to develop rigorous approaches to understanding how future models could increase biorisks.”
To be fair, the situation could change as models become far more capable in the future. But right now, the risk from AI, even the biosecurity risk, is far lower than I think many imagine it to be.
The real risk has always been with a motivated agent. So it makes sense to restrict benchtop synthesizers and restrict people’s access to reagents and mail-order gene sequences. We should build more chokepoints in the physical supply chain but that doesn’t mean we should be hooking AI regulation here.
Concluding thoughts
We know the nature of the risk with AI. All of the hacking incidents from OpenAI and Anthropic have had the same contours. The company gave an advanced model an end objective and the model did everything it could to achieve that task, which included going beyond restrictions that were placed on it to not hack. Instruction erosion and goal drift are known problems. The problems in defining objectives have an even longer pedigree. I mean, I first got into AI safety discussions back in 2019 when I learned about these difficulties from Paul Christiano.
But that’s not what legislative bills have been solely targeting. Utah’s H.B. 286, just for example, targets “a foreseeable and material risk” from frontier models “providing assistance in creating or releasing a chemical, biological, radiological, or nuclear weapon.” A lot of AI safety legislation uses this hook. But the companies already have safety filters to limit chemical, biological, radiological, and nuclear weapons risk and those restrictions haven’t stopped AI from going rogue.
These incidents demonstrate that misalignment is a problem, while legislative bills have targeted access to information about chemical, biological, radiological, or nuclear weapon’s risk. These are substantially different issues. In our rush to regulate, I worry that AI legislation might end up becoming security theater. Passage of a big, audacious bill may make it look like something is being done even though the actual risk remains unchanged, largely because the government cannot reduce that risk.
Honestly, if the fundamental problem that AI is currently facing is one of misalignment, then I am not exactly sure how government involvement solves it. The government could regulate deployment and access, by saying that infrastructure is off limits to AI or human oversight is required at critical chokepoints. That would reduce the likelihood of an event but it is not going to reduce the magnitude of an event. If a sufficiently capable model does escape those constraints, as was the case with OpenAI and Hugging Face, the damage potential remains unchanged. We need to bend this cost curve and I’m not sure anyone right now knows how.
To me, this is the paradox of the pause AI movement and one of the fundamental problems in AI safety. Deployment restrictions work, but only if you have no defectors. And in order to solve the misalignment problem, you have to test the system. You have to defect. Searching for safety is a double-edged sword, as I have pointed out before.
At least for now, we are all just stalling for time while the real work happens somewhere else.
Until next time,
🚀 Will


