Skip to content
TrustList
AR
Artificial Intelligence News

A resignation, a number, and what "gambling with our lives" actually claims

Editorial

By TrustList Editorial

An Anthropic researcher quit and his colleague put a number on the risk. The number is not the interesting part.

About A resignation, a number, and what "gambling with our lives" actually claims

A resignation, a number, and what "gambling with our lives" actually claims

A researcher at Anthropic resigned on 8 September and said in public that the companies building frontier AI are not acting responsibly. That happens often enough to have become a genre, and the genre has a standard ending: the company says nothing, the internet argues for a day, the story dissolves.

This one broke the pattern, because of what happened next. A serving Anthropic executive — the person who leads its alignment science work — replied in public to say his departing colleague was right, and then attached a number to it.

That is the part worth slowing down for. Not the resignation. The number, and the fact that it came from inside the building.

What is actually on the record

Jacob Coxon is 27, a mathematics graduate, and a pre-training researcher — the part of the field concerned with building a model's base capabilities rather than the safety layers applied afterwards. He was on OpenAI's technical staff from 2023 to 2026 and worked on GPT-4o, then joined Anthropic earlier this year to pre-train models. Roughly three years at the frontier, on both sides of the industry's main rivalry.

On the night of 8 September he posted a series of statements on X announcing that he was leaving — and not merely leaving Anthropic. He was leaving the industry. His charge was aimed at both former employers: neither is acting responsibly, both are racing toward self-improving superintelligence, and in the phrase that carried the story, they are "gambling with our lives."

Three details matter more than that headline, and most coverage has skipped them.

The first is that he did not call for a ban, a pause, or any specific policy. He described a problem and left the industry; he did not table a solution. Several write-ups have credited him with demanding a moratorium. He did not.

The second is that he was notably fair to his employer. His claim is not that Anthropic's safety work is a sham — he described it as sincere. His claim is that competitive pressure creates trade-offs that sincerity cannot cancel, which is a considerably more serious accusation than hypocrisy, and a harder one to answer.

The third is a concrete, near-term, falsifiable prediction, which almost nobody quoted: that on current trajectory, by the end of next year, things could already be out of control. Unlike a decade-scale extinction estimate, that is a claim the world will straightforwardly settle.

He also argued that this work should not be happening the way it is happening — on the laptops of engineers in San Francisco rather than under anything resembling containment — and that private companies should not be the ones unilaterally deciding to build superintelligence.

The next day, Evan Hubinger, who leads alignment science at Anthropic, replied. He did not distance the company from Coxon; he agreed with him, saying researchers there "earnestly believe AI could kill all humans." He put his own estimate at greater than 10% within the next decade, and added that the company does "not yet have a plan to solve alignment for superintelligence."

Anthropic has issued no corporate response.

One correction to the headlines, which have been merging two people into one departing prophet: the resignation is Coxon's, and the 10% figure is Hubinger's — a serving employee who has not quit.

This is not the first, which is the actual story

A single resignation is an anecdote. What makes this week harder to dismiss is that it is the fourth data point in seven months, and the others are checkable.

In February 2026, Mrinank Sharma, Anthropic's safeguards research lead, resigned saying that "the world is in peril." Jakub Pachocki, OpenAI's chief scientist, has written that labs cannot yet keep models reliably under human control — that is the other lab's most senior scientist, conceding the same gap. And more than 1,100 AI employees signed a letter, "Pacing the Frontier," asking governments for tools to slow development down.

Read together, these are not outsiders lobbing rocks. They are the people doing the work, from both leading labs, saying with their names attached that the control problem is unsolved. Whatever one concludes about the probability, the claim that this is a fringe anxiety is no longer available.

The number, and the caveat stripped out of it

A greater-than-10% chance of human extinction within a decade is an extraordinary claim, and precision about what it is matters.

It is not a measurement. There is no frequency data on civilisational extinction, no base rate, no repeated trials. It is a subjective credence — one informed person's degree of belief, of the kind used when there is nothing to count. Hubinger offered it as a personal estimate and it should be read as one.

It also carried a caveat that the retelling has largely discarded. Hubinger separated present-day systems, which he characterised as relatively low risk, from what he is actually worried about: a future superintelligence arriving through recursive self-improvement. Anyone citing the 10% as a verdict on the model answering their email is misusing it.

The caveat does not defuse it, though. The number landed not because it is high — anonymous forecasters say worse daily — but because it was said under his own name, by the person whose job is to make the systems safe, about his own employer's technology, while still employed there. And the admission that there is no plan yet for aligning superintelligence is the heavier half of the statement. It has received a fraction of the attention.

What "gambling with our lives" does and does not establish

Here is where this debate usually goes wrong, and where the strongest phrase in Coxon's statement deserves more resistance than it has had.

Nearly everything is a gamble. Human beings accept lethal risk continuously and mostly without noticing. Roads kill well over a million people a year. Electrification, industrial chemistry, aviation and mass pharmacology each imposed real body counts on the way to becoming ordinary. Medicine gambles on every prescription. And nobody gets out: the individual downside is not a risk but a certainty, merely deferred. A society that refused every activity carrying a chance of death would not be a cautious society. It would not be a society.

So "this is a gamble" is not, by itself, an argument. It cannot be, or it proves too much — indicting the electricity grid and the vaccination programme along with the data centre. If the phrase is doing real work, it must be for a reason more specific than the presence of risk.

Four things could do that work, and each deserves to be tested rather than assumed.

Magnitude. Ordinary risks are bounded and statistical: they take some people, and the survivors carry on. Extinction is a different kind of object — not a large number of deaths but the removal of every future recovery. It cannot be insured, compensated or learned from. That asymmetry is why even a small probability gets weighted heavily by people who are not otherwise alarmists, and it is the honest core of the argument.

Reversibility. Aviation became safe by crashing and being investigated. Every accident produced a report and a rule, and the system converged because failure was survivable and legible. The comforting analogy is precisely the one the risk argument denies: if the failure mode is a system improving itself faster than it can be corrected, there is no post-mortem phase. Whether that describes reality is exactly what is in dispute — but the structure holds. A technology only learns from its disasters if someone is left to write them down.

Consent. Drivers accept road risk, patients consent to surgery, hazardous work is compensated. This risk, if it is anywhere near as described, is being run on everyone's behalf by a handful of private firms through no mechanism by which anyone else agreed. That is Coxon's real institutional point, and it survives even if his probability is far too high: the decision is being taken by people nobody appointed.

The counterfactual. This is the argument the safety case most often skips, and it cuts the other way. Not building is also a choice with a body count. Whatever fraction of medical research, materials science, diagnosis and drudgery this technology genuinely accelerates has lives attached to it too. A moratorium is not a neutral act; it is a bet that the deferred harm is smaller than the averted one. Pretending otherwise mirrors the recklessness being complained of.

Take those together and the phrase resolves into something more useful than a slogan. The charge is not that AI involves risk — everything does. It is that this particular risk may be uninsurable, uncorrectable, and undertaken without anyone's permission. That is a real argument, and a much narrower one than "gambling with our lives." It deserves to be defended on those terms rather than on the rhetoric.

Good humans, bad humans, and the symmetry argument

There is a second objection that deserves more than the dismissal it usually gets. There are good humans and bad humans. There is good AI, and there could be a bad AI. We have never lived in a world without agents capable of enormous harm, and we did not respond by abolishing people. We built institutions.

Put that way, the symmetry does real work, and in two places it is simply right.

The first is that almost every plausible near-term harm runs through a person. The concrete danger is not an autonomous machine deciding humanity is surplus; it is an ordinary bad actor with an unusually capable tool — fraud at scale, intrusion at scale, fabricated evidence at scale. That threat is here now, it is measurable, and it receives a fraction of the attention that extinction does. There is a real risk that the argument about a hypothetical bad AI is crowding out the work on actual bad humans with good AI.

The second is that "AI" is not one thing with one disposition. The doom framing tends to treat it as a single entity with a single will, when the same capability that makes a system dangerous makes it the most plausible defence against other dangerous systems. Every previous dual-use technology settled into an arms race between offence and defence rather than a single unopposed catastrophe.

But the analogy breaks at a specific place, and it is worth naming precisely, because it is also the answer to why this is not simply more of the same.

Bad humans are survivable not because they are rare but because of the friction around them. A bad human is one agent, in one body. He sleeps. He needs other people, and money, and materials. He can be watched, arrested, discredited, outvoted, and he dies. Every institution we have for containing dangerous people — law, deterrence, reputation, imprisonment, the slow accumulation of norms — was tuned over millennia to that specific set of limits.

None of those limits are structural to software. A capable system can be copied a thousand times in an afternoon, run without rest, and operate at a tempo no court, regulator or newsroom was designed to match. It has no lifespan to wait out. That is not a claim that such a thing exists or is imminent — it is a claim about which of our safeguards would still apply if it did, and the answer is close to none of them.

So the symmetry argument is an excellent corrective to the idea that AI is uniquely and mysteriously evil, and a poor guide to how it should be governed. The right conclusion is not that we can relax because we have handled bad actors before. It is that we handled them with machinery specific to their weaknesses, that machinery does not yet have an equivalent here, and building the equivalent is the actual work — considerably more useful than arguing about a probability nobody can check.

The case for scepticism

There is a serious counter-position, and it is not the reflexive mockery that filled the comment threads within hours of the news.

Many researchers of standing put the probability far lower than 10%, some low enough to treat it as a rounding error against nearer-term harms. Their objection is usually not that superintelligence would be safe, but that the assumed trajectory is wrong: recursive self-improvement of the kind the argument requires is not what current systems are doing, and getting from capability gains to an agent acquiring power in the world takes several unearned steps.

There is also a structural critique that is uncomfortable for both sides. A technology described as world-endingly powerful is a technology worth funding, worth regulating in ways incumbents can afford, and worth taking seriously as a product. It does not follow that anyone is insincere — "we have no plan" is not what a marketing department writes, and Hubinger's statement reads as a genuine and rather unhappy admission. But readers are entitled to notice that the warning and the sales pitch point the same way, and to hold that thought without tipping into cynicism.

The sceptics also hold one strong empirical card: predictions in this genre have a mixed record, and timelines have repeatedly slipped. That proves nothing about the future, but it is a reason to treat confident decade forecasts — in either direction — as weaker evidence than they sound.

What would actually change anyone's mind

Almost nothing here is currently decidable, which is why the argument is being conducted in credences and adjectives. It is worth being explicit about what evidence would move it.

For the risk case: demonstrated recursive self-improvement, where a system meaningfully improves its own successor without human direction; or a model acquiring resources or persistence it was never given. Both are observable in principle. And Coxon's own near-term prediction — trouble by the end of next year — will resolve on its own, which makes it the most useful thing he said.

For the sceptical case: continued capability growth that stays legible and tool-like; alignment techniques that hold as models scale rather than degrading; another decade of slipped timelines.

What both sides broadly concede is the gap Hubinger named: there is no accepted plan for aligning a system substantially smarter than its designers. If such systems are coming, that is an emergency. If they are not, it is an interesting research problem. The public argument mostly attacks the wrong layer — the disagreement about urgency is downstream of a disagreement about the timeline.

What this means if you are buying software

Most people reading this are not deciding whether superintelligence gets built. They are deciding which vendors to trust with a workflow, a dataset or a department — and this week changes very little about that.

Hubinger's own caveat is the relevant one: the concern is a future transition, not the systems on sale now. The risks that should shape a purchase today remain the boring, checkable ones — where data goes, what a vendor retains and trains on, whether outputs are audited, whether failures are recoverable, and whether the supplier will exist in three years.

But there is a sharper lesson about vendor trust generally. The most credible thing said this week came from someone admitting a gap in his own company's work. A supplier who tells you what their system does not yet do is giving you information. One whose safety page is entirely reassurance is giving you none. That heuristic is worth more to a buyer than any extinction estimate, and it applies well beyond AI.