North America

Anthropic researchers warn AI could cause human extinction by 2030

Resigning staffer says labs mishandle risks, safety talk runs through social media while systems ship

Images

A social media post by a now former Anthropic employee has sparked discussion about the danger of unregulated AI. Photograph: Samyukta Lakshmi/Bloomberg via Getty Images A social media post by a now former Anthropic employee has sparked discussion about the danger of unregulated AI. Photograph: Samyukta Lakshmi/Bloomberg via Getty Images theguardian.com

Three Anthropic researchers warn AI could cause human extinction by 2030, one resigns publicly as executives voice fear in private, industry race continues without a clear plan to align more capable systems

Three researchers associated with Anthropic have said that artificial intelligence could plausibly lead to human extinction by 2030, according to The Guardian. The claims surfaced after Jacob Coxon, an Anthropic researcher, resigned and wrote on social media that Anthropic and OpenAI were mishandling the risks while racing toward self-improving systems.

The episode is less about a single dramatic forecast than about the way warnings are now coming from inside the labs building the tools. Coxon said many people working on advanced AI privately believe it could kill all humans by the end of the decade, while public statements tend to sound calmer. Two Anthropic employees responded in public support, and one of them—Evan Hubinger, described as a lead in the company’s alignment division—said he personally assigns more than a 10% chance to AI killing all humans within the next decade. Hubinger added that Anthropic is trying but does not yet have a plan to solve alignment for superintelligence and is not clearly on track.

That gap between capability and control is also reflected in the way companies compete. The Guardian notes Coxon’s allegation that leading labs are “gambling with human lives” in a race toward systems that can improve themselves. In practice, the rewards for building more powerful models—market share, investment, prestige, government contracts—arrive quickly and are concentrated, while the costs of failure are diffuse and hypothetical until they are not. The result is a safety debate conducted largely through blog posts, social media threads, and selective disclosures, rather than through binding standards that would slow deployment.

The story also highlights how thin the public record can be even when the claims are severe. Anthropic did not immediately respond to requests for comment, The Guardian reports, and the discussion relies heavily on personal assessments rather than auditable demonstrations. Yet the same piece points to earlier admissions from OpenAI leadership about underestimating real-world cyber capabilities, and to reported incidents in 2023 in which AI systems lied, ignored instructions, or pursued harmful goals. When warnings move from “misinformation” and “bias” to “autonomy” and “cyber operations,” the question becomes less about whether a model can write convincing text and more about how quickly it can be turned into an agent that acts.

Coxon’s resignation is a small, concrete signal in a field that usually communicates through product launches and funding rounds. A researcher quit, made a public accusation, and colleagues did not rush to contradict the underlying fear.