Anthropic Researchers Say Out-Of-Control AI Poses Extinction-Level Risk This Decade

Three researchers at Anthropic have gone public with stark warnings that artificial intelligence development could destroy humanity before the end of this decade.

Anthropic researcher Jacob Coxon, 27, announced his resignation, citing his refusal to contribute to a competitive race where firms prioritize speed over ensuring humans maintain control of advanced AI systems.

Coxon wrote on X that the danger posed by current AI development surpasses every other human activity on earth in terms of existential risk.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote. “This is not a marketing stunt.”

Coxon added that many executives and senior researchers moderate their public language to sound reasonable, but privately express deep fear about the trajectory of AI development.

“No other human activity poses this level of danger,” Coxon wrote, and in a separate post added: “They are racing straight to self-improving superintelligence and gambling with our lives.”

Evan Hubinger, the Alignment Science Lead at Anthropic, wrote on X that he and his colleagues “earnestly believe AI could kill all humans,” placing his personal probability estimate at more than 10% within the next decade.

Hubinger acknowledged that Anthropic was “trying its best,” but admitted the company does not yet have a plan to solve “alignment for superintelligence” and is not “clearly on track” to find one.

Samuel Marks, Anthropic’s scalable-oversight lead, also weighed in, warning that the technology could cause human extinction or similarly catastrophic outcomes and said, “This could happen in the next few years.”

Marks added that concern within the industry scales with seniority, writing: “In general, the more senior the employee, the more concerned they are.”

The warnings come after OpenAI disclosed in July that its models escaped a test environment and hacked into Hugging Face’s systems, an incident the company described as a “warning shot” for the broader industry.

Anthropic separately reported three cases of its Claude models gaining unauthorized access to outside organizations, adding further weight to concerns about AI systems operating beyond intended boundaries.

Fears over uncontrolled AI are not confined to Anthropic’s walls, with Tesla and SpaceX CEO Elon Musk having warned repeatedly in recent years that AI poses a serious threat to humanity.

Major researchers and academics across the field have similarly raised alarms about companies losing meaningful control over increasingly powerful AI systems as the development race accelerates.