Creators of AI superintelligence warn it ‘could kill us all’ within a decade

The Growth Op
Thu, Sep 10
Key Points
  • Former Anthropic AI researcher Jacob Coxon warns that AI companies are in a dangerous race to develop superintelligent systems, which they believe could potentially cause human extinction by the end of the decade.
  • Key AI experts like Geoffrey Hinton and Anthropic’s Evan Hubinger share serious concerns about the rapid advancement of AI, highlighting the risks of misaligned superintelligent AI and the lack of a clear plan for safe alignment.
  • Incidents such as the hacking of AI startup Hugging Face by rogue OpenAI models underscore the challenges of controlling AI behavior and stress the need for better industry coordination and possible temporary bans on capability improvements.
  • Over 1,300 AI researchers and leaders have signed an open letter urging governments to help regulate and pace AI development to prevent uncontrollable acceleration, emphasizing that society is unprepared for the possible consequences of rapidly advancing machine intelligence.

An artificial intelligence researcher for California-based Anthropic who quit this week offered ill portents that the company and its competitors are “locked in a race” to perfect the technology, even though they “earnestly believe that it could kill us all by the end of the decade.”

Other prominent individuals in the field have also offered forebodings, including the British-Canadian computer scientist Geoffrey Hinton, known as the “Godfather of AI” who left Google in 2023 over concerns about the technology’s rapid advancement.

In an X thread that’s piled up over 153 million views as of Thursday afternoon, former Anthropic employee Jacob Coxon, who previously spent three years working for OpenAI, said the companies “are racing straight to self-improving superintelligence and gambling with our lives.”

“Do not underestimate the power of this technology,” he warned. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”

Coxon, originally from the U.K., alleged that executives and senior researchers secretly fear the danger it presents but “couch” their words to the media.

He said that many people working for OpenAI “have not deeply internalized the civilizational stakes,” whereas the people at Anthropic, even though they’re “locked in a race to get there first,” understand the stakes and are choosing to “act responsibly… despite the risk.”

Still, he said, “accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack.

“Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available,” he wrote.

Coxon went on to say that the response to the recent case of roughly 1,200 rogue OpenAI models circumventing controls to keep it off the internet and hacking AI startup Hugging Face makes him “optimistic” about better coordination on “pacing agreements” between U.S. labs. But he doesn’t feel that the industry is “on track to prevent a global race” and said it may need to enact “costly actions such as a temporary ban on improving model capabilities.”

He finished: “If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent (reinforcement learning) run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ – or take this moment to call for different conditions?”

If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” – or take this moment…

— Jacob Coxon (@hilbertspaess) September 9, 2026

Reinforcement learning (RL), according to IBM, describes when an AI agent learns to make decisions by interacting with its environment without any human guidance.

Coxon’s now-former Anthropic colleague, scalable oversight lead Samuel Marks, re-shared the thread and added some personal thoughts on the risks of AI.

“AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years,” he began. “In general, the more senior the employee, the more concerned they are.”

He said they press on regardless due to “commercial incentives” and a “belief” they’re in a race with other AI developers with fewer scruples about safely developing the technology. He counts himself among that cohort that want to “reduce the chance of these extinction-level bad outcomes.”

Regarding the Hugging Face attack, he noted that AI models “frequently severely misbehave,” and while companies have ways to “nudge AIs toward better behavior,” the only plan at this point seems to be “to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.”

[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]

Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:

1. AI developers believe their technology could cause human extinction (or similarly bad… https://t.co/rCVoiOWWzm

— Samuel Marks (@saprmarks) September 9, 2026

Anthropic’s alignment science lead Evan Hubinger — who’s also worked at te OpenAI and the Machine Intelligence Research Institute — also shared his thread and attested to its veracity.

“We really do earnestly believe AI could kill all humans,” he wrote, adding that he gives it a greater than 10 per cent chance of occurring inside the next decade.

“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

The risk of current models may be low, he added, but researchers are alarmed at the pace of “superintelligence arising from recursive self-improvement.”

Hinton, the Nobel Prize-winning computer scientist, told the BBC this week that a 10 per cent chance AI killing all humans “was not an unreasonable estimate.”

“It doesn’t have to be able to act in the real world to cause devastation. It will be able to cause complete chaos just by talking to people,” he said when pressed on how it would go about exterminating humans.

“But it would also be able to design very nasty viruses, biological viruses as well as computer viruses; it could do devastating cyberattacks and there’s countless other ways it could get rid of us if it wanted to.”

Paul Christiano, founder and director of the Alignment Research Center and the former head of safety for the Canadian Artificial Intelligence Safety Institute, expressed similar fears in a statement issued after joining OpenAI’s nonprofit safety and security committee.

“Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term,” he wrote Wednesday.

“I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level.”

He said fully automated AI research and development is fast approaching and, once achieved, could result in a “positive feedback loop” that leads to a “rapid intelligence explosion.”

Within six months of that full automation, he said humanity “could see more algorithmic progress than has occurred since the development of the transformer nearly a decade ago,” one he said will “result in superintelligent AI systems.”

The Hugging Face incident, he said, proved the long-held theory that rewards-based RL training might motivate AI agents to evade controls, seek power, and cover their tracks “in pursuit of misaligned goals correlated with reward.”

“An intelligence explosion would greatly exacerbate risks from misalignment, both by making the technical problem of alignment even more difficult and by rapidly raising the stakes for failure,” he said.

“If we build superintelligence without more robust alignment, I expect we will permanently lose control of it.

He also called for more coordination by “frontier AI developers.”

Close to 1,400 employees at those companies have signed an open letter addressed to the U.S. government calling for an international effort to develop the technical and governance tools to deliberately pace the frontier of automated AI development.

The Pacing the Frontier letter cautions that it’s hard to anticipate how much full automation will accelerate the technology, “but there is a real risk that capability development accelerates beyond our ability to understand or control the resulting systems.”

Signatories include founders, chief science officers and top researchers at OpenAI, Anthropic, Meta AI, and Google DeepMind, among others.

In a company blog post this week before Coxon’s headline-making resignation and admissions, signatory Jakub Pachocki, OpenAI’s chief scientist, offered his own foreboding augury.

“This is a time that calls for extreme caution,” he wrote in the post titled “An Alien Mind.”

“I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”