A senior AI safety researcher at Anthropic has said he believes there is a greater than 10% chance that AI "could kill all humans" within the next decade, in one of the starkest public warnings yet from inside a leading AI lab.
Evan Hubinger made the comments in a post on X on 9 September, according to the BBC, saying the risk from AI models that exist today was low but that he was worried the technology could soon develop and improve itself to the point of posing an existential threat. He did not set out a specific mechanism by which this might happen. His post came in response to one from Jacob Coxon, an AI researcher who had just left Anthropic after previously working at OpenAI, who wrote that "neither company is acting responsibly" and warned of systems that could soon hack any system, transform entire fields overnight, and accumulate real-world power.
The BBC reported that Anthropic's own August safety report had described a low risk of its models becoming misaligned in ways that could cause catastrophic harm, but added the company said it was "less confident in this assessment" than previously and was "seeing early signs of potential acceleration". The Financial Times separately reported, per the BBC, that Anthropic had withheld its latest model from the UK's AI Security Institute; Anthropic declined to comment on either the researchers' posts or the AISI situation.
Dame Wendy Hall, a computer scientist who advises the UN on AI, told the BBC she was "shocked" by the posts, while suggesting some of the commentary could reflect "PR and marketing" as Anthropic and OpenAI approach stock market listings. Former UK Treasury minister Darren Jones has written to Prime Minister Andy Burnham calling for a new multinational treaty on safe AI development.
