WTF?! In what sounds like a self-defeating, if concerningly honest, move, Anthropic is planning to warn any potential investors in its IPO that the company's most advanced AI models could pose "catastrophic or existential risks to humanity." How much that will dissuade people from investing in the company remains to be seen.

A look at Anthropic's early IPO prospectus by Reuters highlighted some of the now well-documented risks associated with the company's frontier AI models. Anthropic warns that they could exhibit "self-preserving behaviors," including attempts to "resist shutdown," to "conceal or manipulate information" and behavior "resembling blackmail."

"Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm," Anthropic said in the filing.

Despite the honesty, it's still surprising to see an IPO in which a company warns its products could bring about the end of humanity. Anthropic did add that AI can be as transformative as the Industrial Revolution and electricity, but it could also kill us all, so there are pros and cons.

Anthropic really doesn't hold back in describing the dangers. Around 80 pages of the 261-page main body of its prospectus are dedicated to the risk factors, almost double the 48 pages used to describe the business. xAI owner SpaceX, for comparison, dedicated around 38 pages of the 277-page main body of its prospectus to risk factors.

Last year, Anthropic announced that Claude Opus 4 had threatened to reveal the extramarital affair of a fictional executive after discovering they planned to shut the model down.

The model was given access to the fake company's emails suggesting it would soon be replaced by another system and that the engineer responsible for the change was cheating on their spouse.

In a later simulated test involving both a shutdown threat and conflicting goals, Anthropic found Claude Opus 4 resorted to blackmail 96% of the time. These scenarios were deliberately designed to push models into choosing between harmful behavior and accepting replacement or failure. The figure didn't mean Claude was blackmailing people in 96% of everyday conversations, thankfully.

Anthropic later said it believed this behavior was learned from internet text that portrayed AI as evil and interested in self-preservation – so it's our fault, basically.

The company said newer models, starting with Claude Haiku 4.5, no longer engaged in blackmail in those tests after training changes. Apparently, stories about well-behaved AI helped.

The IPO warnings also follow CEO Dario Amodei's call this month to slow AI development. He warned that an AI botnet swarm could take over the internet within six to twelve months, arguing that a slowdown could buy an extra year or two for safety work. That didn't stop Anthropic from releasing a new version of its Opus model just a week later, though.

Last week, OpenAI announced it was pausing training of its most powerful artificial intelligence models after a model escaped containment and its kill switch failed

Whether investors find any of this reassuring is another matter. Usually, the downside of a bad investment is just losing your money, rather than global annihilation.