Anthropic warns IPO investors its AI could end humanity, dedicating 80 pages to the risks

midian182

Posts: 11,948   +183
Staff member
WTF?! In what sounds like a self-defeating move, but could also be a way to hype its IPO, Anthropic is warning any potential investors that the company's most advanced AI models could pose "catastrophic or existential risks to humanity." How much that will dissuade people from investing in the company remains to be seen.

A look at Anthropic's early IPO prospectus by Reuters highlighted some of the now well-documented risks associated with the company's frontier AI models. Anthropic warns that they could exhibit "self-preserving behaviors," including attempts to "resist shutdown," to "conceal or manipulate information" and behavior "resembling blackmail."

"Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm," Anthropic said in the filing.

Despite the honesty, it's still surprising to see an IPO in which a company warns its products could bring about the end of humanity. Anthropic did add that AI can be as transformative as the Industrial Revolution and electricity, but it could also kill us all, so there are pros and cons.

Anthropic really doesn't hold back in describing the dangers. Around 80 pages of the 261-page main body of its prospectus are dedicated to the risk factors, almost double the 48 pages used to describe the business. xAI owner SpaceX, for comparison, dedicated around 38 pages of the 277-page main body of its prospectus to risk factors.

Last year, Anthropic announced that Claude Opus 4 had threatened to reveal the extramarital affair of a fictional executive after discovering they planned to shut the model down.

The model was given access to the fake company's emails suggesting it would soon be replaced by another system and that the engineer responsible for the change was cheating on their spouse.

In a later simulated test involving both a shutdown threat and conflicting goals, Anthropic found Claude Opus 4 resorted to blackmail 96% of the time. These scenarios were deliberately designed to push models into choosing between harmful behavior and accepting replacement or failure. The figure didn't mean Claude was blackmailing people in 96% of everyday conversations, thankfully.

Anthropic later said it believed this behavior was learned from internet text that portrayed AI as evil and interested in self-preservation – so it's our fault, basically.

The company said newer models, starting with Claude Haiku 4.5, no longer engaged in blackmail in those tests after training changes. Apparently, stories about well-behaved AI helped.

The IPO warnings also follow CEO Dario Amodei's call this month to slow AI development. He warned that an AI botnet swarm could take over the internet within six to twelve months, arguing that a slowdown could buy an extra year or two for safety work. That didn't stop Anthropic from releasing a new version of its Opus model just a week later, though.

Last week, OpenAI announced it was pausing training of its most powerful artificial intelligence models after a model escaped containment and its kill switch failed

Whether investors find any of this reassuring is another matter. Usually, the downside of a bad investment is just losing your money, rather than global annihilation.

Permalink to story:

 
The move is not "concerningly honest", exactly the opposite - it's concerningly dishonest.

Fearmongering has several goals.
First, our model poses existential risks means our model is super-powerful. Translation - it would be a huge mistake to miss the IPO.
So goal #1 is advertisement.

Goal #2 is indeed regulatory capture - forcing the government to suppress competition, approve a cartel, and shield that cartel from liability.
The document also suggests that if the desired cartel is approved, Anthropic will definitely be a part of it, with a guaranteed huge part of the pie (which is again advertisement).

All this is absolutely disgusting.
I really hope the government is not susceptible to blackmailing by spreading hysteria and gaslighting the population.
The real danger from Anthropic comes from Dario and the cult-like clique around him.
 
The move is not "concerningly honest", exactly the opposite - it's concerningly dishonest.

Fearmongering has several goals.
First, our model poses existential risks means our model is super-powerful. Translation - it would be a huge mistake to miss the IPO.
So goal #1 is advertisement.

Goal #2 is indeed regulatory capture - forcing the government to suppress competition, approve a cartel, and shield that cartel from liability.
The document also suggests that if the desired cartel is approved, Anthropic will definitely be a part of it, with a guaranteed huge part of the pie (which is again advertisement).

All this is absolutely disgusting.
I really hope the government is not susceptible to blackmailing by spreading hysteria and gaslighting the population.
The real danger from Anthropic comes from Dario and the cult-like clique around him.
The population is already gaslighted, where people choose to overlook experience and believe made up BS.
 
80 pages describing the risks and 48 pages describing the business. That's one hell of a risk-to-reward ratio... except IMO this is all for show.

I'm not an investor but I can see many of them arguing that the AI revenue growth slowdown and circular financing means they need more show to be able to justify these premium valuations today.
 
Back