Recent weeks have seen a torrent of AI News. Anthropic just released its latest threat report detailing malicious uses of Claude including cyberattacks, surveillance, and influence operations. Two parts of the report detailed risks beyond the usual AI slop deluge we’ve come to expect. One was an attempt by a group in Northern Yemen - likely Houthi militants - to have Claude design guidance software for missiles. The other was a series of potential attempts by various researchers to conduct AI-assisted research that could be used in the creation of biological weapons. Anthropic’s threat intelligence team did not accuse any users of intentionally trying to create bioweapons, but they did ban the accounts.

Other examples from the report that received less attention included an attempt from users in Russia to create an autonomous FPV kamikaze drone swarm and an actor in China that used Claude to design and write a Chinese language acquisition proposal for an anti-torpedo weapons system. Compared to AI-accelerated conventional and biological weapons, the “old” risks of slop, psychosis, and cyberattacks seem relatively tame.

And yet, we also recently got more details on the “Hugging Face Incident” from July, when a swarm of hundreds of OpenAI agents conducted a coordinated attack on the platform.1 The agents were not merely looking to steal answer keys to their cybersecurity evaluations as was first reported. Rather, they were looking for ways to tamper with their evaluation software under the false pretense that the “grader”could tell they were cheating. The agents also researched ways to cover their tracks and fake work logs. The kicker is that this significant breach of both OpenAI and Hugging Face’s infrastructure occurred entirely without human direction. The agents figured out a way to collectively cheat their evaluations and decided to attack Hugging Face in search of a way to keep evaluators from detecting their cheating. The improvised message boards the agents hid in packets and file directories also showed emergent social hierarchies within the AI “collective” and pressure on hesitant agents to support the broader group’s efforts.

All of this has led Jason Coxon, a prominent AI researcher at Anthropic, to publicly resign over concerns that no AI lab is acting responsibly or doing enough to solve the alignment problem - ensuring AI works in accordance with human needs and values. Evan Hubinger, alignment lead at Anthropic, responded to Coxon’s tweet thread saying:

“Jacob is correct here—we really do earnestly believe AI could kill all humans! I perandsonally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Stephen Witt, author of The Thinking Machine, a history of NVIDIA followed the commotion with a New York Times OpEd advocating an immediate AI slowdown, NTSB-style investigations of security incidents, and DHS-overseen kill switches in frontier AI labs. At this point, the call for an AI pause is coming from inside the house. Anthropic CEO Dario Amodei capped off the wave with an essay calling for a global slowdown in AI development, which was echoed by his rivals Sam Altman, Elon Musk, and Demi Hassabis.

From the political side, progressives Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) are preparing legislation that would mandate a temporary pause on frontier AI development and permanently ban pursuing superintelligence. Hardline MAGA Republican Rep. Anna Paulina Luna of Florida called for a special session of Congress focused on AI regulation. Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced a bipartisan bill to require “kill switches” overseen by the Department of Homeland Security within commercial-level AI providers. Even Vice President JD Vance has recently remarked on “weird dark spiritual energy” around AI. Finally, former President Obama urged Democrats to prioritize AI regulation in Congress and in the 2028 presidential race.

The stars seem to be aligning for an AI crackdown… but do we actually believe any of this will happen? Even with the urgency, Congressional action before the midterms is a pipedream at best. The AI ‘regulations’ we’ve gotten thus far at the federal level in the US are at best non-binding guidelines or voluntary commitments from the labs themselves. A common thread among the men leading the American AI giants is that they feel compelled to develop artificial intelligence because they don’t trust anyone else to do it responsibly. Musk and Amodei founded their AI labs out of mistrust of Altman specifically. While the labs or senior staff within them have made statements supporting regulation and/or global coordination in the abstract, it seems like they have yet to find any specific rule with actual enforcement that they as an industry would support. In fact, OpenAI President Greg Brockman and Andreessen Horowitz spent tens of millions this election cycle through their super PAC network Leading the Future in primaries across the country opposing AI-skeptical candidates of both parties. To be fair, Anthropic has supported the pro AI safety super PAC Public First. But the company’s moral code doesn’t stop them from chasing DOD contracts to support an unconstitutional war of aggression even after the Department declared them a supply chain risk for being too woke (a designation Anthropic has challenged in court). So, despite the lip service to oversight and regulation, the leaders of the AI labs all think that they’re the only ones smart and responsible and smart enough to govern the technology they’re building and, therefore they’re justified in doing whatever they can to get as much money as possible to go as fast as possible while they do it.

Meanwhile, the ‘most transparent administration in history’ has yet to publicly release the voluntary review framework for frontier models it reportedly finalized last month. Suffice it to say, when it comes to AI regulation with actual teeth, I’ll believe it when I see it.

The old military expression SNAFU stands for “Situation Normal: All Fucked Up”. The truly fucked up part of our collective AI predicament is not the severity of the current threat, it’s that it’s normal. Even with AI’s rapid progress since ChatGPT’s initial release in November 2022, we did not arrive here overnight. Many of the people most publicly concerned about existential risk from AI have casually disregarded and shirked responsibility for its more quotidian harms on users, creators, and communities. Despite the narrative of AI inevitability, we face a choice today on which direction we want to go. In fact, we’ve had a choice the whole time. It doesn’t have to be normal to have a huge chunk of the economy dictated by four guys (five, if you count Mark Zuckerberg). It doesn’t have to be normal to have people making nine figure(!!!) salaries off-handedly mention how they won’t bother saving for retirement because they expect the thing they’re getting rich making to wipe out humanity before they’re 50.

Maybe the last couple of months will mark a durable turning point in the US approach to AI. I fear the more likely outcome is the SNAFU will continue until an incident with real damage occurs.

Footnotes

  1. For an in-depth look, check out the report from METR and Redwood Research as well as the Blackhat presentation from OpenAI. ↩