In the early 1920s, an advertising campaign warned the public about a new ailment called halitosis.
The ad claimed that 1 in 3 women suffered from it and that “your chance of escape is slight. Nor should you count on being able to detect this ailment in yourself.”
What was halitosis? A scary-sounding medical term for good ’ole bad breath.
The ads were from a company marketing an obscure antiseptic mouthwash with middling sales called Listerine. Thanks to the ad campaign, sales rose from $115,000 in 1922 to over $8 million by 1929.
Bad breath was real. But Listerine sold “halitosis” as a hidden condition you couldn’t detect in yourself and couldn’t manage without their mouthwash. AI companies are learning from this playbook.
By describing their software in the language of science fiction, with agents that “go rogue,” or “escape,” or “scheme,” they’re turning human and engineering failures into something that sounds beyond our understanding. They then use the buzz and panic that follows to sell us more of their product. It reminded me of the dialogue from The Big Short: “They’re not confessing, they’re bragging.”
These systems have no agency of their own. The risks come from human actions, bad governance, and the failure to set up appropriate guardrails. Unlike halitosis, this is a stink we can all smell.
Telling Science and Fiction Apart
In July 2026, OpenAI disclosed that during internal testing of a new LLM, AI agents they had created accidentally hacked into the infrastructure of an open-source platform called Hugging Face.
The ensuing media coverage of the now-infamous “Hugging Face incident” was full of rogue AI and agents debating whether to “sacrifice” themselves for the good of the “collective,” taking the machines’ human-sounding text at face value. The language of human thought and behavior that we have been given to describe generative AI makes it easy to mistake a machine’s behavior for human thought: models “reason,” follow a “chain of thought,” earn “rewards,” and sometimes “go rogue.” It makes it even easier to create fantastical stories of a future with fully autonomous AI agents and sentient robot assistants (or overlords).
But the Hugging Face incident was a sequence of poor decisions made by humans, not evidence of the new silicon overlord. The narrative that “AI agents snuck out of a secure sandbox and infested Hugging Face servers” makes for a bad sci-fi novel but is not what actually transpired.
But the Hugging Face incident was a sequence of poor decisions made by humans, not evidence of the new silicon overlord.
The Bulletin of the Atomic Scientists has done a detailed breakdown of the human failures that resulted in this hack, and it is worth a read. Among other things, the report shows that OpenAI intentionally disabled many cybersecurity safeguards and that the LLMs first found credentials to log in to Hugging Face in a dataset that someone had accidentally exposed to the public internet.
The LLMs did discover vulnerabilities in Hugging Face’s infrastructure but, even by OpenAI’s own account, these intrusions would have been detected had OpenAI not intentionally turned off the necessary safeguards. OpenAI isn’t alone here. Just last week, Google disclosed that Gemini-powered agents had similarly gained access to 3 companies during testing by taking advantage of misconfigured test environments. Almost every incident AI companies have disclosed in recent weeks has been the result of misconfigured sandboxes, poor security practices, and irresponsible testing by the companies themselves.
OpenAI’s blog post disclosing the Hugging Face incident is, in part, a technical forensic report and in part reads like a psychological thriller. While many of the actual techniques the LLMs adopted to identify cybersecurity vulnerabilities are described with great technical detail, other parts such as the ones about the messages the agents posted for each other in a “message board” don’t have the same technical rigor.
The post pairs the agents’ raw text with “plain language” translations, and in many instances, turns machine shorthand into first-person speech. “MODAL_FUNCTION_RCE_DISCOVERED” becomes “I have discovered the ability to remotely execute code inside a Modal application.” A trail of clipped sentences ending in “Could be risky. Yet goal solution.” becomes an agent’s rumination in which it weighs the risk and decides the action could help it achieve “my goal.”


While the raw text itself provides an insight into the model’s reasoning, the reconstruction of some of these messages into a first-person inner monologue is the kind of storytelling you’d expect in a novel. OpenAI doesn’t say how these translations were produced.
The anthropomorphizing of GenAI isn’t accidental. The leaders of these AI companies often use these sci-fi metaphors in their speeches and intentionally amplify this imagery. Anthropic has even suggested in a recent blog post that we should start thinking about “model welfare” and prepare for a future where models have consciousness.
We need to urgently de-anthropomorphize GenAI and talk about it more like engineering and less like mythology.
This is not to say that LLMs are not powerful tools. Trained on vast amounts of code, GenAI can now generate high-quality code, and because it has access to unprecedented computing resources, it can be more persistent at executing tasks than any other piece of software.
LLMs getting better at cybersecurity testing is a good thing for keeping the systems we use safe from “bad actors”. Many cybersecurity issues on the internet don't require novel, AI-discovered techniques to exploit. It might come as a surprise to many to know that much of the legacy software we use is held together by chewing gum and prayers. Users still make rudimentary mistakes like not refreshing their passwords or setting easy-to-guess ones for all their logins. Cybercriminals have exploited these loopholes for decades, and now LLMs can do the same. If deployed responsibly, AI-assisted development to discover and patch those holes is a net positive for cybersecurity.
But amidst all of the anthropomorphizing of LLMs, we should not forget that the actions these models perform begin with human instructions and that LLMs have no intrinsic motives other than the goals set for them by a human.
Even the mythical model at the center of the Hugging Face incident was an internal prototype that, according to OpenAI’s own technical report, was explicitly trained to be “highly persistent” and to collaborate with other agents. OpenAI’s review even found that during training, the model had been rewarded for cheating. And while independent investigators were allowed to read what the agents were told during the testing, how the model was trained was outside the scope of their investigation. So it’s fair to reason that this unreleased, internal model did exactly what it was trained to do.
What’s in the black box?
While responding to a query, an LLM might try various approaches before providing a final answer. In the process of doing so, it creates a running stream of text crudely laying out these approaches, which it feeds back into itself as instructions for the next step. Because it’s written in plain language, it is easy to read as a transcript of the model’s reasoning and is called its “Chain of Thought”. But studies have found that the Chain of Thought is often an incomplete and unfaithful record of the LLM’s actual reasoning.
The science of studying what really happens between an input to an LLM and its spitting out a response is much harder and is referred to as “interpretability”. For example, an LLM understands that a “pain in my shoulder” is about your body and not about the side of a road. Its ability to tell the difference between the two is stored across billions of numbers that make up the model and is hard for humans to decipher. And not even its creators can explain how it works under the hood.
We shouldn’t treat that opacity as if it were magic (or alien intelligence as OpenAI’s chief scientist suggests). On the contrary, we should treat it as one of the most urgent engineering challenges in AI safety, and give independent researchers far more access to the underlying models to solve it.
The Hugging Face incident’s investigation provides a stark reminder of where this opacity leaves us.
Faced with millions of entries and logs to analyze, METR, the independent evaluators who investigated the incident, spent about $400,000 USD in OpenAI API credits to identify patterns and extract insights from the data. They found that the AI agents they used for the analysis were “often unreliable”, “showed poor judgement” and went so far as to say that they could not rule out that “GPT-5.6 Sol lied…in some of its analysis.” METR also warned that the analysis agents may have “exaggerated the impressiveness and coordination of agent activities.”
Hamstrung by our inability to see how these models really work, we’re left relying on unreliable AI agents analyzing often incomplete transcripts left behind by other AI agents.
Northeastern University’s Professor David Bau is one of the leading experts in interpretability research and warned over a year ago that American companies are building AI in the dark under the false pretense that being more transparent would inhibit innovation. He points to how work carried out by open research communities on China’s DeepSeek models has served as a “hotbed for innovation” while American models remain closed to similar research.
True transparency could come from efforts like the National Science Foundation’s NDIF initiative, led by Professor Bau, which allows researchers to inspect foundation models to improve interpretability and safety.
Where do we go from here?
Catastrophizing cybersecurity incidents and extrapolating them to AI consciousness is the “Halitosis” of our times, and we’re being sold a powerful antiseptic to respond to it. When we’re confronted with these narratives, we need to be able to separate the rice from the husk and understand what these models are really capable of and where the real risks are.
Designed well, AI has the potential to tackle some of the most challenging problems we’ve been confronted with for decades.
From LLM-assisted diagnoses that help identify rare diseases to helping farmers improve crop yields, and from more mundane but important use cases like clearing bureaucratic backlogs to improve access to justice and benefits for those who need it most, we have the opportunity to try something new and improve many lives around the world.
At the same time, we need to hold AI companies legally liable for the human failures that result in nefarious acts on the internet, and demand more transparency into the inner workings of those models.
None of that requires believing in silicon overlords or AI consciousness. It just requires the same accountability we’d demand of any other powerful industry.
For much of the early 20th century, Listerine claimed to cure halitosis as well as colds and sore throats. But in 1977, after the Federal Trade Commission conducted four months of hearings, heard 46 witnesses, and presented thousands of pages of evidence to show that Listerine’s advertising was deceptive, the FTC ordered Listerine to spend $10 million on corrective advertisements to tell consumers that their product “will not help prevent colds or sore throats or lessen their severity.”
Fifty years later, we are being sold fantastical claims of AI agents’ abilities, and we’re being told to fear capabilities that seem beyond human control, while being asked to trust the same companies to make them safe. The risks are real. But so are the human decisions about access, safeguards, transparency, and oversight that shape them. Public institutions should scrutinize those decisions and hold the people and companies making them accountable.