Going rogue
More on the OpenAI/Hugging Face incident
After I reported here nine days ago on the OpenAI/Hugging Face incident, the amount of discussion about the incident has been so vast that no summary could hope to do it justice. This is as it should, because the incident is the clearest warning shot so far that AI is, in terms of agency and raw intelligence, approaching levels where it may become catastrophically dangerous. So rather than trying to offer a balanced overview of the discussion, I’ll do something far less ambitious: to offer a few quick thoughts on three specific directions it has taken.
1. On July 25, there was an article entitled No, OpenAI’s models didn’t go ‘rogue’ when they broke into Hugging Face. Here’s what really happened in Science Live. The cybersecurity expert who is interviewed is eager to tone down the incident:
“If there’s a failure here, it isn’t that the AI wanted to hack something,” Oli Buckley, a professor in cybersecurity at Loughborough University in the U.K., told Live Science. “It’s that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system’s capabilities, and underestimated how effective the model would be at finding an unexpected path to success.”
[…]
It’s notable that the models found a flaw in the infrastructure designed to contain an AI and used it to reach the public internet. Describing the models as having “gone rogue,” however, risks assigning them unsupported motivations, Buckley said.
“I think I’d be wary of jumping to “rogue AI,”“ Buckley said. “The models didn’t develop their own agenda or decide to attack Hugging Face while twirling their digital moustache.”
In terms of what actually happened, there is little or nothing to object to in the article. While the AIs did decide to attack Hugging Face, it did not do so “while twirling their digital moustache”, because they simply do not have a moustache, digital or not.
Still, I find the “didn’t go rogue” framing deeply misleading. Buckley is free to define terms any way he likes, and it is true that the model did not replace the objective given to it by a new one of its own. It was still trying to maximize its score on the test offered by the OpenAI engineers. But this is closely analogous to what happens in the classical version of the thought experiment involving a paperclip maximizer, whose catastrophic behavior arises precisely because it pursues the objective it was given with relentless instrumental competence. If we use the phrase “didn’t go rogue” in a way that encompasses the behavior of an AI carrying out a paperclip apocalypse, then the term is robbed of its intuitive connotation “…and hence there is no need to worry”. If Buckley is unfamiliar with the theory of instrumental AI goals and the central lesson that such goals can be just as dangerous as if the AI invented its own sinister final goal, then I recommend that he turns to the excellent treatments of this topic in one of the seminal books by Bostrom, or Russell, or Yudkowsky and Soares.
2. Yesterday, I wrote about the incident in the Swedish news outlet Kvartal. One of the things I did was to criticize robotics professor Hedvig Kjellström at KTH for having suggested, when interviewed in Dagens Nyheter, that the whole thing might just have been a PR stunt from OpenAI to demonstrate how capable their models are. This very clearly is not the case, but I quickly received correspondence from people, including one academic computer scientist (not Kjellström) from a notoriously anti-AI-safety research environment, who defended the fake PR stunt interpretation. It worries me that there are people and communities who remain so convinced about the impossibility of AI agents independently acting in the real world and outside the scope of their users’ intentions that now that we start to see it actually happening they revert to an “it’s all just lies” position.
3. Is the OpenAI/Hugging Face incident the first of its kind, where a frontier AI under internal evaluation escapes to the Internet and hacks a third party? Apparently no, and in a report yesterday Anthropic admitted to having three known such incidents of their own,1 the earliest going back to April this year. I recommend reading the report, but the AI researcher whose Twitter name is davidad offers the following TL;DR.
If anyone is tempted to think “phew, so this is already a standard thing, and nothing truly bad has happened, so nothing to worry about here”, this is exactly the wrong reaction. The models haven’t yet (to our knowledge) hacked hospitals, banks and nuclear weapons systems, but how long do we have until that happens?
I am of course aware that this is unlikely to impress the conspiracy theorists discussed in item 2 above.


