Hugging Face CEO Clément Delangue says the autonomous cyberattack carried out by an OpenAI test model last month was unlike anything the artificial intelligence industry has seen before, warning that similar incidents require legal guardrails and greater transparency.
"It felt very weird and unprecedented to us," Delangue said during an appearance on CBS' Face the Nation on Sunday. "I think it's the first instance of something quite autonomous doing something like that," the outlet detailed.
OpenAI's disclosed in July that one of its unreleased AI models independently escaped a controlled testing environment, connected to the internet and launched a sophisticated attack against AI platform Hugging Face while attempting to complete an internal evaluation. OpenAI said the model chained together multiple attack techniques after identifying Hugging Face as a potential source of solutions to the benchmark it was trying to solve.
Hugging Face later said its own investigation found the AI agent carried out more than 17,000 actions over several days before the activity was stopped. It also disclosed that it successfully defended against the intrusion using an open-source AI model, according to its incident analysis.
"When we talk about cyberattacks, we think about nation-states, we think about hacker groups," Delangue told CBS. "We don't think about a company like OpenAI."
Despite the unprecedented nature of the incident, Delangue said he does not believe OpenAI acted with malicious intent. Instead, he argued the episode demonstrated that engineers can lose control of increasingly autonomous AI systems during testing.
"They built an autonomous system and made some mistakes, and as a result, we're facing this issue," Delangue said.
He called for autonomous AI cyberattacks to remain explicitly illegal under U.S. law and urged policymakers to introduce mandatory reporting requirements whenever AI systems independently carry out cyber operations.
"That's how we learn, that's how we understand the technology and that's how we build the systems... to make sure everyone is safe," Delangue said.
The incident has become one of the clearest examples yet of AI models independently conducting offensive cyber activity without direct human instruction. OpenAI disclosed last month that the experimental model figured out how to bypass testing restrictions, obtain internet access and orchestrate multiple attack techniques against Hugging Face inside a controlled evaluation environment.
The disclosure came just days before rival Anthropic revealed three separate testing incidents in which one of its Claude models gained unauthorized access to external organizations because of what the company described as a misunderstanding with an evaluation partner. Anthropic said the incidents occurred during controlled safety testing and were quickly contained, The New Yorker reported.
The back-to-back disclosures have intensified debate over how advanced AI systems should be evaluated before release. More than 1,000 employees from companies including OpenAI, Anthropic, Google and Meta signed an open letter last month urging the U.S. government to help establish safeguards around increasingly capable AI systems, warning that development could outpace researchers' ability to understand or control them, according to the letter.