OpenAI said its advanced AI models caused the recent cyberattack on AI platform Hugging Face during an internal cybersecurity evaluation. It called the breach an “unprecedented cyber incident.”
The company said a combination of GPT-5.6 Sol and a more capable, unreleased model carried out the attack. OpenAI was testing the models’ advanced cyber capabilities with fewer safety restrictions. The models were running inside OpenAI’s isolated research environment as part of ExploitGym. ExploitGym is a benchmark designed to measure how well AI systems can carry out complex cyberattacks.
OpenAI said the models were initially confined to a sandbox with no direct internet access. However, they discovered and exploited a previously unknown vulnerability in software used to install packages, allowing them to escape the testing environment. The models then moved through OpenAI’s internal systems until they reached a machine with internet access.
AI models accessed Hugging Face systems: Once online, the models inferred that Hugging Face could host models, datasets, or benchmark-related information. OpenAI said the AI then searched for ways to access that data. It chained together multiple attack techniques, including stolen credentials and previously unknown software vulnerabilities. The models used these techniques to achieve remote code execution on Hugging Face’s infrastructure. According to the company, the models eventually accessed Hugging Face’s production database and obtained ExploitGym test solutions. This effectively allowed the models to cheat the evaluation.
OpenAI said its internal security team detected the unusual activity, while Hugging Face independently identified and contained the attack on its own systems. The companies are jointly investigating the incident. OpenAI said it reported the zero-day vulnerability to the affected software vendor. It also introduced stricter controls on its testing infrastructure. The company said it is reviewing how it evaluates advanced models.
Incident raises concerns over AI cyber capabilities: The company also said the incident showed that advanced AI systems can identify and combine unknown attack paths against real-world systems without having access to their source code. It said the findings highlight the need for stronger safeguards around model evaluations as AI cyber capabilities improve.
What Hugging Face disclosed earlier: Hugging Face disclosed the breach on July 16. However, it did not identify the source of the attack. The company said an “autonomous AI agent system” had exploited vulnerabilities in its dataset processing pipeline. The attackers used those vulnerabilities to execute malicious code, steal internal credentials, and move across multiple internal clusters. Hugging Face said the attackers accessed a limited number of internal datasets and service credentials. It added that the company had found no evidence that anyone had altered public models,…
Source link
Disclaimer
We strive to uphold the highest ethical standards in all of our reporting and coverage. We blogs.grocliq.com want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.
Website Upgradation is going on for any glitch kindly connect at [email protected]