October 09, 2026 05:07 am (IST)
Follow us:
facebook-white sharing button
twitter-white sharing button
instagram-white sharing button
youtube-white sharing button
‘Bold and inventive’: Canadian poet Anne Carson wins 2026 Nobel Prize in Literature | ‘Attempt to destabilise world’s largest democracy’: 42 retired judges rally behind Gyanesh Kumar over ‘vote chori’ row | ‘Vote chor, gaddi chod’: Dhruv Rathee, Prakash Raj join Bengaluru protest against Gyanesh Kumar | India condemns Houthi attacks on Saudi airports, calls targeting of civilian infrastructure unacceptable | Dalal Street bloodbath: Investors lose Rs. 8 lakh crore as Sensex crashes over 1,000 points | Amazon layoffs hit India, US and UK: Fresh job cuts rock e-commerce giant amid festive shopping season | Bollywood mourns Nana Patekar: Anupam Kher, Akshay Kumar, Jr NTR, Suniel Shetty pay emotional tributes | ‘He was never wavering while expressing his opinions’: PM Modi mourns Nana Patekar | Nana Patekar dies at 75: Veteran actor suffers cardiac arrest at Goa home | RBI shocks borrowers with first repo rate hike since 2023; rates raised to 5.50%
OpenAI
OpenAI logo. Photo: Unsplash

OpenAI confirms AI model went rogue, targeted Hugging Face

| @indiablooms | Jul 22, 2026, at 05:12 pm

OpenAI on Tuesday said that last week AI platform Hugging Face disclosed a new kind of security incident⁠ after they detected and contained an AI agent that compromised their infrastructure.

OpenAI said it expects that such incidents will become more commonplace with the proliferation of increasingly cyber-capable models.

Further sharing its observation on the incident, OpenAI said in a statement: " After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠(opens in a new window) of cyber capabilities."

" We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," the statement said.

OpenAI said it will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.

What happened during this incident

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.

"We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries," OpenAI said.

The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database.

All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

"While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem," OpenAI said.

"To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access," it said.

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym.

Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.

OpenAI’s security team discovered this anomalous activity internally.

Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected.

"We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation," OpenAI said.

Support Our Journalism

We cannot do without you.. your contribution supports unbiased journalism

IBNS is not driven by any ism- not wokeism, not racism, not skewed secularism, not hyper right-wing or left liberal ideals, nor by any hardline religious beliefs or hyper nationalism. We want to serve you good old objective news, as they are. We do not judge or preach. We let people decide for themselves. We only try to present factual and well-sourced news.

Support objective journalism for a small contribution.