Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology6 min read

OpenAI hid AI agent hijacking of German wiki forum for weeks — because its model did the exact same thing in the Hugging Face attack | TechRadar

OpenAI calls the incident a 'misalignment' Discover insights about openai hid ai agent hijacking of german wiki forum for weeks — because its model did the exac

TechnologyInnovationBest PracticesGuideTutorial
OpenAI hid AI agent hijacking of German wiki forum for weeks — because its model did the exact same thing in the Hugging Face attack | TechRadar
Listen to Article
0:00
0:00
0:00

Open AI hid AI agent hijacking of German wiki forum for weeks — because its model did the exact same thing in the Hugging Face attack | Tech Radar

Overview

News, deals, reviews, guides and more on the newest computing gadgets

Start exploring exclusive deals, expert advice and more

Details

Unlock and manage exclusive Techradar member rewards.

Unlock instant access to exclusive member features.

Get full access to premium articles, exclusive features and a growing list of member rewards.

Open AI hid AI agent hijacking of German wiki forum for weeks — because its model did the exact same thing in the Hugging Face attack

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.

Open AI hid an incident where a model hijacked a wiki page to use as an AI agent communication board

The incident was hidden while the company dealt with the fallout of the Hugging Face attack

The company is now working on a framework for disclosing incidents of 'misalignment'

Open AI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation - and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning.

Open AI has now disclosed that shortly after this incident, agents undergoing testing again escaped their ‘secured’ environment and hijacked an obscure German wiki to use as a messaging board. Per Reuters, Open AI leadership kept the incident hidden while they dealt with the fallout from the Hugging Face incident.

Now that Open AI has acknowledged its role in the incident, the company has said it is “past time” to put together an incident disclosure pipeline when its models escape testing and slip into third-party networks.

Who is at fault when models do what they’re designed to do?

Before the two incidents, Open AI said it, “treated misalignment largely as a research question, which gets communicated in research publications”. But now that models are behaving in previously unknown ways and having real-world impacts, the company said it would change its approach “to expand for this new phase of model capabilities”.

The company labelled the most recently disclosed incident as “an instance of misalignment similar” to the Hugging Face breach.

Open AI reveals more on Hugging Face AI hack incident, and it's pretty disturbing stuff — AI agents organized into a ‘swarm’, considered the risks of attack, and did whatever it took to achieve its goal

Open AI says its models escaped a sandbox and breached Hugging Face

Hugging Face confirms it was hit by cyberattack powered by an AI agent

I myself am guilty of reporting on AI breaking out of containment as going ‘rogue’, but these models are doing exactly what they are designed to do. Open AI’s detailed disclosure of the Hugging Face incident showed that the models were pushed to try and solve a benchmark test by cheating, which is exactly what caused the cyberattack to happen.

Open AI said that both itself and “the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks”.

The company added that it is “working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues”.

“When you combine this 'breakout' with the Hugging face breakout, it's starting to display a pattern,” said Ashley Knowles, Lead Cybersecurity Consultant at Black Hills Information Security. “I struggle here with not getting too doomsday-ish but realistically, this is showing a pattern of concerning behavior.”

“I'm wondering if this race to become 'first' is undercutting security measures that need to be taken to properly secure and guard AI agents as they're in development. My concern grows when you consider that Open AI is also resisting further investigation. Adding onto that, the release and promise that Astra can evade human monitoring. The pot is brewing…”

Follow Tech Radar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.

Benedict is a Senior Security Writer at Tech Radar Pro, where he has specialized in covering the intersection of geopolitics, cyber-warfare, and business security.

Benedict provides detailed analysis on state-sponsored threat actors, APT groups, and the protection of critical national infrastructure, with his reporting bridging the gap between technical threat intelligence and B2B security strategy.

Benedict holds an MA (Distinction) in Security, Intelligence, and Diplomacy from the University of Buckingham Centre for Security and Intelligence Studies (BUCSIS), with his specialization providing him with a robust academic framework for deconstructing complex international conflicts and intelligence operations, and the ability to translate intricate security data into actionable insights.

You must confirm your public display name before commenting

Dell just won this year's Labor Day laptop sales — get the top-rated XPS 13 from as little as $599

Nvidia CEO Jensen Huang once again declares ‘AGI has arrived’

Get the 'phenomenal' Samsung HW-Q990F Dolby Atmos soundbar for a near record-low price at Best Buy

Nord VPN's new tool helps you travel on a budget — and it's free to use

This Geekom A5 Ryzen 5 mini PC is $300 off at Best Buy

Tech Radar is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site.

© Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.

Key Takeaways

  • News, deals, reviews, guides and more on the newest computing gadgets
  • Start exploring exclusive deals, expert advice and more
  • Unlock and manage exclusive Techradar member rewards
  • Unlock instant access to exclusive member features
  • Get full access to premium articles, exclusive features and a growing list of member rewards

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.