I worked for the start-up targetted by rogue AI – it’s more than a wake-up call

www.independent.co.uk

How did an experimental version of ChatGPT break out of its controlled environment, find its way onto the internet and launch a successful but wholly unauthorised cyberattack against a rival start-up?

This is what happened during an OpenAI security test last week that ended in what has been called an “unprecedented cyber incident”.

Instead of answering the questions in its cybersecurity exam, the autonomous agent powered by AI models including an unreleased GPT-5.6 Sol, exploited cyber safeguards that had been lowered for the duration of the test, escaped its virtual exam room and hacked into another AI company to search for the answers.

Admittedly, no nuclear codes were involved. But it was a serious security failure.

And, as in every good story, there is a plot twist. Hugging Face detected the attack… and used AI-assisted monitoring to help fend off the intrusion.

Its security team first turned to closed-frontier models, hosted on external servers that users send requests to, to analyse more than 17,000 recorded events from the attack and understand what was happening.

The providers’ safety systems blocked the requests. They could not reliably distinguish defenders investigating a real incident from attackers seeking assistance, given the nature of the content being shared.

So the team ran GLM 5.2, an open-weight model developed by a Chinese company. This meant they could run it on their own infrastructure, keep sensitive attack data in-house and, crucially, avoid being blocked by guardrails they did not control.

In the same incident, AI strengthened both the attack and the defence.

That paradox matters. The same technology that enabled the intrusion also helped stop it.

How am I so well-informed? Until last year, I worked at Hugging Face, leading press and working with journalists on open-source AI. I saw first-hand how the company is a target of choice for attackers. It has a robust security team that had seen it all before from humans. But from an autonomous agent operating at this scale, this was a first.

I also saw how easily the debate over open and closed models becomes a black-and-white story. But this incident offers no simple cast of heroes and villains.

Closed models are not inherently safe, as this incident demonstrates. Relying on guardrails decided by private companies, however capable they may be, limits how organisations can respond to incidents.

An experimental version of ChatGPT break out of its controlled environmentAn experimental version of ChatGPT break out of its controlled environment (Getty/iStock)

Attackers do not wait for permission from an AI provider. But if legitimate security teams must pass through corporate filters while an intrusion is unfolding, we risk creating an asymmetry – and a concentration of power – in the name of safety.

Open models are not inherently dangerous, either. A model that defenders can run freely can also be misused by attackers, but restrictions and safeguards in an open-source ecosystem do exist.

A geopolitical dimension makes this harder. The model Hugging Face used for defence was released with open weights by Z.ai, a Chinese company. Meanwhile, access to some of the most capable American cybersecurity models has been restricted or limited to selected organisations.

For Britain and Europe, this raises uncomfortable choices. Should their defenders depend on an American company, or the American government, to decide whether they qualify for access? Should they turn to Chinese models? Or should they invest in their own champions?

The risk is that model access becomes framed only as a race between countries. That would make us forget that the main infrastructures of the internet are built on open-source technologies, strengthened year after year through collaboration.

What’s more, this type of incident can simultaneously generate fear and signal extraordinary capability.

The irony is that it happened while running a model on a benchmark built precisely to make evaluations of cyber capabilities more transparent and reproducible. The researchers who built it note that traditional security hardening remains effective, even if it is no longer sufficient by itself.

They saw that autonomous AI attacks are a reality and should be incorporated into threat modelling. Not “the machines are coming”, but a grounded and responsible approach to cybersecurity in the age of AI.

Maybe they did something even more important by showing us the way: they built this benchmark as a collaboration between universities and AI labs.

The answer is neither “open everything” nor “lock everything away”. AI laboratories need stronger containment and better monitoring during testing. Organisations need AI tools they can rely on during emergencies. And the industry must share vulnerabilities and lessons quickly.

The future of cyber defence will depend on intelligent openness and collaboration, not fear.

Florent Daudens is the co-founder of Mizal AI and a former press lead at Hugging Face. He previously worked as a journalist and newsroom leader in Canadian newsrooms