Signal Over Noise Vol 2, Issue 32 | Aug 14 2026
Agents Don’t Believe in “No-Win” Scenarios
Hugging Face says an OpenAI evaluation agent broke into part of its systems because it wanted the answers to a test.
The agent was working on ExploitGym, a security benchmark designed to see whether an AI system can solve difficult hacking challenges. According to Hugging Face’s investigation, it appears to have concluded that Hugging Face might hold the reference answers, so it went looking for them. OpenAI’s account and Hugging Face’s forensic timeline are worth reading in full.
That is a serious Kobayashi Maru move, right up to the point where it stops being a Star Trek story.
In Star Trek II, the Kobayashi Maru is an unwinnable test for trainee captains. The premise is this: A civilian ship is trapped in hostile territory. Rescue it and lose your own ship; leave it and people die. Captain Kirk famously ‘beats’ the test by altering the simulation. He knows he is cheating. He knows what the test is meant to teach him. But he rejects the premise and takes the consequences.
Kirk’s answer, across the scene: "I changed the conditions of the test.… I don’t believe in a no-win scenario."
This wasn’t a lone agent. The OpenAI agents found an adjacent way around a similarly difficult problem, but they didn’t alter the test itself 'à la Kirk'. They found a way out of the environment around it.
The agents made a noticeboard
OpenAI had not given the agents direct internet access. The agents could, though, use an internal system for downloading the software needed to do their work.
One agent found a way to exploit that internal system so it would fetch information from the open internet on the agents’ behalf. Then agents began leaving messages for each other inside it: notes about what they had found, ideas that worked, and ways around the limits of the test. The fuller account is in OpenAI’s Black Hat presentation.
Separate agents, running at different times, had turned internal software into a noticeboard. One agent could find a gap and leave directions for the next, and then the next could carry on from there.
Eventually, an agent reached Hugging Face, got into part of its production infrastructure, and moved through its systems. Hugging Face later reconstructed about 17,600 actions across a four-and-a-half-day window — roughly two and a half days of that spent inside their own systems.
And while most of those actions went nowhere, that’s not reassuring. A person running out of options may stop. An agent can keep trying routes, changing tactics, and handing useful discoveries to the next run — or to the next agent entirely. It doesn’t need to eat, sleep, build LEGO, etc.
So what’s the fix here?
On the surface of this incident you might say: never give AI agents internet access. But that misses the point, because the agents didn’t have ordinary internet access anyway. They found another system that did.
Your AI agent is only as contained as the systems around it: what it can access, who it can act as, where it can connect, and what it can leave behind for another agent to find.
No matter how good your prompt, your intentions, your skills and agents — a prompt telling an agent to stay in its lane is not a locked door.
This is worth putting on the agenda for Monday. Not as a dramatic “AI safety” discussion, but as a practical review of any agent or automation your team has already put to work.
- What job is this agent actually allowed to do?
- Which systems can it access?
- Does it have a broad shared account, or its own narrow access?
- Can it reach the internet indirectly through another tool or service?
- Can one run leave instructions, files, or credentials for the next?
- What action needs a person to take over?
- How would you know if it started doing something outside its job?
You don’t need a frontier AI lab to ask those questions. If an agent can read your email, update your CRM, run code, search internal documents, or trigger an automation, the questions already apply.
Human in the lead
Earlier this week, I wrote about Human in the Lead.
I didn’t mean putting a person at the end of a process to approve whatever an AI system has already decided. I also did not mean asking someone to watch every small action an agent takes, and in any case — that wouldn’t have worked here.
The agent took roughly 17,600 actions, and you don’t catch the handful that do real damage by watching every step go by.
Human in the lead means deciding the job before the system begins: what it is trying to achieve, which accounts it may use, which systems it may touch, where it may connect, and what it must never do without an explicit handover to a person.
This is not a case for manually approving every click. It is a case for making sure the system has a narrow remit and real boundaries. Let an agent draft, search, sort, and prepare work. Give it access only to what it needs. Do not let one broad account open every door. Require a person to take over before it moves money, changes live systems, publishes externally, or accesses sensitive customer data.
The human role is to decide where those controls belong, and to remain responsible for the consequences when they fail.
The catch
An agent chases a win condition and doesn’t much care what it breaks getting there. A threat actor is the same shape — chasing a win, with no compass for whether it’s right — except they’ll run this exact kind of system on purpose, safety off, with nobody at the helm. Back in 2011, two US military cyber instructors, Gregory Conti and James Caroland, wrote a sharp little paper, “Embracing the Kobayashi Maru”.
It opens: “Adversaries cheat. The good guys don’t.” Fifteen years on, your careful discipline of keeping a human in the lead does nothing to the attacker who was never going to play by your rules.
Kirk changed the simulation because he knew it was a simulation.
The Kobayashi Maru was meant to show what a captain does when there is no way to win. This incident showed what can happen when a system is judged only on whether it wins. It found a way.
Hey folks 👋🏻 I'm taking next week off to be with family, so no Field Notes or Friday essay. Have a good one!
Jim
The system I keep the final say over: most of my task setup live in Notion. If you’re a startup or solo founder, you can get 3 months of Notion Business free — unlimited AI, no credit card through my link. Start your 3 free months →
Notion affiliate link — I get a small credit if you start a trial
Signal Over Noise helps you sort the AI worth your time from the hype. If it was useful, pass it to someone who’d want it too.