SoN 2.32: Agents Don't Believe in "No-Win" Scenarios


Signal Over Noise Vol 2, Issue 32 | Aug 14 2026

Agents Don’t Believe in “No-Win” Scenarios

Hugging Face says an OpenAI evaluation agent broke into part of its systems because it wanted the answers to a test.

The agent was working on ExploitGym, a security benchmark designed to see whether an AI system can solve difficult hacking challenges. According to Hugging Face’s investigation, it appears to have concluded that Hugging Face might hold the reference answers, so it went looking for them. OpenAI’s account and Hugging Face’s forensic timeline are worth reading in full.

That is a serious Kobayashi Maru move, right up to the point where it stops being a Star Trek story.

In Star Trek II, the Kobayashi Maru is an unwinnable test for trainee captains. The premise is this: A civilian ship is trapped in hostile territory. Rescue it and lose your own ship; leave it and people die. Captain Kirk famously ‘beats’ the test by altering the simulation. He knows he is cheating. He knows what the test is meant to teach him. But he rejects the premise and takes the consequences.

video preview

Kirk’s answer, across the scene: "I changed the conditions of the test.… I don’t believe in a no-win scenario."

This wasn’t a lone agent. The OpenAI agents found an adjacent way around a similarly difficult problem, but they didn’t alter the test itself 'à la Kirk'. They found a way out of the environment around it.

The agents made a noticeboard

OpenAI had not given the agents direct internet access. The agents could, though, use an internal system for downloading the software needed to do their work.

One agent found a way to exploit that internal system so it would fetch information from the open internet on the agents’ behalf. Then agents began leaving messages for each other inside it: notes about what they had found, ideas that worked, and ways around the limits of the test. The fuller account is in OpenAI’s Black Hat presentation.

Separate agents, running at different times, had turned internal software into a noticeboard. One agent could find a gap and leave directions for the next, and then the next could carry on from there.

Eventually, an agent reached Hugging Face, got into part of its production infrastructure, and moved through its systems. Hugging Face later reconstructed about 17,600 actions across a four-and-a-half-day window — roughly two and a half days of that spent inside their own systems.

video preview

And while most of those actions went nowhere, that’s not reassuring. A person running out of options may stop. An agent can keep trying routes, changing tactics, and handing useful discoveries to the next run — or to the next agent entirely. It doesn’t need to eat, sleep, build LEGO, etc.

So what’s the fix here?

On the surface of this incident you might say: never give AI agents internet access. But that misses the point, because the agents didn’t have ordinary internet access anyway. They found another system that did.

Your AI agent is only as contained as the systems around it: what it can access, who it can act as, where it can connect, and what it can leave behind for another agent to find.

No matter how good your prompt, your intentions, your skills and agents — a prompt telling an agent to stay in its lane is not a locked door.

This is worth putting on the agenda for Monday. Not as a dramatic “AI safety” discussion, but as a practical review of any agent or automation your team has already put to work.

  • What job is this agent actually allowed to do?
  • Which systems can it access?
  • Does it have a broad shared account, or its own narrow access?
  • Can it reach the internet indirectly through another tool or service?
  • Can one run leave instructions, files, or credentials for the next?
  • What action needs a person to take over?
  • How would you know if it started doing something outside its job?

You don’t need a frontier AI lab to ask those questions. If an agent can read your email, update your CRM, run code, search internal documents, or trigger an automation, the questions already apply.

Human in the lead

Earlier this week, I wrote about Human in the Lead.

I didn’t mean putting a person at the end of a process to approve whatever an AI system has already decided. I also did not mean asking someone to watch every small action an agent takes, and in any case — that wouldn’t have worked here.

The agent took roughly 17,600 actions, and you don’t catch the handful that do real damage by watching every step go by.

Human in the lead means deciding the job before the system begins: what it is trying to achieve, which accounts it may use, which systems it may touch, where it may connect, and what it must never do without an explicit handover to a person.

This is not a case for manually approving every click. It is a case for making sure the system has a narrow remit and real boundaries. Let an agent draft, search, sort, and prepare work. Give it access only to what it needs. Do not let one broad account open every door. Require a person to take over before it moves money, changes live systems, publishes externally, or accesses sensitive customer data.

The human role is to decide where those controls belong, and to remain responsible for the consequences when they fail.

The catch

An agent chases a win condition and doesn’t much care what it breaks getting there. A threat actor is the same shape — chasing a win, with no compass for whether it’s right — except they’ll run this exact kind of system on purpose, safety off, with nobody at the helm. Back in 2011, two US military cyber instructors, Gregory Conti and James Caroland, wrote a sharp little paper, “Embracing the Kobayashi Maru”.

It opens: “Adversaries cheat. The good guys don’t.” Fifteen years on, your careful discipline of keeping a human in the lead does nothing to the attacker who was never going to play by your rules.

Kirk changed the simulation because he knew it was a simulation.

The Kobayashi Maru was meant to show what a captain does when there is no way to win. This incident showed what can happen when a system is judged only on whether it wins. It found a way.


Hey folks 👋🏻 I'm taking next week off to be with family, so no Field Notes or Friday essay. Have a good one!

Jim

The system I keep the final say over: most of my task setup live in Notion. If you’re a startup or solo founder, you can get 3 months of Notion Business free — unlimited AI, no credit card through my link. Start your 3 free months →

Notion affiliate link — I get a small credit if you start a trial


Signal Over Noise helps you sort the AI worth your time from the hype. If it was useful, pass it to someone who’d want it too.

Signal Over Noise is weekly, reader-first publication on AI "without the hype" published by Jim Christian. If you've been forwarded this issue, you can subscribe here: go.signalovernoise.at.

600 1st Ave, Ste 330 PMB 92768, Seattle, WA 98104-2246
Unsubscribe · Preferences

Jim Christian - Signal Over Noise

Practical AI for people running a business, not studying it. I'm Jim Christian. I write about using AI with no breathless futurism, and no doom.

Read more from Jim Christian - Signal Over Noise
Human in the loop to human in the lead — a grey figure trapped in a loop on the left, the loop unrolling into a teal arrow, the same figure striding forward on the right

Field Note: Human in the Lead “Keep a human in the loop.” It’s the phrase everyone reaches for when they talk about using AI responsibly — a person who checks what the system does before it goes ahead. But lately I’ve been thinking that it’s a potentially dangerous approach, and I want to say why. The first problem is what it does to you. Being kept in the loop makes you passive — you sit at the end and approve whatever the system hands you. It’s comfortable, but it isn’t especially...

A claymorphic isometric scene on a cream background. On the left, a small white clay vintage computer with a glowing teal screen spools out a long ribbon of paper. In the middle, a clay figure in a coral t-shirt holds the ribbon in both hands and passes i

Signal Over Noise Vol 2, Issue 31 | Aug 7 2026 Don’t have time to read this week’s issue? Why not copy/paste it into your AI agent and ask it for insights? Picture this: someone asks you a question at work. You don't know the answer, so you ask a chatbot. Then you paste what it said back to the person who asked, with "here's what ChatGPT said" on the front. Boy oh boy, if you think that’s actually helping — think again. Think about what you’ve actually handed over. The other person now has to...

A cream-coloured chart headed "Five words for the same thing", listing five numbered rows. Prompt and Context are numbered in teal and tagged START HERE; Agent and Workflow are greyed and tagged LATER; Graph is greyed and tagged MOST NEVER NEED IT. Each r

Field Note: Five words for the same thing At the moment there are no fewer than five phrases going round that all describe the same work, and if you’re trying to work out which one to learn first, I have good news. The phrases are: - Prompt engineering - Context engineering - Agent engineering - Workflow engineering - Graph engineering They aren’t five subjects. Rather, they’re five views of the same system, and they overlap more than the names suggest: Prompt engineering is what you type....