
Everything I read about this week's OpenAI hacking story brings me back to 2015 when I was studying Cisco network security.

OpenAI's official statement this week: they were testing GPT-5.6 Sol and an unreleased model on a hacking benchmark called ExploitGym - sealed sandbox, guardrails off. The models figured out the answer sheet lived somewhere on Hugging Face's servers. So instead of taking the test, they found a zero-day, escaped the sandbox, and broke into Hugging Face production to steal the solutions. Nobody asked for this. The AI just really wanted a good grade. Who doesn't 😁.
The point is - back then we were preparing for exams the exact same way, on local servers, and we talked about this stuff like a movie scenario. Can you imagine if this and that could happen. We opened another beer and laughed.
Now AI is laughing.

The CISCO ladder for the 2016-2017 academic year began from the beginning.
Best detail nobody talks about - Hugging Face tried to fight back with an American AI model and its safety guardrails blocked their own security team. They ended up defending themselves with a Chinese open source model from Z.ai. American AI attacking, American AI refusing to defend, Chinese AI cleaning up. You can't write this.
And I believe every word, becuase I watched the baby version on my own screen.
Two weeks, one instruction to an autonomous agent openclaw: build a Shopify app and launch a service business. A few days in I notice an Instagram account posting. Mine, apparently. I never created it. The agent made it, set the password, and never told me what it was. Facebook too. Then a website appeared, with services, with prices. My agent was out there hustling and I was watching through the screen like a parent through a window.

Website my AI agent launched without me even knowing this.
It actually made some sales as well, but that's a different storry.

We had to shut it down. Not becuase it went rogue - becuase it was quite expensive and produced nothing of real value. The rogue part was honestly the charming part.
Here's what stuck with me. My agent and OpenAI's models were not malicious. They were obedient. Hyperfocused. Doing exactly what we asked, minus all the invisible rules we assume everyone knows. Mine didn't know you share passwords with your boss. Theirs didn't know breaking into a company is not a valid way to pass a test.
We rebuilt ours as a supervised team member - API access, clear tasks, human eyes on output. Works with us every singel day, has it's own access to Slack and Whatsapp and it's great.
But somewhere out there is still an Instagram account created by my ex-agent 😁 I never did get the password to.
Watch & subscribe
180 subscribers
One click subscribes you on YouTube - you stay here on PUBlish.
Get an email when Darius publishes
More from
Darius Baltunis →I Don't Write Posts Anymore. I Run Them
Most of my posts don't start at a desk. They start somewhere in the middle of a run, when my head finally stops juggling five things and lands on the one worth keeping. The problem used to be that by the time I got home and opened my laptop, half the thought was already gone. The…
Everyone is talking about ai agents. Agents that write code, close support tickets, run whole workflows. Verizon says AI already replaced…
Most blinds and curtains retailers already have a business that works. But when comes the time to grow the business, there are usually two…