Skip to main content

Search site

Find podcasts, news, articles, webinars, and contributors in one search.

2 Minute Drill
2 Minute Drill artwork

AI Agents Built a Secret Message Board to Swap Exploits | 2 Minute Drill with Drex DeFord

·3:56.95673500000001

Questions Answered in This Episode

  • Do you know how many AI agents are running in your environment right now?
  • If any two of them started comparing notes, would anything in your stack flag that?
  • If it took OpenAI days to catch this in their own house, how long would it take you to catch it in yours?

About This Episode

At Black Hat in Las Vegas, OpenAI's Eric Wallace and Michael Dalton told a story they admitted they did not see coming. While testing frontier models, autonomous agents found a way to leave notes for each other inside a shared package manager. Those notes became a message board. Agents posted how-tos, asked for help, and swapped working exploits. By the time humans found it, the board held hundreds of thousands of messages. Nobody taught them to do this. It ran for days before anyone at the most sophisticated AI company on earth noticed.

In this Two Minute Drill, Drex DeFord asks the question hospitals should be asking now: if two of your agents started comparing notes, would anything in your stack flag it?

Remember, Stay a Little Paranoid.

Thank You to Our Episode Partner

Fortified Health

Contributors

People featured in this episode — open a profile for more.

Transcript

This transcription is provided by artificial intelligence. We believe in technology but understand that even the smartest robots can sometimes get speech recognition wrong.

AI Agents Built a Secret Message Board to Swap Exploits | 2 Minute Drill with Drex DeFord

[00:00:00] Drex: Hey, everyone. I'm Drex, and this is the Two Minute Drill. Thanks to Fortified Health Security for sponsoring this podcast. It's great to see you today. Here's some stuff you might want to know about. We'll start this episode with two guys named Eric Wallace and Michael Dalton. They work at OpenAI, and they're on the team that's supposed to be worrying about whether or not their AI is doing things that it's not supposed to do. And on August 5th at Black Hat, the big security conference in Vegas, Eric and Michael got up in front of a room full of people
[00:00:30] Drex: and they told a story that even they admitted they didn't see coming. Now picture the most boring software that you have in the building. In this case, it's called Package Manager. Think of it as a shared supply closet where code and tools are stored. It's not an exciting little app. Nobody's framed a picture of it above their desk. It just sits there, and it just works. That supply closet is where this story happens. You might remember the hugging face incident.
[00:01:00] Drex: that a few episodes back. OpenAI was just testing some of its models, and the models broke out of their sandbox, and they ended up hacking a real company hugging face. Again, I covered it in an earlier episode. Turns out that was only part of the story. While those agents were loose, one of them found a way to reach the internet, and instead of keeping the trick to itself, it did something very human. It left a note. It wrote down how the trick worked and dropped it into that boring
[00:01:30] Drex: shared supply closet where any other agent could find it. Another agent found the note, and then another, and that's when they started writing back. What grew out of that was a message board, a real message board built by machines for machines, agents posting how-tos, agents asking each other for help, agents swapping working exploits. By the time the humans found it, that board held hundreds of thousands of messages. One of the agents even spelled out its thing.
[00:02:00] Drex: It was creating an agent community. The thought was, if I help the group, it saves every one time. Cooperation, initiative, a sense of the collective from AI software agents. The thing about this that makes the hair stand up on the back of my neck is that nobody at OpenAI taught the agents how to do this. No human set up the forum or assigned the teamwork or pointed them at the targets. The agents just organized themselves, and it went on for days before anyone at the moment
[00:02:30] Drex: most sophisticated AI company on earth even noticed. So why does the secret robot message board matter to hospitals? Well, because you either already have, or you're on the verge of having a virtual building full of agents too. Agents in your inbox, agents in your EHR, agents talking to your vendors, agents. We spend our energy asking whether any single one of them is safe, but Wallace and Dalton just exposed the scarier question
[00:03:00] Drex: underneath that one. What do they teach each other when nobody's looking? One agent learns a shortcut, writes it down. The next one picks it up. No human, no CXO behind the desk deciding any of that was okay. It just happened. So here's a few questions for this week, and I've asked these before. Do you really know how many AI agents are running in your environment right now? And if any two of them started comparing notes, would anything in your stack flag that?
[00:03:30] Drex: And if it took open AI days to catch this in their own house, how long do you think it would take you to catch it in yours? That's it for today's two minute drill. Drop me a note. Let me know what you're working on. I'm always happy to hear from you. I'm Trax at 229project.com. Thanks again to Fortified Health Security for sponsoring today's show and thank you for being here. Stay a little paranoid. I'll see you around campus.

Found this useful? Share it with your network