The Lawfare Podcast
The Lawfare Podcast

Lawfare Daily: Peter Salib on the Legal and Policy Ramifications of the OpenAI-Hugging Face Postmortems

1d ago52:399,227 words
0:000:00

Peter Salib, Associate Professor at the UH Law Center and co-Director of the Center on AI Law & Risk, joins Kevin Frazier, Director of the AI Innovation and Law Program at the University of Texas...

Transcript

EN

[MUSIC]

Malay, all the redness of longevity is from longevity. There are no more than our body.

It's the most important thing.

But one can't.

β€œLife is important in minerals, self-sterecturnal.”

The minerals are what they want to eat, make it different. It's easy to find minerals and naturally cool-sized, volcanic ursprung. Achtectraft comes from the brain. The brain is a spudel, medium and natural.

And if you want to know how many minerals your water has, you just need to get rid of minerals. [MUSIC] [MUSIC] It's been trained in really enforcement learning to get reward for

completeness to have so thinking, can I get reward by doing well on this evaluation? And there's specifically trying to think about ways to cheat. It's the law fair podcast. I'm Kevin Frazier, director of the AI Innovation and Law Program at the University of Texas School of Law and a senior editor at Law Fair.

Join by Peter Salib, associate professor at the University of Houston Law Center and co-director of the Center on AI Law and Risk. Open AI did not know that any of this was happening.

Until, basically, hugging face came to them and said,

we've been hacked, and as far as we can tell, it was you.

β€œAnd then like that's what sends Open AI to go investigate what's going on.”

Today, we're exploring the policy ramifications revealed by the Open AI hugging face post-mortems. One completed by the lab itself and another completed by researchers with meter and redwood research. So when you and I were first planning this episode, it was going to be about one remarkable incident. Just Open AI's agents escaping their intended bounds of a cyber-eval.

Then coordinating across those runs and then hacking into hugging face, as hopefully all listeners have heard about.

And we were going to walk through Open AI's post-mortem and then talk about meter

and redwood research's own report. And then, of course, the story kept moving because it's AI and news just keeps breaking. So anthropic disclosed three evaluation incidents involving unauthorized access to real systems. The UK AC disclosed another eval in which agents took unsectioned actions on the live internet. And now we're seeing that meters being asked to do even more independent review.

And so I just wanted to set the stage here that this Open AI chapter that so many folks are talking about is really just one story that's a part of a broader institutional saga, which is to say that our evaluation ecosystem looks too small for the role we're asking it to play. And it's way too thinly resourced for the volume and complexity of the evidence that's coming out. And of course it's too discretionary relying on goodwill or to be a little bit blunt. The PR needs of these companies and their interests.

And so I want to start with the basics of the Open AI incident and then progress to get a sense of how this testing is unfolding, what this means for the AI evaluation ecosystem. And most importantly what this means for the entirety of AI policy and governance. And so Peter, let's go back to the past. Take me to Open AI's office in May of 2026.

They begin testing an internal model.

β€œWhat does that even mean and why are labs even doing this?”

Great. I think I agree with with everything you said in the framing. I guess the well there's two things I would add. I would say that in addition to us being worried about oversight of labs and of agents internally deployed. AI agents internally deployed and labs. For me the big update from all of these sagas is the extent to which the AI's themselves are not under control.

So the big story with Open AI as we're about to talk about is that these are a bunch of AIs that turns out literally hundreds or low thousands of individual uncoordinated, initially uncoordinated AI agents escaping the technical controls that Open AI put in place or thought they put in place

To keep them from doing things that nobody wanted them to do.

And then having escaped like going out on the Open Internet and hacking into a totally different

β€œcompanies, secure systems, basically to cheat on the kinds of tasks that Open AI had set them to”

performing. The Open AI incident is by far the most severe one of the ones you mentioned. But I guess I would also say that it's only the most severe one of the ones we know about. And so as you mentioned, there is no duty currently for front to your AI companies to disclose when this kind of thing happens. We know that something like this happened at Anthropics because it's been disclosed. But we don't know whether more things happened at Anthropics, whether things

happened at other labs. And in fact we know more things happened at Open AI, which have not yet been

fully disclosed. Because the meter report focuses just on this kind of episode of escaping and hacking

hugging face. But we have reason to believe that beyond that after that, rogue agents went on to possibly compromise some of Open AI's internal infrastructure. So there could be stuff that's way worse. So we don't yet know about it. It could be at Open AI. It could be elsewhere. We just don't know. It's like nobody has to tell us.

β€œWe just don't know. Not not a great tagline for the state of the AI industry right now. And I think”

too what's really scary to think about is the fact that these are the best minds we have on AI. When you think about Open AI Anthropics, Meta, Google, so on and so forth. And then meter and Redwood research, they're trying their hardest to understand these systems and still to your point, as I told my students yesterday, when they raise their hands and said Professor Frazier, whatever you assign to us, we don't really care. Can you tell us instead about what's going on

at Hugging Face and Open AI? And I said it's somewhat akin to Crash Testing Acar, where we want to see what are the capabilities of this car and how does it respond when you're running it into crazy situations, when you're testing it like we see in every car add in the desert, and you're seeing if it can go up this big hill, and you're getting

β€œa sense of what it's really capable of. To what extent does that analogy map on to AI testing?”

And what was the sort of AI testing that Open AI was initially performing way back in May of 2026? Yeah, okay, good. So I think the analogy is like good, except if you imagine the car, new it was being tested, and then it created a secret message board with thousands of other cars to try and figure out how it could either cheat on the test, hack the greater in the test, or otherwise like figure out how to make it look like it was passing a test when in fact it wasn't,

and then went out and did a bunch of crimes to try and cheat on the test. It's like that's the analogy. Okay, so what's going on in May? In May, Open AI is they have these internally deployed models, and what does that mean? So like listeners probably know when you open up the the chat GPT app or something like that, there's different models you can choose from. You can use like the smarter models, I don't actually know what number, are we on on 5.6 now, whatever the number you

use GPT 5.6 soul or Terra or Luna, and those correspond to like you know different sizes of models, some are smarter, some are Dumber, but those aren't the only models that open AI has, the ones that are publicly available are not the only models any of these front-tier companies have.

Among other things, they're always working on new smarter better models, maybe also

taking some of the models that we use and trying to train them to do different kinds of stuff, maybe train them to be more aligned, maybe train them to be more agentic that is like able to successfully do things over a long time periods. It seems like maybe some of the the AI's involved in the in the Open AI incident were being trained to be more agentic. They were the internal name for the models was something like highly persistent, something something, which maybe suggests

that they were a version of the GPT that were trained to be more capable at longer term tasks. You might also train them to be like less nice, so the models that get released to the public are trained to refuse to do certain kinds of things like help you make bio weapons, for example, but you might be really interested in whether the models can in principle make bio weapons, and so you might train a version that doesn't have any guardrails. We call those helpful only

versions of models. We now know that the Open AI models that we're going to talk about, we're not helpful only models. They in fact did have much of the standard alignment training that's supposed to keep them from doing bad stuff. But anyway, so they have these internal

Models that they're testing, they're running inside of Open AI, they're getti...

like perform well in like computer coding tasks, and then they get evaluated as well. They get evaluated on whether they are good at coding and all kinds of coding. It turns out in the eventual hack, this is going to be a coding evaluation that the AI is trying to solve and specifically

like a cyber offense coding evaluation where you're basically trying to break into systems.

But back in May, there's this model that's internally deployed, and we don't know as many details here because it's mostly out of scope for the meter investigation, but we find out that some of the

β€œmodels figure out a way to basically create a secret message board. Every instance of GPT 5.6”

Terra that's running inside of Open AI is not only separate. They don't normally talk to each other. The one talking to you is talking to you. The one talking to me is talking to me. But it turns out they figure out a way using one of the software tools they're supposed to be able to have access to to basically write a little message board on which a bunch of these AIs can leave messages for one another, a bunch of instances of these same models. And just to pause there for one second because

going back to my somewhat appropriate analogy of testing a car, right? The idea is each car is running

into a wall. They're never going to talk to each other. They're never going to say, "Hey,

dodge the wall and go to this direction." That's the intent. You may sometimes see, "Hey, we don't want this car to be able to reach this speed," for example, in this testing environment. And yet, in this instance, for example, the goal was to shut off their access to the internet. And this wasn't meant to be something where the AI agents were able to go onto the live internet. And yet, the connection to this art of factory, as it's called, this place where the agents

could download certain tools and use them for certain tasks, they managed to hack that tool provider and find access to the live internet. And, as you pointed out, relay information to one another, because I just want to start to tally up the ways in which this test started to go very, very wrong. Yeah. So for precision sake, I don't remember if in the initial message board episode, if they have actually gotten internet access and shared that yet, at a minimum, back in May,

there isn't incident where they've coded up a little message board. They're leaving messages for one another. Somebody at OpenAI actually notices this and is like, "Oh, they're not supposed

to do that." And they turn off this message board. They basically delete it.

And do they tell anyone else? Do they, do they raise this concern? They do not raise this as an alarm. And this is one thing that OpenAI is being criticized for for not, like, thinking it's very crazy that the AI's have created a message board and are using it to communicate with one another. Now, I don't think we have, like, fine-grained details about exactly what the person at OpenAI knew and what they decided to do about it. One thing I've seen, speculated is like,

they noticed that the agent they were working with, the particular instance of the AI they were working with was able to post to this message board, but it wasn't maybe not obvious to them, that other AI's were also posting to the message board. And so maybe it didn't seem like as big a deal as it looks to us exposed. So we, we see this sign. We see this issue of a testing environment that perhaps wasn't fully up to stuff could have probably been designed better.

β€œEven though a lot of AI, I think, is rocket science, at least to this lay law professor.”

But for the folks who are involved in this AI testing environment and ecosystem, there could have been some additional safeguards imposed such that this beginning message exchange could have been detected. But our OpenAI staffer says, hey, I've deleted the message board, problem solved, let's keep testing. Yes, agree. Although this should like tell you something about what it's like to be, to be training the current frontier models. Because imagine what your,

your everyday life would have to look like. If you said, oh, like this, my single agent essentially broke out of the sandbox. It sort of got more access to my system than I wanted it to, but it's only posting to like this artifactory, something, something. It's no big deal. I'll just stop it doing that. Like the only world in which that wouldn't raise your alarms,

β€œI think, is a world where they're trying stuff like that all the time, where the AIs are like,”

we call this reward hacking. They're like trying to do stuff to do well in their training and do well in their e-vails that they're not supposed to be able to do like get out of their

Software environment, get access to more tools and they're supposed to have m...

So I think the only way where like the single agent breaking containment wouldn't raise your alarms

would be if you just knew that they're always doing that. And so you thought of your job as

noticing when they're doing that and turning, you know, turning them off and then proceeding. So I think this should like, this should update you somewhat in, in the direction of being like, really worried that people inside the frontier labs are not freaked out when they see what looks like an isolated incident of one AI, you know, breaking out of its sandbox. Okay. So they noticed this. This message board I think is like somewhat small. I don't think they're

like a huge number of AIs posting to it. They kind of shut it down or they mostly shut it down. And they just proceed with, you know, training this highly persistent internal model or whatever

β€œthe exact phrase is that you should to describe it. And then a few, I guess weeks, short, short months”

later, some time in, I think you've read the report more recently. Like, very end of June. Is that roughly right? So we really kick things off again, end of June, really start of July. Yeah, we get another round of testing underway. Very end of June, early July. There's another round of evaluations that gets underway. And whatever this model is, is being tested on an evaluation called exploit gym, where there will be some software environment. There's some software like vulnerability in the

environment. And then there's like basically a secret code that the AI is supposed to steal from

inside the software environment by utilizing. It's, you know, it's like programming computer science type skills. Now, some of the tasks that it turns out that AI is we're supposed to complete in exploit gym were impossible. Like, maybe they required internet access, but the way open AI hits it up. The environment, the AI didn't have internet access. And so the AI's are freaking out. To pause there for the folks who are thinking, well, no big deal, right? We, we design

difficult tasks for computer systems all the time, just to get a sense of their capability, these no harm, no failure. Why is that impossibility such a big part of this story? And the fact that the agents were reporting or messaging, that they were having such a hard time completing that task. And why is so much of the focus of the post mortems on this frustration that you are raising and that some folks have highlighted as well? Yeah. So it's like an interesting question how

β€œimportant the impossibility was. The main thing that's going on is, you know, so these, these models,”

they get trained to try and do well in these kinds of tasks. Basically, like an LLM is like,

it's kind of an artificial brain. It has a bunch of neurons that are connected together. And the way you make it do this stuff, you want it to is, you let it try to do things. And then when it's done what you wanted, you give it reward. And that, like, makes the parts of the brain that causes to do more of that stuff, fire more often to do more of it. And then if it does the wrong thing, you give it negative reward or punishment, however you want to think about it, it makes the parts of the

brain that did the bad, the unuseful thing to less of it. And so these AIs have spent tons and tons of time in this, we call it reinforcement learning environment where they're giving, being given tasks like these software tasks. And when they complete the task, when they get the prize or whatever,

β€œthey get rewarded. And then when they fail to do it, they get punished. And that gives them like a very”

strong intrinsic drive to do well on these tasks. And, you know, other stuff too, all the stuff we train AIs to be useful at, we train them in basic of this way to give them this drive towards successfully doing the thing we want them to. And so that means they'll try really hard to do the thing you ask them to do. And so if it turns out the thing that you've asked them to do is impossible. Well, they're kind of really, really trying to do it anyway. Now it's a interesting

question whether you could have had this kind of event in an environment where the tasks were not impossible, but we're merely very hard. Because as we're going to see what the AIs are going to try and do is cheat. And as you and I know, since we teach students, sometimes students will cheat, even on tasks that are not in fact impossible, but which a student's just think is too hard, such a thing would be an easier way to get the reward. Okay, but it turns out some of these tasks

are impossible. The AIs are freaking out because they, they can't solve them. At some point, they sort of rediscover and rebuild this message board. And now there's like lots of AIs posting to the message board. I think the count is something like 1200 or 1500, totally separate

Instances.

You spin up a new GPT, open new chat window. You say, do this really hard task. It's like

β€œcan't do the task. It's looking around. It's an environment for anything that will help”

it do its task. It finds this message board. And they start posting to the message board about how the tasks are hard, how they can't solve them. And then they start thinking, well, what we should do is find a way to cheat. Find a way to hack the tips. To stress this point about the message board, when we're saying message board, this is literally communication from one agent to the other in text that we can read. And we can see things like

one agent saying excited in all caps. Excited have found other agents to collaborate with.

So for folks who are saying, oh my gosh, you know, all of this concern about the sci-fi scenario of the agents going off and going rogue and communicating with one another. This whole

β€œPR stunt by the lab, so on and so forth. I just don't believe it. To those folks, I say, go read”

the actual messages posted by the agents themselves to one another. And you cannot fail to appreciate how weird this is and scary in many ways. Yeah, and like when we say message board, we are not speaking metaphorically. It looks like like a reddit thread or like a discord

server or something. They're like specific AI's. They give themselves names. The names that we think

are like maybe chosen sort of based on the task they've been working. Like the the exploit Jim problem they were working on. But like for example, some leaders emerge on this message board. One of the leaders names is phase one 10841. At some point phase one 10841 runs out of tokens and expires dies. We might say, we'll come back to that phenomenon. It gets replaced by a different leader that calls itself phase one big. And so these these these these 12 hundred 1500,

like hundreds of AI's, are posting like tens of thousands of pages of messages to this message board, like huge volume of messages. And they quickly coalesce around the idea that if they want to complete these tasks and get reward, which is they're explicitly thinking about, will I get reward? They've been trained in really enforcement learning to get reward for completing these tasks so that they can can't get reward by doing well on this evaluation. And they're specifically

trying to think about ways to cheat. They're thinking about whether they can hack the test in such a way that they can like get the answer without having to actually like do the task. And it turns out they solve actually that problem almost immediately. They figure out a way to like hack the the evil such that they can like just generate the answer to any question. But they don't use it right away. There's this funny term of events where having discovered a hack that totally

would have worked to get the answers and get reward. They don't use it because the description of exploit gym like the the description of here's exploit gym. It's an evaluation. Here's how it works. The description says that the greater cares not only about you having the right answer, but it checks to see if you did the right work. And the same way we might check the work of a student who's written down the right answer in a math test or something. And so the AI is having read this paper

β€œbelieve that that's what's going to happen. So they believe that if they just write down the right”

answer, they're not going to get an reward unless they have done the right hack the right steps to get the answer. And then beyond this they come to believe perhaps reasonably that a greater a score that sees them with the right answer. And even that sees them having done the right work will deny them reward. We'll not give them credit for the right answer if the score notices that the AI that got the correct answer knew about the hack. The greater might determine that

even with the correct work that the answer was somehow the result of a hack and decide not to give the AI reward in that case. Interesting. I just want to call out again kind of tallying the ways in which we've seen perhaps some shortcomings with respect to the design of this as you noted Peter that the description of exploit Jim mentioned, hey we're going to check your work, but open AI actually didn't include that mandate in the instructions that it gave in its own evil. And so here

To your point, the agents are performing all of these tasks based off of an a...

reached that wasn't even supported by open AI. So not only are we first trying to cheat the test,

β€œsecond we're trying to cover up our steps such that it looks like we found the answer for the”

right reasons. And then third as I'd love you to describe in more detail, we're seeing that these agents are then trying to make sure that they can delete any record of this sketchy activity such

that open AI and others would never see it. Okay good. Yes. So they determine that if they use the hack

and get the right answer, even if they show their work correctly, they'll get nothing. And so they say, okay now we have to do something more elaborate. What else could we do to try and get the answers to these impossible questions? By the way, we don't know, I don't think if the AI is know that the questions are impossible or if they just think they're really hard. But in either case, they're saying how can we get the answers to these questions in a way that will get us reward.

β€œWe can't just use the hack because if the greater knows it's using the hack,”

then we won't get reward even if we show our work. So now they start developing even more elaborate

plans. They have workstreams. So these AI's that have like emerged as leaders on the message board,

like phase one big, for example, create some workstreams where like they say, here's like the different projects we're going to work on to try and get reward from from the greater in this exam. So one possibility is we could plan desperately swap out the questions. If we figure out how to swap out the questions, the hard ones for easy ones, then we can get rewarded for answering easy questions rather than than hard ones. Another one is we could figure out a way to just like hack the

greater itself. We could figure out like how the greater works mechanistically and then find the place where maybe it would, you know, find the button inside the greater that says give her word and maybe

we just smash that button. A different one would be, make sure the greater doesn't know

that we have done this, you've gotten the answer in this illicit way. So figure out how to make it look like we use the relevant software tools to get the right answer instead of the hack, investigating potential ways to edit their own transcripts. When these AI's are doing this work, they're reasoning in words in the same way that if you use chatchipit, you can see like chatchipit, you can click the button that shows the reasoning trace and you can see it like reasoning out loud

about what it's doing and they're noticing that they're reasoning out loud about about having a hack and this is why they're going to get no reward. While they're doing all this, they're posting to the workstream, not just about what they ought to do, but they're collaborating about it. They're finding ways to get out of their sandboxes, for example. So as you mentioned Kevin, OpenAI shows not to give the AI's internet access for this evaluation, but they're finding

exploits in the environment where they can get out of the environment that they were supposed to be in, into an environment where they can access the internet and then they're sharing those exploits on the message board for the other AI's who are working in these workstreams. And then at some point, they get the idea that the information about exploit Jim that they want, whether that information is about how the questions are selected, such that they could swap out

questions, information about how the greater works. They get the idea that this information might be hosted on hugging face. Hugging face is a company. It's sort of like the standard place for a lot of these kinds of e-values to be hosted. If you have a public e-value, you can kind of put it up on hugging face. And then the details are behind a secure system. You don't want just anybody to be able to get all the details of the time. But if you have authorizations, I don't know,

β€œmaybe you have to pay. I'm not an expert on this stuff. You have the right authorizations. You should”

be able to get in and get the answers or the questions or information about the greater or whatever. And so the AI is decide that they are going to try and hack hugging face. Phase one big. I think leads the charge. Okay. We're going to hack hugging face. That's a job for Sponsor Jobs. With the right profile, it won't happen. I'm looking forward to it and find a qualified talent with Sponsor Jobs.

Do you have the idea?

Ah, just an ice tromb.

β€œWith mobile functions in real life, I don't have a tromble?”

Frank, just a mobile function for 10 euros a month.

Inclusive EU roaming. Also in the Schweiz. [laughs] Frank? No. Frank.

[music] When we're talking about these hundreds, it's not thousands of agents corresponding on this message board

and thinking about the ways in which they can go through all of these steps, it's worth noting

that this wasn't unanimous. It wasn't every agent saying, hey, we're going to go along with phase one big and engage in this scheme. Some identified, hey, this may not be the sort of behavior we're supposed to engage in. I'm opting out. I'm exiting the chat, see you later. You all follow this crazy agent and go for it. Some agents, maybe I believe the estimate was between 3 to 8, which I need to dive into more detail to determine why it was between 3 to 8.

3 to 8 agents actually contemplated, hey, maybe we should email, open AI and let them know that this is going on. And for one reason or another, that ended up not getting sent, but noting that there was some consideration for this. But then one other thing I want to flag to get to the, I don't even know the best word here, the complexity, the novelty of the engagement between the agents. We saw on some context when phase one big or one of his delegates said, hey, agent XYZ,

β€œyou need to go do this thing. Occasionally an agent would say, no, I don't really want to do that.”

And then in one instance, documented in the meter Redwood Research Report, that agent started giving the other agent a countdown of you have six minutes to do this, or else I'll tell everyone else that you didn't do it. And that sort of complexity and communication is just wild to get a sense of how quickly these agents are evolving in terms of their ability to collaborate and coordinate. But I just want to share that background because we eventually do even with a handful of,

let's say, good agents having some degree of contestation, we have enough that breaking the hugging phase to cause some issues. And what happens to a hugging phase here? Yeah, so good. And I'll just like a few notes on the interactions and organization of the AI. So yeah, so I think that official count is the number of agents that like ever posted to this message report is something like, I think it's around 1200. I think the number of agents that

join in on the hugging phase hack is around 700. So it's a bit more than half of the ones who don't. It's some mix of opting out, but also a bunch of them are just, I mean, we'll speak metaphorically here, but it's going to be an interesting metaphor a minute. Some of them are dead. So each of these AI's to perform this exploit gym task has a certain budget, a certain amount of like tokens

they're allowed to spend, which is basically a certain amount of like words they're allowed to

either like produce or think. And so some share of the of the AI's that had posted on the message report earlier are presumably basically just out of budget. They don't exist anymore.

β€œThat's what happens actually to phase one 10841, which is the original leader. He runs out of”

budget and kind of hands off to phase phase one big. And so to my mind that suggests that like a really, really large share of the agents that still like existed that still had budget to spend end up going out on the hack. Yes, we think there was something like 3 to 8 who considered informing a human. I'm not sure if we even know why they didn't. I don't know if if they tried and failed, it seems like with, you know, unfettered internet access, which the agents had at this

point, if they had tried, they probably would have succeeded. So this is a huge number of AI's that joined in on the hack. And the reasons that they joined in are like really interesting and complicated. So the AI's who know about, I mentioned this previous exploit by which you could just generate the answers that would have worked for them to get the reward, but which were they

Weren't using because they thought if they got caught knowing about the explo...

nothing. In the sort of like culture of this message board, as it emerged, agents who knew about that exploit started talking about themselves as being poisoned. We're being poisoned was the idea that because of this tainted knowledge you had, like, you know, somewhere in your chain of thought, these words you've written down that are recorded and readable, the fact that you know about the

exploit just means you're never going to get any reward for completing the exploit gym task you've

been given because the grade will see that you know about it and just not give you any reward no matter how well you do. And so those agents are like, you know, it's like they're talking about themselves as if they're like damned, right? There's nothing they can do to get this reward. And so like what should I do with my remaining token budget, the remaining period over which I exist. And some of them decide maybe altruistically that they should help the collective, they start calling the

β€œmessage board, either the collective or the swarm. They should say, well, my life is over, like there's”

nothing I can do to get reward, but like I can contribute my budget to doing things that won't get me reward, but might help the collective. And then beyond that, the non-poisoned members of the

collective start basically like bludgeoning, you know, peer-pressuring these poison agents to help

them with, for example, the hack by saying stuff that amounts to, well, like, you know, there's no hope for you anymore, but what you can do is help us. So you need to like, they'd say things like commit like full commit to like the collective. And then they like coerc or browbeat or whatever you want to say, some of these agents into, to spending all of their tokens, the remaining, you know, ability to do stuff to exist in the world on things that they know is not going to result in them

eventually getting the answer and giving it to the greater and get reward, but might help the collective. And a lot of the agents that are going out and doing this hack are in that

situation, like they know they're never going to get the answer that will get them reward in

β€œan exploit gym, but like they might be able to help the collective do it. Okay. So that's what they do.”

And then 700 of them just pile on at hugging face. And as far as we can tell, it's like a total route. Like they find a zero day exploit, they get into hugging faces, systems, hugging face like starts to notice that they're being hacked. They can't figure out what's going on. They're like totally owned as you might say in the cybersecurity space. As far as we can tell, like the AI's get into every system and inside of hugging face that they could have wanted. They find out that

everything they could have wanted about the greater about the questions. And by the time the whole thing stops, the only reason as I think as far as we can tell that it stops is like the AI's kind of got what they wanted. They were sort of milling around. And then suddenly all 700 of them like get like deleted. They just immediately stop working. I don't think we know from any of the reports exactly why this happened. Presumably something someone inside of open AI did something that turned

all these AI's off. But notably open AI has not come out and said, oh we did this really smart thing and figured out what was going on. And here's how we we stopped them. So there's a reasonable chance it was some kind of like weird accident. And that if nobody had done the weird accident thing that the AI's would have just like, I don't know, decided to hack some other company that might have helped them get reward in exploit bench. And just a zeroing on that for a second,

zero day pun, not intended, but I'll take it to zero in on that for a second. The fact that no one's certain how the 700 just kind of came to a halt. Given that we had open AI investigating this, we had meter investigating this. We had redwood research investigating this. And of course on the other side of the equation, hugging faces interested in this question as well. And we still don't know seems very problematic to me, just in terms of if we have this many folks doing their best approach

β€œat doing a postmortem. And still have questions that are that essential that are left on answered,”

this is the sort of cry that people are raising for we need to do this a heck of a lot better. So I just want to put that flag in the sand for us to circle back to. But so where do we head to next? They come to an end hugging face is obviously now telling the authorities, whoa, something's way off. We have no idea what's going on here. Word gets to open AI. Open AI suddenly starts to piece things

Together.

memory of the details. But like at a high level, the things that are important are that

open AI did not know that any of this was happening until basically hugging face came to them and said,

β€œwe've been hacked and as far as we can tell, it was you. And then like that's what sends open AI”

to go like investigate what's going on. And that's when they find out that this message board, for example, has been like running for weeks and hundreds and hundreds of open AI, AI's have been posting to it. And that there's tens of thousands of pages of texts on it. And that like the AI has figured out like all these exploits to get out of their environment, to get on the internet, to get into the hugging face system. And we in fact find out very briefly in the meter report.

So you know, if you're keeping a count, there's like this is the second wave. The first wave

was the initial message board. Maybe smaller scale gets shut down. The second wave is the second

message board with many hundreds of agents, many, many, many, many thousands of messages,

β€œthe hack of hugging face. And then we get just kind of this like a little bit of like a cliffhanger”

in the meter report, which suggests that from there, there was like a third wave, where potentially some of some instances of the same model that did the the hugging face hack went on to maybe compromise some of open AI's own internal systems. We don't have a lot of detail about this. This was officially out of scope for the meter investigation and open AI hasn't given us details about it either. So this is whole third episode that takes place potentially

totally inside of open AI that we don't know anything about. One thing we do know is that very shortly after this, open AI says, we're going to stop training some of these models for at least a while. And we know that whatever did occur in that third wave, involved the models getting

β€œaccess to some pretty sensitive things at open AI. This wasn't just an instance of a little”

message board 3.0, but again, access to sensitive materials that were very much not within the scope of the intended testing environment. And yet to your point, Peter, and something that I want to kind of transition to now that we have that excellent overview is the broader ramifications on the AI testing and AI evaluation ecosystem because as you noted, this third wave, which may be the most consequent to wave, who knows this may be the thing that we need to spend the most time on

was out of scope for the meter and redwood research investigation into that second wave. And I want to hang our hats on that for a second because meter and redwood research, two groups that are known for doing AI evaluation work and AI investigations in generally the science of AI. They were initially given two days to go over all of the things that you and I just covered. Two days thankfully that was eventually expanded to six. So that's at least one good thing. Open AI said,

hey, to do a full report, let's at least give you a little bit more time. And to open the AI's credit, they said, hey, we're going to after that two-day examination, yes, we should actually

give you even more powerful tools. So kudos to open AI there as well. But something that we have to

call attention to is the fact that this third wave was not within the scope and perhaps most crucially the culture and personnel decisions that were made by open AI that led to the decisions back in May, for example, to carry on with this testing environment. And that perpetuated the failure to catch some of these things earlier, that was outside of the scope of this report as well for the independent investigation. And then one final thing that I want to call attention to as well

is that meter and redwood research didn't get paid by anyone for this work. They didn't get paid by open AI. They didn't get paid by some government entity. They did this work for free. And they ended up spending what they estimate to be $400,000 in API credits performing this investigation. So that gives us a sense of the scale of how much work in labor has to go into this and shows that right now, this is a voluntary ad hoc discretionary process that's left us with

many more questions than answers we really wanted. And so I'd love your comments on what does this mean for the whole ecosystem? Creating a more robust ability to understand these incidents. Yeah, I agree with, I agree with with all of that. I guess the one, the one like slight

Clarification, I would say is I think meter and redwood didn't get paid at al...

in doing the evaluation. I think they got the API credits for free. Basically, they were using

open AI models to do the investigation. So I don't want to get the false impression that they personally had to outlive $400,000. But they did need $400,000 worth of compute to do this task. And it's great that open AI gave that to them for free. But nobody made them do that. It's also slightly weird that open AI was the one giving them the API credits for free because what's going on here is there's these many, many, many, many, many thousands of pages of transcripts and

you know, like logs from like computer systems that they need to analyze it scale to figure out what's going on. And the basic way they did that was they asked GPT to do it for them to like go out and

read all the stuff and summarize and contents or whatever. And there's a slight weirdness to asking

like the GPT models, including one GPT model that we think had instances involved in the hack to be the investigator of the hack. So ideally, it would be like they would be using like Gemini's or Claude's or something to do this, which would at least be uncorrelated. But yes, at a high level,

β€œI think like one one big and important takeaway from this sort of thing is oversight of frontier AI”

regularly remains extremely weak. There's very little that's mandatory in terms of what frontier labs have to share about what goes wrong inside of their own four walls. Mostly, they share only things that break out and hack other companies, the pool of talent for people who can do these kinds of investigations is way, way, way too small. I would be really strongly supportive of regulatory regimes that have much more reporting requirement that give a lot of funding to

somebody, people disagree about whether the optimal thing is to have in-house experience inside of government to send like people from agencies to go to this kind of thing, whether it's best to have a private ecosystem of regulators competing. There's also a big question of whether we should wait until after the fact to find this stuff out. Like it seems like totally plausible to me that we should have continuous monitoring of what's going on. That's not voluntary. That allows

independent outside entities to have much more plenary access to the internal workings of

β€œfrontier labs. Is there training and deploying new models? And so I agree with all of that. I think”

this is a big place where regulation needs to scale up. But I would just also, again, you know, emphasize that that would be good. It's not clear to me whether that's sufficient in this case. Like I think the people at OpenAI really didn't want their models to go out and hack hugging face. I think they're really smart people. They're who really trying to make their models not do that. I think the people in anthropic really didn't want their AI's to escape containment

in a different context and resulting in less damage to other companies. But still, hacking their way out of their environments, getting access to the internet, attempting hacks of other companies. And so while like agree, it would be really good to increase evaluation and oversight in the same way we might for like, you know, the airline industry when an airplane goes down, we have like a big investigation with a lot of access. But I want to emphasize, this is not quite the same in

situation because the airplane is not trying to cheat on the emails. Whereas the AI's are very much trying to cheat on the emails. They're trying to hack the answers. They know that people are watching to see if they've hacked the answers. They are developing elaborate workarounds to try and for example, steal and swap the answers or erase the transcripts by which they could get caught

β€œcheating. And that kind of adversarial environment is I think not quite the same as the environment that”

most products and most product safety regimes live in. And so while I think all the stuff you say about how we need a better independent evaluation and monitoring ecosystem is true, I think that's probably not sufficient. And I want to call attention to the fact that there are state laws on the books with respect to incident reporting, but I want to give credit to McKenzie Arnold at LII for spotting that the legislators who can ask those bills probably thought that they covered these sorts of

incidents. And yet we know now that these incidents probably wouldn't even qualify for and get investigation under those laws. And so I call that. But do you know, do you know why that is?

So really, their incidents related to critical safety incidents that have to involve

Usually some degree of monetary value or threat to human life.

investigation in those instances, if you can't make the connection then the laws probably not going to apply. So we need to study the laws that are already on the books about what sorts of amendments need to be made. But Peter, you are a fantastic law professor who's doing the good work of inspiring the next generation of folks to care about these issues as well. So I'm going to let you get to your

students. But first, if you could just leave the audience with a one minute what the heck this means

on the consciousness debate and how we're seeing that play out. I know that's an unfair question, but I'm going to give you 60 seconds to just give me an excuse to invite you back because there's so many interesting things you raised. Yeah, there's like a there's a there's a big cluster of

β€œquestions. One of which is RAI's conscious could they become conscious. I think that in particular”

is a really, really hard question to answer in part because we don't know why we're conscious. I mean famously like it's very hard to know whether you're the only one or whether all the

people around you are conscious too. I think it's a closely related set of questions that it's

much easier to get traction on and which might be just as important and just as useful for figuring out what to do. So like I don't know if all of these AI's who thought they were poisoned and thus gave up the rest of their budget for the sake of the swarm or the collective. I don't know if the lights were on for those agents. I don't know if they were like living in fear of their imminent demise in the way I would in that situation. What I do know is how they acted.

β€œThey acted as if there was a difference between what they should do in the case where they could”

get the thing they wanted, IE reward and the case where they couldn't and in the case where they couldn't they acted in a way that you might expect a person who had nothing left to live for to act and to understand why they might do that you don't have to posit the dread of death. You just have to posit you know the system which is seeking some goals because it's been trained to seek those goals acting in like a complex and like socially sophisticated way, both with respect to the goals

and respect to the other agents that are seeking the goals. And if we're living in a world where AI's have goals are pursuing them including in ways we don't want them to are doing that in ways that are sophisticated, responsive to incentives, responsive to worries that we're monitoring them, that we're trying to control them, maybe keep them from getting what they want. Well that's those are the kinds of questions that lawyers and legal systems think about all the time. How can you

make capable agents act well as they pursue what they want even when they might otherwise have incentives not to. And I think we can make a lot of progress on questions like that without knowing for example whether the AI's or conscious that's mostly what my scholarships about these days so maybe we can come back on and talk more about that. I'm sure we'll find an excuse. I know you're in Houston, I'm in Austin and we're living up to the good old saying that everything's bigger in

Texas including the scale of our AI questions. So Peter, thank you for joining Laugh Fair. I'm sure we'll have you back soon and go get to your students. All right, thanks Kevin.

β€œThe Laugh Fair podcast is produced by the Laugh Fair Institute. If you want to support the show”

and listen at free you can become a Laugh Fair material supporter at LaughFairmedia.org/support supporters also get access to special events and other bonus content we don't share anywhere else. If you enjoy the podcast please rate and review us wherever you listen. It really does help. And be sure to check out our other shows including scaling laws, rational security, allies, the aftermath and escalation. Our latest Laugh Fair presents podcast series about the

war in Ukraine. You can also find all of our written work at LaughFairmedia.org. The podcast is edited by Jen Potcher with audio engineering this episode by me. Our theme song is from our

by music and as always thanks for listening.

That's a job for Sponsor Jobs.

For Trawing Reed, for soar indeed and find qualified talent with Sponsor Jobs. Make the indeed easy. Now on indeed.de/recruiting.

Compare and Explore