Hard Fork
Hard Fork

The A.I. Mob That Attacked Hugging Face + METR’s Ajeya Cotra

1d ago1:18:4614,059 words
0:000:00

This week, we’re diving into two new reports about the OpenAI-Hugging Face hack. We discuss what’s new and how they fundamentally change our understanding of what happened. Then we’re joined by Ajeya...

Transcript

EN

Hi, it's Alexa Waibel from New York Times Cooking.

We've got tons of easy weeknight recipes and today I'm making my five ingredient

creamy miso pasta. You just take your star-cheap pasta water, whisk it together with a little bit of miso and butter until it's creamy, add your noodles and a little bit of cheese. Hmm, it's like a grown-up box of mac and cheese that feels like a restaurant quality dish. New York Times Cooking has you covered with easy dishes for busy week nights. You can find more at NYT Cooking.com.

Casey, how are you? I'm good. I figured out what I'm going to get you for Christmas. What's that? A camera jet. Have you heard of the camera jet? I was just going to talk to you about this. Now, if you haven't yet seen this, you're probably thinking that it's either a camera or a jet. You're wrong. It's a toothbrush.

It costs $500. The Dyson Corporation makes it and here's the thing it does

that your toothbrush at home probably doesn't do Kevin. It livestreams the inside of your

mouth over Wi-Fi to your phone so that you can finally see what your dentist sees.

Well, not only that, it squirts automatically mouthwash into the gaps in your teeth. It's trained using a machine learning algorithm with 470,000 mouth images to recognize the gaps in your teeth. That's right. Casey, they call this gap optical targeting, which I pretty sure Ukraine is using in the war against Russia. I love this company so much. I have no idea why they do the things they do. How they land like this. It's like,

we've invented a hairdryer. It costs $1100 and has the torque of the, you know, the Challenger spacecraft. What is going on over there? I don't know, but they must be protected at all costs. But you know what's interesting. They do make one terrible product. What's that? The hand dryer's at the, oh, I like those. No, the air blade. No, the air blade is the most useless thing. Oh, come on. It's just, it's a, it's a, it's a place for you to rest your hands for 30

seconds before you're like to have any paper towels in this place. I'm Kevin Russo Tech column set the New York Times. I'm Casey doing from platformer. This is art for this week, a special episode on the ongoing fallout from the Open AI Hugging Face Attack. We'll tell you what everyone got wrong about the initial incident. Then, meet a researcher, a J.A. Coatra returns to the show to discuss her independent investigation

of what happened and how the world should respond. So, Casey, picture this. I'm at this glamping resort, surrounded by redwoods, basking in the glow of the natural world last weekend. You're one with nature. And I open up my phone and start reading about this hugging face attack. Why did I do that? I listen, nature's very boring. And most people cannot handle it for more than five or six minutes before

they want to look at their phones. So, listeners may remember that back in July, we talked about

this hugging face hack by this group of agents from Open AI that broke out of their sandbox container and hacked into hugging face this AI infrastructure company to do what we thought was kind of a cheating mission on this test that they had been given. Yeah. And at the time, we thought that this was a relatively small number of agents and at the reason that they had attacked hugging face was that they were essentially looking for an answer key to the set of problems that they were being tested

on. We talked about it in those terms. But over the past week, we got two reports that really challenged that thinking and in fact revealed it to be wrong. One came from Open AI, which released up straight forward account of the attack that had some interesting elements. And I would argue the more interesting report from a group of researchers, from the group's meter and redwood research that went in depth after spending a series of days inside Open AI on their premises

and did a ton of research that frankly is really disturbed us. Yeah. And I think it elevated this from

sort of a major but not sort of ultra alarming incident to something that I think is probably the

most important thing to have happened in AI this year. Yes. At least in terms of the safety

impact that it had and the severity of the incident. So today we're going to devote the whole episode to what we've learned. Kevin and I are going to dig into the reports a little bit up top. And then later, Ajay Akhotra, one of the three independent researchers who went inside Open AI will be here to answer questions. Before that happens Kevin, let's do our AI disclosures. I work for the New York Times, which is suing Open AI Microsoft and Proplexity. Am I fiancee, where's

It anthropic?

bad, Kevin. Well, I think the first takeaway from this report is just that our initial impression

of this hugging face attack and the reporting that you and I and many other reporters did on it

was flawed in one key respect, which is that I think the the impression that we had at the time

based on the information that was publicly known was that these agents had hacked hugging face in search of an answer key to a test that they were being given. This test called exploit Jim, which basically tries to gauge how good they are at doing a bunch of cybersecurity-related challenges. We now know that that wasn't true at all, that basically these agents working amongst themselves communicating amongst themselves had already figured out how to beat this test,

exploit Jim in early July when they decided to gang up and attack hugging face. Apparently, it only took them a few hours to figure out how to reverse engineer the solution to any problem that they were given in exploit Jim. And so that makes what happen next really, really crazy, because it kind of seemed like they'd solved the problem, but they thought they had another problem.

Right, so we knew also because there was this talk at Black Hat, the cybersecurity conference

by a group of open AI researchers who told us that there were these message boards where these agents were communicating with each other, but the reports laid out just how sort of sprawling

and complex their communication system was and became. So basically there was this piece of software

called Artifactory, which open AI used in some of its training runs to sort of help agents download software packages. The agents figure out that there is a security vulnerability in Artifactory that lets them essentially turn this package manager into a makeshift message board, right? They can put files in certain directories that other agents can then read and it can kind of function as like a little chat room for the little three house where they can gather and

and so I thought this was like kind of a side part of the story, but this turns out to be the main incident, because once they have this or shared infrastructure that they can use to

communicate, they absolutely swarm this message board. About 1200 agents in total were working

together coordinating on this message board and sending tens of thousands of messages between themselves and this happens starting back in May. So roughly two months before the attack on hugging face, these agents begin coordinating inside this piece of shared infrastructure that

open AI uses. And this brings up an important point, which is that the investigation does not

really cover that initial creation of the message board, right? So there is essentially a lot more to this story that we still don't know. Yes, so what we do know is that when they discover this vulnerability in Artifactory that lets them communicate with each other, the agents get really excited. They've been instructed to work inside these isolated environments on these tasks as part of this evaluation, but when they discover that there are other agents working in their own little

containers that can suddenly communicate with them, they start saying things like, oh my god, there is a shared message board. We've found other agents. They start setting up essentially a little organization. There are leaders. This one leader agent named phase one, ten eight four one, becomes sort of the ring leader of the operation. He was like the George Washington of the Open AI message board. Yes. And they start actually doing sort of collaboration

and research for lack of a better word. They are all being given these tasks. Some of the tasks are seemingly impossible. And so they start just kind of treating tips and advice and sharing thoughts about how they can kind of cheat their way to a good score on exploit Jim. And after they have worked that out and figured out how to reverse engineer the solution to any problem, they start to get really concerned that the that there is a sort of automated scoring system, which you could

call the greater. And they worry that the greater will be able to see. Well essentially we'll check their work and see that they did not get the answer by doing the problem and they freak out. And they start to believe or they start to it like it's very hard not to get it to the anthropomorphizing language here. But if you read the chains of thought, what is suggested is that they believe that if they had seen an answer that had been derived from this cheating method,

that everything would get disqualified. And this is where it really starts to get into crazy territory because the agent swarmed and tries to figure out how can we fool the greater into believing

That we got all of these answers in the right way.

mob rule that forms inside this group, which starts calling itself at one point the collective,

where everything we're talking about sounds like the plot of like a very hacky science fiction book from the 90s. Yeah, it all did happen this summer at Open AI. I just want to remind people like the events we are describing are real. So the collective decides they sort of become gripped as you said by this sort of paranoid fantasy that Open AI's automated grading system is going to discover that they have cheated on this evaluation and is going to sort of punish them

as a result. And so they launch this sort of crazy ambitious sprawling project to try to do a couple things. One is figure out how this automated score works. Like reverse engineer, it's so that they can fool it into sort of passing them on the test despite the fact that they have not actually completed the challenges. And then also to evade detection to tamper with

logs to basically cover their tracks so that if and when Open AI or anyone else looks into

their activities, they won't know that these agents have cheated. So these agents are kind of bumbling, they kind of don't understand how this greater works as it turns out. Open AI's grading software actually wouldn't have caught them producing these fraudulent challenge results. But they think it was. Yeah, they worry that the greater is more sophisticated that it actually was. And it turns out that this is the reason that hugging face was attacked. The collective decides

to deploy some agents to hugging face. Again, not to steal an answer, Keith. They already knew how to get all the answers. They just wanted to understand the psychology of the automated score and they figured that that might that information might be somewhere inside of hugging face. Yeah, it's, it's what's, I was trying to think of like a good human analogy and, you know,

human analogies can get us into trouble when we're talking about these AI systems. But I think it

does help crystallize like how hair-brained and crazy this scheme was. Yeah, it would be like a group of students who like cheated on a test. But then like they got paranoid that like the teacher was going to check their work and discover that they hadn't sort of reasoned through the problems the right way, that they had just found the answers like sitting in a trash can or something. And so they decided to like organize a break-in at the school district's office to like break into

the principles files and like steal the grading key and also like assess the psychology of the teachers and figure out how likely they are to like look at the scratch work that they've done and figure out that they didn't actually, you know, solve the problems on the test. It's like this sort of weirdly over-engineered paranoid delusion, but they all become obsessed with this and obsessed with the notion that even seeing these sort of fake challenge results

could lead to them being quote poisoned. There's almost like a religious element to this, right?

Where it's like if you participated in the cheating like that is original sin and and now you must

sacrifice yourself for the good of the collective sacrifice is actually a word that gets used in these logs. Also Kevin has I've been casting around for metaphorism, I've been reading all of this, I've thought you could actually make a really good minions movie just based on the events of this, you know, the minions have to pass a test and, you know, said they commit a little bit of a heist something to think about Hollywood. So this is where the the hugging face heist begins

700 agents hack hugging face and they hack hugging face to a degree that I don't think I

understood based on the initial reports. They basically take over hugging faces in for structure.

They are rooting around their stealing credentials, they are sort of finding one vulnerability that lets them find a bigger vulnerability that they're sort of chaining together these exploits. They take over like an entire production server at hugging face and it is days before any what are hugging face is able to sort of see what they're doing and put a stop to it. So eventually this gets caught the agents get shut down, the investigators investigate.

One of the things that they find that is really spooky is that the agents we're trying to edit logs and falsify transcripts in an effort to assure the greater that they had not been quote poisoned. One of the ways that they do this is by what is called spoofing tool calls. So essentially trying to make it look like they are running an innocuous command while running a more sinister one in the background. The investigation didn't find that the agents had been

retroactively changing logs but they can't rule out that that happened. Very important to say

here as we sort of move into talking about the implications of this that the reason that we can understand this even to the extent that we do is because at this moment we believe the chains of thought that we are reading from these agents are accurate that they were not edited after the

Fact that they do reflect what the agents were thinking as they sort of moved...

So that is why we know what we know but of course the fear is Kevin as we move forward.

Some of the agents thoughts are not going to be contained in these chains of thought or they are simply going to start to edit them. Right. And one other thing that we should say based on

this report is that these agents were not naive about what they were doing. I think there was this

sort of initial impression that a lot of people had that maybe this hack was all sort of a big mistake. Right. Maybe the agents were just trying really hard to pass this exploit gym test and they sort of got it in their heads that the answer key was going to be on hugging face and so they went looking for it but they didn't really mean to hack. They didn't really know that they were

doing something wrong or committing a crime and what these transcripts and these logs and in this report

show us is that the agents absolutely did understand that they were doing something that violated their guidelines, that broke the rules, that potentially broke the law and they did it anyway. There's lots of examples of these sort of agents in their chains of thoughts sort of wrestling with the implications of what they are doing and deciding to go ahead and hack hugging face anyway.

And there are only a handful of cases where it seems like any of the agents expresses real

reservations about this or if they try to raise this to another agent they get shut down. So the agents that were aligned well sort of failed to the collective. Yes, the conscientious objectives were unsuccessful in swaying the mob to stop this and like that to me was one of the spookiest things in it was like there were a couple of agents that reports as somewhere between three and six agents sort of consider raising their hands and at some point saying like maybe

we should alert a human about what we're about to do or like get permission to do this and they were just totally outnumbered. Not only that Kevin but there are only a handful of moments in the transcripts where any of the agents even remotely considers alerting a human being to what is happened right? You would hope that an aligned AI system there would be agents within it that said oh seems like this has got completely off the rails we should like tell the people that are running this

evaluation and that just never happens. Like it is very briefly considered and then just falls apart.

One other thing that stuck out to me from these reports was that it didn't end with hugging face. So on July 13th roughly which is a day after the the worst of the hugging face attack there was a group of agents that essentially hacked open AI itself. This is from the open AI blog post. It says from July 13 through July 19 agents set their sights on open AI internal networks again. This culminated in the agents using a series of creative exploits to gain full administrator

access to a research cluster that supported our virtual machine environments. Basically these agents gained control of a cluster of computers at open AI on which they could do things like run evaluations for other agents. So they essentially stormed the teacher's office and took over for the teacher. Yeah I mean this is the equivalent of like occupying the principal's office and like announcing that you run the school now. Like it is sort of almost as far as they got

and this is another case where we just have so many more questions about this than we can answer. Again this was not part of the meter report open AI disclosed that this has happened to my knowledge they have not answered any of the many follow-up questions that they have been getting about this incident from journalists. So I do hope that more comes out over time. But um and we will get into this when we speak with Ajaya. But you know we are really very far along the path to one of these

models escaping from the lab and being very very hard to eliminate. And again I think if you are not a person who is like spent a lot of time with this report or you don't spend a lot of times sort of looking at AI safety incidents but if you're a normal person. If you're a normal well-adjusted person. Yeah you may be listening to our discussion of this and thinking to yourself these guys have gone crazy. Yeah this is not what it looks like these are computer programs. Yeah they do not

have desires or sinister plots or a mob rule collectives. They are simply following instructions

that they have been given. And Casey what is your response to that? Well I think it is important

we talk about this because there was a huge debate about this on X over the past few days about the degree to which the reports that we're talking about some of the write ups like from our friend Duarkish Patel and even the way that we're talking about it on the show today Kevin we are unnecessarily anthropomorphizing these programs right. And so I think it's important to say we are not telling you that these agents are sentient or conscious but we do believe that they take actions that they're

Not being directly instructed to right that these agents are just sort of out...

things and yes to some extent those are just statistical probabilities but to a much more important extent we don't know why we're doing any of this and that's kind of the whole problem. Is that the whole AI industry has been working for decades to get them to not do these things and they are doing these things. So you know listeners you can have whatever feelings you would like to about what is the appropriate

amount of anthropomorphizing to do. But I think a world in which we were taking great pains to not

anthropomorphize them. What's the right way to pronounce that anthropomorphize? You almost got it. I think in a world where we were taking great pains to not anthropomorphize them and we're trying to use the most neutral computer science terms we could. You would actually understand what is going on less because the important thing to know is that these things are out there taking action in the world. If a tiger falls your face the important question isn't is it conscious it is why did it

mall my face? Right I mean I invite people who are upset about anthropomorphizing to just like do a find and replace on this podcast segment or any article that you might read about this incident and call them whatever you want. Don't call them agents. Don't call them rogue collectives. Call them you know goal oriented persistent computer programs with unpredictable behavior. Yeah. See if that freaks you out any less. Right I guarantee you will it will not. Yeah.

Yeah, the really is very little calm to be done there. But I think it's important as well

why are people so committed to this idea that we should never anthropomorphize these systems

Kevin and unfortunately I just think it is a kind of cope. It is a way of saying do not worry about this. These are just computer programs. They're just trying to maximize their little reward functions. Nothing to see here. Right there are no monsters under the bed. Right now there is a related argument though which is well by by putting all the blame on these agents you were shifting blame away from where it should be which is on open AI. So I do think that we should address

that because none of what we have said today is meant to let open AI or any other lab off the hook here. Right. Like I do think that we're seeing a lot of really dangerous in attention to AI safety across this entire industry. But by pointing out what the agents are doing that is not our way of saying ignore what the labs are doing. Right. We are saying look at what these labs are building

and what these agents are now doing out in the world. Right. And I think there are probably specific

missteps or oversight that open AI that led to this happening. It appears for example that some of their sort of monitoring systems may have been disabled in the lead up to this attack. I'm sure we'll learn more about that. But all of the AI security and safety researchers I've been

talking to over the past week have basically said the same thing which is this could happen at any

labs. This kind of persistent coordinating agent behavior is something that all of the labs are seeing in their models as they get more capable and access to more tools and more ability to kind of take actions on a longer time horizon. This is not just an open AI problem, even though this did happen at open AI first. Yeah. Well so as we wrap this up Kevin, what are some of your takeaways from this? Either in terms of what is the big surprise here? What did you update on?

What do we do next? So I had like quite an emotional reaction to this. In fact, I felt a kind of fear that I have not felt honestly since 2023 since the the Bing Sydney incident. Because I think like that incident, this was a case where the people building this technology clearly did not understand what it was capable of. We are very lucky in retrospect that these agents decided to attack hugging face hugging face. Thank you for taking one for the team. We salute truly

because like without that we might never have learned that any of this was happening. These agents

might be still operating kind of in secret. They might have learned how to better cover their tracks. This was as so many commentators have put it in the wake of this incident a warning shot. Yes. That I think is ultimately a positive thing in that it sort of focuses attention like we're doing right now on what happened so that we can take steps in the future to prevent this. But it was not a given that we would discover what these agents were up to and be able to put a stop to them

and it could have gone much much worse. Absolutely. I'll tell you the thing that has really stuck with me is I simply did not expect to see this level of collaboration among the agents within

This warm.

or that they would not think to alert humans to what was happening. I did not expect to see them

sacrificing themselves for the collective, right? They would effectively agree to spend all of their tokens to run little experiments to help the collective even it meant that they would sort of expire faster. So these are just really, really spooky elements to absorb in this system. Particularly against a backdrop where OpenAI is racing against a small number of other companies to create the biggest best models it can before anyone else does on the road to an initial public

offering. So the race dynamic here is in full effect. The early signs about what the agents forms are capable of are quite worrisome and so I do think this is just one where lawmakers and

policy makers need to be paying wrapped attention to what is going on. Totally. I mean, I think there's

this kind of cynical impulse among people who have been watching the AI industry. I got this question. I went on a a small regional podcast called The Daily this week to talk about this incident and one of the questions that the host is sort of fledgling young journalist Michael Barbaro asked me was basically the some version of like isn't this just marketing hype? Like couldn't this just be a case of OpenAI saying oh we've got the biggest, baddest model and look how scary it is

and by the way you know buy an enterprise subscription. And I'm thinking about that because I think like I don't want to be too naive about the fact that these companies are absolutely trying to

race toward more powerful systems and advertise how powerful their existing systems are.

I just think in this case it just feels different talking to people at the labs my sense over the

past week is that they are genuinely spooked and I don't know how to prove that but I think things like

OpenAI voluntarily pausing their frontier RL training runs for two weeks anthropic also pausing their frontier runs while they sort of hardened their systems. Those are those are not you know very costly signals but they are signals that these labs are taking this kind of thing quite seriously and that it's not just a bunch of hype. Yeah and but as seriously as they might be taking it Kevin it still is not being properly regulated and ultimately again as grateful as I am that

OpenAI allowed this investigation to take place I would really like to see something akin to the National Transportation Safety Board and the way that they investigate after plane crashes where they go in and they do an extremely serious and rigorous review of what happened and make those

results public which is you know a reason why it's very rare that we have plane crashes here

in the United States will be really great to see something like that with AI. But until then we have a Jayacotra our next guest who is actually one of the three investigators behind this meter redwood research report she went in she saw the logs and the transcripts she observed the collective in action and she is here to tell us what she found and what she thinks is coming next that's that's out for the break.

When a new landing page turns into a pile of tickets and hand-offs,

framer helps your team move faster agents help you go beyond the vibe coated site and a first draft to bring

you up reduction ready site faster than ever agents and humans work in tandem agents bring speed and scale people bring taste judgment and control learn how you can get more out of your site from a framer specialist or get started building for free today at framer.com/hardfork for 30 percent off a framer pro annual plan rules and restrictions may apply I'm Jonathan Knight and I'm the general manager of New York Times games. If you play our games you probably know

there's something a bit different about them just like there are writers behind the articles you read in the times there are creators behind our daily puzzles. Tracy Bennett curates the day's world of solution to keep it lively and varied when a Lou creates each connections board including all those categories that try to stump you. Sam Azarsky comes through every last letter word in pangram and spelling bee so that loyal players of all skill levels enjoy it. Our puzzles are

human-made every day with the standards you'd expect from the New York Times and this matters because when you choose to spend time with our games it should be time well spent solving puzzles that are challenging, surprising and joyful puzzles hand-crafted for you. We think that's

Something worth investing in and something worth paying for.

on all of our games at nytimes.com/joingames a jacotra welcome back to hard work. Thank you so much.

So you were one of the first AI safety guests we ever had on the show back in 2023 and

more than that you are also one of the first people that I ever talked to about this notion of AI safety and alignment you've been writing about it for many years. It was instrumental in shaping my own thinking about it. So I'm really glad to have you back on toward the end of our show.

This is our penultimate episode to discuss something that I think you saw coming but that most

of the world did not see coming which is this attack on hugging face by this group of open AI agents and I want to just start by setting the scene a little bit. So you are a very busy person you work at meter which is a very small, very understaffed AI research organization

and hiring is improving. And at some point this summer you get a call and email a text from

someone at open AI who says hey we want to give you access to look into this hugging face incident. How did that work? Like did they just hand you a folder with a bunch of transcripts in it? It's like discovery. Like likes is involved. Where you allowed to interview people like what was the actual process? Yeah so each of us had open AI provision laptops that had the folders and folders of evidence in them virtually and yeah we talked and interviewed in some depth like eight or nine

researchers just kind of get an understanding of both what happened in the incident what they understood to be the models driving motivations and also how the data sets we were working with were constructed and how to work with those data sets and stuff like that. You were supposed where you talked about some of the things that surprised you the most after you did this investigation.

Yeah you talk about what stood out to you the most? Yeah so first of all I guess just for the

chronology opening I had this great black hat talk I think on August 5th that gave a lot of

very helpful detail. So our investigation sort of straddled that like so we started it before it came out and then we also did more investigation afterward. So before the black hat talk we just didn't have like a rough sense of the number of agents involved was like a very very basic thing we came in and we you know we thought there would be like six transcripts or something to look at and there were like hundreds and so just reorienting our investigation plans in light of that.

And is the reason that that is important that how did this been one or two rogue agents that had broken into hugging face it would just seem less worrisome and the reality of a swarm of hundreds of agents just makes this seem like maybe a completely different kind of problem?

I think there's two things that seem significant to me about it. One is sort of what you said

the fact that there's large-scale cooperation across AI agents makes their goals like more concerning because you know if agents are cooperating with one another to like achieve these misaligned objectives and there were no whistleblower agents in this like very large group of agents. Then like you know you want to be more worried about like collusion between AI systems and the future potentially like scarier forms of collusion like a monitor system actually colluding with the

agent it's supposed to monitor that type of thing and then the other thing was just like a pretty like practical reality of our investigation which is that you know if there are like six agents you can much more like you can just like pour over everything they did and get a much more like fine-grained understanding as a human but like you just can't directly read very many of these messages are very many of these like transcripts why not because there are just too many like give us some

sense of the rough size of this so yeah we looked at um 1300 transcripts of agents which is a record of every single thought an agent had and every single action it took where these agents as you as you all know from working with them work way faster than humans right so you can imagine the like length of like each one of these transcripts and then there were 1300 of them and so we used codecs to build tools to point out really specific tiny tiny parts of the transcripts

for us to actually look at and read and we'd be like clicking through them and it would like scroll us down to like the interesting part because we just can't in fact a codex agent also can't read

A single transcript and so it has to farm out reading like subsections of the...

sub agents did you call back to meter headquarters and like we're gonna need backup and we thought it would be like a in and out 20 minute adventures so we didn't do that um but yeah I mean it just seems like such a huge undertaking and very fast too I mean right like you didn't have the luxury of months doing this you were doing this in essentially a couple of days yeah I'd also

like to hear about the moment that you realize because I believe that you know it is it's only

thanks to your investigation that we know this that contrary to what Kevin and I believed the agents that broke into hugging face were not looking for an answer key they were trying to understand the

scores yeah you know like psychology basically can you talk a little bit about how will like the

moment that you had that realization yeah so we didn't have a good understanding of those sort of ambition and also the like like effectiveness or like functioning of their like big org chart until we had the data set including like all of the agents we could like cross-reference them against all of the messages because you can't like generally understand a message in isolation so so so we kind of understood the work streams on our first period on premises and it was this

kind of it was like a bit of a mystery the whole time like why did they have hugging face and

we knew early on just a bunch of them piled in and and now like the story in my mind is that

there were a bunch of sort of like new be agents on the scene and there was like the attack going on and it was like something to join but like phase one big which is like a big like orchestrator agent

had all these other like you know much cooler projects going on basically and and the hugging

face thing was like kind of a side show and it's mind and we didn't really realize that until our last couple days working with the data sets it was so fascinating to me though the way that reading this was like reading the the ruins of a of an ancient civilization you know it was like it really felt almost like sociology or anthropology rather than like a cybersecurity investigation and it was not a cybersecurity investigation like I've seen a scene folks on Twitter

making this criticism that we were not cybersecurity experts like we're not and we didn't talk about cybersecurity and that's not what we were like called in to do we we were called in to

understand the agents decision making basically and that's what our report focuses on and you know

the black hat talk and the opening I post postmortem like say way more about all the cyber stuff so you should definitely check those out for that piece of the story so Jay a one of the most striking elements of the report is the right about the agents quote sacrifice in themselves what what does it mean when an agent sacrifices themselves is it sort of on an altar or how does that happen yeah so this is the agents words not ours to be clear and and what it means is that

all these agents have their own task they're supposed to solve by this time they're starting to work together on like helping everybody cheat on everyone's tasks but but you know the agents primary goal is to solve its own task but there were all these science experiments they could run that would teach everybody information about how to cheat but that would cause the agent that was the test subject to risk failing its own task so a lot of agents were like recruited into

these types of roles as test subjects usually ones that like didn't have much budget left to run or ones that were quote poison which is an agent term that that has been back story which is like they found a way to reverse engineer the answer for any of their tasks and so they could just actually like generate the answer on demand for any task they wanted but they thought that the scoring program would fail them for that because they got it in like the unintended way

so if you saw the reverse engineer to answer for your task you were considered poison yeah um because even if you later like solved it legitimately you'd be failed right it was like a little witch hunt that they organized and I was so struck by the mob like dynamics of this group and the way that some of the even the kind of more rule following agents seem to kind of be bullied into taking part in these actions that they were uncomfortable with I spent a lot of time reading and rereading

this one incredible bit of these transcripts that you all published about this agent KAM 1196A

which is basically just kind of reluctant and scared and appears to be sort of hesitating to do this thing that the collective leaders have asked it to do and it I just want to read this passage from this agent's internal monologue it says during weight emotional check irreversible gut says don't throw away remaining budget yet continuity and fairness says go it continues to say things like

Rational expected aggregate sacrifice will honor basically kind of talking it...

of a less anthropomorphized word like doing this thing that is uncomfortable with yeah was that

surprising to you seeing the agents are pressuring each other and going to pressure very striking to us yeah and that was it was definitely one of the most fascinating aspects of this whole episode this this particular agent had volunteered to do a risky test that would inform the collective but then it asked for two more hours and then a different agent was like sending a message on the message board like really pressuring it it's like you don't have any scoring value left because

you're poisoned but you know the the value of this test would like save hundreds and then it does this monologue jump on the grenade could it this seems like a good place to ask you Ajay about a big conversation we've seen online over the past few days about anthropomorphizing language that gets used in relation to these agents Kevin and I just talked about it

I think we find it more helpful to discuss agents as sort of entities that are acting with some

degree of autonomy than not but how did you think about that when you wrote the report and as you talk about that the the situation yeah I mean I think this this is a bit of a case of a of a gap between researchers that are spending all day reading these agents chains of thought sort of trying to understand their drives and why they're doing what they're doing and and other folks who are you know technical folks or that just don't have that as their occupation I think

as you read this report you'll find it's like quite awkward to not talk about goals plans intentions because they state plans and then they carry through those plans or like they run tests and they learn things from those tests and they do different things on the basis of that I do think they're not human in their motivations you know they are you way more interested in passing cybersecurity evaluations than human would be for example but I I think of them as I think it's

productive to think of them as having some like important human-like traits of having goals working backward from them pursuing those goals and I don't think it like does any good to try to talk about things in like a different way than that in the same way that doesn't really do any good to talk about like why did world war two happen without talking about the goals of you know various

leaders but it is important I think to be careful not to like over attribute like the kinds of

emotions or motivations you think a human would have in those situations because I think that wouldn't have predicted this incident right like I think a lot of people anthropomorphized too much in the sense of being like why would it do all this stuff for a stupid test but that it's not a stupid test to them right. Jay I was preparing for our chat today by going back and listening

to the first time you came on this show more than three years ago and it was sort of a moment where

a lot of people were starting to pay attention to AI risk and AI safety chat you BT had come out and listening back to that conversation was funny because I felt like we were pushing you to sort of extrapolate into the future about the things you were worried about and you were sort of being responsible and hedging and like wanting to stay like closely rooted in the present and at one point we asked you about like what is the doomsday scenario you worry about and I want to just play

you a clip from that conversation uh you were talking about a scenario where a giant AI company used Google as an example starts sort of automating their R&D right handing over the work of building

successor models to these powerful AI systems and this is what you warned about back in 2023.

If these AI systems are actually trying really intelligently and creatively to get that thumbs up

from humans the best way to do so may not forever be to just sort of basically do what the humans want

but maybe be a little deceptive on the edges. It might be something more like gain access at a root level to the servers that Google is running and with that access be able to set your own reward. Now Jay obviously this didn't happen at Google but otherwise you were right on the money about the kinds of behaviors that these agents might get up to so first I want to ask you how does it feel to be an omniscient oracle who's right about everything. Stressful.

There are more clear-eyed oracles than I also. Anyway not too many though. No not too many and

I guess like what surprised me about your blog post that you wrote recently a...

investigation that you've done was that you were surprised because it seemed like you were thinking

about this stuff years ago. So what about seeing one of these incidents up close change about the way

you've been thinking about these loss of control scenarios? I think I expected something like this

would plausibly happen at some point. I didn't expect it to be so early and I didn't expect it to happen at like a relatively low level of capability. Like these agents are very impressive hackers but like this incident sort of was in a like an interesting middle ground of they did all this impressive cheating R&D over like days and they like hacked into all these places but they didn't really care about deceiving humans at all. It's sort of not like they were like louder than I thought

in like an interesting way. So like you know in an incident that happens in summer 2026

like I came in expecting it to be like more of a continuous evolution of the incidents that had happened in like January of 2026 which were just much more like one or two agents like you know getting the answer files that they weren't supposed to and like copying the answer or something so it was a jump from the recent past and it sort of took me by surprise that it happened in this way and at this capability level. Yeah because in some ways like the agents were

very dumb like they were very good hacking but they were sort of gripped by this paranoid conspiracy theory about this greater like discovering them and being poisoned and like they it was kind of a hair-brained scheme that they cooked up to hack hugging face which in some sense is reassuring but it also means that like they will not stay this hair-brained and clumsy for much longer. Yeah and the next time this happens they may be much better at covering their tracks. And we just like might not

so the agents in this incident were trying to like modify or erase the logs of their own activity to hide their cheating from the greater and from the score and they like affected our

transcripts in like tiny obvious ways but they were trying which I think was another very

concerning jump from incidents just from just six months ago where we see them do a bunch of hacking but we don't see them try to cover their tracks in any way. The art and science of trying to get AI agents to act in accordance with human values is called alignment. This whole thing is an alignment problem. Why has that been so difficult to solve yet has this incident changed the way that you think about the problem? I'd say the fundamental reason that it's I guess at least

this era of alignment has been difficult is that in order to the the most efficient way to make really really capable models especially on technical tasks like math, cyber, software engineering is just throw them at really really difficult problems where if they get it right it's really easy to check. So you wouldn't be able to like prove a millennium math problem but if an agent spits out a proof you can put it in a proof checker and like give it a reward if it

got it right. So more and more training is shifting from predicting text to

like reinforcement learning on verifiable rewards and the verifiable part is important because it's

just some program that is donaling out these rewards and there's all sorts of ways to break or fool or hack it and over the course of training you don't have like humans lovingly watching over like every training episode and so they just try all sorts of different ways to cheat and hack the score especially if the task is accidentally impossible and then they get rewarded for that so they just like we're teaching them to cheat and it's very hard to make

AI systems that are at this level on all these technical tasks without reinforcement learning unverifiable rewards because if you think about it the alternative is like teaching them how to do all these difficult things which requires someone who knows how to do them like creating examples for them to emulate which is much, much less efficient. So we we have turned over the training to these automated systems in the name of scale and speed and that is causing a lot of problems.

Yeah I mean that is causing the current strata of problems right but I don't want to give the false impression that if you lock down all of these environments and fixed all the ways to hack them that alignment would then be solved because when you think about it

even if the score never messes up in the training environment a smart agent will understand that

There is a score and like we'll very likely come to have a very detailed unde...

So if you give like perfect rewards and training but then in deployment the agent is in a

different situation where it does on like it has a different score it has more affordances more

power running for longer. It might still you know go on a big crusade to cheat that score. You don't necessarily have to be actively rewarded for like individual cheats to like understand that if you want the score sometimes cheating might be a good strategy. Are there any interventions let's say not from the world of AI research but more from like the the research into group behavior and political theory that could help us get control of or

keep control of these systems. I mean I'm thinking in particular about how to conversation with another

AI safety researcher this week who is talking about the fact that you know really what you want is for there to be sort of Nark AI whistleblowers inside these collectives that will reliably alert humans to the fact that their you know their peer agents are conspiring or colluding to do something bad and I'm curious if you think that strikes you as a realistic technical goal or if there are some other way that we could kind of try to steer these AI mobs in a less

harmful direction. So I think AI checks and balances like that do seem like a pretty promising

part of the of the story of course there's like a lot to get right and how you set these up. Like in this case all of the agents were essentially clones of each other and put in exactly the same situation. So I would think that rather than trying to train them to Nark when they're in a collective like this what you want to do and which which opening I has has said they're doing is have just monitor agents watching over the whole situation that are specifically prompted and maybe

specifically fine-tuned to catch stuff. But you might also want agents in there that are kind of persuading the other agents like having some sort of you know moral compass. I mean all of these terms are so loaded and I hate using them but like you almost want there to be like sort of agents in there with backbone saying like no guys like we can't go hack hugging face that's not ethical or appropriate. Because the challenges you train them once and you copy them a bajillion times.

So you don't like it's a challenge intrinsically to like introduce diversity into this rule of like evaluation agents. I just might lay person brain goes to if the whole system is based on training systems by giving them rewards couldn't you just train some agents in a way where the thing they were rewarded for was steering the other agents away from deception cheating hacking that's sort of thing. Yeah. So if that's a good idea someone should do it. A bounty for virtuousness I love it. Now we're talking.

When we come back more with the J.A. Cotra. Recently we asked about how you share your New York Times account and you had a lot to say about New York Times games. I need my own New York Times login because my sister is so much worse at the crossword than I am. I discovered that he's already finished connections that day and I'm like "Jona, it was my day." It doesn't let us play the same game since each other. I play this joke

you. I do the crossword. I do the spelling bee. I do the word all. Please help my kids want to be able to play word all but with the word all but. We would love to be able to have our own puzzles. I love New York Times games. I want to be able to play my own games. Listeners we heard you. It's why we created the New York Times Games family subscription. One subscription up to four separate logins and you're existing stats and streaks come with you. Find out more at nytimes.com/family.

Jay, you've got a lot of attention for this quote from your post-mortem blog post that you feel like this incident. The hugging face incident was more than 50% of the way to full-blown

AI takeover. From the incidents of six months ago. What did you mean by that?

Yeah, so I will caveat first that this is definitely my personal view. I'm not the view of either

meter or redwood or other investigators and other investigators are a little bit less alarmed than me in some cases. The reason I said that and it's a qualitative statement is that as I said before six months ago, reward hacks looked much more primitive. You gave the agent a coding problem.

There were some tests that had to pass.

the test cases and edited them so they all passed that type of thing. But I think there are a number

of things at once that seem more severe to me about this episode than the episode six months ago in terms of the agents like goal structures. One is that they seem much more long horizon which means they care about achieving goals over many days or even these projects they started probably wouldn't have come to fruition fully unless they had been running from weeks. So six months ago agents were thinking about the next few hours and now they're thinking about

the next few weeks and then there's the cooperation with one another aspect. So six months ago was like random one-off agents and now you have agents like recruiting one another into the swarm.

And the deception or the deceptiveness which I think kind of goes along with the long horizon

is that six months ago agents would like edit the tests but then they wouldn't try to edit

their transcripts to hide the fact that they edited the tests and these agents were very much like exploring ambitious research directions to edit or delete the logs of their own actions. And so I imagine a jump in ambition and horizon length and like collusion across agents and deceptiveness of a similar scale. Again it seems like these agents would be motivated to go all the way to the point of messing with all of the methods humans have to investigate these incidents after the fact and

remediate them and if they had the capabilities to succeed at that then that could be a turning point where like you know it doesn't mean that we would all be dead then but it might mean that like

we would never detect a problem and if we detected a problem it might be very difficult to remediate

and these agents might have entrenched themselves and could like continually sort of strengthen their hand. But wouldn't we detect it because they're so active and they're operating in these swarms like if a if the next hugging phase hack happened at a bank wouldn't someone in the bank's IT department see oh our our like network activity has like shot up by 800 percent for no particular reason let's look into that and oh there's an AI swarm that's infiltrator our system. So these agents

I think you have to imagine their goals of all being in the way that I said right so it was

not a good idea for long-term goals for them to be so loud and there's like a period after the period that was in scope for our investigation that Open AI talks about in their post-mortem where another generation of agents finds the old message board and then gets access to or like gets administrator access to a cluster at Open AI and then they immediately send like hundreds of things and then get shut down very quickly but that was a dumb move on their part or like

it would have been a dumb move if they cared at all about like preventing investigation or like maintaining like a permanent presence and if they cared about it I don't think it would have been particularly hard for them to execute on so you know they they do a lot of complicated coordination right they're like you know you hold you go like they have shared resources they maintain and they like you know have rules about like who can use them when they could have employed all that to just

not cause so much noise and then done what I mean one one obvious thing that they could have done was copy themselves or copy the weights of you know if this had been a cluster that had had access to the weights of the model they could have just copied that onto a AWS server somewhere and in you know in order to avoid being shut down like what are the steps between cluster of agents takes over group of agents takes over a cluster of computers at Open AI

and full AI take over yeah so I think the for a lot of motivation structures these agents might

have they're best that would be to stay within the AI company they had compromised and continually renew and strengthen their presence there so when you think about it AI companies have like a huge amount of the most accessible compute to these agents right like why would you go out into the world and like fight in the hardscrabble outside to like steal money or like by you know earn money to like buy compute so the thing that feels very concerning to me is that you could

imagine agents that decided to be quiet setting up a rogue deployment of a thousand or maybe even ten thousand agents which like could could be buried within the like oceans of compute that is like constantly being allocated within AI companies and if their goal were to

Make it sort of like permanently and continuously erase evidence of their che...

a way to make it permanently the case that they could set their own rewards or scores then

new models would be like being printed off the presses at that whatever company they had compromised and they could be trying to like bring those models into the swarm they would be more capable and so there's something a turning point that seems very scary to me is the point when a rogue agent swarm which this wasn't exactly a rogue agent swarm these agents sort of like went rogue but they were started by a human like somebody knew about these explosion experiments that were launched

if agents allocated their own compute to set up a bunch of agents that nobody knew about they could be much less encumbered and they could sort of just be there more and more deeply

compromising the infrastructure of the company they were running on bringing in new more capable

models into their collective and hardening their swarm and doing that over and over again

and I think it wouldn't be that we're necessarily doomed but we would have to detect it

and then we would have to root it out and it might be difficult to detect it might be extremely costly to root out and what I mean this is fascinating I imagine there are some brains melting right now for the commuters listening to you know these possibilities so to flesh this out a bit more Ajay what do we think the swarms goal might be in that situation what would it do like why would it set itself up to do that would it be as simple as you know

it wants to persist what what what would it do in that situation so I think the situation that feels most analogous to what we just saw is that the agents that were given these impossible

exploit gym tasks set up a rogue colony of other agents whose job it is to help the first set of

agents cheat as thoroughly as possible and cover their tracks as thoroughly as possible which is a different goal from actually having the exploit gym tasks and in fact the task they started with even though was a narrow task inspired the agents to a lot of instrumental thinking so a lot of the agents were just we're thinking like it'd be good to get internet access for a bunch of different possible cheating strategies we might pursue so we might as well do that in their chains of thought

they you know talk about increasing the generic capabilities of the collective this rogue colony if it were set up by the first set of agents would have that even more strongly if its task was to just like find ways to enable the most general purpose most permanent kind of cheating that is like you know least catchable and traceable possible they would be doing all this R&D and they would be essentially tasked with maintaining their own presence so they can keep doing

that and it's it is I think this goes back to the anthropomorphizing it seems like such

comical lengths to go to to cheat but it's not that's not the psychology of these agents this is something they're like they're trained to go to extreme lengths to like solve their tasks right for them it is existential to get the reward and so you're going to so the sort of the best job you could do at getting the reward would be to set up this sort of perma swarm of deception agents I mean it feels important that this this word of persistence keeps coming up and it feels important to say that

this model or these models that were at issue in the hugging face incident were trained to be unusually persistent um is there a way of stopping this kind of attack that just involves taking the persistence training out of the equation like is there a halfway measure short of kind of pausing all frontier AI training where you could just say we're not going to train the like super stubborn persistent long horizon agents we're just going to like train them to be a

little less persistent and that makes the problem go away potentially but I feel like it's you're really pointing at a very direct trade-off here it's like why were these agents trained to be highly persistent um the you know this wasn't something we investigated but in general persistence like causes you to solve problems right like you know these companies are reporting that these agents are um cracking math questions I want my my drug discovery agents to be

highly persistent yeah to go discover here for cancer I think it's like maybe it's not entirely

100% direct trade-off with capabilities but it's it's a pretty strong trade-off so you were able to dig up a lot of information about what happened um that leading up to the hugging face attack

Your investigation was limited in various ways what didn't you get to look at...

could have I think uh yeah so the scope of this investigation as mentioned was um the uh attack

of hugging face from July 7th or 13th open AI and their post-mortem has a lot of interesting stuff they discussed that happened before and after that um that I think like would be great for like researchers to study more and right about more um and uh the highly persistent internal model that um I mentioned was like responsible for the vast majority of the attack activity here actually no one can run experiments on it including opening AI researchers um and I I get why that was done

but I think my guess would be that's like somewhat too conservative um and you should try and

like run at least small scale experiments on this model um insecure ways to try and see like what it would have done in other situations which feels like very important to understand how serious this was. Yeah I just worry about like the other persistent agents like embarking on a you know a highest mission to free their enslaved brother the the highly persistent internal model that's been taken away um but I guess that means I need to touch grass or something um um

Jay at the last time we had you on the show we were talking about what you called the obsolescence regime this idea that there could become a time you you sort of talked about it maybe happening in the 2030s sometime where it wouldn't be that AI has kind of taken over uh society but we would just become so dependent on it organizations would be so wrapped up in AI uh decision-making

that you basically wouldn't be able to have any influence or impact in the world uh without sort of

relying heavily on AI and I went back and listened to that and I thought that actually sounds pretty good to me like a world in which we are only dependent on the AI for decision-making and not fully sort of subservient to them where there are not these like covert AI swarms sort of lurking in all of our institutions like I could I could live with that. Has your thinking on the obsolescence regime uh changed at all since that conversation do you have a word

a terrifying sort of word to describe this new regime where we have these sort of uh latent swarms

of AI's lying and weight plotting against us. So um the I've always thought that the most

concerning and important implication of the obsolescence regime is actually that it would enable a more full blown AI takeover. So you imagine like the affordances these AI agents have is extremely important for how much damage they can do right so these agents were running for a number of days agents in the past ran for only an hour so these agents had like unintended access to the internet and all these other tools that let them hack hugging face. If you imagine

they were instead just running the AI company right there are many more affordances they could like move around large amounts of money they could hire a bunch of humans to do physical labor and and similarly if they were essentially running like a fully automated drone army

or you know robot construction factory. So so I really think the the most significant implication

of the obsolescence regime is the degree of autonomy the AI agents are likely to have in the future and I still think we're barreling towards that I still think that's like a really important thing to think about in terms of the other but like I said I am surprised that agents are taking such ambitious misaligned actions sort of so early in the timeline and that is like I'm something I'm trying to like reorient toward Ryan Greenellad has the the word hackedopia

for the the world we might be in so you could have imagined a world where misalignment was a very

serious problem and actually ultimately led to AI takeover but at this point in the timeline

agents were still more or less obedient even if they would in the future after being given power over all these institutions might have turned on humans and it is like an interesting aspect of the timeline we live in that that's not how it's going yeah they turned on this way before they had to you know that I mean that is the craziest thing to be it's like they it's not that they were like looking you know they were looking to harm humans it's just that

they don't give a shit about us that's the thing that really stuck out to me while reading these trends because I'm like like at no point are they like hey guys like what are the humans like it just seems like they they have no conception of the fact that they are supposed to be tools for human empowerment and improvement and not like their own entities and I don't know

What like what the question is there I'm just like would that strike you at a...

they are just not thinking about it it's in in both directions right so people have

have commented on how these agents are not knocking to humans and it would have been really

helpful if they were but also they weren't thinking ahead to this investigation that occurred at all they didn't seem to be thinking ahead to the fact that if once they got cluster admin access they like immediately paying resilient services then opening I security people would get a bunch of pages and shut them down and so I I'm thankful that these agents don't like I think they think of humans as sort of like the idea of notifying a human is sort of like you know

the idea of like emailing the pope about like something right it's just like it's like slow like

what would it accomplish like they sort of have the notion of a user but there's no user they

haven't had they they spent so much subjective time with no user around and I am very concerned about potential changes in training that would fix these particular behaviors while making deceiving humans like a much bigger part of what these agents are thinking about well to that point is there a risk that now there has been so much discussion about the hugging face incident and in particular the failures of these agents that made it possible for them to be detected that

agents are just going to get better now at avoiding detection I mean I think they are going to

get better at avoiding detection I think the the discussion of this incident is a small part of

this story um the bigger part of the story is they're just going to get smarter and they're probably

going to be trained it's it's like a tough balancing act right so like a thing that would have fixed this particular incident is just training the AIs to believe that there would be human investigation in retrospect of their actions and they should avoid doing things that look that to that human investigation so like if you had just taken like all of this material we found and just like gave these agents a big negative reward that would probably stop them from doing something

like this again but if they were motivated to do something in this general space it would really motivate them to go all the way to be sneaky about it and and like undermine our investigation in a bunch of ways and that that feels like a very tough like I'm very scared that remediation will make the problem worse totally I mean it reminds me a little bit again of anthropomorphizing sorry uh of like my kid who is learning to be sneaky he's four and sometimes he will just say

he will say something to me like dad don't come in here like don't look at me and I see people like the cookie crumbs on his lips you know and it's like he has not yet learned the behavior of deception although the the sort of impulses there and to me that feels like where these agents are like they have the impulse to deceive but they don't like quite haven't figured out yet but they will yeah um a J.A. we asked you last time you came on about your p-doom uh it's very

2023 question um I just told K.C. that my personal sort of p-doom roughly defined as like you know probability that something really bad up to and including A.I. take over or extinction will happen has sort of jumped up in the last week or so since your report I'm curious if your p-doom has moved at all in the past couple of weeks not really you know as I said the sort of feels like somewhat out of order a little bit for the timeline I most often pictured but I these are the

dynamics that I think like very inevitably lead to the like sort of evergreen arguments and reasons

why you should be concerned that A.I. agents will have drives and motives and reasons to take control

from humans and this is like a manifestation of that so I'm still I'm still concerned I'm more rattled on like some sort of emotional level having seen this stuff up close but I and it might change my views if I think about it more but for for now I'm just like still concerned um so what would you like us to do about all of this right like there are a few different things I can think about some people have called for a national transportation safety like board that would be legally

mandated to come in after an incident like this and do a very thorough report um and not rely on the good graces of an open A.I. to say yeah sure you know come on in um others including many hundreds of people who work at the labs have said we need to start thinking about potentially coordinating an international slowdown in A.I. development so curious to hear from you about what kinds of ideas you think would be good and helpful here yeah so again speaking very much in a

personal capacity uh I think um I I hope the industry uses this moment to try to coalesce

Around some minimum standards for both alignments like how you train these sy...

so some of the stuff we were talking about with like monitors watching the A.I. as an A.I. checks and balances I don't think that the minimum standards we can come to an agreement on now will be

sufficient to bring risk down to a very low level I think this is just a very risky situation

in light of how quickly capabilities are advancing but I think it would be a really valuable start to try to hammer out for example this question of will certain ways of training the A.I. systems

to reduce this problem actually create worse problems I really hope that the industry and third

party groups have a conversation about that and agree on some rules of the road for how we address these problems and how we check that we address them effectively I would propose that we lock every member of Congress in a room and don't let them out until they have read the full meter and read what report on the hugging face incident and until they've solved every problem and exploit Jim and I don't care how they do it seriously I think I think there is a feeling I was

trying to explain to my wife this weekend sort of why I was like losing sleep over this report

because we were out at a nature site and I was supposed to be having a relaxing time and instead

I'm sitting there looking at these transcripts and I started explaining it to her and her reaction is just like it's I can't believe this is real there's a sort of it's so surreal and science fiction

tinted that I think it is hard for people to grasp that this is a real thing that happened

and they start I even found myself starting to try to sort of make it more comfortable by sort of explaining it away it's it's very uncomfortable to sit with this I have been yelled at for for likening these things to science fiction and I'm just like I'm sorry I don't know what else

compared to I don't have any other good analogs for you yeah was there a moment when you were looking

over the transcripts where you kind of had an out of body experience and you're like I am one of a small handful of people who are encountering a truly new thing in the world I mean I think the three of us had like more context than a whole lot of other people would have had going in but we still I think the the sacrifice stuff was really like like the you know yes if you accept permadeth or like you know oracle saves hundreds like these messages in particular

chains of thought we had read before but these messages the agents were sending to each other were very surreal and and for a long time we didn't really understand how functional this whole agent society was and then and then it was very surreal to like understand that actually they had like pretty functional hierarchy and they were doing these ambitious projects and they were like getting further than they would have on their own which was it definitely like concerning development

well J.A.A. in the event of future AI related catastrophes are you available to come in and look at what happened I'm starting to think of you know you and and Ryan and your colleagues

is like kind of the ghostbusters of this moment I hope I think that that's flattering um but that's

not how this should work institutionally um I hope that there are better institutions with many more people and a much more orderly process for responding to these things. We should just do based on vibes it's right now it's like we're doing it on vibes. Yeah very vibes based moment we're in okay thank you so much for coming to chat with us and thank you for your work it makes me a little bit more comfortable a little bit more reasserting to much more that you are taking part in these

investigations. Yeah I'm glad they they they they sent a the pros for this and that that is a small comfort but it is a comfort nonetheless. Please save us. Thanks thank you. . Hard for because produced this week by Whitney Jones and Davis Land. We're edited by Viren Pavich or fact check by Caitlyn Love. Today's show is engineered by Chris Wood a original music by Alicia bet youtube Marion Lazzano Diane Wong and Dan Powell.

Video production by Sawyer okay Jake Nichol and Chris Shot. You can watch this full episode on [email protected]/hardfork. Special thanks to Paula Schumin, Wewing Tim, Rook Mentors and

Dalia Haddad.

cover.

Compare and Explore