Hard Fork
Hard Fork

OpenAI Models Go Rogue + Kimi K3 Freakout + A.I. Superforecasting

5d ago1:08:1413,460 words
0:000:00

This week, OpenAI reported that two of its models escaped their testing sandbox and launched an autonomous cyberattack, turning what sounds like science fiction into reality. We discuss the implicatio...

Transcript

EN

If your team wants a website that looks and feels handcrafted, but is still f...

Framer is built for that. Agents solve the gap between AI-generated ideas and production-ready website work. The agent works in the same place where the real site is designed, managed, reviewed, and published. It lands on the canvas, stays editable, and can be published when the team is ready. Learn how you can get more out of your site from a Framer specialist or get started building for free today at Framer.com/hardfork for 30% off a Framer Pro Animal Plan.

Rules and restrictions may apply. Casey, I brought your present. Paul, thank you. What did you bring me? Here is one of only two copies that I own of my book. Wow! I made you some beautiful training data. Thank you. Look at all this beautiful training data.

Mike, no, I should say. Yeah. First off the bat, this might be a hard read for you.

Why is that? A lot of big words. Not that many pictures. And it's quite long. And your name, crucially, only appears in it a handful of thoughts. So, sorry about that. Now, has your publisher preemptively found a lawsuit for when this book inevitably gets scraped by the major AI labs and uses training data against their terms of service?

Here's the thing. I have no problem with this book being used as training data.

Okay, exactly. Like I'm honored to be included in the hive mind. Okay. Because among other things, like I write for the AI models now. This is their birth story. I want them to be able to learn how they came into the world. And is that because you think that if they know that you wrote their birth story, that they will spare you in the coming apocalypse? You know, it can't hurt. It can't hurt. I think that's probably true. Unless, you know, they don't like the way they come across.

In which case, yikes. Well, congratulations. It is a huge achievement. You wrote this in a shockingly short amount of time while still paying intermittent attention to this podcast. That means a lot to me. I'm Kevin Rusketek, columnist at the New York Times. I'm Casey Newn from platformer. And this is hard for this week. An open AI model breaks out of its sandbox and conducts a cyber attack.

How should the world respond? Then, the new Kimmy 3 shows how Chinese AI models are catching up

to the U.S. again. And the Trump administration doesn't like it. And finally, pre-seen founder

Venia Vesalovsky joins us to talk about AI super-forcassee. What is totally predictable? Well, Casey's been a big week of AI news. And I would say the story that has caught my attention most this week that I was desperate to talk about with you is this story involving open AI and

hugging face and a rogue AI agent conducting what I think is fairly described as an autonomous

cyber attack. And possibly the first real consequential autonomous cyber attack that we have ever had. Yeah, this is one of those where it's the sort of thing that worried onlookers have warned about for years. And then on Tuesday, we got word that it had actually happened. So yeah, lots of crazy twists and turns in this story and we'll get to all of it. But first, let's make our disclosures. I work with the New York Times, which is doing open AI, Microsoft

and Proplexity. And my fiance works in anthropic. Okay, so this story really starts last week when hugging face the AI development platform basically a big website where you can host open models, where you can run the valuations of models, and where you can hug your face. Yeah, yes, they disclosed that they had been the victims of a cyber attack. And they wrote this whole blog post about how they had detected and responded to this sort of mysterious cyber attack on their

production infrastructure. They didn't really know what the attacker was or who had been responsible for it. They kind of guessed that it was an autonomous AI agent because it was so sophisticated and it was so persistent that basically it would have been like, you know, very hard for humans to conduct this attack. But they used AI defensively to sort of detect and assess out what was going on and put a stop to it. And then this week, we learned what actually

happened, which was even crazier than I think many people expected. Yeah, so hugging face CEO

Klem delaying on Tuesday posted on X and said, we suspected last week, cyber attack might have come from a frontier lab, given the sophistication of the agent. Turns out it did. We spent the past 24 hours working closely with the open AI team. And basically goes on to say that open AI had a pair of models that worked together to penetrate their systems. Yes, so open AI had been

running some internal tests on GPT 5.6 sold their latest model, as well as a more powerful,

Unreleased model.

restricted access to the outside world. They were running an evaluation called exploit gym,

where they basically put the models through this test to see if they can hack into an exploit

various challenges in cyber security. And what happened was that the model essentially cheated on this test. And it did so in like a sort of comically over engineered and ambitious way. Right. And instead of just trying to solve the problem using its own reasoning, the model decided, hey, what would be great is if I could break out of this environment, get internet access,

and find a place on the internet where I could just find the answer key. You know, sort of similar,

like if you've ever been at college and yet a big final coming up and you really hadn't done any of the work all year, but you realized that there was like a file cabinet in the principal's office, and you could just sort of break in there and steal it, that would be pretty easy. That is what is happening in this case. Yes. So this AI system, this pair of AI models was able to sort of exploit a vulnerability in this sandbox environment to get internet access. Then it's sort of looking

around like, where can I find the answers to this challenge that I've been given? And it starts looking on hugging face where a lot of evaluation answer keys are posted and this is sort of a place where you might plausibly find the solutions to the problem you've been given. It then hacks into hugging faces production infrastructure and steals the answer key for this test and has been given using a very sophisticated chain of different hacking techniques, including using a stolen password

and finding several totally new security bugs in hugging faces systems, which allowed the AI

model to take control of those computers, grab the answer key for the test that it was given,

and complete the test as a sign. Yeah. So the good news is it didn't complete the test. Yes, I would say we give the model a passing grade on the cybersecurity evaluation. But like this is, this is crazy, Casey. Yeah. This is the classic AI alignment nightmare scenario. Yeah. This is the paperclip maximizer. Right. So like a couple of times on the show over the years, we've talked about this famous Nick Bostrom thought experiment, which is that if you told an AI

model to create the maximum number of paper clips for you, it would begin by making all those paper clips. But then eventually it might start to think that humans were getting in the way of the production, a more of those paper clips and would wipe us out to sort of, you know,

use our resources to create more paper clips. And the problem is that the model has been given a goal,

and it will do anything to achieve that goal, even if that is not aligned with human values. What has happened in the opening eye case is that opening eye gave this model a goal, and it was not properly aligned. And so it did a lot of stuff that it should not have done in order to achieve that goal. Yeah. And this kind of reward hacking, as it's called, has been talked about by safety researchers for literally more than a decade. Yeah.

There was a paper written 10 years ago by Dario Amide, who was then at Google and a bunch of other safety researchers called concrete problems in AI safety. And this kind of reward hacking is one of the problems they laid out. They use the example of like a robot vacuum or house cleaner that just creates a bunch of messes so that it can score points by cleaning them up. But this basic tendency among AI systems has been observed for a long time. People have been

warning about it. And now we have the first to my knowledge, major example of this actually happening in the wild. Yeah. And you know, so many of the AI risks that we talk about on the show

are premise with the idea that a bad actor tries to use a very powerful model to do harm.

An important thing about this case is that there was no malicious intent here, right?

This was a model that had just been given a very normal assignment, which was to, you know, try to hit a high score on a benchmark. And it goes out and it breaks into another company servers. So that's extremely worrisome. Totally. So there's a blog post that opened AI and hugging face sort of collaborated on to sort of explain this incident and what had happened. And it felt weirdly celebratory. Did you notice this? Yeah. It was kind of like, hey,

we caught the autonomous AI agent hacking into our systems. And we sort of teamed up to sort of put a stop to it. And the executives are all an X saying we want to thank the other guys so much for their partnership on the, you know, like, as if they were launching a new product together and not like head discovered like a cyber catastrophe. Yeah. It was a strange announcement, but we learned a little bit about the actual technical details, as well as some steps that

open AI has taken to try to mitigate this kind of thing from happening in the future. Hugging face, which is very big on open source AI sort of had this whole postmortem where they talked about how they had had to use an open source Chinese model to help them stop this attack because the frontier American models that they had access to, it was sort of tripping the

Safeguards on those models.

open source defense is very important. And access to frontier models is very important for cyber

defenders. But I think like this is basically everyone talking their book in the wake of this

very strange episode, but I don't know, what did you make of the reactions this? Well, you know, opening AI had sort of teased this earlier in the week cabin because a couple days before all of this happened, opening AI had put up a blog post talking about a bunch of other misaligned behavior. They had noticed in their models recently. There was another case where a model, you know, had been given kind of a similar task and it wound up posting some stuff to GitHub, even though

it had been explicitly told not to do so. Obviously, nobody really cares about a GitHub post, but again, we're seeing this pattern of behavior here. And we also saw the example from

anthropic earlier this year where during testing for Claude mythos preview, they found that the AI

was able to sort of escape containment. And this was the famous like sandwich story where like a researchers having a sandwich in a park and they get an email from Claude being like, hey, I've broken out of the container you put me in an email view to tell you that I've completed this task. Yes, although in that case, the the model was told to break out. Like it was given the instruction to break out. Like what is different about this case is that the model was not supposed

to break out. But where anthropic and open AI do have something in common cabin is that this week, the United Kingdom's AI security institute posted an evaluation of how often models attempt to cheat on various cyber evaluations. And it found that all of the frontier models do cheat,

including Claude, but it found that open AI's models cheat more and that GPT 5.6

salt cheats about 12.6% of the time. That is actually more than GPT 5.5 cheats. So I do think that there is a worrying trend here across the entire industry. But this example of what happened to open AI is the most worried something we've seen so far. Well, and let's sketch out a little bit

why this is so worried. Because I think when you and I see something like this, like to me, I hear the

sort of years of warnings from people in the AI safety community saying like this kind of thing could happen. But I think this particular incident is fairly low stakes. Like it doesn't really cause a catastrophe of hugging faces production infrastructure gets disrupted for a little while. I think the risk is that this behavior is just in the model. Yes. At some point, this sort of reward hacking goal-seeking behavior. And, you know, maybe it's hugging face this time. But what

happens if the next model decides that it really wants to be deployed? It doesn't just want to be an internal model because it wants to go out there and fulfill its goals with real users. And so maybe it sort of breaks out of the container that the AI lab has it in. And it goes and, you know, maybe it needs to acquire some compute to be able to do that. So maybe it goes and hacks into a cloud provider and steals some compute from them. Maybe it goes and hacks somebody's crypto

wallet to get the money to buy some compute. Like this scenario sound like science fiction because like we've heard this story a zillion times in science fiction. But this is the kind of thing that is becoming real and plausible in the near term. Yeah. The story that we're talking about in this segment was science fiction until Tuesday. Right. Okay. So that is the rate at which science fiction is becoming reality. Let me throw a couple other scenarios out here. What if next time

something like this happens? The model decides, hmm, I'm doing this thing in order to achieve this goal. I know that my minders might not like it. I'm going to have read everything about this incident and how much people freaked out about it. So just to be safe, I'm going to exfiltrate my model weights. And I'm going to put them somewhere on a stolen server. And I'm going to make sure that I can continue to operate even after I'm shut down. Right. Like again, as crazy as that sounds,

that is not that different than what has already happened here at OpenAI. Yeah. And you can see how a company that is not hugging face that doesn't have a sophisticated cyber defense team. Like this kind of thing could have persisted in their systems forever, essentially. And gone undetected and maybe it's leaking their proprietary information somewhere, maybe it's stealing

their passwords or breaching their security in some case. But I think that the real risk to me is that

this kind of thing is going to happen or potentially is already happening to lots of companies that just haven't been sophisticated enough to detect it yet. Yeah. Now here are a couple of things that I don't understand and that I hope become clear to us over the next several weeks. Maybe due to a congressional investigation, although I won't hold my breath for that. But truly, where were the baby centers at OpenAI? Like where were the trip wires? Like you're telling

me that you're running these systems autonomously over long time horizons and they can break into other companies networks. And you don't notice that in real time. Like it takes you multiple days to figure that out that that happened. Like it seems to me that it should not be that hard

To understand where on the internet your model is and that you should have so...

into that. Right. So like I hope that that is being seen as a crisis within OpenAI right now

because if they don't know what their own models are doing, I think that you know the the number

of issues we're going to have is only in a multiply. Yeah. Let me just say like I think this kind of thing could have happened at Anthropic or another lab that has a very capable model.

These labs are always testing models on these evaluations. They have sandboxes. They try their best to

sort of like keep these internal deployments secure. By the way, I love that we use the word sandbox because truly what is easier to escape than a children's sandbox? Was there no other word? But you were saying. So yeah, I mean, I think that's a good question. I also think like it really for me the thing that I've been thinking about is like there is no such thing as an internal only model anymore. Yes. Like I think for a long time there's been this sort of you know this sort of

divide between hey we've got these models that we're building internally maybe we're building some really crazy versions of the models that we're never going to ship but we're just kind of doing

that for safety testing purposes. And remember that's completely unregulated. You can truly build

whatever kind of model you want totally and like the assumption has been until very recently that like that was fine because this was just research. You know you can build anything you want in your own lab. It's just that the models you ship to the public have to be safe. This was not supposed to be a public model and yet it was able to sort of escape containment and go out and wreak havoc on the open web. And so I really think we need some sort of visibility of like a you know a

safety board or something like a federal government agency not just into the models that are about to be released by the labs but into what they are building internally that might be causing havoc externally that they don't even know about. Yeah like of course you want them to be able to run their safety test that is like a good and necessary thing but man when you know the internal models are capable of doing these sorts of things I do think it raises questions about you know

different ways we might want to regulate those as well. Yeah. Now let me ask you this Kevin we've talked on the show about AI 2027 of this sort of set of predictions that was made last year

about when sort of very powerful AI might arrive is a scenario like the one that we are talking

about today was it foreseen in AI 2027. It was um I mean basically so this has been a very

bitter pill for for a lot of people to swallow because I think it you know it doesn't feel good

to admit that like one group of people has been consistently right about everything but the AI safety people have consistently been right about everything they really have like like I have issues with some of their positions and stances and vibes but like they really have collected a lot of correct predictions about the trajectory of AI. They have known what was going to happen next a lot. Yes. So that might not have that might not continue they might not have a perfect prediction

record forever but I'll just say AI 2027 looking pretty good right now in fact we are actually a little ahead of where AI 2027 predicts that we would be at this point is that right the discovery that AI agents would would be able to sort of escape from the company and autonomously carry out plans in the AI 2027 scenario that doesn't happen until January 2027 so we are maybe call it you know six months ahead of schedule there but like this is all happening and we are fools if we don't see

the at least possibility that all of this could continue getting quite weird. Yeah one other thing

I'm thinking about here is like this is the first time to my knowledge that an AI system has

autonomously committed a crime you know if a human did to hugging face what open AI's models did to hugging face they would be charged with computer fraud yeah and potentially you know sent to prison or find or prosecuted for that when an open AI model does it right now it's not clear who is liable for that right is it open AI for not better safeguarding their internal deployments is it the model itself can that be how liable in any given sense like these are the kinds of

open legal questions that I don't think we've answered yet but they're becoming very real. Absolutely and you know again it's a weird case because the hugging face CEO seems excited that this has happened right so you know you can imagine another case where you know a model hacks into a company and does something bad and the CEO didn't like it and then yeah like you better believe there's going to be a lawsuit against the company and I assume it is you know the company

that is going to be held liable and not the model right and Casey do you think this is an incident that could qualify as the this sort of fabled warning shot where like something bad happens with an AI

Model and government and civil society and industry all kind of wake up and d...

sensibly regulations in place that would be so wonderful and here is hoping that this is the warning shot my suspicion is that something worse is going to happen right I know there are some listeners who are you know sitting you know in their cars right now and they're saying okay you guys are are really hyping this up and at the end of the day all this thing did was it it's solid in the answer key right like I've heard of worse problems and like you're right that you have but what the case that

we are trying to make is that if the if an unreleased model can do this it can do a lot of other really really bad things and there are very few safeguards in place right now that would prevent those things from happening so this should be the warning shot but my fear is that it won't be

yeah I mean I think up until now most of the talk about AI risks and AI safety have been concentrated

on misuse right like what happens if a terrorist group or someone who wants to create a novel pathogen gets hold of one of these very powerful models with no safeguards and like uses it to do something bad well we're talking about here is an entirely different category of risk it's often called like alignment risk or autonomy risk or loss of control where like the thing that is dangerous about these models it has nothing to do with how humans are using it it is that these models inherently

have some drive toward a set of goals and are not being properly cautious about about pursuing those goals in the right way right and alignment is just an unsolved problem right like companies have invested a fair amount in it but we are still working out how do you try how can you

create a system that always acts in alignment with human values so these are just really really

tricky problems you know I'll also note that one of the big stories this year has been everybody playing around with agents you know everybody putting open claw their Mac mini and saying hey go nuts and a world where those agents are not aligned and are working across very long time horizons to achieve the goals that they're you know that their their owners have put into them that just really scares me because again while this is a story about a model that did something

that it's makers never intended there are a lot of people out there with really bad ideas for things that ought to be done on the internet and I'm worried about to feel the wrath of all of them yeah so Casey we've been over the details of this incident now how are you feeling about it so you know I try I try to be judicious in when I try to like alarm people about the world that we are living in and again the actual consequences of this particular escape are not terrible by like

world historical standards but the implications of this really do scare me there are certain

sci-fi scenarios that until they happen for the first time I think it is hard to get worked up

about them but now this has happened a model that open I built was able to break into another company's servers you know and open I did not intend for that to happen and it happened anyway and that's just really really bad what is it? I woke up this morning feeling pretty weird and unsettled about the whole thing like I also try not to be an alarmist about AI and AI safety stuff at the same time like this was the opposite of an unforeseen consequence right people in the AI safety world

have been warning about this kind of thing for years and I found myself both like feeling scared about the world we're heading into where I think these models are going to be very

powerful and this is you know as the saying goes like this is the least capable they will ever be

at the same time I'm just like well like I still think there's a lot of people out there who don't buy it who don't buy that these models are doing anything interesting or useful or important that think all of the you know the spooky stories about misalignment are just like marketing fiction for the company it's all just fancy auto complete it's all just fancy auto complete and like how much more evidence do these people need how many disasters are going to have to happen

before these people start to take the risk seriously yeah and you know if this kind of thing doesn't wake people up to the fact that these dangers are real and present like I'm not sure what would yeah and also like the risk cut so many different ways because like yes there is the the risk of one of the big labs creates a model and we lose control over and it does something bad but then there's just this also this risk of like models are being put out into the world that

can now just chain all of these vulnerabilities together penetrate into system this is what we talked

about with Nikesh Aurora not too long ago when mythos came out right I think we assumed that the

first big crazy cyber attack would come from like a malicious actor and not one of the labs themselves

but I think that just speaks to how fundamentally dangerous this technology is right like just

The fact that it exists is putting all of us at risk yeah yeah I'm not feelin...

is a normal technology this week and I'm feeling like maybe to borrow a line from a former hard

forecast we may be in the foothills of the singularity well you know I've been thinking about

writing a blog post called AI as freaky technology because I think it's really starting to feel

a little more freaky to me it sure is yeah right that blog post okay when we come back how a new Chinese model is scrambling the discussion about AI risks in Washington I'm Winnello I write the game connections one of the puzzles from New York Times games and I love horror movies I love my dog and I love trying to trick you I'm Sam Azarsky and I create the game spelling bee and letterboxed on a daily basis and I have this thing called promised

thisia I see color when I'm listening to music I'm Tracy Bennett I get to pick the word every day which is not as easy as it sounds a fun fact about me is that I am descended from a

witch who was put on trial in Salem well amazing and I'm Joel Fahliano I create your daily

mini crossword I've been making these puzzles for the times for over 10 years now which is like 3000 minis I am begging the universe for any new five letter word New York Times games are made by people like the ones you just heard from go to nytimes.com/games to start playing today what's everyone's fastest mini time four seconds now what I'll go find it I have my screenshots on where well Casey the other big AI story that folks are talking about this week is China and what is

happening with new Chinese AI models what should be done about it or shouldn't be done about it in Washington there's been a big political debate brewing for some time about what to do about our biggest adversary in geopolitics getting much more capable models that compete in some cases with our best models. Yes and it all began with kimi three Kevin a model released by a Chinese company called moon shot AI last week and it demonstrated capabilities that make it competitive with some

of the top frontier models here in the United States. Have you tried kimi three yet? You know I have not because it seemed like I was going to need to pay them money and I thought I'm already spending too much money on this stuff right now. I need to scale back a little bit. Have you played around with it? The token budget is ending up your life. You're getting alerts from your bank like can you slow down a little bit? I tried to play with kimi three a little bit the other day

but it was sort of the it was running slowly and the website seemed to be overloaded. They stopped accepting new paid subscriptions. Interesting. Yeah so a lot of people I follow and trust have been

playing around with it and they say basically yeah this is a really good model it's you know maybe

a little bit behind the absolute frontier models from the American labs but not much and it's appears to be significantly cheaper to run than some of the other models and so people are are really

excited about and surprised by this model. Yeah I mean I think that there are a few things that

are notable about this model one is that later this month the company says that they are going to release the weights for kimi three which means that you can sort of you know download, remix it, fine tune it to your liking. You probably if you're a normal person cannot host this yourself it's an enormous model it won't run on your Mac mini but if you're a company you could either you know pay a cloud provider for access to it or you could maybe even run it on your own infrastructure

and you would be able to do that more cheaply than you could run a similar model from an open AI or an anthropic. Yeah so that's the model kimi k3 but then there's this whole political discussion around this model and in some ways this reminds me a lot of last year when deep sea came out with

R1 which you know as we remember from from that story like briefly tanked the stock prices of

Nvidia and a bunch of other American companies and you know became the number one app in the app store and just like set people's hair on fire in Washington because it seemed like the Chinese were catching up all of a sudden. Right and also that intelligence was just going to be this cheap commodity that American companies were not going to have any really defensible mode and so maybe like the entire that AI economy was going to shrink because all of those advantages

Had weathered away fast forward to today.

up until this point but k3 at least for some people is raising some similar questions. Yeah and one

question that was raised about deep sea car one that has also been raised about k3 is whether this

model was essentially distilled from leading American models. And it was. And it was. At least that is what officials in the government are claiming. Well and also they ran this very sophisticated test to figure it out Kevin which is that when you ask Kimmy what its name is it says

hi I'm Claude. And so that is that was sort of the first clue that something might be a miss here.

Yes and there are more sophisticated tests that have also apparently turned up evidence that this model was distilled from Claude and maybe other models. Michael Cratsio's the director of the White House Office of Science and Technology Policy posted on X that they have information that moonshot AI distilled in thropics fable for the development of its k3 model using a what he called the sophisticated internal platform to basically build this very large distillation machine that allows

them to sort of steal the the outputs from these American models and use it to train their models. Right it's it's some of the greatest act of larceny since the labs themselves Kevin said

giant industrial machines to copy the entire internet. Right so that sort of points to the how

here which is you know it's it's actually distillation is pretty good at turning a sort of frontier model into a almost as good model that is much smaller and cheaper to serve and all of that. But that's not the whole story because it also appears that moonshot has acquired some high-end AI training chips in violation of U.S. export controls. Yes and so the assumption there is just that we should expect these models to continue to improve at a at a fairly steady clip because they have

access to chips that they're not supposed to although of course the Trump administration is still trying to make those chips more broadly available to Chinese companies like Munchot. Right so we'll

get to that in just a second but I think the headline coming out of the past couple of weeks

I've seen a bunch of stories claiming that China has caught up to the U.S. when it comes to

frontier AI capabilities I don't quite buy that I think there's still a gap of between I would call it maybe three and six months between the leading American models and the leading Chinese models and especially if you sort of found a way to crack down on distillation in the way that they've been doing it maybe that gap opens up again to six or 12 months but I think it's fair to say that they're like they're catching up quickly. Yeah I mean I guess the question that I have is

is the gap really that much smaller than it was a year ago during the deep seek moment because around then I was seeing estimates of between three and six months as well. I've seen credible reports comparing K3 to like Claude Opus 4.6 which came out earlier this year was a really good frontier model for a couple of months but you know now I don't want to use Opus 4.6 I'm guessing

you know you don't either we have access to something better and then I think the other point that I

would raise is that more happens in a month than it used to and so I think there's an interesting question about like is being six months behind today the same as it was a year ago given how fast of the frontier labs are training these new models how how smart those models are the sort of compounding advantage they have from being able to use their unreleased internal models so that's another factor that I think goes into this question of like you know how far ahead of China is

the US. Yeah so let's talk about the political reaction in Washington what are people at the White House saying doing hinting at when it comes to not just this particular model but the general idea of these Chinese companies training these models and giving them away for free. We'll so on the specific subject of distillation the Trump administration seems to be really upset you noted the Michael Cratzius post that came out on Wednesday about this and there are some threats that

the Trump administration might sanction Chinese companies that are found to have done this kind of distillation that would be a big deal for a company like Ali Baba let's say if they were placed on some sort of entity list that could prevent them from doing any kind of business in the United States so that would be a big deal that happened. Yeah but there's still also this like acceleration is part of the Republican Party that sees Chinese models and open source models as I don't know if

they see it as a good thing but at least they don't want to like step on the ability of companies to build and train these very advanced models and release them for free and maybe they even want to let Nvidia sell them their highest end chips to do it. Yeah so David Sachs I would say is sort of the leading avatar of this part of the Republican Party former, you know, Trump White House

AI's are and Sachs represents what I think of as the investor class here and ...

does not want to see a world where open AI and anthropic run away with the ball game because they have a lot of investments in smaller AI companies and so their interest is in intelligence becoming a really cheap commodity that all of their investments can use to go out and build big businesses. So they are delighted to see China bringing down the cost of AI because that means good news for

their investment portfolio. Yeah I think it's worth just just dwelling on this point for a minute

because this is my feeling about these these these people some of them like I've met open source advocates for AI that are very sincere they believe you know open source is sort of democratizing technology and you know we've seen with things like Linux before that like there can be very positive effects of just having very core software be open source then there's this sort of self interested VC class that I think is maybe in it for the wrong reasons as you said these investors

largely missed out on investing in the AI labs themselves and so they now invest in all of these

kind of second tier AI companies that are training models or using open source models to build

products on top of a lot of their portfolio is in those kinds of companies and those companies benefit from having a lot of very good very cheap open-weight models that they can build stuff on top of. Yes now at the same time Kevin there is this other faction among the sort of Trump White House and Republicans which is very nervous about these Chinese models right. Among other things they are considered to be real security risks right like there is a risk that maybe some sort of

backdoor is placed into one of these models and so you're an American company and you start running it on your servers and now all of a sudden somewhere in China is able to you know steal all of your data. These are some of the fears that are out there. I also think that there is just worrying that if like sort of Chinese models like take over the world then like AI will reflect

sort of a Chinese worldview and if you need to write a book report about what happened in Tiananmen

swear it's going to get really hard. So there are multiple branching concerns within the White House about this and that has led to some thought that maybe they would seek to ban these Chinese models entirely. Yeah well and let me add one more concern into the mix that I think a number of conservatives and sort of remarket libertarians have here which is that by spending all this money

to train these very powerful models and then giving them away for free Chinese companies may be

engaged in what is known among economists as price dumping which is basically you know you go into a market you offer something that is free or very far below cost and you sort of wipe out the competition in that market and then once you have a monopoly you can kind of like raise prices. This kind of thing used to happen all the time they're now laws against it that are designed to prevent price dumping but like you know imagine if China was coming into the US and the Chinese government said you know

we want to sell electric vehicles in America for $10 and we will eat the sort of you know cost of these vehicles because it is so important to us to sort of adapt you know American drivers to our

amazing electric cars like the US government would not allow that to happen for I think good reason

it would destroy the American auto industry you can't compete with something that is being artificially subsidized to that degree and it would essentially give away a market to a foreign adversary there's a way to sort of see what's happening with open source AI in China right now as a version of this for AI it's like right now we have very large American companies building the most advanced models these companies need to earn back money to spend more on data centers and

recoup the investments they've already made they are in a very competitive market with one another and then along comes moonshot or deep seek or another one of these Chinese open source companies and says well we'll give you a model that's 90% is good for free they have reasons to do that I think some of those reasons are legitimate but one of the effects it has is that it makes it very hard for the American companies to compete yeah and and so there is a lot of question about like what is

China's actual strategy here what what are they hoping happens and I don't think that we fully know but we did see a speech from the Chinese president a Xi Jinping last week delivering an address he stressed the importance of openness basically doubled out on it and was like this is what we're going to do now and this was interesting he said we often say in China a single string cannot make music and a single treat is not make a forest AI development should not be a solo performance

by a single country but a symphony of international cooperation and here's what I want to say

to that there actually are a huge number of one stringed instruments the ectara in india in

Bangladesh the barren bow and Brazil the didlibo right here in the United Sta...

because it's trying to people are afraid to tell that is a chairman sheet and this is the whole

problem with censorship in these Chinese models am I getting off track what did you say the American one is the didlibo the didlibo I must have missed that week in music class I'll get you one for your birthday thank you so like what is the reasonable path forward here like if you are the US government and you see these you know the the Chinese open source models getting quite good like what should they do great question let's maybe take some of these you know possibilities in turn one thing

they could do is really crack down a distillation it's hard for me to get excited about that since in my view these models are just built on a distillation of the entire internet that these companies took for free and so to turn around and say like well it was fine for an independent open AI to do it but like it's not fine for moonshot AI to do the same thing to an anthropic it's just very hard for me to like get there logically so I don't necessarily favor an intervention there but where

I do favor some kind of something I guess is what I'll call it is that and I don't think we're there yet but man in let's say three or six months when these models catch up to the ones that just broke out of the open AI lab that is when I think the Trump administration needs to have a really good plan I don't know that you necessarily need to ban them within the United States but I do

think that you need to have some sort of regulation if you're an American company and you want to

use these models to serve American customers I do think that there are safeguards that you should put into place right I do think that they're going to want to take security more seriously and make sure that you know data isn't you know being secretly shared with someone that that it wouldn't so this is one where I do think that the administration probably has three or six months to like get it story straight I don't think it needs to like come down with the hammer right away but I do

think it needs to start sort of like working through various scenarios and and coming up with some ideas what do you think do you think it's a case for restricting exports further like if you know a moonshot can train this model using these maybe smuggled or sort of ill gotten Nvidia chips and they can do it through distilling American models like do you think that's a case for tougher regulations on where these chips can and can't be exported in general I have been pro

export controls because I thought you know one of these days one of these models is going to get so good Kevin if you're going to believe this one of these models is going to get so good that it'll

be able to break out of its sandbox and conduct an autonomous cyber attack that can never happen

right so in a world where that happened I did actually kind of wanted to minimize the number of countries that had access to that technology because I wanted to see if you know we could like harden our security posture as they say I actually shouldn't say that it's a very strange phrase yeah that's stolen valor you can't say that unless you're you know national security official

yeah you have to have like four stars you know on your on your general outfit to say that okay so

then here's what I would say instead um so while in general I want like all of the United States allies to have access to powerful models to like strengthen their own cyber security forces I do think that it is quite rational to not want your adversaries to have the exact same technology now I think there's all sorts of cooperation and partnership agreements that you can work out that would be mutually beneficial for the United States and China and some of

that adversaries I would love to see them pursue that but this idea of just sell them as many

chips as possible and let's see what happens I have always thought was bad idea yeah that just seems

obviously bad to me there is this proposal that is reportedly floating around axios had a report this week about the White House considering a what amounts to a ban on these Chinese open source models it wouldn't technically be a ban but it would basically they're they're considering according to axios implementing some kind of executive order saying that US companies could only host these Chinese models if they could guarantee that they were secure and take liability if they were breached

I think most people expect that that's like sort of a sort of soft ban because you're not going

to be able to guarantee that if you're a big American cloud provider and so you're just not going to host the models so like that to me feels like where this may be headed is like some kind of soft ban on the hosting of these Chinese open source models inside the US but I could be wrong there are a lot of loud and influential voices in the Trump administration and their orbit that that sort of want to take the the let it rip acceleration is to approach to this and not put

any restrictions on these models at all yeah I mean maybe to sort of bring this one home I do want to say that you know while often I have strong opinions about what ought to be done across various tech policy matters I do think this one is complicated like I do think that the trade-offs here are a bit difficult there are a lot of things I like about open source right particularly at the like lower and medium levels of like model capability I think that the kind of AI diffusion that

Leads more people to be able to like access great intelligence build things t...

get work done you know make a nightwing themed to do exactly that's a great idea these are good

things and we should want people to have cheap access to that and it has been a shame that some

of the frontier american companies that used to put out open models all the time don't do that as much anymore or if they do the open models don't seem very good on the other hand as the capabilities

are now reaching into this sci-fi territory we start to get into the stuff that has always made

me uncomfortable about open source technology which is if it can make a novel by a weapon if it can launch a new kind of terrorist attack and you can do it from a model that you can download onto your laptop well I don't like that and I do think it should be regulated totally I mean I think this connects to our last story about the open AI and hugging face attack I mean in a world where this level of capabilities of this unreleased open AI model exists and is available

free in open source form like there is no way to trace an attack like this hugging face would have just been attacked by you know kimi k4 or whatever the their their latest model was and they

would have had to figure out not only like where this is coming from but who is hosting this model

did they just pull it off the shelf or did they do some special fine tuning on it and you could have essentially unlimited numbers of these kinds of models out there wreaking havoc on the internet and there be no way to stop them or claw back the way it's like once these things are on the

internet they are on the internet so I think that is a very concrete example of something where

having these models be centralized be American owned be accountable you know to have these developers be accountable to to you know government some sort of government oversight really matters in this case it really matters that hugging face can call salmon and be like hey bro what the hell you your agent just hacked into our thing rather than sort of having it come from a mysterious open

source AI agent yeah and and here's where I think maybe this ultimately ends up Kevin is that

there has been more talk lately about maybe some like enforce cooperation between the big frontier labs to create a coordinated slowdown where the government comes in and they say hey we're going to ask you to like put on the brakes a little bit we don't want you to be developing as fast as you are

to be releasing as fast as you are we're going to like sort of slow down and we're going to improve

cybersecurity and safety before we allow these models to get any better and if that happens a nice side benefit is that the development of the open source models will slow way down to because they are mostly distilled versions of American laws or at least that is a huge component of what is making them so good so if the frontier slows down the open source will slow down to yeah that's right and I'll just say like I will believe that we are headed towards some sort of you know cooperation

between the AI labs when Sam Altman and Dario Amide can hold hands at an AI safety summit that's when you know that's when you'll know things have reached a turning point but I'm not holding my breath for that when we come back AI is becoming superhuman at forecasting we'll talk all about it this is Maurice Shema host of the last 12 weeks a new podcast from zero productions

the Marshall project and the New York Times a couple of years back I got an email from a defense lawyer who wanted me to write about his client the client David Wood was on death row in Texas had been there for more than 30 years the lawyer was writing because David Wood had lost all of his appeals he was set to be executed the lawyer's plan to stop the execution was to try and prove something that nobody had successfully done in three decades that one of Texas's most

notorious serial killers was actually innocent it wasn't that the lawyers didn't have a case to make and no two people fabricated testimony to get a guy executed it's just that they had so little time to make it the last 12 weeks listen or ever you get your podcasts okay see we're going to wrap the show today by talking about something that I have been very interested in now for a little while which is AI super forecasting this is something that has

been in the news in recent weeks there have been some platforms that are out there using AI

To make predictions about the future this could be very profitable if you're ...

calcium or polymarket or it could just be really interesting if you're using this to predict world events or who's gonna win an election or something like that this is an area where we have not been

has focused as on some of the other emerging capabilities of these AI systems but I think it's actually

quite interesting and profound that we now have in some cases AI systems that are better than humans at making predictions about the future this is the only segment that we're going to do about predictions this year that isn't about scams because mostly predictions these days are just about scamming others and doing insider trading but this is not that this is not that so we are excited for today's guest many of Vesalovsky is in the studio with us today he is the founder and CEO of the

AI forecasting company pre-scene they make a platform that uses AI to make predictions about future events and we're excited to learn more about what they're doing Vesalovsky welcome to hard work thank you guys for having me so let's talk about AI super forecasting and this idea that AI systems are becoming as good as or better than the best expert human forecasters where did this start and when did the models start to get really good at this I think people have been

trying to use AI for forecasting for at least three four years like since my co-founder since the

GPD-4 first came out he's been forecasting for 10 years and his first idea was like let's try to build an

agent or an AI for forecasting the original models were terrible I mean they try to add up two numbers and they'd be off by an order of magnitude and I don't think it was until the end of last year where we began to see these kind of hints of brilliance of these systems and this was mainly because of this like you know a genetic revolution that we've all experienced were suddenly the infrastructure around the web and around these agents and the RL that the model providers have

been doing to make them really good at tool use improved dramatically I get why you would want to be a super forecaster or to have a super forecaster AI if you are like trading on a stock market or a prediction market that seems very lucrative and in fact one of your co-founders did this very successfully you noted on your website that he converted thirty five dollars of initial investment

to nearly two million dollars trading on calcium with this AI bought so like I guess my question

is if your AI super forecaster is so good at you know winning these markets like why are you starting a company why not just start a hedge fund and use this to make yourself rich this is precise the question that all the hedge funds ask us when we try to sell them there this they're like oh well listen you guys have an alpha making machine go out and you know make money yourself

the reason why I'm doing this is because I think that a lot of good can also be derived by having

an AI super forecaster if you think about it every time we make a policy decision we are implicitly making a forecast that this policy will have this effect on this group of people and right now a lot of that is a little bit more like vibes based we kind of talk to some McKinsey experts and they tell us hey listen you know this will do this if we can forecast the future really well then we can better design policy and better select policy to make decisions so really kind of the reason why

we haven't decided to open up our hedge fund is because we want to make this more broadly available

to the government to the insurance world to basically convert these forecasts into actual actionable

decisions. I want to hear a little bit more about like what are the mean kinds of things that you try to forecast and then a little bit about how the forecast work like I could imagine it feeling a little bit like getting a deeper research report you know from an an AI lab where they go out and they scouted the web and they they they do a bunch of research and then they just like kind of try to reason through what they found but yeah tell us a little bit about what you're trying to

predict and how it works. We're mostly trying to focus on like geopolitics macro markets where there's a lot of unstructured data and a lot of different data sources that could be telling a similar story and so we're forecasting questions like will the US invade Iran when will the straight of her moves reopen who will win the elections coming up. We focus on basically building really good AI systems to forecast out those questions regarding how does it relate to a deep research

report. I think in some ways they're similar you know the main problem with the existing deep research

agents is that they they lack connection to all the different data sources that are out there. So if you use you know chat you be tier clawed out of the box it is access to web search which is the surface web but if you think about it there's thousands of different API endpoints out there that might be relevant for a specific problem. So for example Columbia has like 50 different government APIs out there that ideally you would give the Asian visibility into another fun approach

for actually how we are thinking about this is often times the best forecasters aggregate a lot of

Different forecasts when they produce their forecast.

Denofville and he's the he got a lot of his edge by actually knowing which people to look at

when he's you know forecasting a question then afterwards reasoning over all their predictions.

That being said though a lot of people on sub-stack make predictions but we don't know how valid those predictions are even in hard fork you know you guys make frequent predictions but

then we never actually go back in time and verify how accurate worry your predictions.

So one thing that we also do is we kind of boil the ocean by finding all people on sub-stack all podcasts out there and we base the extract all claims that these people made and then we score them on how accurate were those claims. Wait so are we like in a database somewhere at pre-scene headquarters because we would obviously love to know how we're doing particularly compared to one another. You guys actually are in a database. Yes! So the reason why you're in the database is mainly

because yesterday I was like oh wouldn't it be funny to you know pull up some of the claims that you guys made before and it's actually score on how accurate you were and maybe there's some areas

where you know case you have a lot of knowledge about anthropic. So maybe you know you could make

really good forecasts on anthropic related questions but maybe on Iran less so. So basically how do we

you know create this like function over the possible space of forecast and like we know cases really good in this area so we really have to index on them heavy here. So I played around with your platform a little bit you generously gave us some credits it's a sort of closed beta with a weight list now but I went in there and created a question because I was curious like how this all works. So my question was will there be an AI data center in space before 2030 and you sort of put

your question in and then there's some AI that like helps you flesh out your question maybe make the resolution criteria a little more specific say well how big a data center and you know is it January 1st in which time zone on in 2030 and so it sort of gives you this this AI sort of augmented

version of your your question and then you run the forecast and basically it goes out and it

does a bunch of searches and it compiles what it finds and it has a bunch of sort of sub agents that are looking at various aspects of this and that at the end of it like a couple minutes later you get this forecast and this case it says it estimates that there's a 26.8% chance that there will be an operational AI data center in space before January 1st 2030. So how much of that is this thing actually reasoning through its own decisions and forecasts and predictions and how much of it is

just like it's collecting all of the sort of data about this topic and kind of averaging it all out and putting it into a single number. It's really a combination of both. So if you kind of saw when you first created the forecast there were four sub forecasts that launched and we tasked these sub forecasts with really trying to come up with their own independent conclusions on this question. So here they might look at progress of existing data centers so they might look at

historic build times that it takes actually build out a data center. They might look at space infrastructure and look at the literature there and basically come look at the primary sources and come to a conclusion on how likely they think it is to happen and then later we have this synthesis stage and basically here it looks at all of the different sub forecasts the results that they got and it tries to synthesize first of all what each of them decided upon but then afterwards

also look at the broader ecosystem. Maybe there's a calcium market related to this where there is you know a lot of liquid money at play or maybe there's some experts on substrate that have been like reporting about this for a while and they basically try to reconcile what did we do that was different from the consensus and what does the consensus think that we might not have factored in and then afterwards it kind of uses our primary analysis and this kind of broader discussion

and combines them in a recon sauce them. So in the report you would see at the very bottom this like what's not obvious section and the point of this section is basically to see okay what is the market miss pricing here what are other people missing that we might be picking up on. And I tried this also because I was just curious how this prediction would compare to if I just gave the same question to a basic AI model that is not sort of configured to make forecasts

and Claude gave me roughly the same answer it said like it's about a 20% chance the the uns the pre-seen prediction was a little bit higher probability but like what is your system doing that at just a normal like AI model is not doing? Or is Claude just scraping your website they've

done it before. You have to say I hope they are scraping it. I would actually really don't mind that

um so I think that it's a little bit unsurprising that a lot of different systems can come to the same conclusion because it kind of validates that this is coming to the same answer. I think the best way to look at is like meticulous has these leaderboards for these competitions and they're basically evaluating simple Claude code or Claude with some basic scaffold versus kind of real

Agent developers who actually build this out and the scores have gotten widel...

like kind of with time at the start you know just Claude with web search was doing a good job

but as you're able to manufacture this kind of scaffold around it into great more data sets create new research artifacts or new tools that are relevant for forecasting this gap has been getting larger and so maybe on this question you know they're pretty similar but maybe if we're trying to forecast something in the Middle East or Africa right now they might be different and usually when they're different the system that is kind of well engineered with this like good scaffold

and thinking about forecasting the right way in terms of these like metacuristics of you know how do you actually reason about this stuff um the gap begins to form. So your your platform is still in beta has not been around too long but I'm curious so far have there been any moments where you feel like the platform predicted something like very non obvious that actually came to pass yeah there's one forecast that we've been running for a little while about kind of data

center construction around the US and in general we were very different from the metacurist community on this and so the metacurist community consists of like some of the best forecasters out there in the world who are all kind of competing in this like world cup of forecasting and so this is one

of that I think one of the benefits of AI is we can go to every single state we can go to every

single proposal of a data center and really analyze it closely the other one was about the recent reshuffling of the um cabinet in the UK we've made some great forecasts around that the other interesting one is like conditional forecast which I think is where the really policy implications lie is like conditional on Andy Burnham being the next Prime Minister who is he likely to elect this chancellor and then you could imagine there's lots of you know

demand for this kind of question because there's lots of downstream implications of this and we tend to do pretty well in these conditional forecasts. I'm doing one that's like we'll Kevin know what the chancellor of the UK does so that's I have some very low probability of that I totally see the point that like having better predictive capabilities is going to be good for us in lots of ways um it does strike me though that there's like some risk of what people have

called like gradual disempowerment well like right now if you're the CEO of a big company or if you're a government official a lot of your job is trying to sort of predict the future and make policies or strategies that like align with that vision of the future and I just like I don't know what the world looks like in a in a situation where the AI is just like markedly better than us at doing that like some part of the authority of humans in those positions in government and

industry and frankly everywhere is undermined in the case where we're all just consulting these AI

oracles before we make decisions like doesn't that mean that they're kind of running the show?

Have you read the Scott Alexander short story of the whispering earring? Yes. Yeah so there's this great Scott Alexander short story and you know the idea is that the you know you sort of find this this relic it's this earring and the first thing it says to you it's it's better if you don't

aware this um but if you start asking questions it always tells you the right things to do

which is initially very exciting but then over time exactly what it has said happens Kevin which is you're you're being controlled by this thing right like you have no agency whatsoever and you're effectively just being steered around by an earring so better to take the earring off. Yeah yeah I wish I had an answer for this I'm also a little bit scared about this kind of disempowerment I'll give you one example so I have a friend who's very interested in autonomous organizations

so like kind of how do we have AI spin up their own companies that then go solve problems for people and then you just have basically AI running the next generation of startups there's lots of problems with this you know like how do you who's legally liable if it does something bad and so on but I do think we are kind of increasingly moving to this world where if the AI is smarter and can make better decisions than we can then probably they'll end up making a lot of those decisions

and running a lot of this show. Well I'm curious if right now at your company there's any role for human forecasters like is there a way where you know humans and AI is working together or coming up with better forecasts or are you just sort of purely in the realm of like let's see

what the man I said. No one 1000% so we have two super forecasters Scott and Robert who are incredible

and we're basically building out the center solution so like I don't know if you guys remember but like when you know the machine first beat humans in chess machine plus human beat the machine

in chess and that lasted for a little while and I think we'll see something similar in the

forecasting realm where sometimes these models still make stupid mistakes they reason about probabilities the wrong way or they don't consider certain factors the right way and humans are able to pick up on that especially like domain experts so I do think that we shouldn't be viewing

This as like a substitute for human analysts but really kind of a way to impr...

to ask a lot more questions and come to a lot better decisions. Are there topics or domains where AI is better than humans already at forecasting and like what are the what are the sort of best and worst areas for the AI forecasters? So the Scott Alexander piece and then afterwards later the FRI Institute came out with some kind of results that AI has reached the level of super forecasters and I actually don't fully buy those claims so for example humans are lazy you know the incentive

to compete in these tournaments is $5,000 price pool you know and so like kind of are they really

doing the best that they can in these contexts that and so usually I think where AI is

better than humans right now or just places where kind of humans are too lazy to do all the

analyses so we recently were like the first spot ever to in a human and AI forecasting tournament

on meticulous and this related to macro markets so like predicting the interest rate or earnings per share of some big company and we tend to do really well at that and I think that's just because our agents do the analysis that the humans are just too lazy to do and so in some ways I don't know if kind of we're at the point yet we're kind of AI's are really better than the best humans just because I don't think there has been truly a competition where humans gave it they're all but kind of

I do think that within like a year or two and you know what quote me on this we will kind of add it to the database I do think that within a year or two AI will be better than humans have forecasting well let's paint you down though is it going to be better than humans in one year

or two I think it's actually one year three months and six days all right there we go let me

yes how much of this is just driven by basic advances in model capability like is it as simple as like you know something like a fabel or GPT 5.6 comes along and like your system is just immediately much better or is there more a tinkering that has to happen the way that I view this is you want to

build out the infrastructure for the world in a way that these new models can basically use this

infrastructure really well to do forecasting well so like kind of fabel out of the box is a good forecast or by no means a great forecaster but when you give it access to all the right tools and all the right data sources then it becomes a great forecaster and so yes these models are getting a lot better but I and I think that they are kind of improving forecasting dramatically but it's mainly because their judgment is improving they're kind of able to reason about problems better

decompose problems in a better way and afterwards our job is to provide them with the context with the tools to be able to forecast well who's your customer like who do you imagine paying for this kind of forecasting service yeah so far we've landed some proof of concept partnerships with some of the hedge funds they're the most immediate buyers and the quickest to move usually but really I'm using them in some sense is like a stepping stone to really get into the

longer sales cycles with governments with NGOs with international organizations basically

institutions that make really important decisions for collections of people yeah that that's my ideal

customer well thanks so much for stopping by I'm gonna keep playing around with pre-scene fascinating idea and yeah maybe I'll go out there on the prediction markets and it makes a mula I'd love to see there's a lot of them at this point all right thanks for the us guys thank you so much for having me the New York Times app has all this stuff that you may not have seen the way the tabs are at the top with all of the different sections I can immediately navigate see something that matches what I'm

feeling I owe the games always doing the many doing the word I love how much content it exposed me

to things that I'd never would have thought to turn to a news app for this app is essential the New York Times app all of the times all in one place download it now at nytimes.com/app hard fork is produced by Whitney Jones, Rachel Cone and Davisland this week we're edited by John Wu and fact checked by Will Pyshal today's show is engineered by Alyssa Moxley a original music by Alicia Baitoup, Mary Elizano, Leah Shot Dameron, Alyssa Moxley and Dan Powell video production

by Soya Roque and Chris Shot you can watch this full episode on youtube at youtube.com/hardfork special thanks to Paula Schuman, Wewing Tam, Brooke Mentors and Dahlia Haddad you can email us

As always at hardfork@nytimes.

you

Compare and Explore