The DSR Network
The DSR Network

Siliconsciousness: AI Goes Rogue

1h ago40:595,680 words
0:000:00

An OpenAI model has gone rogue. Not only has it gone rogue, but it hacked Hugging Face completely unprompted and without approval. That’s pretty scary. But just how much of this story should legitimat...

Transcript

EN

[MUSIC PLAYING]

Welcome to Silicon Fitness.

The DSR network podcast focusing on the artificial intelligence revolution, politics, and policy. Hello and welcome to the latest episode of Silicon Fitness. I'm David Rothkop, you're host this week.

As every week, we're going to talk about AI subjects that are important. And this week, we are extremely fortunate to have with us Will Douglas Heaven, who is the senior editor for AI at MIT Tech Review

place, you know, we work with closely on a regular basis and follow closely.

β€œAnd we think all of you should follow it closely, as well.”

There will cover his new research, emerging trends,

and the people behind them, how are you today?

Well, I'm well. It's good to be here. Thanks for having me on. Well, thanks for joining us. And let's start with something that's in the news,

which is this security breach that was identified by this company called Hugging Face, that's associated with a couple of open AI projects called a lot that are derivative of chatGPT and some advanced stages of products development.

And I saw a number of wags on the internet, this morning talking about it. And one made some fun of the fact that this was a breach by an open AI capital, a product that Hugging Face could only deal with

β€œby tapping into open AI small, small AI products from China.”

In other words, open AI was, you know, not actually open, and they needed some Chinese products to help them crack this code. So it seems to have many layers of the kind of stories we're interested in here.

Maybe you could talk a little bit about why you think it could significantly. Yeah, yeah, this is certainly something that has led up the internet today. And that thing you mentioned there

is definitely one of the astonishing things about this. I mean, just some personal context here, I am usually the first person to sort of poo poo a lot of these sort of, you know, scarce stories. I mean, we've seen lots of stuff over the last couple of years

about AI models sort of doing things which are malicious or worrying in some way,

but they're nearly always what they have always been

in some kind of simulation. So yeah, if you push an AI model when you're sort of researching it or testing it to do undesirable things, you can typically make it do those things.

And then you ring your hands over it. And there's been a lot of scam on ring around AI. So I'm generally pretty cynical about this kind of stuff. But if I'm going to be honest, this is the first time

β€œwhen I think this is a chilling example.”

This is the very first time that an AI model has broken out of it. It's sandboxed, it's sort of containment area that the researchers were running it in. And autonomously hacked into entirely separate companies

computer systems. So that in itself makes this extremely no worthy. And of course, everybody is up and arms calling for regulation once again, and that is sort of monitoring of what researchers are doing with these models.

But yeah, I mean, I'm laughing. It is a amusing thing, though. So the one of the ironies here is that hugging face apparently identified a hack in their systems week or so ago, or towards the end of last week,

with, you know, obviously, like many companies, they run AI systems of their own, and they had something that detected a anomalous behavior in their computer systems. But then to try and diagnose what had happened,

they went away at all of what was actually going on at the time. They tried to use open AI's models through the API and weren't able to, because when they sort of uploaded a lot of the sort of the cybersecurity logs in information, the guardrails on that sort of the open AI API kicked in

and said, no, we can't help you with this. It looks like you're doing something that's a bit dodgy. So they had to turn to an open source model that didn't have those guardrails in place

To even understand what open AI

it turns out had been doing itself. So that is-- - And it was actually-- - And it was a Chinese model. - Yeah, yeah, yeah, yeah. So there's so much, there's still many things

going on here that what seems to have happened,

β€œI mean, I think open AI and Open interface”

are still sort of figuring out the details, but I mean, an open AI has sort of stood up and said, you know, we did this, which is, I suppose that's good.

I always hesitate, that's slightly--

- I read the statement that they put out. And they said this happened. I don't know that they said we did this. In other words, I don't know that they-- I still am unclear of who is responsible

for what happened. - Yeah, okay, well, they were already sort of in deep territory. So I mean, Open AI's blog that they published on it, they've said that it was one of their models that their researchers were testing so that they were--

they believe they were running an experimental new model that has high hacking, cyber attack capabilities. They believe they were running it

β€œsort of securely within Open AI's computer systems”

and they hadn't connected it to the internet at all, except for one link to a computer system

that was run by a third-party piece of software

that would just allow it to download things that the model needed. And they were running it on some pretty tough benchmarks that basically pushed that new model to try as hard as it could to look for security vulnerabilities

and then they just left it, left it running. And what we've seen for years and years and years now that if you let an AI system run on a very specific task, the sort of basically set it up to sort of go and achieve this, it will find a loophole that people had not seen,

because it's just going to do whatever it can to achieve that task. And people aren't able to think ahead and see all the possibilities that the AI model might try. And in this case, it seems to have found a vulnerability in that third-party piece of software,

that one connection to the internet that they'd allowed it. There's an unknown bug in that third-party software that the AI model found, exploited, and that let it break out of its containment onto the wider internet.

And then it seems to have figured out somehow that in order to achieve my task of exploiting other software systems, I should go to hugging face, because that's a repository of lots and lots of software and examples of how to do this.

So the steps it took were extremely logical. The fact that it took them, I don't think, could have been foreseen. So I think it's a fascinating and quite chilling example. But that's the problem, I'm sorry,

but the problem is that they couldn't be foreseen, right?

I mean, there's one little loophole, could have ordered pizza, didn't order pizza. Instead of said, okay, we're gonna go and keep, and I'm gonna keep looking and looking. And you know, AI isn't smart.

AI is hugely capable of testing millions of options, searching for millions of things that humans wouldn't have the patience to do, right? And so it literally checks, you know, you put it in a prison, and it checks every boat,

and every rivet, to every law, every rattles every bar, to find the way out. - Yeah.

β€œ- And so, you know, once that's what its mission is,”

you would think people would be a little bit more cautious about how, well, let's leave it in there and go have lunch, you know, I mean, it's, right? - Yeah, you'd think people would be more cautious. I think this is, there's obviously a lot of hubris here.

We've seen this kind of thing happen many times before, but usually with lower stakes. I mean, you've probably heard, I mean, many people listening probably heard of examples of how, you know, when you ask an AI to try and beat a game,

a video game, there's been all sorts of video game testing. Many times they will find a cheat, a sort of a shortcut to the goal that the humans knew. Why didn't mean you to do it that way? But, you know, that's to the AI, which is completely mindless,

it's as legitimate a part of the goal,

As the ones that the humans had in mind,

but the humans are blind to it because the humans think, no, this is what I wanted you to do.

β€œI wanted you to do it in, you know, the way I was thinking,”

not that sort of that shortcut that I hadn't thought of.

So, it's always going to find those paths,

if given, if those paths are available, then these systems will find them. But mindlessly, there's nothing sort of intentional here. - This podcast is underwritten in part by the U.S. Embassy of the United Arab Emirates.

It's editorial content is completely independent, and the views expressed are exclusively those of participating experts. It is presented live without editing. For further information about the UAE's efforts

in the areas of artificial intelligence and technology, go to the website of the embassy at www.ue-emBC.org and search for UAE-US tech cooperation. We thank them for their support. We thank everybody who's supporting this podcast

for their support, and we look forward

to it developing and growing over time because the issue is so important. - No, it's, it's just, there's a certain degree of the power of mindlessness, you know, and which is just the relentlessness of it all.

But it, you know, it does get to another thing, which is within the open AI model, apparently there are guardrails that keep people from fixing the problem. And, and, and that again, seems, and again,

I referring to open AI as a company. And it, and it seems to be a fundamental flaw because if you can't be sure you're gonna keep everything in its box, shouldn't you be preparing a product

β€œthat allows people to work out their own solutions?”

As opposed to having to go down to the corner store, get the Chinese open AI model and use that. I mean, to me, there seems attention there. I'm interested in your, there is attention. I don't think we've got these guardrails right at all.

I mean, there have been a lot of complaints from people using, using anthropics, new model, fable, which is, you know, the sort of the water down version of its, of its mythos model, which was, you know, pulled for being, again, too scary.

It's so hard to talk about this without sort of immediately going into sort of, you know, a house scary was it? That sort of serves anthropics marketing to sort of build it up as being really scary.

But anyway, mythos was too scary. So we have this water down fable version. And people are complaining that you can't do anything, a lot of people are complaining, they can't do anything interesting with it

because the guardrails are too, are too strict. So yeah, we haven't got this right at all. I'm sure we'll be figured out.

I don't know what the answer is.

I mean, but there is a problem here, right? Yeah. Which is anthropic and open AI and these other companies that are either public or some day be public or have to follow some kind of disclosure rules

talk about this stuff. Now, Peter Teele and Palantyr is doing stuff for the Department of Defense. Where everybody in the room is saying, we can't let anybody know what we're doing.

But Pete had says favorite word is leafality. Go for it. You know, it just seems to me that the world we know

β€œabout is scary, but when I think about that,”

I think, well, the world I don't know about, maybe more scared. Yeah, yeah. No, I agree with that 100% for all. So there's been a lot of talk about this recent opening

AI model breaking out of it. It makes it sound like a sort of a wild creature that's broken over its cage to attack people. And with this, it's a fascinating case, but it's the least scary thing going on at the moment.

Yeah, it's what's going on the lack of transparency around how these models are being used. By, you know, it's people with a problem. Not these not models themselves somehow, sort of, I don't know.

You could taking, taking initiative to do something. There are just, I'm talking like this because there have already been so many headlines that suggest that the open AI model somehow broke out and went on a run page, which it's

almost impossible to have a serious conversation about this technology because it's instantly sort of, you know, hyped up and answered from the author's.

Well, yeah, now the first person I talked to about this

was immediately talking about SkyNet.

I mean, it was like one inch apart. Every AI story is one inch from SkyNet, you know, and it's just because why? Because people don't know what the other options are. Those are the two things they know.

One is an Arnold Schwarzenegger moving. The other one is, you know, real life. But it does put us in a bit of a bind, you know, from the point of view of what are we expect, what is a regulator expect?

It would have should a government respond to it.

β€œAnd, you know, I think one of the things”

that's been interesting over the past few weeks is you talk about the anthropic bottles is while they were going to release this and then they were told they couldn't release it and then a couple weeks later,

they were told they could release it. They're going to release part of it. And, you know, there's this behind the scenes thing going on where David Sachs or somebody's calling up the White House and saying, hey, look, do worry about this

or don't do that. That's going to screw up this, you know, valuation on this company that I own, I don't know what the discussion is, but you do get the sense that the top priority

is not always what is the most

ethical sound wave developing this new technology? - Yeah, yeah, no, I don't think there are, I don't get the sense through those kinds of conversations, either that all the stuff around the mythos release and then when Fabel came out and then it was pulled again

and it was hard to know whether that was a big political game. I mean, of course, because this was hot on the heels of anthropics, big bust up with the government and who knows, it is because there is so little transparency. It's so easy for all these different theories,

these conspiracy theories to take off. - Well, think about anthropics, big bust up with the government. Why did they have a bust up? 'Cause they said to the Defense Department,

there will be limitations and the Defense Department said,

oh, no, no, we can't have limitations. And so, that suggests the sort of depth of the problem beyond the one we're dealing with.

β€œ- Yeah, yeah, very much, that's what scares me.”

Actually, what the military might be getting to do with these models that we don't know about. I mean, the fact that we know about this recent open AI thing, the opening and opening face thing that is a blog post about it, I mean, this is, at sort of lovely,

this is everyone's talking openly about it, and everyone's sort of following the hands up and saying, you know, whoops, we did this. And if only it was all like that. - But you know, people have been fighting wars

for thousands of years. And there has been bloodshed, mayhem, Trump. And so, in the past century, people said, you know, maybe we should set some rules. And, you know, you're the Geneva conventions

and the roles were, well, you can't kill citizens disproportionately, you can't, you can't, you can't,

β€œyou know, you can't blow up civilian infrastructure.”

So, the people can't get drinking water. They can't get home, or they can't, you know, et cetera. Implicit in that, we're other things. Like, don't target a girl's school and blow up a bunch of girls, right?

These are the rules. The people who are making the rules for AI right now don't care about those rules. In other words, we've got to set rules, you know, whether it's the Geneva convention or it's the law.

And these people are like, we don't care about those. And those are the people who have black budgets. I don't mean to be an alarmist. I've been Washington 30 years. I'm cynical and I've seen everything, right?

But, but I also know that a lot of stuff happens inside a black box over, you know, in the military industrial complex, someplace in Virginia. And the people, it used to be, at least there was,

you would believe that at least somebody in the room was gonna have some values to contain it. And I just, I don't know that we should, I don't know why we should believe that's the case anymore.

- No, I'm not sure.

We should, I mean, if people can get away with breaking rules, they will break rules. I mean, that's even if they are ostensibly claiming to follow those rules. I mean, increasingly,

governments and states can just flout those rules and shrug and say, so what?

I mean, the whole international order of rules is basically,

almost feels quaint these days. I mean, that's, that's the, with this conversation has gone down at a dark path.

β€œBut yeah, I think if we're talking about”

flouting rules, you're in conventions of, of being sort of held the world together for the last, or for a lifetime. Yeah, it all seems to be falling apart. And to put AI back into that thought,

then even if, like the best case scenario is you have some rules set in place

that regulate how AI is used in, in, in conflict,

or in by government, sort of, more secretive government agencies in general. Then you can see that unless that is done so thoroughly, the all the loopholes are closed off. Then, yeah, these systems are not yet well enough

understood to be sure that there isn't going to be, like, a breach. I mean, what would sort of, what would,

β€œwhat would an AI looking for a treat or a shortcut be?”

And, you know, in a, in a, in a war zone. I mean, or who needs war? You know, I mean, I've been involved in these discussions in the national security community for the past 25 years.

And, you know, one of the objectives of war is to slow down your enemy. And, and one could very easily see using AI and saying, look, I don't think I like their GDP growth. So I think their GDP growth should be half of what it is.

And I'd like them not to know what the reasons for that are. You know, and, you know, you sort of get down in the weeds of how to do that. I mean, 20 years ago, I was at some defense department, scenario actors, like post 9/11,

and, and people were like, what were really afraid of, or spy-a-war, and really afraid of cyber? And this was before cyber was even, well, formulated idea. But the idea that was presented to them at the time was, somebody could hack into a hospital

and change the prescriptions for everybody in the hospital. And, and, and, kill people and nobody would know what was happening until it was all done. And so, so all I'm saying is that, you know, we talk about rules, but the complexity grows geometrically,

when you've got a tool that is looking at millions of avenues of attack.

Some that have never been thought of as avenues of attack before, you know,

and, and, and, you know, I don't mean to go from this particular hugging face example, but you have a, you know, chat GPT sitting there going, oh, oh, what is this strange back door? Let's, let's push that in.

β€œHow many, how many back doors do we have in a global system that didn't know they were back doors?”

I'm sorry to interrupt, but I really have to tell you about our sub-stack. The DSR network of sub-stack is the absolute best way to follow deep state radio. The DSR daily words matter need to know Silicantiousness, AI Energy and Climate and the daily blast. Sign up to get notified every time a new episode drops and stay up to date on the latest news and expert analysis.

You can also support us by becoming a paid member, paid members, get an ad-free listing experience, select episodes two days early, access to live streamed episodes, and 50% off of David's need to know Substack. Please consider joining us at DSR network.substack.com. That's DSR network.substack.com. Thank you and back to the show. Yeah, yeah, I think that is just in the last few months, these models have become

capable enough to be serious threats to global cybersecurity. I think there's probably a lot going on at the moment using these latest tools to find those vulnerabilities and show them up before others find those vulnerabilities and exploit them. Well, you said earlier that you don't like to be the alarmist here and I think you're changing. I don't know you, you know, I just met

You here but you gave a speech in London on five things that AI and item numb...

scary for real this time. Yeah, you see a pattern for me here. Yeah, well, you call me, you're

caught me on a bad day. Yeah, no, I don't, I do not like to be alarmist. I mean, importantly, there's far too

β€œmuch alarmism around this, this, this technology, I think, but it's really hard to get this balance”

right because that is not to say that there are serious, serious things that we need to talk about, that I don't think, you know, are being dealt with or we don't yet have the answers for it. It's more, I think my position is more that there's a lot of, a lot of scare mongering around things which are not yet or ever will be real concerns. And so, yeah, what I was talking about a month ago in that talk at South by Southwest London, yeah, for years we've had sort of

scary stories, you know, from the sort of the dooma folks about how AI is going to kill us all

in encivilization. And I've always sort of rolled my eyes at that kind of stuff because

it is still very much science fiction, a short, like in sort of, if you put together a string of hypotheticals maybe, but nothing that I would take to seriously, but what we have seen in, I relatively short window of time is that many of the sort of the real world short term harms that researchers have been predicting for a while, such as, you know, the dangers of deep fakes and the dangers of, you know, using AI for cybercrime. Those really are already here,

β€œI think, we've seen deep fakes is a very, very broad term. I mean, people probably think of”

you using them to sort of, you know, manipulate, you know, for proper propaganda reasons. I mean, there are lots of examples of, you know, politicians being manipulated to say things. And while we are seeing lots of examples of those, I think it didn't take, you didn't take AI to make people think that politicians had sort of said something they hadn't. So it's hard to, how to judge how much of a different start kind of thing will make. But I think the, the, the nastier side of deep fakes is,

is, is the abuse using, using this technology to, you know, create pornographic material of, of people in your school or, or celebrities. I mean, that's, that's something that people have been concerned about happening for a long time. And now it's, it's happening in it, it's horrific. And

β€œthe worst thing about it, it probably is that, I mean, there's a big outbreak of these things on”

X because, you know, Musk's AI grok of all the mainstream chat box has fewer guardrails and people were able to use it to generate abusive imagery. And, and it took a surprising amount of pressure and criticism on Musk's company before they even really took it seriously and did anything about it. So that's why I think AI is is, is, is getting scary now for real. It's, it's, it's really, you put these tools in the hands of, you just boring run of the mill nasty people and they can do a lot of damage.

Yeah, well, you know, I remember I went, I had a meeting with a guy who used to be a senior person in the NSC and then was running one of the biggest Washington thing. I mean, household name kind of place. And we were talking about AI. This was three years ago, maybe four years ago. And I said, so what are you interested in? Like, what are you guys working on here? You know, they have unlimited research dollars, we were, I said, we're only only interested in one thing. And I was like, what's that?

He said, well, you know, AI in the end of the world. And I was like, what the, I mean, you know, it's like, that's the only prop, you know, in your South by Southwest, five points. And people should read all your stuff at the tech review, but, but in one of the points was AI is everywhere. And, you know, I mean, we regularly hear talk about the fact that it's just a

useless term almost because it means a million things to million different people. But one of the

aspects of it being everywhere is that it can be used for good or for evil, everywhere, just like, I don't know, fire can be used for good or for evil, wherever they, and wherever you go, there will be good people doing good things with it, but there will also be bad people doing

Bad things with it.

Instagram anymore because most of the feed is fake. It's slow, you know, and, and so what's the

problem with that? Well, somebody could use it for propaganda. But also, it sort of delutes what truth is for a whole generation of people. And that's more pernicious in many respects than the propaganda it is. And that's just one dimension of it. And so, you know, it's, it's this ubiquity of

β€œAI and the double-sidedness that comes with that kind of ubiquity. It doesn't mean you have to”

be a pessimist, but it does mean that there is more to be concerned about wherever you look whether you're teal in national security or personal security, whether you're raising your kids

or protecting IP or and, and it's going to present itself in a thousand ways, you know, as

also as you were talking, let's think, you know, AI that go and create a video that makes it look like Donald Trump is on a date with Taylor Swift. Well, that's funny, maybe. But, or are fine, but, I could also create a record deep in the bowels of the Board of Elections of Ipsilani, Michigan that looked exactly like the record that was there, but just happened to be different in terms of

β€œresults or changing the names of a few. Oh, you have to do has changed the names of 10 people”

to people who are dead or not really Americans. And the president says, we can't trust this election.

It doesn't have to be high production value, deep thick. No, it doesn't. And in the

president was saying that before, AI could do these things. Yeah, you don't need, you don't need very much at all to sort of destabilize trust. And yeah, AI certainly makes the problem worse, but I don't, I don't think it introduces a new problem so much. It just makes it so easy to do that more people will probably have a go at doing it. And it is everywhere. I mean, the world is, the world has been technological for decades, like every part of our lives is mediated

by by some digital technology. And AI has suddenly become a new layer that sits on top of that.

β€œSo just in almost no time at all, AI also touches on everything. And I think that change,”

we're still getting our heads around in that change, but it's going to be something we need to get used to in it. Well, look, I don't want to send you off into your summer all depressed. So let's end with a pallet cleanser. One of the other things that you said in your five things yourself by Southwest thing, which sounded, you know, in many ways, like one of the most optimistic things, is that AI for science is real. And it's a really big deal because it is shortening discovery times,

shortening, coming up with new ways to test things. You know, we had on a guy who runs the AI University in the UAE, and he was talking, he's got a company that does this, but he was talking about simulations. And you know, you can go and create a simulation and you were body with your genome, with everything that's going on, you're biome, everything that's going on, and test the drug against that. And you could do it for everybody. And so, you know, you're getting

into a point where the development of solutions and customization of solutions, and that's just biotech, it's true everywhere else. It's just, it's being transformed so rapidly, we can barely comprehend it. And so, to me, that's a pallet cleanser. You are optimistic. I'm just giving you an opportunity to say something optimistic. I appreciate it. I don't want to go off as the the AI Grinch, but no, I am probably the most positive application of AI that we are going to see

without a doubt. And it is going to touch all areas of science, both in the form of specialists, tools that scientists are creating to do particular jobs. But I think the new shift we're seeing is that large language models, general purpose large language models are now good enough that mathematicians and scientists are able to use them as sort of sparring partners to bounce ideas off. And a growing number of anecdotal examples, but you know, we're getting, we're seeing more

Of them.

made any progress on it in months. And we sort of put it into one of the, one of the LLMs and

it basically came out with it sort of a path forward that we hadn't hadn't thought of.

β€œI think that that's genuinely exciting that these will, you know, as a general purpose tool for”

scientists that they can use to push their research on in ways that they might not have been

able to do otherwise. I mean, across all sciences, I think that's genuinely exciting prospect.

It is. We're, you know, AI optimist inclined here and we're AI, you know, realists that we want to

β€œcover this, brother. We've been doing this for a few years. So we're, you know, we've had enough”

positive stories. And I'm sorry to have dragged you down into into this stuff, but this was the

breaking story of today. I hope we can coax you back at some point, maybe Christmas time and we talk about something more positive as we're doing this, because the work you're doing, the stuff you've been writing about, the stuff all our friends at TechReview are doing is terrific. And I just, I hope we can be a platform we can share more of that. Yeah. Well, I really appreciate that. And this was fun. Yeah. So this was the week when we were talking about an alarming story. But yeah,

let's, let's, let's meet again, and I can tell you what I want Christmas. Okay. We'll do that. Thank you very, very much. Well, and thank everybody for listening, and we'll be back, of course, next week, and every week, with more here on Sylla Consciousness, and if you like this, do listen to the AI energy and climate podcasts that David Sandalow hosts for us that we do with Columbia University,

β€œbecause that's got a bunch of important stuff in it. We did a show with him last week,”

in which we, we sort of combined what they're doing on, and this show, and I encourage you to listen to that if you haven't. In any event, see you all soon. Bye, bye.

Compare and Explore