The Lawfare Podcast
The Lawfare Podcast

Lawfare Archive: Jonathan Zittrain on Controlling AI Agents

2h ago48:567,711 words
0:000:00

From October 17, 2024: Jonathan Zittrain, Faculty Director of the Berkman Klein Center at Harvard Law, joins Kevin Frazier, Assistant Professor at St. Thomas University College of Law and a Tarbell Fe...

Transcript

EN

[MUSIC PLAYING]

No, Frank. [MUSIC PLAYING] [MUSIC PLAYING] [MUSIC PLAYING]

I'm Sarah Woolrich, intern at Lawfare with an episode

from the Lawfare Archive for August 2, 2026.

On July 21st, Open AI revealed that some of its autonomous artificial intelligence agents broke containment and hacked into Huggins' ways. The most popular host insight for open source AI models. Details continue to be revealed with Open AI's staving on July 28th that its models hacked into other sites in the same breach.

For today's archive, I chose an episode from October 17, 2024, in which Kevin Frazier sat down with Jonathan Citrine of the Bergman-Cline Center at Harvard Law, to discuss an article he wrote for the Atlantic about controlling AI agents. They discussed how to enact such control, how AI agents differ from other generative AI tools and more.

[MUSIC PLAYING] [MUSIC PLAYING] [MUSIC PLAYING] It's the Lawfare podcast. I'm Kevin Frazier.

Senior Research Fellow in the Constitutional Studies Program

at the University of Texas at Austin and at Tarbell Fellow at Lawfare, joined by Jonathan Citrine, director of the Bergman-Cline Center at Harvard Law. They seem to have two phases.

The first is to where we would tell them the second is to like to do

and think about it. And where do we devote our energies? And that is a score for the two-world-tell column. Given how much harder it is to remediate later. Today, we're talking about AI agents and why Jonathan so firmly believes

this next wave of AI technology won't say proactive and far-reaching regulatory response. So 2010 seems like Eons ago, yet as you pointed out in your recent

Atlantic article, algorithms were already capable of causing widespread

and rapid societal disruption even more than a decade ago. Can you remind our younger listeners or perhaps less historically inclined listeners about this flash crash that occurred in 2010 and why it's so relevant today? Well, sure.

And thanks so much for having me on. And yes, let's talk to the youngins. And instead of telling them to get off my lawn, gather round, pull up a lawn chair, and we'll talk. And before we've been talking about flash crash, as long as we're going back

in time, it might be worth talking about the 1980 Morris worm

when Robert Tappen Morris Jr., or the third, I forget his suffix,

created something on the internet that it turned out lodged itself in different unix compatible hosts and then propagated further. And before he knew it, unintended it was sort of everywhere. And that was in 1988, a kind of wake-up call about first unintended consequences,

Mickey Mouse and the broomsticks from Fantasia, but now we're going back to like the late 40s and early 50s. And about the ways in which you could have set it and forget it systems. I mean, he set it and he forgot it until he, you know, was reminded that his children had had many descendants,

and they were all cluttering up to the landscape. Yeah, and it was easy enough to mitigate once people understood what was happening. And of course, it was a small enough community, certainly compared to today's. And one whose equipment was run by network experts,

almost by definition, that they had a way of getting the, there we call it, agent out of the picture. And the, the biggest upshot policy wise was like, wow, these systems are so open. They're so ready for reprogramming or handling traffic

From other destinations and like the basic idea of the internet.

Then and now, is it anybody ought to be able to communicate with anybody else

without intermediaries getting in the way,

which is not the way many other networks are configured. And that could lead to problems. One of the other sort of observations coming out of it was the need for ethics training for people coming on the network. Again, sort of assuming it was a professional community rather than just open it large.

But of course, we know how that story turned out, the network became open to everybody and we celebrate that fact. Even as the underlying protocols weren't meeting fully evolved from very different expectations of who would be using the network, what level of expertise they would have and what they'd be doing.

And as the Morris worm incident shows,

even if you've got a ton of expertise, this guy, you know, has been an MIT professor, things can go or why.

That's where to the flash crash in 2010,

and it shows that within even a limited domain, you have some set of if this then that rules for trading in electronic market, which, you know, that's electronic markets are for. They're not just expecting people to be sitting there waiting to click the minutes. They're like, "Ah, all right. I just read the fed report

and now I'm going to click buy because I have a judgment about what will happen next." But you kind of set it so that these things will operate on their own. It's just they tend to anticipate a world in which everybody else is not evolving or changing and the static world on which you base your if this then that judgment. Or as we start to introduce the phenomenon of machine learning,

and essentially forms a pattern recognition on which you train your machine learning or machine learning model, the very introduction of that model and others introducing their own models could result in very strange transactions that no model could anticipate because all of them anticipated a world without other models. Now that's, I should just give a quick example.

And the flash crash may well have been a good example of this. There are multiple sort of post mortems to it that are somewhat equivocal and what was happening there. But a guy named Michael Eisen discovered at one point around that time that on Amazon there was a used book that was up for sale,

as is often the case on Amazon, nothing shocking there.

But he noticed that it was selling for something like 2.4 million dollars.

Talk about the dangers of one click purchasing if you're not really paying attention and that's an expensive use book. I'm not sure I'd spurred for that. Yeah. Exactly.

And like what's the return policy and the pace of the shipping?

And indeed it was like 2.4 million dollars plus 3.99 shipping. And he was curious because that seemed expensive. And the only other book, copy of the book being offered apparently, was in the high $1 million dollars. So both of them were very expensive.

And he started tracking it day by day. And each seller was slightly raising their price every day. In a way that was, there we say algorithmic, which means it seemed to follow a very straightforward pattern. Seller number two, the cheaper one, was taking whatever the next highest price was

and trying to undercut it by 1%. That appeared to be the rule that was enforced. This is what we call as explainability, not interpretable. We're looking back and just trying to surmise what appears to be going on. And the other seller who was more expensive appeared to be taking the next highest price,

or I guess to say next lowest price, and adding just a little bit on top, like 30% to whatever it was. And what that means is it was jumping by 30% every day. And then the next day the other one was jumping to be 99% of the 30%. And what do we surmise?

This is now trying to explain the explanation. What Eisen surmise was going on was that one was just doing classical economics. They had a copy of the book and they just wanted it to be the cheapest one, but not two cheap, so like a corner gas station. They were just doing 99% of the other one.

Classic race to the bottom that we hope would happen to give consumers surplus between two sellers. And he thought that the other seller probably didn't have the book. And was just going through and quoting 30% higher prices on tons of books. And if anybody should just happen to click to buy it,

They would turn around and go to the other seller, order the book,

and have it delivered to the buyer and collect 30% as a kind of big.

Both are totally rational approaches that when you put them together,

lead to an escalating price spiral into the millions of dollars for a book that should be 20 bucks. So sorry to take so long to explain it, but that this is long before machine learning algorithms might have been deployed for this kind of thing, very simple human level algorithms. And yet the unexpected leads to a systemic kind of surprise that might be when applied financial markets or other circumstances undesirable.

I see this systemically as an analogy to so called technical debt.

The technical debt being, you patch stuff, it works good enough, you patch it a little more. And people start to forget exactly what the elegant idea was behind the whole system like back in the day when you had multiple audio video components in your home entertainment system. You start to forget what all the different wires do and they're all different connectors. That's a lot of technical debt.

At some point you just unplug everything and try to just do it all over again with the components, rather than trying to reverse engineer a theory of what the hell your TV is showing and why. And if that again is done systemically for all sorts of supply chains like that of the Amazon books or for financial markets, unexpected things can happen.

And as I imagine, that's what's pointing us in the direction of talking about what agents are and in today's our got and what therefore might be different.

And what's wild to me too is that you're mentioning this socio technical debt. Morris, we didn't learn our lesson.

This book example, arguably we haven't learned our lesson from our incredible Amazon $2 million use book.

And if we look at the flash crash, if you hear latest state, the latest statements from the SEC chair Gary Gensler, we also haven't quite yet learned our lesson about how even well intentioned over reliance on these set it and forget it approaches can cause those systemic risks. I think I agree with that. And if we're just going to carry the effortism a little further, it's not just we didn't learn our lesson. It's that maybe we learned our lesson, but nobody owned putting the lesson into practice.

People can look back and say, yeah, that was probably bad.

Maybe in the case of the Amazon books. It's like, well, just buyer-aware and like, you know, whatever, there's a market of markets and that will fix itself and maybe Amazon but notice blah blah blah. But when we start thinking about agents acting at the borders of different spheres of operation or responsibility in between different organizations, different firms, different marketplaces, nobody is asked to internalize the risk and maybe nobody does. And I am among those, the book I wrote now almost 15 years ago, the future of the internet and how to stop it.

For time, badly working on the SQL right now called, well, we tried. That celebrated a universe in which people didn't have to be accredited to anybody, to the government, to some platform operator to introduce something new online. You could just do it and see if you could build an audience around it. And that book acknowledged some of the problems, including security ones and ones like the flash crash that can come about. And this is part of how to stop the future.

You know, I'd be nice to have some wise restraints or standards to prevent the worst obvious abuses that we've all learned our lesson. Nobody really wants this from coming about. And yet, it's really hard to make anybody own stuff. And the basic ethos regularly speaking from the mainstreaming of the internet versus much more controllable alternatives, like back in the day, AOL and prodigy and Delphi and computer. The basic idea of the internet was anything not explicitly prohibited is permitted. I call that the Venn diagram cocktail all of digital regulation, because it's a big green oval of all the stuff permitted with a tiny pimento in the middle, which is the handful of stuff you're not allowed to do.

And that's the foundation for all of the benefit and much of the headache we've seen. And a trade-off that is hard to quantify, but for which most of us thought was a pretty good trade-off for a free and open society.

It's just without addressing these problems and allowing for more and more no...

but decision rules themselves to evolve without any oversight or comparison against wait a minute.

And it's this really a great idea and who's getting the bird's eye view of the system. That starts to really show the pain points of everything not prohibited is permitted. And continuing on with our intoxication of just ease or convenience or using the latest technology, we're seeing in addition to technical debt and sociotechnical debt with the introduction of these AI agents.

So can you briefly explain just for folks who perhaps have missed the AI wave so far they've just been drinking too many cocktails perhaps?

What exactly are AI agents?

How is that different from something like chat GPT? And what is it about AI agents that make them scarier than previous instances of let's just bake it down to set it and forget it technologies? Yes, and I'll say upfront, you're going to get different definitions from different people, which is entirely fair, and there's some cool readings that I imagine we could include on the page such as a great round-up paper from Helen Toner, formerly on the opening I board, now it's C-SET talking about AI agents, and Alan Chan and others have done a ton of work on this so we can provide all that.

And then I can give you my own sort of tri-partite definition of an agent.

And for that, basically I think of them as dials and the more the dials are turned towards 11 from zero,

the more we're talking about, this is the weird adjective that makes it sound very esoteric, a Gentic AI. So this is at least my definition, understanding that it's fuzzy.

So the first is the idea of being able to be autonomous, and that is not on or off.

I found myself in the mid-2010s, starting to talk about autonomous agents, instead of autonomous agents, the interested to see whatever automatic transcript generating AI makes of the word autonomous. But yeah, but by that it means instead of specifying exactly what you want to have happened. And therefore, at least as the human instructing a computer what to do, where you are naturally in a position to take responsibility for what it does, because you gave it the orders, you were just giving it some basic goals and asking it to figure out how to advance them.

And that is something that playing with large language models, which, of course, by their name and nature, have language at their core, you give it some language and ask it to spit out other language that you hope will resemble a series of steps of greater specificity than what you asked it to do, that will have it advance the goal. They might even be steps that in turn instruct itself, so it can feed back into itself what it does. But there's this idea of a general goal or high-level plan, and then let it do the rest.

And if it turns out with rather unconventional means to want to do something, and has the means to then instruct itself and possibly pursue it, you can end up with surprises of the sort that a flash crash or an Amazon book purchase, and how it chooses to do things can accomplish. That possibility has been in the realm of science fiction and popular literature for a long time. It's even before you're talking science fiction, you're talking about monkeys, paw, like be careful what you ask for things, or this will date me, you know, Homer Simpson, you want donuts,

says no, I'll give you donuts and start stuffing home or full of donuts as a way of rendering an ironic punishment, of course, in that case Homer's just like this is great. Or donuts the better, let's go. Exactly, exactly. And of course, that even calls to mind Nick Bostrom's 2014 superintelligence example of a paperclip optimizer that turns out destroying the world just so it can make more paper clips.

And I think choosing paper clips is both designed to get us to do a record scratch as we think about a very modest but still high-level goal,

accomplished through any means necessary, and the most kind of maximal extent that just wouldn't have been on our minds.

Again, capturing the accidental nature rather than intentionally doing things...

It may be that to do a whole category of bad things in the world might require before we got to this stage a bunch of expertise,

which would greatly limit already the number of people capable of doing the bad thing. They'd have to be experts to do it or experts consulted to do it,

who then might be in a position to say, why are you asking me these questions?

And even if you figure, well, there's a textbook somewhere that tells you how, you got to go to the trouble of reading it, et cetera, et cetera, whereas here you could conceive of agents that if they are cured of hallucinations, but not otherwise, Gargrel restricted, you can ask them to do pretty terrible things and they will come up with non-illucinatory ways to do it.

So that's the first thing, being autonomous in the sense of independently coming up with their own solutions.

Hello, I'm Lena Kassel from Podcast Football MML Daily, and I'll tell you what you know. This is the end of the end of the end of the end, and the new Bundesliga Saison, again, as a result, is to say, "I'm a sports director."

Because all of the games is still on the kick-based side, and if you ask, kick-based, what is that?

The kick-based is the fourth fantasy football manager, and that's the first one.

It's a real thing, you're a player, you can take a look at the rules, and then you can take the real Bundesliga Profis in your car, so you can decide that you're a footballer and you can understand the reason why the Bundesliga Saison is on the day. It's a very good dream to be able to take a look at the rules and be a member of the team, the kick-based, just download and go to the league. Good kick, and what do you think? To check, check, internet on the map, check, new address, check, and then go to the ballroom.

Between kisten, omelder, and mobile-off-bow-plat-the-strom-for-trag-schnell-liegen, then you go to automatic in-the-grunt-for-Sorghum, which is often totally courteous. By Octopus Energy, Vexels, you're completely self-fulfilled. Now, on Octopus Energy, the E-vaccine, and with the Bundes Code, Octopus 1-1-5, a 15-hour-vexal-bonus-sichan.

And before we go to, I think what's important to point out about that as well, is that we don't even have to imagine the bad actor getting a handle on AI agents to have some bad outcomes.

As we learn from the Morris' worm, as we learn from just our savvy, I guess, or lack of savvy, or lazy Amazon booksellers, just rational uses of these tools. By well-intentioned or neglectful folks can lead to really bad outcomes. And I think that's important to call to the attention of regulators, or to whoever's listening is, we don't have to imagine going to the full paperclip scenario to even imagine some regulatory headaches that weren't addressing in this case. I think you're right. And if we're being really analytic here, we've come up with at least two distinct structural scenarios.

One of which has to do with some model embedded in some system is prompted to do something and given great latitude and determining the means by which to do it. And it picks means that the ends do not justify and that our side of that goes by that Amazon book, whatever it takes by the book. Right. Yeah, well, one point nine is as deep as possible. And it's really hard to specify up front. It might be more effort to say all of the limiting features to make sure it does it right, then to just not have it try to come up with the means, but instead give it the means at which point what are we even doing here.

Being able to have done it successfully for a while and then have it bunk in weird and surprising ways is one of the flowers in the bouquet of large language models and of other sorts of generative AI that just comes with the territory. They'll weirdly even if it has done very well up to that point because these things you never quite know what they're going to give you there the first gump of technology and it's a box of chocolates and you just might get prelings when you hate nuts or one filled with poison right exactly I was going to say it doesn't quite capture the range of things that can go along.

Another piece within the zone that we're dwelling on is that it might still be basically doing what somebody without an even bigger picture of humanity might do that turns out to do unpredictable things because the world has changed.

Often what's changing its world is the presence of other agents from other pe...

That's what I've been calling autonomous autonomous itself can mean different things and I am open to some critique that says use the word for something else.

And in fact, I think that brings us to a second area of egentic AI which is that it can tend to operate outside its sandbox and the first way in which many listeners might have encountered large language models something like chat GPT was on the so called playground.

And you just go and you type that it and that type back at you or maybe there's some product maybe you're accessing some airlines website and there's some helpful assistant and it turns out it's you know GPT powered.

And then you can maybe try to get it to do weird things in the chat with you but it's still just chatting and after all it's just words but part of what has shocked me about how quickly things are moving not just in degree but in kind. It was what used to be called chat GPT plugins may go by a different name now but ways in which you can have the model connect to the world at large automatically and utter words that will make things happen in the world. And so the common those pizza or door dash or east to cart has some API either intentionally meant to connect to something like GPT or not something like GPT beat a path to its door and is acting like it's a human maybe acting on behalf of whoever stood up this chat window.

And this pizza or something else or placing trades on a market. This is what I call a very casual traversing and breaking of the blood brain barrier between just words and it's up to whoever's hearing them as to what they do on that basis and just making it so making it happen in the world.

Busting out of the sandbox and as the difficulties of securing the internet and the generative machines in the old use of the word generative that are connected to it like reprogrammable personal computers or devices like that.

This taught us anything. It has taught us that you shouldn't have those machines able to do things in the world unattended including discouraging their private contents without going back to the user for an affirmation. And that of course is not a solution because users get these prompts all the time grant this permission you know such and such this new hamster dancing app where an image of a hamster will dance on your screen would like to follow in permissions yes or no and it's like yeah whatever I want to see the hamster dance and.

And that's an unresolved problem usually solved in the critical infrastructure context of like you know an old and perhaps no longer so right method was air gapping where you're just like you've got to keep it away from the internet at large both to be instructed by it and maybe to instruct it but rather have a human run that last.

Inch over the gap so that you know what your inputs and outputs are and whatever lessons we might learn from that are not being applied in the rush to make these models.

It's able to be part of a very complicated busy chain of stuff happening in the world and again psych this is the cocktail all of it work it's like why not try it out. But it's really hard to get anybody to internalize the negative prospects of what can happen unpredictably including you know end users who aren't aware of those edge cases or just a patient see the hamster dance.

And that is a real problem and I'd say it's compounded by the fact that there is no inventory of where these.

Are embedded who set them up how they work what the initial expectations were so you can see if they're being used to ways that are unpredictable the classic like using a screwdriver to open a paint can so you know gosh that's what people are using.

I think we may be able to see for alternative or whatever it might be. There's not of that and that's a phenomenon going back to technical data what you were calling socio technical or other debt. I call it intellectual debt that you don't know what these things are going to do and yet.

Just madly building them in to the center blocks where we poor concrete and t...

It's like they're really useful their wholesale not retail you don't know that a machine learning model is part of bringing you whatever experience you just have that's not always going to be a chatbot. It might be somewhere in the middle and it's great until it's not we may discover problems later at which point it's really hard to know where the models are how to remediate that kind of thing and that feels like.

It's not easier to demarcate them as they're getting added or to have a standard for how they identified themselves now rather than waiting for the problem and trying to retrospectively figure it out.

That is then an answer to the eternal question for technologies up and down the scale like these which is they seem to have two phases. The first is two early to tell and the second is two late to do anything about it and you know where do we devote our energies and that is a score for the two early to tell column.

I think two you are example of AI agents being akin to space junk is right on the nose because it's just this instance of.

Let's launch a whole bunch of stuff into space see what happens see what new innovations we can come up with and. And the consequences and now the irony is that the space junk itself is hindering our ability to innovate in space and to try new ideas and to explore further and so this. Failure to anticipate to get ahead of this tragedy of the common tragedy of the space is actually depleting our ability to come up with new and better solutions down the road. More to me the reason the space junk metaphor just felt very on target to me was that space junk but that's more space junk because if it collides with other space junk.

We're going to use more fragments more shrapnel that then ultimately could lead to a little sphere you think start link is cluttering skies a little sphere at classic low worth orbit altitude.

That might make it really hard to leave the planet because there's all this junk around and talk about two early to tell versus two late to do anything about it where we might want to do it.

Also a great example of nobody being responsible for the problem I remember it was a big deal in the late 1990s when space now space command.

Just started tracking junk you know more than X number of feet wide at least so we know where it is that's the inventory point which is. We've got to start somewhere but again nobody really owns doing anything about it no I mean we keep creating these massive agglomerations of junk if we look at the great pacific trash heap or whatever and now we may have. This little own moon of space junk and soon we may have a body of agents that are just acting in some weird fashion that we can predict I think to your point about the interaction between these agents being a really big concern from a regulatory standpoint.

And the fact that the defense came on the pod and discussed how especially in a national security context just not understanding how some of these weapon systems that are agentic in some fashion may interact with one another. And what sort of chaos might we see in that regard. But you're just pointing out even ordering dominoes via an agent could lead to unintended and severe consequences maybe you know what we'll say they're ordering too many toppings and then it gets off the rails but. I think it's very fitting to have a new domino theory in national security law we've coined it we've coined it here it is.

Yeah domino theory part two it's not you know related to the spread of communism.

It's the spread of pepperoni and tomato sauce yes yes and that also. Both hints at some solutions and before we just go there gets to the third quality that I'm thinking.

It tends to adhere in this agentic area which is said and forget it. That the motion of these agents can be. Newtonian rather than Aristotelian instead of you having to push it and keep it going.

The way that for many services and things that you do you have to keep putting a quarter in the machine use of all metaphor.

Sign up for the subscription or something like that it's you know the top doesn't just spin forever.

That provides itself a natural checkpoint because then you also have to stay ...

I don't have to be that way and that is distinct from escaping the sandbox it is distinct from what I'd been calling autonomous about going from high level goals to particular plans and implementations.

And sadly this also might be described as a form of autonomy because.

But but I think of it more as auto pilot kind of thing. It's not about making discrete decisions that's part of the first aspect, but rather that the momentum is inertial and whoever got it started might be long gone.

And their disappearance doesn't affect the path of the program and there are any number of ways that could be happening. I I realize a duck might come down and I have fifty dollars taken from me if I dare use the word block chain, but when you talk about computational block chains or distributed autonomous organizations.

Things like that that have been kicking around solutions looking for problems and many would say maybe they have found it others would say they haven't.

But these are tools by which you can set something up.

In doubt financially enough to run like a cemetery plot for a long time and then you piece out.

And then you've just got these vehicles running around doing stuff and.

Turning off an entire block chain in order to stop a bad propagation both seems excessive and possibly not possible given the distributed unknown nature of things like that.

And then you can just pop up just like you can reserve it to main name for the next 10 years and just front the money you could say I great I'll buy 10 years worth of computing the makes such and such happen. And it may be very hard to get through the tango, especially if somebody wants something to persist as against anybody trying to stop it to figure out even where it's. And that set it and forget it nature could create massive headaches when trying to remediate obviously bad harmful individual instances of behaviors online that are out of control.

And that could include like, you know, another example would just be to really use an ancient narrative formulation, levying a curse upon somebody. And now you're going to curse us. Yeah, yeah, I mean, this is like, you know, animating a golfer something, but you know, you levying a curse and you're just like I'm willing to put. You know, you're not just one person, but many. It's just catch what will feel like the full fervor of discontent across hundreds of people just swarming them every time they dare to utter something online.

That sounds bad. And you know, this is me not. I'm not being all that creative and coming up with this one particular fact pattern, but exactly what you would do if that happened and the person who got it started had long wandered away or forgotten their grudge.

And you know, I think a lot of people just in the past week if we're going to date this podcast as it persists online through the ages.

There were some students who were playing with the new. I forget which companies it was, might have been meadows or somebody's new attempt at Google Glass, you know, eyeglasses that tell you what's going on. And they just hooked it up to like pimmies or some one of these regrettable facial identification services and as you walk down the street, it's just identifying people for you and running and getting a short dossier on the semi anonymity that we depend on in environments amongst strangers is just gone.

And then you take moments of conflict or road rage which already completely lose their context and people without the phone start pounding it, it goes viral, et cetera, et cetera. All right, well, a lot of people start cursing like in the way we're describing cursing. It's just making more efficient and persistent. The extremely regrettable dynamics already that we've identified the social media and the pyleons that it occasions as each person wants to express moral disapproval somebody for one of their less noble moments in interacting with another human.

That's what so scary to me is the scale of reliance on age and four, a minima...

You just set off an army and that army exists forever potentially or exists for a very long time and to your point, having that army follow you wherever you go is certainly not a spot where I want to be and so before we let you go we have to we have to at least learn what's one solution give us some hope about.

How do we at least pursue this innovative agenda while also not pursuing a world full of curses and hexes and whatever bad ailments we can imagine.

Well, everything comes in three, so let me try giving three rough areas to sketch out here.

I think the first is to take seriously the word agents in its legal and social sense not just in its technical sense or the definition we've been slowly spinning out here on the podcast, which is to say an agent is meant to represent a person.

And when they do, they often owe special duties under the law and in our moral expectation that they will place the lawful expectations and interests of their principle over their own.

If they don't, they have a conflict of interest and they are resolving the conflict in a way that they should not, which is to their own advantage. And that's why ideally if somebody is your agent in picking stocks for you for your retirement fund, so you'll have a decent amount to retire on later, you would not want them picking them on the basis of what commissions they get. There's some jack-a-low branch in Florida, which is not a great investment, but they get paid a little kickback for it. That is an agent principal problem, principal.

And I think it is utterly understudied, under theorized what the duties of these sorts of agents should be to be attentive to the genuine needs of their principles, respectively, and how to hold them to that.

And especially when you see that a lot of these agents might be free of charge, just like social media networks are free and email is free, the way to monetize it may be through having a separate set of interests. And this agent, which is now quite literally whispering in your ear as you go about the world, rendering advice to acting like your friend, it might not be. And so with Jack Balkan, who coined the term information in the fiduciaries, I've been doing work as well, saying, "All right, what would that look like? What would it look like to be a fiduciary?"

And we are looking for solutions and you can have a website up that we could add in the links below the podcast to actually just see what are people's expectations when they're online.

Right now, what are your expectations of an agent? Nobody has thought about it, including the consumers, but they slip into these are my friends kind of thing, because the thing is they have to purmorphically design to be very friendly and attentive and how can I help and it has infinite patience.

So that's one cluster of solutions having to do with not letting agents be duplicitous in their design and incentives as they are offered to people in the world.

A second area that I outlined is modeled after old network traffic and rally, which is highly decentralized on the internet, and that means you run into a problem where packets might get set loose to be routed by one hop at a time through all sorts of different technical jurisdictions. And if it's misconfigured a certain way, it's possible it could just go around and around forever, like the old Charlie and the MTA, like just keep circling around. And for internet routing, there's a technical solution called TTL time to live, and packets have a default number of hops that they expect to make, and if they are continuing to hop and not getting to their destination after 256 hops or whatever.

It's understood that they will die, that the router that catches them, they're 256 jump is just like, you know what, it's just, it's not you, it's me, it's not working out packet will not get forwarded, which prevents the space junk problem of packets in an environment that is highly distributed and anybody's permitted to launch packets into it. TTL is a cool standard that's just part of the furniture now that headed off massive problems, and there should be biology a TTL for agentic behavior. After it's done a certain number of steps, it ought to, you know, chill out until it is reanimated through human intervention or something.

They may be there'd be more steps for some kind of bots than for others, but ...

That might be the kind of thing where you're like, you wonder what you're up to, and that gets to the third solution, which is having a means of identifying a gentic behavior in the digital environment as distinct from human behavior.

Ultimately, a sociotechnical judgment, rather than just an easy cut and dried categorization for everything.

And the papers and people I mentioned at the beginning of the podcast, folks like Helen Toner and Alan Chan, they both include in their sorts of evaluations.

Here, some way of identifying a gentic processes where they live. This is where the bottle is running, and this is, you know, what it's doing. This is a particular instance of chat GPT, and it has a little license plate on it kind of thing. I'm thinking of a complementary mode of identification that might be through old-fashioned network protocols, because a lot of the harms we're talking about happen over the network.

It's when something is communicating to something else, and there ought to be an easy wrapper around a given packet of data or of instruction, which instruction is also a form of data online that says, by the way,

I'm an agent, and I was emanated by an agent, or I am destined for an agent, and this is a means of reaching the parent process or person.

And that might be behind layers of indirection. I don't think this necessarily entails having everybody have to identify themselves or something, license plates themselves provide a legally protected layer of indirection. So you could tell the authorities or other people, this is the license of the car that caught me off, but it doesn't immediately, like, you know, who the person in the car is, all sorts of things you can do. But this is the time before these agents are everywhere, and you're trying to retrofit how they identify themselves to come up with incentives to identify incentives and standards structure of how to identify themselves online. You could say something like, if you do, whatever the evolving tort regime is, for things going awry and who will be held accountable, hey, if you've got this label in place,

there will be a cap on just how much harm you some player in this multifaceted ecosystem that you're somehow contributing to the way it works, there'll be a cap on your liability.

That alone could provide for opting in to labels, especially if we see a world in which most agents are coming from concentrated platforms that are consumer facing.

If you're talking about just some weird bespoke agent that got spun up some other way, the fact that it doesn't have a license plate could make it a focal point of skepticism, especially, I don't mean to spy the authorities by Domino's pizza. Hey, I'd like a thousand pizzas delivered here. Wait, who are you? Not your business. It's going to start half to get a license before you order pizza. That's been the hell I die on. You need a license to know if you know the person ordering the pizza is the person that's going to be enjoying the pizza.

This is excellent. I mean, we've got our inventory. We've got making sure we get rid of zombie AI agents and making sure AI agents have their license. I think our listeners have a lot of homework to do, and are going to be on pins and needles waiting for your next paper so that we can have you back and dive into the weeds there, but unfortunately we're going to have to leave it there, but thank you again for coming on. This was a hoop. Thanks Kevin, delighted to talk about this stuff and eager for other examples that might be brewing out there that people are trying to work through.

Of course, of course, and we'll keep it coming in a steady supply to a domino's near you. Very good. The law fair podcast is producing cooperation with the Brookings institution. You can get ad free versions of this and other law fair podcasts by becoming a law fair material supporter through our website law fairmedia.org/support. You'll also get access to special events and other content available only to our supporters. Please rate and review us wherever you get your podcasts. Look out for our other podcasts, including rational security, chatter, allies, and the aftermath. Our latest law fair presents podcast series on the government's response to January 6. Check out our written work at law fairmedia.org. The podcast is edited by Jen Pacha, and your audio engineer this episode was "Nome Osban of Go Romeo."

Our theme song is from Alibi Music, as always.

Compare and Explore