The Lawfare Podcast
The Lawfare Podcast

Lawfare Daily: Vinh Nguyen, Elham Tabassi, and Kat Duffy on How to Design a Better AI Regulator

1h ago41:016,938 words
0:000:00

Lawfare Senior Editor Kate Klonick is joined by three guests to discuss their recent article for the Council on Foreign Relations on the FINRA-style AI regulator reportedly under White House review: V...

Transcript

EN

[BEEPING]

[BEEPING]

I heard you will have a mobile phone call, but not just come from here.

Who is with your ex ex-girls dreaming? My second baby? No, I'm going to leave you alone. [BEEPING] Just a little bit.

With mobile phone in real life, Frank, from the beginning of the "Owner Mindest" contract, I would say, "Kunba, 14-0 in Monat." Frank, nine. Frank.

[BEEPING]

β€œIt is not, I think, surprising, that where this administration”

want to lean is in more of an industry-focused industry-friendly mechanism, as opposed to one that leaned into deepening federal government, sort of bureaucracy or agency interactions, or working on international cooperation as the central premise. It's the law fair podcast.

I'm Kate Klemick, senior editor of law fair. And I'm here today with Vin-Wen, former chief AI officer at SA, and now see how far our senior fellow for AI, Elham Tabasi, former chief AI advisor at NIST, and now director Brookings AI and emerging technology initiative.

I'm Kat Duffy, see how far our senior fellow for digital and cybersecurity policy and director of lead AI. You need to have clear information, you know, evaluation, do it independently, so we can actually understand and have a level of

symmetric information flow of sort across, you know, not only the Frontier Labs, the government, you know, the public, the industry. But we need more voices outside of the Frontier Labs and the government, right? Today, we're talking about their recent piece

on the Vin-Wen-Wen style AI regulator, that the White House may actually be about to stand up.

It's one of the first pieces that's been written about

what's wrong with the way an AI regulator is being designed and how to get it right. So I wanted to just kind of set up the stakes for us. The MSISabi is the founder and the former CEO of Google DeepMam published this framework on July 14th.

That called for a US-led Frontier AI standards body that's modeled on Findra. And a few days later, Bloomberg reported that the White House, along with Treasury Secretary Basant being involved, is reviewing a version of it.

And in the same window, we've seen Anthropics, CEO, proposed an FAA style model, Sam Altman, an OpenAI, jumped in with something closer to nuclear governance. And then, you know, obviously, we have, as I said, before DeepMam and CEO backed this Findra model.

β€œSo what is the actual momentum towards a real AI regulator?”

How close is close, basically?

Is that maybe the wrong question? And really about, you know, which model is going to win here for regulation? How close is close? It's a great question.

We're not totally sure. There's also a lot in the news this week, as well, about the Frontier companies and some other AI leaders meeting at the White House. So I might, it might be wise for Vintu sort of explain

that context. But I would say, one thing that's interesting, it's not just the Findra style proposal. It's that you saw the CEO of Anthropics, right, a public proposal for an FAA style type governance model

for artificial intelligence. Then shortly after that, you saw Sam Altman come out with his uped in the Financial Times.

β€œThis CEO of OpenAI calling for a different type of national,”

or his was more of an international regulatory model that was modeled more often in nuclear governance. Then you see the CEO of DeepMam, come out with his proposal that has a Findra style approach. Then we hear tell that inside of Treasury,

that the Secretary of the Treasury is also considering a governance proposal and that that reflects a Findra style model as well. And so I think what we're seeing here, it's not clear yet on the timing, but when you have the three major frontier labs,

there's CEOs all proposing different types of governance models and the Secretary of the Treasury, taking something on, especially in an administration that has been driven by thoughts around finance, right, and financial concerns.

It's clear that we are having a real uptick and an openness to thinking about frontier model governance, which is notable given what an anti-governance narrative this administration began with. So it just feels like we're seeing a shifting tide.

It's not clear exactly where that tide will go. Yeah, I want to add to a cat. Note it is that if you really think about all these models that really designing to solve several problems in some ways, one is how to garner harness the benefits

of AI and the other size, how to manage to risk. And they all have different models, institutional models.

None of these models actually noted the design patterns

and things that are needed to be credible,

to be legitimate, and to provide a level of independence. I think everyone wanted that, but the design is not there.

β€œThat's why we're here to articulate whatever the models”

we need to consider these features to ensure the natural security and economic competitiveness for us and for our allies in the long run. What I think is quite confusing is that currently many of the frontier labs are working with the White House

on the so-called voluntary, but a lot of industry will call that not voluntary, like mandatory model evaluation. And that process and mechanism is solving only a very specific problem. And that is to manage the risk of cybersecurity. It doesn't manage anything else.

It doesn't really calibrate the benefit of cybersecurity for defensive purposes, it's really designed to manage and to evaluate cybersecurity risk, to the ability to produce vulnerabilities, and be able to change them and build exploits at the ability

for AI to autonomously exploit a certain system. So it's quite specific. And so I think when people are thinking of all these models, it's like it will fix everything, but in reality, each of the models

is trying to fix a set of the problems, but none of the models actually laid out the design patterns needed for the models to be successful. That's a great point. And I want to follow up on that specifically.

The way that your piece comes at this, there's this fin-rust style initiative that's getting attention as we talked before. And Cat's point is that this is really just one of several governance moments that is happening at once.

β€œAnd so can I ask a basic question though, why fin-rust specifically?”

When I first read the piece,

I took finner to be less the governance blueprint than a convenient hook kind of for the news moment. So what actually makes fin-re attractive or not attractive as a model here? What is it offered that the other governance proposals

that are on the board don't have right now? This model has been used to my understanding for other cases when it was complex and legislations were immature and the Congress could enact in the right time

to address that. So given that this is all of the things about AI governance as Cat and when we talk about this, there is a lot of new and complex things to handle that that none of the other models origin can address this.

It seems like that this thing is that handing it in industry so they can act faster than any other non-industry standard body or governance body to act on that. What having a government oversight through

so it's not just completely industry and the fact that it has been tried for other complex technical topics. So it became something that we look at. So it brings some sort of the flexibility.

Some of those flexibility are good. Some of those flexibility still apply but I think everybody also agrees that this is not the perfect or the only model and there is still a lot of the questions

and a lot of the adjustment needed to make it work for AI. Kate, I would just add to that. You know, the OpenAI premise really relies on international cooperation. The anthropic premise would rely on highly centralized

governmental control that would be fairly bureaucratic in nature and the deep mind/bessent. We don't know exactly how they're connected to each other versus separate. But it is not I think surprising

that where this administration would want to lean is in more of an industry-focused industry-friendly mechanism as opposed to one that leaned into deepening federal government sort of bureaucracy or agency interactions or working on international

cooperation as the central premise.

β€œSo I think just when you look at those three proposals”

from that lens, it's easy to see why a finrestile model would land better in this administration than other governance proposals. Yeah, I think that those are all great. I don't have five big kind of things that you say are

unresolved questions in using a finrest-type model. And I'm going to kind of, I'm going to put independence

aside for a second because personally, I kind of am the

most interested in that. So we'll close out with kind of that question. You have an independence, national security, concerns, legitimacy, access and talent and measurement to kind of split this groups,

skills from standards, agencies,

Health, and then your time and the NSA and cat.

You're kind of your immense knowledge around at CFR.

β€œLet's kind of like, pardon these off a little bit and like,”

I'm going to come back to you for measurement. Cat, I'm going to come back to you for legitimacy, particularly across kind of international kind of stakeholders and things like that.

But then I really want to dig in for a second

into the national security aspect and your thoughts on the role that a body like this could have. The piece's national security section reads like the one that keeps you up and night. You wrote that Miss Handling this could turn the body into

a bespoke intelligence community subsidiary. But overcorrecting the other way means that the US learns about dangerous capabilities from an adversary first. And from where you were at NSA, what failure mode is kind of

the more likely if this gets built very quickly without much forethought. Well, over the alternative is what will be built and not this one if there is a national security failure.

And for the national security community, there are many considerations.

Number one is the fusion of capability of very powerful

capabilities that China, Russia, Iran or Korea, criminals, cartels, all the bad guys you can use in order to do harm against the United States. If they are using changing different frontier models and be able to scale cyber operations

to augment to enhance the bio defense weapons program or to scale a fraud and scam operations to harm Americans, those will likely be the argument for the National Security community to have a huge say and be able to make decisions on that.

β€œAnd so you have to make sure whatever we build”

that the National Security components need to be getting. Not solving that, we're just getting you an organization institution that has the independence

and legitimacy and the first National Security crisis

will absolutely destroy that set up. The second one is that if the National Security is open to integrate in this new model, right? Because it's not just National Security, but it's really our economic prosperity as well.

And our economic and strategic competitiveness as well. So it's not just one or the other, we need to learn how to balance all of this and integrate to get the optimal outcome for the United States. But if we do this, the first really threat

would be that the adversaries will go after this organization and be able to exploit the data, extract information that would be non-public. I think that would be secondary,

β€œbut the National Security community will have a say”

if adversaries are using capabilities to harm us or that adversaries go after these capabilities and steal them. And so those are the consideration that the National community will have to jump in. Yeah, I wanted to kind of talk about some of those capabilities

and how we measure how good they are. So I wanted to get into the next part that I kind of forecast which is the measurements suggest in that you have and that you focus on. And of course, you are at NIST for a long time.

I'm super fascinated to know kind of how you see the standard setting for this kind of taking place. How we even decide how great these products are or where they're difficult these lie. I know that we call that typically just for listeners.

A lot of that is measured by benchmarks ways that we kind of just over time measure how these models are performing. And as you kind of state in the piece, there's no settled science of frontier assessment on this part. So the benchmarks are still being kind of laid out.

So why is this? This is not just a technical footnote. I kind of want to point out. This is really for to understand what needs to happen with the governance. Can you kind of lay that out for us a little bit?

So you at the beginning you ask why I've been drawn by others and we talked about the complexity of the field. But if the A was stood up, the randomized trial experiments has been more or less settled. Great lieability, so on and so on.

The measurement discipline has existed for the regulator to come and try to figure out what to do. But for this body for AI in general, we are being asked the committee being asked as body or whatever body is being stood up is being asked to build the ruler

and use it both at the same time. And do that for object of measurement that is changing faster than we can figure out how to build the ruler and use that.

That's structural difference that we don't have that underlying ruler

and how to use it compared to the other fields.

β€œThis is not just a measurement problem downstream of the institutional design.”

It is something that the institution design is actually standing on. So you talked about the measurement science, me, mature, benchmark saturates. We have the model strange and all this. I want to just bring up at least two other points. One is when you start certifying and on regulating all this,

a certification may mean different things to the body, to the people that build that and to the public. So when a body or any certifier comes and say that the model is passed, and that was one of my concerns when we were standing up the AC at this, when the test comes and says that the model is passed,

what public here is that this is safe. Another point that we want to make is that, so we talk about the science to be immature, but the science doesn't just stay immature.

The act of regulating it can actually degrade it.

And that's because, you know, we have seen it with all sorts of other testing, benchmarks that the moment a benchmark becomes a gate. It also became a training target. And there's a lot of reasons for that,

β€œbut also optimizing to pass is much cheaper than being safe.”

And contamination is often unsulsifiable from outside, not being able to see and figure out how the test is being done. And the labs and particularly in this case, for the labs, public threshold can measure internally and submit only when they know that they can pass,

and they can be clear. So we have to be careful that how we design all of these things, and all of the testing and benchmark and how to run the tests, in a way that it doesn't erode the same day, that you just got adopted.

One thing that I really appreciate when I'm speaking with measurement experts is that those who really understand how you build measurements, how you build methodologies, like what rigor looks like in that, when they are producing measurements,

what they are implicitly saying is, this is what we could measure. It doesn't reflect the things we couldn't measure, and also hears why either of those things matters. But it's a very nuanced statement.

And there's so much that matters that can't necessarily be measured,

β€œand then we get into what I think we all worry about”

with just sort of like compliance theater or security theater. Another thing I want to add to that, is the whole field of measurement and metrology, the scope of what you measure. The unit of measurement is clear.

The measurement is about measuring specified property in the specified environment. Here we are, we just talk about measuring capability of what, in what conditions, in what interactions. So you add the generalizability on top of that,

but even what we measure is not very clear. So I just kind of want to just take a second and kind of reframe what everyone just said, and then also kind of put this in the context of why this matters for national security,

because I don't think it's necessarily always very clear.

One is that I think both what Kat and Am you were saying is that essentially these measurements are incredibly human to a very real degree. They are reflections of what we can recognize as things that are measurable, that we have measured,

and then it's like itself a filling prophecy. The things that we measure and can measure then become the things that we decide are safe. And the idea, as you said, Kat, that we can measure this, it signals to the public that, okay,

there's only A through D things that need to be measured. Rather than there are just this world of unknowns, that we might want to be measuring. I mean, to kind of give a very real metaphor for kind of what this is.

It's like imagine it's like any drug type of like measurement or any type of drug safety type of testing, but like think of like the litamide, right? Like, okay, like across all of these various metrics, that drug, if you're not familiar with the litamide,

it was a drug for nausea and kind of fatigue that was given to a pregnant women, and it was not actually approved by the FDA, but it was used massively around the world resulting in huge catastrophic birth defects

in children missing limbs and roam. It was just clearly not safe for pregnancy, but there was a period in which people for invariant types of measurements thought that this was safe. Because in the period in which they were testing it, it looked safe.

Women didn't seem to be having negative side effects. Well, weight nine months, and then you will see the negative side effects of the litamide. That is a really kind of like stark will obviously. We should have waited nine months type of thing,

but I want to kind of point that up because maybe we don't know what the obvious period of gestation is, so to speak, for AI. Right, we don't know what the next thing is going to be.

We don't know what the unknown's are going to be.

You can measure all these things,

β€œand there's a false sense of security that you're passing”

on to consumers about what they're taking in. A false sense of security that you're giving back to a national security apparatus about what they can deploy because you're kind of telling them, "Well, we measured these things and everything looks fine."

And that there's like kind of this latent risk that's like kind of like imparted this.

And just finally, to kind of make sure that

I'm understand this correctly, I kind of spell it off for the audience, then to like the point that you're kind of saying about a lot of this is that there's a cat and mouse game. That once you release a lot of these kinds of ideas and where these measurements are taking place,

they're gameable, right? That they're gameable by adversaries and gameable by other types of things. So is that all kind of seem correct to you and if that's I'm you're nodding so I can see you.

But like, I want to also take this to the legitimacy point

β€œbecause I really think that this weaves in well”

to the point that you make about legitimacy and to the national security components and how you develop international standards.

These are really complicated threads

and really complicated things to make neatly fit when you have something that can be both used in an adversary context but which the best practices would counsel that you're sharing information with your adversary is to in order to set kind of these measurement standards.

So cat, can you tell us a little bit about kind of creating some type of legitimacy and what's happening kind of between nations and between kind of cooperating bodies and maybe even between enterprise models

and how this is developing? Sure, I would say, you know, and it's so fundamentally tied into independence, right? These are two sides of the same coin. But if what we end up with

is some sort of standards body

that really only reflects US government conversation with the frontier companies and it is designed among them, then you won't have political legitimacy either domestically or internationally. So on the domestic front, you really need to make sure

that you are bringing in other voices and other areas of expertise that absolutely reflect the public interest and different stakeholders roles. And we talk about that a little bit in the section where we talk about independence

but the challenge with not doing that in alignment with partners and allies is that you then end up with, you know, one of the two leading governments in the world in terms of AI production within their borders, you know, the leading industry players in the world

and you have a technical body or a technical standard. But political legitimacy is what takes that from being a technical standard to a trusted standard or a technical body to a trusted body.

And if that body or whatever standards it might create only reflect the incentives in the inputs of the US government with leading US companies, our partners and allies will not be warm. I think to the embracing of that

because they will have their own concerns that they would want reflected. So I would hope that this would be the beginning of a discussion about what we would need to look at domestically but then also how we would build that out

with partners and allies. Hello, I'm Lena Kassel from podcast "Fusbala-Memail Daily" and I'm going to tell you what it is. So far, this is the new Bundesliga season because of the transition, the mental and sports directors.

Because all of a sudden, it is also the kick-based side, and if you ask kick-based, that's it. The kick-based is the fourth fantasy "Fusbal Manager" and that's the principle. It would be very simple.

You could use your own rules or use the real Bundesliga profile in your body. So it's time for you to know that you're a footballer and the content of the Bundesliga season on the day is one August.

You're going to have to use your own rules and both sides of the table with the kick-based side, just download it and let's go. Good kick and fun. Just an ice-traum.

β€œThe mobile function in the real life is not a tromble.”

Even in the Switzerland. And so can we just dig in a little bit more also to kind of the independent stream of this, which is that you kind of argue that the fix for independence as you just stated is structural.

Like, right now, just compositional.

So things like Fynra require public governors

that outnumber industry ones, right? And so I actually kind of wanted to talk about some of the, you know, when I think of Fynra, I think of that. I really came into knowledge about Fynra, right?

As kind of coming out of college, not to date myself, but like the failures of Fynra. We're really kind of how I came to know Fynra and because of the financial crisis in 2008,

β€œand like why hadn't this kind of Fyn protected from my regulators?”

And so one of the things that really came to the fore, I remember that period was like, well, it's become like, it's become a capture organization. You spend some time talking about how, of course, in any model like Fynra,

there is that concern. And how would you kind of protect that independence in this context? In creating legitimacy and diplomacy, should a body like this have people that are outside the US on it?

Should it have a certain number of people that are in industry? Should industry be kind of the frame in which that role is limited? How do you see this happening? So my short answer to that would just be that ensuring diversification in the composition of the body,

without ensuring that you have separated out. The structural mechanisms from financial mechanisms from areas of expertise, there's a lot more that has to go in to making something be functional, sustainable, and trustworthy.

β€œAnd I think Fynra is a sort of failure mode”

of because public sector representatives, like outnumber industry representatives, I think at this point 10 to 11,

the implicit assumption or the critical assumption underpinning that,

is that public sector representatives and industry representatives wouldn't have aligned incentives. And that's not necessarily true, right? If you've got one state that really relies on a particular industry, then you can't think of that governor

as representing the whole of the public interest when there might be a significant chunk of the public interest that is not represented by an alignment with that industry. It's just not as simple as make sure you've got more public sector actors than private sector actors.

Yeah, I agree with Woodcat. I think this is really about the kind of structural element for independence rather than setting up different bodies is really about the talent as well.

Because right now, the reality is that we just don't have the talent

and talent has scooped up by the labs, question why don't we don't have credible competence third-party auditors because they hire all the competence auditors, right? It makes sense, it makes sense. But we can set up all sorts of boards, right?

And the different construction and constitution of the board, but if the structure of the financial investment, the talent and the binding access not there, the boards will be like a none of the factless committee. And so I think we can solve the board's constitution later

but if we don't solve the structural components, it really doesn't matter.

β€œSo that's what we're really aiming for on what would make”

an organization like this to be credible, legitimate independence in the long term. The idea, I think, is that, again, because of the talent and because of the knowledge is sits in a sort of a osymetric way in industry, how be able to put strap and use that knowledge

and that pipeline of the talent that they know that how to do the testing

and I have more to see on talent in a second.

And the governance of the board can try to kind of achieve the independence and achieve so many of these problems too. So the board is not the one that writes the rule or the one that sort of approves the rule or comes into address the disputes,

the question of open source. Or I should say open rate model too, because the way it's at least written right now, or at this is being talked right now, it's not covering all of the open rate model.

That is also an important point to be considered in all of this. Okay, so actually I'm really glad that you brought up open rates and open source because people can flate these two ideas constantly and so it's worth just kind of taking a beat and explaining both of these for listeners.

A model's weights are the numerical parameters, essentially billions of dials is one way of thinking at it that are set during training that determine how the model turns and input into an output. If the inputs are variable like XYZ,

the weights are the coefficients in front of the variable to put it in like if you remember your algebra. And so that's the mechanism that makes it work. So there's long band a strong open source ethos in tech and tech law.

That's different than an open weights ethos or something.

So let's just keep these separate. So an open source ethos in tech and tech law is the idea that code and information should be shared freely and increasingly, however, that's getting pointed towards AI models. Release the weights, not just the finished product.

This isn't kind of a hypothetical. And, throughout my current accused several Chinese labs recently, including Moonchot, which is the company that's behind if you heard about it, Kimmy, which is the new model that's

basically a clone of clot.

That basically, improperly harvesting clots outputs to train their own models. And the White House made similar allegations about this publicly. Moonchot hasn't confirmed that, so I'll just flag it really kind of as contested rather than subtle,

β€œbut it's pretty clear that's what happened.”

And this part of why open weights has become such a flash point. And separately, kind of several other Chinese labs that are all kind of competing in this space have said outright that they intend to release their weights openly, right? So you have kind of this China versus US narrative

that's starting around open weights. And here's the tension. The open source impulse kind of runs directly into a cybersecurity concern.

The more you disclose about how a powerful model works,

the more you're basically handling adversaries a map to find and exploit its vulnerabilities before anyone can catch them. And so personally, I am instinctively and have been for 20 years instinctively pro open access

and open source, including on copyright. And so this kind of actually is a weird change in my priors to what I need jerk would be pro open weights, but this is put me in kind of an uncomfortable spot because there's a real difference between handy and a library card, right?

And like passing a loaded gun around to a group of friends. Maybe not something everyone should have open access to. So how should a thin restion model or really kind of any governance model handle the open weights problem then and I'm going to start with you?

β€œI think the way to see this clearly is that there are really two levels, right?”

The one level is really about controlling the capabilities. What capabilities being developed is a cyber, bio, coding. You need to have clear information. You know, evaluation, do it independently. So we can actually understand and have a level of symmetric information

flow of sort across, you know, not only the frontty laughs the governments, you know, the public, the industry. But we need more voices outside of the frontty laughs and the governments, right? If you look at the cyber community, the cyber community did not have a big voice. In the capabilities development,

a cyber offensive and defensive, right? And so when they are dealing with like surprises from mythos, a GPT 5.6 or newer models, they realize that they have to deal with all the downstream impact

and effect on things that they didn't have a voice in the first place.

So there's like one layer of sustainability. The second layer is really about the open weight that you're discussing. It's really about the fusion controlling the diffusion of the capabilities.

β€œThere are many toolkits to manage the fusion, right?”

Well, or not well. But there are many ways that you can manage that. And I think like we combine capabilities and the fusion into like one big buckets of stuff. And then we have a very hard way to untangle

because then we have dangerous capabilities. Then we have adversaries, distillations. And then we have the fusions of open weight models that are built by Chinese companies. And suddenly we have like natural security,

economic deployment development in one like big complex, you know, basket of stuff. We should rethink and like break it out and say there are capabilities, controls, and have voices in that place for independence legitimacy in that sector.

And that is that Fundra style component. The Fundra style element will help to inform what to do with the diffusion. It doesn't solve the diffusion. It will help augment and inform that discussion. So this is how I think about it.

Otherwise we just blend up all together. It's like really hard to decipher. You know, as someone who can't like I ran my first symposium on Linux in 97, right? So as someone who who started out

really in olden days open source discussions and then now or in a mode where people are referring to open weighted models as being open source AI. I lose it a little bit because I'm like this is not open source, right? But arguing about what open source isn't is not is a

stalwart of the open source community. I mean, that's like that's just a day that ends in why.

I think where we need to move the discussion is around what openness

is going to look like in artificial intelligence.

I don't think that we will have open source in the way that we've traditionally thought about it in terms of software development. And I don't think that we should conflate the gains that we had from open source in software development. Open weighted models do not magically achieve those same gains

because they have open weights.

β€œSo I think part of this is also just the immaturity of our lexicon”

at the moment in terms of us trying to find where there's correlation. But we're going to be need to become more nuanced. We're going to think about openness about models about vetting in a space like that differently.

Then we are going to think about it when we're looking at

a small localized language model, perhaps for like a minority language, right? Or a medical system that you're going to use for a particular demographic. These are all going to require distinct approaches in some form of fashion. And some elements of that will be extremely sensitive in terms of national security and areas where we have traditionally agreed

that nation states should collaborate and govern that area together because of the immense dangers that it poses to the public. So bioweapons, chemical weapons, nuclear, right? And there's going to be other areas where we have traditionally said, no, the public should be engaged with this.

And there should be a more deliberative and democratic and open process. And there will be other areas where we're leaning in more on that. So I think for me, my thing is open source is not equal open weighted.

β€œAnd neither of these terms is going to be the only thing that we need”

to talk about AI. Great. So I'm going to quickly kind of put you guys on the spot and ask you, what it is you kind of see is maybe low hanging fruit from the things that you mentioned as necessities that this like a finiro type administrator or AI governance system should kind of take into consideration.

I'll help. I will start with you. We don't have clarity on exactly what to do, but that's okay. We should start somewhere and as a starting point and iterate on that. The the Nord story, the object, the purpose that I put in front of myself

to go towards that is basically building the science and practice of AI evaluation.

And that includes strengthening and advancing the scientific foundations of AI measurement, but also turning evaluation into a profession. And I truly believe that if we can do this AI evaluation, but itself can be an economy and a business that can actually have talent to have infrastructure and business can make a lot of money.

Yeah, my one quick action that we all can take is that the current administration processes don't have a way to do with bio coming up, right? And so the executive order is really unsyber, not unenbio. And so any bio surprise will require an adaptation of sort. And so I think today, you know, even using a finiro structure,

we can practice like a rapid respond team absurd to address and provide a level of independence and legitimate assessment evaluation on the bio threat, right? The problems is immense. We need more voices to come in to check each other, right? To create a level balance in the system.

And so a rapid respond team, so if there's a bio thing or bio claim, I think we should have 18 to be able to pull evaluators together, academics, be able to work with the media to say, hey, these are the clients, but here are like the things that we can do it, you know, convert independently.

I think that will bring a lot of trust in their legitimacy to the American people, but also without allies as well. Yeah, and I would, I would maybe, you know, conclude by saying, I hope people will go read the article,

β€œbut I think what was important to all three of us in writing this article,”

it's not, it's not just us critiquing what has been proposed. It is us saying concretely, these are considerations that you could have in mind. This is what you could build in to respond to this critique, right? It is really, really easy to tear other people's ideas down. I think it's more important to come in with a constructive approach to saying,

yes, and, you know, to make that work, you would also need this. And so for me, no matter what get set up, the importance of separating out the safety specifications, from the sustainability of the infrastructure, from who is doing the evaluating,

those things need to live, they work together,

They need to function independently of each other,

those who are evaluating should not be the ones who were determining

what safety spec should exist,

β€œand those who are, like the AI companies in particular,”

should not be the ones who were defining what the safety specification is for any particular subject matter area, right? So you want, you want those things to be standardized, because then that's also testing an ecosystem that can operate across open-weighted models or closed models.

It's not limited to any particular company.

And then for me, the other component of that then becomes, even if you have those components separated, so you do have the ingredients for legitimacy. You still need to have buy-in,

β€œand you need to have political legitimacy as well.”

That requires really embedding allies, partners, and a public sector focus, and experts who are thinking about societal impacts into that architecture from the outset, not trying to retrofit them in later. Yeah, I really like this framing.

Thank you so much for coming on. I think it's great to have your expert voices in this mix, telling kind of the powers to be the administration, and the frontier companies AI industry generally, kind of where the weaknesses is in the lot of this,

that there isn't a need to reinvent the wheel on a lot of these things we can learn from kind of the problems that have served us in the past, when we've tried to create structure for these types of things. And to that, I'm very happy to have been able to pick your brains about it.

Thank you all so much for coming on. Thank you for having us. Thank you. Thank you. The law fair podcast is produced by the law fair institute.

β€œIf you want to support the show and listen at free,”

you can become a law fair material supporter at law fairmedia.org/support. Supporters also get access to special events and other bonus content we don't share anywhere else. If you enjoyed the podcast, please rate and review us,

wherever you listen, it really does help, and be sure to check out our other shows. Scaling laws, rational security, allies, the aftermath and escalation. Our latest law fair presents podcast series about the war in Ukraine.

You can also find all of our written work at lawfairmedia.org. The podcast is edited by Jim Potia with audio engineering by Goat Rodeo. Our theme song is from Alibi Music.

And as always, thanks for listening.

[BLANK_AUDIO]

Compare and Explore