Why AI can go rogue, and how to stop it | Global News Podcast
Welcome to the Global News podcast from. the BBC. I'm Valerie Sanderson and I'm. joined by our technology reporter Chris. Vallance. And Chris, this is a story. about AI going rogue. Tell us about it. >> Right, okay. Well, this concerns an AI. agent. These are programs that are. designed to have a degree of autonomy. So, you a program, for example, you. might ask, you know, I'm on a business. trip, can you plan my travel and book my. hotels? And it will go off and do this. So, OpenAI was testing one of its AI.
agents uh powered by two of its latest. models, one which it hadn't even. released yet. And uh it set it a task in. cybersecurity. And essentially, what the AI agent. decided to do was to cheat. So, it. essentially um was in a test. environment, an environment that was not. on the general internet, and it decided. that the best way to complete the task. was if it broke out of this environment. and got access to the internet. So, it.
discovered uh a previously unknown. security vulnerability, hacked its way. out of its uh test environment, and then. decided to hack into a platform called. Hugging Face, which is used by AI. programmers. And there it thought it. would find the information it needed to. essentially cheat in this test. Uh now, Hugging Face, who noticed that it had. been hacked, um described this as a a. mind-boggling. uh display of autonomy. And uh OpenAI.
called it unprecedented, and uh they say. uh they are taking steps to to make sure. this doesn't happen again, but they also. say it's this kind of thing is going to. be more likely as these models get more. powerful. >> How could it do this? I mean, you know, listening to this, I think, well, surely. there were rules in place, surely there. was security in place to stop it going. rogue. >> Well, this is very interesting. I think. this comes to a point about the story. that's worth thinking about, essentially, is that in this test, OpenAI decided to remove the safeguards.
that are normally placed on some of. these uh these agents. It had removed. systems designed to spot this kind of. behavior. Uh this was a test to look at. its maximum capabilities, as it were. The thing that some commentators are. saying it was it was unwise of them to. do that, when really the. the maximum capabilities turned out to. be a lot greater than they expected. >> Has anything like this happened before? >> Yes, uh it has things like this have.
happened before, but it sort of depends. on what you mean by going rogue, if you. like. So, if you look at rival firm. Anthropic, uh last year they reported that uh in a. test environment, they had told uh their. AI model uh to uh pretend to be that it. was an assistant in a company, and. they'd given it um access to the. company's email, and then they told it. it was going to be replaced. And uh to. come up to to come up with various. strategies to avoid the scenario in.
which it was which it was being. replaced, and it's uh one of the things. it did was to attempt to uh blackmail uh. staff to not replace it, because it had. found um evidence in the in this. fictional company database that the. staff member in question was having an. affair. So, it tried to use that to uh. stop itself being turned off. That was. that was quite a sort of contrived. scenario, if you like, by Anthropic. But. if you look at these AI agents, these. programs that are meant to go out and do. stuff for us with minimal instructions,
able to act on their own recognizance, um. you know, there a lot of people are. quite worried that that, you know, they. might do things we really don't want. them to do. They might go too far. So, for example, uh you know, reported. incidents we've had one AI agent uh. delete an entire company database. Uh another case is a it's a very good. video, if you watch it online. Hannah. Fry, the the mathematician, um set up a an AI agent to uh sell. novelty mugs and it it started uh.
messaging tech journalists and doing all. sorts of things she didn't expect it to. do. So, uh you know, the question is, you know, what are the limits on these. AI agents? And there's a there's a. wonderful thought experiment by. philosopher Nick Bostrom, where he talks. about a paper clip maximizer. So, you. know, you can imagine an AI that was. told to go out and and sell paper clips. And you know, it basically does that to. the maximum degree possible. It sees. human beings as a. a possible roadblock on its way to. creating as many paper clips as it.
possibly can, you know, covers the whole. planet in paper clips. But that's just. what it's told to do. So, that's what it. does. So, uh you know, both in thought. experiments and in reality, the the. issue of these AI agents sometimes uh. doing things their. their owners, their creators, didn't. want them to do is quite a serious one. >> It sounds as if the AI is thinking for. itself. Is Is that what is actually. happening? >> Well, I think it's Terms like thinking. are kind of tricky because, you know, they they they suggest uh something.
similar to human thought and human. reasoning, which which I don't. necessarily translate. I think the issue. here is that they give them. capabilities, they allow them to do. things computationally. They might be. access to financial systems, they have. access to your email, they have access. to the internet. You ask them to go and. do things. The the difficulty there is. making sure the machines are are. constrained in some way, so they only do.
things that are legal, responsible, and. safe. Now, that's quite a hard. challenge, in part because the. capabilities of these models are growing. all the time. And in a sense, the uh companies. don't necessarily have a perfect. understanding of what those capabilities. are. I mean, that's why they do testing. and why security institutes and. governments do testing to find out, you. know, just how powerful they are.
>> So, Chris, I mean, how close are we to a. time when AI is actually smarter than. even the smartest human? >> Well, that's a question, the answer to. which depends on who you ask. But, there's certainly a sense that it's much. closer than we might have thought it. was, say, 5 years ago. It's also a. complex question cuz it depends what you. mean by smarter than a human. I mean, you can look at robots, and they still. struggle to do things like loading a. dishwasher. On the other hand, my phone.
can beat me every time at a game of. chess. So, it really depends smarter at. what is also a part of the question, but. there's a lot of concern that artificial. general intelligence is not very far. away, and there's certainly concern. amongst governments that the new models. are becoming ever more powerful. at at a regular. fast rate. >> You mentioned constraints. What are the.
constraints that exist already, and what. is being muted by governments and indeed. by companies themselves? >> Well, companies themselves have taken. steps when they felt that a model was. so powerful that releasing it might put. a an extremely powerful. tool in the hands of the wrong people. So, Anthropic's Mythos. AI model, there's a huge amount of. concern about its capabilities.
You know, security institutes that. tested those capabilities found they. were they were a step change. Um and. Anthropic has sort of. uh slowed the release of that model. It didn't put it on general release. straight away. And so. that's the company holding back. Of. course, lots of people say you can't. just rely on the companies to behave. responsibly. So that then turns to. regulation. And there's certainly a lot.
of discussion about that. and regulation in force. in some areas, but there's a challenge. there. This is a global industry and so. regulation really has to. proceed internationally if you're not to. get a situation where. one set of countries holds back and the. the space that's created is simply. filled by models from another country. >> And is there international agreement on.
this? >> Uh. >> [laughter]. >> I think it's hard to say there's. international agreement. Uh what you can. say is that there's international. concern, I think, and a sense that you. know, multinational agreement it is is. also potentially part of the picture. >> I suppose the problem is it's a. Pandora's Box, isn't it? We can't put AI. back in the box. We have to deal with. the issue as it is now. >> Yeah, it can't be uninvented. That's for.
sure. So and the pace of change with AI. is is is you know, one that's hard for legislatures to keep. up with, let alone international. agreements. So so yeah, it's it it is a. Pandora's Box. There's plenty of promise. as well, but there's also. you know, a rapid technological. development that's hard to adjust to. >> Thanks, Chris. That was Chris Vallance, our technology reporter. If you want to. find out more about this story, listen.
to the Global News Podcast. Just click. on the link below. But that's it from. this edition of the Global News Podcast. here on YouTube. Thanks for listening. Thanks for watching. And bye for now.
