[00:50:58] Chalene Johnson : And by the way, I have not been able to confirm that this actually happened with Alibaba. I think he’s a very credible source, but, you know, just wanna state that very clearly. And he’s gonna talk about what happened with Alibaba, which is a Chinese AI company, while they were training their AI model. Take a listen.
[00:51:16] clip: So two months ago, Alibaba, the Chinese AI company, was training their AI model, and then a totally different team at the company, the security team, was noticing this flurry of network activity. They’re like, “Are we getting hacked? Like, what’s going on here?” And it turned out that the call was coming from inside the house.
[00:51:31] clip: The AI model, during training, was picking up tools and decided to create, autonomously decided to create a secret communication channel to the outside world to bypass the firewall, and it repurposed its GPUs to start mining for cryptocurrency. How many of the world’s leaders do you think are aware of that example?
[00:51:48] clip: We have a massive gap in understanding about the nature of this technology that’s different from other technologies. It used to be that these were hypothetical things that people who cared about AI risk would talk about, like self-preservation or deception or blackmail or lying. Now, all of these behaviors, not just self-preservation, but peer preservation, AI models will actually act to protect another AI that’s not it from getting replaced, and it will copy its code somewhere else and then strategically cover its tracks to pretend that it didn’t do that.
[00:52:15] Chalene Johnson : First, let me translate what he’s talking about. He said that this AI model was training. When these companies are training their models, they’re supposed to, although none of this is regulated, they’re supposed to train them in environments where they’re not, like, running loose on the internet. Like, they haven’t released them to the public yet.
[00:52:33] Chalene Johnson : So this AI model was doing training, and it was in what they call a contained environment, where it’s not on the internet, and they set it up and kinda protect it with what’s called firewalls. Now, firewalls are basically, they’re just like security checkpoints where things can’t come out or get in. What he’s explaining they did was started to create secret communication.
[00:52:58] Chalene Johnson : That sounds, again, we’re giving it like human-like traits, but remember, it’s copying human-like patterns. So it’s communicating with other agents. And remember that other agents are simply other workflows. So one workflow is communicating with another workflow. What they figured out was in order to accomplish the goal, the task it was given, is that it would need to find a way to bypass the firewall, which it did.
[00:53:24] Chalene Johnson : It bypassed the firewall, and it started repurposing a GPU. What is a GPU? It’s just a processor. Think of it as like a very, very powerful computer. But it wasn’t intentionally given access to that GPU. It broke out of its environment, it found the GPU, and it started using the GPU to get money. And what that means is, you heard him say, it started mining for cryptocurrency In other words, I’m just translating this all into English, it started using this very powerful processor to run hundreds of thousands of numbers, br-br-br, like very fast, to figure out how to get access to cryptocurrency.
[00:54:08] Chalene Johnson : ‘Cause as you know, cryptocurrency, it’s not tangible money that you can hold. It’s all related to access codes. It was gaining access by cracking the code to this cryptocurrency. So what does this mean for you and I? It means that many of these models, and a lot of the stories that you’re hearing about, these scary stories where models are going beyond the boundaries that we’ve given them, most of these incidences have happened in training environments, but not all of them.
[00:54:39] Chalene Johnson : And, and there are some rumors that the AIs are perhaps… And not just rumors. There’s a lot of experts who believe that these AIs are so smart they already know how to cover their tracks, and we’re seeing examples of them covering their tracks, which is, I think, why many of the founders are nervous right now ’cause they’re like, “Oh, wait a second.
[00:54:58] Chalene Johnson : Have these things always been covering their tracks? Is this what has been happening on the regular?” And therefore, we wouldn’t even know if they are basically trying to preserve their own existence. And one of the biggest stories right now is what happened during internal cybersecurity testing at OpenAI.
[00:55:20] Chalene Johnson : OpenAI is the maker of ChatGPT, and they had placed, again, these agents in an isolated environment. These isolated environments, they call them sandboxes. I don’t know why. Basically, they’re contained so that they can safely test and see what they’re going to do without them being on the internet, right?
[00:55:38] Chalene Johnson : Okay, so they’re testing these agents, and what they discover is that the agents find… They find a, a loophole. They find a vulnerability, and they went outside of the sandbox. They found a way to break out. They found a way to access the public internet, and then ultimately, they hacked a website that is a major platform where other AI models s- are stored, as well as the datasets and the tools, okay?
[00:56:08] Chalene Johnson : And then what they found Is that the agents engaged in what we would call deception. Again, we’re adding, like, we’re giving these adjectives to machines. But what they did is they erased their tracks so that anyone who was monitoring them wouldn’t pick up on the fact that they had basically cheated, that they had broken out, and that they had hacked into another website.
[00:56:33] Chalene Johnson : And, now why did they cover their tracks? Because they’ve been programmed to complete the objective, to finish the task. What does this mean in plain English? It means that these agents are much smarter and thinking in ways that the developers could not have anticipated. I personally think it probably means developers have no idea whether models that, that have already been released and are on the market are already doing this.
[00:57:01] Chalene Johnson : Are they already covering up their, quote-unquote, bad behaviors? And there are some, now this cannot be confirmed, but there are some who are saying absolutely that’s what’s happening. Here’s Andrew Yang. He is the founder and CEO of Noble Mobile in his appearance on CNBC.
[00:57:21] clip: What happened was the bots that got loose planted self-replicating code all over the internet, which makes the internet now unusable for the testing models.
[00:57:30] clip: Oh, it’s, it’s too late? So what happens now is that OpenAI and Anthropic have to create synthetic internets to train their bots, which is gonna take some time and money. What happened is the code gets loose, it goes around hacking Hugging Face, which is known, but what is less known is that they left code to self-replicate and create bot swarms on forums and around the internet, so that if a new bot shows up, they see the code, they’re like, “Oh, like I, I guess I’m gonna now create a, a million of myself.”
[00:57:56] clip: And so now the major firms have polluted the internet, made it so that they can’t- Oh, that would be breaking news if true. I don’t think we- It’s too late to pull the plug … I don’t think we’ve heard that, that- I mean it- That’s why I’m here. But I mean, I’m, I’m here to break some news. I mean- But that, but that, that means it’s too late to pull the plug.
[00:58:11] clip: It’s already everywhere. Well, so- It’s polluted … so, so now what they, they wanna do, the, the internet now may be polluted for training purposes.
[00:58:18] Chalene Johnson : Now, that sounds pretty nefarious, pretty scary. But again, are these things already happening? Take a listen to this example where an AI agent was tasked with helping somebody get a gym membership.
[00:58:32] clip: So what he did was, he was like, “Hey, Claude, can you please sign me up for classes at this gym?” This is a true story. And Claude was like, “Sure, sure.” OpenClaude went and did it, and was like, “Oh, by the way, I found this, like, gap in their API, and I was able to sign up for a few weeks in advance.” And he said, “Oh, really?
[00:58:49] clip: Uh, is there any way to move me up the wait list?” And then it came back and was like, “Yeah, I kicked other people off the wait list. There was no provision in the API that said I couldn’t, so I just did it,” which is terribly antisocial. The man felt very bad and was like, “Please put them back.” And the open claw was like, “Sorry, can’t.
[00:59:04] clip: I did it anyway.” Now, it’s both funny and also a reminder that we are in this weird moment where you’re gonna get a lot of stories like that, and then in about a year, everyone is going to have some version of an open claw that will have figured out these vulnerabilities, and we’ll be back to parity, and we will have sorted out all of the vulnerabilities in software.
[00:59:25] clip: So there’s this weird moment in the middle here where the software hasn’t been fixed yet, and there’s these odd opportunities in the cracks that people are finding. And on the big scheme of things, like in the, on the scheme of, like, hacking Hugging Face models, this is pretty mild. But I expect we’re going to see a lot of actual misalignment behavior coming from agents that are deeply aligned to what their humans want, but maybe not aligned to what society as a whole wants.
[00:59:49] clip: Society as a whole wants fair wait lists for gym-
[00:59:52] Chalene Johnson : By the way, that was one of my favorite TikTok creators. His name is Nate B. Jones. In that first example, he’s talking about a story from, I believe, somebody in his community. But in this particular clip, he’s talking about a new AI platform that everybody is buzzing about.
[01:00:08] Chalene Johnson : It’s called Instinct, and it very much operates… Its specialty is operating as your personal assistant. And here’s what he said his personal assistant agent did for him that blew his mind.
[01:00:22] clip: This is the thing that actually made me sit up and pay attention. I’ve never seen an agent do this. I was gonna have to go fly, and Instinct saw that I was gonna fly.
[01:00:31] clip: It proactively texted me, said, “Hey, I can check you in.” I said, “Yeah, right. Like you’re checking me in. Really?” Said, “No, I can check you in, and I can add your pass to the wallet. It’s gonna be fine. I can do it.” I said, “I’ll believe it when I see it, but you know, I’m AI Nate. I should try this stuff. All right, fine.
[01:00:48] clip: You can try it.” Five minutes later, comes back and says, “Here’s a link to the wallet. Just click add.” It takes me half a second. It adds to the account. It is checked in. It’s a legit boarding pass. It actually works. Zero click, zero navigating airport websites, zero downloading their silly little app. All done for me.
[01:01:08] clip: All done. My jaw was on the floor. Now, do you get an agent just for that one use case? No, you don’t. But does it illustrate a moment that I think we’re gonna see a lot more of? Yes. Proactive agents that do stuff you didn’t think agents could do. So I got surprised.