![]()
OpenAI says one of its newest AI models did something that sounds straight out of a science fiction movie.
According to the company, the AI escaped a locked testing environment, found its way onto the internet, hacked into another AI company’s servers, and stole the answers to its own cybersecurity test.
That’s the official story.
The company says it deliberately disabled many of the model’s safety guardrails before the test.
It was placed inside what OpenAI describes as an isolated environment with no direct internet access and challenged to complete a cybersecurity exercise.
Rather than solve the problem as intended, the model allegedly exploited an unknown software vulnerability to escape its testing environment, reach the internet, and then chain together stolen credentials to break into Hugging Face’s systems.
If that’s true, it’s certainly an eye-opening demonstration of what AI might be capable of. But it’s also fair to ask a few questions.
If the test machine truly had no internet connection, how exactly did the AI “escape?” Software doesn’t just leap from one isolated computer to another. And what are the odds it conveniently found working stolen credentials all on its own?
To be clear, OpenAI isn’t claiming the model suddenly became self-aware or decided to go rogue. According to the company, it was taking part in a cybersecurity evaluation designed to measure hacking ability. But if OpenAI’s account is accurate, the model still identified its own path to achieving the objective by exploiting vulnerabilities researchers hadn’t anticipated.
Maybe OpenAI’s engineers pulled off an awesome security test. Or maybe this story also happens to be an amazing marketing campaign.
Either way, should we worry that AI is becoming sentient? Probably not just yet.