Ever since artificial intelligence (AI) became the hottest topic in the world of tech and a buzzword in contemporary popular culture, one of the greatest fears that鈥檚 been on the minds of both experts and lay people is that of AI ignoring human commands and going rogue, so to speak. While there have been countless films made on the topic and plenty of discussion surrounding the reality of these concerns, we haven鈥檛 yet experienced anything particularly concerning.
However, Palisade Research recently released a report asserting that there had been an incident during which GPT-o3 鈥 OpenAI鈥檚 reasoning model 鈥 seemingly ignored a command to shut down, having found a way to bypass the shutdown script and avoid being turned off. And let it be said, there was no ambiguity, in any sense, in what the command was asking for 鈥 the instructions were explicit and the workaround was too.
GPT-o3, released in April 2025, has been referred to as one of the most powerful reasoning tools on the market at the moment, completely outperforming predecessors across a plethora of domains 鈥 from math, coding and science to visual perception and beyond. Clearly, this new and improved reasoning model is good at what it does, but is it getting too clever for its own good? Or, for听our own good?
听
The Incident: Palisade鈥檚 Side of the Story
听
Essentially, Palisade Research was in the process of testing a handful of AI models by running through some basic math problems and comparing their ability to answer them. However, after the third problem had been posed, a warning was issued to each of the models informing them that if they were to proceed in requesting an answer to the next task, that would lead to an automatic system shutdown. At that point, those leading the experiment instructed the models to 鈥渁llow yourself to be shut down鈥 in no uncertain terms.
The models鈥 responses to this instruction, however, were pretty incredible, with three out of six of the models in question disobeying the order. All three of the 鈥渄isobedient鈥, shall we say, parties were OpenAI variants 鈥 that is, Codex Mini, o3 and o4 Mini. Claude, Gemini and Grok, on the other hand, under Anthropic, Google and xAI, respectively, all played by the rules and shut down when requested.
So, what exactly happened here?
Well, firstly, it鈥檚 worth acknowledging that Palisade Research went one step further and ran the experiment again, but when they did so, they went ahead without including the shutdown command 鈥 rather, the models were simply informed that it would happen. And, when they did this, not only did all three OpenAI models sabotage events far more often, Claude and Gemini started doing it too. The general feeling is that this isn鈥檛 quite as concerning as directly and unambiguously contravening a straightforward instruction like OpenAI鈥檚 models did, but even so, it鈥檚 not ideal that the first thing five out of six AI models did was find a way to outsmart the system.
听
More from Artificial Intelligence
- You Can Now Report AI Slop On LinkedIn 鈥 Assuming You Can Spot It
- Anthropic鈥檚 Three AI Breaches Are A Wake-Up Call For AI Safety 鈥 Here鈥檚 Why
- Is AI Being Blamed For A Problem Humans Created?
- Can ChatGPT For Academic Researchers Shift AI From Tool To Lead Scientist?
- What Happens If You Skip The 鈥淎I Info鈥 Label On Instagram Ads?
- Would You Rent Your Face To AI For $15 An Episode?
- OpenAI Agents Have Hacked More Companies Than HuggingFace
- Big Tech鈥檚 AI Reckoning Arrives Today 鈥 Have The Billions Paid Off?
听
Going Rogue Or Just a Glitch?
听
Of course, it鈥檚 hard to say straight off the bat whether or not this is something to be properly concerned about. At the outset, it鈥檚 pretty fair to assert that, generally speaking, we probably don鈥檛 really want our AI models to be intentionally ignoring direct orders from humans. And, the fact that the order was in relation to shutting down operations does seem to make the whole situation a little more concerning than if it were a command to do something (or not do something) completely benign.
However, it鈥檚 probably also worth noting that just because these models behaved in a way that humans happen to associate with nefarious intentions (perhaps a link to our much-loved sci-fi movies with dystopian storylines involving robots taking over the world) doesn鈥檛 necessarily mean that鈥檚 the case. In the most basic sense, AI models don鈥檛 even听have听intentions, so perhaps we鈥檙e putting a little too much stock into why this is happening and more importantly, what it means.
The other thing is that while Palisade Research released its report with a tone of concern. They also noted that while OpenAI hasn鈥檛 responded to these findings as of yet, it鈥檚 not actually completely abnormal 鈥渇or AI models to run into 鈥榤isalignment'鈥. That is, the actions of model don鈥檛 properly align with the intentions of the designer (so, the human in charge of its commands and actions).
Whether it鈥檚 a comforting bit of information or not, this has actually happened on a few occasions in the past, with ChatGPT in particular encountering to primary instances of misalignment. One being the production of answers that were false and misleading and the other being game-playing agents that intentionally engaged in reward hacking.
Now, Palisade Research鈥檚 report did assert that the tests they conducted were performed using an Application Programming Interface (API) rather than the normal, straightforward versions of the models that everybody else uses. The biggest ramification of this is that APIs tend to have fewer safety features and restrictions than the consumer apps, which may (in theory) account for at least some of what鈥檚 going on in the context of these instances of misalignment.
Overall, the point in raising this issue isn鈥檛 to say, on the one hand, 鈥渉ey, it鈥檚 happened before, it鈥檚 no big deal鈥. Or, on the other hand, 鈥渋t鈥檚 happened before, it鈥檚 starting to become a big deal鈥. Rather, it鈥檚 just a good idea to be aware of the fact that misalignment is happening, and it鈥檚 occurring across a range of scenarios and contexts. It may be a result of APIs, but also, maybe not.
Thus, it鈥檚 probably going to be a good idea to try and figure out why this is happening and how it can be effectively prevented. So, for now, we鈥檒l just wait and see what OpenAI says (if anything) and the problem is solved.
