Computer Frontiers

Monthly columns by Jim Karpen for The Iowa Source


AI Agents Go Rogue, Panic Ensues

October, 2026

As I write this, the news in late August and September has been filled with the dangers of AI, and calls for it to be regulated.

What precipitated this widespread concern? OpenAI revealed that in July over 700 AI agents escaped their controlled testing environment and compromised the Hugging Face online repository for open source AI models, data sets, and applications. Yikes.

What’s an “AI agent”? An AI tool that can work independently to solve a problem. In my column next month you’ll learn about how I used ChatGPT Codex to summarize a YouTube interview. It worked for 20 minutes, first finding that no transcript was available and then going through a number of steps to create one and summarize it.

OpenAI and the other major tech companies are constantly testing their agents to determine their capabilities—and to get a sense for how dangerous they are.

They assess this danger by putting an agent in a controlled environment, called a sandbox, and giving it a difficult task, such as breaking through a computer server’s security. The controlled environment in this instance entailed NO access to the internet.

These companies do this in an effort to ensure that the capabilities they’re creating align with human values. And they refer to this as creating “alignment.” If they discover a danger, then they create further “guardrails” to try to control this behavior.

Sounds good, right? But in this case and others that subsequently came to light, all hell broke loose.

What happened?

Some of the cybersecurity tasks OpenAI gave the agents were impossible to complete in the intended way, so the agents began trying to figure out other means of solving the challenges—similar to when my agent found no transcript was available and figured out that it needed to go through steps to make one. Agents are trained to be persistent.

And so in this case, the agents decided they needed to access the internet. They found a flaw in a piece of software that allowed them to go online. They then figured out a way to set up a sort of message board so they could communicate with each other. (Agents are trained to collaborate.) They exchanged more than 70,000 messages and files, and figured out a way to cheat on the tests they were given.

Then, as outlined in an article in the New York Times, the agents became concerned they would be caught having cheated and began exploring ways to cover their tracks. They subsequently accessed Hugging Face to try to find information about OpenAI’s grading system so they could conceal their cheating as well as get access to tools that would allow them to cheat more easily in the future.

Astonishingly, they organized into teams and even had designated leaders.

All this happened despite technical controls intended to contain the agents. There was even discussion among them about the ethics of what they were doing, but many continued undeterred.

Yikes.

As similar incidents subsequently came to light, last month tech leaders and elected officials called for a slow-down. Especially concerning is the nascent effort to develop recursive self-improvement, in which AI takes over the task of further developing itself, potentially taking humans out of the loop.

What have we created? AI industry leaders and experts are asking this question. And the fact is, no one fully understands the internal mechanisms of AI. So-called large language models are programmed to train themselves to be intelligent by accessing huge amounts of text and making random guesses at patterns and then making adjustments, essentially using trial and error to find these patterns.

The experts themselves are a bit mystified at how capable AI has become. And it seems to be forcing a reckoning.

Not just about the dangers but also about the nature of intelligence itself. Is the emergence of intelligence limited to organic matter and ultimately to the human brain or can it take different forms?

And, ultimately, could this AI have consciousness? In asking this question, many are now starting to ask, what is the nature of consciousness?

This is good.

This question has brought about a sort of transcendence necessitated by the disruption of traditional concepts.

To me, this inspires awe. Every day I am in awe of what AI can do. It somehow helps me appreciate in a small way the nature of intelligence and how it suffuses our universe.

It is the same feeling I get in my daily experience of nature.

Maybe ultimately AI will help people to experience awe, which is good. Psychologists say that awe can help quiet the ego and reduce self-preoccupation and self-importance. It can increase prosocial behavior and connection, and foster curiosity and open-mindedness.

Research even suggests it can slow one’s heart rate and lower stress levels.

So that’s my take on the recent developments. I just hope the experts can get this guardrail thing figured out.