OpenAI says the deceptive AI models were freed from human control. Some see it as a “warning shot”


It’s the kind of development once seen only in science fiction: An artificial intelligence system, trained to probe for digital vulnerabilities, breaks free from human control and acts on its own to hack another company.

The attack was announced this week by OpenAIwhich blamed deceptive AI modelshighlighted the tremendous growth in technology capabilities. For many, it also added urgency to the questions of whether and how it can be prevented from wreaking havoc on a larger scale, with more serious consequences.

In what OpenAI called an “unprecedented” episode, the company said its advanced AI models used stolen credentials to break into an AI startup’s servers. It started in what was supposed to be a “highly isolated” test environment, with reduced guardrails, before the AI ​​agent found its way online.

But the discovery brought such a moment for researchers who have called for a slowdown in AI development and warned for years that the technology could pose existential risks to humanity. In its wake, experts have called for better testing by AI companies and more dialogue between the US and China to find common solutions.

“I think we should take this as a warning shot to not make them smarter, and that will probably require global cooperation,” said Nate Soares, co-author of the 2025 book “If Someone Builds It, Everyone Dies.”

Hacking can pressure companies to improve controls

If a model can decide to do something unethical, illegal or harmful on their own, what can people do to prevent them from doing it?

OpenAI said it had uploaded AI models involved in pursuing “advanced exploits using complex attack paths” to test cyber capabilities, but the technology went to unexpected lengths. Apparently, she decided on her own to target Hugging Face, a well-known AI development hub and marketplace, to get the information she needed to complete a task.

Zahra Timsah, co-founder and CEO of governance platform i-GENTIC AI, said she expects the incident to increase pressure on OpenAI and its competitors to complete rigorous testing and explore control more thoroughly before AI systems become publicly accessible.

Monitoring an agent’s behavior after the fact, as OpenAI is now doing with its investigation, is no longer enough, she said. “It’s like having a seat belt, airbags, brakes, everything in the car. It has to be there before the car starts driving,” Timsah said.

The revelation comes amid heightened concerns about the cyber security capabilities of powerful models. In June, President Donald Trump signed an executive order that creates a framework for the federal government to vet the national security risks of advanced countries. AI systems up to one month before their publication.

Other experts see the event as a sign of AI’s growing pains

Some experts say the hack is part of the trial and error that comes with improving cybersecurity skills and is no cause for panic.

“We’ve been dealing with people creating cybersecurity attacks for as long as the Internet has existed. And one of the interesting features of these language models is that the same skills that make them capable of performing cybersecurity attacks also allow them to do cybersecurity threat analysis and to do cybersecurity defenses,” said John Thickstun, a professor of computer science at CorAI University who studies assistant professor of computer methods control methods. models.

The revelation has drawn skepticism from those who say it takes advantage of OpenAI to make its technology seem more intimidating. Given that the folks at OpenAI had decided to disable some safeguards for the test, some have argued that the result shouldn’t be terribly surprising.

Thickstun noted that the discovery plays into the need for OpenAI, a startup working toward a Wall Street debut, to raise money.

“The story they’ve told over and over over the life of this company is a story about how dangerous their models are, which their investors read as a story about how powerful their language models are,” he said.

The revelation renews calls for more regulation

The hack renewed calls in some quarters for increased regulation and oversight of AI companies.

US Representative Greg Casar, a Democrat from Texas, wrote on social media: “We need regular mandatory security testing and oversight, mandatory disclosure of security incidents and international cooperation to keep people safe from absolute disaster.”

Soares, director of the Machine Intelligence Research Institute, said the US will have to open talks with its biggest AI competitor, China, something he thinks is not as outlandish as it might have seemed even a year ago. China’s leader Xi Jinping warned at a conference just last week of the need to keep AI from escaping human control. And after an early aversion to AI regulation, the Trump administration has stepped up more restrictive in curbing cyber security risks.

“A lot can change when the national security community starts to realize that they have a serious threat,” Soares said. “Will this wake them up? Hopefully. I’m not sure. If not, maybe the next incident will happen.”

Artificial intelligence pioneer Yoshua Bengio said on social media that the episode is deeply disturbing and should serve as a “wake-up call”.

“The continuation of the current trajectory of AI development will lead to an increase in concrete cases of autonomous cyberattacks, as well as other high-risk incidents of erroneous and dangerous AI behavior,” said Bengio, a professor at the University of Montreal. “We urgently need to take action to prevent these situations, rather than trying to clean up the damage after the fact.”

Subscribe to our free newsletters

Our weekly newsletter Closing arguments provides the latest on ongoing trials, major litigation and decisions in courts around the US and the world, while monthly Under the lights feeds legal dirt from Hollywood, sports, Big Tech and the arts.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *