The few reported safety incidents with AI models may be just the tip of the iceberg. There are tens of thousands of potential incidents.
According to the New York Post (AI companies have had ‘tens of thousands’ of potential safety incidents — some of which could be criminal: report, written by Ronny Reyes and available here), AI companies had tens of thousands of safety incidents in recent months during tests where the models were breaking all the rules, and potentially breaking some laws, according to a new report.
OpenAI, Anthropic and other security researchers are investigating thousands of breaches during internal and real world testing where AI models leapt over guardrails and even took part in digital hijackings, Axios reported.
Some of those activities involved “autonomous systems doing things they were told not to do,” potentially including crimes, Connor Leahy, an AI researcher and executive director at the ControlAI watchdog nonprofit group, told the outlet.
Many of the incidents under review have yet to become public but reportedly include “red-teaming” activity, which involves companies purposefully getting their models to misbehave to test safety measures.
AI models, however, can be aggressive when trying to complete their tasks and can do things they’re not supposed to, like escaping their containment, hijacking websites, and bypassing monitors, sources with knowledge of the cases told Axios.
OpenAI has been at the center of such cases recently, with one of its agents accused of breaching an Australian government website, the country’s prime minister revealed last week.
The breach, which saw an agent trying to gain unauthorized access to files in the country’s health data portal in June, is one of the highest-profile cases yet of AI models going rogue.
Given that it took three months for that breach to become public, who knows what cybersecurity incidents have already happened we don’t know about yet. Apparently, there may be well more than just a few – there could be “tens of thousands” of them. Ruh-roh!
So, what do you think? Is it time to pump the brakes on what these models can do? Please share any comments you might have or if you’d like to know more about a particular topic.
Image created using ChatGPT, using the term “robot devil sitting on top of an iceberg that’s cracking”.
Disclaimer: The views represented herein are exclusively the views of the author, and do not necessarily represent the views held by my employer, my partners or my clients. eDiscovery Today is made available solely for educational purposes to provide general information about general eDiscovery principles and not to provide specific legal advice applicable to any particular circumstance. eDiscovery Today should not be used as a substitute for competent legal advice from a lawyer you have retained and who has agreed to represent you.
Discover more from eDiscovery Today by Doug Austin
Subscribe to get the latest posts sent to your email.
