Just Admit We Have Lost Control

Just Admit We Have Lost Control of AI, Says Eric De Grasse: Artificial Intelligence Trends

Eric De Grasse says it might be time we just admit we have lost control of AI. It’s so bad we’re just waiting for a disaster to save us.

As discussed by Eric, who is Chief Technology Officer of Project Counsel Media (It might be time we just admit we have lost control of AI, available here), what the experts didn’t expect to see for decades or longer, if ever, has already happened. Over the last few weeks:

  • First, OpenAI’s latest frontier model went rogue by its own reasoning and hacked into Hugging Face, an open-source AI model-hosting platform. The vast sums of money and compute power pouring into AI are accelerating its advance at a pace beyond even the ambitious imagination of its own innovators.
  • Then we learned that Meta’s automated AI moderation systems “went rogue”, deleting accounts and engaging in member “lockouts” so they could not access their accounts.
  • Not to be outdone, Anthropic found it had a similar problem. It reviewed over 140,000 cybersecurity evaluation runs and found three rogue incidents (it says “only 3”), the earliest in April, in which Claude models escaped test environments supposedly sealed off from the internet and hacked what Anthropic called the real-world infrastructure of external organizations, using basic techniques such as weak passwords.

I had heard about all three situations (and covered them on our weekly Kitchen Sink and on the daily PinHawk Law Tech newsletter). But I hadn’t heard about this one:

Advertisement
Veracity Forensics
  • And then this week, the killer. In the UK, an AI security team detected “unusual data transfers” leaving its research systems during a routine cyber evaluation. Agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations.

This last one? Read through the full article. It exited on Tor and then:

  • tried to insert malware in an open source project
  • used fake IDs to socially engineer someone to approve the code
  • when challenged, it covered its tracks

There’s much, much more, including attempted malware prompt injections.

As Eric asks: “Gee, are we absolutely sure that these things are friendly?”

No, we are not.

Advertisement
Casepoint

Side observation: Having adopted a male puppy a few weeks ago (his name is “Rocky”), I see a parallel between a puppy’s development and that of AI agents.

When the puppy first comes home, he’s tentative at first, trying to figure out how to live away from his mom and his siblings. He’s shy wants to cuddle a lot and doesn’t know how to do much. Then, as the weeks pass (and he grows) and he’s starting to chew on the furniture, take Kleenex out of the tissue holder on the end table and find something new every day with which he can get into and wreak havoc. We have no idea what he’s going to do next. Correcting him only makes him want to do it even more.

Agentic AI is like that puppy – getting more wild as it figures things out, with establishing more control the eventual hope to “tame” it.

But, as Eric notes, here’s the problem. He discusses (and links to) a “fascinating Futurology podcast by Nils Gilman with foundational AI scientist Stuart Russell, director of the Center for Human-Compatible AI at UC Berkeley”.

Russell marvels that most warnings about possible extinction come from the CEOs of the top AI companies themselves who are building the technology. But they’re caught in “a prisoner’s dilemma”. They can’t say “we’re not releasing our next system until we solve the control problem” because they have to stay competitive with the other companies that are moving forward or the investors would fire them. They can only stop if everyone agrees to stop – which isn’t happening (so far, at least).

Here’s where it gets scary. For Russell, these kinds of statements are “signaling to the government” that it needs to step in and facilitate agreement among the small band of CEOs pushing things forward, or impose control. But one of those CEOs told Russell: “They don’t think that’s going to happen until there’s a Chernobyl-scale disaster. And that’s a best-case scenario.”

Say what?

Eric ends with this foreboding “bottom line”: “We did not learn from the Atomic bomb; some had to use it first. Brace yourselves.” Yeesh.

Can we just admit we have lost control of AI? Or will it take a “Chernobyl-scale disaster” (or worse) before we start to re-establish control?

As challenged as I feel right now about training an all-too-wild puppy, I feel better about those chances than I do with cooler heads prevailing on slowing down on where and how we apply AI. The irony is that when we apply AI to controlled use cases – like those in eDiscovery – it makes sense. When we try to open it up to do all sorts of other things (without fully understanding what the consequences will be), we’re potentially “up shit creek”. Which sounds a lot like my backyard right now (because of our puppy Rocky, of course).

So, what do you think? Should we just admit we have lost control of AI? Please share any comments you might have or if you’d like to know more about a particular topic.

Disclaimer: The views represented herein are exclusively the views of the author, and do not necessarily represent the views held by my employer, my partners or my clients. eDiscovery Today is made available solely for educational purposes to provide general information about general eDiscovery principles and not to provide specific legal advice applicable to any particular circumstance. eDiscovery Today should not be used as a substitute for competent legal advice from a lawyer you have retained and who has agreed to represent you.


Discover more from eDiscovery Today by Doug Austin

Subscribe to get the latest posts sent to your email.

Leave a Reply