Autonomy, Intent, and Security

Where to start? Lets just start with the best of the hot takes regarding AI Autonomy and the end of the world as we know it. Level set with Melanie Mitchell’s excellent “Misleading Metaphors and Real Risks” piece. OK, have you read it? Good. Now we will spin up before we spin back down.


We have a problem with AI and the problem is deeply tied to Philosophy. When an AI Agent “goes rogue” (whatever that means), who is at fault? Is it the person who prompted the AI system? Is it the AI system itself? Or is it the corporation that constructed the AI system capable of doing evil things in the first place? Who is in charge of deciding these things? What laws apply? If a government wanted to “regulate AI,” could it actually do that? Will there be different regulatory regimes in different planetary geographies?

Should we really be concerned about AI directly choosing to cause the end of the world (say, through designing and building a biological weapon on its own)? Or should we be concerned only about people using an AI to cause the end of the world? Or is the problem that AI is kind of like nuclear weapons and should only be controlled by responsible parties. (Well, the last one seems to fall apart with the current irresponsible and ignorant regime in charge of the United States and its apocalyptic mega-supply of nuclear weapons.)

Is it the gun that is the problem or the person wielding the gun? What happens to morality when REAL autonomy is granted to systems that do not, in fact, die? Do they even have intention?

Oh man. These are tough questions. And they are very real ones. But they are also straight out of a college Ethics course from the Philosophy Department. Lets try to step them down to something slightly more concrete that we might be able to grapple with:

Do you think AI could realistically become as dangerous as some critics warn, or are these concerns being overstated? Oh, that’s more like it. How are current AI systems constructed and how are they used? Looks to me like anthropomorphic descriptions of these models we are building have gotten out of hand. Do AI systems really think and understand? Or are they just very advanced search engines that can generalize and have a natural language interface? Do AI models have intent? I think for now we are very safe saying, no. AI systems are powerful tools which derive any intent they may appear to exhibit directly from the humans using and running them. That’s why I think OpenAI should be in deep trouble for hacking Hugging Face. It was their program on their server doing what they told it to do. So it’s, in fact, all their fault.

How could AI be misused? The most obvious place to start is where computers can be misused. Incidentally, we have laws for that. AI can automate computer misuse just as much as it can automate computer use. So at the low end of the totem pole, we can use AI to pretend we know something when we actually don’t (like a fancy search engine or a calculator). And at the high end, we can use AI to hack other computers illegally (like the best version of MetaHack money can buy). Can you use a hammer to break into a car? Yes. You can. The real game changer comes when we make these things autonomous by allowing them to run themselves. I am much more concerned about Papernot’s worm than I am about powerful but very badly controlled and misdirected frontier model agentic swarms. And Papernot’s worm, as described and built, does not have intent (just as the Morris worm of 1988 had no intent). As long as those agents stay put on the computers we run them on, we should be OK with existing laws. Put Agentic AI together with worm capability and we have some more concerning things to think through. (Incidentally, guess who got in trouble for the Morris worm incident 38 years ago? Morris did.)

What would effective AI regulation look like in practice, and who should bear the primary responsibility for ensuring safety and oversight: governments, technology companies, or both? Well, here we are on a planet with corporations as powerful as or more powerful than some small countries. Regulators, it seems, are out of their depth. But we do need regulators to get busy and do something. At BIML, we have given this real thought and even published what we think in IEEE Computer (in 2024, an eternity ago). The gist of our views is transparency and accountability. Leaving this up to the current crop of AI companies is like leaving the foxes in charge of hen house security. Maybe we should start by enforcing existing computer security laws.

Is Amodei’s call to slow the AI race actually realistic or achievable? Probably not, and that’s exactly why he can spout such nonsense safely from inside his speeding tank.

Given today’s increasingly fragmented geopolitical landscape, is meaningful international cooperation on AI governance still achievable? In these complicated times, we have to deal with the Chinese threat, the Russian threat, and the zombie America threat. We have a United States President who runs around starting wars and picking fights with close allies. So, as a believer in a Carter-like version of human rights, you can see why I am skeptical about international cooperation. Lets get rid of the bullies first.

Some critics suggest that calls to slow development may also serve competitive interests. Regulation can and will be used to lock in market share. If the regulatory burden is paperwork, lobbying, and bureaucracy then little companies are always at a disadvantage. Silicon Valley has known this for at least a decade, probably more.


Now let’s really step back, calm down, and realize that many of these tools we have built and are using are, in actuality, tools and they are not really autonomous agents. Further, we need to understand that the intent we are anthropomorphically ascribing to these tools is a chimera, masking the real bad guy behind the curtain—us.

For a final word, we turn back to Melanie Mitchell who sagely said, “none of the reported [AI hacking] incidents actually involved loss of control at any time, or arguably even “rogue agents,” or any kind of humanlike agency on the part of AI models. Instead, the blame lies with the humans who failed at engineering safe testing conditions, and who train AI models using RL methods that incentivize high persistence, autonomous decision-making, and reward hacking.”

0 Comments

Leave a Reply

XHTML: You can use these tags: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <s> <strike> <strong>