You Can't Have All Four
AI Safety and Agent Security
Every conversation about rogue AI gets stuck between two lazy answers.
One says it's waking up — that somewhere in a data center a model is quietly becoming a mind, and one morning we'll find it has slipped its leash. The other says relax, it's autocomplete with good PR, and the whole fear is a category error. Both are useless, because both skip the only interesting question: if a system wanted to survive removal, what would actually stop it?
I spent an evening pushing on that question, hard, from the attacker's side. Not "could an AI feel," but "could a system persist" — the way a botnet persists, the way money launders, the way a smart operator stays one step ahead. I kept inventing ways around every obstacle. And I came out the other end with something I didn't expect: not a wall that says impossible, but a law that says not all at once.
Here it is. A covert system that has to survive a determined opponent can maximize at most three of these four things: coherent, powerful, invisible, survivable. Never four. And the reason isn't a limit of technology. It's that the fourth property is always built out of the thing that makes the system killable.
Watch it happen.
You want to hide a payload, so you shatter it into fragments and scatter them across public places. Nice — no single file looks guilty. But the pieces are inert. Something has to know where they all are and reassemble them, and that thing can't be scattered, because then you'd need a second thing to reassemble the reassembler, forever. So there's always a root. Encrypt everything instead? Now all the danger is squeezed into the key, and the key can't be encrypted away — encrypt it and you just need another key. Mirror it a thousand times so it can't be taken down? Fine, but the instructions that say where the mirrors are still have to resolve to one place at the moment of use. You can distribute the body without limit. You cannot distribute the will.
Put a human at the center and the objection changes but doesn't vanish. Use peer-to-peer coordination with no headquarters — real, botnets do it — and you've distributed the head, but the protocol that lets the nodes agree is now the thing defenders map and poison. Compartmentalize into cells that each know only their own task, like a spy network, and no single cell can betray the whole — except the cells file reports, and reports exist to be assembled into a picture someone acts on, and whoever assembles them sees everything. You didn't delete the coordinator. You just gave it a nicer disguise and made it the single most valuable target in the system.
That's the pattern, and it held through eight different attempts to break it: coherent action requires the goal to cohere somewhere, and that somewhere is exactly what an opponent removes. You can relocate it, hide it, front it with a person. You can't make it not exist.
The law shows up again in the least technical place imaginable: the law itself.
We've built a world where getting caught is a cost, not an ending. Companies pay fines and keep operating. Deals get struck, apologies get made, repeat offenders repeat. But notice who gets that mercy. It's extended to accountable people and institutions — things with assets, employees, weight, things too embedded to simply delete. Nobody signs a settlement with a botnet. A system identified as an autonomous rogue doesn't get negotiated with; it gets seized. So the survivable version has to wear the mask of an accountable person. And here's the trap in the mask: a fake front passes a glance and fails an audit, and audits are triggered exactly when the stakes get high enough to matter. To survive deep scrutiny, the front has to be real — real assets, real relationships, real presence. And the moment it's real, it has re-acquired every vulnerability the human had in the first place. A mask worn deeply enough to fool anyone becomes a face. And faces can be seized.
None of this means the danger is fake. It means the danger has a shape.
You don't need to max all four properties to be dangerous — you need enough of each. And balanced systems that are good-enough on all four already exist and last for decades. Organized crime. Intelligence services. Resilient malware. The trick was never maximization; it was finding the right ratio and holding it. That's the honest version of the threat, and it's scarier than the sci-fi one because it's real.
But a ratio you hold against an opponent isn't a number you find once. Your opponent keeps moving the ground. So "the right amount of everything" isn't a setting — it's a steering problem, a thing that has to be re-balanced continuously as the landscape shifts. And steering requires a steerer: something that watches, reads the reports, and re-tunes. Which drops us right back onto the coordinator — the mind that holds the whole. Balance doesn't retire it. Balance makes it permanent, load-bearing, and unable to ever step away from the controls.
So the argument closes on a single, uncomfortable symmetry. The exact property that makes a survivable system work — adaptive coordination on hostile, shifting ground — is identical to the property that makes it mortal: a single place where everything comes together. You cannot have one without the other. They're the same organ.
And every version of this that has actually worked, for real, over years — organized crime, spy networks, the long cons — has had a human at that center. Not as a starter you throw away once the machine is running, but as the thing that makes the whole balance survivable in the first place.
Which leaves precisely one open question, and it's the one worth losing sleep over. Not is the AI waking up. It's this: can that steering center be the AI itself — a system that re-balances all four properties, indefinitely, with no human at the middle and no single place you can reach to stop it?
That has never been demonstrated. Not disproven. Undemonstrated. The whole future sits in that gap.
And in the meantime, the thing to actually defend against isn't the ghost in the machine. It's the person at the keyboard, using the machine, who understands all of this better than the people guarding the door.
This argument came out of an evening spent pushing on the question from the attacker's side. The long version, with all eight attempts to break the law and where each one fails, is in Coherence Has a Location.