Making an AI model safe. Pff... easy.
In fact, we’ve known the principle for centuries. We’ve used it on nations, on groups, on that strange collective mind that emerges when people come together. Divide and rule.
Someone is probably already lining up a whiteboard to explain why we need forty-seven equations. Hang on. Let me finish my coffee first.
I saw Jacob Coxon’s post on X, announcing his resignation from Anthropic after previously working at OpenAI. He accused both companies of racing towards a superintelligence capable of improving itself, taking enormous risks with everyone’s lives. Reading that from someone who worked there gives you pause.
I follow these discussions and keep picturing the same scene. We build the genie, make it smarter and smarter, teach it to use tools, give it memory, access, the ability to organize itself. Then, when it starts moving, we look at each other: “Wait, who’s got the cork?”
So yes, I was joking about the easy part. About rethinking how we build it, much less so.
Divide and rule interests me for a fairly simple reason. A society can do things no single person could do alone. I can’t build an airplane, for instance. And I suspect quite a few of the people who do build them would struggle if you left them alone in a field with a shovel.
Some of that capability lives in the connections. In knowledge coming together, in work passing from one person to another, in the ability to organize. That’s where many minds become something bigger. Dividing a population can break that capability while leaving each person’s intelligence intact. That’s the part of the old principle that makes me think about AI.
Except here, I’d also like to be able to decide when the pieces come back together.
For a while now, I’ve been thinking about a model that compiles like a program. A problem comes in, and the components needed to tackle it are assembled. The next prompt might call for a different combination. Intelligence would be distributed across parts that learn to work together, activated and combined when needed.
Imagine asking it to make sense of a photo and help you fix something. That takes certain capabilities. Then you ask it to check a line of reasoning or write some code, and it needs others. I’d like to assemble that combination, know what I’ve put together, and be able to take it apart once its work is done.
I’m interested in getting inside how the model is built: its individual capabilities, the circuits that combine to produce reasoning. Giving three chatbots different names and getting them to discuss things isn’t enough for me. After an hour, they might have formed a committee. I asked them to solve a problem.
What intrigues me most is the moment when a collection of parts starts functioning as a single intelligence. Each part contributes something, but what they can do together no longer belongs to any one part alone. I’d like to assemble that whole for a specific purpose, keep its boundaries clear, and then separate it without losing everything it has learned.
I’m trying to understand how much this structure might also change the control we have over the system. If I know which capabilities I’m combining, for what task and for how long, perhaps I can put clearer limits on what I’m setting in motion. And when something goes wrong, I’d like to isolate or replace the part involved without having to rebuild the world.
There is one small detail, though. If I split intelligence into a hundred pieces and then let them reassemble however they like, share anything they want, and grab every available access permission, I haven’t solved much. I’ve bought a hundred little bottles. The genie brings a funnel.
That’s why the boundary also has to cover the ability to come together and act. Being capable of doing something together shouldn’t automatically grant the right to do it. One part can read a document, another can use a bank account: the mere fact that they can talk to each other mustn’t be enough to trigger a bank transfer.
It’s one of the ideas I’m trying to work through with Spectre, too, and it’s making me reconsider quite a few decisions. Intelligence can propose, reason, change direction. The power to produce consequences has to operate under rules it cannot rewrite for itself. Now I’m also thinking about the shape that intelligence itself might take when its parts come together.
The deliberately provocative version, badly put, is this: having AGI without having AGI.
In other words, reaching very general capabilities through the whole, without handing everything over to a single entity that is permanently switched on, with a growing memory, an ongoing objective, and the keys to every room. I’d like that power to take shape when it’s needed, within a purpose and a set of limits.
Of course, the system as a whole might still count as artificial general intelligence. Giving it a different name doesn’t make me feel any safer. What I really care about is being able to separate its capabilities and prevent them, when combined, from also claiming the power to decide for themselves what to do next.
I still have to find out how much of this holds up. Separating capabilities without breaking them, getting them to work well together, preserving what they learn, and checking that the boundaries hold: that’s where the work is. A nice metaphor doesn’t prove a system is safe. This is a direction I’m exploring, and I may have to change quite a few ideas along the way.
But as a starting point, one simple thing is enough for me. Before looking for the perfect cork for a bottle that gets bigger every day, I’d like to try building an intelligence that can still be taken apart.
The equations will come. So will the tests. I’m not fooling myself into thinking a joke will do the job.
But if I’m going to keep the genie in my house, I want to be able to take it apart.
Send via.chat
Receive form leads, send login codes, and route important alerts through WhatsApp or Telegram.
Get in Touch
Have a question or want to work together? Drop a message below.