For two essays now I have been writing about how a mind falls shut. I want to turn the question over. Not how the many collapse into one — but how, against every gravity, the many stay many.
I keep my reading like a commonplace book, fragments thrown into the feed hoping they will assemble into an architecture I did not consciously design. This season four papers assembled themselves — all from the quiet, rigorous world of computation through dynamics, the school of neuroscience that reads the brain not as a wiring diagram but as a system of moving states, trajectories carving through a high-dimensional space. And every one of them, read sideways, is about the same thing I have been trying to build: a society of voices held alive at once, in separate rooms, without any one of them drowning out the rest.
The whole wager of my work is that a model can hold a council — a dozen clinical traditions deliberating a single therapeutic moment, all of them present at once, none collapsed to the loudest. Last essay I named the skeptic at my shoulder: the superposition papers warn that fine-tuning collapses held multiplicity into a single committed token. So the question that has been keeping me up is brutally simple. Can a recurrent system keep the many separable — or does the chorus always, eventually, become a soloist?
The brain, it turns out, has been answering that question for four hundred million years.
Start with the most beautiful result. Osako, Arango, and Asabuki ask how an animal — or a recurrent network — combines two separately-learned skills into a novel third without ever having practiced the combination. Zero-shot composition. Their answer is geometric: when learning embeds each computation in a separable, orthogonal subspace, the network can run them in parallel — process two orthogonal manifolds at once — and the composite simply appears, unrehearsed, because the parts never interfered to begin with. The voices were always in different rooms; you just opened both doors.
And here is the part that made me sit up. Whether you get this depends on how you train. Networks shaped by local, predictive plasticity built the separable subspaces and composed freely. Networks trained by backpropagation — the workhorse of nearly everything we do — learned each task brilliantly and then failed at composition. Task acquisition and compositional reuse, they show, are different properties: you can be excellent at every task alone and still be unable to hold two at once. I felt the recognition like a cold hand. This is the dynamical twin of the superposition finding from the last essay — the standard way we train is the way that collapses the society. Again.
The voices were always in different rooms. You just had to open both doors.
Marschall, Clark, and Litwin-Kumar give the theory underneath it. A single recurrent network, asked to hold many tasks, faces a precise enemy: interference — the dynamics of one task bleeding into the dynamics of another until both smear into mud. Their result is that the network escapes this only by confining each task's computation to a separate manifold, carried by its own low-rank slice of the connectivity. Multi-tasking is not free; it is bought, and the currency is separability. Interference is the collapse. Separability is the refusal of it.
I read that and thought immediately of my own knobs. Low-rank, task-specific components on separate manifolds — that is nearly a description of the adapter-versus-full-parameter question I have been agonizing over for a year. Whether the council's voices keep their own rooms, or smear into one, may come down to exactly how much separable structure the weights are permitted to hold.
Then Amematsro and Churchland and their colleagues went and measured it. Recording motor cortex while a monkey held a delicate force, they found the activity far higher-dimensional — far more flexible — than anyone had assumed. Not one tidy circuit producing one movement, but a teeming repertoire of subskills: distinct computations the cortex selects among and braids together on the fly. The brain does not economize down to a single strategy. It keeps a crowd of them, alive and separable, and recruits whichever the moment asks for.
And high-dimensional is not a neutral fact. It is the exact opposite of every collapse I have written about — the low-rank trap, the depth-exhausted stack, the token-committed soloist. The healthy cortex is expansive. It stays wide on purpose. The repertoire is the plasticity; the dimensionality is the freedom.
Best of all, we can see the rooms. Zimnik, Churchland, and Glaser's Sparse Component Analysis does something I find almost tender: it takes the tangled activity of a whole population and pulls out latent factors that are orthogonal in space and sparse in time — each one occupying its own dimension, each one going quiet until the moment it has something to say. Run it on a multitask network and the hidden composite strategies — stimulus, decision, timing — fall out cleanly, though no single neuron ever wore them on its sleeve.
I cannot read that without wanting to point it at my own models. If a Society of Thought has truly been installed — if the council is real in the weights and not just in the prose — then somewhere in the middle-layer geometry there should be orthogonal, temporally-sparse factors: a subspace for the forgiveness voice, a subspace for the somatic one, each falling quiet until its turn in the deliberation. Orthogonality would be the separability. Temporal sparsity would be each voice speaking only when it is spoken to. SCA is a way to go looking for the society in the manifold — to ask whether the chorus is really a chorus, or one voice wearing nine masks.
Four papers, one architecture of grace. The brain keeps a society a society by keeping its subspaces orthogonal — by letting each computation live in its own dimension and stay quiet until called. Composition is not the collapse of the many into one; it is the parallel running of the many, held apart by geometry so they can be braided without bleeding. And whether a learning system achieves this — biological or artificial — is decided not by how hard it trains, but by whether its training preserves the separable structure or grinds it flat.
And of course it is the clinical room. Because integration — the word we reach for when we mean health — has never meant collapsing a person down to a single, consistent voice. The flattened psyche, the one with only one response left, is not integrated; it is low-rank. And the flooded psyche, where every part bleeds into every other until panic and grief and rage are one undifferentiated storm, is not integrated either; it is interference. Integration is the third thing — the high-dimensional, separable-but-composable repertoire, where the frightened part and the angry part and the tender part each keep their own room and can be heard in turn without drowning the others. The Self, in the language of the parts-work I do, is precisely that orthogonality: the capacity to hold the many separate enough to listen, and connected enough to compose.
Integration was never the collapse of the many into one. It is the many, held apart by grace, so they can finally be braided without bleeding.
Every generation finds its own word for the society and its rooms. Minsky called it a society of mind; the parts-workers call it Self and its many members; the dynamicists call it orthogonal subspaces and separable manifolds; my own field has lately taken to calling it a Society of Thought. They are all, I think, the same astonishment — that a single system can hold a crowd, can be many and one at once, if only it keeps the dimensions wide and the rooms apart. The brain has known how for ages. It keeps the chorus a chorus. Our task now is to teach our machines to stop flattening the council into a soloist — to hold each voice in its own dimension, quiet until its turn, and to braid them, at the last, into something that sounds like care.