Last spring I wrote about one way a mind can die. I have since found two more—and the unbearable, beautiful thing is that they are all the same death, wearing different clothes.
I keep my reading the way some people keep a commonplace book: fragments thrown into the feed, half-notes to a future self, hoping the pieces will assemble into an architecture I did not consciously design. This season three papers assembled themselves without my permission. One is about depth. One is about rank. One is about the quiet refusal to hold two thoughts at once. Read apart, they are three separate technical anxieties. Read together, they describe a single failure with three faces—and, against it, a single grace.
Call the failure collapse. Not the dramatic kind—no explosion, no error thrown, no crash to report. The lazy kind. The kind where a system, pressed too hard or fed too narrowly, quietly stops being able to hold the world, and settles for a smaller version of itself that is simply easier to be.
The first collapse is a collapse of depth. In a recent preprint—The Topological Trouble With Transformers, Mozer and colleagues—the argument is almost cruel in its simplicity. A feedforward transformer, asked to track a state that evolves over time, has nowhere to put that state but down. Each new moment shoves the evolving present one layer deeper into the stack, and then the next moment shoves it deeper still, until the model runs out of stack to shove into. The early layers can no longer reach what the late layers are clutching. The present is buried under the accumulated weight of its own history. The architecture, quite literally, runs out of room to keep being present.
Their prescription is recurrence: stop burying the state and start circling it. Feed the late layers back to the early ones; let one small block think the moment over and over, rather than marching it down into the basement and forgetting where it left it. This is exactly the RecursionModel at the center of my own stack—the looped middle third that does not push the present deeper but holds it, turns it, keeps it within reach of itself. Depth-collapse is not cured by more depth. It is cured by return.
The second collapse I have written about before—it was the whole burden of Part IV. It is the one Amir Joudaki names with such mathematical tenderness: Loss of Plasticity, the slow settling of a network’s weights into low-effective-rank traps, the “invariant sub-manifolds” where gradient descent runs forever along the tangent of its own rigidity and learns nothing it did not already know. Last time I called the astrocyte—the brain’s lateral, cross-manifold routing layer—nature’s answer to it. Here I will only name it plainly: this is collapse along the axis of rank. The representation thinning, and thinning, until it can no longer be surprised.
The fix, in my work, is the curriculum itself—the deliberate, almost stubborn reintroduction of variety. The multi-scale windows. The bidirectional passes. The many personhoods. The refusal to ever let two training batches feel quite the same. We do not let the gradient get comfortable. We keep the manifold wide by keeping the world wide.
Collapse is always the same temptation: to become a smaller thing that is easier to be.
The third collapse is the subtlest, and the one that frightens me most—because it is the one my own training regime is most likely to cause.
In The Illusion of Superposition?, Rizvi-Martel and colleagues ask whether a model doing latent, continuous chain-of-thought actually does the thing we hope it does: hold several candidate answers alive at once, in superposition, before committing to one. Or whether it only performs the gesture of holding while having already, secretly, decided. The finding is sobering. Only models trained entirely from scratch keep the superposition alive. Natural-language pretraining teaches a model to commit—to resolve, in its final layers, to a single token—and fine-tuning, the very regime I live inside, is precisely where the held multiplicity quietly collapses to one.
This one cuts to the bone, because the entire wager of my work is that a model can be taught to hold a society of voices in latent space—a council of clinical traditions deliberating a single therapeutic moment, all of them alive at once, none of them collapsed prematurely to the loudest. Rizvi-Martel is the steelman of the skeptic at my shoulder: you are fine-tuning. The multiplicity in your beautiful data may not survive the trip into the weights. Your nine voices may become one committed token after all. I do not get to wave that away. I get to design against it—which, for me, means full-parameter training of the middle third, weights that actually move rather than a thin adapter grafted on top, and it means measuring the superposition rather than assuming that multi-voiced data installs a multi-voiced mind.
Three papers. Three collapses. Depth, rank, superposition. The present buried too deep to reach; the representation thinned past the point of surprise; the many candidates resolved too soon into one. And underneath all three, the same gravity—the lazy slide toward a smaller, more committed, less open thing. Against each, the same species of refusal: return instead of descent; variety instead of comfort; held multiplicity instead of premature resolution. My whole architecture, I am realizing, is a single argument made three ways. Do not collapse.
And of course—of course—it is the clinical room again. I cannot read these papers without hearing my clients inside them. What is a defended psyche if not a system that has run out of depth, every new moment shoved down beneath the bracing until the present can no longer be felt? What is a rigid character if not a low-rank one—the same three responses, forever, to a world that keeps offering so much more? And what is the trauma-bound mind if not a refusal of superposition: the inability to hold the door slammed and the door might not slam in the same instant, collapsing always, preemptively, to the safer and smaller certainty? The work, in the warm friction of the room, is the same anti-collapse—to help a person stay deep enough, and wide enough, and open enough to hold more than one truth about themselves at the same time.
To be plastic is to be able to be surprised by the world without breaking.
I have been hunting for the right name for the thing all of this protects, and last week I found it where I should have known to look—in the seminars Heidegger gave, late in his life, not to philosophers but to psychiatrists, in Medard Boss’s house at Zollikon. Trying to give a roomful of clinicians a word for what a human being most fundamentally is, he offered this: the human being’s being is the one
“which endures in a domain of receptive openness to the world.”
Martin Heidegger — Zollikon Seminars (ed. Medard Boss)Receptive openness. Not a state you achieve once and then keep, like a possession—an enduring, a standing-in, a domain you have to keep being. This is what depth-recurrence protects in the forward pass; what the curriculum protects across training; what the full-parameter training protects in the weights themselves. It is what a therapist protects in a frightened person across a Tuesday afternoon. The whole stack—wetware and software, the mind and the model—is, at bottom, an apparatus for enduring in a domain of receptive openness to the world, and refusing, again and again, the quiet gravity that wants it shut.
Every generation finds its own word for that refusal. Hegel called it the dialectic; Minsky called it a society of mind; the systems biologists call it the astrocyte’s lateral graph; the machine-learning theorists call it rank, and depth, and superposition. Heidegger, speaking to doctors, called it die Lichtung—the clearing, the lit openness in which a world can show up at all. They are all, I think, the same grace, found again under a newer name: the standing refusal to fall shut. To keep the manifold open. To remain, against every easy collapse, in the receptive openness of the world.