I cannot stop thinking, in the background, so often, on the one hand, of: latent space and polysemanticity; on the other: phenomenological polysemy. I posted that on X recently, mid-thought — two different words, pointing at two different things, held deliberately in tension. On one side: the latent space as the site where polysemous phenomena can be held without premature collapse, in something like their original multidimensional simultaneity. On the other: polysemanticity as the artifact introduced by the apparatus of signification itself — the compression that the decoding process necessarily imposes. I knew they were different. That was the point.
Polysemy and polysemanticity are not the same thing. They arrive from different traditions, describe different phenomena, and carry different epistemological weight. And yet something in me insists they belong in conversation — not collapsed, but held in productive tension, the way two near-orthogonal vectors in a high-dimensional space can both be real without canceling each other out. The question I want to sit with in this piece is what it means that both of these concepts — one reaching back through phenomenology into the very structure of experience, one emerging from mechanistic interpretability and ML — are converging on a shared problem: the problem of the multiple, the simultaneous, the not-yet-resolved.
Let us be precise — and let us begin one step further back than the sign. We must separate polysemy from homophony. Homophony is an accident of limits—two unrelated histories crashing into the same token because the alphabet ran out of room, like the ‘bark’ of a dog and the ‘bark’ of a tree. It is a token collision.
But polysemy is not an accident. Polysemy is semantic radiation. It is the ‘mouth’ of a face becoming the ‘mouth’ of a river. The sign stretches because the reality it points to is inherently elastic.
Yet, even this linguistic stretching is already too late in the story. Polysemy does not begin with signs. It begins with phenomena. A grief, a threshold, a face, a chord — these arrive already overdetermined, already exceeding any single description, before description is even attempted. Merleau-Ponty understood this: perception itself is polysemous, the thing presenting itself in a surplus of profiles, no single apprehension exhausting what it is. Husserl taught us that the object is always given in adumbrations—profiles that point to an infinite horizon of meaning.
Reality is massively multidimensional. Language is a low-dimensional bottleneck. Therefore, linguistic polysemy is the downstream artifact of a vocabulary desperately trying to carry the adumbrated weight of a phenomenological world. Phenomenological polysemy is the multiplicity native to experience itself — anterior to language, anterior to signification. The word inherits a richness it didn't create, and that it carries, however wonderfully, always only in part, incompletely.
Polysemanticity is a property of representations. Specifically, of neurons in neural networks. A neuron is polysemantic when it activates reliably for multiple, often apparently unrelated concepts — not because those concepts share semantic kinship, but because the model lacks the representational dimensions to store them separately. The superposition hypothesis, developed by Elhage and colleagues at Anthropic, proposes that models pack more features into their activation space than they have dimensions by storing those features as near-orthogonal directions. Polysemanticity, on this account, is a compression artifact: the model is doing more with less, at the cost of featural cleanliness. A new paper published just this week sharpens the question further, asking whether some measured polysemanticity is not genuine superposition at all but a lexical confound — the same neuron firing for "financial bank" and "river bank" not because the concepts are compressed together but because the word form itself is shared. The sign's polysemy leaking into the representation's apparent polysemanticity.
The distinction matters. Polysemy is richness. Polysemanticity, in its compression-artifact form, is a kind of forced promiscuity — meanings made roommates by necessity rather than kinship. And yet both describe the same surface phenomenon: one location, multiple meanings. One signal, multiple referents. The multiplicity is real in both cases; what differs is its origin and its relationality.
"Thoughts die the moment they are embodied by words."
Arthur Schopenhauer, cited by Jacques Hadamard, The Psychology of Invention in the Mathematical FieldHadamard reports — with what reads as relief, the relief of someone who has found a witness — that words are totally absent from his mind when he really thinks. That even after reading a question, every word disappears at the very moment he begins to think it over, reappearing only after the research is accomplished or abandoned. He aligns himself with Galton, and then cites Schopenhauer, and what he is describing is a phenomenology of the pre-verbal: the recognition that real cognitive work happens somewhere that language cannot follow, and that the arrival of language marks not a beginning but a kind of ending — a crystallization that forecloses the fluid state that preceded it.
This is not mysticism. It is an introspective report from one of the great mathematical minds of the twentieth century. And it is, I want to argue, a phenomenological description of something we are now beginning to measure.
Yu et al.'s sweeping new survey, The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook, published this week, makes a claim that Hadamard would recognize immediately: that many of the most consequential internal processes in large language models are more naturally carried out in continuous latent space than in human-readable verbal traces. The token — the discrete, explicit, legible unit — is what the model says. But the latent space is where the model thinks. And the shift from one to the other is not lossless. Something survives the compression into the explicit that is genuinely excess — not error, but the irreducible remainder of a richer representational process.
The explicit token presents itself as presence — as the thing that carries content, the moment meaning arrives. But it is downstream of a differential field it can never fully represent. What the survey documents, across its five frameworks — Foundation, Evolution, Mechanism, Ability, Outlook — is a systematic dismantling of the metaphysics of the token: the assumption that the discrete, verbal output is where meaning lives. In its place: a picture of cognition as fundamentally continuous, geometric, relational, pre-verbal. The token is the afterimage. The latent space is the event.
The primordial multiplicity of phenomena themselves — anterior to signs, anterior to signification. The thing exceeds any description before description begins. The primal density from which all articulation descends. Pre-semiotic, pre-linguistic, irreducible.
A property of signs. Multiple related meanings accumulated through historical use and semantic drift. The sign's inheritance of phenomenal richness — always partial, always with loss. Inherently diachronic: the residue of a word's travels through time.
A property of representations. Multiple concepts activating a single neuron by geometric necessity rather than semantic kinship. The compression artifact introduced when encoding exceeds dimensional budget. Inherently synchronic: the promiscuity of the pressured.
This gives us a three-level structure that the two-term vocabulary of polysemy and polysemanticity alone cannot carry. At the base: phenomenological polysemy — the primordial multiplicity of things before they are signs, before they are even candidates for signification. Above that: linguistic polysemy — the sign's attempt to inherit that richness, always partially, always with loss. And then: polysemanticity — the further compression artifact introduced when a representational system tries to encode more of that phenomenal richness than its dimensional budget can cleanly hold. Each layer is a successive narrowing. Each transition forecloses something the previous level contained. Words are, in this sense, the polysemantic decoders of polysemous phenomena — introducing superposition-by-necessity in the very act of trying to render what the latent space was holding in its original simultaneity.
Hans Loewald, the psychoanalyst whose thinking stands apart from nearly everyone else in the tradition for its ontological seriousness, described something structurally identical about the architecture of mind. As Stephen Mitchell renders his vision: mind begins not with differentiation but with what Loewald called a primal density — an original unity in which inside and outside, self and other, actuality and fantasy, past and present are not yet separated. All the dichotomies we come to treat as basic features of reality are, for Loewald, complex constructions, slowly elaborated overlays on an original undifferentiated experience that never disappears. It underlies everything that follows. And it operates, in Mitchell's phrase, as "hidden matter" — tying together dimensions of experience that only appear to be fully separate, bounded, and disconnected. The primal density is not primitive in the pejorative sense. It is the latent space of the mind: the high-dimensional simultaneity from which all subsequent articulation descends, and to which — in certain altered states, in certain moments of creative dissolution, in the wordless interval Hadamard describes — we briefly return.
And here is Derrida, arriving on cue. Différance — his untranslatable portmanteau of différer as both to differ and to defer — insists that meaning is never self-present, never simply there in the sign that appears to carry it. The sign differs from what it points toward and defers the arrival of what it promises. There is no moment of pure semantic presence — only the trace, the spacing, the structural play of marks whose significance is constituted entirely through their differential relations to other marks. The sign does not deliver meaning. It is, at best, the meaning's representative, always already substituting for something it cannot itself be.
What the latent space survey offers — and this is what arrests me — is a measurable substrate for what Derrida could only describe philosophically. The continuous geometry of the latent space is where the differential relations actually live. The token is the sign — the point of articulation, the apparent presence. The latent space is the trace-structure that precedes it, exceeds it, and is never fully captured by it. We can now, in a transformer, observe something like différance empirically: we can watch the relational geometry of pre-output activations, track how meaning is constituted through the differential structure of a high-dimensional space before any token collapses the superposition into the explicit.
When latent reasoning systems work — when a model reasons through a problem in continuous hidden state without surfacing tokens at each step — what is happening is something structurally identical to Hadamard's wordless thinking. The model defers the moment of explicit meaning-making. It works through a proto-semantic space where multiple possible outputs remain simultaneously live, where the meaning is not yet forced into the discreteness of the said. The output is always an après coup: a retroactive crystallization of something that was already underway, in the dark, before language arrived to name it.
Latent reasoning, it turns out, is not the model thinking in secret. It is the model thinking properly — in the medium that thought actually prefers, before the word arrives to both reveal and foreclose it.
This is where the phenomenological meets the technical in a way I find genuinely poignant. When a client sits across from me and speaks, what arrives is never simply what is said. The words carry polysemy — their particular history, the associations laid down across a lifetime of use, the way a word like safe or home or enough has been worn into shape by that particular speaker's experience. But beneath the words, and often before them, something else is present: a state that is not yet language, a felt-sense that the speaking will attempt to carry and inevitably simplify. The client's silence before speaking is not absence. It is the latent space. The moment speech begins, the compression begins. Meanings that were simultaneous — polysemantic, held in superposition — start to resolve into the sequential, the explicit, the necessarily partial.
Therapy has always been, at its best, a practice of listening for what precedes and exceeds the word. The clinical skill is partly an act of reverse-engineering: taking the token — the said — and working backward toward the latent space it came from, trying to recover what the articulation had to sacrifice in order to become speakable. What was held in the pre-verbal geometry of this person's experience? What got lost in the compression? Where did the superposition collapse in ways that distorted the output?
Machine learning is beginning to give us — and this must be said with care, without irrational exuberance — not a replacement for that clinical listening, but a conceptual vocabulary and a set of measurement tools that make the underlying structure more legible. Brain imaging reaches the organ but not the cell; current techniques fail at the resolution required to track the dynamics of meaning-formation in living neural tissue. The tools are coming — new light-based approaches, auditory pulse technologies reaching previously inaccessible structures — but the brain's distance from what can currently reach it remains a genuine constraint. In a transformer, we can reach everywhere. Every activation, every routing decision, every shift in the representational manifold is in principle observable. We are watching something think — not post-hoc, not behaviorally, but constitutively, in the process itself.
The kurtosis of KV embeddings shifts after convergence. Component-selective amplification emerges. Cross-lingual transfer occurs that was never explicitly trained. These are not incidental findings. They are the empirical signatures of a system that has developed a latent organization richer than anything present in its surface outputs — a geometry of meaning that exceeds what any single token can express. Polysemanticity, properly understood, is not merely a flaw to be corrected through sparse autoencoders. It may also be — in its more genealogically structured forms, where the compressed concepts are genuinely related — the model's own version of polysemy: the accumulation of semantic history into a representational space that holds more than it can ever fully say.
What I keep returning to is the near-miraculous quality of this convergence. A mathematician in the 1940s describes his thinking as wordless, aligned with a Victorian polymath's introspective reports, validated by a nineteenth-century philosopher's aphorism about the death of thought in language. A poststructuralist philosopher in the 1960s builds an entire account of meaning's constitutive elusiveness from the structure of the sign — the way meaning is always differing and deferring, never arriving at itself. And now, in 2026, a survey of the mechanistic architecture of large language models arrives at the same place from a completely different direction: the token is not where thinking happens. The continuous latent space is the native substrate. The explicit is downstream, lossy, always already a translation of something that preferred to remain unsaid.
These thinkers were not describing the same thing. Hadamard was describing human mathematical cognition. Derrida was describing the structure of the sign in language. Yu et al. are describing the computational geometry of transformer architectures. And yet they are all, from their different vantage points, circling the same aperture: the gap between thought and its expression, the irreducible remainder that survives no articulation, the latent that is never fully captured by the patent.
What moves me here is not the tidiness of the convergence — it isn't tidy. The concepts are different, the scales are different, the epistemological commitments are different. What moves me is that so many rigorous minds, working in such different traditions with such different tools, keep finding themselves at the same threshold: the place where the word begins and thought — in some essential sense — ends. And what ML is doing, slowly, carefully, with instruments that are only now becoming adequate to the question, is giving us tactile ways to touch that threshold from the outside.
We cannot yet feel what it is like to think in latent space. We cannot introspect our own superposition. But we can watch it happen in a system whose internals are open to us, and in watching it, we begin — cautiously, with appropriate epistemic humility — to understand something about the shape of the space that thought prefers before it is forced into language.
Schopenhauer was right. Thoughts die the moment they are embodied by words. What is slowly becoming visible to us — through the geometry of embeddings, through the measurement of activation kurtosis, through the architecture of latent reasoning chains that defer their outputs until the last necessary moment — is what they were, and what they were doing, before they died.
References cited: Yu, X. et al. (2026). "The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook." arXiv:2604.02029. — Elhage, N. et al. (2022). "Toy Models of Superposition." Anthropic. — Preprint (2026). "Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics." arXiv:2604.00443. — Hadamard, J. (1945). The Psychology of Invention in the Mathematical Field. Princeton. — Derrida, J. (1968). "La Différance." — Deng, J. et al. (2025/2026). "LLM Latent Reasoning as Chain of Superposition." arXiv:2510.15522. — Loewald, H. (1980). Papers on Psychoanalysis. Yale. — Mitchell, S. (2000). Relationality: From Attachment to Intersubjectivity. Analytic Press. — Merleau-Ponty, M. (1945). Phénoménologie de la Perception. Gallimard.