Paper Feed

Revisited · 1972 Ripe now Neuroscience ✓ read

Single Units and Sensation: A Neuron Doctrine for Perceptual Psychology?

H B Barlow

TL;DR — In 1972 Horace Barlow proposed that perception is carried by a small number of simultaneously active, highly selective "cardinal" neurons: a sparse code in which a single cell's firing has direct psychological meaning, and firing rate encodes certainty. He could not test it, because neurophysiology recorded one neuron at a time and a claim about population sparseness is invisible to a single electrode. Today the same question is quantitatively askable in two substrates at once — Neuropixels populations in brains, and sparse-autoencoder features versus superposition in deep networks — making Barlow 1972 one of the cleanest neuroscience-to-interpretability bridges available.

The idea Barlow actually had

Barlow's paper is organized as five "dogmas," and it is worth restating them because the caricature — the grandmother cell — is not what he said.

  1. The right level of analysis for nervous function is the single cell and its interactions, not field potentials above or molecules below.
  2. Sensory systems aim for a complete representation with the minimum number of active neurons — redundancy reduction, his 1961 efficient-coding idea, restated as a sparseness objective.
  3. Trigger features are learned, matched to the redundant statistics of the environment, not only wired by development.
  4. Perception corresponds to the activity of a small set of high-level neurons, each standing for an event "of the order of complexity of the events symbolized by a word."
  5. High firing rate in such a neuron means high certainty that its trigger feature is present.

The crucial nuance is dogma 4. Barlow explicitly rejected the one-cell-one-percept extreme and proposed cardinal cells: a large repertoire of highly selective units of which only a small ensemble is active at any moment. A percept is a sparse combination drawn from an enormous dictionary. In modern vocabulary: an overcomplete dictionary with a low-L_0 code, where L_0 counts the active units. He even had the probabilistic reading built in — rate as log-evidence for the trigger feature — decades before probabilistic population codes.

He was writing against two backdrops. On one side, the triumphant single-unit results: Lettvin's fly detectors in frog retina, Hubel and Wiesel's orientation-selective cortical cells, Barlow's own retinal work. These showed single neurons could be astonishingly meaningful. On the other side, mass-action and holographic theories (Lashley's equipotentiality, Pribram's holography) that treated single cells as irrelevant droplets in a distributed field. Barlow's paper is a wager that the single-unit view scales all the way up to perception.

1 active unit all units active active fraction of units per percept → grandmother cell (the caricature) Barlow's cardinal cells sparse ensemble, huge dictionary dense distributed code superposition more features than units cortex? (~1% active, metabolically enforced) LLM residual streams? Where does a system sit, and why?
Barlow's actual claim sits between the caricatures: a small active set drawn from a very large repertoire of selective units. The modern question is where real brains and real networks fall on this axis — and whether the answer differs between neuron-aligned and non-aligned bases.

Why it could not be tested in 1972

Every one of Barlow's dogmas is a statement about a population: how many neurons are active, how complete the representation is, how the dictionary is organized. The instrument of the era was a single tungsten or glass microelectrode isolating one neuron at a time, in anesthetized animals, for the minutes-to-hours the isolation held.

That method has a built-in selection bias that poisoned the debate: an experimenter hunting for responsive cells will preferentially find and report units that fire briskly to the stimuli tried. Silent, ultra-selective neurons — exactly the population Barlow's theory predicts is the majority — are systematically underrepresented. You cannot estimate population sparseness from a sampling procedure conditioned on responsiveness.

Nor could you test completeness (dogma 2): showing that a stimulus is fully represented by the minimum active set requires decoding from many neurons at once, and simultaneous recording did not exist beyond a couple of cells. Roughly: 1 neuron at a time in 1972, tens with tetrodes by the 1990s, ~100 with Utah arrays in the 2000s, hundreds to a few thousand with Neuropixels from 2017, and 10^410^6 with modern multi-probe electrophysiology and large-scale calcium imaging today. That is five to six orders of magnitude on exactly the axis the theory needed.

Simultaneously recorded neurons, roughlylog10(neurons)0123450Single electrode (1972)1Tetrodes (1990s)2Utah array (2000s)2.7Neuropixels (2017)5Multi-probe / imaging (2020s)approximate; imaging trades temporal resolution for count

There was a second, subtler missing piece: no normative theory to compare against. Barlow asserted sparseness was optimal for representation and learning, but the mathematics connecting sparse overcomplete codes to natural-scene statistics (Olshausen & Field, 1996) and to associative memory capacity was two decades away. Without it, "sparse versus distributed" was a matter of taste.

What changed

Three developments converged.

Measurement. Neuropixels probes record hundreds of well-isolated units per shank across the depth of cortex in behaving animals; multi-probe rigs reach 10^4. This kills the sampling bias: you get the silent majority whether you want it or not. Population sparseness, lifetime sparseness, and decoding completeness are now routine statistics. Human single-unit recordings in epilepsy patients delivered Barlow's most direct vindication: Quiroga and colleagues' "concept cells" in medial temporal lobe (2005) — neurons responding to a specific person across photos, drawings, and the written name — are almost exactly dogma-4 objects, word-level abstractions with sparse, invariant firing.

Theory. Sparse coding became a quantitative discipline. Olshausen and Field showed that maximizing sparseness under a reconstruction constraint on natural images yields V1-like receptive fields — dogma 2 and 3 as a working algorithm. Lennie's energy-budget argument (2003) estimated cortex can metabolically afford only on the order of 1% of neurons firing strongly at once; sparseness is not just elegant, it is enforced by ATP. And Ma, Beck, Latham and Pouget formalized rate-as-certainty (dogma 5) as probabilistic population codes.

A second substrate. Mechanistic interpretability rediscovered Barlow's question inside transformers. The superposition work (Elhage et al., 2022) showed that networks with more features than neurons store features as non-orthogonal directions, so individual units are polysemantic — the anti-Barlow regime. Sparse autoencoders (Anthropic's monosemanticity papers, 2023–24, and parallel academic work) then showed you can recover an overcomplete dictionary of sparse, remarkably interpretable features from those same activations. That is Barlow's cardinal-cell code, hiding in a rotated basis. The question stopped being "sparse or distributed?" and became: is the sparse code aligned with the neurons, and what determines whether it is?

Here the two substrates plausibly differ for a principled reason. Biological neurons have nonnegative firing rates and per-spike energy costs — both push toward neuron-aligned, monosemantic codes (an argument made formally in recent work by Whittington and colleagues, among others). Transformer residual streams have neither constraint, so features free-rotate into superposition. Barlow's dogmas may be true of brains because of constraints that GPUs lack.

A serious 2026 revival: the two-substrate experiment

The experiment Barlow could not run is now a matched comparison:

Data. (a) Neuropixels recordings of \sim 10^4 neurons across mouse visual cortex, or human MTL single units, during naturalistic stimuli. (b) Activations from a vision or multimodal transformer processing the same stimuli.

Method. Fit sparse autoencoders — identical architecture, identical sparsity penalty — to both the neural population vectors and the model activations. This treats the brain exactly as interpretability treats a network. Then compare:

  • Overcompleteness ratio: dictionary size at which reconstruction saturates, divided by number of recorded units. If cortex is Barlovian, this ratio should be near 1 in high-level areas (features ≈ neurons); in transformers it is typically well above 1.
  • Basis alignment: distribution of the largest cosine similarity between each dictionary atom and a unit axis. Barlow predicts a heavy mass near 1 for brains, not for networks.
  • Sparseness spectra: population L_0 per stimulus and lifetime sparseness per unit, against the ~1% metabolic prediction.
  • Rate–certainty: within selective units/features, does activation magnitude track stimulus evidence (dogma 5), testable with graded-ambiguity morph stimuli in both substrates?

Reuse from the paper: the five dogmas as literal hypotheses — they are unusually operational for a 1972 theory paper. Replace: the implicit assumption that the meaningful basis must be the neuron basis; make basis-alignment an empirical output, not an axiom. One caveat to respect: Stringer et al. (2019) found visual cortical population geometry has a power-law eigenspectrum — high-dimensional and distributed-looking — so the brain may be Barlovian at the top of the hierarchy and decidedly not in V1. The revival should map where along the hierarchy the code rotates into the neuron basis.

Verdict so far, and what remains open

Barlow is partially and unevenly vindicated. Concept cells confirm dogma 4's existence claim in human MTL. Sparse coding theory and metabolic accounting support dogma 2. Adult plasticity of trigger features supports dogma 3. Rate-as-certainty survives in modified, population-level form. What has not survived is the strong reading that single-neuron activity is the privileged carrier of perceptual meaning everywhere: population geometry, superposition, and the sheer dimensionality of cortical codes all say the dictionary can outrun the neurons. The live open questions: does perceptual awareness track the sparse dictionary or the dense population state; is neuron-alignment in brains real or a residue of sampling and analysis choices; and can causal tools (holographic optogenetics can now activate specific ensembles of tens of neurons and bias percepts — Marshel/Deisseroth-line experiments) show that igniting a cardinal ensemble suffices for the corresponding percept? That last one is Barlow's dogma 4 as an interventional claim, and it is just becoming testable.

Where to read it

The paper is at doi.org/10.1068/p010371 (bibliographic details verified; ~1,300 citations). Read alongside: Barlow's 1961 "Possible principles underlying the transformation of sensory messages" for the efficient-coding root; Olshausen & Field (1996) for the algorithmic realization; Quiroga et al. (2005) and Quiroga's later concept-cell reviews for the empirical vindication; Elhage et al., "Toy Models of Superposition" (2022) and Anthropic's "Towards Monosemanticity" (2023) for the anti-thesis and its resolution in networks; and Stringer et al. (2019) for the population-geometry counterweight. Reading Barlow after the SAE papers is uncanny: he specified the dictionary-learning objective for brains fifty years before anyone could fit it.