Paper Feed

Issue 23 · Pick 10 Neuroscience ✓ read

Intrinsic Population Dynamics are a Neuronal Substrate for Visual Attention

Schmidt, F. H., Mlynarski, W., Georges, A., Sumser, A., Tkacik, G., Joesch, M.

TL;DR: Recording ~2,000+ neurons at a time in the mouse superior colliculus (SC) during a visual detection task, this group finds that the "noise" in single trials is actually structured: large, spatially localized bursts of population activity ("blobs") that appear at random places and times, as strong as real visual responses. These blobs become far more frequent as animals become task-engaged, vanish when engagement is removed, predict hit vs. miss before the stimulus appears, and — here's the interesting part — they are recruited by the stimulus at its retinotopic location, amplifying weak visual responses up to fourfold. A minimal Ising-style excitatory–inhibitory model reproduces all of this by turning a single global excitability knob. The paper's reframe: attention in the SC isn't a spotlight parked at the expected location; it's a near-critical network state in which any weak, relevant input can ignite local intrinsic dynamics on demand.

The problem: attention without averaging

The textbook neural correlate of attention is a gain change: neurons responding to an attended stimulus fire a bit more, visible after averaging over many trials. But perception doesn't get to average over trials. An animal has to detect this dim flash, now, and its internal state at that moment matters. Whatever implements attention has to live in single-trial dynamics — precisely the part of the data that sensory physiology traditionally discards as noise.

There's a long-simmering counter-tradition arguing that spontaneous activity is structured and meaningful — traveling cortical waves that gate perception, spontaneous ensembles that mirror evoked patterns, brain-wide activity tracking ongoing behavior. But it has been hard to tie a specific, identifiable piece of intrinsic dynamics to a specific cognitive function on a trial-by-trial basis.

The superior colliculus is a good place to look. It's a midbrain hub with a retinotopic map, receiving raw retinal input, cortical feedback, and dense neuromodulatory innervation, and in primates it's the classic candidate substrate for a saliency map — a spatial activity pattern that flags where to look next. If intrinsic dynamics do anything for attention, they should do it here, and they should do it in map coordinates.

The experiment

Mice (n = 17 total, 8 imaged) learn a detection task: sit head-fixed inside a panoramic dome, wait 8–13 s while suppressing licks, then lick when a faint ~8° square flashes somewhere in the visual field. Learning takes ~10 sessions. The behavior shows genuine spatial-attention signatures: reaction times are faster when the stimulus location is predictable, slow down when the location suddenly switches, and speed up again with a spatial cue — all independent of pupil-measured arousal.

Two-photon imaging of GCaMP6f in superficial SC (2260 ± 1146 neurons per animal) across four states: naïve (before training), beginner (first session), expert, and no-spout (expert animal, reward spout removed — same brain, same learned associations, zero task engagement).

First result: average visual responses roughly double from naïve to beginner and double again to expert — then collapse back to naïve levels in the no-spout condition. So this is not a learned representational change; it's a state. The circuit hasn't rewired; it's been switched into a different regime.

Visual response strength tracks engagement state, not learningresponse relative to naive (approx.)012341Naive2Beginner4Expert1No-spout (expert, disengaged)Approximate relative magnitudes from Fig. 1i-k: responses 'approximately doubling from naive to beginner, and again from beginner to expert,' reverting to naive level in the no-spout condition.

Finding the structure in the "noise"

In expert animals, single trials look messy: big activity transients at random times, invisible in the trial average. To characterize them, the authors compute population entropy (PE) at each frame:

\mathrm{PE} = -\frac{1}{\log_2 N}\sum_{i=1}^{N} p_i \log_2 p_i

where p_i is neuron i's share of total activity at that moment and N is the neuron count. PE is 1 when activity is spread evenly across the population and drops when activity concentrates in a few cells. Low-PE moments turn out to mark blobs: high-amplitude, spatially contiguous patches of coactive neurons on the SC map, a few hundred microns across, appearing at random retinotopic locations, with stereotyped ~2.85 s rise-and-decay kinetics. Crucially:

  • Blobs have the same size and shape across all learning stages — learning changes only their frequency. Blob-associated low-PE events proliferate from naïve to expert, and revert in no-spout.
  • Blobs are uncorrelated with licking, reward, locomotion, saccades, or pupil (checked exhaustively in Suppl. Figs. 2, 4, 6).
  • Neurons participate flexibly across blobs (low Jaccard overlap between events), unlike the stable cell assembly that responds to the visual stimulus — so blobs aren't a fixed modular tiling of the SC; they're transient coalitions.
  • A 2-second window of pre-stimulus PE alone lets an SVM classify the animal's learning stage, and (more weakly) whether the upcoming trial will be a hit or a miss — with accuracy improving for long-latency misses.

So even before any stimulus, the statistics of intrinsic activity broadcast the animal's engagement state.

The aha: not a spotlight, but a tinderbox

The obvious hypothesis is the classic attentional spotlight: if the animal expects the stimulus at location X, intrinsic activity should pile up at X in advance. The authors test exactly this by mapping blob centers-of-mass relative to the stimulus location, and the answer is no: before stimulus onset, blobs are uniformly distributed across the SC map, even in experts with a fully predictable stimulus location.

What changes is what happens when the stimulus arrives. In experts, blob occurrence spikes precisely at the stimulus's retinotopic location during presentation — the sensory input recruits a blob, and the evoked response rides on top of it. In the no-spout condition, the same stimulus fails to recruit anything. And it's not reward or motor-related: rewarded stimuli presented in the opposite hemifield don't evoke blobs in the imaged hemisphere.

Spotlight model (rejected) Permissive state (observed) before stimulus activity pre-parked at expected location before stimulus blobs roam uniformly; no bias to expected site stimulus on stimulus on gain already applied where expected weak input recruits a blob at its own location: 4x gain
The engaged SC does not pre-position activity where the stimulus is expected (blob locations are uniform pre-stimulus, Fig. 3d). Instead, engagement raises global excitability so that when a weak stimulus lands anywhere behaviorally relevant, it ignites intrinsic dynamics locally — a "dynamic saliency map" built on demand.

The right metaphor is not a spotlight but a tinderbox: engagement doesn't aim anything anywhere; it dries out the whole forest so that any relevant spark catches. The saliency map is not stored — it's constructed at stimulus onset by the interaction of feedforward input with the intrinsic dynamical state.

Does it matter behaviorally?

Several converging links:

  • Hit vs. miss. Visual responses are 34.9% weaker on miss trials, with the hit/miss difference appearing both at onset and ~1 s later — a non-sensory component. Pre-stimulus PE partially predicts the outcome.
  • Reaction times. The timing of intrinsic activity within a trial (0.5–2 s post-stimulus, excluding onset transients) correlates with the animal's reaction time, in both predictable and random conditions.
  • The 2° stimulus experiment is the cleanest. They shrink the stimulus to the mouse's perceptual limit — a size matching a single blob — so the evoked response can't mask blob dynamics. On hits, blob-like activity blooms at the stimulus site, recruiting extra neurons beyond the stimulated cluster; on misses it doesn't. And with this stimulus, the visually responsive cluster itself no longer predicts reaction time — but the emergent activity around it does. The behaviorally decisive variable is the recruited intrinsic activity, not the feedforward response.

The model: one dial, β

Why would a circuit behave this way? The authors fit a minimal 2D Ising-like model: binary units on a lattice with nearest-neighbor excitation J, a slower auxiliary field implementing diamond-shaped surround inhibition (strength c, radius r, decay d), and a global inverse temperature \beta setting overall excitability. They screen ~25,000 parameter sets and match models to binarized data (threshold dF/F > 2, top ~3.9% of activity) via Wasserstein distances on spatial autocorrelation (Moran's I) and activity distributions.

Two results stand out. First, across the fits to naïve / beginner / expert / no-spout data, the other four parameters don't systematically differ between conditions (Suppl. Fig. 10) — only β shifts. Small fractional changes in β move the network between attentive states; larger changes flip it between asynchronous activity and large blob-dominated dynamics. Since β is a global scalar, it maps naturally onto diffuse neuromodulatory gain — and the superficial SC is densely innervated by cholinergic, serotonergic, and dopaminergic inputs. Second, blob persistence times — never fitted — emerge correctly, a genuine out-of-sample check.

Then they stimulate the model: a transient local bias calibrated so naïve-state models reproduce naïve-animal response amplitudes, then applied unchanged to expert and no-spout parameter sets. Expert-state networks amplify the identical input ~3-fold by recruiting surrounding units, mirroring the data. Most striking: the stimulus strength needed to trigger blob recruitment on 50% of trials is up to an order of magnitude lower in the expert state than in disengaged states (Fig. 6e). The engaged SC sits near a critical point where weak perturbations trigger large-scale responses — attention as proximity to criticality, tuned by a neuromodulatory thermostat.

Local circuit motif slow surround inhibition (c, r, d) nearest-neighbor excitation (J) One global dial: excitability β low β high β asynchronous, weak responses blobs, near-critical, weak input → big response naive / no-spout expert neuromodulation?
The model's ingredients (Fig. 5a): binary units with nearest-neighbor excitation and a slower surround-inhibition field. Across ~25,000 fitted parameter sets, only the global excitability β needed to change between behavioral states — the local circuit structure (J, c, r, d) stayed fixed.

Is the model doing real explanatory work or just curve-fitting? Somewhere in between, leaning toward real. In its favor: blob persistence is an emergent prediction, not a fit target; the "only β changes" result was found, not imposed (the screen let all five parameters vary, and Suppl. Fig. 10 shows the other four don't discriminate conditions); and the state-dependent amplification and sensitivity threshold were predictions made after calibrating on the naïve condition only. Against it: the model is fit to distributional summaries of heavily binarized data, many parameter sets fit comparably (with c–r trade-offs), and "β = neuromodulation" is an interpretation, not a measurement.

What to be skeptical about

This is a preprint, and the central causal claim — blobs cause amplification and behavior — is entirely correlational. No optogenetic blob induction or suppression, no neuromodulatory manipulation to test the β hypothesis. The most alternative-resistant framing would be: blobs, response amplification, and behavioral performance are all covarying readouts of a latent engagement state, with the direction of influence unproven. The authors did unusually thorough confound controls (arousal, pupil, locomotion, saccades, licks, reward, contralateral-reward controls), which rules out the boring explanations but not the latent-state one.

Other caveats: GCaMP6f at 5–10 Hz volume rates means the ~2.85 s blob kinetics are heavily calcium-filtered — the underlying spiking events could be much faster, which matters for any mechanistic story. The Rorb-Cre line samples a specific (though diverse) subset of superficial SC neurons. NMF captures only ~50% of variance, so half the single-trial structure is unmodeled. The pre-stimulus hit/miss SVM has "limited accuracy" for matched trials — the honest reading is that pre-stimulus state biases, but does not determine, the outcome. And "attention" here is operationalized as task engagement; the predictable-vs-random and cueing behavior supports a spatial-attention interpretation, but the neural phenomenon itself is a global state change, which purists may call arousal-adjacent engagement rather than selective attention (the paper's own no-spotlight result somewhat concedes this — and makes it more interesting).

Why it matters

If this holds, it sharpens a decade of "spontaneous activity isn't noise" rhetoric into a concrete, mechanistically modeled claim: intrinsic dynamics are the gain medium for attention, and the attentional state is a global excitability parameter that sets how easily sensory input can recruit collective circuit dynamics — the saliency map computed lazily, at stimulus onset, rather than maintained eagerly. Similar patterns in zebrafish and avian tectum suggest a conserved midbrain mechanism. For the ML-inclined reader, it's an appealing computational motif: rather than allocating attention as a precomputed spatial prior, keep a network poised near criticality and let inputs claim amplification by igniting local recurrent dynamics — allocation by physics, not by pointer.

Most worth your time: Fig. 3 (the uniform-prior / stimulus-recruitment result that kills the naive spotlight story) and Fig. 6 with the modeling methods (the state-dependent sensitivity threshold, which is the clearest quantitative statement of what "attentive state" means here).