ΒΆPaper Feed

Issue 34 Β· Pick 06 Neuroscience βœ“ read

Efficient coding makes and breaks Webers law

Prat-Carrabin, A., Yamamoto, R., Gershman, S. J.

The full text could not be fetched; this explainer is based on the abstract only.

TL;DR: Weber's law β€” the observation that your just-noticeable difference between two stimuli scales with their magnitude β€” has been treated for 170 years as a near-universal fact about perception. This paper argues it isn't a fact about perception at all; it's a fact about the world, refracted through an efficient perceptual code. Natural stimulus distributions are skewed toward small magnitudes, so an efficient brain spends its representational budget there, making small stimuli easy to discriminate and large ones hard. The authors flip the stimulus distribution β€” making large magnitudes frequent and small ones rare β€” and Weber's law inverts, in three different sensory modalities. That's causal evidence that a foundational psychophysical law is an emergent, adaptable consequence of efficient coding.

A note before we start: only the abstract of this paper was available to me. Everything about the general theory below is standard background in the efficient-coding literature; everything attributed to this paper is grounded in the abstract, and I'll flag where the details I'd want (effect sizes, timescales, per-subject robustness) simply aren't in front of me.

The oldest law in psychology

Weber's law, from Ernst Weber's experiments in the 1830s–40s, says that the smallest detectable change in a stimulus is proportional to the stimulus itself:

\Delta x \propto x,

where x is the stimulus magnitude (weight, brightness, line length, numerosity...) and \Delta x is the just-noticeable difference (JND). You can tell 100g from 105g, but not 1000g from 1005g β€” you need roughly 1050g. Equivalently, discriminability β€” how well you can tell nearby stimuli apart β€” falls off as 1/x. This holds, at least approximately, across a startling range of modalities and species, which is why psychophysics textbooks present it as one of the field's few genuine quantitative laws.

But why does it hold? The classic answer, going back to Fechner, is representational: the brain encodes magnitude on a logarithmic scale, and uniform noise on a log scale produces Weber behavior on a linear scale. That's a description, though, not an explanation. It just pushes the question back one level: why would the brain use a logarithmic code?

Efficient coding: the world sets the ruler

The efficient coding hypothesis (Attneave, Barlow, and a long modern lineage) says sensory systems have limited representational capacity β€” finitely many neurons, spikes, bits β€” and should allocate that capacity according to what actually occurs. If some stimulus values are common and others rare, don't waste resolution on the rare ones.

Modern formalizations make this quantitatively sharp. If a stimulus x occurs with prior probability density p(x), and the encoder has a fixed total budget of Fisher information J(x) (roughly: the local sensitivity of the neural representation, whose square root sets discriminability), then optimizing discrimination performance under the budget constraint yields an allocation where discriminability tracks the prior:

\sqrt{J(x)} \propto p(x), \qquad \text{so} \qquad \Delta x \propto \frac{1}{p(x)}.

Read that second expression carefully, because it's the whole story. The JND at a stimulus value is inversely proportional to how often that value occurs. Common stimuli get fine resolution; rare stimuli get coarse resolution.

Now here's the punchline for Weber's law. Natural magnitude distributions are heavy-tailed β€” small values are common, large values are rare. Sound intensities, object sizes, numerosities in natural scenes: many follow something close to a power law, p(x) \propto 1/x. Plug that into the formula:

\Delta x \propto \frac{1}{p(x)} \propto x.

Weber's law falls out. On this view, Weber's law isn't a design principle of the brain. It's the shadow cast by the statistics of the natural world onto an efficient encoder. The logarithmic representation Fechner posited is just what an efficient code looks like when the input distribution is \propto 1/x.

World: stimulus prior p(x) magnitude x small values common large values rare efficient code Percept: JND Δx ∝ 1/p(x) magnitude x Δx grows with x: Weber's law
The efficient-coding account of Weber's law: a fixed representational budget is allocated in proportion to how often each magnitude occurs. A world skewed toward small magnitudes (left) produces JNDs that grow with magnitude (right). Weber's law is the world's statistics, echoed by the code.

This account has been circulating for over a decade (Ganguli & Simoncelli's optimal-allocation framework; Wei & Stocker's "lawful relation" between prior, bias, and discriminability; Piantadosi's work on numerosity). It elegantly explains Weber's law. But until now the evidence has been largely correlational: measure natural statistics, measure discrimination thresholds, show they match the predicted relationship. Correlational evidence for efficient coding always has an escape hatch β€” maybe the brain's magnitude code is just hardwired to be logarithmic, and the match to natural statistics is coincidental or evolutionary rather than an active, ongoing optimization.

The experiment: run the world in reverse

The way to close that escape hatch is intervention. The theory makes a wild-sounding prediction: the prior distribution is the only thing that matters. If you place someone in a statistical environment where large magnitudes are common and small ones are rare, an efficient encoder should reallocate its budget accordingly β€” and discriminability should now increase with magnitude. The JND should shrink as stimuli get bigger. Weber's law shouldn't just weaken; it should run backwards.

That's the manipulation this paper performs. Per the abstract: the authors ran discrimination tasks in three different sensory modalities (the abstract doesn't name them; magnitude-discrimination staples like numerosity, duration, loudness, or line length are typical in this literature, but I can't confirm which were used). They skewed the stimulus frequency distribution toward large magnitudes, and report three results:

  1. Weber's law broke and inverted. With the skew reversed, the usual pattern of discriminability decreasing with magnitude flipped.
  2. Discriminability tracked the imposed distribution across modalities β€” subjects' perceptual resolution followed where the experimenters put the probability mass.
  3. The adaptation improved task performance. This is the crucial normative check: the reallocation wasn't just some passive aftereffect, it was the useful thing to do given the new statistics, exactly as the efficiency argument requires.
Small-skewed environment x p(x) discriminability Weber's law holds Large-skewed environment x p(x) Weber's law inverted
The causal test. Left: the natural situation β€” small magnitudes frequent, discriminability highest for small stimuli (Weber's law). Right: the experimental manipulation β€” large magnitudes made frequent, and discriminability follows the probability mass, inverting the classic pattern. The abstract reports this inversion across three sensory modalities.

Why this is a bigger deal than "another efficient-coding paper"

The efficient-coding literature is large, and much of it is model-fitting: take psychophysical data, show an efficient-Bayesian observer fits it well. Those papers are valuable but always vulnerable to the "many models fit" critique. This paper is different in kind, for two reasons.

First, it converts a descriptive law into a derived one. There's a meaningful hierarchy of scientific explanation: "the JND is proportional to magnitude" (regularity) β†’ "the brain uses a log scale" (mechanism-flavored redescription) β†’ "a resource-limited encoder matched to environmental statistics must produce this, and produces the opposite when the statistics are opposite" (mechanistic explanation with a controllable knob). Demonstrating the knob works β€” that you can dial Weber's law up, down, and backwards by manipulating stimulus frequencies β€” is the difference between explaining a law and merely restating it.

Second, it implies the magnitude code is plastic on behavioral timescales. If Weber's law were baked in by evolution or development, an experimenter's stimulus distribution over an experimental session (or a few sessions β€” the abstract doesn't say how long adaptation took) shouldn't be able to invert it. That it can suggests the brain is continuously re-estimating environmental statistics and re-allocating representational precision. This connects to a broader picture in which adaptation, serial dependence, central-tendency biases, and now Weber's law are all facets of a single online resource-allocation process. For an ML reader: this is a biological system doing something like continual, unsupervised recalibration of its quantization scheme to the input distribution β€” closer in spirit to online entropy coding or learned codebook adaptation than to a fixed feature extractor.

The three-modality result matters here too. If the inversion happened only for, say, numerosity, you could tell a story about task strategy or numerical cognition specifically. Seeing the same distribution-sensitivity in three modalities argues it's a general property of magnitude coding, which is exactly what an efficiency principle β€” as opposed to a modality-specific mechanism β€” predicts.

What I can't tell you, and what to be skeptical about

Because only the abstract is available, several load-bearing questions are open. If you read the paper, these are what to check:

Effect size and completeness of the inversion. "Inverted the pattern" could mean JNDs that cleanly decrease with magnitude, or a statistically significant flattening-plus-slight-reversal at the group level. The theory predicts full inversion only if adaptation is complete; partial adaptation (a mixture of a long-term natural prior and the short-term experimental one) would produce something in between β€” which would itself be theoretically interesting, but a weaker headline. Look for the figures plotting JND versus magnitude per condition, and whether they show slopes with confidence intervals or just an interaction effect.

Per-subject robustness. Group averages can hide a bimodal population (some adapters, some not). The triage question β€” is the inversion robust across individuals? β€” is exactly right, and only the supplementary per-subject data will answer it.

Timescale and mechanism of adaptation. Did discriminability track the distribution within minutes, or only after extended exposure? Fast adaptation points toward gain-control-like mechanisms; slow adaptation toward learning. This also determines how the result interacts with the older correlational literature: the natural-statistics account of Weber's law assumed a code matched to lifetime statistics, and it's not obvious a priori that the same matching operates over an hour in the lab.

Decision-level confounds. The perennial worry in this literature: is the change in the encoder, or in the decoder/decision rule? Subjects who learn that large stimuli are frequent could adjust criteria or attention in ways that mimic reallocation of sensory precision. The paper's claim that adaptation "improved task performance" is the right kind of evidence (a criterion shift alone typically can't produce the specific quantitative pattern an efficient encoder predicts), but the strength of the argument depends on whether the authors fit the full efficient-coding model β€” including its signature prediction linking discriminability to p(x) with a specific exponent β€” or just show a qualitative flip. The theory's quantitative form (\sqrt{J(x)} \propto p(x), or a generalization with an exponent depending on the loss function) is testable, and whether the data match it quantitatively is the difference between "consistent with efficient coding" and "measures the efficient code."

Ecological reach. An inversion induced in the lab shows the code is adaptable; it doesn't yet tell us how much of everyday Weber-law behavior is set by recent statistics versus lifelong priors. The interesting follow-up is the mixing ratio.

Where this points

If the result holds up quantitatively, the practical upshot for neuroscience is that Weber's law stops being an axiom and becomes a measurement of the observer's internal prior. That's a genuinely useful inversion: psychophysical discrimination thresholds become a readout of what distribution the brain currently believes it's living in. You could imagine using JND-versus-magnitude curves as a probe of statistical learning β€” in development, in clinical populations, even in trained neural networks, where one can ask whether learned representations show the same \Delta x \propto 1/p(x) signature (there is early work suggesting they do, and this paper sharpens the prediction: the signature should follow the training distribution, causally).

It also strengthens a general lesson that keeps recurring across perception research: apparent laws of the perceiver are often laws of the environment in disguise. Efficient coding is the transfer function.

What to read first: without access to the full text I can't point to section numbers, but prioritize (1) the figure showing JND or discriminability as a function of magnitude under both distribution conditions, per modality β€” that's the entire result in one plot β€” and (2) the model comparison or performance analysis backing the claim that adaptation improved task performance, since that's what separates efficient coding from a mere exposure aftereffect.