From the archives · 67 essays and 49 mentions, 1867–1999
Revisited
Old computing papers whose ideas could not work when they were written — for want of compute, data, algorithms, or hardware — and can now. By era; within each era, ranked by how much a 2026 revisit could matter.
This batch spans 1943–1977 and is unusually rich in ideas that hardware, connectomics, or foundation models have only now made testable. The essay picks favor open bets — a canonical cortical algorithm, tractable Solomonoff induction, reversible computing, write-access to visual cortex — over well-trodden lineage pieces, with anchors relegated to mentions. Coverage is deliberately spread across neuroscience, BCI, audio, robotics, computing, and core ML.
Before 1943
This batch reaches back to the era before electronics could carry an idea, and the best of it reads like a to-do list for 2026: self-extending formal reasoners, synergy-based motor control, cortical write-in, thermodynamic computing, and claim-level knowledge graphs. We lead with the ripest open programs and keep two vindicated pieces whose lineage directly instructs current practice in reward modeling and memory.
-
Systems of Logic Based on Ordinals†
Turing's ordinal logics are the only pre-war blueprint for a reasoning system that grows its own trusted foundations, with the ingenuity/intuition split mapping exactly onto today's generator-plus-verifier architectures. Nobody has yet built the explicit loop he described, even though every ingredient (LLM conjecture, Lean verification, expanding axiom bases) now exists.
Now An ordinal-style self-extension system: a model proposes lemmas and definitions, a proof assistant certifies them, and the verified base becomes training and retrieval substrate for the next round—Turing's progression run as an engineering loop.
7 min read·original ↗
-
The Problem of the Interrelation of Coordination and Localization
Bernstein diagnosed exactly the problem current robot foundation models hit: flat muscle-space (or token-space) action outputs cannot cope with context variability, and control must live in task space via flexible synergies with online sensory correction. Modern latent action spaces are rediscovering his architecture piecemeal without the theory.
Now Build dexterous-hand and humanoid controllers explicitly as learned synergy hierarchies with fast sensory-correction loops, and measure the sample-efficiency and robustness gains against flat end-to-end policies.
7 min read
-
Somatic motor and sensory representation in the cerebral cortex of man as studied by electrical stimulation
Penfield established write access to human cortex, but nearly a century later the read side of BCI has raced ahead while patterned, naturalistic stimulation remains primitive. With chronic high-channel arrays now in humans, the encoding model—not the hardware—is the bottleneck, which makes this an ML problem.
Now Learn biomimetic stimulation encoders (percept-conditioned models mapping desired touch, texture, and proprioception to multichannel stimulation patterns) and validate them in closed-loop bidirectional prosthetics.
6 min read·original ↗
-
Über die Entropieverminderung in einem thermodynamischen System bei Eingriffen intelligenter Wesen
Szilard tied information to entropy decades before Shannon, and his thought experiment is now the charter for an entire hardware direction: computing that harnesses thermal noise instead of paying to suppress it. With energy the binding constraint on AI, the gap between current per-op cost and the Landauer limit is the largest untapped efficiency margin in the field.
Now Thermodynamic and probabilistic-bit samplers as native accelerators for the stochastic workloads ML actually runs—sampling, energy-based inference, optimization—benchmarked in joules per effective sample against GPUs.
8 min read·original ↗
-
Traité de documentation : le livre sur le livre, théorie et pratique
Otlet's real idea was not the web but something stronger: decomposing all documents into atomic, linked, provenance-tracked claims, which was impossible until LLM-scale information extraction. Citation graphs and RAG both stop far short of this, and scientific synthesis still lacks a claim-level substrate.
Now An LLM-constructed claim graph over the scientific literature with provenance, contradiction detection, and cumulative synthesis—Otlet's Mundaneum built as infrastructure rather than index cards.
7 min read·original ↗
-
A law of comparative judgment.
Thurstone's 1927 model is literally the core of RLHF reward modeling, but current practice uses only its crudest form, discarding his treatment of per-item discriminal dispersion and correlated judgments. This is a vindicated anchor with a concrete unexploited upgrade path for the most economically important preference-learning pipeline in existence.
Now Heteroscedastic, judge-aware preference models—estimating per-item variance and annotator correlation structure—as drop-in replacements for Bradley-Terry reward models, evaluated on downstream policy quality.
7 min read·original ↗
-
Adaptiveness and Equilibrium
Ashby proposed that adaptive behavior is re-equilibration—keeping essential variables within bounds—rather than maximization, a formalism arguably better suited to long-lived robots and self-maintaining agents than reward. Viability-based objectives remain barely explored despite being computationally tractable now.
Now Agents trained to maintain viability envelopes (energy, integrity, task readiness) as the primary objective with tasks as perturbations, compared head-to-head with reward maximization for long-horizon embodied deployment.
7 min read·original ↗
-
Remembering: a study in experimental and social psychology
Bartlett showed memory is regeneration from a compressed schema, and LLMs are literal Bartlett machines whose hallucinations are his subjects' distortions in kind. His serial-reproduction paradigm is a cheap, quantitative probe that works identically on brains and models, and it is only beginning to be used that way.
Now Iterated-regeneration experiments at scale: measure schema attractors and distortion dynamics in LLMs and humans with the same protocol, turning serial reproduction into a standard interpretability and cognitive-comparison method.
8 min read·original ↗
Also worth knowing
-
The differential analyzer. A new machine for solving differential equations
Ripe now
Bush's differential analyzer is the cleanest historical charter for the analog-accelerator revival: continuous-time analog cores inside a digital wrapper, now with memristors instead of wheel-and-disc integrators.
-
Introduction and removal of reward, and maze performance in rats.
Vindicated
Tolman and Honzik's latent-learning protocol—reward-free pretraining, then instant exploitation when reward appears—deserves to become the standard data-efficiency benchmark for embodied world models.
-
Umwelt und Innenwelt der Tiere
Ripe now
Uexküll's Umwelt is constantly cited and almost never implemented; body- and action-conditioned perceptual objectives for robot foundation models would finally operationalize it.
-
The Croonian lecture.—La fine structure des centres nerveux
Ripe now
Cajal's structural-plasticity hypothesis points at learning by rewiring—dynamic sparse growth and pruning—as an underused alternative to pure weight updates, now testable at scale.
-
On a Type-Reading Optophone
Ripe now
The optophone's fatal flaw was its fixed encoding; ML-optimized, online-adapted sensory-substitution codes have never been seriously tried and could transform assistive interfaces.
-
Handbuch der physiologischen Optik (Vol. III: unconscious inference)
Vindicated
Helmholtz's unconscious inference is broadly vindicated, but the strong claim—cortex literally inverting a learned generative model—is now testable with digital-twin visual cortex models.
1943–1959
These are papers whose central idea was right but starved of compute, data, sensors, or fabrication at birth. We favored entries that are ripe or still open, spread across neuroscience, robotics, audio, hardware, and AI reasoning, with a few vindicated anchors kept only where the lineage teaches something. Best first.
-
Modality and topographic properties of single neurons of cat's somatic sensory cortex
The cortical-column hypothesis—one canonical circuit tiled across cortex—is the single largest open bet in brain-inspired AI, and it directly underwrites mixture-of-experts and thousand-brains style architectures. For sixty years there was no way to say what the column computes; now there is a defined path to find out.
Now Fuse MICrONS-scale EM connectomes with Neuropixels population activity to extract the columnar input-output map, instantiate it as a repeated trainable module, and benchmark against transformers at matched parameters.
8 min read·original ↗
-
What the Frog's Eye Tells the Frog's Brain
The claim that the retina transmits behaviorally relevant events rather than images is now literally realized in event cameras and in-sensor computing, yet the co-design loop the paper implies is barely attempted. It reframes perception as task-specific feature extraction at the sensor, a big lever on latency and energy for robotics.
Now Jointly train 'software-defined retina' front-ends on event hardware with downstream policies, targeting microsecond-latency robotic perception where frame-based pipelines fail.
8 min read·original ↗
-
Electronically controlled manipulator
Force-reflecting master-slave teleoperation is quietly the most important robotics idea of the moment because teleop logs are the training corpus for contact-rich manipulation policies. Goertz built the hardware; modern imitation learning turns it into a data engine for robot foundation models.
Now Deploy fleet-scale, high-fidelity haptic teleoperation as the primary data-collection layer for generalist manipulation policies, treating force feedback as a first-class learning signal.
7 min read·original ↗
-
Production of Reversible Changes in the Central Nervous System by Ultrasound
Reversible, noninvasive modulation of deep brain structures by focused ultrasound is exactly what transcranial FUS now delivers at millimeter scale, and the dosing/mechanism space is still wide open. This is the rare BCI direction that could reach deep circuits without surgery.
Now Build closed-loop noninvasive deep-brain neuromodulation: tFUS steered in real time by EEG/fMRI decoders for depression, pain, and causal circuit mapping.
6 min read·original ↗
-
Some Experiments on the Recognition of Speech, with One and with Two Ears
The cocktail-party problem is well posed and partly solved for clean mixtures, but robust, low-latency, egocentric multi-talker listening remains below human level and is the core unsolved problem for hearables and social robots. It is a distinct, high-value audio frontier rather than another separation tweak.
Now On-device sub-10ms attention-steerable neural hearables conditioned on gaze, speaker enrollment, or EEG-decoded attention, evaluated in real multi-talker rooms.
7 min read·original ↗
-
Probabilistic Logics and the Synthesis of Reliable Organisms From Unreliable Components
Von Neumann's theory of reliable computation from noisy components was shelved because transistors got too good, but near-threshold, analog, and in-memory accelerators bring component noise back as the dominant energy-scaling knob. Statistical redundancy instead of deterministic correctness is a genuinely different design axis.
Now Design ultra-low-voltage or analog accelerators that run components in noisy regimes with multiplexed redundancy, optimizing per-op energy against statistical correction.
8 min read·original ↗
-
Branching dendritic trees and motoneuron membrane resistivity
Cable theory recast neurons as spatially extended computers, and dendritic nonlinearities (NMDA spikes, single-neuron XOR) now show single cells need multi-layer ANNs to emulate. Dendrite-style local nonlinearity and local credit assignment is an underexploited route to efficient learning.
Now Scale dendritic ANNs and two-compartment neuromorphic chips that exploit local nonlinearities and local learning rules, benchmarking parameter and energy efficiency against point-neuron nets.
8 min read·original ↗
-
Programs with Common Sense
McCarthy's Advice Taker defined instructability—changing behavior by being told facts, with consistent updating—which is precisely what LLMs do without any guarantee of persistence or consistency. Closing that gap is a concrete, important research program, not a philosophical one.
Now Build an LLM agent with structured memory and formal scaffolding whose downstream reasoning provably and consistently changes after being told a fact, with benchmarks for advice persistence.
7 min read
-
Realization of a Geometry Theorem Proving Machine
Using a cheap executable model of the domain to prune symbolic proof search is exactly the architecture AlphaGeometry revived to reach olympiad level. Generalizing 'check each inference in a simulator' to LLM chains of thought is a first-class mechanism worth an essay.
Now Filter every chain-of-thought step through executable models—simulators, solvers, calculators, diagrams—making semantic pruning a general reasoning primitive rather than a geometry trick.
7 min read
-
The General and Logical Theory of Automata
Von Neumann's self-reproducing automata anticipated the description/constructor split later confirmed by molecular biology, and macro-scale self-replicating manufacturing is now an engineering target rather than philosophy. The payoff—economies grown from a seed factory—is enormous and almost untried.
Now Design and build a physical self-replicating manufacturing cell (robot arms plus printers plus AI control) that assembles copies of itself from feedstock, with quantified material closure as the metric.
8 min read
-
As We May Think
The memex—lifelong personal memory navigated by associative trails—is finally buildable with cheap storage, embedding retrieval, and LLM summarization, but the 'trails as shareable, provenance-bearing objects' idea remains unexploited. It is the clearest blueprint for augmented knowledge work.
Now A locally stored lifelong multimodal memory with LLM-mediated trail-building, sharing, and provenance, treating associative trails as first-class artifacts.
7 min read·original ↗
-
The chemical basis of morphogenesis
Reaction-diffusion morphogenesis is vindicated in biology and now reappears as neural cellular automata that grow and regenerate target forms, making it a compilation target for self-repairing structure. It is an instructive anchor for the emerging field of growing—rather than building—materials and robot bodies.
Now Learn reaction-diffusion or NCA rules that grow desired structures (materials, tissues, robot morphologies) with built-in self-repair, treating Turing's mechanism as a design compiler.
7 min read·original ↗
Also worth knowing
-
A New Microscopic Principle
Ripe now
Gabor's holography, read as the computing rather than imaging branch, points to photonic/holographic accelerators and associative memories for transformer inference.
-
The limiting information capacity of a neuronal link
Ripe now
The order-of-magnitude information advantage of spike-timing codes argues for temporal-coding-first sensors and neuromorphic hardware that AI still largely ignores.
-
Regularly Occurring Periods of Eye Motility, and Concomitant Phenomena, During Sleep
Still open
REM sleep as fundamental offline generative activity invites engineering explicit consolidation/replay phases into deployed models and testing them against decoded biological replay.
-
A Learning Machine: Part I
Vindicated
Friedberg's program evolution was hopeless with blind mutation but is reborn as LLM-guided program search (FunSearch, AlphaEvolve)—worth revisiting as both method and cautionary baseline.
-
Loss of recent memory after bilateral hippocampal lesions
Vindicated
Patient H.M. defined the fast-episodic/slow-cortical memory split now mirrored in retrieval-augmented and consolidation-based agent memory designs.
-
Zatocoding applied to mechanical organization of knowledge
Ripe now
Zatocoding's superimposed sparse codes are the direct ancestor of hyperdimensional/vector-symbolic memory, a low-power alternative to dense embeddings on in-memory hardware.
-
A Computer Oriented toward Spatial Problems
Vindicated
Unger's 2D SIMD array anticipated GPUs and TPUs and now literally lives inside focal-plane processors for microwatt in-sensor vision.
-
Electrical Simulation of Some Nervous System Functional Activities
Ripe now
Taylor's analog associative network prefigures memristive in-memory learning with local update rules for energy-efficient edge training.
-
Physical Analogues to the Growth of a Concept
Still open
Pask's electrochemically grown 'ear' anticipates physical learning and evolvable hardware where morphology, not just weights, is trained under reward.
-
A universal computer capable of executing an arbitrary number of sub-programs simultaneously
Ripe now
Holland's spatial computer of many concurrent, spawning program-organisms is now buildable on dataflow/wafer-scale hardware and remains untried as a substrate for open-ended software evolution.
1960s
The 1960s laid out an astonishing share of today's agenda decades before the hardware existed: representation change, active vision, brain-machine interfaces, supervisory robotics, and energy-limited computing were all articulated with precision and then shelved. This selection favors ideas that modern compute, sensors, and foundation models finally make executable, with a few vindicated anchors kept where the lineage itself teaches something.
-
On Representations of Problems of Reasoning about Actions
Amarel identified the deepest unsolved problem in AI reasoning: search is trivial once the representation is right, and machines should invent representations, not just search within them. LLMs plus formal verifiers are the first plausible machinery for automating re-representation, making this the single ripest 1960s idea for 2026.
Now Build an agent that iteratively rewrites problem encodings—proposing abstractions, symmetries, and quotient state spaces with an LLM and verifying them with a model checker—scored by search-cost reduction on Amarel's own ladder of missionaries-and-cannibals encodings and on modern planning benchmarks.
8 min read
-
Magnetoencephalography: Evidence of Magnetic Fields Produced by Alpha-Rhythm Currents
Cohen's founding MEG measurement pointed at a contactless, skull-transparent window on neural currents that was physically marginal for fifty years. Wearable optically pumped magnetometers plus modern sequence decoders now make MEG arguably the best noninvasive BCI substrate that nobody has seriously exploited.
Now A whole-head OPM array worn during natural movement, decoded end-to-end with transformer models, benchmarked head-to-head against EEG and implanted interfaces on communication rate and motor decoding.
7 min read·original ↗
-
Augmenting human intellect: a conceptual framework
Engelbart's augmentation framework is more ambitious than anything chat interfaces deliver: explicit shared reasoning structures, measured collective capability, and tools that improve the tools. LLMs supply the missing intelligence half, but the structured-artifact half remains unbuilt, making this a pointed critique of current copilot design.
Now An LLM-backed system maintaining inspectable, typed argument and knowledge structures for teams, evaluated on collective problem-solving throughput rather than per-message chat quality—Engelbart's bootstrapping loop with modern models inside it.
7 min read·original ↗
-
Eye Movements and Vision
Yarbus proved vision is task-conditioned sampling, not uniform image processing, and the quadratic token cost of video in multimodal models has finally made that insight economically urgent. This is a rare case where a 1967 behavioral result directly prescribes an architecture change.
Now Train VLMs with a learned, query-conditioned fixation and tokenization policy targeting 10-100x video-token reduction at equal accuracy, using human scanpaths as both prior and benchmark.
7 min read·original ↗
-
Language identification in the limit
Gold's theorem anchored fifty years of poverty-of-the-stimulus arguments, and LLMs trained on positive-only text are a living counterexample to the way it was interpreted. Revisiting exactly which of Gold's assumptions fail for distributional learners would settle one of the longest debates in cognitive science.
Now Formal characterization of what LLM-style learners provably acquire from positive data under measure-one presentations, paired with controlled experiments on synthetic language classes at the boundary of Gold's negative results.
8 min read·original ↗
-
Supervisory control of remote manipulation
Sheridan and Ferrell diagnosed the exact failure mode of delayed teleoperation and prescribed the fix—humans issue symbolic subgoals to a locally competent executive—sixty years before VLA models made the executive real. This is now the economically decisive architecture for deploying robot fleets and harvesting training data.
Now Language-level supervisory control of many semi-autonomous manipulators per human operator, with delay-tolerant subgoal handoff, measured by robots-per-supervisor and intervention rates in real deployments.
7 min read·original ↗
-
A theory of cerebellar cortex
Marr's cerebellum theory—sparse random expansion plus error-driven plastic readout—was confirmed physiologically and matches modern random-feature theory, yet nobody has built the machine it describes. It offers a concrete, biologically vetted recipe for fast, sample-efficient correction layered on slow foundation policies.
Now A cerebellar sidecar for robot foundation models: massive sparse expansion with a rapidly plastic readout driven by execution error, compared against LoRA-style fine-tuning on adaptation speed and stability.
7 min read·original ↗
-
Irreversibility and Heat Generation in the Computing Process
Landauer identified the only fundamental energy floor in computing and, with it, the loophole: reversible logic pays no erasure cost. With AI energy demand exploding and CMOS within a few orders of magnitude of the limit, reversible computing has shifted from curiosity to plausible engineering frontier.
Now A serious reversible/adiabatic accelerator design study for transformer inference, quantifying achievable energy-per-token against the Landauer bound and against best-case conventional CMOS.
7 min read·original ↗
-
Speech recognition: A model and a program for research
Halle and Stevens framed speech recognition as inference in a generative model of production—explain the waveform, don't classify it—which lost to discriminative methods purely on compute. Generative speech models and amortized inference make synthesis-in-the-loop recognition affordable for the first time, with large expected robustness gains.
Now Use articulatory or diffusion-based speech generators as verifiers and priors inside ASR decoding, evaluated on noisy, accented, and low-resource speech where discriminative models degrade.
7 min read·original ↗
-
Mathematical models for cellular interactions in development I. Filaments with one-sided inputs
L-systems posed the genotype-to-phenotype question—tiny generative programs growing large adaptive structures—that both AI (genomic bottleneck, architecture search) and biology still haven't answered. Differentiable programming and neural cellular automata finally make developmental encodings trainable rather than hand-designed.
Now Learn compact L-system-like developmental genomes whose growth produces networks or robot morphologies, testing whether developmental compression improves generalization and evolvability versus direct parameterization.
7 min read·original ↗
-
Operant Conditioning of Cortical Unit Activity
Fetz showed the brain will learn to control whatever signal you give it feedback on—implying BCI performance is a two-learner problem, not just a decoding problem. The field still mostly optimizes decoders while ignoring the trainable half Fetz demonstrated in 1969.
Now Co-adaptive BCIs where decoder updates and population-level neurofeedback are jointly designed to shape controllable neural subspaces, measured against decoder-only optimization on learning speed and asymptotic control quality.
6 min read·original ↗
-
The Co-ordination and Regulation of Movements
Bernstein's degrees-of-freedom problem and 'repetition without repetition' are exactly the questions 40+ DOF humanoids now force, and his hierarchical-synergy answer is a concrete alternative to flat trajectory cloning. The book is a design document for humanoid learning written before robots existed.
Now Structure humanoid policies as explicit Bernstein levels (tone, synergies, space, action) and test whether variability-rich 'problem-solving' data collection beats trajectory imitation on out-of-distribution generalization.
8 min read
-
Numerical testing of evolution theories
Barricelli ran the first open-ended digital ecologies and saw parasitism and symbiosis emerge before complexity plateaued on IAS-machine memory—a scale failure, not a concept failure. Recent alife results show scale changes what emerges, and open-endedness is now a named goal without a canonical testbed.
Now Datacenter-scale shared-memory digital ecologies with recombining code, instrumented to measure whether symbiogenesis drives sustained complexity growth and whether evolvability itself evolves.
7 min read·original ↗
-
A formal theory of inductive inference. Part II
Solomonoff defined the optimal predictor and left tractability as the only obstacle; LLM-guided program search now makes bounded-resource approximations empirically testable rather than purely philosophical. It also offers the cleanest theoretical lens on why compression-trained models generalize.
Now An explicit bounded-resource Solomonoff predictor using LLM program proposal plus execution-based weighting on ARC-style tasks, testing whether performance tracks description-length compression.
8 min read·original ↗
-
The sensations produced by electrical stimulation of the visual cortex
Brindley's 80-electrode phosphene map is the founding demo of cortical visual prostheses; thousands-channel arrays plus digital-twin stimulation optimization make the full percept-rendering problem newly attackable.
Now Optimize stimulation patterns through a digital twin of the patient's visual cortex so that thousands of channels produce desired percepts, not just points of light.
7 min read·original ↗
-
MH-1, a computer-operated mechanical hand
Kept from the first selection round: The first sensor-guided robot hand: touch and proximity sensors let a computer-controlled arm find, grasp, and stack blocks without vision, using sensor-conditioned programs. Manipulation as feedback, not playback.
Now Train policies primarily on dense tactile feedback for in-hand and occluded tasks where vision fails, treating vision as auxiliary rather than primary.
7 min read·original ↗
Also worth knowing
-
Speech Synthesis (Transmission-Line Vocal Tract Model)
Ripe now
Kelly-Lochbaum physics-based vocal-tract synthesis becomes compelling again as a differentiable, interpretable decoder for speech BCIs and ultra-low-bitrate TTS.
-
Intercommunicating cells, basis for a distributed logic computer
Ripe now
Lee's CPU-less sea of broadcast-and-match memory cells anticipated processing-in-memory; mapping attention and vector search onto such arrays is a live 10-100x energy opportunity.
-
Functional Tuning of the Nervous System with Control of Movement or Maintenance of a Steady Posture
Ripe now
Feldman's equilibrium-point hypothesis suggests an action space of postures and stiffness fields rather than torques—directly testable on today's variable-impedance humanoids.
-
Dual Control Theory. I
Ripe now
Feldbaum's dual control formalized probing-versus-exploiting as optimal control over belief states; learned world models with calibrated uncertainty finally make approximate dual controllers buildable.
-
Intracerebral radio stimulation and recording in completely free patients
Vindicated
Delgado's stimoceiver proposed activity-contingent closed-loop deep-brain stimulation in 1968; biomarker-triggered psychiatric neuromodulation is its overdue serious execution.
-
Vision Substitution by Tactile Image Projection
Ripe now
Bach-y-Rita's sensory substitution proved radical cross-modal plasticity; learned encoders co-trained with users could turn it into sensory addition, not just substitution.
-
Theory of Self-Reproducing Automata
Still open
Von Neumann's universal constructor is the vindicated logical foundation for self-replicating manufacturing, an agenda autonomous labs and self-printing fabs are just beginning to touch.
-
Non-Holographic Associative Memory
Ripe now
Willshaw's near-optimal binary associative memory maps directly onto memristor crossbars as a picojoule episodic-memory layer for agents.
-
Theory of Optical Information Storage in Solids
Ripe now
Van Heerden's volumetric holographic storage with associative recall could become a physical nearest-neighbor layer under retrieval-augmented systems, now that photonics supplies every missing component.
-
Perception of the speech code.
Ripe now
Liberman's motor theory—perception recovers articulatory gestures—predicts that a shared articulatory latent space should improve ASR/TTS generalization across accents and disordered speech, a clean testable claim.
1970s
The 1970s produced blueprints that the decade's hardware could not honor: machine discovery, memory-native computing, cortical prostheses, and perception built for action. Reading the longlist against 2026 capabilities, the striking pattern is how many of these ideas are not merely vindicated but still unfinished, with LLMs, connectomics, photonics, and robot learning finally supplying the missing parts. The selection below favors ideas that remain genuinely open, with a few vindicated anchors whose lineage teaches something current work forgets.
-
AM: An Artificial Intelligence Approach to Discovery in Mathematics as Heuristic Search
AM is the purest early statement of open-ended machine discovery: invent concepts, judge their interestingness, and search onward, rather than solve posed problems. It stalled for exactly the reasons LLMs and formal verifiers now fix—no rich prior over mathematics and no grounding—making it the highest-leverage revisit on the list.
Now An AM-style loop with an LLM proposer, a Lean verifier, and a learned interestingness model, run at FunSearch scale but aimed at concept invention rather than single conjectures.
8 min read
-
Memristor-The missing circuit element
Chua predicted a fundamental circuit element from pure symmetry, and its payoff—analog in-memory matrix multiplication—is precisely the primitive that dominates AI energy budgets today. The race to make memristive inference and on-chip local learning work is live and unwon, which makes the 1971 framing newly practical rather than historical.
Now A serious attempt is a memristor crossbar accelerator benchmarked on LLM inference in joules per token, paired with local plasticity rules that exploit the device's intrinsic memory for on-chip adaptation.
7 min read·original ↗
-
Phosphenes produced by electrical stimulation of human occipital cortex, and their application to the development of a prosthesis for the blind
Dobelle's phosphene maps established that camera-driven cortical stimulation could restore structured vision, half a century before the electrode counts and encoders existed to make it useful. With Neuralink's Blindsight and penetrating arrays now in play, the paper's open problem—how to translate images into stimulation—is the field's central question.
Now Learn the image-to-stimulation encoder end-to-end through a differentiable model of cortical responses, closing the loop on perceptual reports from implanted volunteers.
7 min read·original ↗
-
An Organizing Principle for Cerebral Function: The Unit Module and the Distributed System
Mountcastle's claim that all of neocortex runs one repeated canonical computation was untestable framing for fifty years; MICrONS-scale connectomics and multi-area recordings finally make it an empirical question. It is also the biological bet underlying architecture-uniform AI, so the answer matters in both substrates.
Now Extract the microcircuit's computation from joint connectomic and activity data, then test whether a single learned module transfers across sensory modalities the way the hypothesis demands.
8 min read
-
A truth maintenance system
Doyle built machinery for a problem that did not yet exist: agents that accumulate beliefs, contradict themselves, and need principled, minimal retraction. LLM agents are that system, and bolting justification graphs with provenance onto their memory is a concrete, almost untried attack on hallucination persistence.
Now An agent memory where every stored claim carries a dependency structure, and detected contradictions trigger TMS-style minimal retraction rather than silent overwriting.
7 min read·original ↗
-
The Ecological Approach to Visual Perception
Gibson lost to the reconstruction paradigm partly because 1979 compute favored Marr, but robot learning increasingly shows that action-conditioned perception beats explicit 3D reconstruction. Affordances as the content of vision is the right target representation for robot foundation models, and nobody has built one that takes the thesis fully seriously.
Now Train robot policies whose visual representations are explicitly affordance-structured—graspable, traversable, openable—and evaluate by manipulation generalization instead of reconstruction metrics.
7 min read
-
Organization of the Hearsay II speech understanding system
The blackboard architecture solved multi-expert coordination with principled, uncertainty-driven scheduling, then died because its knowledge sources were too weak; today's multi-agent LLM frameworks have strong components but reinvent the coordination ad hoc. Hearsay-II is the missing theory for the systems everyone is currently building.
Now Rebuild the blackboard with LLM and multimodal knowledge sources plus confidence-propagating opportunistic scheduling, applied to streaming multimodal understanding and agent orchestration.
8 min read·original ↗
-
Model-Directed Learning of Production Rules
Meta-DENDRAL actually induced publishable chemistry rules by constraining hypothesis search with a mechanistic domain model—a discipline most current 'AI scientist' efforts lack. Unlike AM's open-ended search, this is the template for grounded, closed-loop discovery in the wet sciences.
Now LLM hypothesis generation constrained by mechanistic simulators, closed-loop with automated lab experiments in chemistry or biology, with induced rules held to journal-publication standard.
7 min read·original ↗
-
Fully parallel, high-speed incoherent optical method for performing discrete Fourier transforms
Goodman showed light through a fixed mask computes a full matrix-vector product in one pass—an idea that waited fifty years for a workload dominated by exactly that operation. Photonic tensor cores are now commercializing this, but precision, calibration, and on-chip training remain genuinely open.
Now A photonic transformer-inference accelerator with electronic nonlinearities, plus hardware-in-the-loop training that absorbs analog imperfections into the weights.
8 min read·original ↗
-
A versatile system for computer-controlled assembly
Freddy II—vision-guided assembly of products from a jumbled heap, specified declaratively—is nearly the exact mission statement of today's robot foundation-model labs, and it was used by the Lighthill report as evidence AI robotics was hopeless. The gap between its ambition and current VLA capabilities is a precise measure of what remains unsolved.
Now A declarative assembly spec compiled by an LLM into vision-language-action policies with force feedback, evaluated on arbitrary products from cluttered bins.
7 min read·original ↗
-
Adaptive pattern classification and universal recoding: I. Parallel development and coding of neural feature detectors
Grossberg named the stability-plasticity dilemma before mainstream ML knew it had one, and his match-based resonance gating is a concrete mechanism that continual-learning research has largely skipped. Catastrophic forgetting in continually pretrained foundation models makes this the right moment to test it at scale.
Now Add ART-style match/mismatch gating to transformer memory and adapters, and evaluate genuine continual pretraining without replay buffers.
8 min read·original ↗
-
Single Units and Sensation: A Neuron Doctrine for Perceptual Psychology?
Barlow's question—do single, sparse, highly selective units carry perceptual meaning?—is now askable simultaneously in brains (Neuropixels, concept cells) and in networks (sparse autoencoder features, superposition theory). A serious two-substrate comparison would be one of the cleanest neuroscience-interpretability bridges available.
Now Directly compare sparse-coding statistics of large-scale neural recordings against SAE feature geometry in deep networks, testing the neuron doctrine quantitatively in both systems.
7 min read·original ↗
-
Logical Reversibility of Computation
Bennett's proof that computation need not dissipate energy is the endgame for compute efficiency, and adiabatic/reversible accelerators for energy-dominated matmuls are finally being attempted.
Now Design an adiabatic/reversible accelerator around energy-dominated matmuls, benchmarked in joules per token, while pushing reversible-network training as the software analogue.
7 min read·original ↗
Also worth knowing
-
A Computational Model of Skill Acquisition
Ripe now
Sussman's HACKER—an agent that abstracts its own bugs into a growing, transferable library of fix schemata—is the piece today's self-debugging coding agents still lack.
-
Real-time detection of brain events in EEG
Ripe now
Vidal's single-trial EEG control loop anticipates the current push for EEG foundation models that could give noninvasive BCI usable bandwidth.
-
Simple memory: a theory for archicortex
Vindicated
Marr's hippocampal capacity math is the still-unexploited design theory for fast episodic memory modules that write once and distill into the slow weights of foundation models.
-
Planning in a hierarchy of abstraction spaces
Ripe now
ABSTRIPS's open problem—automatically learning which details to ignore at which planning level—remains the crux of long-horizon planning for LLM and world-model agents.
-
Modeling by shortest data description
Vindicated
Rissanen's MDL is newly testable at scale now that LLMs are operationally giant compressors, with prequential codelength as both training signal and theory of in-context learning.
-
Brain Function and Adaptive Systems: A Heterostatic Theory
Ripe now
Klopf's hedonistic-neuron thesis, which seeded reinforcement learning, points to an untried backprop-free credit-assignment paradigm suited to neuromorphic hardware.
-
Responsive environments
Ripe now
Krueger's VIDEOPLACE—deviceless full-body interaction with a responsive generated world—is exactly where camera-based spatial computing plus real-time video generation is converging.
-
Force Feedback in Precise Assembly Tasks
Ripe now
Inoue's taxonomy of compliant contact strategies is still the best framing for today's frontier of tactile, force-conditioned insertion policies.
-
The brain wave equation: a model for the EEG
Ripe now
Nunez's cortical wave equation becomes actionable with OPM-MEG and personalized cortical geometry, offering wave phase and direction as control signals for closed-loop stimulation and BCI priors.
1980s
From the 1980s longlist, these are the papers whose core ideas were starved of compute, sensors, or data rather than wrong, and whose 2026 revival paths are concrete. The essay picks span self-improving AI, memory consolidation, touch and active sensing for robots, energy-limited computing, binding in cortex, silent speech, and agent architectures. Vindicated classics are kept to a minimum in favor of ripe, still-open bets.
-
Eurisko: A program that learns new heuristics and domain concepts
EURISKO is the direct ancestor of FunSearch, AlphaEvolve, and every AI-scientist pipeline: heuristics as inspectable, mutable objects that improve the system's own search. Its 1983 failure was purely a missing proposal distribution, which LLMs now supply, making this the most instructive lineage in the collection.
Now An LLM-driven EURISKO where heuristics are code, evolved and sandbox-evaluated across domains, with the meta-level of which heuristics to keep itself learned—and an honest comparison against the hand-built brittleness that killed the original.
8 min read·original ↗
-
Two-stage model of memory trace formation: A role for “noisy” brain states
Buzsáki's two-stage model—fast labile hippocampal encoding, then prioritized replay into cortex during offline states—is the biological answer to catastrophic forgetting, and it was later causally confirmed. No one has seriously implemented the full architecture at foundation-model scale, despite experience replay being its shallow descendant.
Now A fast episodic store attached to an LLM whose contents are selectively, surprise-prioritized replayed into slow weights during scheduled offline consolidation phases, benchmarked on continual-learning suites against replay-free finetuning.
6 min read·original ↗
-
Design and Implementation of a VLSI Tactile Sensing Computer
Raibert and Tanner put computation in the skin itself in 1982, forty years before manipulation policies became starved for exactly this modality. Modern processes can put thousands of taxels plus neural inference in a fingertip at milliwatts, so the original bottleneck is simply gone.
Now Mass-producible fingertip tactile chips running learned local encoders that stream compressed contact features into VLA policies, closing the touch gap that currently limits dexterous manipulation.
7 min read·original ↗
-
Visual routines
Ullman predicted that relational vision—inside/outside, connectivity, counting—requires serial routines of attention shifts, marking, and tracing, and VLMs fail today on precisely those tasks. This is a rare case of a 1980s framework that diagnoses a 2026 failure mode exactly.
Now Multimodal models equipped with an explicit visual-routine interpreter (callable attention, marking, and tracing operations), evaluated on relational benchmarks where monolithic VLMs plateau.
7 min read·original ↗
-
Conservative logic
Fredkin and Toffoli proved computation need not dissipate energy at all, back when CMOS had a 10^8 kT cushion that made the result academic. With AI datacenters at tens of gigawatts and device energy within a few orders of the Landauer limit, reversibility is now a live engineering bet with real startups.
Now Reversible or adiabatic logic for the energy-dominated inference regime, accepting area and speed penalties for 10–100x energy-per-op reduction, evaluated at accelerator scale.
8 min read·original ↗
-
A learning algorithm for boltzmann machines
The Boltzmann machine failed because Gibbs sampling to equilibrium was hopeless on serial machines—a compute problem, not a conceptual one. Physical hardware that thermalizes natively (p-bit circuits, thermodynamic chips) turns the algorithm's fatal cost into a free physical process.
Now Train full, unrestricted Boltzmann machines on stochastic hardware where equilibration costs picojoules, targeting generative modeling and sampling workloads at a small fraction of GPU energy.
7 min read·original ↗
-
Active perception
Bajcsy argued perception is a control problem—choose where to look and what to touch—decades before the field standardized on passive datasets, and passive internet pretraining has now visibly hit limits for manipulation. The learning machinery to train sensing policies finally exists.
Now Embodied VLMs trained to select viewpoints, zoom, and haptic probes to resolve task-relevant ambiguity, scored against passive models on occlusion-heavy manipulation and inspection.
7 min read·original ↗
-
Stimulus-specific neuronal oscillations in orientation columns of cat visual cortex.
Whether synchrony binds distributed features is a 35-year-old open question that Neuropixels-scale recording and optogenetic entrainment can finally settle, and the binding problem has independently resurfaced in AI via object-centric and synchrony-based networks. A decisive answer would matter on both sides.
Now Causal tests—entrain or disrupt gamma synchrony across areas during perception—paired with synchrony/phase-coding as an explicit binding substrate in neuromorphic and object-centric models.
8 min read·original ↗
-
A Speech Prosthesis Employing a Speech Synthesizer-Vowel Discrimination from Perioral Muscle Activities and Vowel Production
Sugie and Tsunoda decoded vowels from perioral EMG in 1985; high-density surface EMG, deep sequence models, and neural vocoders have since turned silent speech into a solvable end-to-end problem, and wrist-EMG products prove the modality ships at consumer scale. The payoff spans accessibility and a genuinely new input channel.
Now A wearable facial/submental EMG interface with a streaming EMG-to-speech-token model, targeting laryngectomy patients and private voice input for AR glasses.
8 min read·original ↗
-
SOAR: An Architecture for General Intelligence
Today's LLM agent frameworks are reinventing SOAR's ideas—impasse-driven subgoaling, caching successful traces as reusable skills—ad hoc and without its principled control structure. SOAR failed only because every production was hand-coded, precisely the gap LLMs fill.
Now An agent architecture marrying SOAR's impasses, subgoals, and chunking with LLM policies, so successful reasoning traces compile into a growing skill library instead of brittle prompt orchestration.
7 min read·original ↗
Also worth knowing
-
Sparse Distributed Memory
Ripe now
Kanerva's sparse distributed memory is mathematically close to transformer attention and is a ready-made design for write-once episodic agent memory on in-memory or neuromorphic hardware.
-
“Neural” computation of decisions in optimization problems
Ripe now
Hopfield-Tank computation-by-settling is now literal hardware in Ising machines and annealers; the honest benchmark against modern SAT/MIP solvers is still owed.
-
A massively parallel architecture for a self-organizing neural pattern recognition machine
Ripe now
ART's vigilance-gated plasticity is an untested architectural answer to catastrophic forgetting in continually trained foundation models.
-
'Put-That-There': Voice and Gesture at the Graphics Interface
Ripe now
Put-That-There's deictic fusion of speech, gaze, and gesture is finally implementable via multimodal LLMs and is the natural command layer for AR and home robots.
-
Model-based analysis synthesis image coding (MBASIC) system for a person's face
Ripe now
Model-based face coding anticipated generative codecs; neural avatars now make kilobit-rate telepresence technically ready but unstandardized.
-
Recording action potentials from cultured neurons with extracellular microcircuit electrodes
Ripe now
Pine's neurons-on-electrodes program has matured into closed-loop organoid intelligence (DishBrain), a barely explored substrate for studying learning rules in vitro.
-
Studying artificial life with cellular automata
Still open
Langton's search for open-ended evolution in CA universes can now be automated at GPU scale with learned novelty metrics—artificial life's core problem remains unsolved.
-
RF powering of millimeter- and submillimeter-sized neural prosthetic implants
Ripe now
Heetderks derived neural dust from first principles in 1988; sub-mm wireless motes now exist but no full-scale distributed cortical interface has been deployed.
1990s
This batch from the 1990s is unusually rich: the decade wrote down many of the ideas the field is only now equipped to execute. The eight picks below span robot learning from human video, vector-symbolic memory, temporal codes in the brain, analog hardware theory, write-side brain interfaces, body-brain co-evolution, personal agents, and information-theoretic representation learning — each one blocked then by compute, data, or hardware that now exists.
-
Learning by watching: extracting reusable task knowledge from visual observation of human performance
Kuniyoshi built the full learning-from-human-video pipeline — observe once, extract a symbolic action plan, re-execute in a new configuration — in 1994, and it stalled purely on perception. That pipeline is now the central bet of every VLA lab, yet his structured action-segmentation intermediate is largely absent from today's end-to-end approaches.
Now Pretrain manipulation policies on internet-scale human video using Kuniyoshi-style action graphs (grasp/move/place with dependencies) as the intermediate representation, and test whether the structure beats raw behavior cloning on one-shot generalization.
7 min read·original ↗
-
Holographic reduced representations
Plate solved the dimension-blowup problem of symbolic structure in vector spaces: circular convolution stores arbitrarily nested structure in a fixed-width, decodable vector. As context windows balloon and attention costs dominate, a compressed, compositional, fixed-size memory is exactly the alternative nobody has seriously grafted onto LLMs.
Now Build HRR-based episodic memory and pointer mechanisms for LLMs — fixed-width compressed context with decodable structure — and benchmark against long-context attention on retrieval and multi-hop reasoning at matched compute.
9 min read·original ↗
-
Phase Relationship Between Hippocampal Place Units and the EEG Theta Rhythm
Phase precession showed the brain multiplexes sequence information within each oscillatory cycle — spike phase carries time-compressed future trajectories, a fundamentally different scheme from rate coding or token-by-token rollout. It is now confirmed across species and regions, yet no AI architecture exploits oscillation-multiplexed planning sweeps.
Now Design sequence models where a global oscillation phase-codes a compressed lookahead sweep within each step, testing whether phase-multiplexed planning beats autoregressive rollout on latency and horizon.
7 min read
-
Analog Versus Digital: Extrapolating from Electronics to Neurobiology
Sarpeshkar gave a quantitative theory of exactly when analog beats digital and prescribed the hybrid architecture — low-precision distributed analog with periodic digital restoration — that the brain uses. Neural network inference at 4-bit precision is precisely the workload his theory favors, yet accelerator design still lacks this kind of principled per-layer analog/digital allocation.
Now Use the noise-precision-power framework as a co-design tool for transformer accelerators, deciding per-layer and per-operation where analog in-memory compute pays, with hardware-aware training absorbing the mismatch.
8 min read·original ↗
-
Somatosensory discrimination based on cortical microstimulation
Romo showed artificial cortical input is perceptually interchangeable with real sensation — the founding result for the write side of BCIs, which remains far behind decoding. With multichannel stimulation now in human participants, the missing piece is the learned encoder mapping desired percepts to stimulation patterns.
Now Train sensory encoders (percept-to-stimulation) end-to-end in closed loop with human intracortical stimulation, targeting rich artificial touch and proprioception for bidirectional prosthetics.
7 min read·original ↗
-
Evolving virtual creatures
Sims demonstrated that co-evolving morphology and controller produces designs no engineer would find, then the idea died for want of ~10^6x more simulation. GPU physics now supplies exactly that, yet the robotics industry converged on a single humanoid form factor without ever running the search.
Now Foundation-model-guided co-design of robot bodies and policies in massively parallel differentiable simulation, closed with automated fabrication — morphology remains robotics' least exploited optimization lever.
7 min read·original ↗
-
Agents that reduce work and information overload
Maes designed the interaction loop today's agent products lack: agents that learn one user's habits by observation and earn autonomy gradually, with competence and trust calibrated together. LLMs finally supply the competence her memory-based learners lacked, but current assistants skip the learning-from-observation and graduated-trust machinery entirely.
Now Personal agents that continually learn from watching a single user's actual behavior, with Maes' trust model governing which actions they may take unsupervised — evaluated on longitudinal real-user deployments.
7 min read·original ↗
-
The Information Bottleneck Method
The information bottleneck gives representation learning a single clean objective — keep what predicts Y, discard the rest — that predates and could discipline much of modern interpretability. Variational estimators and large models make it operational, and a rigorous IB account of what transformer layers keep and discard is still missing.
Now An IB analysis of LLM layer-by-layer compression (including in-context learning), plus IB-designed objectives for multimodal encoders and agent memory where deciding what to forget is the core problem.
7 min read
Also worth knowing
-
Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks
Vindicated
Fast weight programmers are formally linear transformers thirty years early, and the lineage runs straight to test-time training — the most instructive vindicated anchor in the batch.
-
Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
Ripe now
Rao & Ballard's predictive coding is finally testable both in vivo (layer-resolved recordings of error units) and in silico (local-learning alternatives to backprop on neuromorphic hardware).
-
The PHANToM Haptic Interface: A Device for Probing Virtual Objects
Ripe now
The PHANToM's kHz force-rendering insight points to bilateral force-reflecting teleoperation rigs that capture the contact channel vision-only robot demonstrations discard.
-
Electrostriction of polymer dielectrics with compliant electrodes as a means of actuation
Still open
Dielectric elastomer artificial muscles are one materials breakthrough (lifetime) away from obsoleting gearmotors for compliant humanoid joints.
-
A Possibility for Implementing Curiosity and Boredom in Model-Building Neural Controllers
Ripe now
Compression-progress curiosity computed from a foundation world model has still never been seriously tried, despite ICM/RND partially vindicating the idea.
-
Reactivation of Hippocampal Ensemble Memories During Sleep
Vindicated
Hippocampal replay founded experience replay in RL; the unexploited residue is prioritized, reverse, and generative replay scheduling for continual-learning LLMs and robots.