Issue 25 · Jun 15–21, 2026
Week 2026-W25
3,249 papers scanned 150 shortlisted 10 picked $10.28 spent
This week's strongest signal is a cluster of assumption-breaking results: AI now out-persuades expert human debaters and canvassers, glia turn out to release noradrenaline, medical VLMs often ignore the image, and simple geometry in frozen encoders beats giant world models at spotting physics violations. Robotics contributes real capability jumps (five-ball juggling, human-video-beats-robot-data pretraining) and neuroscience/BCI bring surprising mechanisms and a focal spinal-stimulation method. Many high-upvote LLM efficiency papers are interesting but under-evidenced from abstracts alone; treat their headline numbers skeptically.
-
AI systems out-persuade expert humans
Four preregistered experiments (~19k conversations) showing frontier AI out-persuades laypeople, tournament winners, professional canvassers, and championship debaters—and transfers to real donations—is unusually strong, consequential evidence rather than a benchmark score. It also isolates a mechanism: the edge largely comes from rapidly deploying more information, and vanishes under human speed/length constraints.
Look for Check how persuasion was measured and whether the human-speed 'tie' result generalizes beyond the specific coaching setup.
-
A glial source of noradrenaline shapes synaptic integration and motor adaptation
A direct challenge to the canonical view that noradrenaline comes only from long-range nuclei: Bergmann glia are shown to synthesize and release noradrenaline via VMAT2, shaping Purkinje synaptic integration and motor adaptation. If it holds, it reframes local neuromodulation.
Look for Preprint with limited effect sizes—scrutinize the specificity of the VMAT2/genetic manipulations and whether release is truly glial rather than contamination from sparse axons.
-
Task-Error Residual Learning for Real-Robot Five-Ball Juggling
Stable five-ball juggling on real anthropomorphic arms converging essentially after one failed attempt is a striking real-robot capability, paired with a concrete methodological finding that directional task-error feedback plus an informative prior are jointly necessary (a fixed-Jacobian Newton update wins).
Look for Note how much the idealized analytic prior and hardware calibration carry the result, and whether the approach extends beyond periodic, well-modeled tasks.
-
GEOPHYS: The Geometry of Physical Plausibility
Claims that five geometric properties of frozen image-encoder features detect physical implausibility far better and cheaper than V-JEPA 2, GPT-4o, Gemini, and a dozen video diffusion models—and correlate with human EEG to object-permanence violations. The combination of capability, efficiency, and a neural link is exactly the kind of 'why it works' result worth reading.
Look for Verify that the near-perfect LikePhys/IntPhys2 numbers aren't exploiting dataset artifacts, and how the geometry signals behave outside curated physics benchmarks.
-
Vision-language models for chest radiography do not always need the image
An intervention-based audit shows a 119B multimodal chest-radiograph model is statistically indistinguishable from a 7B text-only baseline, and a text-only model matches radiologist accuracy while grounding at zero. This directly overturns the assumption that high medical-VLM accuracy demonstrates visual reasoning, and offers a reusable causal audit.
Look for Consider whether the finding is specific to finding-name-prior-heavy datasets and how the grounding metrics would transfer to other medical imaging tasks.
-
SPIDER -- Stitched Power-spectra for Inferring Directed information flow from incomplete and asynchronous Experimental Recordings
SPIDER tackles a genuine methodological gap—recovering directed, frequency-specific interactions from asynchronous, partially overlapping recordings with no shared clock—and validates across calcium imaging, Neuropixels, and human iEEG, revealing a theta-band feedforward hierarchy with hippocampal formation at its source across species.
Look for Scrutinize the consistency guarantees' assumptions and whether the matrix-completion step for never-co-observed regions introduces spurious directed flow.
-
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
A controlled comparison finding that egocentric human video, under a good filtering pipeline, outperforms teleoperated robot data for embodied pretraining—24% lower action loss and 90% higher out-of-distribution success—reverses a widely held assumption about the best pretraining source.
Look for Check how much of the gain is due to the filtering/labeling pipeline versus the data source, and whether the data budgets are truly matched.
-
Adaptive Charge Modulation Enables Focal, Selective Spinal Cord Stimulation
Adaptive Charge Modulation achieves focal, deep spinal activation with single-muscle selectivity from epidural surface electrodes—normally requiring implanted contacts—with chronic 68-day stability and dense 2,112-channel brain-spine recordings. A genuinely new neuromodulation strategy relevant to neural interfaces.
Look for Evidence is rat-only; watch for how selectivity was quantified and whether the high-frequency suppression mechanism is established rather than inferred.
-
Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning
Diarization-conditioning the acoustic encoder of a spoken LLM while keeping the decoder frozen yields large speaker-attributed transcription gains over Gemini 3 Flash and Voxtral on far-field multi-talker audio, plus strong long-form multi-speaker QA. A concrete architectural fix for a real speech-model limitation.
Look for Assess fairness of the baseline comparisons and whether gains depend on the specific DiCoW/Voxtral pairing versus the general conditioning idea.
-
Functional segregation of body-brain signals in the area postrema
A functional remapping of area postrema cell types—GFRAL neurons (canonically sickness/nausea) responding to dietary fat independently of GDF15, with GIPR neurons gating them via sugar—connects widely used weight-loss-drug targets to natural nutrient sensing in a non-obvious way.
Look for Mouse physiology only; check the causal specificity of the GDF15-independent fat pathway and how cell-type identity was verified.
Also notable
-
Optimal Deterministic Multicalibration and Omniprediction
AI / ML
Resolves an explicit open problem by giving the first deterministic minimax-optimal multicalibration algorithm, with downstream optimal omnipredictors—specialized but a clean theory result.
-
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
AI / ML
3B model claiming frontier-level verifiable reasoning (94.3 AIME26) and a 'compression-coverage' hypothesis; high upvotes but abstract heavily overclaims parity with much larger models.
-
Current World Models Lack a Persistent State Core
AI / ML
WRBench shows current video 'world models' don't advance a latent physical state while objects are unobserved—a useful, broadly-tested diagnostic of a central world-model claim.
-
A Verifiable Search Is Not a Learnable Chain-of-Thought
AI / ML
Controlled evidence that left-to-right CoT distillation cannot teach backtracking search (cryptarithm stays near chance), with a key-revelation intervention isolating the cause.
-
RHO: Your Coding Agent is Secretly a Roboticist
Robotics
RHO trains coding agents to search over multi-file executable robotics policy repositories at training time, deploying single-turn—large gains on perturbed manipulation benchmarks.
-
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
AI / ML
Argues 81–90% of between-model variance in LLM 'personality' tests is directional response bias, challenging the treatment of LLM psychometric profiles as intrinsic.
-
Visual consequences of saccades explain early cortical response dynamics during natural vision
Neuroscience
Reinterprets the post-saccadic lambda EEG response as driven by saccade-induced retinal motion rather than new input at fixation, supported by matched replay experiments.
-
Ventricular Expansion Couples Hyperosmotic Stress to Thirst
Neuroscience
Proposes a brain-scale mechanical route from hyperosmotic stress to thirst via lateral ventricle expansion and periventricular mechanosensitive channels—unexpected but abstract-thin.
-
Adaptive Neural Reorganization Enables Real-Time Finger-Level Robotic Control in BCI-Naïve Stroke Survivors
BCI
Finger-level robotic hand control from noninvasive EEG in nine BCI-naive stroke survivors (84% two-finger); meaningful capability but small sample and preliminary.
-
Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
AI / ML
Expert Tying shares MoE expert parameters across consecutive layers, claiming ~2x expert-memory reduction with negligible quality loss across OLMoE/Qwen3/DeepSeek-style models.
-
Human Universal Grasping
Robotics
Human Universal Grasping learns natural grasps from 1M egocentric smart-glasses frames and retargets to multiple robot hands, with real-world cross-embodiment gains.
-
Variable-Width Transformers
AI / ML
Variable-width ('X-shaped') transformers reduce fitted FLOPs 22% and KV-cache 15% versus uniform-width baselines from 200M–3B—simple, potentially broadly useful architecture finding.
Projects
A strong week for world models and real-world robotics: sub-minute humanoid RL training, streaming audio-visual generation, and large-scale action-conditioned video models all shipped code. Also a notable debunking of LLM-generated CUDA kernel speedup claims from Meta.
-
kingjulio8238/nanoG1
Training a Unitree G1 walking policy from scratch in ~59 seconds on one GPU, via a robot-specialized compiled physics engine hitting 7.25M steps/s, is a genuine orders-of-magnitude efficiency jump for humanoid RL. It ships the full stack: simulation, training code, browser demo, and real-hardware deployment, making it both a capability result and a reusable tool.
Look for Verify the real-robot gait quality and robustness beyond flat-ground walking, and whether the specialized physics engine generalizes to other tasks or is locked to G1 locomotion.
3 min read ·GitHub ↗ ·C·MIT
-
catnip-ai-tech/MaineCoon
Treating synchronized audio-video generation as a native streaming problem — sub-second interaction and up to 47.5 FPS from a 22B model on a single H100 — is a meaningful shift from adapting offline diffusion, and directly in the reader's omni-modal/interactive-systems lane. If the claims hold, this is the kind of real-time social world model people will build on.
Look for The repo appears to be mostly a technical report and links; check whether weights or runnable code exist and whether the latency/FPS numbers are independently reproducible.
3 min read ·GitHub ↗
-
LogosRoboticsGroup/A2World
An action-conditioned multi-view diffusion world model pretrained on 2.1M manipulation trajectories across 20+ embodiments, with one dynamics prior transferring to both policy-evaluation simulation and instruction-conditioned control. Code and checkpoints are released, making it one of the more substantive robot world-model artifacts this year.
Look for Evidence is mostly from the authors' own benchmarks with near-zero external traction; check rollout fidelity over long horizons and whether the released checkpoints reproduce the reported transfer results.
3 min read ·GitHub ↗ ·Python·Apache-2.0
-
facebookresearch/kernel_bench_verified
Shows that reported LLM-generated CUDA kernel speedups can collapse from 1.43x to 0.88x under realistic TF32 baselines and hidden correctness tests — a result that overturns a widely repeated claim about LLM kernel generation. This is exactly the 'explains why something works (or doesn't)' finding the reader values, from Facebook Research.
Look for Very new with minimal adoption; verify the baseline choices are fair and see whether the original KernelBench authors or kernel-generation groups respond to the methodology.
3 min read ·GitHub ↗ ·Python·MIT
-
ShareLab-SII/UniAR
The claim that the visual tokenizer, not the architecture, is the key to unifying understanding, generation, and editing in one autoregressive model is a clean, testable design insight (ICML 2026) — the model can read its own generated tokens without re-encoding. Code, checkpoints, and a live demo are all released.
Look for Visual-decoder training code is withheld and traction is modest; test the demo yourself on editing chains where re-encoding artifacts would normally accumulate.
3 min read ·GitHub ↗ ·Python
-
Multimedia-Semantic-Analytics-Lab/PerceptionDLM
Using diffusion-LM parallel denoising to caption many image regions simultaneously sidesteps autoregressive latency scaling with region count — a genuinely different decoding regime for perception with a reported 3.4x throughput gain. Full release of code, 8B weights, training data, and a benchmark makes it immediately usable.
Look for Check per-region caption quality against strong autoregressive VLMs at matched compute, since parallel decoding often trades fidelity for throughput.
3 min read ·GitHub ↗ ·Python·Apache-2.0
Also notable
-
boogu-project/Boogu-Image
GitHub
Apache-2.0 image generation/editing family claiming near-closed-source quality with ~10x less training data — worth watching, but the headline claim is not yet independently verified.
-
IliaLarchenko/lehome_solution
GitHub
Prizewinning bimanual garment-folding system (1st in sim, 2nd in the real-world round) with an unusually complete release of Pi0.5-based training pipeline, checkpoints, and data.
-
SJTU-DENG-Lab/Streaming-WAM
GitHub
Reframes robot inference-execution overlap as an action-conditioned modeling problem — feed the executing action prefix into future predictions — a clever fix for stale world-model predictions during asynchronous control.
-
Ingrid789/OmniContact_sim2sim
GitHub
Contact-flow chaining of humanoid meta-skills (carry, push, relocate, kick) into long-horizon loco-manipulation, runnable in MuJoCo but sim-only so far.
-
tgo-app-dev/vpipe
GitHub
Native Apple Silicon runtime with custom Metal kernels that fits large video/audio-generation models like LTX-2.5 on 16GB base Macs — impressive systems work with modest traction.
-
JinPLu/WRBench
GitHub
Diagnostic benchmark isolating whether video world models preserve object and spatial state when the camera looks away and returns, with a frozen 23-model evaluation.
-
kaistmm/SeeandSniff
GitHub
ECCV oral on joint visuo-olfactory representation learning — a genuinely unusual multimodal direction, though the README is sparse on results and released artifacts.
-
InternLM/RNGBench
GitHub
Shows multimodal LLMs fail badly at reconstructing hidden state from visual history in non-Markov games, with strong dependence on textual action traces — a revealing evaluation result.
The shortlist: top candidates that survived triage · Archive