Paper Feed

Issue 25 · Jun 15–21, 2026

Week 2026-W25

3,249 papers scanned 150 shortlisted 10 picked $10.28 spent

This week's strongest signal is a cluster of assumption-breaking results: AI now out-persuades expert human debaters and canvassers, glia turn out to release noradrenaline, medical VLMs often ignore the image, and simple geometry in frozen encoders beats giant world models at spotting physics violations. Robotics contributes real capability jumps (five-ball juggling, human-video-beats-robot-data pretraining) and neuroscience/BCI bring surprising mechanisms and a focal spinal-stimulation method. Many high-upvote LLM efficiency papers are interesting but under-evidenced from abstracts alone; treat their headline numbers skeptically.

  1. AI / ML ✓ read

    AI systems out-persuade expert humans

    Kobi Hackenburg, Caroline Wagner, Luke Hewitt, Ben M. Tappin et al.

    Four preregistered experiments (~19k conversations) showing frontier AI out-persuades laypeople, tournament winners, professional canvassers, and championship debaters—and transfers to real donations—is unusually strong, consequential evidence rather than a benchmark score. It also isolates a mechanism: the edge largely comes from rapidly deploying more information, and vanishes under human speed/length constraints.

    Look for Check how persuasion was measured and whether the human-speed 'tie' result generalizes beyond the specific coaching setup.

    8 min read · arXiv ↗ ·PDF

  2. Neuroscience ✓ read

    A glial source of noradrenaline shapes synaptic integration and motor adaptation

    Mach, S., Royer, J., Niu, W., Li, X. et al.

    A direct challenge to the canonical view that noradrenaline comes only from long-range nuclei: Bergmann glia are shown to synthesize and release noradrenaline via VMAT2, shaping Purkinje synaptic integration and motor adaptation. If it holds, it reframes local neuromodulation.

    Look for Preprint with limited effect sizes—scrutinize the specificity of the VMAT2/genetic manipulations and whether release is truly glial rather than contamination from sparse axons.

    5 min read · bioRxiv ↗ ·PDF

  3. Robotics ✓ read

    Task-Error Residual Learning for Real-Robot Five-Ball Juggling

    Kai Ploeger, Jan Peters

    Stable five-ball juggling on real anthropomorphic arms converging essentially after one failed attempt is a striking real-robot capability, paired with a concrete methodological finding that directional task-error feedback plus an informative prior are jointly necessary (a fixed-Jacobian Newton update wins).

    Look for Note how much the idealized analytic prior and hardware calibration carry the result, and whether the approach extends beyond periodic, well-modeled tasks.

    9 min read · arXiv ↗ ·PDF

  4. AI / ML ✓ read

    GEOPHYS: The Geometry of Physical Plausibility

    Christian Internò, Alexander Pondaven, Habon Issa, Fabio Pizzati et al.

    Claims that five geometric properties of frozen image-encoder features detect physical implausibility far better and cheaper than V-JEPA 2, GPT-4o, Gemini, and a dozen video diffusion models—and correlate with human EEG to object-permanence violations. The combination of capability, efficiency, and a neural link is exactly the kind of 'why it works' result worth reading.

    Look for Verify that the near-perfect LikePhys/IntPhys2 numbers aren't exploiting dataset artifacts, and how the geometry signals behave outside curated physics benchmarks.

    9 min read · arXiv ↗ ·PDF

  5. AI / ML ✓ read

    Vision-language models for chest radiography do not always need the image

    Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn et al.

    An intervention-based audit shows a 119B multimodal chest-radiograph model is statistically indistinguishable from a 7B text-only baseline, and a text-only model matches radiologist accuracy while grounding at zero. This directly overturns the assumption that high medical-VLM accuracy demonstrates visual reasoning, and offers a reusable causal audit.

    Look for Consider whether the finding is specific to finding-name-prior-heavy datasets and how the grounding metrics would transfer to other medical imaging tasks.

    8 min read · arXiv ↗ ·PDF

  6. Neuroscience ✓ read

    SPIDER -- Stitched Power-spectra for Inferring Directed information flow from incomplete and asynchronous Experimental Recordings

    Yisi S. Zhang, Daniel Y. Takahashi

    SPIDER tackles a genuine methodological gap—recovering directed, frequency-specific interactions from asynchronous, partially overlapping recordings with no shared clock—and validates across calcium imaging, Neuropixels, and human iEEG, revealing a theta-band feedforward hierarchy with hippocampal formation at its source across species.

    Look for Scrutinize the consistency guarantees' assumptions and whether the matrix-completion step for never-co-observed regions introduces spurious directed flow.

    9 min read · arXiv ↗ ·PDF

  7. Robotics ▲ 14 ✓ read

    HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

    Juncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai et al.

    A controlled comparison finding that egocentric human video, under a good filtering pipeline, outperforms teleoperated robot data for embodied pretraining—24% lower action loss and 90% higher out-of-distribution success—reverses a widely held assumption about the best pretraining source.

    Look for Check how much of the gain is due to the filtering/labeling pipeline versus the data source, and whether the data budgets are truly matched.

    9 min read · arXiv ↗ ·PDF

  8. BCI ✓ read

    Adaptive Charge Modulation Enables Focal, Selective Spinal Cord Stimulation

    Vatsyayan, R., Khoury, F., Porter, T. S., Sinopoulou, E. et al.

    Adaptive Charge Modulation achieves focal, deep spinal activation with single-muscle selectivity from epidural surface electrodes—normally requiring implanted contacts—with chronic 68-day stability and dense 2,112-channel brain-spine recordings. A genuinely new neuromodulation strategy relevant to neural interfaces.

    Look for Evidence is rat-only; watch for how selectivity was quantified and whether the high-frequency suppression mechanism is established rather than inferred.

    6 min read · bioRxiv ↗ ·PDF

  9. AI / ML ✓ read

    Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

    Alexander Polok, Samuele Cornell, Sathvik Udupa, Jan Černocký et al.

    Diarization-conditioning the acoustic encoder of a spoken LLM while keeping the decoder frozen yields large speaker-attributed transcription gains over Gemini 3 Flash and Voxtral on far-field multi-talker audio, plus strong long-form multi-speaker QA. A concrete architectural fix for a real speech-model limitation.

    Look for Assess fairness of the baseline comparisons and whether gains depend on the specific DiCoW/Voxtral pairing versus the general conditioning idea.

    8 min read · arXiv ↗ ·PDF

  10. Neuroscience ✓ read

    Functional segregation of body-brain signals in the area postrema

    Lopez-Cruz, A., Burgos, N. S. F., Hakimi, A. M., Xie, K. et al.

    A functional remapping of area postrema cell types—GFRAL neurons (canonically sickness/nausea) responding to dietary fat independently of GDF15, with GIPR neurons gating them via sugar—connects widely used weight-loss-drug targets to natural nutrient sensing in a non-obvious way.

    Look for Mouse physiology only; check the causal specificity of the GDF15-independent fat pathway and how cell-type identity was verified.

    7 min read · bioRxiv ↗ ·PDF

Also notable

Projects

A strong week for world models and real-world robotics: sub-minute humanoid RL training, streaming audio-visual generation, and large-scale action-conditioned video models all shipped code. Also a notable debunking of LLM-generated CUDA kernel speedup claims from Meta.

  1. GitHub Robotics ★ 46 ✓ read

    kingjulio8238/nanoG1

    nanoG1 - G1 walking policy trained in < 60s

    Training a Unitree G1 walking policy from scratch in ~59 seconds on one GPU, via a robot-specialized compiled physics engine hitting 7.25M steps/s, is a genuine orders-of-magnitude efficiency jump for humanoid RL. It ships the full stack: simulation, training code, browser demo, and real-hardware deployment, making it both a capability result and a reusable tool.

    Look for Verify the real-robot gait quality and robustness beyond flat-ground walking, and whether the specialized physics engine generalizes to other tasks or is locked to G1 locomotion.

    3 min read ·GitHub ↗ ·C·MIT

  2. GitHub Speech / Video ★ 121 ✓ read

    catnip-ai-tech/MaineCoon

    MaineCoon: Pursuing a Real-Time Audio-Visual Social World Model — technical report & project links. 🌐 https://mainecoon.tech/

    Treating synchronized audio-video generation as a native streaming problem — sub-second interaction and up to 47.5 FPS from a 22B model on a single H100 — is a meaningful shift from adapting offline diffusion, and directly in the reader's omni-modal/interactive-systems lane. If the claims hold, this is the kind of real-time social world model people will build on.

    Look for The repo appears to be mostly a technical report and links; check whether weights or runnable code exist and whether the latency/FPS numbers are independently reproducible.

    3 min read ·GitHub ↗

  3. GitHub Robotics ★ 14 ✓ read

    LogosRoboticsGroup/A2World

    [ECCV 2026] Learning Transferable Dynamics Priors from Action to World Modeling

    An action-conditioned multi-view diffusion world model pretrained on 2.1M manipulation trajectories across 20+ embodiments, with one dynamics prior transferring to both policy-evaluation simulation and instruction-conditioned control. Code and checkpoints are released, making it one of the more substantive robot world-model artifacts this year.

    Look for Evidence is mostly from the authors' own benchmarks with near-zero external traction; check rollout fidelity over long horizons and whether the released checkpoints reproduce the reported transfer results.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

  4. GitHub Tooling ★ 14 ✓ read

    facebookresearch/kernel_bench_verified

    Welcome to KernelBench-Verified. This repository provides a robust, realistic evaluation framework for assessing custom CUDA kernels generated by Large Language Models (LLMs).

    Shows that reported LLM-generated CUDA kernel speedups can collapse from 1.43x to 0.88x under realistic TF32 baselines and hidden correctness tests — a result that overturns a widely repeated claim about LLM kernel generation. This is exactly the 'explains why something works (or doesn't)' finding the reader values, from Facebook Research.

    Look for Very new with minimal adoption; verify the baseline choices are fair and see whether the original KernelBench authors or kernel-generation groups respond to the methodology.

    3 min read ·GitHub ↗ ·Python·MIT

  5. GitHub AI / ML ★ 54 ✓ read

    ShareLab-SII/UniAR

    [ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"

    The claim that the visual tokenizer, not the architecture, is the key to unifying understanding, generation, and editing in one autoregressive model is a clean, testable design insight (ICML 2026) — the model can read its own generated tokens without re-encoding. Code, checkpoints, and a live demo are all released.

    Look for Visual-decoder training code is withheld and traction is modest; test the demo yourself on editing chains where re-encoding artifacts would normally accumulate.

    3 min read ·GitHub ↗ ·Python

  6. GitHub AI / ML ★ 77 ✓ read

    Multimedia-Semantic-Analytics-Lab/PerceptionDLM

    Official Repo For PerceptionDLM Codebase

    Using diffusion-LM parallel denoising to caption many image regions simultaneously sidesteps autoregressive latency scaling with region count — a genuinely different decoding regime for perception with a reported 3.4x throughput gain. Full release of code, 8B weights, training data, and a benchmark makes it immediately usable.

    Look for Check per-region caption quality against strong autoregressive VLMs at matched compute, since parallel decoding often trades fidelity for throughput.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

Also notable

The shortlist: top candidates that survived triage · Archive