Paper Feed

Issue 23 · Jun 1–7, 2026

Week 2026-W23

3,877 papers scanned 150 shortlisted 10 picked $12.78 spent

This week leans heavily toward doing more with less: several papers internalize context into parameters or memory (Frames2LoRA, cartridges, adaptive video codecs) and toward assumption-breaking diagnostics of what our models and benchmarks actually measure. Formal-math agents took a visible jump, physics-constrained generative inference got a foundational correction, and there's a rich neuroscience crop touching directly on computation, coding, and chaos. Note that many of the flashiest claims (Cosmos 3, several agent self-improvement results) rest on thin quantitative evidence in the abstracts, so discount accordingly.

  1. AI / ML ▲ 4 ✓ read

    Frames2LoRA: Parametric Video Internalization for Vision-Language Models

    Manan Suri, Sarvesh Baskar, Dinesh Manocha

    A genuinely new framing: a hypernetwork reads a VLM's layerwise activations while it encodes a video and emits a LoRA adapter in one forward pass, so the video lives in weights rather than context. The reported 6–80x latency and up to 1,500x token reductions with non-inferior quality, plus rank-space composition of independently generated adapters, point at a real new direction for long-video and memory.

    Look for Check whether 'statistical non-inferiority' hides systematic quality loss on harder QA, and how well adapter composition actually holds beyond a few chunks.

    9 min read · arXiv ↗ ·PDF

  2. AI / ML ✓ read

    Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

    Yousef Radwan

    Claims a single valence direction recoverable from just nine emotion anchors that appears in text, vision, audio, AND human EEG encoders never jointly trained, with causal ablation effects in LLMs. If the cross-modal alignment and controls hold, this is an unusually strong statement about shared representational geometry across models and brains.

    Look for Scrutinize the EEG and cross-modal probe alignment for confounds, and note the honest caveats: bounded to continuous attributes and family-specific steering.

    11 min read · arXiv ↗ ·PDF

  3. Neuroscience ✓ read

    Predictable Mean-Field Chaos in Random Recurrent Neural Networks

    Alkesh Yadav, Vladimir Shaidurov, Jonathan Kadmon

    A striking theoretical result: a random recurrent network can have a positive Lyapunov exponent yet be perfectly predictable at the single-neuron level from its continuous past, showing predictive complexity and microscopic instability scale differently. This directly bears on how we interpret neural variability and chaos in both brains and RNNs.

    Look for The result leans on continuous-time DMFT idealizations; look for finite-network, finite-sampling validation and how the log-p horizon degrades under noise.

    9 min read · arXiv ↗ ·PDF

  4. AI / ML ✓ read

    Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

    Jui-Hui Chung, Ziyang Cai, Zihao Li, Qishuo Yin et al.

    Reframes theorem proving around a global dependency-graph blueprint that is refined on failure rather than recursively decomposed, reporting 99.2% MiniF2F, 75.6% PutnamBench, and solutions to recent olympiad problems at a claimed ~500x lower cost. This is a substantial capability-and-efficiency jump in formal math.

    Look for Verify how much comes from the 284B backbone versus the blueprint method, and watch for benchmark contamination on recent competition problems.

    8 min read · arXiv ↗ ·PDF

  5. Robotics ✓ read

    Wave Focusing in Metamaterials: Tactile Displays Beyond the Diffraction Limit

    Gregory Reardon, Max Linnander, Dustin Goetz, Neeli Tummala et al.

    A real fabricated haptic display that uses a locally resonant metamaterial plate to focus tactile waves beyond the plate's diffraction limit, achieving a tenfold reduction in virtual-pixel area with few actuators and validated behaviorally. Genuinely new physics/hardware for distributed touch, backed by builds and human experiments rather than simulation.

    Look for Note bandwidth/frequency constraints and whether independent multi-point control degrades as more simultaneous pixels are demanded.

    9 min read · arXiv ↗ ·PDF

  6. AI / ML ✓ read

    The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

    Jian Xu, Yanning Wu, Delu Zeng, John Paisley et al.

    Argues that the standard recipe of projecting a generative prior onto a hard PDE constraint samples the wrong posterior because it omits a co-area (Fixman) Jacobian, and shows the bias is large (up to 20x the noise floor). This is a potentially foundational correction for the fast-growing area of generative PDE inverse problems.

    Look for Evidence is on controlled problems with an i.i.d. arbiter; watch whether CoCoS remains tractable and accurate on realistic high-dimensional inverse tasks.

    10 min read · arXiv ↗ ·PDF

  7. AI / ML ✓ read

    The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models

    Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin, Bao Li et al.

    A clean, surprising result: LLMs' failure to correct their own reasoning errors is largely an artifact of the chat template's role labeling, not a cognitive deficit—relabeling identical erroneous text as an external role raises correction rates by 23–93 points. This reframes self-correction and how we evaluate instruction tuning.

    Look for Check robustness across chat templates and whether the fix generalizes beyond the tested math/logic tasks or just gates explicit flagging.

    8 min read · arXiv ↗ ·PDF

  8. Robotics ✓ read

    What Are We Actually Benchmarking in Robot Manipulation?

    Tianchong Jiang, Xiangshan Tan, Samuel Wheeler, Luzhe Sun et al.

    Concrete audits showing popular manipulation benchmarks (LIBERO, CALVIN) are weak proxies for capability: a 0.09B language-free probe hits near-SOTA on LIBERO, most gains aren't statistically significant, and modest within-range pose randomization breaks CALVIN policies. Provides reusable diagnostics the field badly needs.

    Look for Consider whether the four diagnostics themselves fully capture real-world manipulation ability, and how the audited benchmarks compare to RoboCasa/RoboTwin.

    9 min read · arXiv ↗ ·PDF

  9. AI / ML ▲ 16 ✓ read

    dots.tts Technical Report

    Shi Lian, Changtao Li, Bohan Li, Hankun Wang et al.

    A strong, fully open continuous-autoregressive TTS system with a prediction-friendly AudioVAE, full-history flow conditioning, reward-free self-correction, and 54–85 ms first-packet latency via MeanFlow distillation. Directly in the reader's speech/voice interest with credible quality and real-time claims plus released checkpoints.

    Look for Comparisons are mostly open-source-SOTA framing; check multilingual robustness and how distillation affects expressiveness versus the non-distilled model.

    9 min read · arXiv ↗ ·PDF

  10. Neuroscience ✓ read

    Intrinsic Population Dynamics are a Neuronal Substrate for Visual Attention

    Schmidt, F. H., Mlynarski, W., Georges, A., Sumser, A. et al.

    Reports structured, stimulus-independent 'blob-like' population dynamics in the superior colliculus that emerge with learning, predict trial-by-trial behavior, and amplify sensory responses up to fourfold—casting intrinsic dynamics as an active attentional substrate rather than noise. A candidate shift in how we think about attention and intrinsic activity.

    Look for Watch how strongly the causal claims are supported and whether the excitatory-inhibitory model is doing explanatory work or just fitting.

    9 min read · bioRxiv ↗ ·PDF

Also notable

Projects

A strong week for open speech and audiovisual generation (a leading open TTS release and a long-horizon audiovisual world model), plus two genuinely new training ideas — RNN pretraining without BPTT and editable KV caches — alongside Meta's non-invasive brain-to-text release.

  1. GitHub Speech / Video ★ 1,286 ✓ read

    studio-dots-ai/dots.tts

    A 2B fully-continuous autoregressive TTS with Apache-2.0 code and weights, strong multilingual zero-shot cloning, 48 kHz output, and streaming/distilled variants — this is exactly the kind of open speech release the reader tracks. Traction and completeness (training + inference code) suggest it will become a widely used baseline.

    Look for Verify the multilingual and cloning benchmarks against Fish/CosyVoice-class systems independently, and check real streaming latency on your hardware rather than reported numbers.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

  2. GitHub Speech / Video ★ 1,972 ✓ read

    jd-opensource/JoyAI-Echo

    JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

    Maintaining cross-shot audiovisual continuity over ~5-minute videos plus an enterable, causally rolled-out world model with joint visual, ambient audio, music, and speech is a real capability step beyond short-clip video generation. Code, checkpoints, and strong early traction back it up.

    Look for Check the actual quality and consistency of the 5-minute samples and the compute/inference profile needed — long-horizon claims often degrade badly past the cherry-picked demos.

    3 min read ·GitHub ↗ ·Python

  3. GitHub BCI ★ 917 ✓ read

    facebookresearch/brain2qwerty

    Non-invasive decoding of typed sentences from MEG and EEG brain recordings using a convolutional encoder, transformer, and character-level language model.

    Sentence-level brain-to-text decoding from non-invasive MEG/EEG is a meaningful BCI capability advance, published in Nature Neuroscience with released code and a Spanish MEG/EEG dataset. Directly in the reader's BCI wheelhouse and reproducible enough to build on.

    Look for Note that MEG requires a shielded room (not wearable), check the character error rates for EEG vs MEG separately, and that the v2 dataset remains embargoed.

    3 min read ·GitHub ↗ ·Python

  4. GitHub AI / ML ★ 71 ✓ read

    akarshkumar0101/smt

    Pretraining Recurrent Networks without Recurrence

    Training nonlinear RNNs without backpropagating through recurrence — using a transformer teacher to generate one-step memory-transition targets — is a genuinely different framing that attacks BPTT's core parallelism and credit-assignment limits. Working PyTorch code and comparisons on language and pixel-sequence tasks make it more than a proposal.

    Look for Check how it scales beyond small models and whether the DAgger-style correction handles compounding state drift at long horizons; the teacher-student dependency may cap final quality.

    3 min read ·GitHub ↗ ·Jupyter Notebook·Apache-2.0

  5. GitHub AI / ML ★ 13 ✓ read

    19PINE-AI/programmable-kv

    Models Take Notes at Prefill: KV Cache Can Be Editable and Composable (arXiv:2606.17107) — paper, code, results, and interactive companion.

    Reframing the KV cache as editable, composable program state — with causal mechanism experiments showing prefill writes conclusions onto downstream tokens — is both an interpretability insight and a potentially important serving primitive (append-only errata, transplantable skills). Very novel, with code and multi-model results.

    Look for It is days old with almost no external validation; test whether edits stay coherent under long generations and across model families before treating the claims as established.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

  6. GitHub AI / ML ★ 22 ✓ read

    divelab/OPDLM

    On-policy distillation for converting pretrained AR language models into block-diffusion models, with released data, code, and 0.6B–8B checkpoints, is a substantive and directly testable contribution to the AR-vs-diffusion LM debate. Training the student on its own diffusion trajectories rather than teacher-forced targets is the interesting methodological piece.

    Look for Compare the converted models' quality/speed tradeoff against the AR originals and other diffusion LMs yourself — conversion papers often hide capability regressions on reasoning tasks.

    3 min read ·GitHub ↗ ·Python·MIT

Also notable

The shortlist: top candidates that survived triage · Archive