Paper Feed

Issue 25 · Jun 15–21, 2026

Every candidate

All 3,249 papers were scored from their abstracts by gpt-5.6-luna; 1,171 were not skipped. Shown here: the top 150 of those, in score order. Picks are marked.

  1. strong Neuroscience picked score 5.5

    A glial source of noradrenaline shapes synaptic integration and motor adaptation

    Mach, S., Royer, J., Niu, W. et al.

    This study reports that Bergmann glia in the cerebellum can synthesize and release noradrenaline, using calcium-dependent VMAT2-mediated release rather than serving only as passive support cells. The glial noradrenaline influences Purkinje-cell synaptic integration and is necessary for motor adaptation, suggesting a local complement to the canonical long-range noradrenergic system.

    It presents a potentially assumption-breaking source of neuromodulation—local noradrenaline release by glia—with cell-specific, imaging, mechanistic, and behavioral evidence, although the preprint abstract provides limited detail on effect sizes and replication.

  2. strong AI / ML picked score 5.4

    Optimal Deterministic Multicalibration and Omniprediction

    Georgy Noarov, Aaron Roth

    The paper gives the first deterministic predictor achieving the minimax-optimal ~O(ε^-3) sample complexity for multicalibration, matching prior randomized methods. It extends the construction to outcome indistinguishability and derives optimal deterministic omnipredictors and panpredictors, resolving several open questions.

    It resolves an explicit open problem by showing that randomization is not required for optimal multicalibration sample complexity, with further consequences for omniprediction and outcome indistinguishability.

  3. strong Neuroscience picked score 5.4

    SPIDER -- Stitched Power-spectra for Inferring Directed information flow from incomplete and asynchronous Experimental Recordings

    Yisi S. Zhang, Daniel Y. Takahashi

    SPIDER estimates directed, frequency-specific interactions when brain regions are observed in separate, partially overlapping recordings with no shared clock. It stitches local power spectra, uses spectral factorization and matrix completion, and reports validation on simulations, calcium imaging, Neuropixels, and human intracranial EEG; notably, it finds a theta-band feedforward hierarchy originating in the hippocampal formation across species and modalities.

    This addresses a major limitation of effective-connectivity analysis with a genuinely new asynchronous multi-session formulation, and supports it with broad real-data validation plus a cross-species, cross-modality biological finding.

  4. maybe AI / ML picked▲ 126 score 5.4

    VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

    Sen Xu, Shixi Liu, Wei Wang et al.

    This report presents a 3B-parameter language model trained with curriculum supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation for verifiable reasoning. It reports unusually strong results on AIME26, LiveCodeBench, unseen LeetCode contests, and instruction following, claiming performance comparable to much larger frontier models, and proposes that verifiable reasoning can be compressed into small models while broad knowledge cannot.

    The claimed small-model parity with frontier systems would be highly consequential and surprising, but the abstract provides too little methodological and comparative detail to trust the unusually large claims without close inspection.

  5. strong Robotics picked score 5.2

    Task-Error Residual Learning for Real-Robot Five-Ball Juggling

    Kai Ploeger, Jan Peters

    The authors use directional task-error feedback and an analytic motion prior to train residual controllers for real-robot juggling, rather than relying on scalar rewards and random exploration. Two anthropomorphic Barrett WAM arms achieve stable three-, four-, and five-ball juggling, with the reported five-ball system recovering after one failed attempt and then improving monotonically; a fixed-Jacobian Newton update was the most reliable learner across the tested variants.

    This is a striking real-robot capability result—five-ball juggling after essentially one failed trial—paired with a concrete finding that informative directional supervision and a useful prior are jointly necessary, rather than merely reporting a small control improvement.

  6. strong AI / ML picked score 5.2

    GEOPHYS: The Geometry of Physical Plausibility

    Christian Internò, Alexander Pondaven, Habon Issa et al.

    GEOPHYS extracts five geometric signals from frame-level embeddings of frozen image encoders to judge whether videos obey basic physical laws, without an external multimodal judge or specialized training. It reports near-perfect discrimination on LikePhys and IntPhys2, where several large video and multimodal models are near chance, and improves MAGI-1 video generation verification from 50.01% to 64.50% while using less time and memory than a V-JEPA 2 verifier. The paper also links these geometric signals to human EEG responses to object-permanence violations.

    The striking result is that simple geometry in frozen image-encoder features reportedly detects physical implausibility far better and more cheaply than much larger video/world models, while also showing correspondence with human neural responses.

  7. strong AI / ML picked score 5.2

    Vision-language models for chest radiography do not always need the image

    Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams et al.

    The paper proposes a causal audit for chest-radiograph vision-language models that tests whether answers change appropriately when relevant or irrelevant image regions are occluded or replaced. Across nine systems and two datasets, image-free or image-ignoring models nearly match multimodal accuracy, including a 119B multimodal model matching a 7B text-only baseline; only some systems and findings show genuine image grounding. It also finds that confidence identifies ungrounded answers only for models that actually use the image.

    This directly challenges the common assumption that high medical VLM accuracy demonstrates visual reasoning, while providing a reusable intervention-based audit and evidence across models, datasets, resolutions, and prompts.

  8. maybe AI / ML ▲ 75 score 5.2

    DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

    Cong Wan, Zeyu Guo, Zijian Cai et al.

    DataClaw0 trains a reusable multimodal model to turn noisy raw streams into intent-conditioned, schema-aligned training examples, avoiding repeated calls to proprietary vision-language models. Across 4B, 9B, and 27B models, joint training across five domains underperforms specialists at small scales but surpasses them near 18B, with gains attributed to cross-domain transfer; it also retains 59% of in-domain performance on entirely unseen domains and improves several downstream tasks.

    The learned, reusable data-tailoring formulation and the reported capacity-dependent crossover from specialization to cross-domain generalization are genuinely interesting, but the abstract gives limited quantitative detail and the broad deployment impact is not yet established.

  9. strong AI / ML picked score 5.1

    AI systems out-persuade expert humans

    Kobi Hackenburg, Caroline Wagner, Luke Hewitt et al.

    Across four preregistered experiments involving nearly 19,000 conversations, frontier AI systems persuaded people more effectively than laypeople, tournament winners, professional canvassers, and world-championship debaters—even when humans selected the topic, prepared, practiced, and received financial incentives. The advantage largely disappeared when AI was restricted to human-like response speed and message length, suggesting that its main edge is rapidly deploying much more information; in a real fundraising test, AI generated nearly three times as many donations as professional canvassers.

    This is unusually broad, preregistered evidence that conversational AI can outperform highly skilled human persuaders and transfer that advantage to consequential real-world behavior, while also identifying a plausible mechanism rather than merely reporting benchmark performance.

  10. maybe AI / ML ▲ 114 score 5.0

    DreamX-World 1.0: A General-Purpose Interactive World Model

    DreamX Team, Yancheng Bai, Rui Chen et al.

    DreamX-World is an interactive text/image-to-video model designed for controllable, long-horizon rollouts with camera movement, revisiting, and prompted events. It combines camera-aware positional encoding, retrieved visual memories, autoregressive distillation, self-generated long-context training, and substantial inference optimizations; the authors report 16 FPS on eight RTX 5090 GPUs and higher scores than two competing world models on a short evaluation.

    The combination of persistent scene memory, controllable camera navigation, and relatively fast interactive generation is a meaningful direction, but the evidence is limited to a five-second evaluation and self-reported aggregate scores, with no demonstrated real-world interaction or broad capability analysis.

  11. maybe Robotics picked▲ 14 score 5.0

    HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

    Juncheng Ma, Jianxin Bi, Yufan Deng et al.

    This paper compares egocentric human video with teleoperated real-robot trajectories for pretraining embodied models, followed by a small amount of robot-data adaptation. Under matched data budgets, the human-video models reportedly achieve 24% lower robot action-prediction loss and substantially higher task success, including 90% higher out-of-distribution success. The main claim is that diverse human video can be a better source for learning general world representations, while limited robot data is mainly needed to align the action space.

    The reported reversal of the usual assumption that robot trajectories are the best pretraining source, especially the large out-of-distribution gain, is genuinely interesting, but the abstract does not establish how broad or well-controlled the comparison is or how much the filtering and labeling pipeline contributes.

  12. maybe AI / ML ▲ 21 score 5.0

    The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

    Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng et al.

    The paper uses a discriminator trained to distinguish real data from a flow-matching model’s samples as an automatic reward for KL-regularized RL, avoiding human preference labels. The discriminator’s representation-space logit is intended to approximate the data/model likelihood ratio, and this substantially improves image fidelity across several flow-matching backbones—for example, reducing SiT FID from 9.38 to 2.62—while also improving later preference-based post-training. The main claim is that RL is correcting a mismatch between training-time velocity regression and the perceptual properties that matter at sampling time.

    The combination of discriminator-derived density-ratio rewards with RL for correcting flow-matching models is a meaningful and broadly tested idea, with unusually large reported gains, but discriminator guidance and likelihood-ratio rewards are established ingredients and the abstract alone does not establish how robust the gains are beyond the reported image-generation setup.

  13. strong Robotics picked score 5.0

    RHO: Your Coding Agent is Secretly a Roboticist

    Karim Elmaaroufi, Justin Svegliato, Sarunas Kalade et al.

    RHO trains coding agents to search for and optimize whole, interpretable robotics policy repositories—prompts, tools, perception, planning, and control code—using execution and environment reward rather than teleoperation demonstrations. The resulting policies run in a single pass at deployment and substantially outperform prior agentic or vision-language baselines on perturbed manipulation tasks, while also improving an LLM-in-the-loop benchmark from 23.5% to 44.3% with fewer calls and lower latency.

    The combination of training-time search over multi-file executable policy harnesses and single-turn deployment addresses a real bottleneck in code-based robotics, with sizable gains across several benchmarks rather than a small isolated improvement.

  14. maybe AI / ML ▲ 210 score 5.0

    LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

    Jian Yang, Shawn Guo, Wei Zhang et al.

    The paper trains 7B coding models with parallel looped Transformers and studies how many shared-computation loops are beneficial. Two loops substantially outperform a non-looped model, including large gains on SWE-bench Verified (43.0 to 64.4) and Multi-SWE (14.0 to 31.0), while three or more loops perform worse. The authors attribute this non-monotonic behavior to useful refinement from the second loop being outweighed by positional mismatch and oscillatory, less diverse updates in later loops.

    The large, broad software-engineering gains and evidence for an optimal rather than monotonically beneficial test-time loop count are notable, but the core architecture is an existing parallel-loop Transformer and the claims are based on one model family and reported benchmarks.

  15. maybe Robotics ▲ 50 score 5.0

    Playful Agentic Robot Learning

    Junyi Zhang, Jiaxin Ge, Hanjun Yoo et al.

    The paper gives an embodied coding agent a pre-task “play” phase in which it invents exploratory tasks, attempts them, diagnoses failures, and stores successful behaviors as reusable code skills. These skills improve held-out robot tasks by 20.6 and 17.0 percentage points over a no-play baseline, and can also be retrieved by other Code-as-Policy agents, yielding gains in simulation and real-world transfer without fine-tuning.

    Self-directed skill acquisition before task specification is a promising and fairly distinct direction, with unusually large reported gains and cross-agent/real-world transfer, though the abstract does not establish how broadly the results generalize beyond the tested environments.

  16. maybe AI / ML ▲ 20 score 4.9

    Native Active Perception as Reasoning for Omni-Modal Understanding

    Zhenghao Xing, Ruiyang Xu, Yuxuan Wang et al.

    OmniAgent treats long-video understanding as an active perception problem: the model iteratively decides what audio-visual evidence to inspect, records it in textual memory, and reasons over the accumulated observations rather than processing the whole video uniformly. It combines trajectory-based supervised fine-tuning with turn-aware reinforcement learning, and reports improving performance with more reasoning turns; its 7B model reaches 50.5% on LVBench versus 47.3% for Qwen2.5-VL-72B across results on ten benchmarks.

    Native, learned active perception for omni-modal video agents is a meaningful direction, and the reported small-model advantage is notable, but the abstract gives limited detail about compute-normalized costs, action quality, and breadth of the claimed state-of-the-art results.

  17. maybe Neuroscience picked score 4.9

    Ventricular Expansion Couples Hyperosmotic Stress to Thirst

    Zhu, T., Jiang, C., Yang, J. et al.

    This study proposes that hyperosmotic stress triggers thirst not only through chemical osmolality signals, but also by expanding the brain’s lateral ventricles. In mice, preventing or relieving this ventricular deformation reduced drinking, and periventricular cells expressed candidate mechanosensitive channels such as Tmem63b and Piezo1.

    The proposed brain-scale mechanical route from osmotic stress to thirst is a genuinely unexpected mechanism, but the abstract provides limited quantitative and causal detail beyond the mouse interventions and candidate-channel evidence.

  18. maybe AI / ML ▲ 19 score 4.9

    HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization

    Zhentao Tan, Wei Chen, Jingyi Shen et al.

    HydraHead mixes full attention and linear attention within individual attention heads rather than assigning one attention type per layer. The authors use interpretability analysis to retain full attention for retrieval-critical heads, normalize the two kinds of outputs, and transfer the design with distillation; they report matching a 3:1 layer-wise hybrid using only a 7:1 linear-to-full-attention ratio and a 69% improvement over baseline at 512K context after training on 15B tokens.

    Head-level hybridization guided by functional specialization is a genuinely interesting design direction with potentially large long-context efficiency gains, but the abstract leaves the baseline definitions, absolute results, and breadth of evaluation too unclear for a strong recommendation.

  19. maybe Robotics ▲ 15 score 4.9

    ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

    Wenli Xiao, Jia Xie, Tonghe Zhang et al.

    ENPIRE builds a closed-loop system in which coding agents reset real-world scenes, run robot policies, inspect failures, modify training code or algorithms, and repeat the process across one or more robots. The authors report that this setup autonomously reaches 99% success on dexterous tasks including pin-box organization, zip-tie fastening, and tool use, with faster improvement when using a robot fleet.

    The notable idea is automating the full policy-improvement loop on physical robots rather than merely using agents for offline code generation, but the abstract gives too little detail about baselines, human involvement, costs, and generalization to justify a strong verdict.

  20. maybe AI / ML ▲ 35 score 4.8

    JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

    Lanxiang Hu, Zhaoxiang Feng, Yulun Wu et al.

    JetSpec is a speculative-decoding method that uses a single parallel draft pass while preserving causal, branch-specific conditioning in the proposed token tree. On Qwen3 dense and MoE models, it reportedly turns larger draft budgets into longer accepted prefixes, reaching up to 9.64× speedup on MATH-500 and 4.58× on open-ended chat, with additional vLLM serving results.

    The combination of one-pass drafting with causally consistent tree construction addresses a real bottleneck in speculative decoding, but the abstract gives limited comparative and systems detail, so the unusually large speedups need verification.

  21. maybe AI / ML ▲ 141 score 4.8

    Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

    Kangsheng Duan, Ziyang Xu, Wenyu Liu et al.

    Moebius is a 0.22B-parameter diffusion inpainting specialist designed to match the quality of an 11.9B-parameter FLUX.1-Fill-Dev model. It combines a compact local/global interaction block with latent-space, multi-granularity distillation, and claims comparable or better quality with over 15× faster inference. The main result is extreme task-specific compression rather than a new general-purpose generation capability.

    The claimed 50× parameter reduction and 15× speedup would be highly useful if validated, but the abstract provides no quantitative results or evidence about fairness of the comparison, so it merits a look rather than strong prioritization.

  22. maybe AI / ML ▲ 51 score 4.8

    Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

    Shengyuan Ding, Xilin Wei, Xinyu Fang et al.

    The paper introduces RNG-Bench, an interactive benchmark for testing whether multimodal language models can remember briefly seen visual information and use it later in non-Markov games. Its card-matching and 3D-maze tasks reach roughly 128K-token contexts and 350 images, and the authors report that frontier models remain unsaturated; a proposed Memory Gap suggests forgetting, rather than poor action choice, causes most errors. Fine-tuning on demonstrations improves both the benchmark and existing tasks.

    It offers a useful, relatively clean formulation of long-horizon multimodal memory in closed-loop interaction and reports the non-obvious finding that forgetting dominates decision errors, though it is still primarily a benchmark paper and the abstract gives few quantitative comparisons.

  23. maybe AI / ML ▲ 10 score 4.8

    ProCUA-SFT Technical Report

    Jaehun Jung, Ximing Lu, Brandon Cui et al.

    The paper presents ProCUA-SFT, a 3.1M-sample synthetic dataset for training computer-use agents, generated by running tasks on live desktops and filtering them with automated precondition checks. Fine-tuning UI-TARS 7B on one epoch reportedly raises OSWorld success from 26.3% to 45.0%, while training on the larger human-trajectory AgentNet dataset causes severe negative transfer; the authors attribute the gain to grounded task generation and matching training contexts to inference-time layouts.

    The large, reproducible-looking gain and the finding that more human trajectory data can substantially hurt performance make this worth a look, although the evidence is centered on one benchmark and a synthetic-data pipeline with limited ablation detail in the abstract.

  24. maybe AI / ML ▲ 10 score 4.8

    Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

    Qian Zhao, Kunlong Chen, Changxin Tian et al.

    The paper argues that FP4 training instability is not just a hardware or recipe issue: the geometrically nonuniform E2M1 format introduces a systematic downward rounding bias that compounds through layers and is worsened by random Hadamard transforms. It proposes UFP4, using a uniform E1M2/INT4 grid with Hadamard transforms and selective stochastic rounding, and reports lower loss degradation than E2M1 baselines across dense and MoE models up to 124B parameters.

    The proposed explanation that FP4 training failures arise from a systematic geometric shrinkage bias, rather than merely insufficient precision, is a potentially important and non-obvious insight, but the abstract gives no quantitative margins or independent validation to justify a strong recommendation.

  25. maybe Robotics ▲ 28 score 4.8

    Guava: An Effective and Universal Harness for Embodied Manipulation

    Haowen Liu, Xirui Li, Shaoxiong Yao et al.

    Guava is a model-agnostic framework for embodied manipulation that repeatedly cycles through multimodal perception, reasoning, and semantic actions rather than directly predicting low-level controls. The authors distill this design into a 4B open-source model using fewer than 2,000 simulated trajectories, reporting comparable performance to frontier proprietary models in simulation and real-world tasks, including novel objects, instructions, and long-horizon behavior.

    The combination of a reusable tool-use harness, very small simulated training set, and claimed transfer to real-world generalization is notable, but the abstract provides no quantitative comparisons or task scale to substantiate the broad performance claims.

  26. maybe AI / ML picked▲ 16 score 4.8

    Variable-Width Transformers

    Zhaofeng Wu, Oliver Sieberling, Shawn Tan et al.

    This paper tests giving transformer layers different widths instead of using the same width throughout. Its “> <former” keeps the early and late layers wide while narrowing the middle, and reportedly improves language-modeling loss over parameter-matched uniform models from 200M to 3B parameters while reducing fitted FLOPs by 22% and KV-cache costs by 15%.

    The consistently observed efficiency gains across dense and MoE scales make this more than a minor architecture tweak, though the abstract lacks enough absolute results and downstream evaluation to justify a strong verdict.

  27. maybe Robotics ▲ 24 score 4.7

    ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Yuyang Zhang, Wenyao Zhang, Zekun Qi et al.

    ImageWAM replaces video prediction in world action models with a pretrained image-editing model: the model denoises an instruction-conditioned target-frame edit, then uses its KV caches as context for an action expert without actually generating the image. The authors report better performance than VLA and comparable WAM baselines across simulation and real-robot experiments, while reducing compute to one-sixth and latency to one-quarter of video-based WAMs; attention patterns concentrate on task-relevant visual changes.

    The shift from dense video forecasting to image-editing representations for action prediction is a plausible and useful new direction, with substantial reported efficiency gains and real-robot validation, though the abstract lacks quantitative task details and independent evidence that the gains generalize.

  28. maybe AI / ML ▲ 13 score 4.7

    MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

    Lichen Bai, Tianhao Zhang, Shitong Shao et al.

    MaineCoon proposes a 22B-parameter autoregressive audio-visual model intended for interactive social-world simulation. It claims real-time streaming at up to 47.5 FPS with sub-second interaction on one GPU, plus long-horizon generation using cache management and several training and alignment techniques. The potentially new aspect is treating human-centric social interaction, rather than physical or game environments, as the target for a real-time world model, but the abstract gives little quantitative evidence beyond throughput.

    The claimed combination of real-time audio-visual generation, long-horizon streaming, and social interaction could be important, but the abstract relies heavily on broad “first” and SOTA claims without reporting quality, latency, hardware details, or comparative results.

  29. maybe AI / ML ▲ 4 score 4.7

    When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

    Xuanfei Ren, Tengyang Xie

    This paper develops a theory of offline reinforcement learning when each trajectory has only a scalar outcome label rather than step-by-step rewards. It gives a pessimistic actor-critic method with matching statistical rates for recovering the usual cumulative-reward objective, extends the analysis to preference feedback, and shows that more general nonlinear outcome objectives can be exponentially hard in horizon unless specific information-preservation conditions hold.

    The combination of matching upper and lower bounds with an exponential impossibility result identifies a meaningful boundary between learnable and fundamentally information-starved outcome-supervised offline RL, though the abstract provides no empirical validation and the impact is primarily theoretical.

  30. maybe AI / ML ▲ 35 score 4.7

    EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

    Chang Nie, Chaoyou Fu, Junlan Feng et al.

    EvoEmbedding replaces static, segment-independent embeddings with a recurrent latent memory that updates as a long context is processed, allowing the same query to retrieve different information depending on preceding context. The authors train it on 180K examples, add a memory queue to reduce recurrent representation collapse, and report better long-context retrieval and agent-memory performance than larger embedding models and dedicated memory systems.

    The stateful, context-evolving embedding formulation is a genuinely interesting departure from standard semantic retrieval and could matter for agent memory, but the abstract gives no quantitative margins or detail sufficient to establish that the claimed gains are broad and robust.

  31. maybe AI / ML picked▲ 19 score 4.7

    Current World Models Lack a Persistent State Core

    Jinpeng Lu, Dexu Zhu, Haoyuan Shi et al.

    The paper introduces WRBench, a benchmark that tests whether video world models maintain and advance a latent physical state while an object is temporarily out of view, rather than simply resuming the last observed appearance. Evaluating 9,600 videos from 23 models, it finds that this failure is widespread across model families, control paradigms, and model scales: systems track what was seen but do not reliably simulate unseen events to completion.

    This is a useful and non-obvious diagnostic of a central world-model claim, supported by broad cross-model evaluation, though it is primarily a benchmark and diagnosis rather than a demonstrated solution.

  32. maybe AI / ML picked▲ 3 score 4.7

    A Verifiable Search Is Not a Learnable Chain-of-Thought

    Harsh Patel

    The paper studies whether a model can learn a short, verifiable program simply by imitating its chain of thought. Across nine synthetic reasoning tasks, forward computations transfer reasonably well, but backtracking cryptarithm search remains near chance despite strong arithmetic accuracy, large models, reinforcement learning, and self-training; exposing the hidden key largely removes the failure. The authors argue that left-to-right chain-of-thought distillation cannot faithfully teach search over information-free structure, whereas precomputing the search and training recall-plus-verification can work.

    The controlled comparison between forward computation and search, including the key-revelation intervention and scaling across backbones, offers a genuinely non-obvious explanation for a common reasoning failure, though the evidence is primarily from synthetic cryptarithm-like tasks and may not generalize broadly.

  33. maybe AI / ML ▲ 77 score 4.7

    Learning from the Self-future: On-policy Self-distillation for dLLMs

    Yifu Luo, Zeyu Chen, Haoyu Wang et al.

    This paper adapts on-policy self-distillation to diffusion language models, which generate tokens in arbitrary order rather than left to right. Its d-OPSD method uses self-generated suffixes as privileged information and applies supervision at denoising-step level; on four reasoning benchmarks, it reportedly beats SFT and RLVR while using about 10% as many optimization steps as RLVR.

    The suffix-conditioned, step-level formulation is a meaningful new post-training direction for dLLMs, but the abstract provides limited evidence beyond four benchmarks and a relative efficiency claim without absolute gains or broader validation.

  34. maybe Robotics ▲ 119 score 4.7

    Geometric Action Model for Robot Policy Learning

    Jisang Han, Seonghu Jeon, Jaewoo Jung et al.

    The paper introduces GAM, a language-conditioned robot policy built by splitting a pretrained geometric foundation model: early layers encode observations, while a causal predictor forecasts future geometric representations that are decoded into actions. It aims to combine 3D geometric priors, temporal prediction, language, and control in one backbone, and reports improved accuracy, robustness, speed, and model size across simulated and real-robot manipulation tasks.

    The integration of a geometric foundation model as both a future world model and action policy is a substantive architectural direction for contact-rich manipulation, but the abstract gives no quantitative results or details sufficient to judge whether the claimed broad gains are real.

  35. maybe AI / ML ▲ 15 score 4.7

    SP$^3$: Spherical Priors for Plug-and-Play Restoration

    Sean Man, Ron Raphaeli, Matan Kleiner et al.

    The paper uses Spherical Encoders as generative image priors inside a plug-and-play restoration algorithm, alternating latent-space projection with a closed-form data-consistency update. It claims sharp outputs from the first iteration and perceptual quality comparable to zero-shot diffusion and flow-based restoration methods while running 3–630 times faster across several restoration tasks.

    The combination of a structured spherical latent prior with fast, gradient-free plug-and-play restoration could be practically important, especially if the large speedups hold across diverse degradations, but the abstract provides no quantitative quality results or details sufficient for a strong recommendation.

  36. maybe AI / ML ▲ 65 score 4.6

    Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

    Byung-Kwan Lee, Ximing Lu, Shizhe Diao et al.

    ZPPO uses a stronger model’s answers as prompt content rather than injecting them directly into the student’s policy-gradient updates. For failed questions, it creates candidate-discrimination prompts containing correct and incorrect answers, plus prompts summarizing recurring student failure modes, and replays these examples until performance improves or they are evicted. Across Qwen3.5 students from 0.8B to 9B and a 31-benchmark vision-language, language, and video suite, it reportedly outperforms several distillation and GRPO baselines, especially for the smallest students, though the abstract gives no numerical gains.

    The teacher-in-the-prompt formulation and replay mechanism are a nontrivial alternative to standard distillation and on-policy RL, with broad claimed evaluation, but the lack of quantitative results makes the size and reliability of the improvement unclear.

  37. maybe AI / ML ▲ 13 score 4.6

    The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation

    Nicolas Dufour, Alexei A. Efros, Patrick Pérez

    This paper measures FID as a distribution over both training seeds and sampling seeds, using several hundred class-conditional ImageNet diffusion/flow models. It finds that retraining randomness affects FID much more than resampling, that larger models and more compute do little to reduce this variation, and that guidance tuning can reduce—but also reorder—the apparent ranking of seeds. It proposes reporting FID error bars across training seeds and treating gaps below roughly 1.3% as inconclusive.

    The large-scale quantification of training-seed effects exposes a potentially important flaw in standard generative-model comparisons, though the results are primarily an evaluation correction rather than a new capability or method.

  38. maybe Tech score 4.6

    Solid-state transcapacitor, a new gain element for logic, memory and interconnects

    Amrita Mathuriya, Roza Kotlyar, Neal Reynolds et al.

    The paper proposes a three-terminal solid-state transcapacitor in which a gate changes a channel’s capacitance through electromechanical coupling, producing gain and inversion without dissipative transport current. The authors argue that TCAP circuits could serve as both logic and 1T-1C-like memory elements, with voltage scaling and energy recovery potentially reducing energy use by up to 100× versus scaled CMOS, though the abstract provides limited evidence for that system-level projection.

    This is a genuinely unconventional post-transistor computing direction with potentially major energy implications, but the abstract does not establish the claimed advantage beyond an early device concept and projected scaling benefits.

  39. maybe Neuroscience picked score 4.6

    Functional segregation of body-brain signals in the area postrema

    Lopez-Cruz, A., Burgos, N. S. F., Hakimi, A. M. et al.

    Using recordings and causal experiments in behaving mice, the study maps distinct area postrema neuron types to different physiological signals: GFRAL neurons promote fat-specific satiation, GIPR neurons respond to sugar and inhibit GFRAL cells, CALCR neurons sense intestinal hyperosmolality, and PRLHR neurons track blood volume and pressure. The most unexpected finding is that GFRAL neurons respond to dietary fat through a pathway independent of GDF15 and canonical gut–brain signaling, despite GFRAL’s usual association with sickness and nausea.

    It provides a non-obvious functional reorganization of area-postrema circuits and connects drug-targeted neurons to natural nutrient sensing, but the evidence is limited to the abstract and mouse physiology rather than a demonstrated translational or broadly computational advance.

  40. maybe Robotics ▲ 10 score 4.6

    Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

    Tongyan Fang, Siyuan Huang, Naiyu Fang et al.

    The paper proposes HABC, an online RL fine-tuning method for vision-language-action policies that turns sparse episode success/failure labels into transition-level training weights. It separates learning whether a trajectory remains viable from learning whether it is efficient, uses a state-dependent gate to balance them, and avoids assigning outcomes to human-intervention segments. On three real-robot bimanual manipulation tasks, it improves success over SFT from 36/44/12% to 92/88/38%.

    The combination of hierarchical viability-versus-efficiency credit assignment and intervention-aware online RL addresses a real bottleneck in sparse-outcome VLA fine-tuning, with large gains on physical robots, but the abstract lacks comparisons to strong online-RL baselines and the third task remains weak.

  41. maybe AI / ML ▲ 17 score 4.6

    Discretizing Reward Models

    Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika et al.

    This paper argues that continuous reward models often distinguish between responses that are equally good, creating oversensitive rewards that can produce poor or reward-hacked policies. It proposes evaluating reward models separately for discriminative ability and specificity, and uses Monte Carlo dropout to cluster scores into discrete rewards without retraining; controlled and natural RL experiments reportedly show improved policies.

    The separation of reward discriminativeness from specificity and the training-free discretization approach offer a non-obvious perspective on a central RLHF failure mode, but the abstract provides no quantitative scale or comparison details sufficient for a stronger verdict.

  42. maybe AI / ML ▲ 43 score 4.6

    MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization

    Guangyi Liu, Pengxiang Zhao, Gao Wu et al.

    MobileForge adapts a vision-language model to mobile apps without human-written tasks, demonstrations, or reward labels. It combines automatic app exploration and curriculum generation with hierarchical feedback from outcomes, intermediate steps, and corrective hints, improving Qwen3-VL-8B to 67.2% Pass@3 on AndroidWorld and a further-trained model to 77.6%, with 41.0% on an out-of-domain MobileWorld split.

    The annotation-free adaptation pipeline and step-level feedback optimization are a meaningful combination with credible benchmark gains, but the abstract does not establish how much comes from the individual components or whether the results generalize beyond GUI benchmarks.

  43. maybe AI / ML ▲ 65 score 4.5

    PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

    Yueyi Sun, Yuhao Wang, Jason Li et al.

    PerceptionDLM adapts multimodal diffusion language models to describe multiple image regions simultaneously, using prompting and structured attention masks so regions and their tokens can be decoded in parallel rather than sequentially. The authors introduce ParaDLC-Bench and report competitive caption quality with substantially faster multi-region inference, although the abstract gives no concrete speed or accuracy figures.

    Parallel region-level perception is a plausible but nontrivial use of diffusion decoding, and could matter for efficient visual understanding, but the abstract does not quantify the claimed speedup or establish how broadly the advantage holds.

  44. maybe AI / ML ▲ 96 score 4.5

    PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems

    Jiayu Liu, Qihan Lin, Cheng Qian et al.

    PlanBench-XL is an interactive benchmark for testing LLM agents that must discover and use tools across long task sequences, with up to 1,665 tools and optional missing, failing, or distracting functions. Across ten models, the best reported model reaches 51.90% accuracy without disruptions but falls to 11.36% under severe blocking, especially when failures are silent or recovery requires long alternative plans.

    This is primarily a benchmark paper, but its retrieval-limited setting and sharp degradation under realistic, weakly signaled tool failures expose a non-obvious weakness in current long-horizon agents that may be useful for future planning research.

  45. maybe AI / ML ▲ 54 score 4.5

    MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction

    Jianing Zhang, Chenhao Zheng, Yajun Yang et al.

    This paper proposes predicting the future 3D trajectories of queried object points from video, conditioned on a natural-language goal. It introduces a 1.16-million-video trajectory corpus, a human-verified benchmark, and a model supporting both autoregressive and flow-matching trajectory generation, with claimed benefits for robot manipulation and video synthesis.

    The combination of language-conditioned, object-level 3D point forecasting and large-scale training data is a meaningful direction with plausible robotics and video-generation value, but the abstract gives no quantitative results and makes broad transfer claims that need verification.

  46. maybe Neuroscience score 4.5

    A hippocampal neuroimaging signature of neurovascular insulin signalling links metabolism to mood

    Cherix, A., Godlewska, B., Lazari, A. et al.

    The study combines human neuroimaging with a mouse model to examine how hippocampal insulin signalling relates to depression and anxiety. Hippocampal GABA and lactate levels, along with hippocampus–default-mode connectivity, tracked glycaemic control and mood symptoms in people; unexpectedly, reducing insulin signalling at the mouse blood–brain barrier increased neuronal metabolism and reduced anxiety-like behaviour, with similar metabolite–behaviour relationships across species.

    The counterintuitive causal result that hippocampal insulin resistance can improve anxiety-like behaviour, alongside a conserved human–mouse neurometabolic signature, is genuinely interesting, but the abstract provides limited detail on sample sizes, controls, and causal specificity.

  47. maybe Robotics ▲ 6 score 4.5

    EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

    Ganlin Yang, Zhangzheng Tu, Yuqiang Yang et al.

    EventVLA is a vision-language-action policy with a learned memory that predicts which future visual moments will matter and stores those keyframes before the evidence disappears. It combines this event-driven memory with persistent initial and recent visual context, and reports a 40% average success-rate gain on 17 simulated long-horizon tasks plus four real-world bimanual tasks.

    The predictive, task-conditioned keyframe memory is a plausible new approach to non-Markovian manipulation and the reported gain is large, but the abstract provides too little detail about baselines, task difficulty, and ablations to justify a strong recommendation.

  48. maybe AI / ML ▲ 25 score 4.5

    BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

    Qizhi Pei, Zhimeng Zhou, Yi Duan et al.

    BioMatrix is a 1.7B/4B decoder-only model that represents molecular and protein sequences, 3D structures, and natural language in one discrete token space. It is trained on 304.4 billion mixed-domain tokens and reportedly matches or beats specialized systems on 77 of 80 understanding and generation tasks, including cross-modal tasks; the main novelty is native generation across all modalities without external encoders or modality-specific heads.

    The unified, genuinely generative treatment of molecules, proteins, structures, and language is a meaningful systems direction, but the abstract gives no task-level numbers or comparisons, so the broad state-of-the-art claim is not yet strong evidence of a major advance.

  49. maybe Robotics picked▲ 2 score 4.4

    Human Universal Grasping

    Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu et al.

    The paper uses egocentric smart-glasses recordings of people picking up objects to learn a distribution of natural human grasps. Its flow-matching model predicts a 3D wrist pose and MANO hand configuration from a single RGB-D view, then retargets the grasp to different robot hands; on a 30-object real-world test set it reportedly improves over prior grasping methods by 23% and 34% across settings. The main contribution is combining a large human-grasp dataset with a generative model aimed at zero-shot, cross-embodiment grasping.

    This is a plausible new route to general-purpose grasping with meaningful real-world and cross-robot evaluation, but the abstract does not give absolute results or enough detail to establish a breakthrough rather than a strong benchmark-and-method advance.

  50. maybe AI / ML score 4.4

    Teaching agentic AI to learn expert reasoning for rare disease diagnosis

    Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler et al.

    The paper introduces Policy Iteration with Human Feedback (PIHF), a clinician-gated, in-context process that turns expert corrections and model failures into a reusable diagnostic policy for an off-the-shelf LLM. On 1,243 rare-disease cases, the system raised top-1 accuracy from 26.5% to 59.3%, with similar gains on diseases and cases withheld during policy development, transferred across model families, and also improved results on 515 real Undiagnosed Diseases Network patients.

    The combination of expert-guided policy learning without model retraining, cross-model transfer, and substantial out-of-distribution gains is genuinely notable, though the abstract alone does not establish how robust or clinically meaningful the evaluation is.

  51. maybe AI / ML picked score 4.4

    Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models

    Martin Jaggi

    The paper shares the same expert parameters across consecutive transformer layers in a mixture-of-experts language model, while keeping routing and attention layer-specific. Across OLMoE, Qwen3, and DeepSeek-style models, this reportedly nearly halves expert memory requirements with little change in perplexity or downstream performance.

    A roughly 2× reduction in MoE parameter memory without an apparent quality cost is a meaningful efficiency result, but the abstract does not provide enough quantitative detail to establish whether the claim holds broadly across scales and training settings.

  52. maybe AI / ML score 4.4

    Beyond the GUI Paradigm: Do Mobile Agents Need the Phone Screen?

    Li Gu, Zihuan Jiang, Linqiang Guo et al.

    This paper argues that mobile agents should use Android’s command-line interface, not only interact through the phone’s graphical interface. Across AndroidWorld and MobileWorld, a coding agent using CLI access outperforms the tested GUI baselines, while oracle CLI programs solve roughly 86–89% of CLI-solvable tasks; a new task suite also shows especially large gains for bulk operations, filtering, aggregation, cross-app workflows, and hidden device state.

    The strong result is a credible, non-obvious reframing of mobile-agent interaction—with substantial reported gains and a useful task suite—but it is still primarily a benchmark comparison rather than a demonstrated broadly general new agent architecture.

  53. maybe AI / ML picked score 4.4

    Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

    Alexander Polok, Samuele Cornell, Sathvik Udupa et al.

    The paper adds speaker-diarization masks to the acoustic encoder of a spoken language model, allowing a frozen decoder to focus on a selected speaker in far-field, multi-talker audio. Its Dixtral system reportedly improves speaker-attributed transcription substantially across four datasets and matches or exceeds larger general-purpose models on long-form multi-speaker question answering, including after fine-tuning.

    The combination of explicit diarization conditioning with a frozen spoken-LLM decoder addresses a real limitation of multi-speaker speech models, and the reported cross-dataset gains are large, but the abstract does not establish how fair or reproducible the comparisons are.

  54. maybe AI / ML score 4.4

    Toward Simultaneously Optimal Regret in U-Calibration

    Rafael Frongillo, Haipeng Luo, Nishant A. Mehta et al.

    The paper gives one online forecasting algorithm that adapts to the difficulty of the downstream loss: it retains near-optimal square-root regret for arbitrary bounded proper losses while achieving logarithmic regret for smooth proper losses, including some non-Lipschitz cases. The key idea is an FTPL method with self-concordant perturbations applied directly in prediction space, together with a new analysis.

    This removes a previously observed incompatibility between universal U-calibration and fast rates, with a technically novel algorithm and a broad theoretical guarantee, though the result concerns a relatively specialized online-learning setting.

  55. maybe AI / ML score 4.4

    Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

    Pengyu Li, Zhitao Gao, Lingling Zhang et al.

    The paper argues that visual-thought tokens in unified multimodal models are expensive because of their generation process, while the rendered visual content itself contributes little. It distills a teacher conditioned on privileged visual-thought traces into a text-only student using on-policy token-level JSD, reportedly improving accuracy by 3.4 percentage points while providing a 14.3× inference speedup across nine benchmarks. The controls suggest the benefit comes from information encoded in the generation pathway rather than merely adding noise or visual context.

    The combination of an unusual finding about visual-thought utility and a potentially large speedup through privileged-trace distillation is worth checking, but the abstract gives insufficient detail to validate the unusually large benchmark comparison and the interpretation of the KL diagnostics.

  56. maybe AI / ML score 4.4

    TokenMem: Faithful Knowledge Injection for Frozen LLMs

    Chengzhang Yu, Chenyang Zheng, Zening Lu et al.

    TokenMem adds a separate cross-attention pathway to inject retrieved knowledge into a frozen LLM, rather than forcing it through the same residual stream as the model’s stored knowledge. A small gating adapter trained in two stages substantially improves compliance with deliberately counterfactual retrieved facts: 69–70% versus 20–52% for vanilla RAG across five models from three families, with ablations suggesting the second training phase is essential.

    The dedicated injection pathway and large, cross-model gains on knowledge-conflict tests are genuinely interesting, but the abstract does not establish whether the method generalizes beyond controlled counterfactual benchmarks or improves ordinary RAG reliability.

  57. maybe AI / ML score 4.4

    Learning When to Denoise: Optimizing Asynchronous Schedules for Latent Diffusion

    Bingshuo Qian, Xiang Cheng

    The paper learns the asynchronous denoising schedule used across multiple image representations in a latent diffusion/flow-matching model, rather than fixing that schedule by hand. On ImageNet 256×256, the learned schedule reaches comparable or better FID with roughly one-quarter the training time, and at longer training outperforms a larger baseline while using a 675M-parameter model; the schedule itself adds less than 1% training compute.

    Learning the cross-representation denoising schedule is a substantive efficiency idea with strong reported ImageNet gains, but the evidence is limited to one benchmark and a specialized diffusion setup rather than a broadly demonstrated advance.

  58. maybe Robotics score 4.4

    EquiVLA: A General Framework for Rotationally Equivariant Vision-Language-Action Models

    Thien-Loc Ha, Quang-Tan Nguyen, Trong-Bao Ho et al.

    EquiVLA adds rotational geometry to vision-language-action models by converting frozen ViT features into approximately SO(2)-equivariant representations and using an exactly equivariant flow-matching action head. On LIBERO, CALVIN, and five Mobile ALOHA tasks, it reports sizable gains over a GR00T N1.5 baseline, including 92.6% vs. 78.1% LIBERO success and 72% vs. 54% real-robot success.

    The combination of equivariant visual processing and action generation for general VLA models is a meaningful direction, and the reported gains span simulation and real robots, but the abstract does not establish how much comes from the equivariance itself or how broadly the approach transfers beyond planar rotations.

  59. maybe AI / ML picked score 4.4

    Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

    Jelena Meyer, David Garcia, Dirk U. Wulff

    The authors test personality and risk-preference questionnaires on 56 instruction-tuned LLMs and compare the results with human reference data. They find that 81–90% of between-model variation is explained by directional response bias rather than the intended traits, and that apparent profiles can change—or be deliberately manufactured—by changing the items. They propose measuring how often an instrument’s trait signal conflicts with this bias (“response orthogonality”) instead of directly importing human psychological tests.

    This is a potentially important, quantitatively supported challenge to treating LLM psychometric profiles as intrinsic properties, but the abstract does not establish how broad or robust the finding is across model families, prompting conditions, and instruments.

  60. maybe Robotics score 4.4

    Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

    Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha et al.

    The paper proposes compressing vision-language-action models by removing redundant layers identified from a single forward pass using Centered Kernel Alignment, without retraining a layer-selection mechanism or loading the full model for optimization. Across three simulation benchmarks and 10 real-world tasks on four robot types, removing up to half the layers reportedly cuts fine-tuning time by 40–50% and inference time by up to 30% while preserving or improving performance.

    A training-free, architecture-level compression method with substantial reported speedups and real-robot validation is worth examining, but the abstract gives no detailed task-by-task results or evidence that the claims generalize beyond the tested VLA families.

  61. maybe Robotics score 4.4

    Robot Self-Improvement via Human-Video Dynamics Models

    Hanzhi Chen, Anran Zhang, Simon Schaefer et al.

    The paper uses human videos to learn embodiment-agnostic action, dynamics, and value representations, then applies them to a robot’s own failed rollouts. Its Dynamics-Guided Action Correction method proposes and ranks corrective actions without additional training, improving success from 40% to 81% across seven real-world manipulation tasks on two robot platforms and multiple policy backbones.

    This is a credible and non-obvious step toward robots improving from their own failures using large-scale human video priors, with substantial real-world gains, but the abstract does not establish how broadly the method generalizes beyond seven tasks or how much depends on the particular setup.

  62. maybe AI / ML score 4.4

    Accelerated and Stable Convergence with Anchored Generalized Optimistic Method

    Motahareh Sohrabi, Jianxin You, Simon Lacoste-Julien et al.

    The paper introduces GOMA, a family of anchored optimistic first-order methods for monotone variational inequalities and min-max optimization. It claims an accelerated deterministic last-iterate rate of O(1/k^2) using the squared gradient norm, and a single-gradient-call stochastic variant achieving O(1/sqrt(k)) under unbounded variance without variance reduction or increasing batch sizes.

    The combination of anchoring and optimistic updates appears to yield unusually strong last-iterate guarantees, especially the claimed first unconstrained stochastic result under unbounded variance, but the abstract provides theory claims without enough detail to establish broad practical impact.

  63. maybe AI / ML score 4.4

    Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers

    Vatsal Ananthula, Adarsh Kumarappan

    The paper trains models to align natural-language reasoning spans with outputs from programmatic verifiers, then tests whether verifier information is actually reflected faithfully in the generated explanations. Across theorem proving, Go, code, and FEVER, verifier-relevant information becomes decodable—and sometimes causally influential—from rationale representations, but models can still produce fluent explanations describing the wrong algorithm or evidence. The main result is that representation-level consistency and decodability are not sufficient guarantees of faithful explanations.

    The clear separation between decodability, causal influence, and faithful generation is a useful and non-obvious challenge to a common interpretability assumption, though the workshop status and abstract-level evidence do not establish a broadly decisive result.

  64. maybe AI / ML score 4.4

    Fixed RAG Compression Collapses Measured Reader Scaling

    Sugam Panthi, Rabab Abdelfattah

    This paper argues that evaluating a fixed RAG compressor with only a few readers can give misleading conclusions: compression may help weak readers while removing details that stronger readers could exploit. Across 20 readers, multiple compression methods, and five benchmarks, it reports declining compression benefit with reader strength, frequent model-ranking reversals, and substantial hiding of upgrades such as Qwen 7B to GPT-4.1-mini; it also releases a toolkit for auditing this effect.

    The cross-reader scaling effect is a non-obvious and potentially important correction to how RAG compression is evaluated, with unusually broad empirical support, but it is primarily an evaluation methodology result rather than a new capability or algorithm.

  65. maybe AI / ML score 4.4

    Spectrally Safe Neural Operator Warm-Starts for Large-Scale Newton Solvers

    Jaemin Oh, Youngkyu Lee, Jerome Darbon et al.

    The paper shows that a neural operator can have low average prediction error yet produce physically invalid local states that make the Newton Jacobian indefinite, undermining Krylov-based PDE solves. It introduces a short label-free fine-tuning stage that penalizes discrete energy, restores a positive-definite spectrum, and yields up to 5.4× speedup over continuation on a 6.4-million-DOF 3D hyperelasticity problem.

    The important contribution is identifying spectral definiteness—not prediction error—as a practical failure mode for neural-operator warm starts and addressing it without additional solution labels, though the evidence appears concentrated on one PDE application rather than broad validation.

  66. maybe AI / ML score 4.4

    One-Bit Clustering for Two Component Sub-Gaussian Mixture Models

    Junren Chen, Yun Yang

    The paper studies clustering two-component sub-Gaussian mixtures when each data entry is reduced to a single dithered bit. A Lloyd-style method retains an exponentially decaying error rate comparable to unquantized data under a non-spikiness condition, and exact recovery needs only a logarithmic-factor larger separation; random rotations can make the condition hold in high dimensions. Matching minimax lower bounds and numerical experiments support the claims.

    This is a genuinely new and potentially useful result showing that extreme one-bit quantization can preserve clustering performance nearly at the unquantized statistical limit, but the contribution is a specialized theoretical advance rather than a demonstrated broad practical capability.

  67. maybe Robotics score 4.4

    Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

    Yangtao Chen, Zixuan Chen, Peiyang Wang et al.

    Wh0 uses a generative video world model to create 50,000 egocentric human-hand manipulation episodes, then reconstructs hand motion and edits the visuals into supervision suitable for robot training. Combined with a small amount of real robot data, this raises zero-shot success on 18 real-world dexterous manipulation tasks from 8.3% to 38.9% relative to robot-only post-training, suggesting generated human manipulation data can help bridge both scale and embodiment mismatch.

    The combination of controllable generative video, hand-motion reconstruction, and robot adaptation is a genuinely interesting data-scaling direction with a large reported gain, but the abstract does not establish how robust the results are or whether generated-video artifacts and reconstruction quality limit generalization.

  68. maybe Robotics score 4.4

    Do Rigid-Body Simulators Dream of Soft Robots? Learning Contact-Rich Manipulation for Tendon-Driven Continuum Robots

    Chengnan Shentu, Nicholas Baldassini, Tongjia Zheng et al.

    The paper embeds tendon-driven continuum robots into MuJoCo using a continuum-mechanics-informed discretization, allowing tendon actuation, contacts, and dynamics to be learned in one simulation pipeline. The simulator agrees with Cosserat-rod simulations and real hardware, and imitation policies trained in simulation transfer zero-shot to a physical three-segment robot for two contact-rich manipulation tasks.

    A unified, learning-oriented simulator enabling what the authors claim is the first zero-shot sim-to-real contact-rich manipulation with continuum robots is a substantial robotics advance, but the evidence is limited to two tasks and one robot configuration.

  69. maybe Neuroscience picked score 4.4

    Visual consequences of saccades explain early cortical response dynamics during natural vision

    Schweitzer, R., Dimigen, O., Huber-Huber, C.

    The study argues that the lambda response after a saccade is driven primarily by the retinal image motion caused by the eye movement, rather than by new visual input at the fixation point. EEG from thousands of natural-viewing saccades matched responses to replayed retinal shifts during fixation, and a model using only eye trajectories, natural-scene statistics, and visual sensitivity reproduced the response dynamics.

    It offers a mechanistic reinterpretation of a major natural-vision EEG response, supported by matched retinal-motion experiments and a parsimonious model, but the abstract alone does not establish how broadly this overturns existing accounts.

  70. maybe Neuroscience score 4.4

    A Structural Principle for Macroscopic Neural Dynamics Correlations

    Wu, Q., Wen, Q., Liu, C.

    The paper proposes that the similarity between regions’ incoming connectivity profiles (“coupling correlation”) is a primary determinant of correlated large-scale neural activity. Using dynamical mean-field theory and random-network simulations, it derives an approximately linear structure–function relationship and argues that long-tailed structural spectra are needed to preserve substantial, size-invariant correlations; the prediction is tested across human, mouse, and fruit-fly datasets.

    This is a potentially useful mechanistic principle linking connectome structure to functional dynamics across species, but the abstract lacks quantitative validation details and the strength of the empirical causal claim is unclear.

  71. maybe BCI picked score 4.4

    Adaptive Charge Modulation Enables Focal, Selective Spinal Cord Stimulation

    Vatsyayan, R., Khoury, F., Porter, T. S. et al.

    The authors introduce Adaptive Charge Modulation, a charge-balanced multipolar stimulation pattern intended to activate deep spinal circuits while suppressing unwanted surface and dorsal-root activation. In rats, epidural ACM reportedly selected a single muscle out of 14 monitored muscles, with high-resolution brain–spine recordings supporting focal recruitment and stable operation for 68 days. The main novelty is achieving relatively selective deep spinal activation using surface electrodes rather than implanted intramedullary contacts.

    This is a potentially important new neuromodulation strategy with unusually strong preclinical selectivity and chronic-recording evidence, but the claims remain limited to rats and the abstract does not establish comparative advantages, mechanism, or translation to humans.

  72. maybe Neuroscience score 4.4

    Focal astrocyte Kir4.1 loss drives seizures, spreading depolarizations and postictal impairments

    Codadu, N. K., Gao, Y., Tyurikova, O. et al.

    The study selectively removed the astrocytic potassium channel Kir4.1 from the adult mouse hippocampus and found impaired potassium buffering, spontaneous recurrent seizures, and increased seizure-associated spreading depolarizations. Combining chronic DC-coupled telemetry with graphene transistor recordings, it showed that seizures accompanied by spreading depolarizations were longer, stronger, and followed by more severe behavioral and electrographic suppression than seizures without them. The work also identifies conventional AC-recording signatures that may allow retrospective detection of these events.

    This provides a fairly direct causal link between focal astrocytic potassium-buffering failure, spontaneous epilepsy, and postictal spreading depolarizations, supported by unusually capable chronic DC recordings, but it remains a preclinical mouse study.

  73. maybe AI / ML ▲ 29 score 4.4

    OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

    Guibin Zhang, Xun Xu, Yanwei Yue et al.

    OPD-Evolver trains an agent to manage its own memory rather than merely retrieve stored experiences. It uses a fast interaction loop over a hierarchical memory system and a slower on-policy distillation loop to teach a deployable policy when to read, use, write, and maintain memories; the abstract reports gains over several prior memory-agent approaches and competitive results from a 9B model against much larger systems.

    The holistic, self-distilled training of memory-management behavior is a meaningful step beyond retrieval-based memory agents, but the abstract gives too little detail about benchmarks, baselines, and evaluation breadth to justify a stronger recommendation.

  74. maybe Robotics ▲ 13 score 4.3

    PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

    Youngjoon Jeong, Jihwan Yu, Minsoo Jo et al.

    PoLAR structures latent actions so that distance from the origin represents how much a scene changes, while angular direction represents the type of change. It uses temporal separation as a weak extent signal and implements the representation in hyperbolic space, reporting better downstream policy learning than latent-action baselines and pretrained vision-language-action models in simulation and real-robot experiments.

    The explicit factorization of action extent and mode through hyperbolic geometry is a meaningful design idea with real-robot validation, but the abstract gives no quantitative gains or evidence that the geometric interpretation is genuinely learned rather than merely useful.

  75. maybe AI / ML ▲ 4 score 4.3

    OpenBioRQ: Unsolved Biomedical Research Questions for Agents

    Minbyul Jeong

    OpenBioRQ is a benchmark of 12,553 genuinely unresolved biomedical research questions designed to test whether agents retrieve and verify evidence rather than merely reproduce known answers and citations. It reports that frontier agents solve only about 29–60% of the hardest questions, and that some agents largely stop using tools on these questions, making tool access provide little benefit. A structured judging checklist also substantially improves agreement between evaluators.

    The combination of open-ended biomedical questions, evidence-grounded agent interaction, and the finding that tool use collapses precisely on the hardest problems is more informative than a standard benchmark, though it remains primarily an evaluation resource rather than a capability advance.

  76. maybe Robotics ▲ 2 score 4.3

    EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video

    Hyunjin Kim, Ri-Zhao Qiu, Guangqi Jiang et al.

    EgoPhys infers a deformable object’s physical properties from egocentric RGB video, distilling expensive per-object inverse-physics solutions into a reusable codebook that predicts stiffness fields for new objects without test-time optimization. The authors report better reconstruction, prediction, and zero-shot generalization, and show an xArm6 using a digital twin initialized from one human-interaction video for deformable-object planning.

    The combination of egocentric observation, amortized inverse physics for deformable objects, and a real-robot planning demonstration is a substantive direction, but the abstract gives no quantitative results or details sufficient to establish a major advance.

  77. maybe AI / ML ▲ 21 score 4.3

    FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows

    Daniel Gilo, Sven Elflein, Ido Sobol et al.

    FlowBender trains conditional flow models to use their own constraint error during generation. It performs a look-ahead prediction, measures the mismatch through a task-specific forward operator, and feeds that feedback into a correction step, with gradient-based and zero-order variants plus a cheaper shortcut. The abstract claims better condition fidelity without sacrificing visual plausibility across image translation, restoration, and 3D texturing, but gives no quantitative results.

    The closed-loop training formulation and support for non-differentiable operators are a meaningful departure from static conditioning and hand-designed guidance, but the broad superiority claims lack numbers and detailed evidence in the abstract.

  78. maybe AI / ML ▲ 12 score 4.3

    Sumi: Open Uniform Diffusion Language Model from Scratch

    Mengyu Ye, Keito Kudo, Wataru Ikeda et al.

    Sumi is an openly released 7B uniform diffusion language model pretrained from scratch on 1.5T tokens, with weights, checkpoints, and a reproducible training recipe. It reportedly matches comparable autoregressive models on knowledge, reasoning, and coding tasks, but trails on commonsense benchmarks, offering a large-scale reference for studying diffusion-language-model scaling and generation behavior.

    Pretraining a uniform diffusion LM at this scale fills an important missing reference point, but the abstract gives no detailed results showing a capability, efficiency, or controllability advantage over autoregressive models.

  79. maybe AI / ML score 4.3

    On the Geometry of Separation in Finite Gaussian Mixtures

    Huy Nguyen, Dung Le, Alessandro Rinaldo et al.

    This paper develops a geometric theory connecting Hellinger distance between finite Gaussian-mixture distributions to Wasserstein distance between their mixing measures, with explicit dependence on component separation and weights. It shows that estimation difficulty depends on the spatial arrangement of components, and that over-specifying the number of components changes the geometry enough to remove minimum-weight dependence from the rates.

    The configuration-dependent separation laws and the claimed change from first- to second-order geometry in over-specified mixtures are substantive theoretical insights, but the abstract provides no concrete rates or empirical/theorem details to justify a stronger verdict.

  80. maybe AI / ML score 4.3

    Embedded Arena: Iterative Optimization via Hardware Feedback

    Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu et al.

    The paper presents an LLM agent that repeatedly edits models and firmware, compiles and flashes them onto real microcontrollers, and uses measured hardware behavior to guide further optimization. The system reportedly reaches deployment in three iterations, exceeds human-designed results within seven, and achieves extreme compression on vision and audio workloads, including elk detection and a battery-free phonetic-transcription wearable.

    Closed-loop LLM optimization using real MCU feedback is a genuinely interesting direction with unusually concrete deployment results, but the abstract provides too little detail to validate the very strong compression, expert-comparison, and generality claims.

  81. maybe AI / ML score 4.3

    The Faithfulness Gap: Certifying Semantic Equivalence Between Natural-Language and Formal Mathematical Statements

    Noor Islam S. Mohammad, Tamim Sheikh

    The paper proposes checking whether an autoformalized theorem preserves the meaning of its natural-language source by comparing forward and backward consequences under targeted probes. Its system combines contrastive counterfactual probes, adaptive probe allocation, a graded faithfulness score, and faithfulness-guided decoding; on a 2,183-pair Lean benchmark it reportedly detects 89.6% of deliberate semantic drift at 3.0% false positives and reduces drift during generation by 47%.

    Faithfulness certification addresses a real weakness of autoformalization and the bidirectional consequence-neighborhood framing is potentially useful, but the abstract provides limited detail about probe validity, benchmark construction, and comparison against stronger semantic-checking baselines.

  82. maybe AI / ML score 4.3

    User as Code: Executable Memory for Personalized Agents

    Bojie Li

    User as Code represents a user’s long-term memory as typed Python state plus executable rules, rather than only retrieved text or facts. An append-only event log is periodically compiled into this code, allowing the agent to aggregate history and trigger rule-based alerts when state changes. It reportedly matches prior systems on recall (78.8% on LOCOMO), but reaches 99% on aggregate-history questions where retrieval systems score 6–43%.

    The executable-memory framing enables genuinely different capabilities—reliable aggregation and unsolicited contradiction alerts—but the abstract provides limited detail on implementation, evaluation breadth, and robustness beyond a few benchmark results.

  83. maybe AI / ML score 4.3

    How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel

    Chang Liu, Chaoyang Ning, Dayi Jiang et al.

    OneModel replaces a multi-component business-agent pipeline with a single model trained to absorb business rules and standard operating procedures through continued pretraining and logic-focused supervised fine-tuning. In an online financial-services A/B test, it reportedly reduced end-to-end latency from 18.7 to 8.0 seconds while raising intelligent resolution rate from 64.3% to 83.3%.

    The substantial simultaneous latency and resolution gains from internalizing workflow logic are notable, but the abstract does not establish how the comparison was controlled, how much engineering or model scaling was involved, or whether the approach generalizes beyond one financial-service deployment.

  84. maybe AI / ML score 4.3

    TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization

    Weiliang Chen, Yuanhui Huang, Xuebo Wang et al.

    TivTok factorizes video representations into time-invariant tokens shared across frames and time-variant tokens carrying frame-specific changes. Its attention design and decoder reuse the shared tokens across frames and chunks, reportedly achieving 2.91× better compression efficiency on 128-frame videos and using only 1.1% as many tokens as downsample-based tokenizers in the evaluation.

    The explicit reuse of persistent visual content across time is a meaningful architectural direction for scalable long-video modeling, but the abstract lacks enough baseline, quality, and experimental detail to establish how substantial the headline compression gains are.

  85. maybe Robotics score 4.3

    MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

    Xingyuming Liu, Ruichun Ma, Heyu Guo et al.

    MuseVLA lets a robot choose sensing modalities such as temperature, audio, or radar on demand, then converts the selected measurement into a common grounded image representation for a VLA policy. On real-robot dexterous manipulation tasks—including temperature-guided placement, audio-based search, and radar-assisted retrieval—it reports an average 80.6% success rate and zero-shot performance on unseen tasks, using synthetic sensor augmentations to reduce the need for multisensory demonstrations.

    The adaptive sensor-selection interface and unified representation address a real limitation of RGB-only VLAs and are demonstrated across several physical sensing modalities, but the abstract lacks detailed task scales, baseline gaps, and ablations needed to establish a major advance.

  86. maybe Robotics score 4.3

    Ghost Attractor Networks: Basin-Structured Dynamical Decoders for Closed-Loop Sequential Generation

    Tianyu Wang, Ying Wang, Zhihao Liu et al.

    The paper introduces Ghost Attractor Networks, small dynamical action decoders whose learned potential creates stable latent basins for multimodal, phase-conditioned, closed-loop generation while using constant memory and single-pass mode switching. In robotic behavioral cloning, a 2.3M-parameter model reportedly matches a 1.07B-parameter diffusion Transformer offline at 462× fewer parameters and 32× lower latency, while improving LIBERO-10 success through basin-based phase conditioning and persistent latent state.

    The basin-attractor formulation for efficient closed-loop action decoding is a genuinely interesting direction with unusually large claimed efficiency gains, but the evidence is mainly one robotic benchmark and the abstract does not establish how broadly the result generalizes.

  87. maybe AI / ML score 4.3

    Small Initialization Matters for Large Language Models

    Liangkai Hang, Junjie Yao, Zhiyu Li et al.

    The paper studies how the scale of parameter initialization affects language-model pretraining, finding that smaller initialization improves training and especially reasoning-oriented performance across model sizes. It argues that small initialization produces an early low-complexity, compressed phase followed by richer representations, and proposes making initialization scale an explicit training knob.

    A potentially high-leverage and surprisingly simple training intervention, supported by claimed cross-scale experiments and a mechanistic account, but the abstract gives no quantitative gains or detail sufficient to establish that the effect is robust and broadly important.

  88. maybe AI / ML score 4.3

    Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction

    Marc-Andre Schulz, Kerstin Ritter

    This paper argues that noisy biomedical measurements preferentially erase higher-order and nonlinear signal, so flexible models can lose their advantage even when the underlying biology is strongly nonlinear. It derives an excess-risk identity linking interaction order to feature reliability and reports that patterns across 140 UK Biobank prediction tasks match this prediction; increasing sample size or model flexibility cannot recover information removed by measurement noise.

    The potentially valuable contribution is a precise, testable explanation for the recurring parity between linear and nonlinear biomedical models, though the core measurement-error effects are classical and the abstract gives limited quantitative detail about the empirical validation.

  89. maybe AI / ML score 4.3

    Beyond Prediction: Tail-Aware Scheduling for LLM Inference

    Yueying Li, Yuanfan Chen, Jiayang Chen et al.

    This paper argues that minimizing average latency with predicted output lengths does not reliably minimize the tail latency users experience, especially under workload shifts, bursty arrivals, and GPU memory pressure. It proposes prediction-free, distribution-aware priority scheduling combined with cache-aware preemption, reporting 35–50% lower P99 total latency than SRPT even with perfect length information and 34–47% lower TTFT across production and open-source traces.

    The notable claim is that prediction-free, tail-aware scheduling can substantially outperform idealized length-based SRPT on P99 latency, but the abstract provides no absolute latency values or enough experimental detail to justify a strong recommendation.

  90. maybe AI / ML score 4.3

    Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs

    Jaeyeon Lee, Shunjie Wen, Dong-Wan Choi

    The paper proposes SPARE, a training-free visual-token pruning method for vision-language models that selects tokens to minimize reconstruction error rather than simply maximizing angular diversity. It also reports that some low image-text-relevance tokens preserve useful context, and uses this anti-relevance signal during selection; on LLaVA, it removes up to 94% of visual tokens while retaining 95% of baseline performance.

    The reconstruction-based formulation and counterintuitive use of low-relevance tokens are meaningful ideas for efficient VLM inference, but the abstract provides no detailed comparisons, compute measurements, or broad quantitative evidence sufficient for a strong verdict.

  91. maybe Robotics score 4.3

    Scaling Learning-based AEB with Massive Unlabeled Data

    Xiangyu Wang, Yang Zhan, Mengxiang Hao et al.

    The paper presents a semi-supervised framework for training automatic emergency-braking models from up to 1 billion unlabeled driving windows, using a small labeled anchor set to correct pseudo-labeling errors. It adds noise-aware anchor selection and kinematics-based gating, and reports deployment in hundreds of thousands of vehicles over 1 billion km, with a claimed >100:1 positive-to-false-activation ratio and 35% more accident-free mileage than a rule-based baseline.

    Large-scale production validation of learning-based AEB and the claimed billion-window/billion-kilometer results are genuinely notable, but the core method is an incremental semi-supervised-learning refinement and the abstract leaves the safety metrics and baseline comparison insufficiently defined.

  92. maybe AI / ML score 4.3

    A Solver-Free Training Method for Predict-then-Optimize

    Beichen Wan, Mo Liu

    The paper proposes training predictive models for downstream linear or combinatorial optimization without calling an optimization solver during training. It uses a measure-transformation principle to construct a surrogate loss, with Fisher-consistency and excess-risk guarantees; experiments reportedly match decision-focused baselines while reducing training time by orders of magnitude.

    Solver-free training would remove a major scalability bottleneck in predict-then-optimize, but the abstract gives no quantitative results or details about the assumptions and range of optimization problems, so the strength of the claimed speedup is not yet clear.

  93. maybe Robotics score 4.3

    VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents

    Marcus Hoerger, Rishikesh Joshi, Rahul Shome et al.

    VOiLA learns transition and observation models for partially observed robotic tasks using conditional diffusion models, then distills the samplers into fast generators for GPU-parallelized online POMDP planning. The authors report nearly 1,000-fold lower sampling cost, comparable or better performance than Recurrent Soft Actor-Critic with under 10% of its training data, stronger generalization to unseen configurations, and successful sim-to-real deployment in 10/10 physical trials.

    The combination of learned generative POMDP models, aggressive sampler distillation, and data-efficient sim-to-real planning is promising, but the abstract provides too few details about task difficulty, baselines, and the robustness of the 10/10 robot result to justify a strong recommendation.

  94. maybe Robotics score 4.3

    Start Right, Arrive Right: Asynchronous Execution via Initial Noise Selection

    Trong-Bao Ho, Quang-Tan Nguyen, Thien-Loc Ha et al.

    PAINT addresses a practical problem in flow-based robot policies: while one action chunk is executing, the next chunk may be generated too late or become inconsistent at the boundary. Instead of modifying the policy or steering the generation trajectory, it selects an initial noise vector using backward-Euler inversion and a repainting rule so the ordinary flow ODE generates a compatible continuation. The training-free method is evaluated on 12 simulation benchmarks and 6 real manipulation tasks across single-arm, bimanual, and humanoid robots.

    The reframing of asynchronous action-chunk generation as initial-noise selection is a genuinely interesting and potentially general idea, with unusually broad real-robot evaluation, but the abstract gives no quantitative gains or comparisons strong enough for a top-tier recommendation.

  95. maybe AI / ML score 4.3

    Code evolution for link prediction in complex networks

    Alexey Vlaskin, Eduardo G. Altmann

    The paper uses automated code evolution—apparently combining large language models with genetic search—to discover link-prediction algorithms rather than manually designing them. Across 580 networks, the evolved methods reportedly reach average AUC 0.915 versus 0.783 for human-designed methods, while also being efficient enough for networks with millions of links; their main novelty is how they select and combine existing node and link features.

    The broad, large-margin improvement and evidence of machine-discovered algorithmic structure are genuinely interesting, but the abstract does not establish how leakage, dataset overlap, baselines, or generalization to unseen network types were controlled.

  96. maybe Robotics score 4.3

    Motor Angular Speed Preintegration for Multirotor UAV State Estimation

    Matěj Petrlík, Filip Novák, Robert Pěnička et al.

    The paper uses motor angular speeds to estimate a multirotor’s translational acceleration and preintegrates them for state propagation, avoiding reliance on vibration-corrupted IMU measurements. Combined with LiDAR in the MAS-LO factor-graph system, this reportedly improves position accuracy by 28%, velocity accuracy by 65%, and reduces lag by 14% versus LIO-SAM, while remaining robust to parameter errors.

    Using propulsion measurements as a substitute for high-rate inertial sensing is a genuinely nonstandard and potentially useful estimation direction, but the abstract does not establish how broadly the gains generalize beyond the evaluated multirotor setup.

  97. maybe AI / ML score 4.3

    Stochastic Linear Contextual Bandits with Bounded Noise: A Set-Membership Approach

    Haonan Xu, Yingying Li

    The paper develops SME-OFU, a linear contextual bandit algorithm that uses set-membership estimation to exploit bounded reward noise rather than treating it merely as sub-Gaussian. Under its assumptions, it claims logarithmic regret in the horizon, versus the usual square-root dependence, and reports better simulations than a sub-Gaussian-noise baseline. The main novelty is turning hard reward bounds into tighter parameter uncertainty sets.

    The claimed transition from square-root to logarithmic regret is potentially important and non-obvious, but the abstract gives no details about the additional gap, realizability, or noise assumptions that likely make this possible, and empirical support is limited to simulations.

  98. maybe AI / ML score 4.3

    SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation

    Shilong Xiang, Zirui Zhang, Lijun Yu et al.

    The paper proposes Spatially Speculative Decoding for autoregressive image generators: instead of predicting only the next token in a flattened sequence, the model also predicts spatially adjacent tokens to exploit 2D locality. It reports up to 13.3× faster generation while preserving fidelity on DPG-Bench and GenEval, though the abstract does not clarify how broadly this speedup holds or what additional model/training costs are required.

    The potentially large inference-speed gain from changing autoregressive decoding to respect image geometry is genuinely interesting, but the abstract gives only an 'up to' result on two benchmarks and insufficient detail to establish that the improvement is robust.

  99. maybe Robotics score 4.3

    Generating Robot Hands from Human Demonstrations

    Sha Yi, Nicklas Hansen, Xueqian Bai et al.

    The paper uses over 4 million frames of human fingertip motion to optimize robot-hand morphology, rather than treating the hand design as fixed and learning a controller for it. It generates tree-structured hands controlled by inverse kinematics, including a 6-DoF general-purpose hand and simpler task-specific hands with mimic joints, then fabricates them as one-piece print-in-place mechanisms. The resulting hands reportedly track teleoperated fingertip motions accurately, while an RL-based design proposer reduces optimization time from hours to minutes.

    Generating physical robot morphology directly from large-scale human motion data is a genuinely interesting direction with real hardware validation, but the abstract gives limited quantitative evidence for the claimed superiority and generality.

  100. maybe Robotics score 4.3

    OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving

    Juntong Peng, Juanwu Lu, Yupeng Zhou et al.

    OmniV2X is a generative end-to-end planner for cooperative driving that consumes multimodal observations from vehicles and other agents through standardized V2X tokens. It pretrains on large single-vehicle planning datasets, then adapts to cooperative driving with under 10% of the cooperative fine-tuning data and under 1% of the communication bandwidth used by prior methods, while reportedly improving performance on DAIR-V2X-Seq.

    The combination of single-agent foundation-model pretraining, lightweight standard-compliant V2X conditioning, and very large claimed reductions in cooperative data and bandwidth is notable, but the abstract provides no numerical comparisons and evidence is limited to one benchmark.

  101. maybe Robotics score 4.3

    FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction

    Guangcheng Chen, Lihuang Fang, Huaqi Tao et al.

    FLM-Occ treats indoor occupancy prediction as estimating a probability distribution over occupied voxels rather than independently classifying voxels. It trains a feed-forward mixture model whose primitives are relocated over long distances, reducing spurious geometry; on Occ-ScanNet it reportedly reaches higher accuracy with 32 superquadrics—2.7% as many as the prior method—and runs 3.7× faster.

    The global likelihood-based formulation and very large reduction in primitives are potentially important, but the abstract provides evidence from only one benchmark and does not quantify the accuracy tradeoff or generalization.

  102. maybe AI / ML score 4.3

    Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability

    Mathieu Cyrille Simon, Pascal Frossard, Christophe De Vleeschouwer

    The paper proposes defining disentangled factors as latent variables whose effects on observations are locally orthogonal, expressed through the Jacobian of the generative model. It claims this geometric condition makes general nonlinear representations identifiable without independence or causal assumptions, assuming the data contain all combinations of factor values, and supports the result with theory and normalizing-flow experiments.

    The claimed identifiability result would substantially challenge the standard impossibility view of unsupervised disentanglement, but the strong full-combinatorial-support assumption and lack of quantitative experimental detail make it worth checking rather than an automatic strong recommendation.

  103. maybe AI / ML score 4.3

    Breaking Chains with Trees: Model-Parallel Deep Learning with $\mathcal{O}(\log N)$ Time Complexity

    Neeraj Mohan Sushma, Aditya Nagarsekar, Cabrel Teguemne Fokam et al.

    TreeProp replaces the usual layer-by-layer forward pass and backpropagation with a tree-structured variational learning procedure, aiming for logarithmic parallel depth in the number of layers. The authors report performance comparable to ordinary end-to-end training on vision and autoregressive language modeling, better results than prior contrastive methods, and an extension to recurrent networks without backpropagation through time. It also implicitly trains subnetworks with different effective depths through multiple tree paths.

    The proposed logarithmic-depth alternative to both forward computation and gradient propagation is potentially important and non-obvious, but the abstract gives no quantitative speedups and may conceal substantially greater computation or communication, so the scalability claim needs close inspection.

  104. maybe AI / ML score 4.3

    Learning the ARTS of Search for Automated Discovery

    Gurusha Juneja, Arnav Kumar Jain, Deepak Nathani et al.

    The paper proposes ARTS, an LLM-guided tree search method that separates the quality of a scientific hypothesis from the quality of its current implementation. The agent reviews execution histories, diagnoses whether failures reflect bad code or bad ideas, and uses test-time training to compress the search tree into its weights; across 22 ML research and engineering tasks, a 4B Qwen3 agent reportedly matches frontier-model performance at up to 5× lower inference cost and can recover a recurrent-memory RL solution that heuristic search prunes.

    The combination of failure diagnosis, hypothesis-preserving search, and test-time training for long search histories is a substantive direction with an unusually strong small-model cost claim, but the abstract provides limited detail for judging the breadth and robustness of the comparisons.

  105. maybe AI / ML score 4.3

    Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

    Xuyang Wang, Zhenyu Li, Jian Ding et al.

    Artic-O reconstructs an articulated object’s full 3D shape, movable parts, and joint parameters from sparse multi-state images in one feed-forward model. It performs this jointly in a pretrained latent geometry space, using a shape prior and image-grounded part reasoning, and reports similar or better articulation accuracy than LARM while cutting inference time from about 9 minutes to 0.3 seconds per object on PartNet-Mobility.

    The combination of joint latent-space geometry and articulation prediction is a meaningful design, and the claimed roughly 1,800× speedup is notable, but evidence is limited to one benchmark and the abstract gives no exact comparative results.

  106. maybe AI / ML score 4.3

    Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

    Li Kong, Qi Qi, Yinyu Ye et al.

    The paper models LLM serving as a two-dimensional geometric scheduling problem, accounting for the time evolution and memory volume of each request’s KV cache. It introduces Smallest Volume First (SVF) and a one-bit-information variant, claiming to improve the worst-case competitive ratio from 48 toward 3 at high concurrency, with reductions in average and tail latency when integrated into vLLM.

    The geometric formulation and claimed large improvement in online-scheduling guarantees are genuinely interesting, but the abstract gives few concrete experimental numbers and the strongest theoretical claim needs careful verification.

  107. maybe AI / ML score 4.3

    Escaping the Variance Trap: Jacobian-Free Dynamics for Root-Finding Bilevel Optimization

    Zhiyu Li, Xi Xuan, Davide Carbone

    The paper argues that converting stochastic root-finding problems into squared-residual minimization creates a “variance trap”: implicit-Jacobian terms can amplify noise and destabilize bilevel optimization. It formulates root-finding bilevel optimization separately and uses two-time-scale stochastic approximation to update from root errors without explicit Jacobians, providing Markovian-noise convergence guarantees and reporting gains across contrastive learning, ODE control, reinforcement learning, and generative modeling.

    The root-finding-versus-minimization framing and Jacobian-free TTSA approach could be broadly useful, but the abstract gives limited detail about assumptions, baselines, and whether the large cross-domain gains are robust enough to warrant a strong recommendation.

  108. maybe AI / ML score 4.3

    SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion

    Ruoyu Feng, Jinming Liu, Yuqi Wang et al.

    SeFi-Image proposes a semantic-first latent diffusion approach for text-to-image generation and studies it at 1B, 2B, and 5B parameters. The authors claim the 5B model reaches performance comparable to or better than larger baselines while using only about 10–20% as much training compute, and also release few-step distilled variants for faster inference.

    The potentially large training-efficiency gain and scaling study make the method worth checking, but the abstract gives no benchmark numbers or independent evidence to substantiate its comparisons and compute accounting.

  109. maybe Neuroscience score 4.3

    Parvalbumin-expressing interneurons improve sensory discrimination by shaping noise geometry in primate V1

    Cole, S., Hildebrand, D. G. C., Nurminen, L.

    Using optogenetic activation of parvalbumin interneurons alongside dense recordings in awake marmoset V1, the authors tested how inhibition affects sensory coding. PV-cell stimulation improved population discriminability not by increasing stimulus-related responses, but by reducing shared trial-to-trial noise and rotating that noise away from the stimulus-coding direction. The work provides causal primate evidence that inhibition can improve perception by reshaping population-noise geometry.

    This is a non-obvious, causally tested mechanism for how cortical inhibition improves coding in primate cortex, but the abstract gives no quantitative effect sizes or evidence of behavioral improvement, so it falls short of a strong recommendation.

  110. maybe Neuroscience score 4.3

    FMR1 gene therapy restores activity-driven inhibition and prevents audiogenic seizures in Fmr1-/y mice

    Maio, B., Singh, A., Hector, R. et al.

    In an FMR1-deficient mouse model of Fragile X syndrome, AAV delivery of FMR1 restored sound-evoked molecular and inhibitory responses in the inferior colliculus and prevented audiogenic seizures. Rescue occurred even after adult treatment, and targeting the inferior colliculus alone was sufficient, implicating impaired activity-dependent Npas4-related translation as a reversible cause of circuit hyperexcitability.

    The combination of adult seizure rescue, localization to the inferior colliculus, and a mechanistic link from FMRP-dependent translation to recruited inhibition is substantially more informative than a generic mouse gene-therapy result, but the evidence remains limited to a preclinical model and the abstract gives few quantitative or generalization details.

  111. maybe Neuroscience score 4.3

    Stimulus identity rather than emotion drives EEG classification on the FACED dataset

    Gerster, M., Sirotina, E., Orlovskii, A. et al.

    The paper argues that high EEG emotion-classification performance on the FACED benchmark mainly reflects which video was shown, rather than the participant’s experienced emotion. Using both a linear classifier and a deep model, the authors show that performance is similar whether subjects reported feeling the assigned emotion or not, worsens with self-reported labels, and can improve when fewer videos are used—evidence of stimulus-identity and temporal-split confounds.

    This is a potentially important benchmark failure analysis with an assumption-challenging result for EEG emotion decoding, but the abstract provides limited quantitative detail and the contribution is primarily diagnostic rather than a new decoding capability or method.

  112. maybe Neuroscience score 4.3

    Reactivating a schizophrenia risk gene-enriched prefrontal ensemble suppresses decision noise

    Iino, Y., Narita, H., Shimizu, C. et al.

    The study identifies a prefrontal cortical neuron ensemble enriched for schizophrenia risk genes that helps stabilize value-guided decisions in mice. NMDA-receptor hypofunction or chemogenetic inhibition disrupted this ensemble and increased decision noise, while clozapine or a rationally designed multi-receptor antagonist cocktail restored ensemble activity and behavior; the cocktail did so without clozapine-like increases in NREM sleep or delta power. This links genetic risk to a causal circuit computation and suggests a route toward less-sedating pharmacological interventions.

    The causal connection between a schizophrenia-risk-enriched ensemble, decision noise, and a potentially less-sedating multi-receptor treatment is a substantial and non-obvious mechanistic result, but the abstract provides no effect sizes and appears limited to preclinical evidence.

  113. maybe Neuroscience score 4.3

    Distinct cortical patches for syntactic and semantic composition in the human brain

    Dighiero-Becht, T., Friedmann, N., Rizzi, L. et al.

    Using 7T fMRI, the authors tested whether syntax and meaning are processed in separate parts of the language network. Across 20 participants, they found distinct voxel populations: one responding to syntactic structure even in semantically impoverished phrases, and another supporting semantic composition, though the exact anatomy varied between people. The result suggests a partially intermixed dual-network organization rather than a single undifferentiated language system.

    This addresses a live question with a potentially important functional dissociation and unusually high-resolution, subject-specific evidence, but the sample is modest and fMRI-based localization makes the strong computational interpretation worth verifying.

  114. maybe Neuroscience score 4.3

    Human striatal population state dynamics

    Korponay, C., stein, e. a., Ross, T. J. et al.

    The authors use a neurobiologically informed analysis of more than 3 billion voxel-frame-wise fMRI coactivation profiles to identify human striatal population states analogous to canonical SPN down-like and up-like states. They report that these states and rare, high-magnitude corticostriatal bursts reorganize with task demands, arousal, reaction time, reward response, and engagement, suggesting a human systems-level counterpart to state dynamics observed in animal electrophysiology.

    The potentially important contribution is linking animal-model SPN state dynamics to human fMRI at population scale, with state transitions and bursts carrying behavioral information, though the abstract provides limited validation beyond modeling observational fMRI data.

  115. maybe Neuroscience score 4.3

    Neuronal Activity-Dependent Electroosmosis and Its Potential Role in Interstitial Fluid Flow in the Glymphatic System

    Hemmati, P., Wang, A. C., Prins, M. L. et al.

    This paper proposes that endogenous neuronal electric fields drive electroosmotic fluid flow through the brain’s narrow extracellular spaces, potentially explaining glymphatic transport and its increase during sleep. Computational models using reconstructed microstructure and local field recordings reportedly produce realistic flow speeds and brain-state effects, but the mechanism is not directly experimentally demonstrated.

    The electroosmosis-based explanation is a genuinely nonstandard candidate mechanism linking neural activity to glymphatic flow, but the evidence is primarily modeling-based and therefore insufficient for a strong verdict.

  116. maybe Neuroscience score 4.3

    Programming Brain Cell-Type-Selective Delivery In Vivo with Transporter-Guided Therapeutics

    Gunasekara, R. W., Zhang, L., Tong, L. et al.

    The authors develop ExACT, a drug-delivery platform that uses endogenous membrane transporters to direct fluorescent compounds and therapeutic cargoes into particular brain and retinal cell types. In mice and human-cell models, transporter-dependent uptake enabled selective delivery to endothelial cells, oligodendrocytes, neurons, and other populations; introducing SLCO1A2 into neurons created an artificial “entry port” for otherwise inaccessible conjugates. The approach worked with antisense oligonucleotides and small-molecule drugs while retaining activity.

    This presents a genuinely interesting alternative to broadly distributed CNS delivery—programming cell selectivity through transporter chemistry and even installing synthetic uptake pathways—but the abstract provides mainly uptake and proof-of-concept evidence, not yet therapeutic efficacy or safety in disease models.

  117. maybe BCI picked score 4.3

    Adaptive Neural Reorganization Enables Real-Time Finger-Level Robotic Control in BCI-Naïve Stroke Survivors

    Ding, Y., Karrenbach, M., Johnson, Z. et al.

    The study used EEG and deep-learning decoders to let nine BCI-naive stroke survivors control individual fingers of a robotic hand through motor imagery. Participants reached average decoding accuracies of 84% for two-finger control and 61% for three-finger control, while neural analyses indicated stroke-related reorganization that still preserves discriminable fine-motor signals.

    Finger-level robotic control from noninvasive EEG in BCI-naive stroke survivors is a meaningful capability result, but the small sample and limited abstract-level evidence make it too preliminary for a strong verdict.

  118. maybe Neuroscience score 4.3

    Terminal Schwann Cells Regulate Presynaptic Vesicle Homeostasis but Not Neuromuscular Junction Integrity in Mice

    Kim, H., Kim, S.-Y., Yoo, K. et al.

    The authors identify Col20a1 as a specific marker of terminal Schwann cells and create an inducible mouse model to label or ablate them without targeting other Schwann cells. Removing these cells during postnatal maturation leaves NMJ structure and motor behavior largely intact, but disrupts presynaptic vesicle availability and short-term transmission; later Schwann-cell repopulation appears to preserve NMJ integrity.

    A specific genetic tool plus the unexpected separation of presynaptic vesicle homeostasis from gross NMJ maintenance makes this a meaningful neuroscience result, though the abstract does not provide enough experimental detail for a stronger recommendation.

  119. maybe AI / ML ▲ 9 score 4.3

    LooseControlVideo: Directorial Video Control using Spatial Blocking

    Shariq Farooq Bhat, Niloy J. Mitra, Kalyan Sunkavalli

    LooseControlVideo uses sparse, oriented 3D boxes as a high-level “blocking” interface for controlling text-to-video scenes, instead of requiring dense frame-by-frame depth or flow guidance. Fine-tuning a Wan 2.2 model with an occlusion-aware 3D representation reportedly improves trajectory, rigid-motion, and occlusion metrics on several benchmarks, while supporting local edits with less disruption to the rest of the scene.

    The sparse 3D blocking interface is a meaningful and plausibly useful control abstraction for multi-object video generation, but the abstract provides limited detail about evaluation breadth and real-world authoring gains, so it does not yet warrant strong priority.

  120. maybe AI / ML ▲ 26 score 4.3

    Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

    Xuanming Zhang, Sining Zhoubian, Yuxuan Chen et al.

    The paper argues that an LLM’s final layer can sometimes move a good intermediate prediction toward generic or alignment-favored tokens. It proposes a training-free decoder that uses entropy to select a reliable near-final layer on each step, reporting gains on reasoning benchmarks for dense and MoE models with no memory cost and under 2% extra latency.

    Dynamically bypassing harmful final-layer changes is a plausible and somewhat surprising extension of early-exit and intermediate-layer decoding, but the abstract gives no quantitative gains or broad ablations to establish that it is more than a useful variant of existing layer-selection methods.

  121. maybe AI / ML ▲ 15 score 4.3

    Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

    Wujian Peng, Lingchen Meng, Yuxuan Cai et al.

    UniAR uses one discrete visual tokenizer for both image understanding and generation, so the autoregressive model can consume its own visual outputs without re-encoding them. It combines multi-level vision features, bitwise quantization, parallel prediction of spatial visual codes, and a diffusion decoder, claiming strong image generation/editing and competitive multimodal understanding after large-scale training.

    The genuinely unified shared-tokenizer design and efficient bitwise visual coding are worth a look, but the abstract gives no quantitative results and the overall direction builds on an active line of unified multimodal models.

  122. maybe AI / ML ▲ 59 score 4.2

    GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

    Tongxu Luo, Rongsheng Wang, Jiaxi Bi et al.

    The paper introduces GameCraft-Bench, a benchmark of 140 Godot game-generation tasks that evaluates whether coding agents produce complete, playable games rather than merely plausible code. Testing frontier agents, it finds that even the best reaches only 41.46%, with common failures in content completeness, visual feedback, and coherent presentation despite implementing basic mechanics.

    The interaction-grounded evaluation of end-to-end game creation is a useful, relatively novel test of coding-agent capability, and the low scores expose a meaningful gap between implementing mechanics and producing finished interactive artifacts, though it is still primarily a benchmark paper with limited evidence in the abstract.

  123. maybe AI / ML score 4.2

    Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

    Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi et al.

    The paper compresses activations exchanged during context-parallel language-model training by constraining them to dynamically learned mixtures of low-dimensional subspaces. It reports over 95% communication reduction, enabling billion-parameter models with contexts beyond 100K tokens over 300 Mbps links while reportedly preserving convergence speed relative to centralized training.

    The proposed communication mechanism and claimed ability to perform very long-context decentralized training over ordinary networks are potentially important, but the abstract provides too few experimental details to validate the unusually strong systems and convergence claims.

  124. maybe Neuroscience score 4.2

    Learning Hybrid Biophysical Neuron Models with Neural ODEs

    Jonas Beck, Michael Deistler, Dóra Viktória Molnár et al.

    The paper embeds neural ODEs inside conductance-based neuron models to learn unknown ion-channel currents or kinetics while retaining interpretable gating variables. It recovers gating dynamics from voltage recordings, generalizes across stimulus regimes and model misspecification, and compresses a multicompartment cortical-neuron model into a single-compartment model with up to 10× lower computational cost.

    This is a substantive hybrid mechanistic–learned modeling framework with broad simulation tests and a potentially useful order-of-magnitude speedup, though the abstract does not establish a major new biological finding or a decisive capability leap.

  125. maybe Robotics score 4.2

    TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations

    Zikang Xiong, Weixin Li, Zhouchonghao Wu et al.

    TerraTransfer trains a driving policy through self-play in a fast vectorized simulator, then transfers that policy into image inputs by aligning its latent representations with a pretrained vision model. This avoids expert driving demonstrations and instead uses paired images and simulator state, while exposing the learner to collisions, near-misses, and recoveries that are rare in logged data. The paper reports comparable or better closed-loop performance than prior end-to-end driving methods, though the abstract gives no quantitative results.

    The separation of learning to drive from learning to see, using self-play actions rather than expert trajectories to supervise visual grounding, is a genuinely interesting direction, but the abstract provides insufficient quantitative or real-world evidence for a stronger verdict.

  126. maybe Tech score 4.2

    A 399uW 114.3 dB DR Companding Readout ASIC for MEMS Microphones Employing a Multirate Time-Domain ADC

    Javier Granizo, Ruben Garvi, Ricardo Carrero et al.

    The paper presents a measured MEMS-microphone readout ASIC that combines companding with a VCO-based, multirate time-domain ADC. A 0.13 μm prototype achieves 114.3 dB dynamic range and over 112 dBc peak SFDR at under 400 μW, while producing standard single-bit PDM and reducing the audible boundary artifacts common in companding microphones.

    The combination of artifact-resistant companding and multirate time-domain conversion, backed by strong measured dynamic-range-per-power results in a complete microphone ASIC, is a meaningful hardware advance, though it is specialized engineering rather than a broadly new direction.

  127. maybe AI / ML score 4.2

    NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment

    Jisung Hwang, Yunhong Min, Jaihoon Kim et al.

    The paper proposes Noise-Tilted Reverse Kernels, a diffusion sampling method that applies reward gradients by modifying the noise distribution rather than shifting the reverse-process mean. This is intended to preserve the pretrained model’s sampling trajectory while still steering outputs toward a reward, and the authors report matching or exceeding prior reward-guidance methods with far fewer function evaluations—claiming a 20× reduction on aesthetic generation.

    The noise-space guidance mechanism and reported 20× inference reduction are potentially important, but the abstract gives no quantitative results, task breadth, or quality comparisons sufficient for a strong recommendation.

  128. maybe AI / ML ▲ 486 score 4.2

    Looped World Models

    Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang et al.

    LoopWM applies a parameter-shared transformer block repeatedly to refine latent environment states, allowing the effective computation depth to vary per prediction while keeping the parameter count fixed. The abstract claims up to 100× parameter efficiency and presents iterative latent depth as an additional scaling axis for world models, but gives no experimental details or quantitative comparisons beyond that headline.

    The shared-weight iterative architecture could be a useful new direction for efficient long-horizon world models, but the abstract provides too little evidence to substantiate the unusually large efficiency claim.

  129. maybe Robotics score 4.2

    SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

    Wei-Cheng Tseng, Gashon Hussein, Yuzhu Dong et al.

    SC3-Eval adapts a pretrained video model into a scalable evaluator for robot manipulation policies, using forward/inverse dynamics, cross-camera consistency, and an inference-time drift check. Across seven real-world vision-language-action policies, its simulated closed-loop scores correlate strongly with real-world performance (Pearson 0.929), while also reproducing specific failure modes and transferring to new tasks.

    The combination of action-manifold consistency, multi-view coherence, and uncertainty-based rollout termination addresses real limitations of video-world-model evaluation, with unusually strong real-world correlation, but it is still primarily an evaluation recipe rather than a demonstrated leap in robot capability.

  130. maybe Tech score 4.2

    Optimal Ansatz-free Hamiltonian Learning In Situ

    Taiqi Zhou, Weiyuan Gong

    The paper gives a control-free, ancilla-free method for learning an otherwise unspecified quantum Hamiltonian using only Pauli-product state preparation and measurement. It achieves total evolution time Θ(Λ/ε² log(Λ/ε)), proves this scaling is unavoidable for control-free protocols, and avoids the extremely fine time resolution required by earlier Heisenberg-limited methods; it also argues robustness to SPAM noise for local Hamiltonians.

    The matching lower bound and experimentally simpler protocol establish a meaningful optimality result for in-situ quantum characterization, though its relevance is mainly to quantum-device researchers rather than the reader’s central AI/robotics/neuroscience interests.

  131. maybe AI / ML score 4.2

    Online Dynamic Batching with Formal Guarantees for LLM Training

    Dian Li, Zekun Wang, Yaoru Wang et al.

    The paper forms training batches after preprocessing and tokenization, when the actual text and visual-token lengths are known, rather than relying on stale or expensive length caches. It adds distributed synchronization guarantees so dynamically sized batches do not deadlock or lose coverage, and reports roughly 1.6–3.8× higher emitted-sample throughput than fixed batching across Qwen3-VL fine-tuning runs, while approaching offline token-budget baselines and preserving quality.

    The combination of online, drop-in length-aware batching with formal DDP alignment guarantees addresses a real systems bottleneck and shows large throughput gains, but the contribution appears primarily engineering-focused and the abstract does not establish comparable gains beyond the tested multimodal fine-tuning settings.

  132. maybe AI / ML score 4.2

    HEPTv2: End-to-End Efficient Point Transformer for Charged Particle Reconstruction

    Siqi Miao, Shitij Govil, Jack P. Rodgers et al.

    HEPTv2 is an end-to-end point-transformer for charged-particle tracking that avoids graph construction, clustering, and filtering. On TrackML, it reports 98.6% tracking efficiency at 0.8% fake rate, with roughly 15 ms latency and 0.4 GB memory per event on an A100, while scaling approximately linearly to 500,000 hits.

    The combination of locality-aware hashing, direct sectorized track decoding, and unusually strong accuracy-latency results is substantially more than an incremental model tweak, but the evidence is limited to a specialized benchmark and the method’s broader applicability is unclear.

  133. maybe Robotics ▲ 12 score 4.2

    Vesta: A Generalist Embodied Reasoning Model

    Johan Bjorck, Zhiqi Li, Yunze Man et al.

    Vesta is a single multimodal model intended to handle localization, spatial reasoning, navigation, and long-horizon planning, rather than using separate specialist models. It combines large-scale spatially grounded training data with a multimodal memory mechanism, and reports substantially higher benchmark and real-robot task success than specialist baselines and ensembles.

    The unified generalist-plus-memory approach and claimed advantage over specialist ensembles could be important, but the abstract gives no concrete benchmarks, model scale, task details, or ablations to substantiate unusually large gains.

  134. maybe AI / ML score 4.2

    Equilibrium with Internal Transfers

    Mingyang Liu, Gabriele Farina, Asuman Ozdaglar

    The paper introduces transfer-based equilibrium concepts in which players make budget-balanced payments contingent on others following prescribed strategies. For polymatrix games, it shows that socially optimal profiles can be supported and computed efficiently, while a mediated version extends this implementation result to any finite game and produces an equilibrium of the original augmented game. The key distinction is between ordinary peer transfers, which have an agent-normal-form limitation, and binding mediation, which supplies stronger enforcement.

    The welfare-optimality and polynomial-time equilibrium results are a substantial mechanism-design direction, but their force depends on enforceable contingent transfers and, for general games, a binding mediator, making the result less broadly surprising than an unrestricted improvement to decentralized equilibrium computation.

  135. maybe AI / ML score 4.2

    SwarmX: Agentic Scheduling for Low-Latency Agentic Systems

    Yeqi Huang, Yanwei Ye, Guomin Chen et al.

    SwarmX is a serving and scheduling system for agentic applications whose latency depends on prompt semantics, model choice, tools, and heterogeneous CPU/GPU resources. It uses learned, distributional latency predictors to make tail-aware routing, scaling, and scheduling decisions, with online adaptation; production and testbed experiments report up to 61.5% lower tail latency and 2× throughput under the same SLO across coding, research, and multimodal workflows.

    The combination of semantic latency prediction and tail-aware scheduling is a substantive systems direction, backed by unusually large-scale deployment and sizable reported gains, though the abstract does not establish how much is due to generalizable ideas versus engineering and workload-specific tuning.

  136. maybe Robotics score 4.2

    Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models

    Tianyou Wang, Anson Lei, Joe Watson et al.

    The paper learns a shared latent action space across demonstrations that may come from different action spaces or have no action labels. It grounds this representation by requiring latent actions to predict the same future visual outcomes, then uses it for behavioral cloning and action transfer. On five simulated and real manipulation tasks, it reports a 48% average success-rate improvement over BC and prior latent-action methods in data-scarce settings.

    Grounding cross-source action representations through predicted environmental effects is a meaningful alternative to heuristic alignment, and the multi-task real-world gains are substantial, though the abstract does not establish how broad or robust the improvement is.

  137. maybe AI / ML score 4.2

    Test-Time Training with Next-Token Prediction

    Xuan Ouyang, Zefan Cai, Junjie Hu

    The paper proposes a drop-in test-time training method for pretrained language models that uses next-token prediction itself to update fast weights during inference. Instead of learning a separate local value target, each update uses a projection of the next-position hidden state; across four released models from 0.6B to 8B parameters, this improves long-context RULER scores by 2.9–4.1 points and LongBench-v2 by 3.7–5.6 points without hurting general knowledge performance.

    This is a relatively clean, broadly applicable connection between the training objective and test-time fast-weight adaptation, supported by results across multiple pretrained backbones and long-context benchmarks, but the abstract does not establish a large enough capability jump or sufficiently strong baseline advantage for a strong verdict.

  138. maybe AI / ML score 4.2

    Steer, Don't Solve: Training Small Critic Models for Large Code Agents

    Shubham Gandhi, Yiqing Xie, Atharva Naik et al.

    The paper freezes a code agent and trains a much smaller critic to give feedback during execution, rather than only scoring completed trajectories. On SWE-bench Verified, the critic transfers across agents and improves performance by roughly 3–5 points; on one 80B agent it raises success from 20.8% to 25.2% while reducing inference cost through shorter trajectories.

    The intra-trajectory steering setup and cross-agent transfer are a meaningful alternative to expensive end-to-end agent training, with credible cost and accuracy results, but the evidence is concentrated on one coding benchmark and does not yet establish broad generality.

  139. maybe AI / ML score 4.2

    SPOTR: Spatio-temporal Pooling One-Token Reconstruction for Universal Physiological Signal Self-supervised Learning

    Yiyu Gui, Mingzhi Chen, Yuesheng Zhu et al.

    SPOTR pretrains a single representation token for an entire physiological recording, then reconstructs the signal from that bottleneck rather than modeling a long flattened spatiotemporal token sequence. Across 20 EEG, iEEG, ECG, and PPG datasets, it reports substantially better linear-probe AUC than the strongest baseline, while reducing latency by about 78% and peak GPU memory by 52% versus a general-purpose time-series foundation model.

    The cross-modality one-token bottleneck combined with large reported linear-probe and efficiency gains is more than a routine SSL tweak, but the abstract does not establish whether comparisons, datasets, or ablations support such unusually broad improvements.

  140. maybe Neuroscience score 4.2

    Deep Cellular and Spatial Profiling of the Mouse Spinal Cord Reveals Sex-Specific Neuron Types and the Ascending Projection Neuron Repertoire

    Cano-Gomez, L., Poldsam, H., Ishishita, S. et al.

    The authors combine retrograde tracing, spatial transcriptomics, and multiomic profiling across more than 750,000 mouse spinal-cord neurons. They identify 78 ascending projection-neuron classes, over 500 anatomically and transcriptionally distinct neuron types, strong rostro-caudal specialization, and previously unreported sex-specific spinal neurons, linking cell types to brain targets and behavioral phenotypes.

    This is an unusually deep, anatomically linked spinal-cord atlas that connects molecular cell types to long-range outputs and sex differences, but it is primarily a reference resource and the abstract does not establish a major mechanistic discovery.

  141. maybe Neuroscience score 4.2

    Striatal activity maintains a short-term action-outcome memory to guide future choice

    Girasole, A. E., Mandelbaum, G., Murray, L. C. et al.

    Using a mouse task where the previous reward or omission determines whether to repeat or switch an action, the authors tracked dopamine and striatal projection-neuron activity. They find that ventrolateral striatal activity represents recent action-outcome associations and that manipulating direct- or indirect-pathway neurons during the association or delay period causally biases the next choice. The proposed new point is that the striatum may actively maintain a short-lived action-outcome memory, rather than merely evaluating outcomes or selecting actions.

    This is a mechanistically interesting, causally tested account of how striatal circuits bridge past outcomes to future choices, though it remains a task-specific mouse result rather than a demonstrated broad revision of decision-making theory.

  142. maybe AI / ML ▲ 6 score 4.2

    Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

    Tolga Şakar

    Morpheus is a neural tokenizer for Turkish that predicts character-level morpheme boundaries while preserving exact reversibility, and it produces a structured word embedding in the same pass. It improves morphological alignment and lexical retrieval over standard subword and contextual baselines, while using less GPU memory, though contextual models still perform better on tasks requiring sentence context.

    The differentiable boundary-to-segment mechanism and joint reversible tokenization/embedding objective are a meaningful idea for agglutinative languages, but the evidence is largely confined to Turkish lexical tasks and does not yet establish broad language-modeling or generation benefits.

  143. maybe AI / ML ▲ 26 score 4.2

    From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning

    Chao Chen, Chengzu Li, Zhiwei Li et al.

    The paper uses the current policy LLM to inspect failures and training statistics, then redesign the next RL environment rather than relying on a fixed curriculum or manual changes. On a configurable FrozenLake-style multi-agent testbed, this procedure outperforms fixed environments and larger proprietary LLMs, and the authors find that failure evidence and retaining successful settings are especially useful. An interesting result is that a trained checkpoint is better at proposing environments for itself than the original base model.

    Automating environment redesign from a policy’s failure modes is a plausible new training direction, and the checkpoint-versus-base-model finding is non-obvious, but the evidence is limited to a synthetic benchmark and does not yet establish broad gains over existing automatic curriculum methods.

  144. maybe AI / ML ▲ 5 score 4.2

    Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation

    Zhongzhu Zhou, Qingyang Wu, Junxiong Wang et al.

    Taylor-Calibrate initializes Gated DeltaNet layers converted from pretrained Transformer attention using Taylor-approximated teacher statistics, setting recurrent memory decay, write gates, value projections, and output gating, followed by brief layerwise alignment. Across four teachers and three layer-retention strategies, it reportedly produces much stronger initial students and reaches comparable distillation targets with 4.9–9.2× fewer training tokens than naive parameter copying.

    This is a potentially useful and non-obvious solution to a practical bottleneck in converting Transformers to efficient recurrent attention, with substantial reported reductions in distillation cost, but the abstract lacks absolute quality results and broader evidence needed for a stronger verdict.

  145. maybe Robotics ▲ 5 score 4.2

    Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

    Rishit Dagli, Donglai Xiang, Vismay Modi et al.

    AdaVoMP predicts dense, spatially varying material properties—Young’s modulus, Poisson’s ratio, and density—from 3D shapes. It replaces fixed-resolution voxel grids with a learned sparse adaptive voxel representation and autoregressive sparse transformer, reportedly reaching 16^3 higher effective resolution with lower test-time compute and producing more realistic deformable simulations.

    The adaptive output representation and large claimed resolution/memory improvement are meaningful for simulation-ready 3D assets, but the abstract provides no quantitative results or evidence that the predicted materials are physically accurate beyond comparisons to prior methods.

  146. maybe AI / ML ▲ 4 score 4.1

    RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

    Ruishan Fang, Siyuan Lu, Chenyi Zhuang et al.

    RODS uses the variance of rollout rewards already produced during GRPO to identify tasks near the agent’s capability boundary, where successes and failures are mixed. It then synthesizes structurally similar multi-turn tool-use tasks and continually refreshes a replay buffer; in a controlled experiment, about 400 human seeds and an active pool of 800 samples matched a 17K-sample offline pipeline with roughly 20× fewer trajectories.

    The boundary-focused online data-generation loop is a meaningful idea with potentially large data-efficiency benefits, but the abstract reports only controlled-setting results and does not establish broad generality beyond the claimed comparison.

  147. maybe Robotics ▲ 55 score 4.1

    ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

    Hao Li, Ganlong Zhao, Yufei Liu et al.

    ACE-Ego-0 pretrains a vision-language-action model jointly on robot trajectories and large amounts of egocentric human video. It converts human videos into noisy robot-style pseudo-actions, aligns them through camera-space and time-chunked representations, and weights supervision by estimated reliability; the resulting model reportedly improves benchmark and real-world bimanual manipulation performance.

    The potentially useful contribution is a practical scheme for turning abundant human video into auxiliary VLA supervision, but the abstract gives no quantitative gains or ablations, so it is unclear whether this is a substantial advance beyond existing human-video pretraining approaches.

  148. maybe AI / ML score 4.1

    Polynomial-Time Mistake-Bounded Language Generation

    Héctor Jimenez, Alexander Kozachinskiy, Vicente Opazo

    This paper studies when a Boolean-function class can generate hypotheses efficiently while making only polynomially many mistakes, extending the mistake-bounded language-generation framework. It gives efficient results for parities, symmetric functions, 2CNFs, and some monotone classes, while cryptographic lower bounds show separations from PAC learnability and failure of closure under union. The results suggest that efficient mistake-bounded generation has a substantially different structure from standard efficient learnability.

    The cryptographic separations between polynomial-time MBLG and PAC learning, including the non-closure result, are a genuinely non-obvious theoretical contribution, but the impact is mainly specialized learning theory rather than an immediate advance in practical ML.

  149. maybe Tech score 4.1

    Rhythm of the Deep: Two-Tier Combinatorial Structure in Sperm Whale Codas Revealed by Acoustic Unit Induction

    Mudit Sinha, Sanika Chavan

    The paper uses frozen audio encoders and controlled waveform tests to argue that sperm-whale codas have layered structure: recurring acoustic click units combine with rhythm into coda units, followed by weaker sequence-level dependencies between codas. Across 1,483 recordings, click composition and rhythm predict induced coda representations, while expert timing features do not fully recover the same structure; however, the discovered units remain encoder-induced rather than directly validated behavioral categories.

    The combination of acoustic unit induction, counterfactual controls, and evidence for two-tier whale vocal structure is genuinely interesting, but the biological significance is limited by reliance on frozen encoders and induced categories rather than independent behavioral or communicative validation.

  150. maybe AI / ML score 4.1

    Closing the Approximation Gap in Simulation-free Latent SDEs

    Henry D. Smith, Brian L. Trippe, Scott W. Linderman

    The paper identifies a limitation of existing simulation-free variational inference for latent SDEs: specifying posterior marginals can unnecessarily restrict the dynamics that the approximate posterior can represent. Helmholtz-SDE optimizes over a broader set of path laws consistent with those marginals, reportedly recovering dynamics more accurately—especially under high uncertainty—and matching simulation-based VI while running much faster.

    The method appears to remove a fundamental expressivity tradeoff in an increasingly useful class of latent-dynamics inference algorithms, but the abstract gives no quantitative results or details sufficient to justify a strong recommendation.