16 September 2026
20 papers shortlisted this week — most discussed: Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation (628 upvotes).
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
This paper introduces a method for post-training attention sparsification that optimizes context ranking end-to-end. By avoiding the hard Top-K selection used in existing methods, it prevents gradient blocking from the language modeling loss, allowing for more efficient reduction of quadratic attention costs in Transformers.
arXiv · 16 September 2026 · Read the original →
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
The authors critique the popular On-Policy Self-Distillation (OPSD) paradigm, finding that it degrades performance on complex reasoning by forcing models to imitate artificially compressed paths. They propose Negative Self-Distillation, which focuses on teaching models to improve by identifying and avoiding reasoning flaws.
arXiv · 16 September 2026 · Read the original →
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
This research addresses the 'vision-action shortcut' in robotics, where models exploit task-irrelevant visual cues that fail under distribution shifts. It introduces Latent Interface Training to decouple visual representations from action generation, improving the generalization of foundation models.
arXiv · 16 September 2026 · Read the original →
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
X-AuT is a progressive compression framework for speech LLMs that reduces audio-encoder depth without causing deletion or end-of-sequence errors. It uses short behavioral probes to select optimal layer combinations for pruning and restores performance through cross-scale representation alignment.
arXiv · 16 September 2026 · Read the original →
SenseNova-U1.5: Towards Native Unified Visual Intelligence
SenseNova-U1.5 is an 8B Mixture-of-Transformers model that achieves native multimodal intelligence using an encoder-free and VAE-free architecture. The model relies on spatially coherent patch reconstruction and structural prompt enhancement to handle visual understanding and generation within a single framework.
arXiv · 16 September 2026 · Read the original →
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
ZGCM-1 is a 7B dense foundation model designed to overcome the parametric capacity limits of small models. It shifts the paradigm from passive web-data memorization to a system that couples deliberate internal thinking with active external tool use across a 256K context window.
arXiv · 16 September 2026 · Read the original →
A lot of effort this week spent shrinking things we only just got to work.
6 stories, every Wednesday
Published here every week. Follow by RSS to get it as it lands.