Newsdesk
Engineering
Data & AI
Industries
Enterprise Systems
Go-to-Market
Longform
DesignIndiaAll stories

arXiv Picks

16 September 2026

20 papers shortlisted this week — most discussed: Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation (628 upvotes).

One week of arXiv Picks, 6 stories, as published.

Systems & Inference

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

This paper introduces a method for post-training attention sparsification that optimizes context ranking end-to-end. By avoiding the hard Top-K selection used in existing methods, it prevents gradient blocking from the language modeling loss, allowing for more efficient reduction of quadratic attention costs in Transformers.

Why it matters — End-to-end optimization of context ranking allows for trainable attention sparsification without the gradient-blocking issues of traditional Top-K selectors.

arXiv · 16 September 2026 · Read the original →

LLMs & Reasoning

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

The authors critique the popular On-Policy Self-Distillation (OPSD) paradigm, finding that it degrades performance on complex reasoning by forcing models to imitate artificially compressed paths. They propose Negative Self-Distillation, which focuses on teaching models to improve by identifying and avoiding reasoning flaws.

Why it matters — Training models to recognize and avoid flawed reasoning steps is more effective for complex tasks than forcing them to imitate 'perfect' but artificially compressed reasoning paths.

arXiv · 16 September 2026 · Read the original →

Research

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

This research addresses the 'vision-action shortcut' in robotics, where models exploit task-irrelevant visual cues that fail under distribution shifts. It introduces Latent Interface Training to decouple visual representations from action generation, improving the generalization of foundation models.

Why it matters — Decoupling visual representation from action generation through a latent interface prevents robotics models from exploiting spurious visual correlations that do not generalize.

arXiv · 16 September 2026 · Read the original →

Systems & Inference

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

X-AuT is a progressive compression framework for speech LLMs that reduces audio-encoder depth without causing deletion or end-of-sequence errors. It uses short behavioral probes to select optimal layer combinations for pruning and restores performance through cross-scale representation alignment.

Why it matters — Pruning audio encoders by selecting layers via short behavioral probes, rather than simple block removal, preserves embedding stability and prevents premature sequence termination.

arXiv · 16 September 2026 · Read the original →

Multimodal

SenseNova-U1.5: Towards Native Unified Visual Intelligence

SenseNova-U1.5 is an 8B Mixture-of-Transformers model that achieves native multimodal intelligence using an encoder-free and VAE-free architecture. The model relies on spatially coherent patch reconstruction and structural prompt enhancement to handle visual understanding and generation within a single framework.

Why it matters — Native multimodal intelligence can be achieved without separate encoders or VAEs by using spatially coherent patch reconstruction within a unified transformer architecture.

arXiv · 16 September 2026 · Read the original →

Training & Efficiency

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1 is a 7B dense foundation model designed to overcome the parametric capacity limits of small models. It shifts the paradigm from passive web-data memorization to a system that couples deliberate internal thinking with active external tool use across a 256K context window.

Why it matters — Small-parameter models can match larger ones by substituting internal memorization with a combination of deliberate chain-of-thought and external tool-augmented search.

arXiv · 16 September 2026 · Read the original →

A lot of effort this week spent shrinking things we only just got to work.

6 stories, every Wednesday

Published here every week. Follow by RSS to get it as it lands.

← Previous issue