arXiv Picks
Papers Worth Reading. Published every Wednesday, here and by RSS.
48 stories 8 issues 8 categories
LLMs & Reasoning Agents & Tools Training & Efficiency Multimodal Evaluation & Benchmarks Safety & Alignment Systems & Inference Research
Top stories
RSSSANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
SANA-Video 2.0 is introduced as a hybrid video diffusion transformer, scaled to 5B and 14B parameters, designed for high-quality 720p video generation on a single GPU. It achieves quality comparable to full-softmax video DiTs while maintaining the favorable long-sequence scaling of linear attention by employing Hybrid Linear-Softmax Attention and Attention Residuals.
arXiv · 29 July 2026 · Read the original →
SenseNova-U1.5: Towards Native Unified Visual Intelligence
SenseNova-U1.5 is an 8B Mixture-of-Transformers model that achieves native multimodal intelligence using an encoder-free and VAE-free architecture. The model relies on spatially coherent patch reconstruction and structural prompt enhancement to handle visual understanding and generation within a single framework.
arXiv · 16 September 2026 · Read the original →
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
ZGCM-1 is a 7B dense foundation model designed to overcome the parametric capacity limits of small models. It shifts the paradigm from passive web-data memorization to a system that couples deliberate internal thinking with active external tool use across a 256K context window.
arXiv · 16 September 2026 · Read the original →
Kimi K3: Open Frontier Intelligence
This paper introduces Kimi K3, a massive 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Its architecture incorporates Kimi Delta Attention and Attention Residuals to enhance information flow across sequence length and model depth, alongside Stable LatentMoE for efficient expert activation.
arXiv · 29 July 2026 · Read the original →
DiffusionGemma Technical Report
DiffusionGemma introduces discrete diffusion to LLM decoding by iteratively refining blocks of 256 tokens in parallel. Converted directly from pre-trained Gemma weights rather than trained from scratch, the model bypasses sequential token-by-token autoregressive decoding bottlenecks while maintaining generation quality.
arXiv · 05 August 2026 · Read the original →
A lot of effort this week spent shrinking things we only just got to work.
Weekly editions
Each edition is one week's digest exactly as it was published — a summary of the week, then the stories.
- 16 September 202620 papers shortlisted this week — most discussed: Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation (628 upvotes).6 stories →
- 09 September 202620 papers shortlisted this week — most discussed: Compile by Training: Turning Natural-Language Specifications into Loca (378 upvotes).6 stories →
- 02 September 202620 papers shortlisted this week — most discussed: StudentSim: Training LLM-based Student Simulators (191 upvotes).6 stories →
- 26 August 202620 papers shortlisted this week — most discussed: EnvHarness: Awakening Static Worlds for Agent Learning (263 upvotes).6 stories →
- 19 August 202620 papers shortlisted this week — most discussed: Can We Defend Against AI-Generated Video Attacks on Real-World Crisis (268 upvotes).6 stories →
- 12 August 202620 papers shortlisted this week — most discussed: Macaron-V1: Towards Open Continual Learning with Self-Improvement and (282 upvotes).6 stories →
- 05 August 202620 papers shortlisted this week — most discussed: SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instru (142 upvotes).6 stories →
- 29 July 202620 papers shortlisted this week — most discussed: Kimi K3: Open Frontier Intelligence (291 upvotes).6 stories →