Newsdesk
Engineering
Data & AI
Industries
Enterprise Systems
Go-to-Market
Longform
DesignIndiaAll stories

arXiv Picks

Papers Worth Reading. Published every Wednesday, here and by RSS.

48 stories 8 issues 8 categories

LLMs & Reasoning Agents & Tools Training & Efficiency Multimodal Evaluation & Benchmarks Safety & Alignment Systems & Inference Research

Top stories

RSS
Multimodal

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

SANA-Video 2.0 is introduced as a hybrid video diffusion transformer, scaled to 5B and 14B parameters, designed for high-quality 720p video generation on a single GPU. It achieves quality comparable to full-softmax video DiTs while maintaining the favorable long-sequence scaling of linear attention by employing Hybrid Linear-Softmax Attention and Attention Residuals.

Why it matters — This paper details architectural innovations, specifically hybrid attention and attention residuals, that enable efficient and high-quality video generation, addressing a key scaling challenge in multimodal models.

arXiv · 29 July 2026 · Read the original →

Multimodal

SenseNova-U1.5: Towards Native Unified Visual Intelligence

SenseNova-U1.5 is an 8B Mixture-of-Transformers model that achieves native multimodal intelligence using an encoder-free and VAE-free architecture. The model relies on spatially coherent patch reconstruction and structural prompt enhancement to handle visual understanding and generation within a single framework.

Why it matters — Native multimodal intelligence can be achieved without separate encoders or VAEs by using spatially coherent patch reconstruction within a unified transformer architecture.

arXiv · 16 September 2026 · Read the original →

Training & Efficiency

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1 is a 7B dense foundation model designed to overcome the parametric capacity limits of small models. It shifts the paradigm from passive web-data memorization to a system that couples deliberate internal thinking with active external tool use across a 256K context window.

Why it matters — Small-parameter models can match larger ones by substituting internal memorization with a combination of deliberate chain-of-thought and external tool-augmented search.

arXiv · 16 September 2026 · Read the original →

LLMs & Reasoning

Kimi K3: Open Frontier Intelligence

This paper introduces Kimi K3, a massive 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Its architecture incorporates Kimi Delta Attention and Attention Residuals to enhance information flow across sequence length and model depth, alongside Stable LatentMoE for efficient expert activation.

Why it matters — It provides insight into the architectural innovations and scaling strategies behind a new frontier-level multimodal LLM, detailing specific attention mechanisms and MoE techniques.

arXiv · 29 July 2026 · Read the original →

Systems & Inference

DiffusionGemma Technical Report

DiffusionGemma introduces discrete diffusion to LLM decoding by iteratively refining blocks of 256 tokens in parallel. Converted directly from pre-trained Gemma weights rather than trained from scratch, the model bypasses sequential token-by-token autoregressive decoding bottlenecks while maintaining generation quality.

Why it matters — Adapting autoregressive model weights into block-parallel discrete diffusion bypasses sequential KV-cache bottlenecks without requiring full pre-training from scratch.

arXiv · 05 August 2026 · Read the original →

A lot of effort this week spent shrinking things we only just got to work.

Weekly editions

Each edition is one week's digest exactly as it was published — a summary of the week, then the stories.