Newsdesk
Engineering
Data & AI
Industries
Enterprise Systems
Go-to-Market
Longform
DesignIndiaAll stories

arXiv Picks

29 July 2026

20 papers shortlisted this week — most discussed: Kimi K3: Open Frontier Intelligence (291 upvotes).

One week of arXiv Picks, 6 stories, as published.

LLMs & Reasoning

Kimi K3: Open Frontier Intelligence

This paper introduces Kimi K3, a massive 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Its architecture incorporates Kimi Delta Attention and Attention Residuals to enhance information flow across sequence length and model depth, alongside Stable LatentMoE for efficient expert activation.

Why it matters — It provides insight into the architectural innovations and scaling strategies behind a new frontier-level multimodal LLM, detailing specific attention mechanisms and MoE techniques.

arXiv · 29 July 2026 · Read the original →

Agents & Tools

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

This work proposes StateAct, an approach for computer-use agents that prioritizes underlying program state over pixel-based screenshots for perception. It argues that screenshots are a lossy rendering, while program state (files, backends, DOM) offers direct, lossless access to task data, enabling more robust inspection and modification.

Why it matters — It presents a fundamental design decision for improving computer-use agents by shifting perception from visual rendering to direct program state interaction, offering a more robust and less lossy approach.

arXiv · 29 July 2026 · Read the original →

Multimodal

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

SANA-Video 2.0 is introduced as a hybrid video diffusion transformer, scaled to 5B and 14B parameters, designed for high-quality 720p video generation on a single GPU. It achieves quality comparable to full-softmax video DiTs while maintaining the favorable long-sequence scaling of linear attention by employing Hybrid Linear-Softmax Attention and Attention Residuals.

Why it matters — This paper details architectural innovations, specifically hybrid attention and attention residuals, that enable efficient and high-quality video generation, addressing a key scaling challenge in multimodal models.

arXiv · 29 July 2026 · Read the original →

LLMs & Reasoning

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

This research explores interaction-driven self-evolution for LLM training, moving beyond manual design and annotation. It introduces Skill Self-Play to address the dilemma between task diversity and verification reliability, allowing LLMs to co-evolve skills and broaden their task space while maintaining reliable verification.

Why it matters — It presents an advanced training paradigm for LLMs that leverages self-evolution and co-evolving skills to overcome fundamental limitations in scaling LLM capabilities and task diversity.

arXiv · 29 July 2026 · Read the original →

Systems & Inference

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

This paper addresses the dominant attention bottleneck in diffusion transformers for high-fidelity video generation by proposing Sol-Attn, an on-the-fly attention sparsification method. It aims to overcome the limitations of existing training-free dynamic sparse attention methods, which struggle with efficient and accurate sparsification due to rigid or costly routing.

Why it matters — It offers a crucial systems and inference optimization technique for video generation models, tackling the computational cost of attention through dynamic sparsification.

arXiv · 29 July 2026 · Read the original →

Research

Data Pyramid for Embodied Manipulation

This work proposes organizing the embodied data ecosystem as a 'pyramid' of five complementary sources to address the unique data requirements of embodied agents. Unlike multimodal foundation models that leverage internet-scale data, embodied agents need data coupling observations with physical states and actions, which this pyramid structure aims to provide.

Why it matters — It offers a structured approach to understanding and leveraging diverse data sources for embodied AI, which is critical for overcoming the data scarcity bottleneck in learning physical interactions.

arXiv · 29 July 2026 · Read the original →

Read the abstract. Skim the method. Steal the idea.

6 stories, every Wednesday

Published here every week. Follow by RSS to get it as it lands.

Next issue →