29 July 2026
20 papers shortlisted this week — most discussed: Kimi K3: Open Frontier Intelligence (291 upvotes).
Kimi K3: Open Frontier Intelligence
This paper introduces Kimi K3, a massive 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Its architecture incorporates Kimi Delta Attention and Attention Residuals to enhance information flow across sequence length and model depth, alongside Stable LatentMoE for efficient expert activation.
arXiv · 29 July 2026 · Read the original →
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
This work proposes StateAct, an approach for computer-use agents that prioritizes underlying program state over pixel-based screenshots for perception. It argues that screenshots are a lossy rendering, while program state (files, backends, DOM) offers direct, lossless access to task data, enabling more robust inspection and modification.
arXiv · 29 July 2026 · Read the original →
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
SANA-Video 2.0 is introduced as a hybrid video diffusion transformer, scaled to 5B and 14B parameters, designed for high-quality 720p video generation on a single GPU. It achieves quality comparable to full-softmax video DiTs while maintaining the favorable long-sequence scaling of linear attention by employing Hybrid Linear-Softmax Attention and Attention Residuals.
arXiv · 29 July 2026 · Read the original →
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
This research explores interaction-driven self-evolution for LLM training, moving beyond manual design and annotation. It introduces Skill Self-Play to address the dilemma between task diversity and verification reliability, allowing LLMs to co-evolve skills and broaden their task space while maintaining reliable verification.
arXiv · 29 July 2026 · Read the original →
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
This paper addresses the dominant attention bottleneck in diffusion transformers for high-fidelity video generation by proposing Sol-Attn, an on-the-fly attention sparsification method. It aims to overcome the limitations of existing training-free dynamic sparse attention methods, which struggle with efficient and accurate sparsification due to rigid or costly routing.
arXiv · 29 July 2026 · Read the original →
Data Pyramid for Embodied Manipulation
This work proposes organizing the embodied data ecosystem as a 'pyramid' of five complementary sources to address the unique data requirements of embodied agents. Unlike multimodal foundation models that leverage internet-scale data, embodied agents need data coupling observations with physical states and actions, which this pyramid structure aims to provide.
arXiv · 29 July 2026 · Read the original →
Read the abstract. Skim the method. Steal the idea.
6 stories, every Wednesday
Published here every week. Follow by RSS to get it as it lands.