Newsdesk
Engineering
Data & AI
Industries
Enterprise Systems
Go-to-Market
Longform
DesignIndiaAll stories

arXiv Picks

19 August 2026

20 papers shortlisted this week — most discussed: Can We Defend Against AI-Generated Video Attacks on Real-World Crisis (268 upvotes).

One week of arXiv Picks, 6 stories, as published.

Agents & Tools

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Introduces StateM, an agent-native runtime designed to improve execution systems around an agent without changing model weights, a concept termed harness scaling. It addresses common long-horizon agent failures such as losing track of mutable state, skipping known procedures, or stopping prematurely.

Why it matters — Managing mutable state and execution history in an external runtime ('harness scaling') can significantly improve long-horizon task accuracy without the need for expensive model fine-tuning.

arXiv · 19 August 2026 · Read the original →

Training & Efficiency

Self-Supervised Visual On-Policy Distillation

Proposes a method for visual on-policy distillation that eliminates the need for privileged teacher models, reference answers, or ground-truth annotations. It achieves this by inverting the typical teacher-student informational asymmetry, enabling self-supervised distillation.

Why it matters — Visual distillation can be performed in a self-supervised manner by structurally inverting the informational asymmetry between models, removing the dependency on privileged teachers or labeled data.

arXiv · 19 August 2026 · Read the original →

Evaluation & Benchmarks

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Introduces a framework that automates the evaluation of visual world models by using agents to judge physical and causal consistency. This moves beyond simple scalar metrics to provide reasoned justifications for rollout quality.

Why it matters — Evaluating complex world models requires agentic reasoning to detect physical and causal violations, which traditional scalar metrics fail to capture.

arXiv · 19 August 2026 · Read the original →

Training & Efficiency

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Identifies a fundamental flaw in multi-reward reinforcement learning for language models, where fixed-weight scalarization of reward vectors allows mastered objectives to dominate gradients. It introduces Saturation Aware Advantage Reweighting to dynamically adjust advantages based on objective mastery.

Why it matters — Dynamically reweighting advantages based on objective saturation prevents already-mastered tasks from dominating the gradient in multi-reward RL post-training.

arXiv · 19 August 2026 · Read the original →

Training & Efficiency

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

Investigates on-policy distillation from long-context teacher models to short-context student models. It addresses key implementation challenges including tokenizer mismatch, distribution mismatch, response length explosion, and training instability.

Why it matters — Transferring reasoning capabilities across different context lengths requires addressing tokenizer and distribution mismatches to stabilize training and prevent response length explosion.

arXiv · 19 August 2026 · Read the original →

Systems & Inference

Agentic Transaction: Towards ACID-Compliant Agent Systems

Explores the application of database transaction concepts—specifically ACID compliance—to LLM agent systems. It aims to ensure reliable execution, state consistency, and error recovery for agents operating over persistent environments and multi-step workflows.

Why it matters — Applying ACID transaction principles to agent workflows provides a structured framework for handling execution failures and maintaining state consistency in persistent environments.

arXiv · 19 August 2026 · Read the original →

Applying ACID compliance to agents suggests we have reached the bargaining stage of reliability.

6 stories, every Wednesday

Published here every week. Follow by RSS to get it as it lands.

← Previous issue Next issue →