19 August 2026
20 papers shortlisted this week — most discussed: Can We Defend Against AI-Generated Video Attacks on Real-World Crisis (268 upvotes).
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Introduces StateM, an agent-native runtime designed to improve execution systems around an agent without changing model weights, a concept termed harness scaling. It addresses common long-horizon agent failures such as losing track of mutable state, skipping known procedures, or stopping prematurely.
arXiv · 19 August 2026 · Read the original →
Self-Supervised Visual On-Policy Distillation
Proposes a method for visual on-policy distillation that eliminates the need for privileged teacher models, reference answers, or ground-truth annotations. It achieves this by inverting the typical teacher-student informational asymmetry, enabling self-supervised distillation.
arXiv · 19 August 2026 · Read the original →
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Introduces a framework that automates the evaluation of visual world models by using agents to judge physical and causal consistency. This moves beyond simple scalar metrics to provide reasoned justifications for rollout quality.
arXiv · 19 August 2026 · Read the original →
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Identifies a fundamental flaw in multi-reward reinforcement learning for language models, where fixed-weight scalarization of reward vectors allows mastered objectives to dominate gradients. It introduces Saturation Aware Advantage Reweighting to dynamically adjust advantages based on objective mastery.
arXiv · 19 August 2026 · Read the original →
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
Investigates on-policy distillation from long-context teacher models to short-context student models. It addresses key implementation challenges including tokenizer mismatch, distribution mismatch, response length explosion, and training instability.
arXiv · 19 August 2026 · Read the original →
Agentic Transaction: Towards ACID-Compliant Agent Systems
Explores the application of database transaction concepts—specifically ACID compliance—to LLM agent systems. It aims to ensure reliable execution, state consistency, and error recovery for agents operating over persistent environments and multi-step workflows.
arXiv · 19 August 2026 · Read the original →
Applying ACID compliance to agents suggests we have reached the bargaining stage of reliability.
6 stories, every Wednesday
Published here every week. Follow by RSS to get it as it lands.