ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 12 days ago • 306
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue Paper • 2607.01213 • Published Jul 1 • 5
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing Paper • 2606.30599 • Published Jun 30 • 12
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts Paper • 2606.05922 • Published Jun 4 • 70