StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Paper • 2608.05703 • Published 7 days ago • 16
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Paper • 2607.28590 • Published 14 days ago • 46
MMSkills: Towards Multimodal Skills for General Visual Agents Paper • 2605.13527 • Published May 14 • 122
HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents Paper • 2605.07177 • Published May 8 • 63