TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 12 days ago • 139
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published 21 days ago • 198
MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models Paper • 2607.11594 • Published 28 days ago • 7
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies Paper • 2606.15768 • Published Jun 14 • 6
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data Paper • 2606.13432 • Published Jun 11 • 113
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253