TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 6 days ago • 136
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published 22 days ago • 44
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 19 days ago • 206
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Paper • 2607.11849 • Published 22 days ago • 33
hyzhang01/IISC_parallel_franka_all_fix_tracks_fixed Viewer • Updated about 15 hours ago • 333k • 120 • 1
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Paper • 2605.28132 • Published May 27 • 25
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
UniT: Unified Geometry Learning with Group Autoregressive Transformer Paper • 2605.21131 • Published May 20 • 8
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 274