Inflect v2 Collection Complete local text-to-waveform speech models at 3.96M and 9.36M parameters, with official PyTorch and ONNX Runtime releases. • 5 items • Updated 4 days ago • 8
Kimi-VL-A3B Collection Moonshot's efficient MoE VLMs, exceptional on agent, long-context, and thinking • 6 items • Updated Mar 2 • 84
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 3 days ago • 22
Parallel Decoding Distillation for Fast Image and Video Generation Paper • 2607.26004 • Published 2 days ago • 6
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 7 days ago • 37
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 8 days ago • 71
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 14 days ago • 170
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Paper • 2607.09581 • Published 20 days ago • 6
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 12 days ago • 138
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 16 days ago • 102
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published 22 days ago • 139
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Paper • 2607.07608 • Published 22 days ago • 56
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published 23 days ago • 95
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots Paper • 2607.02501 • Published 28 days ago • 59
TurboServe: Serving Streaming Video Generation Efficiently and Economically Paper • 2606.19271 • Published Jun 17 • 37