MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Paper • 2607.25948 • Published 2 days ago • 8
Parallel Decoding Distillation for Fast Image and Video Generation Paper • 2607.26004 • Published 2 days ago • 6
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding Paper • 2607.24743 • Published 3 days ago • 7
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published 3 days ago • 70
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 7 days ago • 14
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion Paper • 2607.20417 • Published 8 days ago • 9
Self-Improvements in Modern Agentic Systems: A Survey Paper • 2607.13104 • Published 16 days ago • 31
4D Human-Scene Reconstruction from Low-Overlap Captures Paper • 2607.09125 • Published 20 days ago • 53
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Paper • 2607.12752 • Published 15 days ago • 20
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors Paper • 2607.12000 • Published 17 days ago • 39
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 20 days ago • 85