TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Paper • 2607.21606 • Published May 16 • 3
Spectral Prior for Reducing Exposure Bias in Diffusion Models Paper • 2607.22091 • Published 8 days ago • 6
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 8 days ago • 47
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Paper • 2607.13429 • Published 17 days ago • 16
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation Paper • 2607.13365 • Published 17 days ago • 20
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 11 days ago • 58
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry Paper • 2607.18227 • Published 12 days ago • 50
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 10 days ago • 305
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Paper • 2607.19344 • Published 11 days ago • 4
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 11 days ago • 75
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 14 days ago • 137
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation Paper • 2607.15686 • Published 15 days ago • 16
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 16 days ago • 71
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 14 days ago • 138
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation Paper • 2607.09362 • Published 22 days ago • 12
Latent-Identity Tuning in Text-to-Image Personalization Models Paper • 2607.11885 • Published 19 days ago • 14
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Paper • 2607.04461 • Published 27 days ago • 11
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published 19 days ago • 44