MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors Paper • 2607.12000 • Published 13 days ago • 39
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 7 days ago • 165
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Paper • 2607.12752 • Published 11 days ago • 20
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 4 days ago • 291
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders Paper • 2607.14088 • Published 11 days ago • 14
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published 10 days ago • 68
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Paper • 2607.11683 • Published 13 days ago • 143
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 10 days ago • 139
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published 9 days ago • 12
GRASP: GRanularity-Aware Search Policy for Agentic RAG Paper • 2607.10463 • Published 15 days ago • 8
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Paper • 2607.15273 • Published 10 days ago • 17
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 10 days ago • 69
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 10 days ago • 53
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering Paper • 2603.28583 • Published 12 days ago • 17
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published 13 days ago • 23
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 13 days ago • 84
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published 17 days ago • 87
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 16 days ago • 84