See2Think: Do Multimodal Models Really Use Intermediate Visual States? Paper • 2607.26769 • Published 11 days ago • 25
DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Paper • 2606.31537 • Published Jun 30 • 29
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Paper • 2606.14502 • Published Jun 12 • 117
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models Paper • 2505.19108 • Published May 25, 2025
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Paper • 2606.14502 • Published Jun 12 • 117
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding Paper • 2603.18472 • Published Mar 19 • 20
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention Paper • 2602.05847 • Published Feb 5 • 12