MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 16 days ago • 76
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents Paper • 2604.07430 • Published Apr 8 • 181
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Paper • 2604.11784 • Published Apr 13 • 143
Cephalo Collection Cephalo is a series of multimodal vision large language models (V-LLMs) designed to integrate visual and linguistic reasoning in materials science. • 17 items • Updated Apr 16, 2025 • 5