Running on Zero MCP Featured 42 LTX-Best-Face-ID 🎬 42 Distilled LTX-2.3 identity video from a reference photo
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 16 days ago • 76
Running on Zero MCP Featured 1.56k FireRed Image Edit 1.0 Fast 🔥 1.56k FireRed-Image-Edit × Qwen-Image-Edit-Rapid (Transformers)
Running on Zero MCP 2.1k Wan2.2 14B Fast Preview 🐌 2.1k generate a video from an image with a text prompt
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents Paper • 2604.07430 • Published Apr 8 • 181
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Paper • 2604.11784 • Published Apr 13 • 143
HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive Image-Text-to-Text • 8B • Updated Apr 6 • 601k • 942
Cephalo Collection Cephalo is a series of multimodal vision large language models (V-LLMs) designed to integrate visual and linguistic reasoning in materials science. • 17 items • Updated Apr 16, 2025 • 5