Qwen3.5-4B Ariadna — доменный ML/AI-эксперт Collection Qwen3.5-4B, дообучен на 124K концептов arXiv. safetensors (GRPO v10) + GGUF Q4_K_M с PDF-отчётами и кейсами. • 2 items • Updated 3 days ago
Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF Image-Text-to-Text • 27B • Updated 26 days ago • 89.5k • 678
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning Paper • 2603.21065 • Published Mar 22 • 78
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training Paper • 2602.10693 • Published Feb 11 • 221