trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 Text Generation • 2.43M • Updated Dec 19, 2025 • 16.1M • 17
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Paper • 2606.30124 • Published 30 days ago • 8
Yuivdldk/gsm8k-results-g4b-lora-lr0.0001-ep2-s42-ntr4000-mn512-n1319-d1step Preview • Updated 25 days ago • 95 • 1
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks Paper • 2606.12344 • Published Jun 10 • 71
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration Paper • 2605.28184 • Published May 27 • 6
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171