On-Policy Self-Adaptation Collection Checkpoints of OPSA on different base models • 3 items • Updated about 11 hours ago • 1
On-Policy Self-Adaptation Collection Checkpoints of OPSA on different base models • 3 items • Updated about 11 hours ago • 1
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control Paper • 2604.26326 • Published May 10 • 14
TEMPO: Scaling Test-time Training for Large Reasoning Models Paper • 2604.19295 • Published Apr 21 • 35
Learning Self-Correction in Vision-Language Models via Rollout Augmentation Paper • 2602.08503 • Published Feb 9 • 3
Learning Self-Correction in Vision-Language Models via Rollout Augmentation Paper • 2602.08503 • Published Feb 9 • 3
Octopus Collection RL checkpoints of Octopus-8B and baselines of paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation • 6 items • Updated Feb 9