Composition-RL Datasets and trained checkpoints of Composition-RL: https://github.com/XinXU-USTC/Composition-RL xx18/Polaris-Composition-1323K Viewer • Updated Apr 27 • 1.32M • 41 • 1 xx18/Composition-RL-EVA Viewer • Updated Apr 27 • 12.8k • 105 • 1 xx18/MATH-Composition-199K Viewer • Updated Apr 27 • 199k • 36 • 1 xx18/Composition-RL-4B Text Generation • 4B • Updated Apr 27 • 13
TFPI ICLR2026: Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners https://arxiv.org/abs/2509.26226 xx18/TFPI-EVA Preview • Updated Sep 28, 2025 • 50 • 1 xx18/TFPI-DeepSeek-Qwen-1.5B-Stage1 Text Generation • 2B • Updated Feb 12 • 2 xx18/TFPI-DeepSeek-Qwen-1.5B-Stage2 Text Generation • 2B • Updated Feb 12 • 8 xx18/TFPI-DeepSeek-Qwen-1.5B-Stage3 Text Generation • 2B • Updated Feb 12 • 4
Composition-RL Datasets and trained checkpoints of Composition-RL: https://github.com/XinXU-USTC/Composition-RL xx18/Polaris-Composition-1323K Viewer • Updated Apr 27 • 1.32M • 41 • 1 xx18/Composition-RL-EVA Viewer • Updated Apr 27 • 12.8k • 105 • 1 xx18/MATH-Composition-199K Viewer • Updated Apr 27 • 199k • 36 • 1 xx18/Composition-RL-4B Text Generation • 4B • Updated Apr 27 • 13
TFPI ICLR2026: Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners https://arxiv.org/abs/2509.26226 xx18/TFPI-EVA Preview • Updated Sep 28, 2025 • 50 • 1 xx18/TFPI-DeepSeek-Qwen-1.5B-Stage1 Text Generation • 2B • Updated Feb 12 • 2 xx18/TFPI-DeepSeek-Qwen-1.5B-Stage2 Text Generation • 2B • Updated Feb 12 • 8 xx18/TFPI-DeepSeek-Qwen-1.5B-Stage3 Text Generation • 2B • Updated Feb 12 • 4