Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 4 days ago • 28
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 3 days ago • 55
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published 10 days ago • 8
AutoIndex: Learning Representation Programs for Retrieval Paper • 2607.18603 • Published 5 days ago • 9
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations Paper • 2607.20379 • Published 4 days ago • 4
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization Paper • 2607.10169 • Published 15 days ago • 12
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models Paper • 2607.19604 • Published 5 days ago • 14
SLPO: Scaling Latent Reasoning via a Surrogate Policy Paper • 2607.19691 • Published 4 days ago • 4
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation Paper • 2605.29502 • Published May 28 • 1
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes Paper • 2509.24945 • Published Sep 29, 2025 • 7
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training Paper • 2607.19058 • Published 5 days ago • 6
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Paper • 2607.18110 • Published 6 days ago • 14
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Paper • 2607.17524 • Published 6 days ago • 6
Distilled Reinforcement Learning for LLM Post-training Paper • 2607.17247 • Published 7 days ago • 9
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Paper • 2607.14431 • Published 11 days ago • 12