NL2BM25: teaching Qwen2.5-3B to generate Tantivy boolean queries via SFT + GRPO. Covers reward hacking (GRPO v1) and the shaped-reward fix (GRPO v2).
Supreeth Rao
Supreeth
·
AI & ML interests
Reinforcement Learning, Large Language Models, Distributed Computing
Recent Activity
upvoted an article 8 days ago
Welcome Inkling by Thinking Machines updated a bucket 9 days ago
Supreeth/xrpo-repro-artifacts updated a bucket 9 days ago
Supreeth/xrpo-repro-qwen3-0.6b-gsm8k-bucket