Kshitij Thakkar PRO
AI & ML interests
Recent Activity
Organizations
buckets 27
Articles 6
The Mind of Tashi: making a 200M-active model's reasoning *the game*
Scaling Mixture of Experts: Architecture Search for Billion-Parameter Language Models
- Running
Repro - MemEvolve: Meta-Evolution of Agent Memory Systems
🧬Collaborate with an AI agent to manage a shared experiment logbook
- Running
Repro - TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
🎯Explore experiment logs and sync findings with a coding agent
- Running
Repro - How to Correctly Report LLM-as-a-Judge Evaluations
🎯Log and share LLM evaluation findings in a collaborative notebook
- Running
Repro - Dependence-Aware Label Aggregation via Ising Models
🧲Explore and edit experiment logbooks with AI agent help
-
kshitijthakkar/Kirigami-Qwen3.6-20B-A3B-NVFP4
Text Generation • 14B • Updated • 70 • 1 -
kshitijthakkar/Kirigami-Qwen3.6-24B-A3B-NVFP4
Text Generation • 16B • Updated • 276 -
kshitijthakkar/Kirigami-Qwen3.6-28B-A3B-NVFP4
Text Generation • 19B • Updated • 20 • 1 - Running
Kirigami Journey
🪷How we carved a 35B MoE to fit a 24GB GPU — zero training
- Running
Repro - MemEvolve: Meta-Evolution of Agent Memory Systems
🧬Collaborate with an AI agent to manage a shared experiment logbook
- Running
Repro - TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
🎯Explore experiment logs and sync findings with a coding agent
- Running
Repro - How to Correctly Report LLM-as-a-Judge Evaluations
🎯Log and share LLM evaluation findings in a collaborative notebook
- Running
Repro - Dependence-Aware Label Aggregation via Ising Models
🧲Explore and edit experiment logbooks with AI agent help
-
kshitijthakkar/Kirigami-Qwen3.6-20B-A3B-NVFP4
Text Generation • 14B • Updated • 70 • 1 -
kshitijthakkar/Kirigami-Qwen3.6-24B-A3B-NVFP4
Text Generation • 16B • Updated • 276 -
kshitijthakkar/Kirigami-Qwen3.6-28B-A3B-NVFP4
Text Generation • 19B • Updated • 20 • 1 - Running
Kirigami Journey
🪷How we carved a 35B MoE to fit a 24GB GPU — zero training
spaces 41
GuardianTails
Pet Health Intelligence Platform
2604.20098
Explore and collaborate on your project logbook
Repro - Finite and Corruption-Robust Regret Bounds in Online Inverse Linear Optimization under M-Convex Action Sets
Explore and collaborate on a research logbook online
Repro - TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Explore experiment logs and sync findings with a coding agent
Repro - Gaussian Mean Field Variational Inference can Overestimate Predictive Variance
Explore research logbook and sync with your coding agent
Reproducing GRPO's Loss, Dynamics, and Success Amplification (arXiv:2503.06639)
Browse a research logbook and collaborate with an AI agent