view post Post 1924 š The reasoning backbone quadruples from 8B to 32B , while the action expert remains at 2.3B! š We took a closer look at the architectural evolution from nvidia/Alpamayo-1.5-10B to nvidia/Alpamayo2-Super .Read the analysis here:https://huggingface.co/blog/JonnaMat/alpamayo2-superOur analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. š§ See translation 3 replies Ā· š„ 4 4 + Reply
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper ⢠2608.01964 ⢠Published 5 days ago ⢠158
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking Paper ⢠2306.05426 ⢠Published Jun 8, 2023 ⢠1
IvmeLabs/ExpIvme-DiffusionConversate-v1-Instruct Text Generation ⢠0.1B ⢠Updated 3 days ago ⢠347 ⢠1