Post
84
Boris-2 coming soon!
The Models:
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th.
Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th.
Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We hope to end up in the ballpark of AxiomicLabs/GPT-X2.5-135M or BananaMind/BananaMind-2-Pro-Preview
Following this, we will release the Pro, Instruct and Pro-Instruct variants. More info will be coming soon!
The Models:
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th.
Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th.
Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We hope to end up in the ballpark of AxiomicLabs/GPT-X2.5-135M or BananaMind/BananaMind-2-Pro-Preview
Following this, we will release the Pro, Instruct and Pro-Instruct variants. More info will be coming soon!