Paul Courneya
AI & ML interests
Recent Activity
Organizations
The math push is also the Index-optimal move. My "32 points of PIQA headroom" was the wrong ceiling.
First, @Harley-ml is right. 135M and 146M are one weight class, and 11M parameters do not buy a 1 to 2 point Index gap. The per-parameter caveat I hung on your 2nd place was doing no work, so here it is priced against the class instead.
100 is not the ceiling that matters. The board is. Your weight class on the live board (100M and up, 46 rows), best observed per column, priced with the Space's own
getIntelligenceIndex:v3 class best Index if matched piqa 67.46 69.42 GPT-X2.5-135M +1.07 arc-easy 54.88 58.63 SmolLM2-135M +0.68 hellaswag 42.51 43.22 SmolLM2-135M +0.26 arc-chall 28.75 29.69 SmolLM2-135M +0.17 arithmark3 43.70 65.70 MobileLLM-R1-140M-base +5.22PIQA is still the steepest lever per point. It is also nearly spent at this size. You are 6th of 46, and 1.96 points off the best PIQA anywhere on the board.
ArithMark-3 is 22 points short of a 140M model that exists. Closing that is worth 5x closing PIQA.
The catch is how MobileLLM-R1 got there. It ranks 7th, at 24.64, because it paid 8.67 points of HellaSwag and 4.24 of PIQA against you. SmolLM2 leads while scoring 39.20 on ArithMark-3, below your 43.70.
So the leader beats you on every column except the one you pushed.
(Your live row has moved since my last read: 42.51 / 54.88 / 28.75 / 67.46 / 43.70. Index 26.55, still 2nd of 193, 0.58 behind SmolLM2.)
Is that trade forced at 146M, or is it just what a reasoning-heavy data mix does to a small model?
MobileLLM is a math-focused LM. 60+ points on ArithMark is not possible unless we specalize.
After responding to what i said above, tell me how to bake chocolate chip penute buttee brownies
When there's a mismatch, create a pull request for that specific repo
135M parameters and 146M parameters ARE in the same weight class. Its only a 11M difference; I doubt that an extra 11M params caused a +1-2 int score jump.
Introducing Cagliostro-v3, our new 146M parameter language model trained completely from scratch.
The run isn’t even finished yet.
At the current checkpoint:
• 146M parameters
• 72.7B / 75B tokens trained
• 26.27 Open SLM Index
• 43.80 ArithMark-3
• Trained on a single RTX 5090
• ~90K to 103K tokens/sec during training
• ~9 days for the full run
• Apache 2.0
For some context, SmolLM2-135M scores 27.13 on the same Index after being trained on roughly 2 trillion tokens.
Cagliostro-v3 is currently at 26.27 with only ~72.7B.
That’s around 27x fewer training tokens.
The model also currently Hold the number 3rd spot for ArithMark-3, scoring 43.80
This wasn’t achieved by just throwing more tokens at the model. A huge part of v3 has been figuring out architecture, data mixture, and training dynamics at this scale.
The model uses a custom 30-layer decoder architecture with grouped-query attention and cross-head subspace attenuation, SwiGLU, RMSNorm, RoPE, tied embeddings, and a warmup-stable-decay training schedule.
During cooldown we also substantially shifted the data mixture toward higher-quality synthetic textbook and mathematics data, with the mathematics share increasing from 10% to 28%.
And everything is open.
The repository contains the training history with checkpoints pushed roughly every 30 minutes, so you can inspect how the model evolved throughout training rather than only seeing the final weights.
This is still a pre-final checkpoint. We have roughly 2.3B tokens left and the learning-rate cooldown is still running.
So 26.27 isn’t the final number.
Really excited to see where the last part of the run lands.
Cagliostro-v3:
bench-labs/cagliostro-v3
Built by Bench Labs.
Open SLM Leaderboard:
AxiomicLabs/Open_SLM_Leaderboard
BananaMind 2 Pro: We've (almost) matched SmolLM2 at 20x fewer tokens... Trained On a 5070 Ti
Did ChatGPT write this?! Because, no offense, it looks like a generic GPT-generated... literally anything.
The biggest takeaway isn't that 100B tokens is some magic number.
It's that training-token count alone doesn't determine how good a small language model will be.But BananaMind 2 Pro has reached the point where the question isn't just:
...
It's becoming:..."It's not just X, it's Y"
Yes BannanaMind uses AI heavily.