Open SLM Leaderboard
Open Small Language Model Leaderboard
The Open SLM Leaderboard is a leaderboard for sub 150M models, our own models are on it, alot of SLM models are on it. But since the last few weeks there has been a influx of specialist models specifically targeting Arithmark 2 on it. Heres how they worked and what AxiomicLabs and we did.
Atom 2.7M Releases, 2.7M Parameters, 69.40% Arithmark 2.0. 2.7M Parameters, #6 on the entire leaderboard beating models with 50x its size. And all because of Arithmark 2, it exploited a weakness in the leaderboards ranking system.
The leaderboard defaults to ranking by Average, and because it has 69.40% In arithmark 2, its average was high too.
Other AI labs see the weakness and release models too, they get insanely high average scores with no other capabillity.
After all of this, AxiomicLabs decides to rank all specialist models at the bottomn. This should have broke them. It did. But temporarially.
Ideoa Labs realized, that because of a weakness in the specialist algorithm, if they increase their parameter count and still train on synthetic arithmetic, they dont get classified as specialist. So they release Nexus-Erebus-50M and it gets #1 for <100M. Then Nexus-Erebus-135M #1 for All.
Currently Nexus-Erebus-135M holds #1 on the leaderboard. We dont think these models should get #1 if they just are a specialist. So we just started sending PRs to fix this weakness and the first is PR 56 This PR changes the specialist classfier system to a simpler, but more effective one. If the Arithmark 2 score is more than 30% higher IN points than Hellaswag, it gets classified. This makes Nexus-Erebus-135M classified as a specialist while also adding a exclusion list for any models (none at the time) that actually have this. This PR is definitly not perfect and still leaves Nexus-Erebus-50M as none but its the first step towards fixing this problem. We will keep sending PRs soon to also fix Nexus-Erebus-50M and other models. We aren't saying these models are bad. We are saying they shouldn't be compared to normal ones.
Open Small Language Model Leaderboard