Activity Feed

AI & ML interests

None defined yet.

Recent Activity

arthu1ย  updated a Space 9 days ago
north-previews/README
arthu1ย  published a Space 9 days ago
north-previews/README
View all activity

Banaxi-Techย 
posted an update about 3 hours ago
view post
Post
75
We're excited to release BananaMind 2 Pro, our final version of the Pro model.
Trained on 100B tokens it performs extremely good for its token and size class.
The training took 22 days on one RTX 5070 Ti.
Check it out at
BananaMind/BananaMind-2-Pro
We did not release a Chat version yet because it regressed. Release Later.
Follow us to know when BananaMind 2 Ultra releases and support us at
BananaMind

@Banaxi-Tech
@vovaRL
@DedeProGames
  • 2 replies
ยท
Banaxi-Techย 
posted an update 2 days ago
view post
Post
2681
Today we wanted to release BananaMind 2 Pico, our smallest model yet at ~0.9M parameters. Instead, we accidentally ran a very expensive experiment on what happens when you push a tiny model way past its useful token budget.

Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 โ€” a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through.

The useful compute-optimal range for Pico-tier models looks like ~22Kโ€“30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine โ€” TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting.

Follow us for more:
BananaMind

@vovaRL
@Banaxi-Tech


Full writeup with all checkpoints, the Chinchilla-ratio control run, and the schedule-vs-overtraining analysis: https://huggingface.co/blog/Banaxi-Tech/ovdadadadd


And if anyone, i dont know the reason why you would, wants the 20B token checkpoint reply and ill upload it as BananaMind 2.1 Pico EXP
  • 3 replies
ยท
Banaxi-Techย 
posted an update 3 days ago
view post
Post
1865
We're excited to release BananaMind 2 SLMoE, an experimental sequence-level mixture-of-experts model.
It uses only 8M parameters per message but has 25M total parameters, 13 experts (out of 64) are selected based on the message prefix and reused for the entire response.
We're testing with this sequence-level architecture to find out how big the capability loss actually is and how much of it can be fixed.
The long-term idea is that this could make very large sparse models usable on machines that can't fit them in RAM by putting the entire model (which is big) on disk and only loading the active parts into VRAM.
This architecture is still in research and shouldn't be used for production models.


We trained it on 60B tokens (of FineWeb-HQ, FineWeb-Edu, DCLM ,Cosmopedia v2, FineMath and NPSet-2) on 8 RTX Pro 6000s.

Check it out at BananaMind/BananaMind-2-SLMoE
Follow us for future models:
BananaMind

@vovaRL
@Banaxi-Tech
@DedeProGames

BananaMind 2 Pro in a few days. You've been waiting 22 days for it.
  • 3 replies
ยท
LH-Tech-AIย 
posted an update 3 days ago
view post
Post
2125
Hey community and sponsors!

We are announcing the Supra3 family with four core SLM models:
- Supra3 Flash Lite: 25M parameters, ~60B pretraining tokens
- Supra3 Flash: 50M parameters, ~100B pretraining tokens
- Supra3 Pro: 75M parameters, ~150B pretraining tokens
- Supra3 Ultra: 100M parameters, ~200B pretraining tokens

For Supra3 Pro and Ultra, we search for sponsors who give us free access to compute like RTX 5090 32GB or so.

We estimate the total cost of the pro and ultra models at around $600.

For Supra3 Flash Lite and Flash, we do not need sponsors.

If anyone would apply for helping us, we would be really thankful and this person would get early access to new modele, insider information, credit and more!

Contact: here or on discord: lh_tech_ai
  • 14 replies
ยท
Banaxi-Techย 
posted an update 4 days ago
view post
Post
2800
We're excited to release BananaMind 2 Micro, our smallest model yet.
It fits a compact architecture in only 2.9M parameters achieving the highest parameter efficiency on BananaMind Base Bench against comparable models.
It achieves comparable performance to GPT S2 5M and GPT S 5M at almost half the size while beating CMA 1M Mini.
BananaMind 2 Micro achieved the #1 spot on the Open SLM Leaderboard for the sub 3M category (not added yet but it achieves #1)
For the training we used Muon + the XSA Refresh Gate with a 5e-2 lr for Muon and 4e-3 for the 1D weights.
Its score on our efficiency measure is 0.326 getting the first place with Syn 2.6M on the second place scoring 0.291 and GPT S 5M at 0.235*
Check it out at BananaMind/BananaMind-2-Micro and follow us at:
@vovaRL
@DedeProGames
@Banaxi-Tech
BananaMind


Our new releases aren't stopping ๐Ÿš€ August 13-14 BananaMind 2 Pro
  • 9 replies
ยท
Dorman11ย 
posted an update 5 days ago
view post
Post
125
I've been thinking this for a long time, other than OpenMythos will we see similar type of open source models be released? The more and more top frontier models get good at finding zero day vulnerabilities and security weaknesses how will the open source model community respond. My take is we will see a open source model very soon be as good at finding security holes as the top models, at some point in the future. Is it just me or anyone seeing this?
Banaxi-Techย 
posted an update 5 days ago
view post
Post
2684
Today.
  • 9 replies
ยท
Banaxi-Techย 
posted an update 6 days ago
view post
Post
1945
We're exited to announce BananaMind OS, our OS specically for running BananaMind models!
Its able to run BananaMind 2 Nano at 4 bit on only 7-8MB of ram, the 2 bit on 6MB of ram and the 8 bit version on 14MB of RAM!
It runs on a 486 or newer!
Check out this video and image running BananaMind 2 Nano 4 Bit on 9
MB of RAM and a emulated 486 in QEMU at ~1TPS!
We asked it: "What is the first letter of the alphabet?"
The response is:
"The first letter of the alphabet is:
- A.
"
And if you're asking because of the video, yes I am a arch btw.
Comment and like this post for a GitHub link and comment for adding other models!
  • 25 replies
ยท
Banaxi-Techย 
posted an update 7 days ago
arthu1ย 
updated a Space 9 days ago
arthu1ย 
published a Space 9 days ago
Banaxi-Techย 
posted an update 9 days ago
view post
Post
168
u guys want bananamind 2 ultra?
reply for bananamind 2 ultra want
  • 1 reply
ยท
Banaxi-Techย 
posted an update 10 days ago
view post
Post
3239
We did an experiment, we wanted to see if AI is good enough to train models.
We used GPT 5.6 Sol Max for this because its one of the most powerful ones right now.
Our instructions were, it should write the training code, and start the training process and monitor it by itself.
We also gave it a link to BananaMind 2 Mini to get our architecture right.
The result: It worked, it made the working BananaMind 2 Nano, and even beat our previous MiniBananaMind v4 9M.
Its getting way easier to develop your own models now!
  • 8 replies
ยท
LH-Tech-AIย 
posted an update 11 days ago
Banaxi-Techย 
posted an update 11 days ago
view post
Post
2319
We're excited to release BananaMind 2 Pro Preview, our best model yet.
Trained on ~52B tokens it performs extremely good for its token and size class.
We trained it on a single 5070 Ti in about 11 days.
Check it out at BananaMind/BananaMind-2-Pro-Preview.
Sadly we need to delay BananaMind 2 Micro until the launch of the final BananaMind 2 Pro.
We will release the final checkpoint with 100B tokens in ~11 days.
Go and fine-tune it!
We've also released BananaMind 2 Pro Preview Chat which is the instruct version of it!
BananaMind/BananaMind-2-Pro-Preview-Chat

Follow us to know when the final releases and support us at
BananaMind

@Banaxi-Tech
  • 5 replies
ยท
AxionLab-officialย 
posted an update 12 days ago
view post
Post
3462
Tommorow we are releasing Supra2-Pro(100M params), our flagship, get ready!

(follow SupraLabs to get tuned in!)
SupraLabs
  • 9 replies
ยท
Banaxi-Techย 
posted an update 12 days ago
view post
Post
2688
BananaMind 2 pro Has BEEN RELEASED! BananaMind/BananaMind-2-Pro-Preview


Previous content:


BananaMind 2 Pro Preview will launch tomorrow.
Give us a follow:
BananaMind

Lets get 70 or 75 followers before it releases.
It takes 5 seconds.
August 3, 1PM in Austria time
  • 21 replies
ยท
Banaxi-Techย 
posted an update 14 days ago
view post
Post
1811
BananaMind 2 Pro Preview will release when we hit 75 followers on BananaMind!
Follow us for the release.
We only need 13 more
BananaMind

@Banaxi-Tech
On August 3 (preview date) we will be at 90k-100k
Early Access at
BananaMind-Model-Previewers
if your known in the community
The benchmarks for 80k are very good
  • 1 reply
ยท
LH-Tech-AIย 
posted an update 14 days ago
view post
Post
2394
Announcing The Supra2 Family And Supra2-100M

Today, we are announcing a brand-new series of SupraLabs models: Supra2
This series will feature various models, including such as:
- ๐Ÿœ Supra2-Nano (0.4M) โ†’ The smallest Supra2 model.
- ๐Ÿค Supra2-Small (1.4M) โ†’ The tiny model that runs everywhere.
- ๐Ÿ’ช Supra2-Medium (25M) โ†’ Our medium class model in the Supra2 family. The powerful midsizer.
- ๐Ÿ”ฅ Supra2-Pro (100M): base, instruct, reasoning, code, math and more! โ†’ The most capable model yet! A real allrounder for all your everyday tasks.
- ๐ŸŽจ Supra2-IMG โ†’ our generative text-to-image model
...and many more...

Current progress:
- Nano (0.4M) and Small (1.4M): in training; almost done. Baseline set.
- Medium (25M): coming soon...
- Pro (100M): in training; finishes in 66 hours - Monday, 3rd August 2026, 12:00AM
- IMG: coming soon...

You can support us with a like and follow if you want!
Don't miss our next release! Stay tuned...
  • 30 replies
ยท
AxionLab-officialย 
posted an update 14 days ago
view post
Post
2248
Thanks for 300 followers in SupraLabs!!! That means alot to SupraLabs!



@lesageethan
@LH-Tech-AI
  • 3 replies
ยท