This is very good to know! I appreciate you taking the time and effort to study this and release it to the community
Joseph Jones PRO
KlondikeDev
AI & ML interests
Restoring the Digital Millennium. https://kunix.org/digital-millennium.html
Recent Activity
commentedon an article about 1 hour ago
Extreme Overtraining in Tiny Language Models new activity about 1 hour ago
opencerebral/Lucy-2:So cool updated a collection about 2 hours ago
LucyOrganizations
commented on Extreme Overtraining in Tiny Language Models about 1 hour ago
So cool
2
#1 opened about 1 hour ago
by
Datdanboi25
posted an update about 2 hours ago
Post
9
Lucy out now!
Lucy is a series of Image Generation models created by OpenCerebral.
Lucy: Prototype, faces only -- unconditional, no prompts.
Lucy-2: Text-to-image, general purpose (though still small)
More Lucy models will be arriving in the future!
opencerebral/Lucy
opencerebral/Lucy-2
Lucy is a series of Image Generation models created by OpenCerebral.
Lucy: Prototype, faces only -- unconditional, no prompts.
Lucy-2: Text-to-image, general purpose (though still small)
More Lucy models will be arriving in the future!
opencerebral/Lucy
opencerebral/Lucy-2
replied to their post 3 days ago
That's not a bad idea -- I was planning on 4096 context, but I might scale up, there's not a reason not to
replied to their post 3 days ago
That's not too far off from our original plan, which was FineWeb-Edu, Cosmopedia V2, and DCLM, then some FineMath, but I will definitely look into those other ones, I've been wanting to find some better datasets
OpenCerebral
6
#6 opened 4 days ago
by
KlondikeDev
OpenCerebral
4
#5 opened 4 days ago
by
KlondikeDev
posted an update 4 days ago
Post
2131
Boris-2 coming soon!
The Models:
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th.
Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th.
Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We hope to end up in the ballpark of AxiomicLabs/GPT-X2.5-135M or BananaMind/BananaMind-2-Pro-Preview
Following this, we will release the Pro, Instruct and Pro-Instruct variants. More info will be coming soon!
The Models:
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th.
Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th.
Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We hope to end up in the ballpark of AxiomicLabs/GPT-X2.5-135M or BananaMind/BananaMind-2-Pro-Preview
Following this, we will release the Pro, Instruct and Pro-Instruct variants. More info will be coming soon!
upvoted an article 4 days ago
Article
TutorMoments: Do AI tutors know when to help and when to hold back?
allenai
• • 29Making SLMs is so fun! I really enjoy it, personally. Given how little compute is needed, I think everybody should try it!