Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
49.8
TFLOPS
Dean Byrne
PRO
Quazim0t0
19
7
71
Follow
UMCU's profile picture
gc1999xxx's profile picture
ArtelTaleb's profile picture
152 followers
·
1,572 following
https://huggingface.co/DaisyChainAI
quzi93
dean-byrne-02a28b191
AI & ML interests
DaisyChainAI🌼 / SmallLM's / San Francisco / Open Source
Recent Activity
liked
a model
about 17 hours ago
AxiomicLabs/GPT-X2.5-135M
reacted
to
AbstractPhil
's
post
with 🔥
about 21 hours ago
The AlephLM results are rolling in and I'm very excited for the possibilities. I am very much looking forward to the coming weeks as I train the first AlephLM distillations from MANY teachers into AMOE arms. The AMOE arms hook cleanly to AlephLM structures and provide pos/neg learning elements. Hard positive and hard negatives coalesce to extend the capacity. https://huggingface.co/AbstractPhil/alephlm-0 https://huggingface.co/AbstractPhil/alephlm-adopt-0 As it stands they are structurally sound enough to fully pretrain. As or more stable than a standard Bert experimentally to distill using InfoNCE. AMOE legs improve these structures substantially. Structural behavior can be expanded in many ways on distilled and pretrained models alike. Attaching the AMOE to any model I've tried has created expanded or improved behavioral accumulations. They do have downsides but their upsides are very experimentally exciting. I've distilled multiple vits, multiple berts, and have begun distilling berts into AlephLM structures successfully. This is overall very exciting for me. I've begun formatting larger variants such as including GPT-2 and Qwen 3.5 4b as a paired combinator utilizing pathological T5 learned distilled encodings. It sounds odd, but the results show everything can be expanded and even be taught to cooperate. The CaptionBert-8192-v2 and v2-b are both structurally collapsing after token 480 or so, which is expected due to the small train. By distilling an AMOE arm to V2 by training with a longformer expert, the results are cutting through like butter. V2 has begun stabilizing rapidly for considerably longer token chains and sequences, the structure is repairing and building reusable capacity. I have discovered an improved methodology for sampling the AlephLM for text encoder benchmarks, which is predominantly L2 normalized outputs. Upcoming large paper for the distillation experiments and results within the next week or two. It's going to be a big one.
updated
a model
about 22 hours ago
Quazim0t0/CS2Dream
View all activity
Organizations
Quazim0t0
's datasets
2
Sort: Recently updated
Quazim0t0/flightdeck-traces
Updated
9 days ago
•
72
Quazim0t0/Warcraft-Simulations-5x
Viewer
•
Updated
Jun 1
•
61k
•
24