Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
49.8
TFLOPS
Dean Byrne
PRO
Quazim0t0
19
7
71
Follow
Abyrne19's profile picture
ysdede's profile picture
UnstableLlama's profile picture
150 followers
·
1,571 following
https://huggingface.co/DaisyChainAI
quzi93
dean-byrne-02a28b191
AI & ML interests
DaisyChainAI🌼 / SmallLM's / San Francisco / Open Source
Recent Activity
liked
a model
about 4 hours ago
AxiomicLabs/GPT-X2.5-135M
reacted
to
AbstractPhil
's
post
with 🔥
about 8 hours ago
The AlephLM results are rolling in and I'm very excited for the possibilities. I am very much looking forward to the coming weeks as I train the first AlephLM distillations from MANY teachers into AMOE arms. The AMOE arms hook cleanly to AlephLM structures and provide pos/neg learning elements. Hard positive and hard negatives coalesce to extend the capacity. https://huggingface.co/AbstractPhil/alephlm-0 https://huggingface.co/AbstractPhil/alephlm-adopt-0 As it stands they are structurally sound enough to fully pretrain. As or more stable than a standard Bert experimentally to distill using InfoNCE. AMOE legs improve these structures substantially. Structural behavior can be expanded in many ways on distilled and pretrained models alike. Attaching the AMOE to any model I've tried has created expanded or improved behavioral accumulations. They do have downsides but their upsides are very experimentally exciting. I've distilled multiple vits, multiple berts, and have begun distilling berts into AlephLM structures successfully. This is overall very exciting for me. I've begun formatting larger variants such as including GPT-2 and Qwen 3.5 4b as a paired combinator utilizing pathological T5 learned distilled encodings. It sounds odd, but the results show everything can be expanded and even be taught to cooperate. The CaptionBert-8192-v2 and v2-b are both structurally collapsing after token 480 or so, which is expected due to the small train. By distilling an AMOE arm to V2 by training with a longformer expert, the results are cutting through like butter. V2 has begun stabilizing rapidly for considerably longer token chains and sequences, the structure is repairing and building reusable capacity. I have discovered an improved methodology for sampling the AlephLM for text encoder benchmarks, which is predominantly L2 normalized outputs. Upcoming large paper for the distillation experiments and results within the next week or two. It's going to be a big one.
updated
a model
about 9 hours ago
Quazim0t0/CS2Dream
View all activity
Organizations
Quazim0t0
's models
49
Sort: Recently updated
Quazim0t0/CS2Dream
Updated
about 9 hours ago
Quazim0t0/Positronic-144M-Anti-DEG
Text Generation
•
Updated
7 days ago
Quazim0t0/Escarda-86M-Identity-Anti-DEG
Text Generation
•
97.3M
•
Updated
7 days ago
•
151
•
1
Quazim0t0/Byrne-86M-Base-JL-Anti-DEG
Text Generation
•
96.9M
•
Updated
16 days ago
•
47
Quazim0t0/Escarda-86M-Base-JL-Anti-DEG
Text Generation
•
97.3M
•
Updated
16 days ago
•
23
Quazim0t0/Wheeler-DeWitt-62M
Text Generation
•
Updated
16 days ago
•
1
Quazim0t0/BE-ImgGen-Portrait-187M
Text-to-Image
•
Updated
17 days ago
Quazim0t0/Spikewhale-SNN-Brain2Qwerty
Text Generation
•
Updated
22 days ago
Quazim0t0/Byrne-DFINE-N
Object Detection
•
Updated
25 days ago
•
62
Quazim0t0/Onnx-Library
Text Generation
•
Updated
25 days ago
Quazim0t0/neural-photonic
Updated
26 days ago
•
2
Quazim0t0/Escarda-86M-Base-JL
Text Generation
•
97.3M
•
Updated
27 days ago
•
989
Quazim0t0/Byrne-86M-Base-JL
Text Generation
•
96.9M
•
Updated
27 days ago
•
957
Quazim0t0/neural-cd-preserve
Updated
27 days ago
•
1
Quazim0t0/neural-storage
Updated
27 days ago
•
2
Quazim0t0/neural-ddr
Updated
27 days ago
•
1
Quazim0t0/Chimera-64M
Text Generation
•
Updated
27 days ago
•
2
Quazim0t0/Mycel-LM-79M
Text Generation
•
Updated
27 days ago
•
5
Quazim0t0/SpikeWhale-SNN-216M
Text Generation
•
Updated
27 days ago
•
5
Quazim0t0/Positronic-144M
Text Generation
•
Updated
27 days ago
•
6
Quazim0t0/physgait-weights
Reinforcement Learning
•
Updated
28 days ago
Quazim0t0/Byrne-VE
Image Feature Extraction
•
Updated
28 days ago
Quazim0t0/Byrne-VLM-131M
Image-to-Text
•
Updated
28 days ago
Quazim0t0/Byrne-Docling-131M
Image-to-Text
•
Updated
28 days ago
•
2
Quazim0t0/Escarda-Docling-126M
Image-to-Text
•
Updated
28 days ago
•
1
Quazim0t0/Escarda-86M-Identity
Text Generation
•
97.3M
•
Updated
29 days ago
•
221
•
1
Quazim0t0/Byrne-TriAtn-86M
Text Generation
•
96.9M
•
Updated
29 days ago
•
88
•
1
Quazim0t0/Escarda-86M
Text Generation
•
97.3M
•
Updated
29 days ago
•
105
•
2
Quazim0t0/Escarda-TriAtn-86M
Text Generation
•
97.3M
•
Updated
29 days ago
•
77
•
1
Quazim0t0/Byrne-86M
Text Generation
•
96.9M
•
Updated
29 days ago
•
90
•
1
Previous
1
2
Next