Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
49.3
TFLOPS
Syafiq Kamarul Azman
syaffers
1
5
Follow
MishaGGG's profile picture
rodoviario's profile picture
NILKNARFGonzo's profile picture
3 followers
·
10 following
https://syaffers.xyz
syaffers
syafiqkamarulazman
AI & ML interests
Inference and model servers
Recent Activity
liked
a model
4 days ago
tiny-random/qwen2.5-omni
liked
a model
4 days ago
hmellor/tiny-random-LlamaForCausalLM
reacted
to
Banaxi-Tech
's
post
with 🔥
5 days ago
Today we wanted to release BananaMind 2 Pico, our smallest model yet at ~0.9M parameters. Instead, we accidentally ran a very expensive experiment on what happens when you push a tiny model way past its useful token budget. Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 — a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through. The useful compute-optimal range for Pico-tier models looks like ~22K–30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine — TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting. Follow us for more: https://huggingface.co/BananaMind @vovaRL @Banaxi-Tech Full writeup with all checkpoints, the Chinchilla-ratio control run, and the schedule-vs-overtraining analysis: https://huggingface.co/blog/Banaxi-Tech/ovdadadadd And if anyone, i dont know the reason why you would, wants the 20B token checkpoint reply and ill upload it as BananaMind 2.1 Pico EXP
View all activity
Organizations
None yet
models
5
Sort: Recently updated
syaffers/gemma-3-4b-it-NVFP4
Image-Text-to-Text
•
4B
•
Updated
13 days ago
•
65
syaffers/Atom-350M-GGUF
Text Generation
•
0.4B
•
Updated
23 days ago
•
261
syaffers/tiny-random-storywriter-base
Updated
Jun 4
syaffers/Atom-350M-NVFP4
Text Generation
•
0.2B
•
Updated
Jun 3
•
12
syaffers/tiny-random-llama-lora
Text Generation
•
Updated
Dec 24, 2025
•
18
datasets
1
syaffers/yepyep
Updated
Aug 22, 2025
•
9