Speedrun architecture: 8 pretraining seeds/size, anneal marks every 10 TPP as revisions. Plain ladder: “Scaling Ladder — 4 sizes, 135M–973M, 200 TPP”
Julian Minder
jkminder
AI & ML interests
None yet
Recent Activity
updated a model 3 days ago
jkminder/d20_477m_seed4_sft updated a model 3 days ago
jkminder/d20_477m_seed3_sft updated a model 3 days ago
jkminder/d20_477m_seed1_sftOrganizations
Scaling Ladder (optimized) — 3 sizes, 286M–897M, 200 TPP
Speedrun architecture: 8 pretraining seeds/size, anneal marks every 10 TPP as revisions. Plain ladder: “Scaling Ladder — 4 sizes, 135M–973M, 200 TPP”
Scaling Ladder — 4 sizes, 135M–973M, 200 TPP
The whole ladder, 200 TPP on ClimbMix; plain GPT (nanochat mechanisms disabled). Bases, then chat-SFTs (_sft).
Pretraining priors: pirate-2x2 planted prior (d26)
d26 models with the pirate-2x2 corpora inserted in pretraining: exp-056 anchor + exp-074 dose/window sweep, base and instruction-SFT pairs.
Lorentz-Forcing-EM