Hikari07jp/Ternary-Bonsai-27B-Abliterated-LowDeg-GGUF Text Generation โข 27B โข Updated 11 days ago โข 7.85k โข 21
JBrightmanAI/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED 9B โข Updated 15 days ago โข 31 โข 1
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Paper โข 2607.07769 โข Published 25 days ago โข 10
view post Post 7724 Frontier models use distillation as a step of their post-training pipelines. In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL! See translation 3 replies ยท ๐ 14 14 ๐ฅ 7 7 โค๏ธ 3 3 + Reply
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2 Text Generation โข 12B โข Updated Jun 30 โข 238k โข 81