Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published 2 days ago • 2
MultiverseComputingCAI/Qwen3-Next-80B-A3B-Thinking-Uncensored Text Generation • 80B • Updated Jun 2 • 156 • 16