Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 3 days ago • 84
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation Paper • 2603.11137 • Published Mar 11 • 1
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths Paper • 2601.19320 • Published Jan 27 • 2
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 3 days ago • 84
Flex-Judge Collection Collections of models and papers for works: "Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators " • 4 items • Updated Aug 7, 2025 • 2
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation Paper • 2507.10524 • Published Jul 14, 2025 • 76
Flex-Judge Collection Collections of models and papers for works: "Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators " • 4 items • Updated Aug 7, 2025 • 2
DistiLLM: Towards Streamlined Distillation for Large Language Models Paper • 2402.03898 • Published Feb 6, 2024 • 5
Flex-Judge Collection Collections of models and papers for works: "Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators " • 4 items • Updated Aug 7, 2025 • 2