Checkpoints for Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
-
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Text Generation • 2B • Updated • 456k • 1.57k -
Qwen/Qwen3-4B
Text Generation • 4B • Updated • 6.28M • • 691 -
deepseek-ai/DeepSeek-R1-Distill-Llama-8B
Text Generation • 8B • Updated • 374k • • 875 -
polaris-73/DeepSeek-R1-Distill-Qwen-1.5B-RLVR-Science-step-100
2B • Updated • 7