-
shufanshen/Qwen3-4B-GRPO-DeepMath-50-steps
Reinforcement Learning • 4B • Updated • 16 -
shufanshen/Qwen3-4B-GRPO-DeepMath-100-steps
Reinforcement Learning • 4B • Updated • 14 -
shufanshen/Qwen3-4B-GRPO-DeepMath-150-steps
Reinforcement Learning • 4B • Updated • 19 -
shufanshen/Qwen3-8B-GRPO-DeepMath-50-steps
Reinforcement Learning • 8B • Updated • 17
Shufan Shen
shufanshen
AI & ML interests
Interpretable machine learning, parameter-efficient fine-tuning.
Recent Activity
upvoted a paper about 16 hours ago
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training submitted a paper about 16 hours ago
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training authored a paper 4 days ago
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training