arxiv:2609.39687
Tom Lu
eigentom
AI & ML interests
MLLM, Reinforcement Learning, Agentic RL
Recent Activity
updated a dataset about 9 hours ago
eigentom/minicpm5-swe-native-eval-archive published a dataset about 17 hours ago
eigentom/minicpm5-swe-native-eval-archive authored a paper about 24 hours ago
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation