Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 12 days ago • 110
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published 15 days ago • 35