Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published Aug 20 • 87
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published Aug 17 • 35