OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published 18 days ago • 72
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 21 days ago • 84
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation Paper • 2605.25874 • Published May 25 • 106
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening Paper • 2605.19597 • Published May 19 • 21