Running Featured 789 Agent Memory Leaderboard 🧠 789 Unified memory evaluation · Results expected August 12.
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 7 days ago • 207
TempCloze: Can Video-LLMs Identify the Missing Middle? Paper • 2609.01515 • Published 16 days ago • 32
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 151
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 312
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 315
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published Jul 23 • 157