PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published 7 days ago • 44
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 19 days ago • 259
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Paper • 2608.05137 • Published 10 days ago • 27
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 14 days ago • 46
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published 19 days ago • 16
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 15 days ago • 33
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 17 days ago • 59