OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 16 days ago • 253
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 6 days ago • 10
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published 11 days ago • 44
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Paper • 2608.01851 • Published 14 days ago • 12
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 11 days ago • 40
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 17 days ago • 96
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Paper • 2608.02603 • Published 14 days ago • 34
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published 18 days ago • 37