HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 5 days ago • 123
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published 10 days ago • 285
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Paper • 2608.10692 • Published 11 days ago • 12
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 23 days ago • 307
Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies Paper • 2607.06815 • Published Jul 7 • 5
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 311
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Paper • 2607.12625 • Published Jul 15 • 81
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published Jul 20 • 198
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras Paper • 2607.12993 • Published Jul 14 • 131
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 235
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 108
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published Jul 7 • 95
OpenRath: Session-Centered Runtime State for Agent Systems Paper • 2606.19409 • Published Jun 17 • 78
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Paper • 2606.18837 • Published Jun 17 • 58
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 172
Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? Paper • 2606.08063 • Published Jun 6 • 82
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Paper • 2606.03979 • Published Jun 2 • 29
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation Paper • 2605.22355 • Published May 21 • 178
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise Paper • 2602.12783 • Published Feb 13 • 246