Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Paper • 2607.15263 • Published 6 days ago • 5
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 7 days ago • 95
PalmClaw: A Native On-Device Agent Framework for Mobile Phones Paper • 2607.13027 • Published 9 days ago • 14
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 9 days ago • 209
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published 9 days ago • 107
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents Paper • 2607.08093 • Published 14 days ago • 5
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering Paper • 2603.28583 • Published 9 days ago • 17
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness Paper • 2607.08046 • Published 14 days ago • 14
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published 8 days ago • 85
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 9 days ago • 102
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Paper • 2607.07740 • Published 15 days ago • 23
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 15 days ago • 64
Autonomous Scientific Discovery via Iterative Meta-Reflection Paper • 2607.01131 • Published 22 days ago • 8
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models Paper • 2607.05803 • Published 16 days ago • 10
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training Paper • 2607.05804 • Published 16 days ago • 18
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation Paper • 2607.02642 • Published 21 days ago • 38
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Paper • 2607.01642 • Published 21 days ago • 39
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents Paper • 2606.20023 • Published Jun 18 • 5