Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 7 days ago • 59
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem Paper • 2608.11654 • Published Aug 12 • 1
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Paper • 2507.06892 • Published Jul 9, 2025 • 1
Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid Algorithms Paper • 2401.11963 • Published Jan 22, 2024 • 2
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 24 days ago • 144
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 171
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 174