BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 8 days ago • 24
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 6 days ago • 11
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 6 days ago • 127
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 6 days ago • 29
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 10 days ago • 145
AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data Viewer • Updated Feb 1 • 5.28k • 172 • 5
jarguello76/reinforcement_learning_lunar_landing Reinforcement Learning • Updated Aug 17, 2025 • 3 • 4
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 7 days ago • 56
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 7 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 7 days ago • 42
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 9 days ago • 44
andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 131 • 6