MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 6 days ago • 32
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Paper • 2609.39071 • Published 8 days ago • 63
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization Paper • 2609.32993 • Published 12 days ago • 76
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 8 days ago • 64
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics Paper • 2609.35259 • Published 10 days ago • 198
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation Paper • 2610.02196 • Published 7 days ago • 61
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? Paper • 2609.39578 • Published 8 days ago • 68
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 9 days ago • 580
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 8 days ago • 293
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation Paper • 2609.35347 • Published 10 days ago • 180
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 14 days ago • 29
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 11 days ago • 46
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 15 days ago • 16
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 9 days ago • 42
Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents Paper • 2607.11433 • Published 14 days ago • 30
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 9 days ago • 64