MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 7 days ago • 32
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 8 days ago • 275
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 10 days ago • 138
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 9 days ago • 64
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 13 days ago • 325
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 10 days ago • 42
AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data Viewer • Updated Feb 1 • 5.28k • 251 • 11
imflash217/proximal_policy_optimization_huggy_unity Reinforcement Learning • Updated Jan 14, 2023 • 55 • 3
imflash217/proximal_policy_optimization_lunar_lander_v2 Reinforcement Learning • Updated Jan 13, 2023 • 2 • 4
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents Paper • 2609.33848 • Published 12 days ago • 46
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue Paper • 2609.31948 • Published 14 days ago • 89
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 15 days ago • 13
Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors Paper • 2609.24385 • Published 18 days ago • 24
crumb/bloom-560m-RLHF-SD2-prompter-aesthetic Text Generation • 0.6B • Updated Mar 19, 2023 • 326 • 27
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 16 days ago • 14