HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 24 days ago • 340
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 24 days ago • 151
PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection Paper • 2607.04690 • Published Jul 6 • 3
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
timaeus/rl-lm-pythia70m-formality-pos-beta0-grpo-nostd-gs4-tp1-tk0-pt80000-steerDotL6c64s8-seed20 Updated Jul 3 • 1
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 127