AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 5 days ago • 214
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 5 days ago • 37
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation Paper • 2608.25936 • Published 18 days ago • 15
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published 11 days ago • 19
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 10 days ago • 292
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 13 days ago • 153
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation Paper • 2608.30935 • Published 13 days ago • 32
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 12 days ago • 104
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 9
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry Paper • 2608.30457 • Published 13 days ago • 9
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 13 days ago • 383
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation Paper • 2608.19098 • Published 25 days ago • 23
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 19 days ago • 148
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning Paper • 2608.23318 • Published 20 days ago • 32
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling Paper • 2608.21839 • Published 22 days ago • 4
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 19 days ago • 143
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 18 days ago • 270
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 24 days ago • 111