Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails Paper • 2609.09134 • Published 11 days ago • 10
DriveZero: End-to-End Driving Beyond Human Demonstrations Paper • 2609.06055 • Published 14 days ago • 56
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection Paper • 2609.07670 • Published 12 days ago • 18
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal Paper • 2609.04482 • Published 16 days ago • 11
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 17 days ago • 61
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 18 days ago • 89
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Paper • 2608.27456 • Published 23 days ago • 86
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Paper • 2609.01404 • Published 18 days ago • 28
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered Paper • 2608.29464 • Published 21 days ago • 15
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 24 days ago • 59
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 24 days ago • 88
4DAnyone: Create Anyone in 4D from a Casual Monocular Video Paper • 2608.20335 • Published 30 days ago • 84
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 30 days ago • 86
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published 25 days ago • 30