RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 274
From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation Paper • 2610.02179 • Published 6 days ago • 15
Receiver-Conditioned Latent Communication gives 94% CacheBack Paper • 2609.32046 • Published 12 days ago • 11
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Paper • 2609.33463 • Published 10 days ago • 11
From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders Paper • 2610.02486 • Published 6 days ago • 11
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 11 days ago • 324
Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents Paper • 2609.37236 • Published 8 days ago • 40
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 8 days ago • 63
APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants Paper • 2609.37559 • Published 8 days ago • 46