RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 262
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models Paper • 2609.39820 • Published 7 days ago • 14
Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers Paper • 2608.26762 • Published 11 days ago • 16
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Paper • 2609.33463 • Published 10 days ago • 11
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence Paper • 2609.34563 • Published 9 days ago • 216
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? Paper • 2609.39578 • Published 7 days ago • 67
Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation Paper • 2609.38660 • Published 8 days ago • 36
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training Paper • 2609.40111 • Published 7 days ago • 50
RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement Paper • 2609.39045 • Published 7 days ago • 90
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models Paper • 2609.38827 • Published 7 days ago • 60
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering Paper • 2609.38177 • Published 8 days ago • 71