KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published 8 days ago • 42
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 10 days ago • 83
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published 9 days ago • 107
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 12 days ago • 70
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 9 days ago • 102
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Paper • 2607.07608 • Published 15 days ago • 56
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 15 days ago • 64
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Paper • 2607.02963 • Published 20 days ago • 28
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Paper • 2607.05691 • Published 17 days ago • 4
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation Paper • 2607.02642 • Published 21 days ago • 38
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning Paper • 2606.31825 • Published 23 days ago • 29
Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training Paper • 2607.01763 • Published 21 days ago • 10
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Paper • 2607.01191 • Published 22 days ago • 19
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising Paper • 2607.00407 • Published 22 days ago • 10
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published 27 days ago • 111
RedVox: Safety and Fairness Gaps in Speech Models Across Languages Paper • 2606.26968 • Published 28 days ago • 16
ReFreeKV: Towards Threshold-Free KV Cache Compression Paper • 2502.16886 • Published 27 days ago • 48
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do Paper • 2606.22565 • Published Jun 21 • 9