CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 4 days ago • 22
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published 3 days ago • 24
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Paper • 2608.03812 • Published 3 days ago • 25
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Paper • 2608.03457 • Published 3 days ago • 27
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Paper • 2608.03979 • Published 3 days ago • 45
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 4 days ago • 79
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published 4 days ago • 74
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 3 days ago • 81