Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published 10 days ago • 118
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published 14 days ago • 52
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization Paper • 2608.12314 • Published 29 days ago • 27
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published Aug 3 • 93
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 105
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 77
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published Jul 16 • 75
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published Jul 18 • 141
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 89
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Paper • 2606.19531 • Published Jun 17 • 25
DreamX-World 1.0: A General-Purpose Interactive World Model Paper • 2606.16993 • Published Jun 15 • 116
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 217
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 85