On the Generalization Capacities of MLLMs for Spatial Intelligence Paper • 2603.06704 • Published Mar 5
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 6 days ago • 298
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance Paper • 2510.22684 • Published Oct 26, 2025
GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation Paper • 2509.15733 • Published Sep 19, 2025
Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining Paper • 2401.08407 • Published Jan 16, 2024
MMRel: A Relation Understanding Dataset and Benchmark in the MLLM Era Paper • 2406.09121 • Published Jun 13, 2024 • 1
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Paper • 2607.14614 • Published 20 days ago • 13
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 6 days ago • 298 • 7
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 6 days ago • 298
Scaling Language-Centric Omnimodal Representation Learning Paper • 2510.11693 • Published Oct 13, 2025 • 109
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework Paper • 2411.06176 • Published Nov 9, 2024 • 45