SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Abstract
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.
Community
We introduce SpatialCLI, a framework that teaches vision-language models to reason with spatial tools and internalize their capabilities for tool-free inference.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning (2026)
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence (2026)
- OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping (2026)
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators (2026)
- AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models (2026)
- PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking (2026)
- Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.27703 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
ZYT-MFM/SpatialCLI-Data
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper