DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
Abstract
Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error-prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object-centric egocentric trajectories, each paired with a fine-grained natural-language instruction rather than a coarse verb-noun label. We further propose DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image-to-video diffusion model at an early denoising step. A lightweight flow-matching Reader decodes query-key attention tracks and pooled hidden states into relative 6-DoF poses. To our knowledge, this is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels. DreamTraj sets a new state of the art on both translation and rotation against forecasters that consume multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines.
Community
DreamTraj predicts a 6-DoF object trajectory from a single RGB image and one language instruction, by reading motion directly from the internal representations of a frozen image-to-video diffusion model at an early denoising step. The paper also introduces MOVE, 5,038 egocentric 6-DoF trajectories with language annotations.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction (2026)
- MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction (2026)
- TriMotion: Modality-Agnostic Camera Control for Video Generation (2026)
- ActWorld: From Explorable to Interactive World Model via Action-Aware Memory (2026)
- HandFlow: Fully Generative 4D Hand Recovery with Flow Matching (2026)
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report (2026)
- Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.00486 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper