Instructions to use ddz16/CamDistill-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ddz16/CamDistill-8B with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ddz16/CamDistill-8B") model = AutoModelForMultimodalLM.from_pretrained("ddz16/CamDistill-8B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CamDistill-8B
Camera-movement understanding model trained with Camera Token Distillation on top of
Qwen/Qwen3-VL-8B-Instruct. A lightweight Camera Token Module learns geometry-aware camera
tokens (distilled from VGGT) and injects them into the language model. Given a video, it outputs
structured JSON describing every camera-movement segment.
- Paper: Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
- Project page: https://ddz16.github.io/cammotion.github.io
- Code: https://github.com/ddz16/CamDistill
⚠️ This model cannot be loaded with plain 🤗 Transformers. It contains an extra Camera Token Module and a patched forward pass. Loading it as a standard
Qwen3VLForConditionalGenerationwould silently drop those weights and produce incorrect results. Use the CamDistill repo, which registers the required custom model type through a plugin.
Usage
Clone the CamDistill repo, then run (camera tokens are generated internally — no online VGGT required):
python camera_movement_sft/infer_single.py \
--model ddz16/CamDistill-8B \
--video /path/to/video.mp4 \
--variant camdistill
See the repo's README for environment setup and batch evaluation.
- Downloads last month
- 56
Model tree for ddz16/CamDistill-8B
Base model
Qwen/Qwen3-VL-8B-Instruct