Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
Abstract
Vidu S2 introduces real-time interactive avatar and video editing models that support high-resolution spatial video generation and dynamic reference updates.
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.
Community
๐ Thrilled to introduce Vidu S2: real-time AI video you can talk to, direct, and real-time editing.
๐๏ธ Vidu S2-Avatar
Interactive characters at 720p and 25โ42 FPS, with stronger instruction following, expressive full-body motion, and dancing. Introduce new reference images anytime to change outfits, interact with objects, or switch scenes mid-conversation.
๐จ Vidu S2-Editing
Transform an incoming video stream in real time while preserving motion and timing: style transfer, virtual try-on, character replacement, and background replacement.
๐ฅฝ Spatial video
Bring generated characters and live video edits into stereoscopic video for immersive VR experiences.
๐ SOTA results across all five public benchmarks in our evaluation, spanning avatar generation and video editing.
๐ฎ Try it: https://vidu.com/vidu-stream
๐ ๏ธ API: http://platform.vidu.com/vidu-stream/doc
๐ Technical report: https://arxiv.org/pdf/2609.11638
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion (2026)
- EditaLive! Unified Character Video Editing for Live Streaming (2026)
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation (2026)
- EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing (2026)
- Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds (2026)
- LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time (2026)
- InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.11638 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper