Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation Paper • 2610.05608 • Published 6 days ago • 160
view article Article FLUX 3 Action: a world action model you can fine-tune black-forest-labs • 16 days ago • 17
LTX-2.5 Collection LTX-2.5 base models, quantized models and accompanying LoRAs and IC-LoRAs • 5 items • Updated about 1 month ago • 70
KVAE: Family of Tokenizers for Multimodal Generative Models Paper • 2608.05798 • Published Aug 6 • 31
Kandinsky WM 1.0 Collection Image-to-Video models for Physical AI: autonomous driving, robotics, general physics. • 3 items • Updated 4 days ago • 6
Laguna S 2.1 Collection Our most capable model to date, designed for long-horizon work. • 13 items • Updated Aug 3 • 51
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation Paper • 2607.02515 • Published Jul 2 • 21
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Paper • 2606.10029 • Published Jun 8 • 12
A Geometric Account of Activation Steering through Angle-Norm Decomposition Paper • 2606.06735 • Published Jun 4 • 27
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders Paper • 2606.07473 • Published Jun 5 • 15
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 91
KVAE 2.0 Collection KVAE 2.0 is a family of image and video tokenizers with a time compression ratio of 4 and spacial compression ratio of 8 and 16 • 3 items • Updated 4 days ago • 5
Interpreting CLIP with Hierarchical Sparse Autoencoders Paper • 2502.20578 • Published Feb 27, 2025 • 1