Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency Paper • 2610.04318 • Published 6 days ago • 18
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs Paper • 2508.18264 • Published Aug 25, 2025 • 25
nyu-dice-lab/VeriThoughts-Reasoning-7B Text Generation • 8B • Updated May 13, 2025 • 40 • • 3
Running 4.07k The Ultra-Scale Playbook 🌌 4.07k The ultimate guide to training LLM on large GPU Clusters