StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 24 days ago • 446
view article Article NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset nvidia • Aug 20, 2025 • 19
On-Policy Adversarial Flow Distillation for Autoregressive Video Generation Paper • 2605.26105 • Published May 25 • 20
Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated 28 days ago • 191
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models Paper • 2511.16668 • Published Nov 20, 2025 • 56
Tulu 2 Llama 3 Update Collection Llama 3 models trained on the tulu dataset, following https://arxiv.org/abs/2311.10702 (tulu 2) and https://arxiv.org/abs/2406.09279 (tulu 2.5). • 12 items • Updated Mar 4, 2025 • 2