Running Featured 1.43k FineWeb: decanting the web for the finest text data at scale 🍷 1.43k Explore and download the FineWeb web‑scale text dataset
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 10 days ago • 292
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 10 days ago • 94
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics Paper • 2603.13391 • Published Mar 11 • 21
On the Design Fundamentals of Pixel Text Representation Learning Paper • 2609.01147 • Published 12 days ago • 32
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation Paper • 2608.29846 • Published 14 days ago • 16
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 11 days ago • 11
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 12 days ago • 267
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Paper • 2602.10604 • Published Feb 11 • 202
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 16 days ago • 106
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 13 days ago • 59
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published about 1 month ago • 282
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 343
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 20 days ago • 64
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction Paper • 2605.29341 • Published May 28 • 20
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 20 days ago • 207
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published 23 days ago • 63
Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses Paper • 2608.08466 • Published Aug 9 • 13
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Paper • 2608.18580 • Published 25 days ago • 121