The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Paper • 2609.18063 • Published 8 days ago • 18
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 10 days ago • 50
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Paper • 2609.06490 • Published 18 days ago • 9
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 21 days ago • 113
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published 16 days ago • 70
ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation Paper • 2609.03756 • Published 21 days ago • 25
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 28 days ago • 40
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss Paper • 2609.00591 • Published 23 days ago • 18