view article Article Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original MultiverseComputingCAI • 7 days ago • 41
view article Article Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers +5 ariG23498, sergiopaniego, reach-vb, pcuenq, ArthurZ, SaylorTwift, cyrilvallez • Sep 11, 2025 • 189
view article Article From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels drbh, danieldk • Aug 18, 2025 • 111
view article Article Accelerate ND-Parallel: A guide to Efficient Multi-GPU Training +3 smohammadi, siro1, winglian, marcsun13, djsaunde • Aug 8, 2025 • 100
MrezaPRZ/realmath_2025-2025-02-with-gemini-2_5-flash-context-2025-06-23-hard-cleaned Viewer • Updated Jul 25, 2025 • 105 • 11
MrezaPRZ/realmath_2025-2025-02-with-gemini-2_5-flash-context-2025-06-23-hard-cleaned Viewer • Updated Jul 25, 2025 • 105 • 11