Accurate Expert Predictions in MoE Inference via Cross-Layer Gate Paper • 2502.12224 • Published Feb 17, 2025 • 1
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference Paper • 2411.01433 • Published Nov 3, 2024 • 2
MoE-Infinity: Activation-Aware Expert Offloading for Efficient MoE Serving Paper • 2401.14361 • Published Jan 25, 2024 • 3
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference Paper • 2308.12066 • Published Aug 23, 2023 • 5
ExpertFlow: Optimized Expert Activation and Token Allocation for Efficient Mixture-of-Experts Inference Paper • 2410.17954 • Published Oct 23, 2024 • 1
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models Paper • 2402.07033 • Published Feb 10, 2024 • 20
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup Paper • 2511.11248 • Published Nov 14, 2025 • 1
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification Paper • 2601.07892 • Published Jan 12 • 5
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated Paper • 2407.10969 • Published Jul 15, 2024 • 24
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs Paper • 2402.04291 • Published Feb 6, 2024 • 51
ParetoQ: Scaling Laws in Extremely Low-bit LLM Quantization Paper • 2502.02631 • Published Feb 4, 2025 • 5
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks Paper • 2407.09527 • Published Jun 24, 2024 • 1
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations Paper • 2502.05003 • Published Feb 7, 2025 • 44
Codebook Configuration for 1-bit RIS-aided Systems Based on Implicit Neural Representations Paper • 2306.00544 • Published Jun 1, 2023 • 2
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models Paper • 2602.06694 • Published Feb 6 • 22
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation Paper • 2412.11741 • Published Dec 16, 2024 • 1
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 21 days ago • 151
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation Paper • 2407.13481 • Published Jul 18, 2024 • 11
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression Paper • 2306.03078 • Published Jun 5, 2023 • 4