HARP Quantized Models Quantized Llama 2 (7B/13B/70B) at 2-bit via HARP, a learnable orthogonal preprocessor for extreme LLM quantization. EMNLP 2026. brain-lab/Llama-2-7b-QuIP-HARP-2Bit Text Generation • 0.6B • Updated 11 days ago • 160 brain-lab/Llama-2-13b-QuIP-HARP-2Bit Text Generation • 0.9B • Updated 11 days ago • 154 brain-lab/Llama-2-70b-QuIP-HARP-2Bit Text Generation • 3B • Updated 11 days ago • 159
Gradient-Faithful Surrogates paper models Final checkpoints of Llama3 and Qwen3 models from Gradient-Faithful Surrogates for KV-Cache Quantization Recovery paper brain-lab/kivi-llama-gradient-faithful 8B • Updated 22 days ago • 8 brain-lab/qjl-llama-gradient-faithful 8B • Updated 22 days ago • 15 brain-lab/kivi-qwen3-gradient-faithful 4B • Updated 22 days ago • 13 brain-lab/qjl-qwen3-gradient-faithful 4B • Updated 22 days ago • 10
HARP Quantized Models Quantized Llama 2 (7B/13B/70B) at 2-bit via HARP, a learnable orthogonal preprocessor for extreme LLM quantization. EMNLP 2026. brain-lab/Llama-2-7b-QuIP-HARP-2Bit Text Generation • 0.6B • Updated 11 days ago • 160 brain-lab/Llama-2-13b-QuIP-HARP-2Bit Text Generation • 0.9B • Updated 11 days ago • 154 brain-lab/Llama-2-70b-QuIP-HARP-2Bit Text Generation • 3B • Updated 11 days ago • 159
Gradient-Faithful Surrogates paper models Final checkpoints of Llama3 and Qwen3 models from Gradient-Faithful Surrogates for KV-Cache Quantization Recovery paper brain-lab/kivi-llama-gradient-faithful 8B • Updated 22 days ago • 8 brain-lab/qjl-llama-gradient-faithful 8B • Updated 22 days ago • 15 brain-lab/kivi-qwen3-gradient-faithful 4B • Updated 22 days ago • 13 brain-lab/qjl-qwen3-gradient-faithful 4B • Updated 22 days ago • 10