Papers
arxiv:2607.13511

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

Published on Jul 15
Authors:

Abstract

We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix A in R^{m times n} into A approx B diag(D) C with ternary factors B in {-1,0,+1}^{m times k}, C in {-1,0,+1}^{k times n} and a real scale vector D in R^k. The inner rank k = μmin(m,n) is deliberately expanded beyond full rank (μ> 1), so that components past full rank correct the quantization error of earlier ones. We prove the residual decreases monotonically in k and can be driven below any varepsilon > 0: ExTernD approaches bf16 accuracy arbitrarily closely, which no ternary scheme with a fixed plane count can do. Memory and compute scale continuously with μ, and factor sparsity continuously with a threshold τ, so an accuracy target is hit exactly rather than rounded to the next bit-width. ExTernD matches Q4_K's per-matrix accuracy at 5.2-5.5 effective bpw (5.1-5.5 with importance weighting) on Gemma-4-E2B and Qwen3.5-4B, and a full Qwen3.5-4B conversion at μ= 3 reaches 10.10 wikitext-2 perplexity against 9.78 for bf16 (+3.2%), placing it near the Q4_K/Q5_K accuracy band at ~5.7 effective bpw.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.13511
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.13511 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.13511 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.