Inspect model quantization in the browser, without download
Curious how models on Hugging Face spend their "bit budget"? A few days ago I shared the first version of a little tool I built out of my own curiosity (and for fun). Since then, thanks in large part to great feedback from people here, a lot has changed.
New in the last few days: - GGUF support - handy with all the new great GGUF quants - Decode for AWQ, GPTQ, NF4, mxfp4 packed experts, packed-int32, and additive-codebook formats - Improved comparison view for diffing two quants of the same model - Built-in anonymous report-issue button connecting a report to specific model - plus many small fixes
After my first post I got great feedback from several community members, and some issues were fixed within hours. I'm planning an acknowledgments section on the site, and when you report an issue you get a receipt ID you can keep to claim credit later. (reports are anonymous by design; I store no identity, so the receipt hash works like a bearer token for your find)
It's still very much a side project I hope others find useful. Explore any HF model in the browser without downloading it, the webpage reads from the safetensors header via a range request, and only tensors you click stream, and large ones are sampled, not downloaded in full. And there is a report button right in the tool when things don’t look right.
Feedback very welcome, especially models that break it :) Or ideas on what is missing.
I work on quantizing models to run efficiently on local hardware, and kept being curious how existing quants spend their "bit budget" during optimization and built a local tool to explore. Many quants apply one setting across all tensors, but some do more interesting things: the model in the screenshot holds attention K at 4.5 bits while Q/V/O get 8.5, and protects layer 0 MLP.
Explore any HF model in the browser without downloading it. The anatomy map is read from the safetensors header via a range request, and only tensors you click ever stream. Large tensors are sampled rather than streamed in full.
Limitations: safetensors only (no GGUF yet), some exotic variants don't work yet, and gated repos aren't supported yet.
Feedback very welcome, especially models that break it.