Sakura EmbeddingGemma 2 — AutoRound W4A16

Community Quantization. Not official from Google.
Quantized and evaluated by Sakura (webmp3) using Intel AutoRound optimization on top of Google's official embeddinggemma-2 architecture.


Overview

This repository provides an optimized AutoRound W4A16 (4-bit weights, 16-bit activations) quantization of google/embeddinggemma-2.

It is packaged as a standard Hugging Face repository containing packed Safetensors weights that run natively and out of the box with sentence-transformers and transformers.

  • Upstream Model: google/embeddinggemma-2
  • Pinned Upstream Commit: 914f7f89142e33e77833254d9c9b90c3cef7303b
  • Base Architecture: EmbeddingGemma2ForSequenceClassification / EmbeddingGemma2TextModel
  • Model File Size: model.safetensors: 1,299,010,632 bytes / 1,299.01 MB / 1,238.83 MiB. The pinned BF16 original model file is 1,488,915,288 bytes / 1,488.92 MB / 1,419.94 MiB: 12.75% smaller on disk. This compares model files, not total directories or measured runtime memory.
  • Verification Status: Public Hub access verified; local release smoke test passed; packaged for Transformers & SentenceTransformers.
  • Quantization Framework: AutoRound 0.16.0, as recorded in the published quantization_config.json.
  • Quantization Scheme: W4A16 symmetric (bits=4, group_size=64, iters=200), calibrated on 234 mixed retrieval texts (build/calibration_dual256.json); recorded packing format auto_round:auto_gptq.
  • Version: v2 (2026-10-07). Re-quantized with a broader calibration set, more iterations and a smaller group size. The v1 release (group size 128, 100 iterations, 64 calibration texts) is kept in this repository's git history. v2 improves BF16 fidelity and external retrieval modestly; see the tables below.
  • Release Decision: RELEASE GO (Early community release based on empirical fidelity retention)

Architectural Breakdown: What is W4 vs. What Remains BF16

We explicitly do not claim "Full INT4/Q4". High-fidelity embedding models require careful treatment of sensitive components:

Component Precision Details
Text Backbone Linear Layers W4A16 216 Linear layers in language_model.layers.0 through language_model.layers.23 (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj). Group size 64, symmetric.
Token Embeddings BF16 language_model.embed_tokens (256,000 vocab) preserved in bfloat16 to avoid semantic vocabulary collapse.
Embedding Projection Head BF16 language_model.embedding_projection (768d output head) preserved in bfloat16 to preserve precise directional geometry.
Normalization Layers BF16 All RMSNorm and LayerNorm modules preserved in bfloat16 to prevent activation scale clipping.
Vision & Audio Towers BF16 Multimodal encoders (vision_tower, audio_tower) and multimodal projection heads are preserved unquantized.

Fidelity & Benchmark Results

All evaluations were conducted against an unquantized bfloat16 reference across 100 query/document retrieval pairs (English and German) and 50 code retrieval pairs.

Release Gate Status

The candidate was evaluated against strict quality criteria:

  • Empirical Gate Result: RELEASE GO / Early Community Release
  • Gate Context: The initial theoretical target of $\ge 0.99$ Mean Cosine Similarity and $\ge 95%$ Top-5 Retrieval Agreement is not fully reached (v2 achieved: 0.9897 Mean Cosine, 89.0 % Top-5 Agreement; v1: 0.9878 / 87.6 %). With 100.0 % Top-1 Agreement, 100.0 % Recall@5 and 0.9814 Spearman Correlation, these benchmark results show retained retrieval success alongside measurable BF16-fidelity loss; no NaN or Inf anomaly was observed.

The companion GGUF HQ releases achieve higher reported BF16 fidelity on the same existing BF16 reference and text/code benchmark. They are separate quantization runs; their results do not change this Safetensors model's release-gate outcome.

GGUF companion tier MiB Combined Mean Cosine (768d) Combined Spearman (768d) Text Top-5 Text/Code Recall@5
Q4 HQ 169.79 0.99400270 0.98857239 90.00% 100% / 100%
Q5 HQ — recommended balance 199.88 0.99828082 0.99664205 95.40% 100% / 100%
Q6 HQ 231.75 0.99940187 0.99879852 96.80% 100% / 100%

Values are quoted from the GGUF README's Side-by-Side Comparison. Aggregation differs: this Safetensors section reports alignment over 200 text query/document vectors; the GGUF combined Mean Cosine/Spearman include all 300 text-and-code vectors. Text Top-5 uses the 100-pair text retrieval benchmark in both. These are not identically aggregated mean scores, and no all-metric comparison with Unsloth is implied.

MRL (Matryoshka Representation Learning) Dimensions

EmbeddingGemma 2 supports dimension truncation followed by L2-renormalization. The W4A16 model demonstrates robust stability across all standard MRL truncations:

MRL Dimension Mean Cosine Sim Min Cosine Sim Spearman Rank Corr Mean L2 Drift Top-1 Agreement Top-5 Agreement Recall@5 nDCG@10
768d (Full) 0.9897 0.9662 0.9814 0.1404 100.0 % 89.0 % 100.0 % 0.9645
512d 0.9899 0.9670 0.9815 0.1390 99.0 % 88.6 % 100.0 % 0.9590
256d 0.9910 0.9701 0.9832 0.1309 100.0 % 85.6 % 100.0 % 0.9579
128d 0.9934 0.9807 0.9828 0.1120 99.0 % 86.6 % 100.0 % 0.9454

Language & Domain Subsets (768d)

Subset Mean Cosine Sim Top-1 Agreement Top-5 Agreement Recall@5
English Queries & Docs 0.9897 100.0 % 88.4 % 100.0 %
German Queries & Docs 0.9898 100.0 % 87.6 % 100.0 %
Code Retrieval 0.9893 100.0 % 83.2 % 100.0 %

External retrieval (BEIR / MTEB test sets, measured)

Same texts and prefixes for all rows, texts truncated to 128 tokens; 500-document corpora (all positives + sampled distractors). BF16 = unquantized google/embeddinggemma-2.

Benchmark Model MRR nDCG@10 Recall@10 MRR vs BF16 nDCG@10 vs BF16
SciFact (300 queries) BF16 reference 0.8971 0.9137 0.9700 100 % 100 %
W4A16 v2 (this release) 0.8921 0.9105 0.9733 99.4 % 99.7 %
W4A16 v1 (previous release) 0.8892 0.9051 0.9667 99.1 % 99.1 %
NFCorpus (249 queries) BF16 reference 0.4609 0.3099 0.6386 100 % 100 %
W4A16 v2 (this release) 0.4509 0.2971 0.6265 97.8 % 95.9 %
W4A16 v1 (previous release) 0.4466 0.2973 0.6145 96.9 % 95.9 %

The differences between v1 and v2 are small and not uniform. v2 is better in mean cosine and Spearman correlation in every MRL dimension and subset, in Top-5 agreement at 768d (+1.4 points), 128d, English and German, and in SciFact/NFCorpus MRR. It is slightly worse in Top-5 agreement at 256d (-0.6 points) and on the code subset (-4.0 points: 83.2 % vs. 87.2 %), and nDCG@10 on the text benchmark is 0.001-0.003 lower; NFCorpus nDCG@10 is unchanged. NFCorpus has only 249 queries, so differences of about one point are within noise. Overall v2 is a modest improvement, not a step change.


Runtime & Community Format Comparison

Distribution / Format Runtime Compatibility Direct Transformers / ST Support Notes
Sakura AutoRound W4A16 (This repo) Python, PyTorch, Transformers, SentenceTransformers Yes (Plug-and-play) Runs directly in existing Python AI pipelines without needing custom binary builds.
GGUF Community Releases (unsloth, ggml-org) llama.cpp No (requires llama-server or bindings) EmbeddingGemma 2 support is upstream in llama.cpp via PR #30054, merged 2026-10-06. The earlier unknown model architecture result came from an older local build; use a build that includes this support. Community GGUF controls have since been loaded and benchmarked with a supported build.
ONNX Community Releases (onnx-community) ONNX Runtime / Transformers.js No (ONNX graph format) Modular multi-graph export tailored for WebGPU/JavaScript execution.
Sakura AutoRound GGUF HQ (Companion) llama.cpp Via llama-server / bindings Q4 HQ 169.79 MiB, Q5 HQ 199.88 MiB (recommended balance), Q6 HQ 231.75 MiB. This Safetensors/ST package is for native Python pipelines; the standalone text GGUF variants are for llama.cpp.

Safetensors vs. GGUF HQ: separate quantization runs

These are separate optimization runs, not two exports of one quantized state. The published Safetensors quantization_config.json records a different scheme and iteration count from the preserved GGUF cal256 build configurations/logs. Both use AutoRound 0.16.0 and the same base model, but that does not imply identical optimized weights or calibration.

Setting This Safetensors W4A16 release GGUF HQ cal256 runs
Transformer weight types 4-bit symmetric W4A16 Q4_K/Q5_K/Q6_K base with 48 Q8_0 block-PLE linears
Weight grouping Group size 64 Q4_K/Q5_K: 32 weights per subgroup, 8 subgroups per 256-weight superblock; Q6_K: 16, 16 subgroups per 256-weight superblock. Q8_0 uses 32-weight blocks.
Iterations 200, published configuration 50, saved build configurations/logs
Calibration selection 234 real retrieval texts from 13 public datasets (queries with the EmbeddingGemma query prefix, passages with the document prefix; EN/DE/code/science/web), manifest build/calibration_dual256.json (SHA256 d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd), built with build/requant.py 256 synthetic retrieval samples: 48 EN queries, 48 EN docs, 48 DE queries, 48 DE docs, 32 code queries and 32 code snippets; zero exact benchmark overlap
Optimization/export path AutoRound W4A16; published packing format auto_round:auto_gptq; Safetensors for Transformers/ST Native AutoRound SignRoundV2 (enable_alg_ext=True), matching GGUF optimized-state packing, then selected embedding-path precision overrides
AutoRound version 0.16.0, published configuration 0.16.0, preserved builds
Scope Full multimodal checkpoint, vision/audio and sensitive components retained in BF16 Standalone text GGUF; CPU/native llama.cpp runtime

Provenance of this release (v2): the calibration texts, the quantization script (build/requant.py, run with AutoRound 0.16.0 on CPU) and the evaluation results are published or reproducible from this repository. The v1 release (git history) was built by the earlier pipeline script with up to 64 texts from a local calibration file whose exact historical content was not preserved. The calibration texts come from public datasets (some, e.g. MS MARCO, carry research-use terms); only the texts used for calibration are included.

The GGUF cal256 runs do retain explicit sample count, calibration SHA256 b65d22bad14722ab02d316ddae2ea6f860b92c94669f5717321b66e81abb91b8, resolved per-layer configuration, native optimized-state packing evidence and training logs. These differences are enough to rule out a shared optimized quantization state. Public GGUF settings and runtime provenance are linked from the companion README; the inspected W4A16 pipeline settings are summarized here without claiming historical inputs that were not preserved.


Quickstart & Usage

1. With SentenceTransformers (Recommended)

from sentence_transformers import SentenceTransformer
import torch

# Load the quantized model
model = SentenceTransformer(
    "webmp3/Sakura-EmbeddingGemma-2-AutoRound",
    model_kwargs={"torch_dtype": torch.bfloat16}
)

# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")

# Documents (unprompted)
docs = [
    "Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
    "The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)

# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)

2. Matryoshka Dimension Truncation (MRL)

To reduce memory and storage footprint, simply slice the vector and re-normalize:

import torch
import torch.nn.functional as F

# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)

3. With Hugging Face Transformers

from transformers import AutoModel, AutoTokenizer
import torch

model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)

inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

Limitations & Honest Disclosure

  1. Text Precision and File Size: The 216 text-backbone linear layers use 4-bit packed weights. model.safetensors is 1,299,010,632 bytes (1,299.01 MB / 1,238.83 MiB), compared with 1,488,915,288 bytes (1,488.92 MB / 1,419.94 MiB) for the pinned full BF16 original. The 12.75% model-file saving is modest: token embeddings and multimodal towers remain BF16. These on-disk measurements do not establish a runtime-memory reduction; the former ~5.5 GB BF16 comparison is not supported by the checked model files.
  2. Multimodal Towers: The vision and audio towers remain in 16-bit precision. If you do not use vision or audio inputs, memory consumption can be minimized by only loading the text language model.
  3. Execution Device: Optimal execution requires modern CPU (AVX-512 / VNNI) or GPU environments supporting accelerated bfloat16 and packed int4 kernels.

SHA256 Checksums

Every file in this release has been cryptographically verified:

1be8f9083684a77b24317163310782efd0df30e5ce5cc2c097d447b3a5402ce1  model.safetensors
48d4e0ca37c329ae9fc6582f7de5224f6d5ce9a3b81ea00dfe61eaf7a622c45c  config.json
072b3e5dc502e1beabac6414e8a663516995b75abe1c03fedf1ddb696eb49483  quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5  config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228  sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd  modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78  preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c  processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4  tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874  tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e  chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9  1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524  2_Normalize/config.json
d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd  build/calibration_dual256.json

Citation & Acknowledgements

Downloads last month
18
Safetensors
Model size
0.6B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound

Quantized
(26)
this model