Winnow-E4B

Winnow-E4B is a compact EldanRing fine-tune of Gemma 4 E4B IT for typed decisions, chat, and image input. Give it a shared state and questions with known answer options; it scores the candidates without generating an explanation. Typical uses include routing requests, choosing actions, checking conditions, and rating urgency.

The GGUF downloads contain the merged fine-tune. No separate adapter or base-model download is needed. Image input uses the matching vision projector.

Quickstart Β· Inference code Β· API reference Β· Benchmarks

Decision results

These published Q8 results compare the models on the same questions. Jev and Kev report top-choice accuracy; Typed reports agreement with synthetic teacher labels.

Evaluation Winnow-E4B Q8 Winnow-12B Q8
JevBench public subset, 231 questions 80.52% (186/231) 85.71% (198/231)
Kev v9 clean, 1,046 questions 72.66% 81.45%
Kev v9 additional, 390 questions 55.90% 68.97%
Local typed decisions, 2,000 decisions 72.30% 70.10%

Winnow-12B leads on the public Jev and Kev suites. E4B leads by 2.2 points on the local Typed teacher-agreement panel. These frozen panels were used during development. JevBench reports public-subset accuracy rather than the composite leaderboard score.

The 12B Kev-clean comparator is 852/1,046, distinct from its separate 853/1,046 release campaign. Evaluation details cover uncertainty, calibration, and scope.

Winnow-E4B and Winnow-12B Q8 decision scores on three evaluation panels

Downloads

File Role Size
Winnow-E4B-Q8_0.gguf Q8 language model; measured decision, chat, and vision setup 8.01 GB / 7.46 GiB
Winnow-E4B-BF16.gguf Higher-precision language model 15.05 GB / 14.02 GiB
mmproj-Winnow-E4B.gguf F16 vision projector 990 MB / 0.922 GiB

Choose Q8 for a smaller memory footprint, or BF16 for higher-precision weights. Add the matching projector for images.

Use an exact target filename or a Winnow preset. Checksums identify every file.

Running the model

The quickstart covers installation, downloads, and launch commands. The server provides /v1/systemone for decisions and /v1/chat/completions for ordinary chat, streaming, and images.

F16 is the default and recommended target K/V cache. Cache precision is separate from the GGUF weight format; use --cache q8_0 for an explicit Q8 override. The pinned MTP assistant uses the target's shared K/V cache.

Question type Input Result
noul A yes/no question Probability of true
choice Named options with optional descriptions Selected option, probabilities, confidence
score Ordered levels Level probabilities, expected score, confidence

Questions share a state prefill and can reuse a cached prefix. Candidate probabilities are conditional on the supplied options; confidence measures concentration, not guaranteed correctness.

Measured performance

On an RTX 5070 Ti 16 GB, the matched 8K text workload delivered 147.8 decisions/s for E4B Q8 versus 67.2 decisions/s for 12B Q8 at 64 questions per request. These are medians of ten warm requests with Q8 weights and KV cache; results depend on workload and batching.

Q8 configuration Observed peak device memory Suggested GPU capacity
8K text-only decision workload 8.56 GiB 10 GB+ VRAM
64K vision and chat operational probes 10.72 GiB across the probes 12 GB+ VRAM

Capacity suggestions are inferred from measurements. Six cached questions over a 65,023-position image state took 129–131 ms after an 11.59 s first request. A separate 62,431-position multimodal prompt generated 512 tokens at 71.2 tokens/s. See timing and memory details.

Calibration and adaptive reasoning

For the direct text profile, Q8's fitted winnow.temperature is 1.2574172017327816, based on 778 calibration questions. BF16's separately fitted value is 1.3331553765162731. Temperature changes probabilities without changing the winning option; the values belong to their respective runtime profiles.

Direct decisions are the default. Adaptive reasoning adds generated context before rescoring a question. A calibrated 50:50 evaluation improved Brier 0.524807 β†’ 0.481402 and NLL 1.591870 β†’ 1.385153, while source-equal agreement changed +0.083 pp. This profile combines calibration with generated context.

A separate historical uncalibrated recipe gained on Jev and Kev but lost 57 Typed teacher agreements (1,447 β†’ 1,390/2,000), for a net loss of five across all 3,277 decisions. Mean direct HTTP time inside that workflow was 37 ms; the full router averaged 234 ms, including bypasses. Reasoning evaluations give the policies and timing definitions.

MTP separately drafts ordinary chat tokens with the matching E4B assistant. The 8K Q8 vision+MTP4 preset sampled 10,197 MiB peak memory. Direct serving needs no assistant. See runtime profiles and assistant attribution.

Training and limitations

Winnow-E4B is a rank-32, alpha-64 LoRA fine-tune. Language tensors were trained; vision and audio modules remained frozen. The adapter was merged in FP32 before Q8_0 and BF16 export. The private training set combines contrastive synthetic decisions, verified labels, teacher distributions, semantic tasks, and targeted hard cases. Training, development, calibration, and reserved tests were separated; the training data and pipeline are not released.

Credits and license

Winnow-E4B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 E4B IT, released under Apache 2.0. See LICENSE and NOTICE.

The separate inference code builds on llama.cpp and preserves its MIT license. Jev-style describes the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.

Downloads last month
34,651
GGUF
Model size
78M params
Architecture
gemma4-assistant
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for EldanRing/Winnow-E4B

Finetuned
(402)
this model

Spaces using EldanRing/Winnow-E4B 4

Collection including EldanRing/Winnow-E4B