Instructions to use EldanRing/Winnow-E4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EldanRing/Winnow-E4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E4B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E4B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: ./llama-cli -hf EldanRing/Winnow-E4B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EldanRing/Winnow-E4B:BF16
Use Docker
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- LM Studio
- Jan
- vLLM
How to use EldanRing/Winnow-E4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EldanRing/Winnow-E4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EldanRing/Winnow-E4B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- Ollama
How to use EldanRing/Winnow-E4B with Ollama:
ollama run hf.co/EldanRing/Winnow-E4B:BF16
- Unsloth Desktop
- Docker Model Runner
How to use EldanRing/Winnow-E4B with Docker Model Runner:
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- Lemonade
How to use EldanRing/Winnow-E4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EldanRing/Winnow-E4B:BF16
Run and chat with the model
lemonade run user.Winnow-E4B-BF16
List all available models
lemonade list
- Atomic Chat
Winnow-E4B
Winnow-E4B is a compact EldanRing fine-tune of Gemma 4 E4B IT for typed decisions, chat, and image input. Give it a shared state and questions with known answer options; it scores the candidates without generating an explanation. Typical uses include routing requests, choosing actions, checking conditions, and rating urgency.
The GGUF downloads contain the merged fine-tune. No separate adapter or base-model download is needed. Image input uses the matching vision projector.
Quickstart Β· Inference code Β· API reference Β· Benchmarks
Decision results
These published Q8 results compare the models on the same questions. Jev and Kev report top-choice accuracy; Typed reports agreement with synthetic teacher labels.
| Evaluation | Winnow-E4B Q8 | Winnow-12B Q8 |
|---|---|---|
| JevBench public subset, 231 questions | 80.52% (186/231) | 85.71% (198/231) |
| Kev v9 clean, 1,046 questions | 72.66% | 81.45% |
| Kev v9 additional, 390 questions | 55.90% | 68.97% |
| Local typed decisions, 2,000 decisions | 72.30% | 70.10% |
Winnow-12B leads on the public Jev and Kev suites. E4B leads by 2.2 points on the local Typed teacher-agreement panel. These frozen panels were used during development. JevBench reports public-subset accuracy rather than the composite leaderboard score.
The 12B Kev-clean comparator is 852/1,046, distinct from its separate 853/1,046 release campaign. Evaluation details cover uncertainty, calibration, and scope.
Downloads
| File | Role | Size |
|---|---|---|
| Winnow-E4B-Q8_0.gguf | Q8 language model; measured decision, chat, and vision setup | 8.01 GB / 7.46 GiB |
| Winnow-E4B-BF16.gguf | Higher-precision language model | 15.05 GB / 14.02 GiB |
| mmproj-Winnow-E4B.gguf | F16 vision projector | 990 MB / 0.922 GiB |
Choose Q8 for a smaller memory footprint, or BF16 for higher-precision weights. Add the matching projector for images.
Use an exact target filename or a Winnow preset. Checksums identify every file.
Running the model
The quickstart covers installation, downloads, and launch commands. The server provides /v1/systemone for decisions and /v1/chat/completions for ordinary chat, streaming, and images.
F16 is the default and recommended target K/V cache. Cache precision is separate from the GGUF weight format; use --cache q8_0 for an explicit Q8 override. The pinned MTP assistant uses the target's shared K/V cache.
| Question type | Input | Result |
|---|---|---|
noul |
A yes/no question | Probability of true |
choice |
Named options with optional descriptions | Selected option, probabilities, confidence |
score |
Ordered levels | Level probabilities, expected score, confidence |
Questions share a state prefill and can reuse a cached prefix. Candidate probabilities are conditional on the supplied options; confidence measures concentration, not guaranteed correctness.
Measured performance
On an RTX 5070 Ti 16 GB, the matched 8K text workload delivered 147.8 decisions/s for E4B Q8 versus 67.2 decisions/s for 12B Q8 at 64 questions per request. These are medians of ten warm requests with Q8 weights and KV cache; results depend on workload and batching.
| Q8 configuration | Observed peak device memory | Suggested GPU capacity |
|---|---|---|
| 8K text-only decision workload | 8.56 GiB | 10 GB+ VRAM |
| 64K vision and chat operational probes | 10.72 GiB across the probes | 12 GB+ VRAM |
Capacity suggestions are inferred from measurements. Six cached questions over a 65,023-position image state took 129β131 ms after an 11.59 s first request. A separate 62,431-position multimodal prompt generated 512 tokens at 71.2 tokens/s. See timing and memory details.
Calibration and adaptive reasoning
For the direct text profile, Q8's fitted winnow.temperature is 1.2574172017327816, based on 778 calibration questions. BF16's separately fitted value is 1.3331553765162731. Temperature changes probabilities without changing the winning option; the values belong to their respective runtime profiles.
Direct decisions are the default. Adaptive reasoning adds generated context before rescoring a question. A calibrated 50:50 evaluation improved Brier 0.524807 β 0.481402 and NLL 1.591870 β 1.385153, while source-equal agreement changed +0.083 pp. This profile combines calibration with generated context.
A separate historical uncalibrated recipe gained on Jev and Kev but lost 57 Typed teacher agreements (1,447 β 1,390/2,000), for a net loss of five across all 3,277 decisions. Mean direct HTTP time inside that workflow was 37 ms; the full router averaged 234 ms, including bypasses. Reasoning evaluations give the policies and timing definitions.
MTP separately drafts ordinary chat tokens with the matching E4B assistant. The 8K Q8 vision+MTP4 preset sampled 10,197 MiB peak memory. Direct serving needs no assistant. See runtime profiles and assistant attribution.
Training and limitations
Winnow-E4B is a rank-32, alpha-64 LoRA fine-tune. Language tensors were trained; vision and audio modules remained frozen. The adapter was merged in FP32 before Q8_0 and BF16 export. The private training set combines contrastive synthetic decisions, verified labels, teacher distributions, semantic tasks, and targeted hard cases. Training, development, calibration, and reserved tests were separated; the training data and pipeline are not released.
Credits and license
Winnow-E4B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 E4B IT, released under Apache 2.0. See LICENSE and NOTICE.
The separate inference code builds on llama.cpp and preserves its MIT license. Jev-style describes the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.
- Downloads last month
- 34,651
8-bit
16-bit
