Instructions to use kirp/jpt-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kirp/jpt-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="kirp/jpt-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("kirp/jpt-9b") model = AutoModelForMultimodalLM.from_pretrained("kirp/jpt-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kirp/jpt-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kirp/jpt-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kirp/jpt-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/kirp/jpt-9b
- SGLang
How to use kirp/jpt-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kirp/jpt-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kirp/jpt-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kirp/jpt-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kirp/jpt-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use kirp/jpt-9b with Docker Model Runner:
docker model run hf.co/kirp/jpt-9b
JPT-9B
JPT-0.8B ยท JPT-4B ยท JPT-9B ยท JPT-35B-A3B ยท llm2jev
What is JPT-9B
JPT-9B is a fast, open decision model: give it a situation and typed questions, get a calibrated probability for every option from one forward pass. No generated explanation, no reasoning tokens โ latency is one prefill.
The large sibling of JPT-4B: the same recipe and the same data.
It implements the typed-decision interface introduced by Jev from TypeSafe AI [1]: a caller sends a state plus questions, and each question is one of three types. JPT is an independent model, not derived from Jev and not trained on Jev outputs; it is an open alternative behind the same interface.
| Question type | What it answers | Options |
|---|---|---|
choice |
pick one | 2โ255 labels |
score |
a level on an ordered scale | the scale's levels |
noul |
yes / no | true, false |
Built on Qwen/Qwen3.5-9B: a LoRA fine-tune merged into full weights. The vision tower is unchanged.
Probabilities use one temperature T = 1.087, fit once on a held-out split โ never per benchmark.
โฏโฏ Benchmarks
โฏ JevBench v1.4.0
JevBench [2] scores general typed decisions; JPT-9B reaches 0.853 public accuracy, level with JevK5 and just under Winnow-12B and Jev 1.13.0 โ but below the smaller JPT-4B (0.879). v1.4.0 has 231 public items and 308 sealed items only the maintainer can run, so JPT-9B has no official v1.4 score yet.
Table
| System | Params | Public accuracy (231) | Sealed accuracy (308) | v1.4 score |
|---|---|---|---|---|
| JPT-35B-A3B (ours) | 35B-A3B | 0.892 | pending | pending |
| JPT-4B (ours) | 4B | 0.879 | pending | pending |
| Jev 1.13.0 (TypeSafe AI, API) | closed | 0.866 | 0.367 | 63.3 |
| Winnow-12B Q8 | 12B | 0.857 | 0.331 | 55.6 |
| JPT-9B | 9B | 0.853 | pending | pending |
| JevK5 v0.2.0 | 27B | 0.853 | 0.331 | 62.0 |
| decider-35b-a3b | 35B-A3B | 0.831 | 0.315 | 41.2 |
| Hopper | โ | 0.823 | 0.341 | 59.4 |
| openjev 4B v5ยน | 4B | 0.814 | โ | โ |
| SemIf, formerly OpenJev (Qwen3.5-4B) | 4B | 0.810 | 0.263 | 47.7 |
| local-jev Qwen3.5-4B | 4B | 0.805 | 0.260 | 46.8 |
| reflex 4B | 4B | 0.792 | 0.282 | 54.0 |
| kev 4B (research preview) | 4B | 0.662 | 0.224 | 36.1 |
Source: other rows from results/v1.4/jevbench-v1.4-results.json at jevbench commit 2fa63fa (2026-09-23).
Ours measured with JevBench's CLI at 2fa63fa (SGLang 0.5.18, kirp/jpt-9b@7114b0c): 197/231, easy 48 ยท standard 68 ยท
hard 81. Our own harness gets 198: on original-intent-02-0 two options tie exactly (0.4864 each) and the two tools
break the tie differently.
ยน Not in the v1.4 results; its own card's number on the 231 public items. A same-prompt Qwen3.5-9B zero-shot run on
the public items has not been done yet.
โฏ Decision Index 0.2.1
Decision Index [3] 0.2.1 (2026-09-27) is the broadest test: the full frozen suite, 38 scored benchmarks in five areas, chance-corrected โ JPT-9B scores 46.89, the best 9B on the board and 0.2 behind Decider 35B-A3B. Run through llm2jev over SGLang and submitted as apolinario/decision-index#7; 0.2.1 rescores the same run (42.73 under 0.2).
Table
| Model | Params | Decision Index 0.2.1 |
|---|---|---|
| Jev 1.13.0 (TypeSafe AI, API) | closed | 57.91 |
| JPT-35B-A3B (ours, not on the board yet) | 35B-A3B | 52.89 |
| Decider 35B-A3B | 35B-A3B | 47.11 |
| JPT-9B | 9B | 46.89 |
| Decision 1.0 Lux | 9B | 43.49 |
| Bespoke Nimble 9B v2 | 9B | 39.57 |
| Kev 9B | 9B | 38.48 |
Source: live board data/index-v0.2.1.json (generated 2026-09-27 16:59 UTC); full run and scores.json in
kirp/decision-index-results-jpt-9b
(gated: it carries the suite's GPQA/HLE item text).
By area, against Jev 1.13.0 on the same items and scorer: JPT-9B is behind Jev in all five areas and ahead of it on 4 of the 38 index benchmarks.
Per-area and per-benchmark skill vs Jev 1.13.0
| Area | Jev 1.13.0 | JPT-9B |
|---|---|---|
| Knowledge | 51.4 | 31.7 |
| Language | 62.0 | 56.7 |
| Retrieval | 55.4 | 44.6 |
| Tools | 75.1 | 67.0 |
| Arts | 37.7 | 28.6 |
| Area | Benchmark | Jev 1.13.0 | JPT-9B |
|---|---|---|---|
| Arts | BPoMP | 81.8 | 72.4 |
| Arts | ForecastBench | 30.6 | 27.4 |
| Arts | Habermas Machine | 21.5 | 13.2 |
| Arts | Humicroedit | 23.7 | 19.5 |
| Arts | New Yorker | 62.6 | 52.6 |
| Arts | POP909-CL | 15.9 | 3.5 |
| Arts | cfcolor | 28.8 | 11.4 |
| Games | ChessBench | 9.8 | 7.3 |
| Knowledge | BBH | 89.7 | 55.6 |
| Knowledge | CLadder | 45.3 | 33.8 |
| Knowledge | CRUXEval | 57.1 | 29.6 |
| Knowledge | GPQA Diamond | 71.4 | 21.8 |
| Knowledge | GSM8K | 75.6 | 69.2 |
| Knowledge | HLE | 4.7 | 0.0 |
| Knowledge | MMLU-Pro | 80.5 | 49.1 |
| Knowledge | MuSR | 46.1 | 28.1 |
| Knowledge | SATA-Bench | 25.4 | 22.6 |
| Language | ACOS | 27.3 | 20.3 |
| Language | ANLI | 62.2 | 50.0 |
| Language | ContractNLI | 59.1 | 74.2 |
| Language | FinEntity | 80.8 | 90.0 |
| Language | HellaSwag | 92.7 | 81.2 |
| Language | NLI4CT | 69.0 | 59.1 |
| Language | RAGTruth | 51.3 | 50.6 |
| Language | VAST | 46.9 | 46.6 |
| Language | WinoGrande | 83.9 | 49.2 |
| Language | iSarcasmEval | 36.3 | 43.9 |
| Retrieval | Amazon ESCI | 43.8 | 40.6 |
| Retrieval | BANKING77 | 79.5 | 65.0 |
| Retrieval | BRIGHT | 40.6 | 39.9 |
| Retrieval | CLINC150+OOS | 89.2 | 35.8 |
| Retrieval | HoVer | 45.7 | 32.7 |
| Retrieval | PhishNChips phishing decisions | 25.1 | 52.1 |
| Tools | API-Bank | 88.0 | 78.3 |
| Tools | BFCL | 94.3 | 93.1 |
| Tools | Home appliance simulator | 52.3 | 43.2 |
| Tools | ToolRet | 59.9 | 55.8 |
| Tools | When2Call | 74.6 | 57.3 |
Chance-corrected skill ร 100 (0 = random, 100 = perfect). Jev's numbers are its official entry on the live board
(jev-1.13.0); ours are from the same kit and suite.
โฏ More benchmarks
Against JPT-4B on the same evals: JPT-9B is better on NLI, intent classification and EnvBench public, and worse on the JevBench hard tier. Same-prompt base-model (Qwen3.5-9B zero-shot) rows exist only for EnvBench; the image evals (ScreenSpot-v2, Screen2Words, ERQA) have not been run on this checkpoint.
| Benchmark (version, n) | What it tests | JPT-9B | JPT-4B | Jev 1.13.0 |
|---|---|---|---|---|
| JevBench v1.4.0 public hard tier [2] (111) | hardest general decisions | 0.730 (ECE 0.097, Brier 0.382) | 0.784 | โ |
| Typed decisions test (ours, 2,000) | in-distribution typed decisions | 0.806 (ECE 0.149, Brier 0.316) | 0.796 | โ |
| ANLI r1 / r2 / r3 [4] (dev) | adversarial NLI | 0.737 / 0.647 / 0.677 | 0.697 / โ / 0.613 | โ |
| Banking77 [5] / MASSIVE 1.1 [6] (en / de / zh) | intent classification | 0.787 / 0.897 / 0.853 / 0.830 | 0.757 / 0.857 / 0.833 / 0.837 | โ |
| AG News [7] / Emotion [8] / SST-5 [9] (test) | text classification (train splits in the mix) | 0.917 / 0.557 / 0.603 | โ | โ |
| EnvBench v0.1 (ours) public / held-out [10] (skill 0โ100) | sequential decisions in game envs | 50.5 / 46.8 | 47.7 / 47.0 | โ |
Jev 1.13.0 has no official score on these splits (our own test/dev cuts, EnvBench, and the image sets), so its column is "โ"; its official scores on the Decision Index versions of ANLI and BANKING77 are in the per-benchmark table above. Running the Jev API on these splits would fill them.
Qwen3.5-9B, same prompt, zero-shot on EnvBench v0.1: 24.2 public / 25.3 held-out. Banking77, MASSIVE, AG News, Emotion, SST-5 and typed rows are in-distribution (train splits in the mix, test items not).
โฏโฏ Quick Start
Two pieces: an engine that holds the weights, and llm2jev (>= 0.6.1) in front of it, reading option probabilities off the engine.
โก SGLang (recommended)
python -m sglang.launch_server --model-path kirp/jpt-9b --port 30000 \
--context-length 32768 --mamba-scheduler-strategy extra_buffer & # Qwen3.5's DeltaNet layers need this flag
llm2jev --model kirp/jpt-9b --backend sglang --url http://127.0.0.1:30000 --port 8080 --temperature 1.087
Tested with SGLang 0.5.9; install cuDNN 9.15+ over its pinned 9.10:
pip install "sglang==0.5.9" && pip install "nvidia-cudnn-cu12>=9.15".
๐ vLLM
vllm serve kirp/jpt-9b --max-logprobs 256 --return-tokens-as-token-ids --enable-scale-out --port 8000
llm2jev --model kirp/jpt-9b --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.087
The three vLLM flags are required: without them every request is a bare HTTP 400.
๐งช No engine (quick check only)
pip install "llm2jev[hf,vision]"
llm2jev --model kirp/jpt-9b --backend hf --port 8080 --temperature 1.087
Serializes requests; for traffic use SGLang or vLLM.
๐จ Ask it a question
import requests
r = requests.post("http://127.0.0.1:8080/v1/systemone", json={
"state": "Refund policy: full refund within 30 days of purchase; 50% until day 60; none after.\n"
"Order 1182 was bought on 3 March and returned on 20 April.",
"questions": {
"refund": {"type": "choice", "instructions": "What refund does order 1182 get?",
"criteria": {"full": "Full refund", "half": "50% refund", "none": "No refund"}},
"late": {"type": "noul", "instructions": "Was the return made after day 30?",
"criteria": {"true": "Yes", "false": "No"}}}})
print(r.json()["answers"]) # each answer has the per-option probabilities
๐ผ๏ธ With images
state = [{"role": "user", "content": [
{"type": "image", "image": "https://example.com/screen.png"},
{"type": "text", "text": "Task: open the settings page. Numbered boxes mark clickable elements."}]}]
questions = {"click": {"type": "choice", "instructions": "Which box should be clicked?",
"criteria": {"1": None, "2": None, "3": None, "4": None, "5": None}}}
โฏโฏ Training
The JPT-4B recipe and data (mix_train_env_v11, 49,221 typed questions) on a larger base, one epoch, merged into full weights.
| Part | What it is |
|---|---|
| Method | LoRA r=16 on every attention, DeltaNet and MLP projection of the language model, lr 5e-5; vision tower untouched |
| Loss | multi-class Brier over the option labels, on llm2jev's chat prompt with thinking disabled |
| Batch | 8 GPUs ร 1 ร 5 gradient-accumulation steps = 40 questions per step |
| Data | 49,221 questions in 32,835 records; one epoch over two option-shuffled copies โ sources on the JPT-4B card |
| Held out | no item from JevBench, EnvBench held-out seeds, the Decision Index frozen suite or our typed test split |
โฏโฏ Limitations
- Not strictly better than JPT-4B. The JevBench hard tier (0.730 vs 0.784) and the EnvBench held-out
gamearea (0.288 vs 0.315) are lower, despite lower training and validation loss throughout. Not yet root-caused. - Arithmetic and dates are its weakest area: it answers from the evidence given and has no reasoning phase by design.
- Up to 255 options are accepted; training covered up to 77 (Banking77).
- English first. Other languages come only from a few multilingual classification sets.
- Images are untested on this checkpoint; they go zero-shot through the base vision tower.
โฏโฏ References
- TypeSafe AI. Jev. https://typesafe.ai
- F. Standhartinger. JevBench, v1.4.0. https://github.com/fstandhartinger/jevbench
- Decision Index, edition 0.2.1. https://huggingface.co/spaces/multimodalart/jev-decision-index
- Nie et al. Adversarial NLI. ACL 2020.
- Casanueva et al. Efficient Intent Detection with Dual Sentence Encoders (Banking77). NLP4ConvAI 2020.
- FitzGerald et al. MASSIVE. ACL 2023.
- Zhang et al. Character-level Convolutional Networks for Text Classification (AG News). NeurIPS 2015.
- Saravia et al. CARER: Contextualized Affect Representations for Emotion Recognition. EMNLP 2018.
- Socher et al. Recursive Deep Models for Semantic Compositionality (SST). EMNLP 2013.
- EnvBench, v0.1 (ours, frozen 2026-09-23; not yet public): programmatically solved game, planning and rule decisions with exact gold answers.
โฏโฏ License
CC BY-NC 4.0. The weights derive from Qwen3.5-9B (Apache-2.0), but some training datasets allow only non-commercial or research use, so the model is released for non-commercial use.
JPT-9B is an independent open model that implements a typed-decision interface (noul, choice and score questions answered with probabilities). It is not affiliated with, endorsed by or derived from TypeSafe AI or its Jev model, and it was not trained on Jev outputs.
- Downloads last month
- 783


