š„ Imajev-4b is #1 of 50 on Image JevBench (v0.1.4, 29 Sep 2026), the leaderboard for AI models that make decisions from images.
A small demo built on imajev-4b: a closet stylist š Request you to star it here so we can make it better - https://github.com/mohit67890/imajev.
Tap a piece and it reads the photo (red 75%, checked 99%), then the app picks bottoms, shoes and a bag from your own closet in the colours you like. Change your colours and the outfit changes.
Under the hood it's one request with one photo and 4 typed questions. Every option gets a probability, so the app applies its rules (one pattern per outfit) and ranks what's left. About 1.1 s per outfit on a Mac (MLX). Every % in the video is the model's real answer.
Also, thanks to @zenmagnets for the FP8 version for Blackwell GPUs: same calibration, 0.1 pt less accuracy on all 23,900 DecisionBench rows, ~1.6Ć faster š zenmagnets/Imajev-4B-FP8-SM120
š„ Imajev-4b is #1 of 50 on Image JevBench (v0.1.4, 29 Sep 2026), the leaderboard for AI models that make decisions from images.
A small demo built on imajev-4b: a closet stylist š Request you to star it here so we can make it better - https://github.com/mohit67890/imajev.
Tap a piece and it reads the photo (red 75%, checked 99%), then the app picks bottoms, shoes and a bag from your own closet in the colours you like. Change your colours and the outfit changes.
Under the hood it's one request with one photo and 4 typed questions. Every option gets a probability, so the app applies its rules (one pattern per outfit) and ranks what's left. About 1.1 s per outfit on a Mac (MLX). Every % in the video is the model's real answer.
Also, thanks to @zenmagnets for the FP8 version for Blackwell GPUs: same calibration, 0.1 pt less accuracy on all 23,900 DecisionBench rows, ~1.6Ć faster š zenmagnets/Imajev-4B-FP8-SM120
Imajev-4b is #1 of 91 on JevBench and #3 of 56 on DecisionBench š
Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.
When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.
Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.
Results this week, both run by the benchmarks' own maintainers:
š„ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29 š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models. ā Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set
Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.
What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.
Imajev-4b is #1 of 91 on JevBench and #3 of 56 on DecisionBench š
Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.
When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.
Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.
Results this week, both run by the benchmarks' own maintainers:
š„ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29 š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models. ā Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set
Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.
What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.