clef-vision-0.8b (Core ML)

A 0.92B-parameter image decision model for Apple silicon, distilled from Cloudflare's Clef-flash (Qwen3.5-9B) into a Qwen3.5-0.8B backbone with a copy of Clef's joint schema head. Same contract as Clef and the Jev/SystemOne API: a state (text or JSON), optional images, and typed questions (noul yes/no, choice, ordered score) go in; one probability for every allowed option of every question comes out of a single forward pass. Nothing is generated.

Runtime: ClefVisionManager in FluidUse (Swift, macOS 15+). Training, evaluation and conversion code: Fluid Inference model-lab (models/clef-vision-0.8b); conversion scripts also in mobius.

Packages

Package Role Size Precision
Vision_P784/VisionTower_fp32.mlpackage Qwen3.5 vision tower, one image ≀ 784 patches (448Β² cap) β†’ 196 tokens 377 MB fp32 (fp16 and mixed precision are numerically wrong on this ViT's activation range)
LM_L512, LM_L1024, LM_L2048/LMRows_fp16.mlpackage Qwen3.5-0.8B decoder + final norm, prefill only, one bucket per sequence length ~960 MB each fp16
Head_L1024_Q16_O64/Head_fp32.mlpackage joint schema head: ≀ 16 questions, ≀ 64 options, one logit per option 264 MB fp32
embeddings.f16 tied token embedding table (host-side gather, also the head's lexical rows) 508 MB fp16

The host does the gathers between packages: tokenization and Clef's record rendering, image smart-resize / bicubic / patchify, learned position-embedding interpolation and 2-D rope for the vision tower, token embedding gather with the vision tokens spliced in, M-RoPE tables, and the head's span-mean matrices. config.json lists every shape and constant; the Swift runtime is the reference host.

Quality

Held-out test, 1,500 records / 4,094 questions over Oxford-IIIT Pets, Food-101 and COCO val images (classification, reference matching, tile grids), PyTorch student:

Teacher agreement Gold accuracy KL
this model (0.8B) 92.5% 92.2% 0.095
teacher Clef-flash 9B (MLX 8-bit) β€” 94.1% β€”

Core ML pipeline vs the PyTorch student on held-out records: max logit difference 6.3e-3, probability drift 2.3e-4, argmax agreement 100%.

Per task (gold): COCO object presence 98%, Food-101 dish 97%, pets breed 95%, reference match grids 96–100%, tile grids 88–93%. "Select all tiles with cats / dogs" 3Γ—3 grids scored one tile per record: 90% whole-grid pass, 98.9% tile accuracy.

Speed (M5 Pro, GPU, CPU_AND_GPU)

Stage Median
vision tower, one image 35–51 ms
LM rows, 1,024 tokens 190 ms
head 14 ms
one record, 1 image, ~500 tokens β‰ˆ 260 ms from Python (coremltools), 136–157 ms from the Swift runtime (vision 39, LM 95–113, head 5 ms); PyTorch MPS: 604 ms
4-image / 1,111-token grid, Swift runtime 640 ms

The Swift runtime (FluidUse ClefVisionCheck parity) matches the PyTorch student on the bundled fixtures to a max probability drift of 8.2e-3 with identical argmax. The Qwen3.5 decoder does not run on the Neural Engine (same as Kev / Cua-S1); the vision tower is fp32, so this is a GPU model.

Training

14,246 records (train 5,648 / dev 2,167 / test 6,431, hash-split by image) labelled by Clef-flash running locally in MLX (8-bit backbone, exact MLX port of the head; validated at ARC-Easy 100% and BANKING77 96.7% against Cloudflare's published 99.5 / 90.9). Student: LoRA r64 on the language model + full head, loss = KL(T=2) to the teacher + 0.3 Γ— label-smoothed CE on dataset gold, 530 steps Γ— 16 records, batch 1 with gradient checkpointing on a 24 GB M5 Pro (6.2 h). Vision tower frozen.

Licenses

Apache-2.0 (this model, Qwen3.5-0.8B, Clef). Training images: Oxford-IIIT Pet (CC BY-SA 4.0), Food-101 (research use), COCO 2017 val (CC BY 4.0).

Downloads last month
157
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FluidInference/clef-vision-0.8b-coreml

Finetuned
(455)
this model