clef-vision-0.8b (Core ML)
A 0.92B-parameter image decision model for Apple silicon, distilled from Cloudflare's
Clef-flash (Qwen3.5-9B) into a
Qwen3.5-0.8B backbone with a copy of Clef's joint schema head.
Same contract as Clef and the Jev/SystemOne API: a state (text or JSON), optional images, and typed questions
(noul yes/no, choice, ordered score) go in; one probability for every allowed option of every question
comes out of a single forward pass. Nothing is generated.
Runtime: ClefVisionManager in FluidUse (Swift, macOS 15+).
Training, evaluation and conversion code: Fluid Inference model-lab (models/clef-vision-0.8b); conversion
scripts also in mobius.
Packages
| Package | Role | Size | Precision |
|---|---|---|---|
Vision_P784/VisionTower_fp32.mlpackage |
Qwen3.5 vision tower, one image β€ 784 patches (448Β² cap) β 196 tokens | 377 MB | fp32 (fp16 and mixed precision are numerically wrong on this ViT's activation range) |
LM_L512, LM_L1024, LM_L2048/LMRows_fp16.mlpackage |
Qwen3.5-0.8B decoder + final norm, prefill only, one bucket per sequence length | ~960 MB each | fp16 |
Head_L1024_Q16_O64/Head_fp32.mlpackage |
joint schema head: β€ 16 questions, β€ 64 options, one logit per option | 264 MB | fp32 |
embeddings.f16 |
tied token embedding table (host-side gather, also the head's lexical rows) | 508 MB | fp16 |
The host does the gathers between packages: tokenization and Clef's record rendering, image smart-resize /
bicubic / patchify, learned position-embedding interpolation and 2-D rope for the vision tower, token embedding
gather with the vision tokens spliced in, M-RoPE tables, and the head's span-mean matrices. config.json lists
every shape and constant; the Swift runtime is the reference host.
Quality
Held-out test, 1,500 records / 4,094 questions over Oxford-IIIT Pets, Food-101 and COCO val images (classification, reference matching, tile grids), PyTorch student:
| Teacher agreement | Gold accuracy | KL | |
|---|---|---|---|
| this model (0.8B) | 92.5% | 92.2% | 0.095 |
| teacher Clef-flash 9B (MLX 8-bit) | β | 94.1% | β |
Core ML pipeline vs the PyTorch student on held-out records: max logit difference 6.3e-3, probability drift 2.3e-4, argmax agreement 100%.
Per task (gold): COCO object presence 98%, Food-101 dish 97%, pets breed 95%, reference match grids 96β100%, tile grids 88β93%. "Select all tiles with cats / dogs" 3Γ3 grids scored one tile per record: 90% whole-grid pass, 98.9% tile accuracy.
Speed (M5 Pro, GPU, CPU_AND_GPU)
| Stage | Median |
|---|---|
| vision tower, one image | 35β51 ms |
| LM rows, 1,024 tokens | 190 ms |
| head | 14 ms |
| one record, 1 image, ~500 tokens | β 260 ms from Python (coremltools), 136β157 ms from the Swift runtime (vision 39, LM 95β113, head 5 ms); PyTorch MPS: 604 ms |
| 4-image / 1,111-token grid, Swift runtime | 640 ms |
The Swift runtime (FluidUse ClefVisionCheck parity) matches the PyTorch student on the bundled fixtures to a
max probability drift of 8.2e-3 with identical argmax. The Qwen3.5 decoder does not run on the Neural Engine (same
as Kev / Cua-S1); the vision tower is fp32, so this is a GPU model.
Training
14,246 records (train 5,648 / dev 2,167 / test 6,431, hash-split by image) labelled by Clef-flash running locally in MLX (8-bit backbone, exact MLX port of the head; validated at ARC-Easy 100% and BANKING77 96.7% against Cloudflare's published 99.5 / 90.9). Student: LoRA r64 on the language model + full head, loss = KL(T=2) to the teacher + 0.3 Γ label-smoothed CE on dataset gold, 530 steps Γ 16 records, batch 1 with gradient checkpointing on a 24 GB M5 Pro (6.2 h). Vision tower frozen.
Licenses
Apache-2.0 (this model, Qwen3.5-0.8B, Clef). Training images: Oxford-IIIT Pet (CC BY-SA 4.0), Food-101 (research use), COCO 2017 val (CC BY 4.0).
- Downloads last month
- 157