Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SeaWolf-AIΒ 
posted an update 2 days ago
Post
3868
πŸ“± POCKET β€” a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU

We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud β€” it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.

Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0):
β€’ CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s β†’ 2.69Γ— faster
β€’ GPU generate (H100): 197 vs 89 tok/s β†’ 2.22Γ— faster
β€’ GPU prompt processing (H100): 753 vs 1816 β†’ 0.41Γ— (Bonsai wins this one β€” MoE prefill wakes every expert, so sparsity stops helping there. We say so.)
β€’ Quality (HellaSwag, 400 q): 61.0% vs 60.0% β†’ a tie (confidence intervals overlap)

On a real consumer laptop β€” MacBook M3 Pro (18 GB) β€” POCKET wins every axis, prompt processing included:
β€’ Metal generate: 25.4 vs 12.8 β†’ 1.99Γ—
β€’ CPU generate: 13.8 vs 4.4 β†’ 3.13Γ—
β€’ Metal prompt: 240.7 vs 73.4 β†’ 3.28Γ—

One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all β€” it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.

πŸ“– Full story (tech, measurements, recipes): https://huggingface.co/blog/FINAL-Bench/pocket

Models:
πŸ“¦ POCKET-35B-GGUF (PC / server, no GPU): FINAL-Bench/POCKET-35B-GGUF
πŸ‡°πŸ‡· POCKET-KR-GGUF (Android): FINAL-Bench/POCKET-KR-GGUF
🍎 POCKET-KR-MLX (iPhone / Mac): FINAL-Bench/POCKET-KR-MLX
🌍 POCKET-EN-GGUF (English phone / PC): FINAL-Bench/POCKET-EN-GGUF
πŸ–₯️ Live demo (answering on a CPU, no GPU): FINAL-Bench/POCKET-35B-CPU
πŸ“š Collection: FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6

The prefill number is the one that actually matters, and it flips depending on hardware.

On H100, POCKET loses prompt processing 0.41x, MoE prefill wakes every expert so sparsity buys nothing there. But on the M3 Pro, POCKET wins prompt processing too, 3.28x. Same architecture, opposite verdict, depending only on where it runs.

Is that Metal handling sparse-expert prefill differently than H100's batched prefill, or is the H100 loss really about batch size, the dense model gets to batch prefill and Metal single-stream doesn't?