GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

FINAL-Bench/POCKET-35B-GGUF overview

πŸ“š Collections β–Ά POCKET Models https://huggingface.co/collections/FINAL Bench/pocket models 6a618ee5d23eafb7e185a5c6 β€” this family on device, no GPU Darwin Fam…

llama.cppggufconversationalon-devicemobileiphoneandroidcpulocal-llmedgemixture-of-expertsmoequantizedpocketvidraftqwen3_5_moedarwintext-generationbase_model:FINAL-Bench/Darwin-36B-Opusbase_model:quantized:FINAL-Bench/Darwin-36B-Opuslicense:apache-2.0endpoints_compatibleregion:usimatrix

Runs locally from ~7.67 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6,077
Likes
41
Pipeline
text-generation

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
POCKET-35B-IQ1_M.ggufGGUFIQ1_M7.67 GBDownload
POCKET-35B-Q2_K.ggufGGUFQ2_K12.05 GBDownload
POCKET-35B-Q3_K_M.ggufGGUFQ3_K_M15.61 GBDownload
POCKET-35B-Q4_K_M.ggufGGUFQ4_K_M19.71 GBDownload

Model Details

Model IDFINAL-Bench/POCKET-35B-GGUF
AuthorFINAL-Bench
Pipelinetext-generation
Licenseapache-2.0
Base modelFINAL-Bench/Darwin-36B-Opus
Last modified2026-07-26T05:44:53.000Z

Model README

---

license: apache-2.0

library_name: llama.cpp

pipeline_tag: text-generation

base_model:

  • FINAL-Bench/Darwin-36B-Opus

tags:

  • gguf
  • llama.cpp
  • conversational
  • on-device
  • mobile
  • iphone
  • android
  • cpu
  • local-llm
  • edge
  • mixture-of-experts
  • moe
  • quantized
  • pocket
  • vidraft
  • qwen3_5_moe
  • darwin

---

> ### πŸ“š Collections

> β–Ά POCKET Models β€” this family (on-device, no GPU)

> Darwin Family Β· Aether Foundation Β· VKAE Accelerated Β· Metacognition Adapters

!POCKET

POCKET-35B-GGUF

A 35B model that runs on your PC with no GPU β€” and on your phone. Just stock llama.cpp. No fork, no CUDA, no cloud.

> πŸš€ Try it live, no install β†’ ![POCKET-35B demo](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) ![POCKET-26B demo](https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU) β€” both answering on a CPU-only box (no GPU). POCKET-26B is Gemma4-based.

![License](https://www.apache.org/licenses/LICENSE-2.0) ![Runtime](https://github.com/ggml-org/llama.cpp) ![No GPU]() ![Base]()

Pick your build β†’ ![35B](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) ![KR GGUF](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) ![KR MLX-0f6e56)](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) ![EN GGUF](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) ![26B](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF)

The POCKET lineup β€” pick by your device

| Repo | File | Size | Runs on | Best for | Korean PPL* |

|---|---|---|---|---|---|

| POCKET-35B-GGUF | Q4_K_M | 21 GB | PC / server (32 GB RAM) | top quality | 5.79 |

| POCKET-35B-GGUF | Q2_K ⭐ | 13 GB | mini-PC, no GPU | daily driver | 6.49 |

| POCKET-35B-GGUF | IQ1_M | 8.2 GB | 16 GB RAM box | smallest full model | 9.69 |

| POCKET-KR-GGUF | IQ2_M | 5.1 GB | Android 8 GB+ | πŸ‡°πŸ‡· Korean phone | 7.95 |

| POCKET-KR-MLX | 2-bit | 5.1 GB | 🍎 iPhone / iPad / Mac | πŸ‡°πŸ‡· Korean, Apple-native | 7.95 |

| POCKET-EN-GGUF | iPhone-mix | 5.3 GB | 🍎 iPhone (PocketPal) | 🌍 English phone | β€” |

| POCKET-EN-GGUF | PC-mix | 6.8 GB | PC / Android | 🌍 English, best quality | β€” |

*Wikipedia-Korean perplexity, lower is better. Q4_K_M = 5.79 baseline. English builds are tuned on English; see each repo.

> 🍎 Why MLX for Korean but GGUF for English on iPhone? Apple-native MLX only does uniform quantization. Korean survives it (96 experts hold up); English needs our proprietary quantization, which only GGUF supports β€” so the English iPhone build ships as a GGUF you run with PocketPal. Honest, not lazy.

> πŸ†• POCKET-26B β€” a Gemma4-26B-A4B-based sibling that loads in any app today (Ollama Β· LM Studio Β· PocketPal Β· MLX), no bleeding-edge runtime needed: GGUF (Q2_K 11 GB Β· Q4_K_M 17 GB Β· GPQA-Diamond 67%). Universal compatibility for 12 GB phones, PC, and browser.

!Speed vs Bonsai

Benchmarks β€” what is measured, what is not

We measure Bonsai on the same machine with the same stock llama.cpp, and we tell you where we lose.

[measured] Generation speed β€” POCKET wins on both CPU and GPU:

| | POCKET-35B IQ1_M | Bonsai-27B Q1_0 | |

|---|---|---|---|

| CPU generate (Xeon, 16t) | 27.0 tok/s | 10.1 | 🟒 2.69Γ— |

| GPU generate (H100) | 197 tok/s | 89 | 🟒 2.22Γ— |

| GPU prompt (H100) | 753 | 1816 | πŸ”΄ 0.41Γ— |

| Quality (HellaSwag, 400q) | 61.0% | 60.0% | βšͺ tie (CI overlaps) |

[measured on a MacBook M3 Pro, 18 GB] β€” and on a laptop, POCKET wins every axis, including prompt processing:

| | POCKET-35B IQ1_M | Bonsai-27B Q1_0 | |

|---|---|---|---|

| Metal generate (tg64) | 25.4 tok/s | 12.8 | 🟒 1.99Γ— |

| CPU generate (8 threads) | 13.8 tok/s | 4.4 | 🟒 3.13Γ— |

| Metal prompt (pp128) | 240.7 tok/s | 73.4 | 🟒 3.28Γ— |

| CPU prompt (pp128) | 45.5 tok/s | 9.6 | 🟒 4.75Γ— |

On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board. POCKET-35B-Q2_K runs on the M3 Pro's CPU at 19.5 tok/s β€” on an 18 GB Mac, run Q2_K on CPU (-ngl 0); its 13 GB exceeds the recommended Metal budget.

[measured β€” GPQA Diamond, 198q, greedy] reasoning quality vs quantization:

| Model | GPQA-Diamond (greedy) |

|---|---|

| Qwen3.6-35B-A3B | 73.2% |

| POCKET-35B Q4_K_M | 68.7% |

| POCKET-35B Q2_K | 60.1% |

[pending β€” community reports welcome] on-device iPhone and Strix Halo throughput. We publish only what we ran ourselves; help us fill the rest.

> The same-size rival Ternary-Bonsai-27B-Q2_0 (7.2 GB) fails to load in upstream llama.cpp β€” it needs the PrismML fork. POCKET runs on the tools you already have.

Files in this repo

| File | Size | bpw | Runs on | Korean PPL |

|---|---|---|---|---|

| POCKET-35B-Q4_K_M.gguf | 21 GB | 4.5 | PC 32 GB RAM | 5.79 (top) |

| POCKET-35B-Q3_K_M.gguf | 16 GB | 3.4 | PC 24 GB | 6.06 |

| POCKET-35B-Q2_K.gguf ⭐ | 13 GB | 2.6 | mini-PC 16–24 GB | 6.49 (best value) |

| POCKET-35B-IQ1_M.gguf | 8.2 GB | 1.9 | 16 GB RAM | 9.69 (smallest) |

Quickstart β€” no fork needed

# any recent llama.cpp β€” brew / winget / apt, or LM Studio / Ollama
llama-cli -m POCKET-35B-Q2_K.gguf -p "μ•ˆλ…•ν•˜μ„Έμš”" -ngl 0 -t 8
# reproduce our CPU numbers:
llama-bench -m POCKET-35B-IQ1_M.gguf -p 128 -n 64 -ngl 0 -t 16

Use physical-core count for -t (max ~32). Do not pass all threads β€” it can slow down sharply.

Lineage β€” where POCKET comes from

POCKET is quantized from Darwin-36B-Opus, VIDRAFT's flagship β€” a model bred and evolved over several generations on the Darwin platform (crossbreeding, healing, expert surgery). Darwin-36B-Opus itself traces back to a Qwen3.5-family MoE architecture.

| Component | Origin |

|---|---|

| Starting checkpoint | Darwin-36B-Opus β€” VIDRAFT, multi-generation Darwin evolution |

| Base architecture | Qwen3.5-family MoE (256 experts, top-8), unchanged |

| Quantization (Q4_K_M…IQ1_M) | stock llama.cpp β€” no custom format |

| Runtime | upstream llama.cpp / Apple MLX β€” unmodified |

| Proprietary language-specific tuning (KR/EN builds) | ours (VIDRAFT) |

The CPU/GPU speed comes from the sparse-MoE architecture plus ordinary quantization β€” reproducible with the same base and the same tools. What we add is the Darwin-evolved weights, the honest measurement, the Korean tuning, and the pruning that makes the 5 GB phone builds.

Limitations

  • The iPhone/Mac speed is not yet measured by us β€” community reports welcome.
  • Extreme quants (IQ1_M) hurt Korean ~2.8Γ— more than English; use Q2_K or larger for quality.
  • English phone builds trade quality for size; the PC build (PC-mix) is much closer to full quality.

License

Apache-2.0.

---

POCKET is a VIDRAFT model family. 35B, in your pocket. No GPU.

Learn more

<!-- POCKET-FAMILY -->

---

🧩 The POCKET Family β€” On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

πŸ“š Full POCKET collection

<!-- /POCKET-FAMILY -->

Run FINAL-Bench/POCKET-35B-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models