FINAL-Bench/POCKET-26B-GGUF overview
π Collections βΆ POCKET Models https://huggingface.co/collections/FINAL Bench/pocket models 6a618ee5d23eafb7e185a5c6 β this family on device, no GPU Darwin Famβ¦
Runs locally from ~10.36 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | FINAL-Bench/POCKET-26B-GGUF |
|---|---|
| Author | FINAL-Bench |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | google/gemma-4-26B-A4B-it |
| Last modified | 2026-07-26T05:45:00.000Z |
Model README
---
license: apache-2.0
library_name: llama.cpp
pipeline_tag: text-generation
base_model:
- google/gemma-4-26B-A4B-it
tags:
- gguf
- llama.cpp
- conversational
- on-device
- mobile
- korean
- korean-llm
- cpu
- local-llm
- edge
- gemma
- gemma4
- mixture-of-experts
- moe
- vidraft
---
> ### π Collections
> βΆ POCKET Models β this family (on-device, no GPU)
> Darwin Family Β· Aether Foundation Β· VKAE Accelerated Β· Metacognition Adapters
POCKET-26B-GGUF Β· νκ΅μ΄
A Gemma4-26B-A4B-based pocket model that loads in any app today β Ollama, LM Studio, PocketPal β with no bleeding-edge runtime needed. Korean-tuned, GPU-optional.
> π Try it live on a CPU (no GPU), no install β  
  ![Compat]() 
Pick your build β    
Why this one?
POCKET-26B takes Google's Gemma4-26B-A4B (25.2B total, ~4B active MoE, Apache-2.0) and re-quantizes it with our proprietary Korean-tuned quantization β unpruned, so quality holds. Unlike our Qwen-based POCKET (which needs a very recent llama.cpp build for its qwen35moe architecture), Gemma4 loads in every mainstream runtime today: Ollama, LM Studio, PocketPal, koboldcpp, and the browser.
Quality β GPQA-Diamond, greedy, 198 questions (our harness)
| Build | GPQA-Diamond | vs base |
|---|---|---|
| Gemma4-26B-A4B (base) | 67.7% | β |
| POCKET-26B Q4_K_M | 67.7% | = base (lossless) |
| POCKET-26B Q2_K (mixed) β | 67.2% | β0.5pp (β lossless) |
Single greedy pass, 198 items β Β±~3 pp noise. Our proprietary Korean-tuned quantization is statistically lossless vs the base.
Files in this repo
| File | Size | Runs on | Best for |
|---|---|---|---|
| POCKET-26B-Q4_K_M.gguf | 17 GB | PC / high-RAM | top quality |
| POCKET-26B-Q2_K.gguf β | 11 GB | 12 GB phone / PC / browser | universal daily driver |
> Our mixed-precision quantization keeps the most quality-critical weights at higher precision β that is why Q2_K holds 67.2% while a plain uniform Q2 collapses to ~44%.
Quickstart β loads anywhere
# stock llama.cpp β brew / winget / apt, or LM Studio / Ollama / PocketPal
llama-cli -m POCKET-26B-Q2_K.gguf -p "λνλ―Όκ΅μ μλλ?" -ngl 0 -t 8
No fork, no bleeding-edge build β Gemma4 support has shipped in every mainstream runtime since April 2026.
Lineage (honest)
Based on google/gemma-4-26B-A4B-it (Apache-2.0). We do not re-host it unchanged β we add our proprietary Korean-tuned quantization (VIDRAFT). We deliberately do not prune it: Gemma4's low-bit robustness collapses under pruning (measured), so we keep all 128 experts and win on quality + universal compatibility instead.
Limitations
- For 8 GB phones (~5 GB budget), use POCKET-KR-GGUF (5.1 GB) β Gemma4 cannot be shrunk that far without collapse.
- On-device iPhone/Mac throughput not yet measured by us β community reports welcome.
Learn more
- Why on-device LLMs matter, and how POCKET measures up: Can you run a large LLM without a GPU?
- What model quantization is, and why a 4-bit model stays smart: What is model quantization?
License
Apache-2.0 β use, modify, redistribute freely.
---
POCKET is a VIDRAFT model family. Runs anywhere, no GPU.
<!-- POCKET-FAMILY -->
---
π§© The POCKET Family β On-device AI by VIDRAFT
Big models, small hardware. No GPU, no cloud.
Models
- π¦ POCKET-35B-GGUF β flagship, PC / server, no GPU
- π¦ POCKET-26B-GGUF β compact 26B
- π°π· POCKET-KR-GGUF β Korean, Android
- π POCKET-KR-MLX β Korean, iPhone / Mac
- π POCKET-EN-GGUF β English, phone / PC
Demos & tools (Spaces)
- πΌοΈ POCKET-Image β character-perfect Korean text in images
- π₯οΈ POCKET-35B-CPU β 35B answering on a CPU
- π₯οΈ POCKET-26B-CPU β 26B on a CPU
<!-- /POCKET-FAMILY -->
Run FINAL-Bench/POCKET-26B-GGUF with guIDE
Download guIDE β the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face Β· Compare models