pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF overview
Gemma 4 E4B — Calibrated GGUFs and MTP Assistants This repository contains GGUF quantizations of base and QAT Gemma 4 E4B, calibrated on 15M tokens of pearsonk…
Runs locally from ~94.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf | GGUF | Q8_0 | 94.1 MB | Download |
| MTP/qat/mtp-gemma-4-E4B-it-qat-Q8_0.gguf | GGUF | Q8_0 | 94.1 MB | Download |
| gemma-4-E4B-it-IQ3_M-imatrix.gguf | GGUF | IQ3_M | 4.39 GB | Download |
| gemma-4-E4B-it-IQ4_XS-imatrix.gguf | GGUF | IQ4_XS | 4.72 GB | Download |
| gemma-4-E4B-it-qat-IQ3_M-imatrix.gguf | GGUF | IQ3_M | 4.37 GB | Download |
| gemma-4-E4B-it-qat-IQ4_XS-imatrix.gguf | GGUF | IQ4_XS | 4.69 GB | Download |
Model Details
| Model ID | pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF |
|---|---|
| Author | pearsonkyle |
| Pipeline | — |
| License | apache-2.0 |
| Base model | google/gemma-4-E4B-it,google/gemma-4-E4B-it-qat-q4_0-unquantized |
| Last modified | 2026-10-07T00:27:28.000Z |
Model README
---
library_name: gguf
license: apache-2.0
base_model:
- google/gemma-4-E4B-it
- google/gemma-4-E4B-it-qat-q4_0-unquantized
tags:
- gemma-4
- quantization
- imatrix
- gguf
- llama.cpp
- iq3_m
- iq4_xs
- speculative-decoding
- mtp
datasets:
- pearsonkyle/llmtk-sft-corpus-v2
- pearsonkyle/broad-domain-supplement
model-index:
- name: gemma-4-E4B-it-IQ3_M-imatrix
results: []
- name: gemma-4-E4B-it-IQ4_XS-imatrix
results: []
- name: gemma-4-E4B-it-qat-IQ3_M-imatrix
results: []
- name: gemma-4-E4B-it-qat-IQ4_XS-imatrix
results: []
---
Gemma-4 E4B — Calibrated GGUFs and MTP Assistants
This repository contains GGUF quantizations of base and QAT Gemma-4 E4B, calibrated on 15M tokens of pearsonkyle/llmtk-sft-corpus-v2 at 32k context. The original study compared imatrix and AWQ at IQ2_M, IQ3_M, and IQ4_XS on both models.
On 2026-10-06, 8 target GGUFs were removed using KLD > 1 in either of the two evaluation distributions (general and instruct/tools). Lower KLD indicates closer agreement with the F16 reference. The full study is preserved below; removed results are historical and their GGUFs are no longer offered on the default branch.
Available target files
| File | Size (GiB) | KLD (general) | KLD (instruct) |
|---|---|---|---|
| gemma-4-E4B-it-IQ3_M-imatrix.gguf | 4.391 | 0.342402 | 0.221945 |
| gemma-4-E4B-it-IQ4_XS-imatrix.gguf | 4.723 | 0.148996 | 0.093848 |
| gemma-4-E4B-it-qat-IQ3_M-imatrix.gguf | 4.365 | 0.255145 | 0.205045 |
| gemma-4-E4B-it-qat-IQ4_XS-imatrix.gguf | 4.691 | 0.057307 | 0.048981 |
The lowest measured KLD among the available files is gemma-4-E4B-it-qat-IQ4_XS-imatrix.gguf (0.057307 general, 0.048981 instruct). The retained IQ3_M files are smaller alternatives. These fidelity measurements compare each quant to its matching model reference; they do not establish task accuracy or superiority of one base model over another.
MTP speculative decoding
Google publishes matching assistants for base E4B and QAT E4B. This repository includes their Q8_0 GGUF conversions from Unsloth, approximately 94 MiB each. Use the base assistant with base targets and the QAT assistant with QAT targets.
| Target | Draft file |
|---|---|
| gemma-4-E4B-it | MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf |
| gemma-4-E4B-it-qat | MTP/qat/mtp-gemma-4-E4B-it-qat-Q8_0.gguf |
The target and assistant are separate GGUFs (gemma4 and gemma4-assistant architectures). No target weights were changed. Select the assistant explicitly: this repository contains two target families, and filename-based draft auto-discovery does not establish the correct pairing.
With a current llama.cpp build supporting E4B MTP (support added in PR #24282), download your chosen target and its matching assistant. For example:
hf download pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF \
gemma-4-E4B-it-IQ3_M-imatrix.gguf MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf \
--local-dir gemma-e4b
llama-server \
-m gemma-e4b/gemma-4-E4B-it-IQ3_M-imatrix.gguf \
--model-draft gemma-e4b/MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf \
--spec-type draft-mtp --spec-draft-n-max 3 \
-ngl 99 --gpu-layers-draft 99 -c 8192 -np 1 -fa off
For a QAT target, replace the target filename with gemma-4-E4B-it-qat-IQ3_M-imatrix.gguf (or its IQ4_XS counterpart), and replace the assistant path in both commands with MTP/qat/mtp-gemma-4-E4B-it-qat-Q8_0.gguf. The example uses -fa off because the upstream E4B draft instructions report a CUDA flash-attention failure. Metal was tested here with -fa on.
Validation: both assistants loaded and actively drafted against base and QAT IQ3_M/IQ4_XS on Apple Silicon Metal with llama.cpp b9692 (f3e182816). Two short prompts per target, 96-token cap, greedy sampling, 2k context, and draft depth 3 gave acceptance of 46–53%. Seven of eight generated answers matched the corresponding baseline exactly; one QAT IQ3_M answer differed. These are smoke checks, not a quality or long-context benchmark. No consistent speed improvement was observed on this Mac; measure throughput with your workload before enabling MTP by default.
Source revisions, SHA-256 checksums, and validation counts are in MTP/validation.json. The converted drafts are copied unchanged from Unsloth base E4B and Unsloth QAT E4B.
Original benchmark study
| Model | Quant | Method | Size (GiB) | BPW | KLD (general) | KLD (instruct) | Top-P (general) | Top-P (instruct) | Decode (tok/s) | Availability |
|---|---|---|---|---|---|---|---|---|---|---|
| gemma-4-E4B-it | IQ2_M | imatrix | 3.550 | 4.056 | 1.461 | 0.978 | 50.0 | 61.8 | 50.3 | Removed |
| gemma-4-E4B-it | IQ2_M | AWQ | 3.527 | 4.030 | 12.266 | 13.830 | 0.1 | 0.1 | 69.9 | Removed |
| gemma-4-E4B-it | IQ3_M | imatrix | 4.391 | 5.017 | 0.342 | 0.222 | 74.9 | 81.1 | 68.8 | Available |
| gemma-4-E4B-it | IQ3_M | AWQ | 4.365 | 4.988 | 16.393 | 17.650 | 0.3 | 0.3 | 63.5 | Removed |
| gemma-4-E4B-it | IQ4_XS | imatrix | 4.723 | 5.396 | 0.149 | 0.094 | 82.8 | 87.6 | 59.5 | Available |
| gemma-4-E4B-it | IQ4_XS | AWQ | 4.691 | 5.360 | 16.443 | 17.562 | 0.03 | 0.02 | 40.8 | Removed |
| gemma-4-E4B-it-qat | IQ2_M | imatrix | 3.527 | 4.060 | 4.888 | 4.940 | 21.9 | 23.8 | 76.2 | Removed |
| gemma-4-E4B-it-qat | IQ2_M | AWQ | 3.527 | 4.060 | 18.363 | 20.090 | 0.1 | 0.0 | 65.1 | Removed |
| gemma-4-E4B-it-qat | IQ3_M | imatrix | 4.365 | 5.025 | 0.255 | 0.205 | 76.6 | 80.1 | 56.6 | Available |
| gemma-4-E4B-it-qat | IQ3_M | AWQ | 4.365 | 5.025 | 15.942 | 17.118 | 0.2 | 0.1 | 59.5 | Removed |
| gemma-4-E4B-it-qat | IQ4_XS | imatrix | 4.691 | 5.400 | 0.057 | 0.049 | 88.0 | 89.8 | 86.1 | Available |
| gemma-4-E4B-it-qat | IQ4_XS | AWQ | 4.691 | 5.400 | 14.067 | 15.494 | 0.9 | 0.7 | 68.4 | Removed |
The archived measurements are also available as original-benchmarks.csv. The original model card remains accessible at the pre-cleanup revision.
Methodology
- Calibration: 15M tokens from
pearsonkyle/llmtk-sft-corpus-v232k split (seed-42 shuffle, 9,597 sessions, 32k context). Imatrix collected viallama-imatrix -c 32768 --parse-special; hybrid_custom variant re-weights per-tensor. AWQ α-search via proxy quantizer, imatrix collected on folded F16. - GQA fix: gemma-4 has 4G GQA (8 Q / 2 KV heads, 42 layers, 24 KV-grouped layers).
llama-imatrixonly collects K/V stats for layers 0–23; a_backfill_missing_kv_layerspatch copies per-channel mean of collected K/V vectors onto the 18 missing layers. - Evaluation: KLD + perplexity + top-p agreement via
llama-perplexityonpearsonkyle/broad-domain-supplement(general-domain: 30,710 tokens; instruct/tools: 30,853 tokens). Speed viallama-bench. - Machine: Apple Silicon (MPS), llama.cpp vendored build.
Run pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models