GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF overview

Gemma 4 E4B — Calibrated GGUFs and MTP Assistants This repository contains GGUF quantizations of base and QAT Gemma 4 E4B, calibrated on 15M tokens of pearsonk…

ggufgemma-4quantizationimatrixllama.cppiq3_miq4_xsspeculative-decodingmtpdataset:pearsonkyle/llmtk-sft-corpus-v2dataset:pearsonkyle/broad-domain-supplementbase_model:google/gemma-4-E4B-itbase_model:quantized:google/gemma-4-E4B-itlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~94.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
3,637
Likes
1
Pipeline
—

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
MTP/base/mtp-gemma-4-E4B-it-Q8_0.ggufGGUFQ8_094.1 MBDownload
MTP/qat/mtp-gemma-4-E4B-it-qat-Q8_0.ggufGGUFQ8_094.1 MBDownload
gemma-4-E4B-it-IQ3_M-imatrix.ggufGGUFIQ3_M4.39 GBDownload
gemma-4-E4B-it-IQ4_XS-imatrix.ggufGGUFIQ4_XS4.72 GBDownload
gemma-4-E4B-it-qat-IQ3_M-imatrix.ggufGGUFIQ3_M4.37 GBDownload
gemma-4-E4B-it-qat-IQ4_XS-imatrix.ggufGGUFIQ4_XS4.69 GBDownload

Model Details

Model IDpearsonkyle/gemma-4-E4B-imatrix-awq-GGUF
Authorpearsonkyle
Pipeline—
Licenseapache-2.0
Base modelgoogle/gemma-4-E4B-it,google/gemma-4-E4B-it-qat-q4_0-unquantized
Last modified2026-10-07T00:27:28.000Z

Model README

---

library_name: gguf

license: apache-2.0

base_model:

  • google/gemma-4-E4B-it
  • google/gemma-4-E4B-it-qat-q4_0-unquantized

tags:

  • gemma-4
  • quantization
  • imatrix
  • gguf
  • llama.cpp
  • iq3_m
  • iq4_xs
  • speculative-decoding
  • mtp

datasets:

  • pearsonkyle/llmtk-sft-corpus-v2
  • pearsonkyle/broad-domain-supplement

model-index:

  • name: gemma-4-E4B-it-IQ3_M-imatrix

results: []

  • name: gemma-4-E4B-it-IQ4_XS-imatrix

results: []

  • name: gemma-4-E4B-it-qat-IQ3_M-imatrix

results: []

  • name: gemma-4-E4B-it-qat-IQ4_XS-imatrix

results: []

---

Gemma-4 E4B — Calibrated GGUFs and MTP Assistants

This repository contains GGUF quantizations of base and QAT Gemma-4 E4B, calibrated on 15M tokens of pearsonkyle/llmtk-sft-corpus-v2 at 32k context. The original study compared imatrix and AWQ at IQ2_M, IQ3_M, and IQ4_XS on both models.

On 2026-10-06, 8 target GGUFs were removed using KLD > 1 in either of the two evaluation distributions (general and instruct/tools). Lower KLD indicates closer agreement with the F16 reference. The full study is preserved below; removed results are historical and their GGUFs are no longer offered on the default branch.

Available target files

| File | Size (GiB) | KLD (general) | KLD (instruct) |

|---|---|---|---|

| gemma-4-E4B-it-IQ3_M-imatrix.gguf | 4.391 | 0.342402 | 0.221945 |

| gemma-4-E4B-it-IQ4_XS-imatrix.gguf | 4.723 | 0.148996 | 0.093848 |

| gemma-4-E4B-it-qat-IQ3_M-imatrix.gguf | 4.365 | 0.255145 | 0.205045 |

| gemma-4-E4B-it-qat-IQ4_XS-imatrix.gguf | 4.691 | 0.057307 | 0.048981 |

The lowest measured KLD among the available files is gemma-4-E4B-it-qat-IQ4_XS-imatrix.gguf (0.057307 general, 0.048981 instruct). The retained IQ3_M files are smaller alternatives. These fidelity measurements compare each quant to its matching model reference; they do not establish task accuracy or superiority of one base model over another.

MTP speculative decoding

Google publishes matching assistants for base E4B and QAT E4B. This repository includes their Q8_0 GGUF conversions from Unsloth, approximately 94 MiB each. Use the base assistant with base targets and the QAT assistant with QAT targets.

| Target | Draft file |

|---|---|

| gemma-4-E4B-it | MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf |

| gemma-4-E4B-it-qat | MTP/qat/mtp-gemma-4-E4B-it-qat-Q8_0.gguf |

The target and assistant are separate GGUFs (gemma4 and gemma4-assistant architectures). No target weights were changed. Select the assistant explicitly: this repository contains two target families, and filename-based draft auto-discovery does not establish the correct pairing.

With a current llama.cpp build supporting E4B MTP (support added in PR #24282), download your chosen target and its matching assistant. For example:

hf download pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF \
  gemma-4-E4B-it-IQ3_M-imatrix.gguf MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf \
  --local-dir gemma-e4b

llama-server \
  -m gemma-e4b/gemma-4-E4B-it-IQ3_M-imatrix.gguf \
  --model-draft gemma-e4b/MTP/base/mtp-gemma-4-E4B-it-Q8_0.gguf \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  -ngl 99 --gpu-layers-draft 99 -c 8192 -np 1 -fa off

For a QAT target, replace the target filename with gemma-4-E4B-it-qat-IQ3_M-imatrix.gguf (or its IQ4_XS counterpart), and replace the assistant path in both commands with MTP/qat/mtp-gemma-4-E4B-it-qat-Q8_0.gguf. The example uses -fa off because the upstream E4B draft instructions report a CUDA flash-attention failure. Metal was tested here with -fa on.

Validation: both assistants loaded and actively drafted against base and QAT IQ3_M/IQ4_XS on Apple Silicon Metal with llama.cpp b9692 (f3e182816). Two short prompts per target, 96-token cap, greedy sampling, 2k context, and draft depth 3 gave acceptance of 46–53%. Seven of eight generated answers matched the corresponding baseline exactly; one QAT IQ3_M answer differed. These are smoke checks, not a quality or long-context benchmark. No consistent speed improvement was observed on this Mac; measure throughput with your workload before enabling MTP by default.

Source revisions, SHA-256 checksums, and validation counts are in MTP/validation.json. The converted drafts are copied unchanged from Unsloth base E4B and Unsloth QAT E4B.

Original benchmark study

| Model | Quant | Method | Size (GiB) | BPW | KLD (general) | KLD (instruct) | Top-P (general) | Top-P (instruct) | Decode (tok/s) | Availability |

|---|---|---|---|---|---|---|---|---|---|---|

| gemma-4-E4B-it | IQ2_M | imatrix | 3.550 | 4.056 | 1.461 | 0.978 | 50.0 | 61.8 | 50.3 | Removed |

| gemma-4-E4B-it | IQ2_M | AWQ | 3.527 | 4.030 | 12.266 | 13.830 | 0.1 | 0.1 | 69.9 | Removed |

| gemma-4-E4B-it | IQ3_M | imatrix | 4.391 | 5.017 | 0.342 | 0.222 | 74.9 | 81.1 | 68.8 | Available |

| gemma-4-E4B-it | IQ3_M | AWQ | 4.365 | 4.988 | 16.393 | 17.650 | 0.3 | 0.3 | 63.5 | Removed |

| gemma-4-E4B-it | IQ4_XS | imatrix | 4.723 | 5.396 | 0.149 | 0.094 | 82.8 | 87.6 | 59.5 | Available |

| gemma-4-E4B-it | IQ4_XS | AWQ | 4.691 | 5.360 | 16.443 | 17.562 | 0.03 | 0.02 | 40.8 | Removed |

| gemma-4-E4B-it-qat | IQ2_M | imatrix | 3.527 | 4.060 | 4.888 | 4.940 | 21.9 | 23.8 | 76.2 | Removed |

| gemma-4-E4B-it-qat | IQ2_M | AWQ | 3.527 | 4.060 | 18.363 | 20.090 | 0.1 | 0.0 | 65.1 | Removed |

| gemma-4-E4B-it-qat | IQ3_M | imatrix | 4.365 | 5.025 | 0.255 | 0.205 | 76.6 | 80.1 | 56.6 | Available |

| gemma-4-E4B-it-qat | IQ3_M | AWQ | 4.365 | 5.025 | 15.942 | 17.118 | 0.2 | 0.1 | 59.5 | Removed |

| gemma-4-E4B-it-qat | IQ4_XS | imatrix | 4.691 | 5.400 | 0.057 | 0.049 | 88.0 | 89.8 | 86.1 | Available |

| gemma-4-E4B-it-qat | IQ4_XS | AWQ | 4.691 | 5.400 | 14.067 | 15.494 | 0.9 | 0.7 | 68.4 | Removed |

The archived measurements are also available as original-benchmarks.csv. The original model card remains accessible at the pre-cleanup revision.

Methodology

  • Calibration: 15M tokens from pearsonkyle/llmtk-sft-corpus-v2 32k split (seed-42 shuffle, 9,597 sessions, 32k context). Imatrix collected via llama-imatrix -c 32768 --parse-special; hybrid_custom variant re-weights per-tensor. AWQ α-search via proxy quantizer, imatrix collected on folded F16.
  • GQA fix: gemma-4 has 4G GQA (8 Q / 2 KV heads, 42 layers, 24 KV-grouped layers). llama-imatrix only collects K/V stats for layers 0–23; a _backfill_missing_kv_layers patch copies per-channel mean of collected K/V vectors onto the 18 missing layers.
  • Evaluation: KLD + perplexity + top-p agreement via llama-perplexity on pearsonkyle/broad-domain-supplement (general-domain: 30,710 tokens; instruct/tools: 30,853 tokens). Speed via llama-bench.
  • Machine: Apple Silicon (MPS), llama.cpp vendored build.

Run pearsonkyle/gemma-4-E4B-imatrix-awq-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models