JMingo/gemma-4-E2B-it-qat-GGUF overview
Near lossless GGUF quants of: google/gemma 4 E2B it qat q4 0 unquantized https://huggingface.co/google/gemma 4 E2B it qat q4 0 unquantized google/gemma 4 E2B i…
Runs locally from ~56.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gemma-4-E2B-it-qat-BF16.gguf | GGUF | BF16 | 8.64 GB | Download |
| gemma-4-E2B-it-qat-Q4_0.gguf | GGUF | Q4_0 | 2.44 GB | Download |
| mmproj-gemma-4-E2B-it-qat-BF16.gguf | GGUF | BF16 | 941.1 MB | Download |
| mmproj-gemma-4-E2B-it-qat-Q2_K_XL.gguf | GGUF | Q2_K_XL | 350.3 MB | Download |
| mtp-gemma-4-E2B-it-qat-BF16.gguf | GGUF | BF16 | 162.3 MB | Download |
| mtp-gemma-4-E2B-it-qat-Q4_0.gguf | GGUF | Q4_0 | 56.5 MB | Download |
Model Details
Model README
---
license: apache-2.0
base_model:
- google/gemma-4-E2B-it-qat-q4_0-unquantized
tags:
- gemma4
---
Near-lossless GGUF quants of:
Uses the chat template from google/gemma-4-E2B-it (2026-05-18), replacing the outdated chat template that was originally bundled with the base QAT model.
---
Original QAT Specifications
The base QAT model was trained with the following specifications:
- Text & Draft Models: 4-bit
- Multimodal Projector (
mmproj):
- Audio Encoder: Mostly 2-bit (with some tensors in 4-bit or BF16)
- Vision Encoder: 8-bit
- Other parts: BF16
---
Quantization Error Evaluation
- Dequantization: Quantized weights were dequantized to F32 for evaluation.
- Error Calculation: Quantization error metrics (compared to the BF16 baseline) were computed and accumulated in F64 precision to prevent numerical underflow and precision loss.
Metrics:
- MAE (Mean Absolute Error)
- RMSE (Root Mean Squared Error)
- Max Error
> [!NOTE]
> Interpretation of Metrics:
> These error metrics do not directly reflect actual model performance. This is because imatrix optimization intentionally increases measured error by downweighting less important parameters to improve overall performance. However, these metrics remain useful for estimating how close the quantized model is to the original, lossless quality.
gemma-4-E2B-it-qat
| Model | MAE | RMSE | Max Error |
| :------------------------------------------------------------------------------------------------------------------------------------------------------ | :------------- | :------------- | :------------- |
| JMingo/gemma-4-E2B-it-qat-BF16.gguf (Baseline) | 0.00000000 | 0.00000000 | 0.00000000 |
| JMingo/gemma-4-E2B-it-qat-Q4_0.gguf | 0.00003548 | 0.00006434 | 0.00256348 |
| idkwhattoputherenow/gemma-4-E2B-it-qat-q4_0-unquantized-q4_0-maxerr.gguf | 0.00004027 | 0.00007371 | 0.00299072 |
| unsloth/gemma-4-E2B-it-qat-UD-Q4_K_XL.gguf | 0.00005024 | 0.00008256 | 0.00341797 |
| google/gemma-4-E2B_q4_0-it.gguf | 0.00052373 | 0.00093318 | 0.06982422 |
| lmstudio-community/gemma-4-E2B-it-QAT-Q4_0.gguf | 0.00052373 | 0.00093318 | 0.06982422 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-Q4_0.gguf | 0.00032190 | 0.00060744 | 0.03930664 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.Q4_K_S.gguf | 0.00054283 | 0.00086672 | 0.03857422 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.Q4_K_M.gguf | 0.00053254 | 0.00086113 | 0.03857422 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-Q4_1.gguf | 0.00048704 | 0.00084594 | 0.13549805 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-Q4_K_S.gguf | 0.00053921 | 0.00086299 | 0.14585495 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-Q4_K_M.gguf | 0.00052917 | 0.00085745 | 0.14585495 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.IQ4_XS.gguf | 0.00078202 | 0.00116219 | 0.05517125 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-IQ4_XS.gguf | 0.00075868 | 0.00114195 | 0.31855154 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-IQ4_NL.gguf | 0.00075354 | 0.00113171 | 0.35578442 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.Q3_K_L.gguf | 0.00113207 | 0.00216116 | 0.12817383 |
| mradermacher/gemma-4-E2B-it-qat-q4_0-unquantized.i1-Q3_K_L.gguf | 0.00109405 | 0.00208034 | 0.31637573 |
mtp-gemma-4-E2B-it-qat
| Model | MAE | RMSE | Max Error |
| :------------------------------------------------------------------------------------------ | :------------- | :------------- | :------------- |
| JMingo/mtp-gemma-4-E2B-it-qat-BF16.gguf (Baseline) | 0.00000000 | 0.00000000 | 0.00000000 |
| JMingo/mtp-gemma-4-E2B-it-qat-Q4_0.gguf | 0.00005037 | 0.00008262 | 0.00219727 |
| unsloth/mtp-gemma-4-E2B-it.gguf | 0.00007164 | 0.00010655 | 0.00781250 |
mmproj-gemma-4-E2B-it-qat
| Model | MAE | RMSE | Max Error |
| :------------------------------------------------------ | :------------- | :------------- | :------------- |
| JMingo/mmproj-gemma-4-E2B-it-qat-BF16.gguf (Baseline) | 0.00000000 | 0.00000000 | 0.00000000 |
| JMingo/mmproj-gemma-4-E2B-it-qat-Q2_K_XL.gguf | 0.00001210 | 0.00003352 | 0.00271225 |
Run JMingo/gemma-4-E2B-it-qat-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models