JMingo/gemma-4-12B-it-qat-GGUF overview
Near lossless GGUF quants of: google/gemma 4 12B it qat q4 0 unquantized https://huggingface.co/google/gemma 4 12B it qat q4 0 unquantized google/gemma 4 12B i…
Runs locally from ~167.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gemma-4-12B-it-qat-BF16.gguf | GGUF | BF16 | 22.20 GB | Download |
| gemma-4-12B-it-qat-Q4_0.gguf | GGUF | Q4_0 | 6.26 GB | Download |
| mmproj-gemma-4-12B-it-qat-BF16.gguf | GGUF | BF16 | 167.0 MB | Download |
| mtp-gemma-4-12B-it-qat-BF16.gguf | GGUF | BF16 | 821.6 MB | Download |
| mtp-gemma-4-12B-it-qat-Q4_0.gguf | GGUF | Q4_0 | 242.0 MB | Download |
Model Details
Model README
---
license: apache-2.0
base_model:
- google/gemma-4-12B-it-qat-q4_0-unquantized
tags:
- gemma4
---
Near-lossless GGUF quants of:
---
Original QAT Specifications
The base QAT model was trained with the following specifications:
- Text & Draft Models: 4-bit
- Multimodal Projector (
mmproj): BF16
---
Quantization Error Evaluation
- Dequantization: Quantized weights were dequantized to F32 for evaluation.
- Error Calculation: Quantization error metrics (compared to the BF16 baseline) were computed and accumulated in F64 precision to prevent numerical underflow and precision loss.
Metrics:
- MAE (Mean Absolute Error)
- RMSE (Root Mean Squared Error)
- Max Error
> [!NOTE]
> Interpretation of Metrics:
> These error metrics do not directly reflect actual model performance. This is because imatrix optimization intentionally increases measured error by downweighting less important parameters to improve overall performance. However, these metrics remain useful for estimating how close the quantized model is to the original, lossless quality.
gemma-4-12B-it-qat
| Model | MAE | RMSE | Max Error |
| :------------------------------------------------------------------------------------------------------------------------------------------------------ | :------------- | :------------- | :------------- |
| JMingo/gemma-4-12B-it-qat-BF16.gguf (Baseline) | 0.00000000 | 0.00000000 | 0.00000000 |
| JMingo/gemma-4-12B-it-qat-Q4_0.gguf | 0.00001161 | 0.00001969 | 0.00122070 |
| idkwhattoputherenow/gemma-4-12B-it-qat-q4_0-unquantized-q4_0-maxerr.gguf | 0.00001319 | 0.00002255 | 0.00122070 |
| unsloth/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf | 0.00001645 | 0.00002537 | 0.00390625 |
| google/gemma-4-12b-it-qat-q4_0.gguf | 0.00032522 | 0.00067853 | 0.01940918 |
| lmstudio-community/gemma-4-12B-it-QAT-Q4_0.gguf | 0.00032522 | 0.00067853 | 0.01940918 |
| mradermacher/gemma-4-12B-it-qat-q4_0-unquantized.i1-Q4_0.gguf | 0.00002872 | 0.00009840 | 0.04531860 |
| ggnoy/gemma-4-12B-it-qat-q4-bim-v2.gguf | 0.00002873 | 0.00009699 | 0.03878784 |
| dahara1/gemma-4-12B-it-qat-ja-Q4_0.gguf | 0.00005232 | 0.00021027 | 0.04718876 |
| dahara1/gemma-4-12B-it-qat-ja-Q8_0.gguf | 0.00005324 | 0.00006941 | 0.00268555 |
| dahara1/gemma-4-12B-it-qat-ja-UD-Q8_K_XL.gguf | 0.00004886 | 0.00006646 | 0.00268555 |
mtp-gemma-4-12B-it-qat
| Model | MAE | RMSE | Max Error |
| :------------------------------------------------------------------------------------------ | :------------- | :------------- | :------------- |
| JMingo/mtp-gemma-4-12B-it-qat-BF16.gguf (Baseline) | 0.00000000 | 0.00000000 | 0.00000000 |
| JMingo/mtp-gemma-4-12B-it-qat-Q4_0.gguf | 0.00002353 | 0.00003937 | 0.00219727 |
| unsloth/mtp-gemma-4-12B-it.gguf | 0.00003323 | 0.00005117 | 0.00390625 |
Run JMingo/gemma-4-12B-it-qat-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models