FreedomAISVR/Gemma-4-E4B-it-QAT-MXFP4-GGUF overview
Gemma 4 E4B Instruct QAT + MXFP4 Hybrid GGUF QAT optimized weights preserved at Q4 0, overhead tensors quantized to MXFP4. What Makes This Different This is a …
Runs locally from ~945.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | FreedomAISVR/Gemma-4-E4B-it-QAT-MXFP4-GGUF |
|---|---|
| Author | FreedomAISVR |
| Pipeline | — |
| License | apache-2.0 |
| Base model | google/gemma-4-E4B-it-qat-q4_0-unquantized |
| Last modified | 2026-06-19T00:32:26.000Z |
Model README
---
language:
- en
- multilingual
tags:
- gemma
- gemma-4
- qat
- mxfp4
- blackwell
- gguf
- vision
- multimodal
license: apache-2.0
base_model: google/gemma-4-E4B-it-qat-q4_0-unquantized
---
Gemma 4 E4B Instruct - QAT + MXFP4 Hybrid GGUF
QAT-optimized weights preserved at Q4_0, overhead tensors quantized to MXFP4.
What Makes This Different
This is a hybrid quantization of Google official QAT (Quantization-Aware Training) model. Instead of requantizing the Q4_0 weights (which breaks QAT benefits and vision quality), we:
- Kept all weight tensors at Q4_0 - attention, FFN, embeddings - exactly as Google trained them
- Quantized only the F32 norm/bias tensors to MXFP4 - these are the overhead tensors (layer norms, RMS norms, etc.)
- Used Google QAT mmproj - the vision projector trained alongside the QAT model
Why Standard MXFP4 from QAT Breaks Vision
Google QAT model was specifically trained to be resilient to Q4_0 quantization patterns. The weight values learned during QAT compensate for Q4_0 rounding. When you requantize Q4_0 -> F32 -> MXFP4, a second round of quantization error is introduced that QAT training did not account for. Vision tokens flow through the same attention/FFN layers - precision loss disproportionately degrades vision.
How the Hybrid Approach Works
Using llama-quantize --tensor-type-file with --allow-requantize:
llama-quantize --allow-requantize --tensor-type-file keep_q4.txt input.gguf output.gguf MXFP4
The tensor-type-file lists all Q4_0/Q4_K tensors to keep at their current type. When the quantizer sees cur_type == new_type, it copies the tensor data as-is - zero precision loss. Only the remaining F32 tensors are quantized to MXFP4.
Usage
# llama.cpp
llama-server -m gemma-4-E4B-it-qat-mxfp4.gguf --mmproj mmproj-gemma-4-E4B-it-qat.gguf -ngl 99
Source
- Base model: google/gemma-4-E4B-it-qat-q4_0-unquantized
- Quantized with: llama.cpp build 537 (commit d2c6795)
- Chat template: Native Gemma 4 (thinking enabled by default)
- Vision: Full multimodal support via QAT mmproj
Files
| File | Description |
|------|-------------|
| gemma-4-E4B-it-qat-mxfp4.gguf | Q4_0 weights + MXFP4 norms |
| mmproj-gemma-4-E4B-it-qat.gguf | QAT vision projector |
License
Apache 2.0 (same as base model)
Run FreedomAISVR/Gemma-4-E4B-it-QAT-MXFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models