GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

FreedomAISVR/Gemma-4-E2B-it-QAT-NVFP4-GGUF overview

Gemma 4 E2B Instruct QAT + NVFP4 Hybrid GGUF QAT optimized weights preserved at Q4 0, overhead tensors quantized to NVFP4. What Makes This Different This is a …

ggufgemmagemma-4qatnvfp4blackwellvisionmultimodalenmultilingualbase_model:google/gemma-4-E2B-it-qat-q4_0-unquantizedbase_model:quantized:google/gemma-4-E2B-it-qat-q4_0-unquantizedlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~941.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-E2B-it-qat-nvfp4.ggufGGUFGGUF2.46 GBDownload
mmproj-gemma-4-E2B-it-qat.ggufGGUFGGUF941.1 MBDownload

Model Details

Model IDFreedomAISVR/Gemma-4-E2B-it-QAT-NVFP4-GGUF
AuthorFreedomAISVR
Pipeline
Licenseapache-2.0
Base modelgoogle/gemma-4-E2B-it-qat-q4_0-unquantized
Last modified2026-06-19T00:32:11.000Z

Model README

---

language:

  • en
  • multilingual

tags:

  • gemma
  • gemma-4
  • qat
  • nvfp4
  • blackwell
  • gguf
  • vision
  • multimodal

license: apache-2.0

base_model: google/gemma-4-E2B-it-qat-q4_0-unquantized

---

Gemma 4 E2B Instruct - QAT + NVFP4 Hybrid GGUF

QAT-optimized weights preserved at Q4_0, overhead tensors quantized to NVFP4.

What Makes This Different

This is a hybrid quantization of Google official QAT (Quantization-Aware Training) model. Instead of requantizing the Q4_0 weights (which breaks QAT benefits and vision quality), we:

  1. Kept all weight tensors at Q4_0 - attention, FFN, embeddings - exactly as Google trained them
  2. Quantized only the F32 norm/bias tensors to NVFP4 - these are the overhead tensors (layer norms, RMS norms, etc.)
  3. Used Google QAT mmproj - the vision projector trained alongside the QAT model

Why Standard NVFP4 from QAT Breaks Vision

Google QAT model was specifically trained to be resilient to Q4_0 quantization patterns. The weight values learned during QAT compensate for Q4_0 rounding. When you requantize Q4_0 -> F32 -> NVFP4, a second round of quantization error is introduced that QAT training did not account for. Vision tokens flow through the same attention/FFN layers - precision loss disproportionately degrades vision.

How the Hybrid Approach Works

Using llama-quantize --tensor-type-file with --allow-requantize:

llama-quantize --allow-requantize --tensor-type-file keep_q4.txt input.gguf output.gguf NVFP4

The tensor-type-file lists all Q4_0/Q4_K tensors to keep at their current type. When the quantizer sees cur_type == new_type, it copies the tensor data as-is - zero precision loss. Only the remaining F32 tensors are quantized to NVFP4.

Usage

# llama.cpp
llama-server -m gemma-4-E2B-it-qat-nvfp4.gguf --mmproj mmproj-gemma-4-E2B-it-qat.gguf -ngl 99

Source

  • Base model: google/gemma-4-E2B-it-qat-q4_0-unquantized
  • Quantized with: llama.cpp build 537 (commit d2c6795)
  • Chat template: Native Gemma 4 (thinking enabled by default)
  • Vision: Full multimodal support via QAT mmproj

Files

| File | Description |

|------|-------------|

| gemma-4-E2B-it-qat-nvfp4.gguf | Q4_0 weights + NVFP4 norms |

| mmproj-gemma-4-E2B-it-qat.gguf | QAT vision projector |

License

Apache 2.0 (same as base model)

Run FreedomAISVR/Gemma-4-E2B-it-QAT-NVFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models