GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

FreedomAISVR/Gemma-4-E2B-it-NVFP4-GGUF overview

Gemma 4 E2B IT — NVFP4 GGUF Base Model Model: google/gemma 4 E2B it https://huggingface.co/google/gemma 4 E2B it Architecture: Gemma4ForConditionalGeneration V…

llama.cppggufconversationalnvfp4visionbase_model:google/gemma-4-E2B-itbase_model:quantized:google/gemma-4-E2B-itlicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~940.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
661
Likes
1
Pipeline

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-e2b-it-nvfp4.ggufGGUFGGUF3.13 GBDownload
mmproj-gemma-4-E2B-it-f16.ggufGGUFF16940.0 MBDownload

Model Details

Model IDFreedomAISVR/Gemma-4-E2B-it-NVFP4-GGUF
AuthorFreedomAISVR
Pipeline
Licenseapache-2.0
Base modelgoogle/gemma-4-E2B-it
Last modified2026-09-17T06:14:51.000Z

Model README

---

tags:

- gguf

- conversational

- nvfp4

- vision

library_name: llama.cpp

license: apache-2.0

base_model: google/gemma-4-E2B-it

---

Gemma 4 E2B IT — NVFP4 GGUF

Base Model

  • Model: google/gemma-4-E2B-it
  • Architecture: Gemma4ForConditionalGeneration (Vision + Text)
  • Parameters: ~3B (E2B = Efficient 2B-class)
  • Context Length: 131,072 tokens (128K)
  • License: Apache 2.0
  • Update Date: July 20, 2026

Quantization Details

  • Format: NVFP4 (NVIDIA Blackwell FP4)
  • BPW: 5.76 bits per weight
  • File Size: 3.4 GB
  • Non-expert tensors: F16

Vision Support

  • mmproj: mmproj-gemma-4-E2B-it-f16.gguf (985 MB, F16)
  • Supports image and video input

Performance (RTX 5060 Ti 16GB)

  • Generation Speed: ~153 t/s at 128K context
  • Context: 128K with Q8_0 KV cache fits in 16GB VRAM

Usage (llama.cpp)

llama-cli -m gemma-4-e2b-it-nvfp4.gguf -ngl 99 --flash-attn on -c 131072 --cache-type-k q8_0 --cache-type-v q8_0

Why No MTP?

The Gemma 4 E2B model does not include MTP (Multi-Token Prediction) heads.

Hardware Target

  • NVIDIA RTX 50 series (Blackwell) for NVFP4 acceleration
  • 16GB+ VRAM recommended

Run FreedomAISVR/Gemma-4-E2B-it-NVFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models