FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF overview
Gemma 4 12B it Uncensored Heretic NVFP4 GGUF NVFP4 GGUF quantization of llmfan46/gemma 4 12B it uncensored heretic https://huggingface.co/llmfan46/gemma 4 12B …
Runs locally from ~116.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF |
|---|---|
| Author | FreedomAISVR |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | llmfan46/gemma-4-12B-it-uncensored-heretic |
| Last modified | 2026-07-06T13:10:11.000Z |
Model README
---
license: apache-2.0
language:
- en
library_name: gguf
tags:
- gguf
- gemma4
- nvfp4
- fp4
- blackwell
- vision
- multimodal
- uncensored
- heretic
- abliterated
base_model: llmfan46/gemma-4-12B-it-uncensored-heretic
pipeline_tag: image-text-to-text
inference: false
quantized_by: FreedomAISVR
---
Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF
NVFP4 GGUF quantization of llmfan46/gemma-4-12B-it-uncensored-heretic - an uncensored/heretic (abliterated) finetune of Google's Gemma 4 12B with vision support.
About NVFP4
NVFP4 is NVIDIA's native 4-bit floating point format (E4M3) designed for Blackwell architecture GPUs (RTX 50-series, B100/B200). It provides:
- Native tensor core acceleration on Blackwell GPUs
- Better dynamic range than INT4 formats due to floating point representation
- No dequantization overhead - processed directly in FP4
When to use NVFP4 vs other formats:
- NVFP4 - Best for Blackwell GPUs (RTX 5060 Ti, 5070, 5080, 5090, B100, B200)
- Q4_K_M - Best for pre-Blackwell GPUs and CPU inference
- MXFP4 - Open standard, works on any GPU with MX support
Files
| File | Type | Size | Description |
|------|------|------|-------------|
| gemma4-12b-heretic-nvfp4.gguf | NVFP4 | ~6.5 GB | Text model (4.68 BPW) |
| mmproj-gemma-4-12b-heretic-f16.gguf | F16 | ~116 MB | Vision encoder (mmproj) |
Quantization Details
| Property | Value |
|----------|-------|
| Format | NVFP4 (E4M3) |
| Bits Per Weight | 4.68 BPW |
| Source Model | llmfan46/gemma-4-12B-it-uncensored-heretic |
| Architecture | Gemma4UnifiedForConditionalGeneration |
| Layers | 48 |
| Hidden Size | 3840 |
| Context Length | 262144 |
| Vision | Yes (Gemma4V projector) |
| Thinking | Enabled by default (opt-out via enable_thinking=false) |
Model Description
This is an abliterated (uncensored/heretic) finetune of Google's Gemma 4 12B, a multimodal model with both text and vision capabilities. The original model was finetuned to remove safety alignment restrictions while maintaining the model's core capabilities.
Gemma 4 features a hybrid attention architecture with alternating sliding window and full attention layers, native vision encoding, and tool calling support.
Usage
llama.cpp CLI
# Text only
./llama-cli -m gemma4-12b-heretic-nvfp4.gguf -p "Hello" -n 100
# With vision (requires mmproj)
./llama-server -m gemma4-12b-heretic-nvfp4.gguf \
--mmproj mmproj-gemma-4-12b-heretic-f16.gguf \
--host 0.0.0.0 --port 8080 -ngl 99
LM Studio
- Download both files
- Load
gemma4-12b-heretic-nvfp4.ggufas the model - Load
mmproj-gemma-4-12b-heretic-f16.ggufas the mmproj - The model supports image inputs via the vision encoder
huggingface-hub
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(
repo_id="FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF",
filename="gemma4-12b-heretic-nvfp4.gguf"
)
mmproj_path = hf_hub_download(
repo_id="FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF",
filename="mmproj-gemma-4-12b-heretic-f16.gguf"
)
Quantization Pipeline
- Download source:
llmfan46/gemma-4-12B-it-uncensored-heretic - Convert to F16 GGUF:
convert_hf_to_gguf.py --outtype f16 - Extract mmproj:
convert_hf_to_gguf.py --mmproj --outtype f16 - Quantize text:
llama-quantize input-f16.gguf output-nvfp4.gguf NVFP4
Hardware Requirements
| Component | Requirement |
|-----------|-------------|
| GPU | NVIDIA Blackwell (RTX 50-series) for full acceleration |
| VRAM | ~7 GB minimum |
| RAM | ~16 GB recommended |
| Storage | ~7 GB |
License
Apache 2.0 (inherited from base model)
Run FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-NVFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models