GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-MXFP4-GGUF overview

Gemma 4 12B it Uncensored Heretic MXFP4 GGUF MXFP4 GGUF quantization of llmfan46/gemma 4 12B it uncensored heretic https://huggingface.co/llmfan46/gemma 4 12B …

ggufgemma4mxfp4fp4visionmultimodaluncensoredhereticabliteratedimage-text-to-textenbase_model:llmfan46/gemma-4-12B-it-uncensored-hereticbase_model:quantized:llmfan46/gemma-4-12B-it-uncensored-hereticlicense:apache-2.0region:usconversational

Runs locally from ~116.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
image-text-to-text

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma4-12b-heretic-mxfp4.ggufGGUFGGUF6.18 GBDownload
mmproj-gemma-4-12b-heretic-f16.ggufGGUFF16116.4 MBDownload

Model Details

Model IDFreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-MXFP4-GGUF
AuthorFreedomAISVR
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelllmfan46/gemma-4-12B-it-uncensored-heretic
Last modified2026-07-06T13:10:15.000Z

Model README

---

license: apache-2.0

language:

  • en

library_name: gguf

tags:

  • gguf
  • gemma4
  • mxfp4
  • fp4
  • vision
  • multimodal
  • uncensored
  • heretic
  • abliterated

base_model: llmfan46/gemma-4-12B-it-uncensored-heretic

pipeline_tag: image-text-to-text

inference: false

quantized_by: FreedomAISVR

---

Gemma-4-12B-it-Uncensored-Heretic-MXFP4-GGUF

MXFP4 GGUF quantization of llmfan46/gemma-4-12B-it-uncensored-heretic - an uncensored/heretic (abliterated) finetune of Google's Gemma 4 12B with vision support.

About MXFP4

MXFP4 (Microscaling FP4) is an open standard (OCP) 4-bit floating point format (E2M1) supported by NVIDIA, AMD, Microsoft, and Meta.

  • Works on any GPU with MXFP4 support
  • Open standard - not vendor locked
  • Block-scaled format with shared scale factors

When to use MXFP4 vs other formats:

  • MXFP4 - Open standard, broad hardware support
  • NVFP4 - NVIDIA Blackwell native, best performance on RTX 50-series
  • Q4_K_M - Best for pre-Blackwell GPUs and CPU inference

Files

| File | Type | Size | Description |

|------|------|------|-------------|

| gemma4-12b-heretic-mxfp4.gguf | MXFP4 | ~6.3 GB | Text model (4.45 BPW) |

| mmproj-gemma-4-12b-heretic-f16.gguf | F16 | ~116 MB | Vision encoder (mmproj) |

Quantization Details

| Property | Value |

|----------|-------|

| Format | MXFP4 (E2M1) |

| Bits Per Weight | 4.45 BPW |

| Source Model | llmfan46/gemma-4-12B-it-uncensored-heretic |

| Architecture | Gemma4UnifiedForConditionalGeneration |

| Layers | 48 |

| Hidden Size | 3840 |

| Context Length | 262144 |

| Vision | Yes (Gemma4V projector) |

| Thinking | Enabled by default (opt-out via enable_thinking=false) |

Model Description

This is an abliterated (uncensored/heretic) finetune of Google's Gemma 4 12B, a multimodal model with both text and vision capabilities. The original model was finetuned to remove safety alignment restrictions while maintaining the model's core capabilities.

Gemma 4 features a hybrid attention architecture with alternating sliding window and full attention layers, native vision encoding, and tool calling support.

Usage

llama.cpp CLI

# Text only
./llama-cli -m gemma4-12b-heretic-mxfp4.gguf -p "Hello" -n 100

# With vision (requires mmproj)
./llama-server -m gemma4-12b-heretic-mxfp4.gguf \
  --mmproj mmproj-gemma-4-12b-heretic-f16.gguf \
  --host 0.0.0.0 --port 8080 -ngl 99

huggingface-hub

from huggingface_hub import hf_hub_download

model_path = hf_hub_download(
    repo_id="FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-MXFP4-GGUF",
    filename="gemma4-12b-heretic-mxfp4.gguf"
)
mmproj_path = hf_hub_download(
    repo_id="FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-MXFP4-GGUF",
    filename="mmproj-gemma-4-12b-heretic-f16.gguf"
)

Quantization Pipeline

  1. Download source: llmfan46/gemma-4-12B-it-uncensored-heretic
  2. Convert to F16 GGUF: convert_hf_to_gguf.py --outtype f16
  3. Extract mmproj: convert_hf_to_gguf.py --mmproj --outtype f16
  4. Quantize text: llama-quantize input-f16.gguf output-mxfp4.gguf MXFP4

Hardware Requirements

| Component | Requirement |

|-----------|-------------|

| GPU | Any with MXFP4 support, or CPU fallback |

| VRAM | ~7 GB minimum |

| RAM | ~16 GB recommended |

| Storage | ~7 GB |

License

Apache 2.0 (inherited from base model)

Run FreedomAISVR/Gemma-4-12B-it-Uncensored-Heretic-MXFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models