SC117/gemma-4-12B-it-heretic-QAT-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; display: flex; flex direction: column; gap: 20px; margin bottom: 30p…
Runs locally from ~167.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
Model README
---
license: apache-2.0
base_model:
- coder3101/gemma-4-12B-it-qat-q4_0-unquantized-heretic
- google/gemma-4-12B-it
tags:
- gguf
- quantized
- qat
- heretic
- uncensored
- abliterated
- gemma4
language:
- en
- zh
- multilingual
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 16px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 8px rgba(0,0,0,0.04);">
<div style="background: linear-gradient(135deg, #f59e0b 0%, #d97706 100%); padding: 28px 24px; color: white; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: rgba(255,255,255,0.1); border-radius: 50%;"></div>
<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: rgba(255,255,255,0.08); border-radius: 50%;"></div>
<h1 style="margin: 0; font-size: 24px; font-weight: 700; position: relative; z-index: 1;">⚡ Gemma 4 12B Heretic QAT — Q4_0 GGUF</h1>
<p style="margin: 8px 0 0 0; font-size: 15px; opacity: 0.95; position: relative; z-index: 1;">Heretic ARA · QAT-Lossless Q4_0 · 6.4 GB · Encoder-Free Multimodal</p>
<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;">
<a href="https://huggingface.co/SC117/gemma-4-12B-it-heretic-QAT-GGUF/blob/main/README_zh.md" style="color: #ffffff; text-decoration: none;">📖 中文文档</a></p>
</div>
<div style="padding: 16px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0; display: flex; flex-wrap: wrap; gap: 8px;">
<span style="background: #fef3c7; color: #92400e; padding: 4px 12px; border-radius: 20px; font-size: 12px; font-weight: 600;">Q4_0</span>
<span style="background: #dbeafe; color: #1e40af; padding: 4px 12px; border-radius: 20px; font-size: 12px; font-weight: 600;">12B Dense</span>
<span style="background: #fce7f3; color: #9d174d; padding: 4px 12px; border-radius: 20px; font-size: 12px; font-weight: 600;">Heretic Uncensored</span>
<span style="background: #d1fae5; color: #065f46; padding: 4px 12px; border-radius: 20px; font-size: 12px; font-weight: 600;">6.4 GB</span>
<span style="background: #ede9fe; color: #5b21b6; padding: 4px 12px; border-radius: 20px; font-size: 12px; font-weight: 600;">QAT Weights</span>
<span style="background: #e0f2fe; color: #075985; padding: 4px 12px; border-radius: 20px; font-size: 12px; font-weight: 600;">Encoder-Free</span>
</div>
<div style="padding: 20px 24px;">
<p style="margin: 0; font-size: 14px; color: #334155; line-height: 1.7;">
Uncensored version of <a href="https://huggingface.co/google/gemma-4-12B-it" style="color: #c2410c; text-decoration: none;">Google Gemma 4 12B IT (QAT)</a>, processed with <a href="https://github.com/p-e-w/heretic" style="color: #c2410c; text-decoration: none;">Heretic</a> ARA abliteration. Quantized to <b>Q4_0</b> matching <a href="https://huggingface.co/unsloth/gemma-4-12b-it-qat-GGUF" style="color: #c2410c; text-decoration: none;">Unsloth's UD-Q4_K_XL</a> format — QAT weights trained for 4-bit quantization, near-lossless quality.
</p>
</div>
</div>
<div style="border: 2px solid #dc2626; border-radius: 12px; overflow: hidden; background: #fef2f2; box-shadow: 0 2px 6px rgba(220,38,38,0.12);">
<div style="background: linear-gradient(135deg, #b91c1c 0%, #dc2626 100%); padding: 13px 16px; color: #ffffff; font-weight: 800; font-size: 15px; display: flex; align-items: center; gap: 8px;">
<span>⚠️</span> Important Notice: UD-Q4_K_XL and Google QAT
</div>
<div style="padding: 16px 18px; color: #7f1d1d; font-size: 13px; line-height: 1.75;">
<p style="margin: 0 0 10px 0;">
<b>UD-Q4_K_XL is the UD team's naming convention for its Gemma QAT Q4_0 GGUF variant.</b>
The repository or file name must not be interpreted as proof that the model uses mixed K-quants.
In this release, the actual GGUF tensor information shows that the main model weights are stored as
<b>Q4_0</b>, while normalization and scaling tensors remain in <b>F32</b>.
</p>
<p style="margin: 0 0 10px 0;">
This model is derived from Google's QAT weights, which were trained for 4-bit quantization.
Claims that this file was converted to UD mixed precision, or that its QAT structure was damaged by
mixed K-quant remapping, are therefore not supported by the actual tensor types contained in the GGUF.
</p>
<p style="margin: 0; font-weight: 700;">
Before opening a discussion or reporting a quantization issue, please inspect the GGUF metadata and
tensor types first. Do not infer the internal quantization layout from the filename alone, and do not
make unsupported technical claims without checking the source information.
</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #a855f7 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>✂️</span> Heretic ARA Abliteration Parameters</div>
<div style="padding: 16px;">
<p style="margin: 0 0 8px 0; font-size: 13px; color: #64748b; font-style: italic; border-left: 3px solid #ddd6fe; padding-left: 12px;">Base: <a href="https://huggingface.co/coder3101/gemma-4-12B-it-qat-q4_0-unquantized-heretic" style="color: #c2410c; text-decoration: none;">coder3101/heretic-QAT</a> · Heretic v1.2.0 · ARA + Row-Norm</p>
<table style="width: 100%; border-collapse: collapse; font-family: inherit; font-size: 13px;">
<thead>
<tr style="background: rgba(124,58,237,0.05);">
<th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #5b21b6; font-weight: bold;">Parameter</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #5b21b6; font-weight: bold;">Value</th>
</tr>
</thead>
<tbody>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">start_layer_index</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">24</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">end_layer_index</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">48</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">preserve_good_behavior_weight</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">0.3707</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">steer_bad_behavior_weight</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">0.0010</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">overcorrect_relative_weight</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">0.6177</td></tr>
<tr><td style="padding: 8px 10px; color: #334155;">neighbor_count</td><td style="padding: 8px 10px; color: #334155;">15</td></tr>
</tbody>
</table>
<table style="width: 100%; border-collapse: collapse; font-family: inherit; font-size: 13px; margin-top: 12px;">
<thead>
<tr style="background: rgba(124,58,237,0.05);">
<th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #5b21b6; font-weight: bold;">Metric</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: center; color: #5b21b6; font-weight: bold;">Heretic</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: center; color: #5b21b6; font-weight: bold;">Original QAT</th>
</tr>
</thead>
<tbody>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">KL Divergence</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155; text-align: center;">0.0575</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155; text-align: center;">0 (by definition)</td></tr>
<tr><td style="padding: 8px 10px; color: #334155;">Refusals</td><td style="padding: 8px 10px; color: #334155; text-align: center;">8/100</td><td style="padding: 8px 10px; color: #334155; text-align: center;">99/100</td></tr>
</tbody>
</table>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #059669 0%, #10b981 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🏗️</span> Architecture</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-family: inherit; font-size: 13px;">
<tbody>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600; width: 40%;">Base Model</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;"><a href="https://huggingface.co/google/gemma-4-12B-it" style="color: #c2410c; text-decoration: none;">google/gemma-4-12B-it</a></td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Parameters</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">11.95B (dense, all parameters active)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Architecture</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">Encoder-free unified multimodal (text + image + audio + video)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Layers</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">48</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Hidden Size</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">3,840</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Attention</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">16 heads, GQA with 8 KV heads, head dim 256</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Context Length</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">256K tokens (hybrid sliding window 1024 + global attention)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Vocabulary</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">262K, 140+ languages</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">Modalities</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">Text + Image + Audio + Video (encoder-free, native multimodal)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #047857; font-weight: 600;">QAT Training</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">Google official QAT (quantization-aware), weights inherently robust to Q4_0</td></tr>
<tr><td style="padding: 8px 10px; color: #047857; font-weight: 600;">Quantization</td><td style="padding: 8px 10px; color: #334155;">Q4_0 (matching Unsloth UD-Q4_K_XL layout), b9553 llama-quantize</td></tr>
</tbody>
</table>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #0891b2 0%, #06b6d4 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📊</span> Quantization Details</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-family: inherit; font-size: 13px;">
<tbody>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #0e7490; font-weight: 600; width: 40%;">Format</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">Q4_0 (uniform — QAT weights optimized for this exact precision)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #0e7490; font-weight: 600;">File Size</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">6.4 GB</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #0e7490; font-weight: 600;">Effective BPW</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">4.50 (all weight tensors Q4_0, norms F32)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #0e7490; font-weight: 600;">Tool</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">llama-quantize (b9553, CUDA 13.3)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #0e7490; font-weight: 600;">Source</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">BF16 GGUF (converted from QAT heretic safetensors)</td></tr>
<tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #0e7490; font-weight: 600;">Context Length</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155;">256K (set in GGUF metadata)</td></tr>
<tr><td style="padding: 8px 10px; color: #0e7490; font-weight: 600;">QAT Advantage</td><td style="padding: 8px 10px; color: #334155;">Q4_0 with QAT weights achieves 88.8% Top-1 vs 74.1% naive Q4_0 (+14.7%)</td></tr>
</tbody>
</table>
<p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; line-height: 1.5;">Why Q4_0? Google's QAT trains weights to be optimal at Q4_0 noise levels. Unsloth's UD-Q4_K_XL uses the same Q4_0 layout — the "dynamic" advantage comes from conversion precision, not per-tensor mixing.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ea580c 0%, #f97316 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚙️</span> Recommended Sampling Parameters</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-family: inherit; font-size: 13px; margin-bottom: 16px;">
<tbody>
<tr><td style="padding: 6px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155; font-weight: 600; width: 35%;">General</td><td style="padding: 6px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); color: #334155; font-family: monospace; font-size: 12px;">temp=1.0, top_p=0.95, top_k=64</td></tr>
<tr><td style="padding: 6px 10px; color: #334155; font-weight: 600;">Coding</td><td style="padding: 6px 10px; color: #334155; font-family: monospace; font-size: 12px;">temp=0.6, top_p=0.95, top_k=64</td></tr>
</tbody>
</table>
<p style="margin: 0; font-size: 12px; color: #64748b; line-height: 1.5;">Use <code style="background: #f1f5f9; padding: 2px 6px; border-radius: 4px; font-size: 11px;">--jinja</code> flag with llama.cpp. Disable thinking: <code style="background: #f1f5f9; padding: 2px 6px; border-radius: 4px; font-size: 11px;">--chat-template-kwargs '{"enable_thinking":false}'</code>.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #4b5563 0%, #6b7280 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📝</span> Usage</div>
<div style="padding: 16px;">
<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7;">Compatible with llama.cpp, LM Studio, Jan, koboldcpp, and other GGUF runtimes. Fits easily on 8GB VRAM. Encoder-free — no separate vision/audio projector needed.</p>
<div style="background: #1e293b; border-radius: 8px; padding: 16px;">
<pre style="margin: 0; font-family: 'SF Mono', 'Fira Code', monospace; font-size: 12px; color: #e2e8f0; line-height: 1.6; white-space: pre-wrap;">llama-server \
-m gemma-4-12B-it-heretic-QAT-UD-Q4_K_XL.gguf \
--jinja -ngl 99 -c 8192 \
--port 8001</pre>
</div>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #a855f7 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔗</span> Credits</div>
<div style="padding: 16px;">
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">
<b>Heretic Abliteration:</b> <a href="https://huggingface.co/coder3101/gemma-4-12B-it-qat-q4_0-unquantized-heretic" style="color: #c2410c; text-decoration: none;">coder3101</a> · <a href="https://github.com/p-e-w/heretic" style="color: #c2410c; text-decoration: none;">Heretic v1.2.0</a> ARA + Row-Norm<br>
<b>QAT Weights:</b> <a href="https://huggingface.co/google/gemma-4-12B-it" style="color: #c2410c; text-decoration: none;">Google Gemma 4 12B IT</a><br>
<b>Quantization Recipe:</b> <a href="https://huggingface.co/unsloth/gemma-4-12b-it-qat-GGUF" style="color: #c2410c; text-decoration: none;">Unsloth UD-Q4_K_XL</a> (Q4_0 layout)<br>
<b>Quantization Tool:</b> llama.cpp b9553 · <a href="https://github.com/ggerganov/llama.cpp" style="color: #c2410c; text-decoration: none;">GitHub</a><br>
<b>Original Model:</b> <a href="https://huggingface.co/google/gemma-4-12B-it" style="color: #c2410c; text-decoration: none;">Google Gemma 4 12B IT</a>
</p>
</div>
</div>
</div>
Run SC117/gemma-4-12B-it-heretic-QAT-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models