GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

soyaakinohara/gemma-4-31b-it-heretic-3.68bpw-14gb.gguf overview

Gemma 4 31B IT Heretic 3.68bpw 14GB GGUF A mixed precision GGUF quantization of DavidAU/gemma 4 31B it The DECKARD HERETIC UNCENSORED Thinking https://huggingf…

ggufllama.cppgemma4gemma-431bmixed-quantizationimatrixabliteratedhereticuncensorediq2_siq3_sq4_kq5_kq6_ktext-generationjaenmultilingualbase_model:DavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinkingbase_model:quantized:DavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinkinglicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~13.18 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
929
Likes
2
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-31b-it-heretic-3.68bpw-14gb.ggufGGUFGGUF13.18 GBDownload

Model Details

Model IDsoyaakinohara/gemma-4-31b-it-heretic-3.68bpw-14gb.gguf
Authorsoyaakinohara
Pipelinetext-generation
Licenseapache-2.0
Base modelDavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking
Last modified2026-09-18T00:51:04.000Z

Model README

---

license: apache-2.0

base_model:

  • DavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking

language:

  • ja
  • en
  • multilingual

library_name: gguf

pipeline_tag: text-generation

tags:

  • gguf
  • llama.cpp
  • gemma4
  • gemma-4
  • 31b
  • mixed-quantization
  • imatrix
  • abliterated
  • heretic
  • uncensored
  • iq2_s
  • iq3_s
  • q4_k
  • q5_k
  • q6_k
  • text-generation

---

Gemma 4 31B IT Heretic 3.68bpw 14GB GGUF

A mixed-precision GGUF quantization of

DavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking.

This is an independent quantization of the DavidAU derivative. It is not an

official Google release and is not a reproduction of a publisher's private

quantization artifact.

The quantization keeps selected attention and output paths at higher precision

while applying stronger compression to most FFN weights. The goal is to retain

useful quality while reducing the memory and storage requirements for local

llama.cpp inference.

> The model is an uncensored / refusal-removed derivative. Read the responsible

> use section before deploying it for other users.

Files

gemma-4-31b-it-heretic-3.68bpw-14gb.gguf

| Property | Value |

|---|---|

| Architecture | Gemma 4 31B (gemma4) |

| Parameters | 31B class |

| Format | GGUF |

| Nominal quantization | 3.68 bpw |

| File size | 14,152,729,312 bytes (~13.18 GiB) |

| GGUF tensor count | 833 |

| Context metadata | 262,144 tokens |

| MTP | Not included |

| Vision / audio | Not included in this text GGUF |

| License | Apache-2.0, inherited from the source model |

Mixed-precision layout

The final GGUF contains the following tensor types:

| GGML type | Tensor count | Main use |

|---|---:|---|

| F32 | 422 | norms, scalar tensors, and other sensitive small tensors |

| Q6_K | 31 | output, token embedding, and protected attention tensors |

| Q5_K | 50 | selected sliding-attention tensors |

| Q4_K | 154 | selected attention and edge-layer FFN tensors |

| IQ3_S | 64 | selected FFN tensors kept above the most aggressive level |

| IQ2_S | 112 | heavily compressed FFN tensors |

The map is architecture-aware rather than a flat quantization preset. Gemma 4

has 60 text layers, including full-attention layers at 5, 11, 17, 23, 29, 35,

41, 47, 53, and 59. These layers received more conservative treatment than

most sliding-attention layers.

Calibration and provenance

The importance matrix was generated from the exact BF16 GGUF derived from the

DavidAU checkpoint. The local calibration corpus contains:

  • English WikiText material
  • Japanese Wikipedia material
  • Japanese technical and reasoning prose
  • llama.cpp Gemma 4 implementation and conversion code
  • the Gemma 4 chat template and model-card material
context length:     512
batch size:         512
calibration chunks: 40
process output:     enabled
importance entries: 410

The source checkpoint was pinned to:

70faf534ff040eb4762b7a51d82192de8a83293b

The source files were BF16 safetensors. The weights were converted directly to

BF16 GGUF before quantization; this file was not produced by requantizing an

already low-bit GGUF.

The conversion and quantization used llama.cpp commit:

030ebb558a5820b444a8f836ed5cdd46c9b4bd7a

A small compatibility addition was made to the local copy of config.json so

the converter could read global_head_dim and

num_global_key_value_heads from the Gemma 4 configuration. No model weights

were changed by this compatibility step.

llama.cpp usage

The following command was used for the local server test:

./llama-server \
  -m '/mnt/hitachi1tb/gemma4-deckard/models/gemma-4-31b-it-heretic-3.68bpw-14gb.gguf' \
  --split-mode layer \
  --tensor-split 2,1 \
  --host 0.0.0.0 \
  --port 8080 \
  -ctk q4_0 \
  -ctv q4_0 \
  -c 100000 \
  -ngl 999

For a local-only server, prefer:

--host 127.0.0.1

0.0.0.0 exposes the server to the network. Add authentication, firewall

rules, and access controls before using it outside a trusted local network.

The model metadata advertises a 262,144-token context, while the example

command intentionally uses a 100,000-token context. Actual usable context is

also limited by available RAM, KV-cache settings, and frontend behavior.

Reported local performance

Hardware and runtime:

OS:       Ubuntu 24.04
CPU:      Intel Core i7-10700K
GPU 0:    NVIDIA GeForce RTX 5060 Ti 16GB
GPU 1:    NVIDIA GeForce RTX 3070 8GB
RAM:      32GB
Runtime:  llama.cpp CUDA build
KV cache: q4_0 / q4_0
Context:  100,000

The repository owner reports approximately 19 tokens/second with the server

command above. This is a local measurement, not a guaranteed benchmark. Actual

speed varies with prompt length, generated-token length, context occupancy, KV

cache usage, sampling settings, llama.cpp version, background load, and GPU

memory state.

Quality note

The model was spot-checked locally with Japanese instruct generation after

quantization. The model retained general Japanese knowledge and conversational

behavior in the tested prompts. This is not a formal benchmark, and no claim is

made that the quantization is bit-identical to any other Gemma 4 GGUF.

Source links

AI assistance disclosure

The local model download, conversion, calibration-data preparation, imatrix

generation, mixed quantization, validation, and this model card were performed

with assistance from GPT-5.6-Luna via Hermes Agent. The repository owner

reviewed the result and requested the publication.

Responsible use

This model is an uncensored / refusal-removed derivative and may produce

unsafe, illegal, or harmful content that aligned models would refuse. It should

not be treated as having a reliable built-in safety layer.

For public or multi-user deployment, add appropriate access control, input and

output moderation, audit logging, rate limiting, and human review. Users are

responsible for prompts, outputs, and downstream actions based on the outputs.

The model is provided as-is under the applicable Apache-2.0 license and source

model terms.

---

日本語

概要

これは、DavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinkingを元に作成した、Gemma 4 31Bの混合精度GGUF量子化モデルです。

重要度の高いAttentionや出力経路には比較的高い精度を残し、パラメータの大部分を占めるFFNには強い圧縮を適用しています。単純に全テンソルを同じ量子化方式へ落とすのではなく、Gemma 4の層構成を考慮したマップとimatrixを使用しています。

Google公式モデルやDavidAU公式GGUFではなく、独立して作成した派生量子化モデルです。

> 無検閲・拒否除去済みの派生モデルです。公開運用前に「利用上の注意」を確認してください。

基本仕様

ファイル名: gemma-4-31b-it-heretic-3.68bpw-14gb.gguf
形式:       GGUF
サイズ:     14,152,729,312 bytes(約13.18 GiB)
量子化:     3.68 bpw
テンソル数: 833
コンテキスト: 262,144 tokens(メタデータ)
MTP:        なし
Vision/Audio: なし(テキストGGUF)

混合量子化の内訳

F32    422 tensors  norm / scalarなど
Q6_K    31 tensors  output、embedding、保護したAttention
Q5_K    50 tensors  主にSliding Attentionの一部
Q4_K   154 tensors  Attentionの一部、端の層のFFN
IQ3_S   64 tensors  精度を少し残したFFN
IQ2_S  112 tensors  強く圧縮したFFN

Gemma 4の60層のうち、Full Attention層は以下です。

5, 11, 17, 23, 29, 35, 41, 47, 53, 59

これらの層は、主にSliding Attentionを使う層より保守的な精度配分にしています。

キャリブレーション

対象のDavidAU派生モデルをBF16 GGUFへ変換した後、そのモデル自身からimatrixを作成しました。

context:     512
batch:       512
chunks:      40
process_output: 有効
importance entries: 410

キャリブレーションデータには以下を含めています。

  • 英語WikiText
  • 日本語Wikipedia
  • 日本語の技術文章・推論文章
  • llama.cppのGemma 4実装・変換コード
  • Gemma 4のチャットテンプレート
  • 元モデルカードの内容

元モデルのrevisionは以下に固定しています。

70faf534ff040eb4762b7a51d82192de8a83293b

BF16 safetensorsから直接BF16 GGUFを作成してから量子化しており、低ビットGGUFを再量子化したものではありません。

使用したllama.cppのcommit:

030ebb558a5820b444a8f836ed5cdd46c9b4bd7a

Gemma 4 converterが設定を読み込めるよう、ローカルのconfig.jsonglobal_head_dimnum_global_key_value_headsを補完しました。この互換性対応によってモデルの重みは変更していません。

llama.cppでの起動

ローカルで使用したサーバー起動コマンドです。

./llama-server \
  -m '/mnt/hitachi1tb/gemma4-deckard/models/gemma-4-31b-it-heretic-3.68bpw-14gb.gguf' \
  --split-mode layer \
  --tensor-split 2,1 \
  --host 0.0.0.0 \
  --port 8080 \
  -ctk q4_0 \
  -ctv q4_0 \
  -c 100000 \
  -ngl 999

ローカルマシンからのみ接続する場合は、次を推奨します。

--host 127.0.0.1

0.0.0.0はネットワーク上へサーバーを公開します。LAN外や複数ユーザー向けに公開する場合は、認証、Firewall、アクセス制御を追加してください。

モデルメタデータ上のコンテキスト長は262,144 tokensですが、上記コマンドでは意図的に100,000 tokensを指定しています。実際の利用可能長は、RAM、KV cache、フロントエンドの仕様にも左右されます。

実測速度

OS:       Ubuntu 24.04
CPU:      Intel Core i7-10700K
GPU 0:    NVIDIA GeForce RTX 5060 Ti 16GB
GPU 1:    NVIDIA GeForce RTX 3070 8GB
RAM:      32GB
KV cache: q4_0 / q4_0
Context:  100,000
Runtime:  llama.cpp CUDA build

上記の起動設定で、リポジトリ所有者の環境では約19 tokens/secondを記録しました。これはローカル環境での報告値であり、保証されたベンチマークではありません。プロンプト長、生成長、コンテキスト使用量、KV cache、サンプリング設定、llama.cppのバージョン、バックグラウンド負荷などで変動します。

品質について

量子化後に日本語のinstruct生成を実機で確認し、今回のテスト範囲では日本語の一般知識と会話挙動を維持していました。これは簡易確認であり、正式なベンチマークではありません。また、他のGemma 4 GGUFとビット単位で同一であるという意味ではありません。

関連リンク

AI利用の開示

モデルの取得、変換、キャリブレーションデータ作成、imatrix作成、混合量子化、動作確認、およびこのモデルカードの作成は、Hermes Agent経由のGPT-5.6-Lunaの支援を受けて行われました。最終的な確認と公開はリポジトリ所有者が行っています。

利用上の注意

このモデルは無検閲・拒否除去済みの派生モデルです。通常のアライン済みモデルが拒否する危険・違法・有害な内容を出力する可能性があります。信頼できる安全機構を内蔵しているとは考えないでください。

公開運用や複数ユーザー向けの運用では、アクセス制御、入力・出力フィルタ、監査ログ、レート制限、人間による確認などを用途に応じて実装してください。プロンプト、出力、および出力を利用したdownstreamの行為については利用者が責任を負います。

モデルは現状のまま提供されます。適用されるApache-2.0ライセンスおよび元モデルの条件を確認して利用してください。

Run soyaakinohara/gemma-4-31b-it-heretic-3.68bpw-14gb.gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models