pearsonkyle/gemma-4-31b-it-imatrix-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid 99f6e4; border radius: 16px; box shadow: 0 10px 15…
Runs locally from ~10.17 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | pearsonkyle/gemma-4-31b-it-imatrix-GGUF |
|---|---|
| Author | pearsonkyle |
| Pipeline | text-generation |
| License | gemma |
| Base model | google/gemma-4-31B-it |
| Last modified | 2026-06-25T03:35:24.000Z |
Model README
---
library_name: gguf
base_model:
- google/gemma-4-31B-it
tags:
- gguf
- llama.cpp
- text-generation
- text-generation-inference
- transformers
- quantization
- quantized
- imatrix
- low-bit
- 2-bit
- iq2_m
- gemma
- gemma-4
- 31b
- coder
- tool-use
- function-calling
- agentic
- swe-bench
- long-context
license: gemma
language:
- en
pipeline_tag: text-generation
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">
<div style="background: linear-gradient(135deg, #0d9488 0%, #134e4a 100%); padding: 24px; color: white;">
<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">
<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 Google/Gemma-4-31B-it · imatrix · GGUF</h1>
<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix (hybrid)</span>
</div>
</div>
<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">
<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 10.17 GiB</span>
<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;">IQ2_M · 2.845 BPW</span>
<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ llama.cpp f3e1828</span>
<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 SWE-rebench 47% pass · 100% patch</span>
</div>
<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">
<div style="background: #f0fdfa; border-left: 5px solid #0d9488; padding: 16px; border-radius: 0 8px 8px 0;">
<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #115e59; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">An aggressively compressed (under 3 bpw) <b>IQ2_M</b> quantization of <a href="https://huggingface.co/google/gemma-4-31B-it"><b>google/gemma-4-31B-it</b></a>, calibrated with a <b>hybrid imatrix</b> built from real coding/tool-use logs. Runs in vanilla <code>llama.cpp</code> / Ollama / LM Studio — <b>no custom runtime, no extra inference cost</b>. Higher-bit <b>IQ3_M</b> (3.76 bpw) and <b>IQ4_XS</b> (4.36 bpw) builds are also included for users with more VRAM.</p>
</div>
<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~5.6× smaller</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">10.17 GiB on disk vs 57.2 GiB FP16, at ~2.85 bits/weight.</span></div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🤖 Actually agentic</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">47% pass / 100% patch on a 10-instance agentic SWE-rebench holdout (IQ4_XS). IQ2_M still resolved 40% — best of every sub-3-bpw arm tested.</span></div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🛠️ Standard GGUF</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Loads anywhere llama.cpp runs. No patches, kernels, or forks.</span></div>
</div>
</div>
</div>
📊 Unified benchmark & quality table
Agentic metrics from a SWE-rebench holdout run through the OpenAI Agents SDK (10 instances × 3 reps). Static metrics (PPL / KLD / top-p) measured against FP16 on a held-out eval corpus at ctx=4096. KLD column is median for robustness to per-token tails.
| Quant | BPW | Size (GiB) | File | 🤖 Pass Rate | 🤖 Patch Rate | 🤖 Tool Errors | 🤖 Mean Tokens | 📐 PPL | 📐 KLD (med) | 📐 same_top_p |
|:---|---:|---:|:---|---:|---:|---:|---:|---:|---:|---:|
| FP16 (ref) | 16.0 | 57.20 | — | — | — | — | — | 215.5 | 0.000 | 100.0% |
| IQ4_XS ⭐⭐⭐ | 4.36 | 15.59 | IQ4_XS.gguf | 47±5% | 100% | 10±3% | 575K±70K | 520±32 | 319.4 | 0.073 | 78.8% |
| IQ3_M ⭐⭐ | 3.76 | 13.43 | IQ3_M.gguf | 33±12% | 100% | 16±2% | 483K±75K | 980±43 | 734.1 | 0.435 | 63.1% |
| IQ2_M ⭐ (this repo) | 2.85 | 10.17 | IQ2_M.gguf | 40±8% | 100% | 16±1% | 558K±94K | 1007±61 | 1958.7 | 1.571 | 46.6% |
<details>
<summary><b>📌 Sampling & methodology details</b></summary>
> Sampling: temperature=0.25, top_p=0.95, top_k=20, max_tokens=32768, ctx=131072, thinking=false. Run on Apple Silicon (Metal); SWE-rebench linux/amd64 images under emulation, so wall-clock is relative, not absolute.
>
> Pass Rate = gold tests pass after agent's patch (real resolution). Patch Rate = non-empty diff produced.
</details>
> 💡 Takeaway: If you have the VRAM, IQ4_XS is the clear winner — best pass rate, lowest KLD, and 2× faster wall-clock. IQ2_M is the play for ≤12 GB GPUs: it still resolved 40% of issues at under 3 bpw, beating IQ3_M despite half the bits.
---
🔬 How it was made
- Hybrid imatrix — activation energy
E[a²]mixed with weight-column energy‖W[:,i]‖²·E[a²]per tensor, collected over real coding/tool-use logs +wiki.test.rawvia quant-tuner. - IQ2_M codebook — 2-bit E8-lattice non-uniform codes with per-tensor tier bumps (attention output, early
ffn_downget more bits).llama-quantizedecides the mix. - Disjoint splits — calibration (imatrix), validation (per-tensor α gate), and eval (PPL/KLD) come from different corpora; the SWE-rebench holdout never appears in any calibration set.
- Toolchain: quant-tuner for imatrix calibration, llama.cpp
@ f3e1828for final quantization. Calibration logs mined with LogMiner.
---
🚀 Usage
Ollama
ollama run hf.co/pearsonkyle/gemma-4-31b-it-imatrix-GGUF:IQ2_M
llama.cpp (GPU)
# Build with CUDA (-DGGML_CUDA=OFF for CPU/Metal)
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server
cp llama.cpp/build/bin/llama-* llama.cpp/
# Run the server
./llama-server \
--model gemma-4-31B-it-IQ2_M.gguf \
--ctx-size 16384 --n-gpu-layers 999 --split-mode layer \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--parallel 1 --batch-size 2048 --ubatch-size 512 \
--host 0.0.0.0 --port 1234
OpenAI-compatible API (Python)
import json, urllib.request
def ask(content, max_tokens=256):
body = {
"messages": [{"role": "user", "content": content}],
"max_tokens": max_tokens,
# Gemma 4 is a thinking model — disable or raise max_tokens
"chat_template_kwargs": {"enable_thinking": False},
}
req = urllib.request.Request(
"http://127.0.0.1:1234/v1/chat/completions",
json.dumps(body).encode(),
{"Content-Type": "application/json"},
)
return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"])
print(ask("What is 1+1?"))
---
🪪 License & attribution
- Inherits the Gemma Terms of Use from the base model.
- Base weights:
google/gemma-4-31B-it. - Calibration + quantization: Quant-Tuner with vendored llama.cpp
@ f3e1828. - Calibration logs mined with LogMiner.
Run pearsonkyle/gemma-4-31b-it-imatrix-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models