GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

pearsonkyle/gemma-4-31b-it-imatrix-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid 99f6e4; border radius: 16px; box shadow: 0 10px 15…

ggufllama.cpptext-generationtext-generation-inferencetransformersquantizationquantizedimatrixlow-bit2-bitiq2_mgemmagemma-431bcodertool-usefunction-callingagenticswe-benchlong-contextenbase_model:google/gemma-4-31B-itbase_model:quantized:google/gemma-4-31B-itlicense:gemma

Runs locally from ~10.17 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
12,447
Likes
0
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-31B-it-IQ2_M.ggufGGUFIQ2_M10.17 GBDownload
gemma-4-31B-it-IQ3_M.ggufGGUFIQ3_M13.43 GBDownload
gemma-4-31B-it-IQ4_XS.ggufGGUFIQ4_XS15.59 GBDownload

Model Details

Model IDpearsonkyle/gemma-4-31b-it-imatrix-GGUF
Authorpearsonkyle
Pipelinetext-generation
Licensegemma
Base modelgoogle/gemma-4-31B-it
Last modified2026-06-25T03:35:24.000Z

Model README

---

library_name: gguf

base_model:

  • google/gemma-4-31B-it

tags:

  • gguf
  • llama.cpp
  • text-generation
  • text-generation-inference
  • transformers
  • quantization
  • quantized
  • imatrix
  • low-bit
  • 2-bit
  • iq2_m
  • gemma
  • gemma-4
  • 31b
  • coder
  • tool-use
  • function-calling
  • agentic
  • swe-bench
  • long-context

license: gemma

language:

  • en

pipeline_tag: text-generation

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">

<div style="background: linear-gradient(135deg, #0d9488 0%, #134e4a 100%); padding: 24px; color: white;">

<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">

<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 Google/Gemma-4-31B-it · imatrix · GGUF</h1>

<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix (hybrid)</span>

</div>

</div>

<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">

<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 10.17 GiB</span>

<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;">IQ2_M · 2.845 BPW</span>

<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ llama.cpp f3e1828</span>

<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 SWE-rebench 47% pass · 100% patch</span>

</div>

<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">

<div style="background: #f0fdfa; border-left: 5px solid #0d9488; padding: 16px; border-radius: 0 8px 8px 0;">

<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #115e59; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">An aggressively compressed (under 3 bpw) <b>IQ2_M</b> quantization of <a href="https://huggingface.co/google/gemma-4-31B-it"><b>google/gemma-4-31B-it</b></a>, calibrated with a <b>hybrid imatrix</b> built from real coding/tool-use logs. Runs in vanilla <code>llama.cpp</code> / Ollama / LM Studio — <b>no custom runtime, no extra inference cost</b>. Higher-bit <b>IQ3_M</b> (3.76 bpw) and <b>IQ4_XS</b> (4.36 bpw) builds are also included for users with more VRAM.</p>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~5.6× smaller</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">10.17 GiB on disk vs 57.2 GiB FP16, at ~2.85 bits/weight.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🤖 Actually agentic</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">47% pass / 100% patch on a 10-instance agentic SWE-rebench holdout (IQ4_XS). IQ2_M still resolved 40% — best of every sub-3-bpw arm tested.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🛠️ Standard GGUF</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Loads anywhere llama.cpp runs. No patches, kernels, or forks.</span></div>

</div>

</div>

</div>

📊 Unified benchmark & quality table

Agentic metrics from a SWE-rebench holdout run through the OpenAI Agents SDK (10 instances × 3 reps). Static metrics (PPL / KLD / top-p) measured against FP16 on a held-out eval corpus at ctx=4096. KLD column is median for robustness to per-token tails.

| Quant | BPW | Size (GiB) | File | 🤖 Pass Rate | 🤖 Patch Rate | 🤖 Tool Errors | 🤖 Mean Tokens | 📐 PPL | 📐 KLD (med) | 📐 same_top_p |

|:---|---:|---:|:---|---:|---:|---:|---:|---:|---:|---:|

| FP16 (ref) | 16.0 | 57.20 | — | — | — | — | — | 215.5 | 0.000 | 100.0% |

| IQ4_XS ⭐⭐⭐ | 4.36 | 15.59 | IQ4_XS.gguf | 47±5% | 100% | 10±3% | 575K±70K | 520±32 | 319.4 | 0.073 | 78.8% |

| IQ3_M ⭐⭐ | 3.76 | 13.43 | IQ3_M.gguf | 33±12% | 100% | 16±2% | 483K±75K | 980±43 | 734.1 | 0.435 | 63.1% |

| IQ2_M ⭐ (this repo) | 2.85 | 10.17 | IQ2_M.gguf | 40±8% | 100% | 16±1% | 558K±94K | 1007±61 | 1958.7 | 1.571 | 46.6% |

<details>

<summary><b>📌 Sampling & methodology details</b></summary>

> Sampling: temperature=0.25, top_p=0.95, top_k=20, max_tokens=32768, ctx=131072, thinking=false. Run on Apple Silicon (Metal); SWE-rebench linux/amd64 images under emulation, so wall-clock is relative, not absolute.

>

> Pass Rate = gold tests pass after agent's patch (real resolution). Patch Rate = non-empty diff produced.

</details>

> 💡 Takeaway: If you have the VRAM, IQ4_XS is the clear winner — best pass rate, lowest KLD, and 2× faster wall-clock. IQ2_M is the play for ≤12 GB GPUs: it still resolved 40% of issues at under 3 bpw, beating IQ3_M despite half the bits.

---

🔬 How it was made

  • Hybrid imatrix — activation energy E[a²] mixed with weight-column energy ‖W[:,i]‖²·E[a²] per tensor, collected over real coding/tool-use logs + wiki.test.raw via quant-tuner.
  • IQ2_M codebook — 2-bit E8-lattice non-uniform codes with per-tensor tier bumps (attention output, early ffn_down get more bits). llama-quantize decides the mix.
  • Disjoint splits — calibration (imatrix), validation (per-tensor α gate), and eval (PPL/KLD) come from different corpora; the SWE-rebench holdout never appears in any calibration set.
  • Toolchain: quant-tuner for imatrix calibration, llama.cpp @ f3e1828 for final quantization. Calibration logs mined with LogMiner.

---

🚀 Usage

Ollama

ollama run hf.co/pearsonkyle/gemma-4-31b-it-imatrix-GGUF:IQ2_M

llama.cpp (GPU)

# Build with CUDA (-DGGML_CUDA=OFF for CPU/Metal)
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server
cp llama.cpp/build/bin/llama-* llama.cpp/

# Run the server
./llama-server \
    --model gemma-4-31B-it-IQ2_M.gguf \
    --ctx-size 16384 --n-gpu-layers 999 --split-mode layer \
    --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
    --parallel 1 --batch-size 2048 --ubatch-size 512 \
    --host 0.0.0.0 --port 1234

OpenAI-compatible API (Python)

import json, urllib.request

def ask(content, max_tokens=256):
    body = {
        "messages": [{"role": "user", "content": content}],
        "max_tokens": max_tokens,
        # Gemma 4 is a thinking model — disable or raise max_tokens
        "chat_template_kwargs": {"enable_thinking": False},
    }
    req = urllib.request.Request(
        "http://127.0.0.1:1234/v1/chat/completions",
        json.dumps(body).encode(),
        {"Content-Type": "application/json"},
    )
    return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"])

print(ask("What is 1+1?"))

---

🪪 License & attribution

Run pearsonkyle/gemma-4-31b-it-imatrix-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models