GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

pearsonkyle/gemma-4-31B-it-awq-2bit-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid 99f6e4; border radius: 16px; box shadow: 0 10px 15…

ggufllama.cpptext-generationtext-generation-inferencetransformersquantizationquantizedimatrixlow-bit2-bitiq2_mgemmagemma-431bcodertool-usefunction-callingagenticswe-benchlong-contextenbase_model:google/gemma-4-31B-itbase_model:quantized:google/gemma-4-31B-itlicense:gemma

Runs locally from ~10.17 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
12,187
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-31B-it-IQ2_M.ggufGGUFIQ2_M10.17 GBDownload

Model Details

Model IDpearsonkyle/gemma-4-31B-it-awq-2bit-GGUF
Authorpearsonkyle
Pipelinetext-generation
Licensegemma
Base modelgoogle/gemma-4-31B-it
Last modified2026-06-21T18:59:03.000Z

Model README

---

library_name: gguf

base_model:

  • google/gemma-4-31B-it

tags:

  • gguf
  • llama.cpp
  • text-generation
  • text-generation-inference
  • transformers
  • quantization
  • quantized
  • imatrix
  • low-bit
  • 2-bit
  • iq2_m
  • gemma
  • gemma-4
  • 31b
  • coder
  • tool-use
  • function-calling
  • agentic
  • swe-bench
  • long-context

license: gemma

language:

  • en

pipeline_tag: text-generation

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">

<div style="background: linear-gradient(135deg, #0d9488 0%, #134e4a 100%); padding: 24px; color: white;">

<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">

<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 Gemma-4-31B-it · IQ2_M · GGUF</h1>

<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix (hybrid)</span>

</div>

</div>

<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">

<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 10.17 GiB</span>

<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;">IQ2_M · 2.845 BPW</span>

<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ llama.cpp f3e1828</span>

<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 SWE-rebench 30% pass · 80% patch</span>

</div>

<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">

<div style="background: #f0fdfa; border-left: 5px solid #0d9488; padding: 16px; border-radius: 0 8px 8px 0;">

<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #115e59; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">A single, aggressively compressed (under 3 bits per weight) <b>IQ2_M</b> quantization of <b>google/gemma-4-31B-it</b>, calibrated with a <b>hybrid imatrix</b> built from real coding/tool-use logs. This is the variant that <b>actually solved real GitHub issues</b> in an agentic SWE-rebench harness — it was the only sub-3-bpw arm we tested that resolved any instances at all (see §1). It is a plain GGUF with <b>no custom runtime and no extra inference cost</b>.</p>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~5.6x smaller</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">10.17 GiB on disk vs 57.2 GiB for FP16, at ~2.85 bits per weight.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🤖 Actually agentic</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">30% pass rate / 80% patch rate on a 10-instance agentic SWE-rebench holdout — the best of every 2-bit arm tested.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🛠️ Standard GGUF</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Loads in vanilla llama.cpp / llama-server / LM Studio / Ollama with no custom runtime, kernels, or patches.</span></div>

</div>

</div>

</div>

🤖 1. Agentic benchmark (why this one ships)

Every other quant was dropped from this repo because it lost where it matters: can the model actually drive an agent and resolve a real GitHub issue? We ran each candidate through an agentic SWE-rebench harness — mini-swe-agent driving a local llama-server, one clean Docker container per instance, graded by running the gold FAIL_TO_PASS / PASS_TO_PASS tests — over a 10-instance held-out slice.

IQ2_M (imatrix) was the only sub-3-bpw arm that resolved any instances, and it also produced the most valid patches:

| Model | Pass Rate | Patch Rate | Avg Tools | Avg Tokens |

|---|---|---|---|---|

| IQ2_M (imatrix) ⭐ ships here | 3/10 (30%) | 8/10 (80%) | 13.7 | 127K |

| IQ2_M (AWQ cv-gate) | 0/10 (0%) | 5/10 (50%) | 31.5 | 450K |

| QAT Q2_K_S (imatrix) | 0/10 (0%) | 7/10 (70%) | 22.9 | 367K |

| Q2_K_S (AWQ cv-gate) | 0/10 (0%) | 4/10 (40%) | 8.1 | 68K |

| QAT Q2_K_S (AWQ cv-gate) | 0/10 (0%) | 4/10 (40%) | 8.9 | 50K |

| Q2_K_S (imatrix) | 0/10 (0%) | 3/10 (30%) | 2.7 | 31K |

| IQ2_XS (AWQ cv-gate) | 0/10 (0%) | 3/10 (30%) | 0.9 | 12K |

| IQ2_XS (imatrix) | 0/10 (0%) | 3/10 (30%) | 0.7 | 11K |

| QAT IQ2_XS (imatrix) | 0/10 (0%) | 3/10 (30%) | 0.5 | 10K |

| QAT IQ2_XS (AWQ cv-gate) | 0/10 (0%) | 3/10 (30%) | 0.3 | 10K |

| Q2_K (plain, no calibration) | — | — | — | — (crashed) |

Pass Rate = instances whose gold tests pass after the agent's patch (real resolution). Patch Rate = instances where the agent produced a non-empty diff. Avg Tools = mean bash tool-calls per instance. Avg Tokens = mean total tokens per instance.

The takeaways that decided the release:

  • AWQ helped PPL/KLD but hurt the agent. The AWQ cv-gate IQ2_M looks competitive on static metrics (see §2) yet resolved zero instances and burned ~3.5× the tokens looping. Static top-token agreement does not predict agentic competence at this bit budget.
  • More tools/tokens ≠ better. The AWQ and QAT-Q2_K_S arms made far more tool calls and produced patches, but never converged on a correct one — they thrash. IQ2_M-imatrix uses ~14 tools and the fewest tokens among the patch-producing arms while landing the only passes.
  • Calibration is load-bearing. Plain Q2_K (no imatrix) crashed the harness outright; the uncalibrated low-bit arms barely call tools at all.

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #fde68a; border-radius: 12px; background: #fffbeb; padding: 16px; margin: 16px 0; color: #92400e; font-size: 13px; line-height: 1.7;">

<b>⚠️ Caveat.</b> This is a sub-3-bpw quant of a 31B reasoning model. It is meaningfully better than the alternatives at the same size, but it is <b>not</b> a substitute for FP16 / Q4_K_M / Q5_K_M when you have the VRAM. A 30% agentic pass rate is strong <i>for 2 bits</i> — use it when memory is the binding constraint.

</div>

---

📊 2. Static metrics (PPL / KLD / top-p)

Benched on corpus.eval.txt (~90k tokens of external code+math+tools text, not drawn from the calibration data), same llama.cpp build, ctx=4096. KLD and top_p are measured against google/gemma-4-31B-it FP16. KLD is the median per-token value for robustness to tails.

| Metric | FP16 (ref) | IQ2_M (imatrix) |

|---|---|---:|

| File | — | gemma-4-31B-it-IQ2_M.gguf |

| Quant | FP16 | IQ2_M |

| Technique | none (reference) | hybrid imatrix |

| Size (GiB) | 57.20 | 10.17 |

| BPW | 16.005 | 2.845 |

| PPL | 215.50 | 1958.74 |

| KLD (median) | 0.00000 | 1.571 |

| same_top_p | 100.00% | 46.57% |

Note that PPL and KLD only loosely track agentic competence at this bit budget — the AWQ variant scored better KLD (1.602) and top_p (48.1%) than this file, yet resolved zero issues in §1. We ship the model that does the job, not the one that wins the static leaderboard.

Corpora used to build and evaluate this model live under calibration_data/:

| File | Role |

|---|---|

| corpus.cal.txt | imatrix collection (wiki.test.raw + logtrain TRAIN slice, windowed packer) |

| corpus.val.txt | held-out validation slice (logtrain TEST split, ~10k tokens) |

| corpus.eval.txt | PPL / KLD eval (external code+math+tools, ~90k tokens) |

---

🔬 3. How it was made

This file is an IQ2_M quantization guided by a hybrid imatrix — an importance matrix that tells llama-quantize which channels deserve more of the tiny 2-bit budget.

  • Importance matrix (imatrix). Collected with llama-imatrix over the calibration corpus (real coding/tool-use logs + wiki.test.raw), the imatrix records E[a²]ᵢ per input channel so the codebook spends precision on the channels that actually drive each layer's output. This release uses a hybrid variant that mixes activation energy E[a²] with weight-column energy ‖W[:, i]‖² · E[a²] per tensor.
  • IQ2_M codebook. A 2-bit non-uniform (E8-lattice) codebook with a per-tensor mix that bumps the most sensitive tensors (attention output, early ffn_down) up a tier — llama-quantize decides the mix per tensor.
  • AWQ was evaluated and rejected. We also built activation-aware (AWQ) variants of this and other quants. AWQ folds a per-channel rescale into the preceding RMSNorm to flatten outlier channels before quantizing — it improved static PPL/KLD, but the AWQ IQ2_M resolved zero agentic instances while looping for 3.5× the tokens (§1). The plain hybrid-imatrix IQ2_M won the only metric that mattered, so that is what ships.

Data slices are kept disjoint — calibration (imatrix), validation (held-out gate during the AWQ study), and eval (PPL/KLD) come from different sources, and the agentic SWE-rebench holdout is a wholly separate dataset whose issues never appear in any calibration corpus.

Toolchain: imatrix calibration orchestrated by quant-tuner; final quantization with llama-quantize from llama.cpp pinned to commit f3e1828.

---

🚀 4. Usage

Ollama

Pull and run directly from Hugging Face — no manual download needed:

ollama run hf.co/pearsonkyle/gemma-4-31B-it-awq-2bit-GGUF:IQ2_M

Building llama.cpp from source (GPU)

# Install dependencies and clone repo
apt-get update && apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp

# Build with CUDA (set -DGGML_CUDA=OFF for CPU/Metal)
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server

# Move binaries
cp llama.cpp/build/bin/llama-* llama.cpp/

Running the server

./llama-server \
    --model gemma-4-31B-it-IQ2_M.gguf \
    --ctx-size 16384 \
    --n-gpu-layers 999 \
    --split-mode layer \
    --flash-attn on \
    --cache-type-k q8_0 \
    --cache-type-v q8_0 \
    --parallel 1 \
    --batch-size 2048 \
    --ubatch-size 512 \
    --host 0.0.0.0 \
    --port 1234

Querying via the OpenAI-compatible API

import json, urllib.request

def ask(content, max_tokens=256):
    body = {
        "messages": [{"role": "user", "content": content}],
        "max_tokens": max_tokens,
        # Gemma 4 is a thinking model. Set this to False (or raise max_tokens),
        # otherwise the reply lands in reasoning_content and "content" is empty.
        "chat_template_kwargs": {"enable_thinking": False},
    }
    req = urllib.request.Request("http://127.0.0.1:1234/v1/chat/completions",
                                 json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"]

print(ask("What is 1+1?"))

---

🪪 5. License & attribution

</content>

Run pearsonkyle/gemma-4-31B-it-awq-2bit-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models