GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ManniX-ITA/gemma-4-A4B-98e-v7-coder-it-GGUF overview

gemma 4 A4B 98e v7 coder it GGUF GGUF quantizations of ManniX ITA/gemma 4 A4B 98e v7 coder it https://huggingface.co/ManniX ITA/gemma 4 A4B 98e v7 coder it , t…

ggufimatrixquantizedbase_model:ManniX-ITA/gemma-4-A4B-98e-v7-coder-itbase_model:quantized:ManniX-ITA/gemma-4-A4B-98e-v7-coder-itlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.11 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
22,341
Likes
4
Pipeline

Repository Files & Downloads

15 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-A4B-98e-v7-coder-it-CD-Q2_K.ggufGGUFQ2_K8.22 GBDownload
gemma-4-A4B-98e-v7-coder-it-F16.ggufGGUFF1637.06 GBDownload
gemma-4-A4B-98e-v7-coder-it-IQ4_NL.ggufGGUFIQ4_NL10.63 GBDownload
gemma-4-A4B-98e-v7-coder-it-IQ4_XS.ggufGGUFIQ4_XS10.25 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q3_K_L.ggufGGUFQ3_K_L10.19 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q3_K_M.ggufGGUFQ3_K_M9.79 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q4_K_L.ggufGGUFQ4_K_L12.50 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q4_K_M.ggufGGUFQ4_K_M12.33 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q4_K_S.ggufGGUFQ4_K_S11.37 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q5_K_L.ggufGGUFQ5_K_L14.20 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q5_K_M.ggufGGUFQ5_K_M14.04 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q6_K.ggufGGUFQ6_K16.58 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q6_K_L.ggufGGUFQ6_K_L16.75 GBDownload
gemma-4-A4B-98e-v7-coder-it-Q8_0.ggufGGUFQ8_019.71 GBDownload
mmproj-gemma4.ggufGGUFGGUF1.11 GBDownload

Model Details

Model IDManniX-ITA/gemma-4-A4B-98e-v7-coder-it-GGUF
AuthorManniX-ITA
Pipeline
Licenseapache-2.0
Base modelManniX-ITA/gemma-4-A4B-98e-v7-coder-it
Last modified2026-07-14T17:33:20.000Z

Model README

---

base_model: ManniX-ITA/gemma-4-A4B-98e-v7-coder-it

tags:

- gguf

- imatrix

- quantized

license: apache-2.0

---

gemma-4-A4B-98e-v7-coder-it-GGUF

GGUF quantizations of ManniX-ITA/gemma-4-A4B-98e-v7-coder-it,

the loop-fixed code prune of Gemma 4 26B-A4B (128→98 experts/layer, ~20.8B).

All quants made using imatrix with calibration data v5.

The imatrix.dat used is included in this repo for reproducibility/audit; mmproj-gemma4.gguf

is the shared Gemma 4 SigLIP vision tower (untouched by pruning) for multimodal use.

Quantizations — score, size & bits-per-weight

Each published tier was scored on HumanEval+ (164) and MultiPL-E-100 (llama.cpp,

deployment samplermin_p), as a quant-vs-quant comparison of how each tier holds the

model's code ability. The table shows, per tier, the code score, the exact size, and the **true

bits-per-weight (bpw = 8 × bytes ÷ 19,877,953,946). ⭐ marks a recommended pick**.

| Tier | Size | bpw | HE+ % | MPE-100 % |

|---|---:|---:|---:|---:|

| Q8_0 | 21.16 GB | 8.52 | 91.46 | 91.00 |

| Q6_K_L | 17.98 GB | 7.24 | 91.46 | 89.67 |

| Q6_K | 17.81 GB | 7.17 | 90.85 | 89.00 |

| Q5_K_L | 15.25 GB | 6.14 | 91.46 | 90.00 |

| Q5_K_M | 15.07 GB | 6.07 | 91.46 | 90.33 |

| Q4_K_L | 13.42 GB | 5.40 | 92.07 | 89.33 |

| Q4_K_M ⭐ | 13.24 GB | 5.33 | 92.68 | 89.00 |

| Q4_K_S | 12.21 GB | 4.91 | 92.68 | 89.33 |

| IQ4_NL | 11.42 GB | 4.60 | 89.63 | 88.67 |

| IQ4_XS | 11.01 GB | 4.43 | 91.46 | 88.33 |

| Q3_K_L | 10.94 GB | 4.40 | 90.85 | 89.00 |

| Q3_K_M | 10.51 GB | 4.23 | 91.46 | 87.00 |

| CD-Q2_K ⭐ | 8.82 GB | 3.55 | 89.63 | 86.00 |

Recommended picks:

  • Q4_K_M ⭐ (13.24 GB) — recommended default — ties the best HE+ in the sweep (92.68%) at 13.2 GB.
  • Q4_K_S ⭐ (12.21 GB) — same 92.68% HE+ at 12.2 GB if you want the smaller 4-bit K-quant.
  • CD-Q2_K ⭐ (8.82 GB) — smallest tier still in the ~90% band — 89.63% HE+ / 86.0% MPE at 8.8 GB.

Reads: HE+ holds in the 89–93% band and MPE in the 86–91% band across the whole K-quant /

CD ladder — the 4-bit K-quants (Q4_K_M / Q4_K_S at 92.68%) actually top the sweep, and even the

3.55-bpw CD-Q2_K stays at ~90%. The CD-Q2_K ContribDynamic per-layer body is the

recommended low-bit path; the pruned MoE degenerates on a plain 2-bit or an IQ-family body at this

size (token-salad), so those tiers are not offered.

Head-to-head by file size — v7-coder vs Qwen2.5-Coder-14B (iso-disk)

Pairing by tier name is misleading — v7-coder is a ~20.8B-total MoE and Qwen2.5-Coder-14B is a

14.7B dense model, so the same tier name lands at a different file size. The fair comparison is

iso-disk: at a given GB budget, which model scores higher on HumanEval+? Qwen GGUFs are

bartowski's Qwen2.5-Coder-14B-Instruct-GGUF;

its ladder sits at 83–85% across the whole stack. v7-coder HE+ is the deployment-sampler

per-tier sweep above; Qwen HE+ is the same-stack reference.

| Disk band | Qwen2.5-Coder-14B (size / bpw / HE+) | v7-coder best (size / bpw / HE+) | Δ HE+ |

|---|---|---|---:|

| ~21.2 GB | (none — Qwen ceiling is Q8_0 15.70 GB) | Q8_0 21.16 / 8.52 / 91.46% | new top |

| ~17.8 GB | (none — Qwen ceiling is Q8_0 15.70 GB) | Q6_K 17.81 / 7.17 / 90.85% | new top |

| ~15.1 GB | Q8_0 15.70 / 8.54 / 84.76% | Q5_K_M 15.07 / 6.07 / 91.46% | +6.70 |

| ~13.2 GB | Q6_K 12.12 / 6.60 / 84.76% | Q4_K_M 13.24 / 5.33 / 92.68% | +7.92 |

| ~12.2 GB | Q6_K 12.12 / 6.60 / 84.76% | Q4_K_S 12.21 / 4.91 / 92.68% | +7.92 |

| ~11.0 GB | Q5_K_M 10.51 / 5.72 / 83.54% | IQ4_XS 11.01 / 4.43 / 91.46% | +7.92 |

| ~10.5 GB | Q5_K_M 10.51 / 5.72 / 83.54% | Q3_K_M 10.51 / 4.23 / 91.46% — iso-disk (same 10.5 GB) | +7.92 |

| ~8.8 GB | Q4_K_M 8.99 / 4.89 / 85.37% | CD-Q2_K 8.82 / 3.55 / 89.63% — ⭐ smallest ~90% | +4.26 |

Reads:

  1. Iso-disk ~10.5 GB. v7-coder Q3_K_M (10.51 GB / 4.23 bpw / 91.46%) vs Qwen Q5_K_M (10.51 GB / 5.72 bpw / 83.54%): +7.92pp at the exact same file size, −1.49 bpw.
  2. Sub-9 GB code-grade. CD-Q2_K (8.82 GB / 3.55 bpw / 89.63%) holds the ~90% HE+ band ~0.2 GB smaller and ~1.3 bpw lower than Qwen's best (85.37%, Q4_K_M, 8.99 GB).
  3. Every band wins at lower bpw. Across the ladder the MoE uses 1.5–4 bpw less than the

dense Qwen tier at the same disk and still scores higher on HumanEval+ — the point of the iso-disk framing.

> CD recipes are open-source — generator at

> omnimergekit/scripts/generate_cd_maps.py.

How to Use

With llama.cpp:

llama-server -m gemma-4-A4B-98e-v7-coder-it-Q4_K_M.gguf -c 32768 -ngl 99 \
    --reasoning-format deepseek --reasoning-budget 8192 \
    --temp 1.0 --top-k 64 --top-p 0.95 --min-p 0.05 \
    --repeat-penalty 1.02 --repeat-last-n 2048

The --reasoning-budget flag is required for Gemma 4 thinking — without it the

model emits malformed channel tokens. Keep -c several times larger than the

reasoning budget — with -c equal to the budget the thinking phase can fill the

whole window and the answer degenerates.

Recommended sampling--repeat-penalty 1.02 --repeat-last-n 2048 is the

field-tested sweet spot for long agentic / tool-calling sessions (validated with

opencode driving multi-turn coding work): it suppresses the intermittent

repetition loops without side effects. Stronger penalties (1.05–1.1) also stop

the loops but cause premature end-of-turn — the model announces a step and then

stops mid-task. Narrower windows (64–1024) at 1.02 let long-period loops through;

keep the full 2048. The remaining sampler values above are the Gemma 4 vendor

defaults baked into the GGUF, stated explicitly so CLI defaults don't override them.

With ollama: ollama pull mannix/gemma4-98e-v7-coder:Q4_K_M

(:latest = Q4_K_M; :vision-<tier> pairs the tier with the SigLIP vision tower).

---

Gemma 4 A4B 98-Expert v7-coder — loop-fixed code prune (~20.8B)

> Eval complete (Q6_K / llama.cpp, greedy, same host). Every cell in the

> scoreboard is read from summary.json under the cohort-pinned greedy recipe

> (temperature 0.0, top_p 1.0, top_k 0). The 128e, v6-coder and v7-coderx columns

> are the matching same-host Q6_K runs. The GGUF and NVFP4A16 formats are deployment

> targets and are not separately benchmarked (cohort policy) — the Q6_K column

> is representative.

>

> Headline — the cohort's balanced code build. v7-coder leads the cohort on the

> easier-but-broad LiveCodeBench-medium slices (LCB-55-v4 98.18%, **LCB-100-v4

> 94.0%) and on HumanEval (98.17%), ties the cohort top on MATH-500 (95.0%)**,

> and takes AIME (80.0%). On the all-hard LiveCodeBench-77 set — the most

> discriminating LCB slice — it scores 84.42% (128e 79.22%, v7-coderx 85.71%),

> just behind the code-maximal sibling. This is the loop-fixed build: it

> force-keeps the agentic loop-protection experts and **replaces the earlier looping

> fs2440 prune**. Like its sibling it spends the prune budget on graduate science —

> GPQA-diamond sits at 51.52% (no targeted_gpqa term; ≈ v7-coderx 51.01%). If you

> need the hardest-code lean (all-hard LCB + HE+), see the sibling

> v7-coderx.

A research checkpoint that prunes the unpruned

Gemma 4 26B-A4B-it

(128 experts/layer, top-8 + shared, 30 layers) down to 98 experts per layer. The

fkbroad drop map (generate_drop_map_v5) up-weights generic-code (3×) and

LiveCodeBench-medium (2×) with no science or multilingual targeting, and

force-keeps the agentic loop-protection experts (agentic_eog, 46 experts, 0/46 dropped)

so the served model does not loop. Same 98e shape, same router, same attention, same norms as

the rest of the cohort, plus the mandatory shared-FFN α=1.2 upweight all coder

variants carry. No per-layer floor clamp and no DERN fold.

Benchmarks

Q6_K · llama.cpp · greedy (temperature 0.0, top_p 1.0, top_k 0), all four

models scored on the same host from summary.json. Row-max in bold.

This repo = v7-coder.

| Benchmark | 128e (unpruned) | v6-coder | v7-coder | v7-coderx |

|---|---:|---:|---:|---:|

| GPQA-diamond (198q) | 67.17 | 61.11 | 51.52 | 51.01 |

| AIME (30q) | 73.33 | 56.67 | 80.00 | 76.67 |

| MATH500 (100q) | 92.00 | 89.00 | 95.00 | 95.00 |

| GSM8K (100q) | 89.00 | 88.00 | 91.00 | 93.00 |

| ARC-Challenge (full) | 96.50 | 95.39 | 92.15 | 86.60 |

| IFEval (100q, strict) | 97.00 | 92.00 | 92.00 | 92.00 |

| HumanEval (164) | 97.56 | 98.17 | 98.17 | 96.95 |

| HumanEval+ (164) | 92.07 | 92.68 | 92.07 | 93.29 |

| LCB-medium-55 v4 | 96.36 | 92.73 | 98.18 | 92.73 |

| LCB-medium-100 v4 | 97.00 | 94.00 | 94.00 | 91.00 |

| MultiPL-E (100) | 90.00 | 89.00 | 89.67 | 89.00 |

<sub>Metrics: GPQA & GSM8K = exact_match flexible-extract · MATH500 = math_verify ·

ARC & AIME = exact_match · IFEval = prompt_level_strict_acc · HumanEval/+ = pass@1

chat-extract · LCB-55/100 & MultiPL-E = pass@1. 128e uses the lcb_medium_55/100

templates; the prunes use lcb_medium_*_v4 (corrected harness, equivalent task).</sub>

v7-coder is the balanced code sibling: it tops the cohort on LCB-medium and

HumanEval and ties on MATH/AIME, while v7-coderx leads the all-hard LCB-77 and HE+.

Both pay the budget on graduate science (GPQA ≈ 51) and the easier ARC axis.

LiveCodeBench across problem sets

The code score depends on the LiveCodeBench slice. All cells are the **same greedy

Q6_K / imat-Q6 llama.cpp stack** (build provenance verified per run); v4-55/100 mirror the

9-bench above. The all-hard 77q set is the most demanding and the most discriminating

across the cohort.

| LCB problem set | 128e | v7-coder | v7-coderx |

|---|---:|---:|---:|

| LCB-medium-55 (v4, 55q) | 96.36% | 98.18% | 92.73% |

| LCB-medium-100 (v4, 100q) | 97.00% | 94.00% | 91.00% |

| LCB-hard-77 (all-hard, 77q) | 79.22% | 84.42% | 85.71% |

Coder-field comparison — v7-coder vs Qwen2.5-Coder-14B / 7B + Qwen3.5-9B (Q6_K, llama.cpp, greedy)

The 9 canonical benches + MultiPL-E-100, all on the identical llama.cpp Q6_K / greedy

recipe (reasoning models served with `--reasoning-format deepseek --reasoning-budget 12288

--parallel 2`). Architectures differ — this is a same-harness comparison, not a same-class one:

  • v7-coder — Gemma-4 26B-A4B MoE pruned to 98 experts (~20.8B total, ~A4B active), reasoning.
  • Qwen2.5-Coder-14B / 7B-Instruct — dense, non-reasoning code specialists (bartowski Q6_K).
  • Qwen3.5-9B — dense reasoning model (bartowski Q6_K).

| Bench (n) | v7-coder Q6_K | Qwen2.5-Coder-14B | Qwen2.5-Coder-7B | Qwen3.5-9B |

|---|---:|---:|---:|---:|

| ARC-Challenge-chat (1172) | 92.15% | 90.53% | 85.58% | 96.76% |

| GPQA Diamond flex (198) | 51.52% | 34.85% | 26.26% | 73.74% |

| GSM8K-100 flex | 91.00% | 89.00% | 80.00% | 79.00% |

| MATH-500-100 math_verify | 95.00% | 62.00% | 66.00% | 59.00% |

| AIME 2024 (30) | 80.00% | 10.00% | 10.00% | 56.67% |

| IFEval-100 (prompt_strict) | 92.00% | 68.00% | 54.00% | 93.00% |

| HumanEval-164 chat | 98.17% | 90.85% | 87.20% | 89.02% |

| HumanEval+-164 chat | 92.07% | 84.76% † | 83.54% | 80.49% |

| LCB-medium-55 v4 | 98.18% | 18.18% † | 12.73% | 58.18% |

| MultiPL-E-100 (macro) | 89.67% | 84.67% | 80.67% | 80.33% |

† Qwen2.5-Coder-14B HumanEval+ / LCB-medium-55 are the same-stack GGUF HE+ sweep numbers

(not re-run in this chain). All Qwen cells are the same-host reference runs used on the

v6-coder card — Qwen is a

fixed reference, so the columns are identical across the cohort; only the Gemma column changes.

> Note on Qwen3.5-9B. Qwen3.5-9B is a verbose, slow thinking model: it emits long

> <think> reasoning chains (often ≥1900 tokens even on a trivial GSM8K question), so it runs

> several× slower per question than the non-reasoning Qwen2.5-Coder models — well beyond what

> its 9B size would suggest. Its GSM8K / MATH-500 / GPQA cells were re-run after a harness fix

> (under batched, reasoning-parsed serving the verbose thinking intermittently left the final

> answer inside the reasoning block, mis-scored as empty content).

Recipe (summary)

98e prune from 128e via the fkbroad code recipe (generate_drop_map_v5: generic_code 3×,

targeted_lcb_medium_55 2×, all other targeting 0; target=98, protect_top=16, alpha=2.0,

strategy=max + breadth_bonus, no per-layer floor clamp), then the **agentic loop-protection

force-keep** (agentic_eog, 46 experts, 0/46 dropped — the loop fix that replaces fs2440), then

the mandatory shared-FFN α=1.2 upweight. No targeted_gpqa term and no DERN fold. v7-coder uses code/LCB weight

3×/2×; the sibling v7-coderx uses 4×/3×. Full recipe is on the

bf16 card.

Intended use & limitations

Compact (~13 GB at Q4_K_M, single 12–16 GB GPU) Gemma 4 checkpoint for **agentic

coding and code reasoning**. Serve with the reasoning parser enabled. A research prune,

not an official Google release; generic_multilingual is de-weighted (0×) and graduate

science (GPQA) is a budget axis (51.52% vs 128e 67.17%). Prefer Q4_K_M or higher.

Lineage

128e → (v4 → v5 → v6-coder code line) → v7 competence-map rebuild → **fkbroad code3/lcb2

selection + agentic loop-protection force-keep** = v7-coder (loop-fixed; supersedes fs2440).

Built and evaluated on the omnimergekit toolchain.

Run ManniX-ITA/gemma-4-A4B-98e-v7-coder-it-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models