mannix-ita/gemma-4-a4b-109e-v3-it-gguf Q5_K_M GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
mannix-ita/gemma-4-a4b-109e-v3-it-gguf overview
GGUF quantizations of ManniX-ITA/gemma-4-A4B-109e-v3-it. All quants made using imatrix with calibration data v5.
Downloads
15,377
Likes
1
Pipeline
—
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
29 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gemma-4-A4B-109e-v3-CD-Q3_K_M.gguf | GGUF | Q3_K_M | 10.37 GB | Download |
| gemma-4-A4B-109e-v3-CD-Q4_K_M.gguf | GGUF | Q4_K_M | 11.08 GB | Download |
| gemma-4-A4B-109e-v3-CD-Q5_K_M.gguf | GGUF | Q5_K_M | 13.51 GB | Download |
| gemma-4-A4B-109e-v3-CD-Q6_K.gguf | GGUF | Q6_K | 15.80 GB | Download |
| gemma-4-A4B-109e-v3-F16.gguf | GGUF | F16 | 40.72 GB | Download |
| gemma-4-A4B-109e-v3-IQ2_M.gguf | GGUF | IQ2_M | 8.39 GB | Download |
| gemma-4-A4B-109e-v3-IQ2_S.gguf | GGUF | IQ2_S | 7.99 GB | Download |
| gemma-4-A4B-109e-v3-IQ2_XS.gguf | GGUF | IQ2_XS | 7.94 GB | Download |
| gemma-4-A4B-109e-v3-IQ2_XXS.gguf | GGUF | IQ2_XXS | 7.52 GB | Download |
| gemma-4-A4B-109e-v3-IQ3_M.gguf | GGUF | IQ3_M | 10.03 GB | Download |
| gemma-4-A4B-109e-v3-IQ3_XS.gguf | GGUF | IQ3_XS | 9.41 GB | Download |
| gemma-4-A4B-109e-v3-IQ3_XXS.gguf | GGUF | IQ3_XXS | 9.14 GB | Download |
| gemma-4-A4B-109e-v3-IQ4_NL.gguf | GGUF | IQ4_NL | 11.67 GB | Download |
| gemma-4-A4B-109e-v3-IQ4_XS.gguf | GGUF | IQ4_XS | 11.25 GB | Download |
| gemma-4-A4B-109e-v3-Q3_K_L.gguf | GGUF | Q3_K_L | 11.18 GB | Download |
| gemma-4-A4B-109e-v3-Q3_K_M.gguf | GGUF | Q3_K_M | 10.74 GB | Download |
| gemma-4-A4B-109e-v3-Q3_K_S.gguf | GGUF | Q3_K_S | 9.88 GB | Download |
| gemma-4-A4B-109e-v3-Q3_K_XL.gguf | GGUF | Q3_K_XL | 10.90 GB | Download |
| gemma-4-A4B-109e-v3-Q4_0.gguf | GGUF | — | 11.67 GB | Download |
| gemma-4-A4B-109e-v3-Q4_1.gguf | GGUF | — | 12.89 GB | Download |
| gemma-4-A4B-109e-v3-Q4_K_L.gguf | GGUF | Q4_K_L | 13.71 GB | Download |
| gemma-4-A4B-109e-v3-Q4_K_M.gguf | GGUF | Q4_K_M | 13.54 GB | Download |
| gemma-4-A4B-109e-v3-Q4_K_S.gguf | GGUF | Q4_K_S | 12.48 GB | Download |
| gemma-4-A4B-109e-v3-Q5_K_L.gguf | GGUF | Q5_K_L | 15.59 GB | Download |
| gemma-4-A4B-109e-v3-Q5_K_M.gguf | GGUF | Q5_K_M | 15.42 GB | Download |
| gemma-4-A4B-109e-v3-Q5_K_S.gguf | GGUF | Q5_K_S | 14.51 GB | Download |
| gemma-4-A4B-109e-v3-Q6_K.gguf | GGUF | Q6_K | 18.23 GB | Download |
| gemma-4-A4B-109e-v3-Q6_K_L.gguf | GGUF | Q6_K_L | 18.40 GB | Download |
| gemma-4-A4B-109e-v3-Q8_0.gguf | GGUF | — | 21.65 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": "ManniX-ITA/gemma-4-A4B-109e-v3-it",
"tags": [
"gguf",
"imatrix",
"quantized"
],
"license": "apache-2.0",
"frontmatter": {
"base_model": "ManniX-ITA/gemma-4-A4B-109e-v3-it",
"tags": [
"gguf",
"imatrix",
"quantized"
],
"license": "apache-2.0"
},
"hero_image_url": "",
"summary": "GGUF quantizations of ManniX-ITA/gemma-4-A4B-109e-v3-it. All quants made using imatrix with calibration data v5.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model: ManniX-ITA/gemma-4-A4B-109e-v3-it\ntags:\n - gguf\n - imatrix\n - quantized\nlicense: apache-2.0\n---\n\n# gemma-4-A4B-109e-v3-GGUF\n\nGGUF quantizations of [ManniX-ITA/gemma-4-A4B-109e-v3-it](https://huggingface.co/ManniX-ITA/gemma-4-A4B-109e-v3-it).\n\nAll quants made using imatrix with [calibration data v5](https://gist.github.com/bartowski1182/82ae9b520227f57d79ba04add13d0d0d).\n\n## Available Quantizations\n\n### Full precision\n| Filename | Quant | Size |\n|---|---|---|\n| `gemma-4-A4B-109e-v3-F16.gguf` | F16 | 40.72 GB |\n\n### Standard bartowski-style quants (recommended for most users)\n| Filename | Quant | Size | Notes |\n|---|---|---|---|\n| `gemma-4-A4B-109e-v3-Q8_0.gguf` | Q8_0 | 21.65 GB | near-lossless, large |\n| `gemma-4-A4B-109e-v3-Q6_K_L.gguf` | Q6_K_L | 18.40 GB | Q6_K with Q8_0 embed/output |\n| `gemma-4-A4B-109e-v3-Q6_K.gguf` | Q6_K | 18.23 GB | very high quality |\n| `gemma-4-A4B-109e-v3-Q5_K_L.gguf` | Q5_K_L | 15.59 GB | Q5_K_M with Q8_0 embed/output |\n| `gemma-4-A4B-109e-v3-Q5_K_M.gguf` | Q5_K_M | 15.42 GB | high quality |\n| `gemma-4-A4B-109e-v3-Q5_K_S.gguf` | Q5_K_S | 14.51 GB | high quality |\n| `gemma-4-A4B-109e-v3-Q4_K_L.gguf` | Q4_K_L | 13.71 GB | Q4_K_M with Q8_0 embed/output |\n| `gemma-4-A4B-109e-v3-Q4_K_M.gguf` | Q4_K_M | 13.54 GB | **recommended for most hardware** |\n| `gemma-4-A4B-109e-v3-Q4_K_S.gguf` | Q4_K_S | 12.48 GB | slightly worse than Q4_K_M, a bit smaller |\n| `gemma-4-A4B-109e-v3-Q4_1.gguf` | Q4_1 | 12.89 GB | legacy 4-bit |\n| `gemma-4-A4B-109e-v3-Q4_0.gguf` | Q4_0 | 11.67 GB | legacy 4-bit |\n| `gemma-4-A4B-109e-v3-IQ4_NL.gguf` | IQ4_NL | 11.67 GB | i-quant, similar to Q4_0 but better |\n| `gemma-4-A4B-109e-v3-IQ4_XS.gguf` | IQ4_XS | 11.25 GB | i-quant, smaller than Q4_K_S |\n| `gemma-4-A4B-109e-v3-Q3_K_XL.gguf` | Q3_K_XL | 10.90 GB | Q3_K_L with Q8_0 embed/output |\n| `gemma-4-A4B-109e-v3-Q3_K_L.gguf` | Q3_K_L | 11.18 GB | OK quality at 3-bit |\n| `gemma-4-A4B-109e-v3-Q3_K_M.gguf` | Q3_K_M | 10.74 GB | lower quality |\n| `gemma-4-A4B-109e-v3-Q3_K_S.gguf` | Q3_K_S | 9.88 GB | low quality, not recommended |\n| `gemma-4-A4B-109e-v3-IQ3_M.gguf` | IQ3_M | 10.03 GB | i-quant, surprisingly good at 3-bit |\n| `gemma-4-A4B-109e-v3-IQ3_XS.gguf` | IQ3_XS | 9.41 GB | i-quant, aggressive |\n| `gemma-4-A4B-109e-v3-IQ3_XXS.gguf` | IQ3_XXS | 9.14 GB | i-quant, very aggressive |\n| `gemma-4-A4B-109e-v3-IQ2_M.gguf` | IQ2_M | 8.39 GB | i-quant, lossy but usable |\n| `gemma-4-A4B-109e-v3-IQ2_S.gguf` | IQ2_S | 7.99 GB | i-quant, low quality |\n| `gemma-4-A4B-109e-v3-IQ2_XS.gguf` | IQ2_XS | 7.94 GB | i-quant, very low quality |\n| `gemma-4-A4B-109e-v3-IQ2_XXS.gguf` | IQ2_XXS | 7.52 GB | i-quant, smallest — expect quality loss |\n\n### ContribDynamic (CD) experimental quants\n\nPer-layer dynamic quantization driven by our **clean fp32 expert-contribution analysis**: tensors in high-importance layers (L0, L1-6, L10) get higher precision; low-importance layers (L11-29) get lower precision. Inspired by Unsloth's Dynamic (UD) approach but using our own profiling data. See the CD block below for per-layer details.\n\n| Filename | Quant | Size | Notes |\n|---|---|---|---|\n| `gemma-4-A4B-109e-v3-CD-Q6_K.gguf` | CD-Q6_K | 15.80 GB | hybrid Q6_K / Q5_K / Q4_K per layer — **same footprint as Q5_K_L, Q6_K quality in the layers that matter** |\n| `gemma-4-A4B-109e-v3-CD-Q5_K_M.gguf` | CD-Q5_K_M | 13.51 GB | hybrid Q5_K / Q4_K / Q3_K per layer |\n| `gemma-4-A4B-109e-v3-CD-Q4_K_M.gguf` | CD-Q4_K_M | 11.08 GB | hybrid Q4_K / Q3_K per layer |\n| `gemma-4-A4B-109e-v3-CD-Q3_K_M.gguf` | CD-Q3_K_M | 10.37 GB | hybrid Q3_K / Q2_K per layer |\n\nCD-Q2_K was attempted but not shipped — the low-tier IQ2_S assignment in its tensor-type map requires `--imatrix` at quantize time, and the initial build pipeline missed that guard. The bug is now fixed in `scripts/quantize_gguf.py` for future builds.\n\n## How to Use\n\nWith [llama.cpp](https://github.com/ggml-org/llama.cpp):\n```bash\nllama-server -m gemma-4-A4B-109e-v3-Q4_K_M.gguf -c 8192 -ngl 99 \\\n --reasoning-format deepseek --reasoning-budget 16384\n```\n\n> **Important**: always pass `--reasoning-format deepseek --reasoning-budget 16384` when serving Gemma 4. Without the budget, chemistry-heavy prompts trigger unbounded chain-of-thought that crashes the server. See the inference settings section below.\n\nWith [ollama](https://ollama.ai) (requires Modelfile or HF direct load).\n\n---\n\n## Original Model Card\n\n# Gemma 4 A4B 109-Expert v3 (22.4B)\n\nExpert-pruned version of [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it), reduced from 128 to 109 experts per MoE layer using a **clean fp32 teacher-force analysis** on GPQA Diamond.\n\n| | Original (128e) | This model (109e v3) | Delta |\n|---|---|---|---|\n| **Total params** | 26B | 22.4B | -14% |\n| **Experts per layer** | 128 | 109 | -19/layer |\n| **Top-k routing** | 8 | 8 | — |\n| **GGUF Q6_K size** | 18 GB | 18 GB | — |\n| **bf16 size** | 50 GB | 42 GB | -16% |\n| **GPQA Diamond (Q6_K, lm-eval + patches)** | **81.82%** | **80.30%** | **−1.52 pp** |\n\n**The v3 clean teacher-force drop only costs 1.52 percentage points vs the full 128-expert reference at Q6_K.**\n\n## Why v3?\n\nEarlier `109e` / `109e-v2` releases were built from a drop map that turned out to have **~43% wrong expert selections** due to a `bf16 → .norm()` overflow bug in the teacher-force analysis script. The bug produced `NaN`/`inf` entries in the per-expert norms of layers 11-29 and corrupted the ranking used to pick which experts to drop.\n\n**v3 fixes it**:\n- `.float().norm()` for the bf16→fp32 reduction (eliminates overflow)\n- NaN guard on hidden states (skip upstream-NaN tokens)\n- 4096-token truncation on calibration inputs\n- `attn_implementation=\"eager\"` (Gemma 4 `head_dim=512` is not supported by FlashAttention 2)\n- Analysis re-run over the full GPQA Diamond 196-question set on the 128e reference, producing a new clean drop map (`teacher_force_109e_p16_clean.json`).\n\nThe v3 clean drop map differs on roughly **43% of expert selections** compared to v2 — it is effectively a different model. And at Q6_K it beats v2 by **+1.51 pp** on GPQA Diamond (80.30% vs 78.79%).\n\n## Pruning Method\n\n### Teacher-Force Expert Analysis\n\nThe pruning decision is based on measuring the **actual contribution** of each expert during teacher-forced inference on GPQA Diamond prompts using the **full 128-expert reference model** as the teacher.\n\n**Process** (`scripts/teacher_force_analysis.py`):\n1. **Teacher passes**: For each of 196 GPQA Diamond questions, run the 128e reference model through the complete `prompt + correct CoT + answer` sequence in a single forward pass with teacher forcing.\n2. **Per-expert norms**: At every MoE layer, hook the experts module and recompute each activated expert's output `||routing_weight · expert_output||_2` on GPU, aggregated in fp32.\n3. **Per-question top-16 protection**: For each question independently, rank experts per layer by weighted norm and mark the top 16 as \"protected for this question\".\n4. **Aggregate across questions**: Union the per-question top-16 sets across all 196 GPQA questions. The top 109 experts per layer by coverage are kept; the bottom 19 are dropped.\n\nThis gives a drop map specifically optimized for the GPQA Diamond task domain while remaining deterministic and reproducible.\n\n### Key Findings (clean fp32 analysis)\n\n- **Experts are NOT topic-specialized**: confirmed across domains.\n- **The bf16 bug mattered**: ~2% of per-expert norm entries in layers 11-29 were `NaN` or `inf` in the corrupted analysis, dragging those experts to artificially extreme ranks. Fixing to fp32 changes 43% of the \"protected top-16 per question\" decisions.\n- **Expert weight similarity is near zero**: cosine similarity between expert weight matrices maxes at ~0.05 — merging experts by averaging destroys the model. Expert *dropping* (what we do here) is the only viable structural compression.\n\n### Pruning Decision\n\n**Uniform 109 experts per layer** (19 dropped per layer), based on the clean teacher-force aggregate ranking. The router `proj.weight` is resized from `[128, hidden]` to `[109, hidden]`, keeping only the rows corresponding to retained experts. The router naturally adapts — removed experts simply become unavailable and the top-8 selection falls through to the next-best available expert. No fine-tuning needed.\n\nThe drop map used to build this model is deterministic and stored in `expert_drop_metadata.json`.\n\n## GPQA Diamond Evaluation\n\n### Setup\n\nAll variants are evaluated with the **same canonical script** ([`scripts/eval_gpqa_v3.sh`](https://github.com/ManniX-ITA)) for apples-to-apples comparison:\n\n- **Quantization**: GGUF Q6_K via llama.cpp `llama-quantize` (imatrix calibration)\n- **Inference**: llama.cpp `llama-server` (OpenAI-compatible API)\n- **Evaluation**: lm-evaluation-harness, task `gpqa_diamond_cot_zeroshot`\n- **Backend**: `local-chat-completions` against llama.cpp API\n- **GPU**: NVIDIA RTX 3090 (24 GB), 99 layers offloaded\n\n### Configuration (locked, all variants identical)\n\n| Parameter | Value |\n|---|---|\n| Context size | 32768 tokens |\n| Reasoning format | `deepseek` (separates thinking into `reasoning_content`) |\n| Reasoning budget | **16384 tokens** |\n| max_gen_toks | 24576 |\n| Temperature | **1.0** (Gemma 4 official sampling) |\n| top_p | 0.95 |\n| top_k | 64 |\n| Seed | 42 |\n| DRY multiplier | 0.5 (anti-degenerate-loop sampler, proven to fix Q53 \"re-re-re\" crash) |\n| Tokenizer | `google/gemma-4-26B-A4B-it` (original, unmodified) |\n\nThe **reasoning budget** is critical: without it, Gemma 4 enters overthinking loops on hard questions and exhausts the full context without committing to an answer. This is base-model behavior — the 128-expert reference does the same.\n\n### Results (Q6_K, full 198-question GPQA Diamond)\n\n| Model | Drop map | Score | flex-extract % | vs 128e |\n|---|---|---|---|---|\n| **gemma-4-26B-A4B-it** (reference) | — | **162/198** | **81.82%** | — |\n| **gemma-4-A4B-109e-v3** (this) | clean fp32 teacher-force | **159/198** | **80.30%** | **−1.52 pp** |\n| gemma-4-A4B-109e-v2 (old v2) | corrupted bf16 teacher-force | 156/198 | 78.79% | −3.03 pp |\n| gemma-4-A4B-120e-v4 | corrupted bf16 teacher-force | 154/198 | 77.78% | −4.04 pp |\n| gemma-4-A4B-120e-v3 | old greedy | 152/198 | 76.77% | −5.05 pp |\n| gemma-4-A4B-109e (older) | old greedy contribution | 148/198 | 74.75% | −7.07 pp |\n\n### Patch Methodology\n\nThe raw lm-eval run on **this model** scored **154/198 = 77.78%**. 10 questions returned empty or truncated responses because of a llama.cpp PEG parser bug (upstream issue #21418, merged but not fully fixed) that triggers on some chemistry/reasoning questions when `reasoning-format deepseek` is active.\n\nThese 10 were re-run using the llama.cpp `/completion` endpoint with a short prefilled continuation (`\"The answer is (\"`) at `n_predict=10`. This bypasses the reasoning parser entirely — the model commits to a single letter in 1-2 tokens, with no channel/thought markers to trip up the parser.\n\n**Patch result: 5 of 10 were correct** (expected 2.5 if random; the model has real signal on the \"missing\" questions when given a path to answer them). Final patched score: **159/198 = 80.30%**.\n\nThe same patch method was applied consistently to all compared variants above (128e, 109e-v2, 120e-v4) — so the ranking is fair.\n\n### Wrong-answer breakdown (Q6_K, flexible-extract)\n\n- **128 correct** from raw lm-eval + **5 patched** = 159 correct\n- 29 wrong\n- 10 would-be-invalid → patched (5 correct, 5 wrong)\n\n## Architecture\n\nUnchanged from the original except `num_experts: 109` (was 128):\n\n- **Layers**: 30\n- **Hidden size**: 2816\n- **Expert intermediate size**: 704 per expert\n- **Dense MLP intermediate size**: 2112 (always active)\n- **Top-k routing**: 8\n- **Attention**: Hybrid sliding (5) + global (1) pattern, `head_dim=512` (requires `attn_implementation=\"eager\"` — FlashAttention 2 does not support head_dim > 256)\n- **Vocabulary**: 262,144\n\n## Files\n\n- `config.json` — Model config with `num_experts: 109`\n- `model-0000N-of-00009.safetensors` — Model weights (9 shards, 41.7 GB total bf16)\n- `expert_drop_metadata.json` — Per-layer keep/drop expert indices and methodology\n- `tokenizer.json` / `tokenizer_config.json` / `chat_template.jinja` — from the base 26B-A4B-it (unchanged)\n\n## Usage\n\n### Transformers\n\n```python\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\nimport torch\n\nmodel = AutoModelForCausalLM.from_pretrained(\n \"ManniX-ITA/gemma-4-A4B-109e-v3-it\",\n torch_dtype=torch.bfloat16,\n device_map=\"auto\",\n attn_implementation=\"eager\", # Gemma 4 head_dim=512 is not supported by FA2\n)\ntok = AutoTokenizer.from_pretrained(\"ManniX-ITA/gemma-4-A4B-109e-v3-it\")\n\nmsgs = [{\"role\": \"user\", \"content\": \"Explain the Heisenberg uncertainty principle.\"}]\ninputs = tok.apply_chat_template(msgs, return_tensors=\"pt\", add_generation_prompt=True).to(model.device)\nout = model.generate(inputs, max_new_tokens=512, do_sample=True, temperature=1.0, top_p=0.95, top_k=64)\nprint(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))\n```\n\n### llama.cpp (recommended for consumer hardware)\n\nGGUF quantizations are available at [ManniX-ITA/gemma-4-A4B-109e-v3-it-GGUF](https://huggingface.co/ManniX-ITA/gemma-4-A4B-109e-v3-it-GGUF). The **Q6_K** quant fits comfortably in ~19 GB VRAM and was used for all benchmarks above.\n\n```bash\nllama-server -m gemma-4-A4B-109e-v3-Q6_K.gguf \\\n --port 8099 -c 32768 -ngl 99 \\\n --reasoning-format deepseek --reasoning-budget 16384 \\\n --temp 1.0 --top-p 0.95 --top-k 64 --dry-multiplier 0.5 --seed 42\n```\n\nOr convert locally:\n\n```bash\npython llama.cpp/convert_hf_to_gguf.py gemma-4-A4B-109e-v3-it --outfile model-f16.gguf --outtype f16\nllama.cpp/build/bin/llama-quantize model-f16.gguf model-Q6_K.gguf Q6_K\n```\n\n## Reproduction\n\nThe full pipeline is deterministic and bit-reproducible. Same base model + same drop map + same script = bit-identical safetensors (verified across two independent rebuilds at different sites, all 9 shards SHA256-matched):\n\n1. `scripts/teacher_force_analysis.py` — Clean fp32 per-expert contribution analysis on GPQA Diamond\n2. `scripts/generate_drop_map.py` — Aggregate per-question top-16 protections into a global drop map\n3. `scripts/expert_drop.py` — Deterministic expert pruning from the drop map\n4. `scripts/eval_gpqa_v3.sh` — Canonical locked-methodology evaluation via llama.cpp + lm-eval\n\n## License\n\nThis model inherits the [Gemma license](https://ai.google.dev/gemma/terms) from the base model.\n\n## Acknowledgements\n\n- Google for the base Gemma 4 26B-A4B-it model\n- The GPQA Diamond benchmark (Rein et al., 2023)\n- bartowski for the calibration data v5 used in imatrix-based GGUF quantization\n",
"related_quantizations": []
},
"tags": [
"gguf",
"imatrix",
"quantized",
"base_model:ManniX-ITA/gemma-4-A4B-109e-v3-it",
"base_model:quantized:ManniX-ITA/gemma-4-A4B-109e-v3-it",
"license:apache-2.0",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 1,
"downloads": 15377,
"gated": false,
"private": false,
"last_modified": "2026-04-11T18:41:57.000Z",
"created_at": "2026-04-11T09:00:17.000Z",
"pipeline_tag": "",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69da0da12525d7abb95ce96f",
"id": "ManniX-ITA/gemma-4-A4B-109e-v3-it-GGUF",
"modelId": "ManniX-ITA/gemma-4-A4B-109e-v3-it-GGUF",
"sha": "c45d350f25f22d1dcc09e022166b86c8d7f41701",
"createdAt": "2026-04-11T09:00:17.000Z",
"lastModified": "2026-04-11T18:41:57.000Z",
"author": "ManniX-ITA",
"downloads": 15377,
"likes": 1,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "",
"siblings_count": 31
}