mannix-ita/gemma-4-a4b-109e-it-gguf Q3_K_S GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
mannix-ita/gemma-4-a4b-109e-it-gguf overview
GGUF quantizations of ManniX-ITA/gemma-4-A4B-109e-it — an expert-pruned Gemma 4 26B-A4B (128 to 109 experts, 26B to 22.4B params). All standard quants made using imatrix with calibration data v5.
Downloads
9,866
Likes
4
Pipeline
—
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
30 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gemma-4-A4B-109e-it-CD-Q3_K_M.gguf | GGUF | Q3_K_M | 10.37 GB | Download |
| gemma-4-A4B-109e-it-CD-Q4_K_M.gguf | GGUF | Q4_K_M | 11.08 GB | Download |
| gemma-4-A4B-109e-it-CD-Q5_K_M.gguf | GGUF | Q5_K_M | 13.51 GB | Download |
| gemma-4-A4B-109e-it-CD-Q6_K.gguf | GGUF | Q6_K | 15.80 GB | Download |
| gemma-4-A4B-109e-it-F16.gguf | GGUF | F16 | 40.72 GB | Download |
| gemma-4-A4B-109e-it-IQ2_M.gguf | GGUF | IQ2_M | 8.39 GB | Download |
| gemma-4-A4B-109e-it-IQ2_S.gguf | GGUF | IQ2_S | 7.99 GB | Download |
| gemma-4-A4B-109e-it-IQ2_XS.gguf | GGUF | IQ2_XS | 7.94 GB | Download |
| gemma-4-A4B-109e-it-IQ2_XXS.gguf | GGUF | IQ2_XXS | 7.52 GB | Download |
| gemma-4-A4B-109e-it-IQ3_M.gguf | GGUF | IQ3_M | 10.03 GB | Download |
| gemma-4-A4B-109e-it-IQ3_XS.gguf | GGUF | IQ3_XS | 9.41 GB | Download |
| gemma-4-A4B-109e-it-IQ3_XXS.gguf | GGUF | IQ3_XXS | 9.14 GB | Download |
| gemma-4-A4B-109e-it-IQ4_NL.gguf | GGUF | IQ4_NL | 11.67 GB | Download |
| gemma-4-A4B-109e-it-IQ4_XS.gguf | GGUF | IQ4_XS | 11.25 GB | Download |
| gemma-4-A4B-109e-it-Q2_K.gguf | GGUF | Q2_K | 8.57 GB | Download |
| gemma-4-A4B-109e-it-Q3_K_L.gguf | GGUF | Q3_K_L | 11.18 GB | Download |
| gemma-4-A4B-109e-it-Q3_K_M.gguf | GGUF | Q3_K_M | 10.74 GB | Download |
| gemma-4-A4B-109e-it-Q3_K_S.gguf | GGUF | Q3_K_S | 9.88 GB | Download |
| gemma-4-A4B-109e-it-Q3_K_XL.gguf | GGUF | Q3_K_XL | 10.90 GB | Download |
| gemma-4-A4B-109e-it-Q4_0.gguf | GGUF | — | 11.67 GB | Download |
| gemma-4-A4B-109e-it-Q4_1.gguf | GGUF | — | 12.89 GB | Download |
| gemma-4-A4B-109e-it-Q4_K_L.gguf | GGUF | Q4_K_L | 13.71 GB | Download |
| gemma-4-A4B-109e-it-Q4_K_M.gguf | GGUF | Q4_K_M | 13.54 GB | Download |
| gemma-4-A4B-109e-it-Q4_K_S.gguf | GGUF | Q4_K_S | 12.48 GB | Download |
| gemma-4-A4B-109e-it-Q5_K_L.gguf | GGUF | Q5_K_L | 15.59 GB | Download |
| gemma-4-A4B-109e-it-Q5_K_M.gguf | GGUF | Q5_K_M | 15.42 GB | Download |
| gemma-4-A4B-109e-it-Q5_K_S.gguf | GGUF | Q5_K_S | 14.51 GB | Download |
| gemma-4-A4B-109e-it-Q6_K.gguf | GGUF | Q6_K | 18.23 GB | Download |
| gemma-4-A4B-109e-it-Q6_K_L.gguf | GGUF | Q6_K_L | 18.40 GB | Download |
| gemma-4-A4B-109e-it-Q8_0.gguf | GGUF | — | 21.65 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": "ManniX-ITA/gemma-4-A4B-109e-it",
"tags": [
"gguf",
"imatrix",
"quantized",
"gemma4",
"moe",
"expert-pruning"
],
"license": "gemma",
"frontmatter": {
"base_model": "ManniX-ITA/gemma-4-A4B-109e-it",
"tags": [
"gguf",
"imatrix",
"quantized",
"gemma4",
"moe",
"expert-pruning"
],
"license": "gemma"
},
"hero_image_url": "",
"summary": "GGUF quantizations of ManniX-ITA/gemma-4-A4B-109e-it — an expert-pruned Gemma 4 26B-A4B (128 to 109 experts, 26B to 22.4B params). All standard quants made using imatrix with calibration data v5.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model: ManniX-ITA/gemma-4-A4B-109e-it\ntags:\n - gguf\n - imatrix\n - quantized\n - gemma4\n - moe\n - expert-pruning\nlicense: gemma\n---\n\n# gemma-4-A4B-109e-it-GGUF\n\nGGUF quantizations of [ManniX-ITA/gemma-4-A4B-109e-it](https://huggingface.co/ManniX-ITA/gemma-4-A4B-109e-it) — an expert-pruned Gemma 4 26B-A4B (128 to 109 experts, 26B to 22.4B params).\n\nAll standard quants made using imatrix with [calibration data v5](https://gist.github.com/bartowski1182/82ae9b520227f57d79ba04add13d0d0d).\n\n## ContribDynamic (CD) Quants\n\nCD quants use **per-layer dynamic quantization** based on actual expert contribution analysis of the model. Important layers (early layers that contribute more to the residual stream) get higher precision, while less important layers get lower precision.\n\nFor a CD-Q4_K_M target:\n- **Layer 0** (highest importance): Q5_K precision\n- **Layers 1-6, 10** (medium importance): Q4_K precision \n- **Layers 7-29** (lower importance): Q3_K precision\n- **Output/embeddings**: Q8_0 precision\n\nThis approach is inspired by Unsloth's UD quantization but uses our own expert contribution profiling data derived from measuring actual norms across 40 calibration prompts.\n\n## Available Quantizations\n\n| Quantization | Size |\n|---|---|\n| Q8_0 | 21.65 GB |\n| Q6_K_L | 18.40 GB |\n| Q6_K | 18.23 GB |\n| Q5_K_L | 15.59 GB |\n| Q5_K_M | 15.42 GB |\n| Q5_K_S | 14.51 GB |\n| Q4_K_L | 13.71 GB |\n| Q4_K_M | 13.54 GB |\n| Q4_1 | 12.89 GB |\n| Q4_K_S | 12.48 GB |\n| Q4_0 | 11.67 GB |\n| IQ4_NL | 11.67 GB |\n| IQ4_XS | 11.25 GB |\n| Q3_K_XL | 10.90 GB |\n| IQ3_M | 10.03 GB |\n| Q3_K_L | 11.18 GB |\n| Q3_K_M | 10.74 GB |\n| Q3_K_S | 9.88 GB |\n| IQ3_XS | 9.41 GB |\n| IQ3_XXS | 9.14 GB |\n| Q2_K | 8.57 GB |\n| IQ2_M | 8.39 GB |\n| IQ2_S | 7.99 GB |\n| IQ2_XS | 7.94 GB |\n| IQ2_XXS | 7.52 GB |\n| CD-Q6_K | 15.80 GB |\n| CD-Q5_K_M | 13.51 GB |\n| CD-Q4_K_M | 11.08 GB |\n| CD-Q3_K_M | 10.37 GB |\n\n\nAll quants passed a 3-question sanity check (capital cities in JSON format) via llama.cpp before upload.\n\n## How to Use\n\n\n\n## Original Model\n\nSee [ManniX-ITA/gemma-4-A4B-109e-it](https://huggingface.co/ManniX-ITA/gemma-4-A4B-109e-it) for the full model card, pruning methodology, and benchmark results (71.7% GPQA Diamond).\n\n## License\n\n[Gemma license](https://ai.google.dev/gemma/terms)\n",
"related_quantizations": []
},
"tags": [
"gguf",
"imatrix",
"quantized",
"gemma4",
"moe",
"expert-pruning",
"base_model:ManniX-ITA/gemma-4-A4B-109e-it",
"base_model:quantized:ManniX-ITA/gemma-4-A4B-109e-it",
"license:gemma",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 4,
"downloads": 9866,
"gated": false,
"private": false,
"last_modified": "2026-04-05T19:30:14.000Z",
"created_at": "2026-04-05T16:13:34.000Z",
"pipeline_tag": "",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69d28a2ef1a967ac148400f7",
"id": "ManniX-ITA/gemma-4-A4B-109e-it-GGUF",
"modelId": "ManniX-ITA/gemma-4-A4B-109e-it-GGUF",
"sha": "136c4d065cb0ad9db817020dc3b1f6d0d6c3aaba",
"createdAt": "2026-04-05T16:13:34.000Z",
"lastModified": "2026-04-05T19:30:14.000Z",
"author": "ManniX-ITA",
"downloads": 9866,
"likes": 4,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "",
"siblings_count": 32
}