GraySoft
Projects Models About FAQ Contact Download guIDE →

mannix-ita/gemma-4-a4b-109e-it-gguf - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

mannix-ita/gemma-4-a4b-109e-it-gguf overview

GGUF quantizations of ManniX-ITA/gemma-4-A4B-109e-it — an expert-pruned Gemma 4 26B-A4B (128 to 109 experts, 26B to 22.4B params). All standard quants made using imatrix with calibration data v5.

ggufimatrixquantizedgemma4moeexpert-pruningbase_model:ManniX-ITA/gemma-4-A4B-109e-itbase_model:quantized:ManniX-ITA/gemma-4-A4B-109e-itlicense:gemmaendpoints_compatibleregion:usconversational
mannix-ita/gemma-4-a4b-109e-it-gguf visual
Downloads
9,866
Likes
4
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

30 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
gemma-4-A4B-109e-it-CD-Q3_K_M.gguf GGUF Q3_K_M 10.37 GB Download
gemma-4-A4B-109e-it-CD-Q4_K_M.gguf GGUF Q4_K_M 11.08 GB Download
gemma-4-A4B-109e-it-CD-Q5_K_M.gguf GGUF Q5_K_M 13.51 GB Download
gemma-4-A4B-109e-it-CD-Q6_K.gguf GGUF Q6_K 15.80 GB Download
gemma-4-A4B-109e-it-F16.gguf GGUF F16 40.72 GB Download
gemma-4-A4B-109e-it-IQ2_M.gguf GGUF IQ2_M 8.39 GB Download
gemma-4-A4B-109e-it-IQ2_S.gguf GGUF IQ2_S 7.99 GB Download
gemma-4-A4B-109e-it-IQ2_XS.gguf GGUF IQ2_XS 7.94 GB Download
gemma-4-A4B-109e-it-IQ2_XXS.gguf GGUF IQ2_XXS 7.52 GB Download
gemma-4-A4B-109e-it-IQ3_M.gguf GGUF IQ3_M 10.03 GB Download
gemma-4-A4B-109e-it-IQ3_XS.gguf GGUF IQ3_XS 9.41 GB Download
gemma-4-A4B-109e-it-IQ3_XXS.gguf GGUF IQ3_XXS 9.14 GB Download
gemma-4-A4B-109e-it-IQ4_NL.gguf GGUF IQ4_NL 11.67 GB Download
gemma-4-A4B-109e-it-IQ4_XS.gguf GGUF IQ4_XS 11.25 GB Download
gemma-4-A4B-109e-it-Q2_K.gguf GGUF Q2_K 8.57 GB Download
gemma-4-A4B-109e-it-Q3_K_L.gguf GGUF Q3_K_L 11.18 GB Download
gemma-4-A4B-109e-it-Q3_K_M.gguf GGUF Q3_K_M 10.74 GB Download
gemma-4-A4B-109e-it-Q3_K_S.gguf GGUF Q3_K_S 9.88 GB Download
gemma-4-A4B-109e-it-Q3_K_XL.gguf GGUF Q3_K_XL 10.90 GB Download
gemma-4-A4B-109e-it-Q4_0.gguf GGUF 11.67 GB Download
gemma-4-A4B-109e-it-Q4_1.gguf GGUF 12.89 GB Download
gemma-4-A4B-109e-it-Q4_K_L.gguf GGUF Q4_K_L 13.71 GB Download
gemma-4-A4B-109e-it-Q4_K_M.gguf GGUF Q4_K_M 13.54 GB Download
gemma-4-A4B-109e-it-Q4_K_S.gguf GGUF Q4_K_S 12.48 GB Download
gemma-4-A4B-109e-it-Q5_K_L.gguf GGUF Q5_K_L 15.59 GB Download
gemma-4-A4B-109e-it-Q5_K_M.gguf GGUF Q5_K_M 15.42 GB Download
gemma-4-A4B-109e-it-Q5_K_S.gguf GGUF Q5_K_S 14.51 GB Download
gemma-4-A4B-109e-it-Q6_K.gguf GGUF Q6_K 18.23 GB Download
gemma-4-A4B-109e-it-Q6_K_L.gguf GGUF Q6_K_L 18.40 GB Download
gemma-4-A4B-109e-it-Q8_0.gguf GGUF 21.65 GB Download

Model Details Live

Model Slug
mannix-ita/gemma-4-a4b-109e-it-gguf
Author
ManniX-ITA
Pipeline Task
Library
Created
2026-04-05
Last Modified
2026-04-05
Gated
No
Private
No
HF SHA
136c4d065cb0ad9db817020dc3b1f6d0d6c3aaba
License
gemma
Language
Unknown
Base Model
ManniX-ITA/gemma-4-A4B-109e-it

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": "ManniX-ITA/gemma-4-A4B-109e-it",
    "tags": [
      "gguf",
      "imatrix",
      "quantized",
      "gemma4",
      "moe",
      "expert-pruning"
    ],
    "license": "gemma",
    "frontmatter": {
      "base_model": "ManniX-ITA/gemma-4-A4B-109e-it",
      "tags": [
        "gguf",
        "imatrix",
        "quantized",
        "gemma4",
        "moe",
        "expert-pruning"
      ],
      "license": "gemma"
    },
    "hero_image_url": "",
    "summary": "GGUF quantizations of ManniX-ITA/gemma-4-A4B-109e-it — an expert-pruned Gemma 4 26B-A4B (128 to 109 experts, 26B to 22.4B params). All standard quants made using imatrix with calibration data v5.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model: ManniX-ITA/gemma-4-A4B-109e-it\ntags:\n  - gguf\n  - imatrix\n  - quantized\n  - gemma4\n  - moe\n  - expert-pruning\nlicense: gemma\n---\n\n# gemma-4-A4B-109e-it-GGUF\n\nGGUF quantizations of [ManniX-ITA/gemma-4-A4B-109e-it](https://huggingface.co/ManniX-ITA/gemma-4-A4B-109e-it) — an expert-pruned Gemma 4 26B-A4B (128 to 109 experts, 26B to 22.4B params).\n\nAll standard quants made using imatrix with [calibration data v5](https://gist.github.com/bartowski1182/82ae9b520227f57d79ba04add13d0d0d).\n\n## ContribDynamic (CD) Quants\n\nCD quants use **per-layer dynamic quantization** based on actual expert contribution analysis of the model. Important layers (early layers that contribute more to the residual stream) get higher precision, while less important layers get lower precision.\n\nFor a CD-Q4_K_M target:\n- **Layer 0** (highest importance): Q5_K precision\n- **Layers 1-6, 10** (medium importance): Q4_K precision  \n- **Layers 7-29** (lower importance): Q3_K precision\n- **Output/embeddings**: Q8_0 precision\n\nThis approach is inspired by Unsloth's UD quantization but uses our own expert contribution profiling data derived from measuring actual  norms across 40 calibration prompts.\n\n## Available Quantizations\n\n| Quantization | Size |\n|---|---|\n| Q8_0 | 21.65 GB |\n| Q6_K_L | 18.40 GB |\n| Q6_K | 18.23 GB |\n| Q5_K_L | 15.59 GB |\n| Q5_K_M | 15.42 GB |\n| Q5_K_S | 14.51 GB |\n| Q4_K_L | 13.71 GB |\n| Q4_K_M | 13.54 GB |\n| Q4_1 | 12.89 GB |\n| Q4_K_S | 12.48 GB |\n| Q4_0 | 11.67 GB |\n| IQ4_NL | 11.67 GB |\n| IQ4_XS | 11.25 GB |\n| Q3_K_XL | 10.90 GB |\n| IQ3_M | 10.03 GB |\n| Q3_K_L | 11.18 GB |\n| Q3_K_M | 10.74 GB |\n| Q3_K_S | 9.88 GB |\n| IQ3_XS | 9.41 GB |\n| IQ3_XXS | 9.14 GB |\n| Q2_K | 8.57 GB |\n| IQ2_M | 8.39 GB |\n| IQ2_S | 7.99 GB |\n| IQ2_XS | 7.94 GB |\n| IQ2_XXS | 7.52 GB |\n| CD-Q6_K | 15.80 GB |\n| CD-Q5_K_M | 13.51 GB |\n| CD-Q4_K_M | 11.08 GB |\n| CD-Q3_K_M | 10.37 GB |\n\n\nAll quants passed a 3-question sanity check (capital cities in JSON format) via llama.cpp before upload.\n\n## How to Use\n\n\n\n## Original Model\n\nSee [ManniX-ITA/gemma-4-A4B-109e-it](https://huggingface.co/ManniX-ITA/gemma-4-A4B-109e-it) for the full model card, pruning methodology, and benchmark results (71.7% GPQA Diamond).\n\n## License\n\n[Gemma license](https://ai.google.dev/gemma/terms)\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "imatrix",
    "quantized",
    "gemma4",
    "moe",
    "expert-pruning",
    "base_model:ManniX-ITA/gemma-4-A4B-109e-it",
    "base_model:quantized:ManniX-ITA/gemma-4-A4B-109e-it",
    "license:gemma",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 4,
  "downloads": 9866,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-05T19:30:14.000Z",
  "created_at": "2026-04-05T16:13:34.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69d28a2ef1a967ac148400f7",
  "id": "ManniX-ITA/gemma-4-A4B-109e-it-GGUF",
  "modelId": "ManniX-ITA/gemma-4-A4B-109e-it-GGUF",
  "sha": "136c4d065cb0ad9db817020dc3b1f6d0d6c3aaba",
  "createdAt": "2026-04-05T16:13:34.000Z",
  "lastModified": "2026-04-05T19:30:14.000Z",
  "author": "ManniX-ITA",
  "downloads": 9866,
  "likes": 4,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 32
}