ddh0/glm-4.5-air-gguf imatrix GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
ddh0/glm-4.5-air-gguf overview
This repository contains several custom GGUF quantizations of GLM-4.5-Air, to be used with llama.cpp. The naming scheme for these custom quantizations is as follows: ModelName-DefaultType-FFN-UpType-GateType-DownType.gguf Where DefaultType refers to the default tensor type, and UpType, GateType, and DownType refer to the tensor types used for the ffnupexps, ffngateexps, and ffndownexps tensors respectively.
Downloads
481
Likes
18
Pipeline
—
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
15 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-4.5-Air-Q8_0-FFN-IQ3_S-IQ3_S-Q5_0.gguf | GGUF | IQ3_S | 57.44 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ3_S-IQ4_NL-v2.gguf | GGUF | IQ4_XS | 56.76 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-IQ4_NL-v2.gguf | GGUF | IQ4_XS | 59.98 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0-v2.gguf | GGUF | IQ4_XS | 63.93 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0.gguf | GGUF | IQ4_XS | 63.87 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q8_0-v2.gguf | GGUF | IQ4_XS | 75.79 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-Q4_K-Q4_K-Q5_1.gguf | GGUF | Q4_K | 67.83 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-Q4_K-Q4_K-Q8_0.gguf | GGUF | Q4_K | 77.72 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-Q5_K-Q5_K-Q8_0-v2.gguf | GGUF | Q5_K | 85.67 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-Q5_K-Q5_K-Q8_0.gguf | GGUF | Q5_K | 85.64 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-Q6_K-Q6_K-Q8_0-v2.gguf | GGUF | Q6_K | 94.07 GB | Download |
| GLM-4.5-Air-Q8_0-FFN-Q6_K-Q6_K-Q8_0.gguf | GGUF | Q6_K | 94.05 GB | Download |
| GLM-4.5-Air-Q8_0.gguf | GGUF | — | 109.39 GB | Download |
| GLM-4.5-Air-bf16.gguf | GGUF | BF16 | 205.82 GB | Download |
| GLM-4.5-Air-ddh0_v2-imatrix.gguf | GGUF | — | 217.81 MB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": [
"zai-org/GLM-4.5-Air"
],
"base_model_relation": "quantized",
"quantized_by": "ddh0",
"license": "mit",
"frontmatter": {
"base_model": [
"zai-org/GLM-4.5-Air"
],
"base_model_relation": "quantized",
"quantized_by": "ddh0",
"license": "mit"
},
"hero_image_url": "",
"summary": "This repository contains several custom GGUF quantizations of GLM-4.5-Air, to be used with llama.cpp. The naming scheme for these custom quantizations is as follows: > **ModelName-DefaultType-FFN-UpType-GateType-DownType.gguf** Where DefaultType refers to the default tensor type, and UpType, GateType, and DownType refer to the tensor types used for the ffn_up_exps, ffn_gate_exps, and ffn_down_exps tensors respectively.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model:\n- zai-org/GLM-4.5-Air\nbase_model_relation: quantized\nquantized_by: ddh0\nlicense: mit\n---\n\n# GLM-4.5-Air-GGUF\n\nThis repository contains several custom GGUF quantizations of [GLM-4.5-Air](https://huggingface.co/zai-org/GLM-4.5-Air), to be used with [llama.cpp](https://github.com/ggml-org/llama.cpp).\n\nThe naming scheme for these custom quantizations is as follows:\n\n> **`ModelName-DefaultType-FFN-UpType-GateType-DownType.gguf`**\n\nWhere `DefaultType` refers to the default tensor type, and `UpType`, `GateType`, and `DownType` refer to the tensor types used for the `ffn_up_exps`, `ffn_gate_exps`, and `ffn_down_exps` tensors respectively.\n\n## Original quantizations\n\nThese quantizations use Q8_0 for all tensors by default - only the dense FFN block and conditional experts are downgraded. The shared expert is always kept in Q8_0. They were quantized using [bartowski's imatrix](https://huggingface.co/bartowski/zai-org_GLM-4.5-Air-GGUF/blob/main/zai-org_GLM-4.5-Air-imatrix.gguf).\n\n| Filename | Size (GB) | Size (GiB) | Average BPW | Direct link |\n| -------------------------------------------- | --------- | ---------- | ----------- | ------------------------------------------------------------------------------------------------------------------ |\n| GLM-4.5-Air-Q8_0-FFN-IQ3_S-IQ3_S-Q5_0.gguf | 61.66 | 57.43 | 4.47 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-IQ3_S-IQ3_S-Q5_0.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0.gguf | 68.56 | 63.86 | 4.97 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-Q4_K-Q4_K-Q5_1.gguf | 72.82 | 67.82 | 5.27 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-Q4_K-Q4_K-Q5_1.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-Q4_K-Q4_K-Q8_0.gguf | 83.44 | 77.71 | 6.04 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-Q4_K-Q4_K-Q8_0.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-Q5_K-Q5_K-Q8_0.gguf | 91.94 | 85.63 | 6.66 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-Q5_K-Q5_K-Q8_0.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-Q6_K-Q6_K-Q8_0.gguf | 100.97 | 94.04 | 7.31 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-Q6_K-Q6_K-Q8_0.gguf) |\n| GLM-4.5-Air-Q8_0.gguf | 117.45 | 109.39 | 8.50 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0.gguf) |\n| GLM-4.5-Air-bf16.gguf | 220.98 | 205.81 | 16.00 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-bf16.gguf) |\n\n## v2 quantizations\n\nThese quantizations use Q8_0 for all tensors by default, including the dense FFN block. Only the conditional experts are downgraded. The shared expert is always kept in Q8_0. They were quantized using [my own imatrix](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/blob/main/GLM-4.5-Air-ddh0_v2-imatrix.gguf) (the calibration text corpus can be found [here](https://huggingface.co/ddh0/imatrices/blob/main/ddh0_imat_calibration_data_v2.txt)).\n\n| Filename | Size (GB) | Size (GiB) | Average BPW | Direct link |\n| ------------------------------------------------- | --------- | ---------- | ----------- | ----------------------------------------------------------------------------------------------------------------------- |\n| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ3_S-IQ4_NL-v2.gguf | 60.94 | 56.76 | 4.41 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ3_S-IQ4_NL-v2.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-IQ4_NL-v2.gguf | 64.39 | 59.97 | 4.66 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-IQ4_NL-v2.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0-v2.gguf | 68.63 | 63.92 | 4.97 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0-v2.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q8_0-v2.gguf | 81.36 | 75.78 | 5.89 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-IQ4_XS-IQ4_XS-Q8_0-v2.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-Q5_K-Q5_K-Q8_0-v2.gguf | 91.97 | 85.66 | 6.66 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-Q5_K-Q5_K-Q8_0-v2.gguf) |\n| GLM-4.5-Air-Q8_0-FFN-Q6_K-Q6_K-Q8_0-v2.gguf | 100.99 | 94.06 | 7.31 | [Download](https://huggingface.co/ddh0/GLM-4.5-Air-GGUF/resolve/main/GLM-4.5-Air-Q8_0-FFN-Q6_K-Q6_K-Q8_0-v2.gguf) |\n",
"related_quantizations": []
},
"tags": [
"gguf",
"base_model:zai-org/GLM-4.5-Air",
"base_model:quantized:zai-org/GLM-4.5-Air",
"license:mit",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 18,
"downloads": 481,
"gated": false,
"private": false,
"last_modified": "2025-11-03T03:43:29.000Z",
"created_at": "2025-08-04T20:25:03.000Z",
"pipeline_tag": "",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "6891171f6dbfb56d82d42aa8",
"id": "ddh0/GLM-4.5-Air-GGUF",
"modelId": "ddh0/GLM-4.5-Air-GGUF",
"sha": "f6c903c5fc5676d502c16826f876b976da4b6e2c",
"createdAt": "2025-08-04T20:25:03.000Z",
"lastModified": "2025-11-03T03:43:29.000Z",
"author": "ddh0",
"downloads": 481,
"likes": 18,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "",
"siblings_count": 17
}