ddh0/glm-4.5-iceblink-v2-106b-a12b-gguf - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
ddh0/glm-4.5-iceblink-v2-106b-a12b-gguf overview
This repository contains several custom GGUF quantizations of zerofata/GLM-4.5-Iceblink-v2-106B-A12B, to be used with llama.cpp. The naming scheme for these custom quantizations is as follows: ModelName-DefaultType-FFN-UpType-GateType-DownType.gguf Where DefaultType refers to the default tensor type, and UpType, GateType, and DownType refer to the tensor types used for the ffnupexps, ffngateexps, and ffndownexps tensors respectively.
Downloads
230
Likes
8
Pipeline
—
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
9 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ3_S-IQ4_NL.gguf | GGUF | IQ4_XS | 56.76 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ4_XS-IQ4_NL.gguf | GGUF | IQ4_XS | 59.98 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0.gguf | GGUF | IQ4_XS | 63.93 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q4_K-Q4_K-Q8_0.gguf | GGUF | Q4_K | 77.76 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q5_K-Q5_K-Q8_0.gguf | GGUF | Q5_K | 85.67 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q6_K-Q6_K-Q8_0.gguf | GGUF | Q6_K | 94.07 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0.gguf | GGUF | — | 109.39 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-bf16.gguf | GGUF | BF16 | 205.82 GB | Download |
| GLM-4.5-Iceblink-v2-106B-A12B-ddh0_v2-imatrix.gguf | GGUF | — | 217.81 MB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": [
"zerofata/GLM-4.5-Iceblink-v2-106B-A12B"
],
"base_model_relation": "quantized",
"quantized_by": "ddh0",
"license": "mit",
"frontmatter": {
"base_model": [
"zerofata/GLM-4.5-Iceblink-v2-106B-A12B"
],
"base_model_relation": "quantized",
"quantized_by": "ddh0",
"license": "mit"
},
"hero_image_url": "",
"summary": "This repository contains several custom GGUF quantizations of zerofata/GLM-4.5-Iceblink-v2-106B-A12B, to be used with llama.cpp. The naming scheme for these custom quantizations is as follows: > **ModelName-DefaultType-FFN-UpType-GateType-DownType.gguf** Where DefaultType refers to the default tensor type, and UpType, GateType, and DownType refer to the tensor types used for the ffn_up_exps, ffn_gate_exps, and ffn_down_exps tensors respectively.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model:\n- zerofata/GLM-4.5-Iceblink-v2-106B-A12B\nbase_model_relation: quantized\nquantized_by: ddh0\nlicense: mit\n---\n\n# GLM-4.5-Iceblink-v2-106B-A12B-106B-A12B-GGUF\n\nThis repository contains several custom GGUF quantizations of [zerofata/GLM-4.5-Iceblink-v2-106B-A12B](https://huggingface.co/zerofata/GLM-4.5-Iceblink-v2-106B-A12B), to be used with [llama.cpp](https://github.com/ggml-org/llama.cpp).\n\nThe naming scheme for these custom quantizations is as follows:\n\n> **`ModelName-DefaultType-FFN-UpType-GateType-DownType.gguf`**\n\nWhere `DefaultType` refers to the default tensor type, and `UpType`, `GateType`, and `DownType` refer to the tensor types used for the `ffn_up_exps`, `ffn_gate_exps`, and `ffn_down_exps` tensors respectively.\n\n## Quantizations\n\nThese quantizations use Q8_0 for all tensors by default, including the dense FFN block. Only the conditional experts are downgraded. The shared expert is always kept in Q8_0. They were quantized using [my own imatrix](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/blob/main/GLM-4.5-Iceblink-v2-106B-A12B-ddh0_v2-imatrix.gguf) (the calibration text corpus can be found [here](https://huggingface.co/ddh0/imatrices/blob/main/ddh0_imat_calibration_data_v2.txt)).\n\n| Filename | Size (GB) | Size (GiB) | Average BPW | Direct link |\n| ---------------------------------------------- | --------- | ---------- | ----------- | -------------------------------------------------------------------------------------------------------------------- |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ3_S-IQ4_NL.gguf | 60.94 | 56.76 | 4.41 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ3_S-IQ4_NL.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ4_XS-IQ4_NL.gguf | 64.39 | 59.97 | 4.66 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ4_XS-IQ4_NL.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0.gguf | 68.63 | 63.92 | 4.97 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-IQ4_XS-IQ4_XS-Q5_0.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q4_K-Q4_K-Q8_0.gguf | 83.49 | 77.76 | 6.05 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q4_K-Q4_K-Q8_0.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q5_K-Q5_K-Q8_0.gguf | 91.97 | 85.66 | 6.66 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q5_K-Q5_K-Q8_0.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q6_K-Q6_K-Q8_0.gguf | 100.99 | 94.06 | 7.31 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0-FFN-Q6_K-Q6_K-Q8_0.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-Q8_0.gguf | 117.45 | 109.38 | 8.51 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-Q8_0.gguf) |\n| GLM-4.5-Iceblink-v2-106B-A12B-bf16.gguf | 220.98 | 205.81 | 16.00 | [Download](https://huggingface.co/ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF/resolve/main/GLM-4.5-Iceblink-v2-106B-A12B-bf16.gguf) |",
"related_quantizations": []
},
"tags": [
"gguf",
"base_model:zerofata/GLM-4.5-Iceblink-v2-106B-A12B",
"base_model:quantized:zerofata/GLM-4.5-Iceblink-v2-106B-A12B",
"license:mit",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 8,
"downloads": 230,
"gated": false,
"private": false,
"last_modified": "2025-11-03T23:20:07.000Z",
"created_at": "2025-11-03T03:40:28.000Z",
"pipeline_tag": "",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "6908242c35a8039ffc433472",
"id": "ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF",
"modelId": "ddh0/GLM-4.5-Iceblink-v2-106B-A12B-GGUF",
"sha": "34ba11e5d84944d1c23e27d9c11928f7d5f37c72",
"createdAt": "2025-11-03T03:40:28.000Z",
"lastModified": "2025-11-03T23:20:07.000Z",
"author": "ddh0",
"downloads": 230,
"likes": 8,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "",
"siblings_count": 11
}