groxaxo/qwen3.5-9b-glm5.1-distill-v1-gguf Q5_K_S GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
groxaxo/qwen3.5-9b-glm5.1-distill-v1-gguf overview
All quantizations were produced using an importance matrix (imatrix) computed on 128 chunks of WikiText-2-raw-v1 for optimal quality at every bit-width.
Downloads
1,553
Likes
1
Pipeline
—
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
13 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.5-9B-GLM5.1-Distill-v1-F16.gguf | GGUF | F16 | 16.69 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-IQ3_M.gguf | GGUF | IQ3_M | 4.11 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-IQ3_XXS.gguf | GGUF | IQ3_XXS | 3.67 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-IQ4_XS.gguf | GGUF | IQ4_XS | 4.84 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q2_K.gguf | GGUF | Q2_K | 3.56 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q3_K_M.gguf | GGUF | Q3_K_M | 4.31 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q3_K_S.gguf | GGUF | Q3_K_S | 3.97 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q4_K_M.gguf | GGUF | Q4_K_M | 5.24 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q4_K_S.gguf | GGUF | Q4_K_S | 4.98 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q5_K_M.gguf | GGUF | Q5_K_M | 6.02 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q5_K_S.gguf | GGUF | Q5_K_S | 5.87 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q6_K.gguf | GGUF | Q6_K | 6.85 GB | Download |
| Qwen3.5-9B-GLM5.1-Distill-v1-Q8_0.gguf | GGUF | — | 8.87 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": "Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1",
"license": "apache-2.0",
"language": [
"en"
],
"tags": [
"gguf",
"quantized",
"qwen3.5",
"glm5.1",
"llama.cpp",
"imatrix"
],
"quantized_by": "groxaxo",
"inference": false,
"frontmatter": {
"base_model": "Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1",
"license": "apache-2.0",
"language": [
"en"
],
"tags": [
"gguf",
"quantized",
"qwen3.5",
"glm5.1",
"llama.cpp",
"imatrix"
],
"quantized_by": "groxaxo",
"inference": "false"
},
"hero_image_url": "",
"summary": "All quantizations were produced using an **importance matrix (imatrix)** computed on 128 chunks of WikiText-2-raw-v1 for optimal quality at every bit-width.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model: Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1\nlicense: apache-2.0\nlanguage:\n- en\ntags:\n- gguf\n- quantized\n- qwen3.5\n- glm5.1\n- llama.cpp\n- imatrix\nquantized_by: groxaxo\ninference: false\n---\n\n# Qwen3.5-9B-GLM5.1-Distill-v1 — GGUF Quantized\n\nAll quantizations were produced using an **importance matrix (imatrix)** computed on 128 chunks of WikiText-2-raw-v1 for optimal quality at every bit-width.\n\n## Quantization Variants\n\n| Variant | Size | BPW | Notes |\n|---------|------|-----|-------|\n| **F16** | 17.0 GB | 16.00 | Lossless half-precision |\n| **Q8_0** | 8.9 GB | 8.50 | Near-lossless, 8-bit round quant |\n| **Q6_K** | 6.9 GB | 6.57 | 6-bit K-quant, excellent quality |\n| **Q5_K_M** | 6.1 GB | 5.77 | 5-bit K-medium, recommended sweet spot |\n| **Q5_K_S** | 5.9 GB | 5.62 | 5-bit K-small, slight size savings |\n| **Q4_K_M** | 5.3 GB | 5.02 | 4-bit K-medium, best quality/size ratio |\n| **Q4_K_S** | 5.0 GB | 4.77 | 4-bit K-small, good balance |\n| **IQ4_XS** | 4.9 GB | 4.63 | 4-bit importance matrix, extra small |\n| **Q3_K_M** | 4.4 GB | 4.12 | 3-bit K-medium |\n| **IQ3_M** | 4.2 GB | 3.94 | 3-bit importance matrix, medium |\n| **Q3_K_S** | 4.0 GB | 3.80 | 3-bit K-small |\n| **IQ3_XXS** | 3.7 GB | 3.51 | 3-bit importance matrix, extra-extra small |\n| **Q2_K** | 3.6 GB | 3.41 | 2-bit K-quant, smallest size |\n\n## Recommendations\n\n- **Best overall:** `Q4_K_M` or `Q5_K_M` — excellent quality-to-size ratio\n- **Maximum quality:** `Q6_K` or `Q8_0`\n- **Tight VRAM:** `IQ4_XS` or `Q3_K_M`\n- **Minimum size:** `Q2_K`\n\n## Usage\n\nCompatible with [llama.cpp](https://github.com/ggml-org/llama.cpp), [LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.com/), [text-generation-webui](https://github.com/oobabooga/text-generation-webui), and any GGUF-compatible inference engine.\n\n```bash\nllama-cli -m Qwen3.5-9B-GLM5.1-Distill-v1-Q4_K_M.gguf -p \"Hello, world!\"\n```\n\n## Source Model\n\n[Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1](https://huggingface.co/Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1) — a Qwen3.5-9B variant distilled with GLM5.1.\n",
"related_quantizations": []
},
"tags": [
"gguf",
"quantized",
"qwen3.5",
"glm5.1",
"llama.cpp",
"imatrix",
"en",
"base_model:Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1",
"base_model:quantized:Jackrong/Qwen3.5-9B-GLM5.1-Distill-v1",
"license:apache-2.0",
"region:us",
"conversational"
],
"likes": 1,
"downloads": 1553,
"gated": false,
"private": false,
"last_modified": "2026-04-16T00:59:47.000Z",
"created_at": "2026-04-16T00:19:01.000Z",
"pipeline_tag": "",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69e02af5416a0c2d850acdde",
"id": "groxaxo/Qwen3.5-9B-GLM5.1-Distill-v1-GGUF",
"modelId": "groxaxo/Qwen3.5-9B-GLM5.1-Distill-v1-GGUF",
"sha": "c469f115681e3f9a3b9cb21a7c9cabee81f895c1",
"createdAt": "2026-04-16T00:19:01.000Z",
"lastModified": "2026-04-16T00:59:47.000Z",
"author": "groxaxo",
"downloads": 1553,
"likes": 1,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "",
"siblings_count": 16
}