AesSedai/GLM-5.3-Flash-GGUF overview
Notes WIP, requires this PR https://github.com/ggml org/llama.cpp/pull/27773 to run This repo contains specialized MoE quants for zai org/GLM 5.3 Flash BF16. T…
Runs locally from ~9.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| IQ2_S/GLM-5.3-Flash-IQ2_S-00001-of-00004.gguf | GGUF | IQ2_S | 9.0 MB | Download |
| IQ2_S/GLM-5.3-Flash-IQ2_S-00002-of-00004.gguf | GGUF | IQ2_S | 46.40 GB | Download |
| IQ2_S/GLM-5.3-Flash-IQ2_S-00003-of-00004.gguf | GGUF | IQ2_S | 46.08 GB | Download |
| IQ2_S/GLM-5.3-Flash-IQ2_S-00004-of-00004.gguf | GGUF | IQ2_S | 13.33 GB | Download |
| IQ3_S/GLM-5.3-Flash-IQ3_S-00001-of-00004.gguf | GGUF | IQ3_S | 9.0 MB | Download |
| IQ3_S/GLM-5.3-Flash-IQ3_S-00002-of-00004.gguf | GGUF | IQ3_S | 46.09 GB | Download |
| IQ3_S/GLM-5.3-Flash-IQ3_S-00003-of-00004.gguf | GGUF | IQ3_S | 45.89 GB | Download |
| IQ3_S/GLM-5.3-Flash-IQ3_S-00004-of-00004.gguf | GGUF | IQ3_S | 24.16 GB | Download |
| IQ4_XS/GLM-5.3-Flash-IQ4_XS-00001-of-00005.gguf | GGUF | IQ4_XS | 9.0 MB | Download |
| IQ4_XS/GLM-5.3-Flash-IQ4_XS-00002-of-00005.gguf | GGUF | IQ4_XS | 46.41 GB | Download |
| IQ4_XS/GLM-5.3-Flash-IQ4_XS-00003-of-00005.gguf | GGUF | IQ4_XS | 46.23 GB | Download |
| IQ4_XS/GLM-5.3-Flash-IQ4_XS-00004-of-00005.gguf | GGUF | IQ4_XS | 46.24 GB | Download |
| IQ4_XS/GLM-5.3-Flash-IQ4_XS-00005-of-00005.gguf | GGUF | IQ4_XS | 9.35 GB | Download |
| Q4_K_M/GLM-5.3-Flash-Q4_K_M-00001-of-00006.gguf | GGUF | Q4_K_M | 9.0 MB | Download |
| Q4_K_M/GLM-5.3-Flash-Q4_K_M-00002-of-00006.gguf | GGUF | Q4_K_M | 46.34 GB | Download |
| Q4_K_M/GLM-5.3-Flash-Q4_K_M-00003-of-00006.gguf | GGUF | Q4_K_M | 45.22 GB | Download |
| Q4_K_M/GLM-5.3-Flash-Q4_K_M-00004-of-00006.gguf | GGUF | Q4_K_M | 45.35 GB | Download |
| Q4_K_M/GLM-5.3-Flash-Q4_K_M-00005-of-00006.gguf | GGUF | Q4_K_M | 46.34 GB | Download |
| Q4_K_M/GLM-5.3-Flash-Q4_K_M-00006-of-00006.gguf | GGUF | Q4_K_M | 4.85 GB | Download |
| Q5_K_M/GLM-5.3-Flash-Q5_K_M-00001-of-00006.gguf | GGUF | Q5_K_M | 9.0 MB | Download |
| Q5_K_M/GLM-5.3-Flash-Q5_K_M-00002-of-00006.gguf | GGUF | Q5_K_M | 45.02 GB | Download |
| Q5_K_M/GLM-5.3-Flash-Q5_K_M-00003-of-00006.gguf | GGUF | Q5_K_M | 46.03 GB | Download |
| Q5_K_M/GLM-5.3-Flash-Q5_K_M-00004-of-00006.gguf | GGUF | Q5_K_M | 46.02 GB | Download |
| Q5_K_M/GLM-5.3-Flash-Q5_K_M-00005-of-00006.gguf | GGUF | Q5_K_M | 46.02 GB | Download |
| Q5_K_M/GLM-5.3-Flash-Q5_K_M-00006-of-00006.gguf | GGUF | Q5_K_M | 41.19 GB | Download |
| imatrix.gguf | GGUF | GGUF | 488.9 MB | Download |
| mmproj-GLM-5.3-Flash-BF16.gguf | GGUF | BF16 | 1.08 GB | Download |
| mmproj-GLM-5.3-Flash-F16.gguf | GGUF | F16 | 1.05 GB | Download |
| mmproj-GLM-5.3-Flash-F32.gguf | GGUF | F32 | 2.10 GB | Download |
| mmproj-GLM-5.3-Flash-Q8_0.gguf | GGUF | Q8_0 | 622.6 MB | Download |
Model Details
Model README
---
base_model:
- zai-org/GLM-5.3-Flash-BF16
---
Notes
- WIP, requires this PR to run
This repo contains specialized MoE-quants for zai-org/GLM-5.3-Flash-BF16. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.
| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |
| :----- | :-------------------- | :------------------------------- | :------------------ | :------------------------ | :------------------ |
| Q5_K_M | 224.28 GiB (6.01 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 3.589877 ± 0.019865 | +0.5529% | 0.027859 ± 0.000207 |
| Q4_K_M | 188.10 GiB (5.04 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 3.635356 ± 0.020204 | +1.8267% | 0.050181 ± 0.000333 |
| IQ4_XS | 148.24 GiB (3.97 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 3.819227 ± 0.021423 | +6.9770% | 0.117358 ± 0.000727 |
| IQ3_S | 116.14 GiB (3.11 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 4.387061 ± 0.025595 | +22.8821% | 0.283438 ± 0.001596 |
| IQ2_S | 105.81 GiB (2.83 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 4.761384 ± 0.028305 | +33.3669% | 0.375406 ± 0.001984 |
Run AesSedai/GLM-5.3-Flash-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models