GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AesSedai/GLM-5.3-GGUF overview

This repo contains specialized MoE quants for zai org/GLM 5.3 BF16. The idea being that given the huge size of the FFN tensors compared to the rest of the tens…

ggufbase_model:zai-org/GLM-5.3-BF16base_model:quantized:zai-org/GLM-5.3-BF16endpoints_compatibleregion:usimatrixconversational

Runs locally from ~9.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
1,111
Likes
7
Pipeline
Author

Repository Files & Downloads

48 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
IQ2_S/GLM-5.3-IQ2_S-00001-of-00007.ggufGGUFIQ2_S9.0 MBDownload
IQ2_S/GLM-5.3-IQ2_S-00002-of-00007.ggufGGUFIQ2_S46.48 GBDownload
IQ2_S/GLM-5.3-IQ2_S-00003-of-00007.ggufGGUFIQ2_S45.82 GBDownload
IQ2_S/GLM-5.3-IQ2_S-00004-of-00007.ggufGGUFIQ2_S45.82 GBDownload
IQ2_S/GLM-5.3-IQ2_S-00005-of-00007.ggufGGUFIQ2_S45.78 GBDownload
IQ2_S/GLM-5.3-IQ2_S-00006-of-00007.ggufGGUFIQ2_S45.82 GBDownload
IQ2_S/GLM-5.3-IQ2_S-00007-of-00007.ggufGGUFIQ2_S11.61 GBDownload
IQ3_S/GLM-5.3-IQ3_S-00001-of-00007.ggufGGUFIQ3_S9.0 MBDownload
IQ3_S/GLM-5.3-IQ3_S-00002-of-00007.ggufGGUFIQ3_S46.54 GBDownload
IQ3_S/GLM-5.3-IQ3_S-00003-of-00007.ggufGGUFIQ3_S46.23 GBDownload
IQ3_S/GLM-5.3-IQ3_S-00004-of-00007.ggufGGUFIQ3_S46.39 GBDownload
IQ3_S/GLM-5.3-IQ3_S-00005-of-00007.ggufGGUFIQ3_S46.03 GBDownload
IQ3_S/GLM-5.3-IQ3_S-00006-of-00007.ggufGGUFIQ3_S46.23 GBDownload
IQ3_S/GLM-5.3-IQ3_S-00007-of-00007.ggufGGUFIQ3_S34.51 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00001-of-00009.ggufGGUFIQ4_XS9.0 MBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00002-of-00009.ggufGGUFIQ4_XS45.70 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00003-of-00009.ggufGGUFIQ4_XS45.35 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00004-of-00009.ggufGGUFIQ4_XS45.46 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00005-of-00009.ggufGGUFIQ4_XS46.51 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00006-of-00009.ggufGGUFIQ4_XS45.61 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00007-of-00009.ggufGGUFIQ4_XS46.51 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00008-of-00009.ggufGGUFIQ4_XS45.65 GBDownload
IQ4_XS/GLM-5.3-IQ4_XS-00009-of-00009.ggufGGUFIQ4_XS21.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00001-of-00011.ggufGGUFQ4_K_M9.0 MBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00002-of-00011.ggufGGUFQ4_K_M44.93 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00003-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00004-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00005-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00006-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00007-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00008-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00009-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00010-of-00011.ggufGGUFQ4_K_M45.22 GBDownload
Q4_K_M/GLM-5.3-Q4_K_M-00011-of-00011.ggufGGUFQ4_K_M30.23 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00001-of-00013.ggufGGUFQ5_K_M9.0 MBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00002-of-00013.ggufGGUFQ5_K_M46.56 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00003-of-00013.ggufGGUFQ5_K_M45.16 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00004-of-00013.ggufGGUFQ5_K_M45.34 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00005-of-00013.ggufGGUFQ5_K_M45.54 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00006-of-00013.ggufGGUFQ5_K_M45.14 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00007-of-00013.ggufGGUFQ5_K_M45.34 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00008-of-00013.ggufGGUFQ5_K_M45.54 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00009-of-00013.ggufGGUFQ5_K_M45.14 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00010-of-00013.ggufGGUFQ5_K_M45.34 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00011-of-00013.ggufGGUFQ5_K_M45.54 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00012-of-00013.ggufGGUFQ5_K_M45.14 GBDownload
Q5_K_M/GLM-5.3-Q5_K_M-00013-of-00013.ggufGGUFQ5_K_M23.27 GBDownload
imatrix.ggufGGUFGGUF1.05 GBDownload

Model Details

Model IDAesSedai/GLM-5.3-GGUF
AuthorAesSedai
Pipeline
License
Base modelzai-org/GLM-5.3-BF16
Last modified2026-09-03T16:27:08.000Z

Model README

---

base_model:

  • zai-org/GLM-5.3-BF16

---

This repo contains specialized MoE-quants for zai-org/GLM-5.3-BF16. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |

| :----- | :-------------------- | :------------------------------- | :------------------ | :------------------------ | :------------------ |

| Q8_0 | 745.77 GiB (8.50 BPW) | Q8_0 | 2.674779 ± 0.013836 | +0.0174% | 0.013080 ± 0.000123 |

| Q5_K_M | 523.07 GiB (5.96 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 2.680808 ± 0.013885 | +0.2428% | 0.020301 ± 0.000174 |

| Q4_K_M | 436.94 GiB (4.98 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 2.699454 ± 0.013914 | +0.9400% | 0.037094 ± 0.000294 |

| IQ4_XS | 342.01 GiB (3.90 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 2.841042 ± 0.014851 | +6.2344% | 0.101583 ± 0.000741 |

| IQ3_S | 265.93 GiB (3.03 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 3.257199 ± 0.017856 | +21.7956% | 0.275247 ± 0.001717 |

| IQ2_S | 241.32 GiB (2.75 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 3.512105 ± 0.019532 | +31.3273% | 0.374183 ± 0.002137 |

!kld_graph

!ppl_graph

Run AesSedai/GLM-5.3-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models