GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

0ppxnhximxr/GLM-5.3-Flash-GGUF overview

GLM 5.3 Flash GGUF GGUF quantizations of zai org/GLM 5.3 Flash BF16 https://huggingface.co/zai org/GLM 5.3 Flash BF16 , converted directly from the official BF…

ggufllama.cppglm5_nextimatrixtext-generationenzhbase_model:zai-org/GLM-5.3-Flash-BF16base_model:quantized:zai-org/GLM-5.3-Flash-BF16license:mitendpoints_compatibleregion:usconversational

Runs locally from ~4.83 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

48 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00001-of-00016.ggufGGUFIQ2_XS5.65 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00002-of-00016.ggufGGUFIQ2_XS5.40 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00003-of-00016.ggufGGUFIQ2_XS5.33 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00004-of-00016.ggufGGUFIQ2_XS4.83 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00005-of-00016.ggufGGUFIQ2_XS5.33 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00006-of-00016.ggufGGUFIQ2_XS5.39 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00007-of-00016.ggufGGUFIQ2_XS5.39 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00008-of-00016.ggufGGUFIQ2_XS5.33 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00009-of-00016.ggufGGUFIQ2_XS5.38 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00010-of-00016.ggufGGUFIQ2_XS5.40 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00011-of-00016.ggufGGUFIQ2_XS5.33 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00012-of-00016.ggufGGUFIQ2_XS5.39 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00013-of-00016.ggufGGUFIQ2_XS5.40 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00014-of-00016.ggufGGUFIQ2_XS5.33 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00015-of-00016.ggufGGUFIQ2_XS5.40 GBDownload
IQ2_XS/GLM-5.3-Flash-IQ2_XS-00016-of-00016.ggufGGUFIQ2_XS5.39 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00001-of-00016.ggufGGUFIQ3_XXS7.06 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00002-of-00016.ggufGGUFIQ3_XXS7.09 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00003-of-00016.ggufGGUFIQ3_XXS7.03 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00004-of-00016.ggufGGUFIQ3_XXS6.34 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00005-of-00016.ggufGGUFIQ3_XXS7.02 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00006-of-00016.ggufGGUFIQ3_XXS7.10 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00007-of-00016.ggufGGUFIQ3_XXS7.10 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00008-of-00016.ggufGGUFIQ3_XXS7.02 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00009-of-00016.ggufGGUFIQ3_XXS7.11 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00010-of-00016.ggufGGUFIQ3_XXS7.09 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00011-of-00016.ggufGGUFIQ3_XXS7.03 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00012-of-00016.ggufGGUFIQ3_XXS7.10 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00013-of-00016.ggufGGUFIQ3_XXS7.09 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00014-of-00016.ggufGGUFIQ3_XXS7.03 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00015-of-00016.ggufGGUFIQ3_XXS7.09 GBDownload
IQ3_XXS/GLM-5.3-Flash-IQ3_XXS-00016-of-00016.ggufGGUFIQ3_XXS7.10 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00001-of-00016.ggufGGUFQ4_K_M10.82 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00002-of-00016.ggufGGUFQ4_K_M11.00 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00003-of-00016.ggufGGUFQ4_K_M10.91 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00004-of-00016.ggufGGUFQ4_K_M9.91 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00005-of-00016.ggufGGUFQ4_K_M10.90 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00006-of-00016.ggufGGUFQ4_K_M11.00 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00007-of-00016.ggufGGUFQ4_K_M10.42 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00008-of-00016.ggufGGUFQ4_K_M11.48 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00009-of-00016.ggufGGUFQ4_K_M11.00 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00010-of-00016.ggufGGUFQ4_K_M11.00 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00011-of-00016.ggufGGUFQ4_K_M10.90 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00012-of-00016.ggufGGUFQ4_K_M11.58 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00013-of-00016.ggufGGUFQ4_K_M11.59 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00014-of-00016.ggufGGUFQ4_K_M12.06 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00015-of-00016.ggufGGUFQ4_K_M11.02 GBDownload
Q4_K_M/GLM-5.3-Flash-Q4_K_M-00016-of-00016.ggufGGUFQ4_K_M10.43 GBDownload

Model Details

Model ID0ppxnhximxr/GLM-5.3-Flash-GGUF
Author0ppxnhximxr
Pipelinetext-generation
Licensemit
Base modelzai-org/GLM-5.3-Flash-BF16
Last modified2026-08-28T17:58:18.000Z

Model README

---

base_model: zai-org/GLM-5.3-Flash-BF16

base_model_relation: quantized

library_name: gguf

license: mit

pipeline_tag: text-generation

language:

- en

- zh

tags:

- gguf

- llama.cpp

- glm5_next

- imatrix

---

GLM-5.3-Flash GGUF

GGUF quantizations of zai-org/GLM-5.3-Flash-BF16, converted directly from the official BF16 weights.

Quantizations

| Quantization | Notes |

|---|---|

| Q4_K_M | Recommended quality/size balance; importance-matrix calibrated |

| IQ3_XXS | Mid-size 3.08 BPW build; importance-matrix calibrated |

| IQ2_XS | Compact 2.35 BPW build; importance-matrix calibrated |

The GLM-5.3-sensitive router, shared-expert, mHC, and linear-attention state tensors are preserved at higher precision by the GLM-5.3 llama.cpp quantization rules. The unsupported NextN/MTP draft head is excluded.

Usage

Use a recent GLM-5.3-compatible llama.cpp build and point to the first shard:

llama-cli -m Q4_K_M/GLM-5.3-Flash-Q4_K_M-00001-of-00016.gguf

All shards in the selected quantization directory are required.

License

MIT, inherited from the original model.

Run 0ppxnhximxr/GLM-5.3-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models