GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF overview

GLM 5.3 Flash W4A16 NVFP4 GGUF GGUF conversion of axiomofmind/GLM 5.3 Flash W4A16 NVFP4 https://huggingface.co/axiomofmind/GLM 5.3 Flash W4A16 NVFP4 , based on…

ggufglmglm5nextnvfp4modeloptmoetext-generationenzhbase_model:axiomofmind/GLM-5.3-Flash-W4A16-NVFP4base_model:quantized:axiomofmind/GLM-5.3-Flash-W4A16-NVFP4license:mitendpoints_compatibleregion:usconversational

Runs locally from ~190.04 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
GLM-5.3-Flash-W4A16-NVFP4-BF16attn-MTP.ggufGGUFBF16190.04 GBDownload

Model Details

Model IDaxiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF
Authoraxiomofmind
Pipelinetext-generation
Licensemit
Base modelaxiomofmind/GLM-5.3-Flash-W4A16-NVFP4
Last modified2026-08-27T15:09:41.000Z

Model README

---

base_model: axiomofmind/GLM-5.3-Flash-W4A16-NVFP4

pipeline_tag: text-generation

license: mit

language:

- en

- zh

tags:

- gguf

- glm

- glm5next

- nvfp4

- modelopt

- moe

---

GLM-5.3-Flash W4A16 NVFP4 GGUF

GGUF conversion of

axiomofmind/GLM-5.3-Flash-W4A16-NVFP4,

based on zai-org/GLM-5.3-Flash-BF16.

The main-model routed experts use W4A16 NVFP4 with group size 16. Attention,

shared experts, routers, embeddings, the output head, MTP weights, and other

retained tensors remain in BF16 or F32.

Files

| File | Description | Size |

| --- | --- | ---: |

| GLM-5.3-Flash-W4A16-NVFP4-BF16attn-MTP.gguf | Text model with MTP weights | 204.1 GB |

Requirements

A llama.cpp build with GLM5Next and NVFP4 GGUF support is required.

License

This model is distributed under the

MIT License.

Refer to the official model card

for architecture details, usage guidance, and limitations.

Run axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models