axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF overview
GLM 5.3 Flash W4A16 NVFP4 GGUF GGUF conversion of axiomofmind/GLM 5.3 Flash W4A16 NVFP4 https://huggingface.co/axiomofmind/GLM 5.3 Flash W4A16 NVFP4 , based on…
Runs locally from ~190.04 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.3-Flash-W4A16-NVFP4-BF16attn-MTP.gguf | GGUF | BF16 | 190.04 GB | Download |
Model Details
| Model ID | axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF |
|---|---|
| Author | axiomofmind |
| Pipeline | text-generation |
| License | mit |
| Base model | axiomofmind/GLM-5.3-Flash-W4A16-NVFP4 |
| Last modified | 2026-08-27T15:09:41.000Z |
Model README
---
base_model: axiomofmind/GLM-5.3-Flash-W4A16-NVFP4
pipeline_tag: text-generation
license: mit
language:
- en
- zh
tags:
- gguf
- glm
- glm5next
- nvfp4
- modelopt
- moe
---
GLM-5.3-Flash W4A16 NVFP4 GGUF
GGUF conversion of
axiomofmind/GLM-5.3-Flash-W4A16-NVFP4,
based on zai-org/GLM-5.3-Flash-BF16.
The main-model routed experts use W4A16 NVFP4 with group size 16. Attention,
shared experts, routers, embeddings, the output head, MTP weights, and other
retained tensors remain in BF16 or F32.
Files
| File | Description | Size |
| --- | --- | ---: |
| GLM-5.3-Flash-W4A16-NVFP4-BF16attn-MTP.gguf | Text model with MTP weights | 204.1 GB |
Requirements
A llama.cpp build with GLM5Next and NVFP4 GGUF support is required.
License
This model is distributed under the
Refer to the official model card
for architecture details, usage guidance, and limitations.
Run axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models