emwesoft/GLM-5.3-NVFP4-MTP-GGUF overview
GLM 5.3 753B — NVFP4 GGUF with MTP head Native GGUF of incoai/GLM 5.3 NVFP4 https://huggingface.co/incoai/GLM 5.3 NVFP4 , the vendor's ModelOpt NVFP4 repack of…
Runs locally from ~2.44 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.3-NVFP4-MTP-00001-of-00011.gguf | GGUF | GGUF | 42.20 GB | Download |
| GLM-5.3-NVFP4-MTP-00002-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00003-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00004-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00005-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00006-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00007-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00008-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00009-of-00011.gguf | GGUF | GGUF | 42.19 GB | Download |
| GLM-5.3-NVFP4-MTP-00010-of-00011.gguf | GGUF | GGUF | 40.25 GB | Download |
| GLM-5.3-NVFP4-MTP-00011-of-00011.gguf | GGUF | GGUF | 13.17 GB | Download |
| dflash/GLM-5.3-DFlash2-BF16.gguf | GGUF | BF16 | 4.59 GB | Download |
| dflash/GLM-5.3-DFlash2-Q8_0.gguf | GGUF | Q8_0 | 2.44 GB | Download |
Model Details
| Model ID | emwesoft/GLM-5.3-NVFP4-MTP-GGUF |
|---|---|
| Author | emwesoft |
| Pipeline | text-generation |
| License | mit |
| Base model | zai-org/GLM-5.3,incoai/GLM-5.3-NVFP4 |
| Last modified | 2026-08-29T10:42:19.000Z |
Model README
---
license: mit
base_model: [zai-org/GLM-5.3, incoai/GLM-5.3-NVFP4]
pipeline_tag: text-generation
library_name: gguf
tags: [gguf, nvfp4, glm, glm-dsa, mtp, speculative-decoding, llama.cpp]
---
GLM-5.3 753B — NVFP4 GGUF with MTP head
Native GGUF of incoai/GLM-5.3-NVFP4, the
vendor's ModelOpt NVFP4 repack of zai-org/GLM-5.3,
including block 78 (the MTP / NextN head) so the checkpoint's own head can drive speculative
decoding. The NVFP4 trunk is kept as NVFP4 — not a requantisation.
465 GB, 11 shards. block_count 79, nextn_predict_layers 1, 1974 tensors,
225 NVFP4 tensors (identical to the no-MTP file — the trunk is not upcast).
Using the MTP head — read this
blk.78's experts are BF16, not NVFP4: 27 tensors, 18.54 GiB, roughly 3.7x a normal
NVFP4 routed bank (5.06 GiB). If you use -ot with a CPU catch-all, **pin block 78 to a GPU
before the catch-all** — -ot is first-match-wins, so otherwise blk.78 lands on CPU and every
drafted token pays a CPU MoE pass, which is worse than not speculating.
-ot 'blk\.(...|78)\.ffn_.*_exps.*=CUDA1,\.ffn_.*_exps.*=CPU'
--spec-type draft-mtp --spec-draft-n-max 3
Related repos
| | |
|---|---|
| Without MTP | emwesoft/GLM-5.3-NVFP4-GGUF — 20 GB smaller, no block 78 |
| DFlash2 drafters | emwesoft/GLM-5.3-DFlash2-GGUF — the alternative to the MTP head |
Engine requirements
llama.cpp with glm-dsa + GGML_TYPE_NVFP4. Three fixes are not yet upstream:
- jinja numeric attribute access (
obj.0) — GLM-5.3's chat template uses
m.content.0.output. Without it the template throws, caps_get() swallows it,
supports_tool_calls reports false, and every tool call comes back as plain text.
glm-dsalayer-input exposure — DFlash needsres->t_layer_inp[il]; without it
attaching a drafter aborts on the first decode with GGML_ASSERT(t_layer_inp[il] != nullptr).
- Whitespace tolerance before
</tool_call>— a stray newline makes the streaming
parser recognise a tool call then lose it, aborting from compute_diffs.
Measured throughput
2x RTX PRO 6000 Blackwell + 4x RTX 3090 + 251 GB RAM, 400K context, -t 36 -tb 40,
experts partly CPU-resident (the weights do not fit in 288 GB of VRAM):
| config | acceptance | decode |
|---|---|---|
| MTP head, n-max 3 | 74.9% (mean len 3.24) | 9.3-12.2 tok/s |
| DFlash2 Q8_0, n-max 4 | 67.7% (mean len 3.69) | 7.6-12.8 tok/s |
| no speculation | - | ~10 tok/s |
Throughput is prompt-dependent because acceptance is. Threads matter: on a 24-core/48-thread
CPU, -t 48 collapsed decode to 0.5 tok/s — the ggml threadpool busy-spins and starves the CUDA
submission thread. Leave headroom.
Sampling
From generation_config.json: temperature 1.0, top_p 0.95. The template exposes
low/high/max reasoning effort only; anything else becomes max, and thinking cannot be
disabled.
Run emwesoft/GLM-5.3-NVFP4-MTP-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models