neuralll author hub
GLM 5.3 Flash MTP draft head GGUF, Q4 K The GLM 5.3 Flash multi token prediction MTP / NextN block as a small standalone GGUF 4.3 GiB , for speculative decoding next to a GLM 5.3 Flash model you already have. No need to download a whole model again just to get MTP. sh llama serv…
Models
2
Downloads
5,602
Run models locally with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.