AliceThirty/GLM-5.3-Flash-UNCENSORED-V2-GGUF overview
A better uncensored version of GLM 5.3 Flash than my previous version. I merged MorinoNushi's lora with the fp16 model weights of GLM 5.3 Flash. Then I quantiz…
Runs locally from ~5.55 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| UD-Q4_K_XL/GLM-5.3-Flash-Uncensored-V2-UD-Q4_K_XL-00001-of-00005.gguf | GGUF | Q4_K_XL | 44.47 GB | Download |
| UD-Q4_K_XL/GLM-5.3-Flash-Uncensored-V2-UD-Q4_K_XL-00002-of-00005.gguf | GGUF | Q4_K_XL | 45.42 GB | Download |
| UD-Q4_K_XL/GLM-5.3-Flash-Uncensored-V2-UD-Q4_K_XL-00003-of-00005.gguf | GGUF | Q4_K_XL | 45.17 GB | Download |
| UD-Q4_K_XL/GLM-5.3-Flash-Uncensored-V2-UD-Q4_K_XL-00004-of-00005.gguf | GGUF | Q4_K_XL | 45.59 GB | Download |
| UD-Q4_K_XL/GLM-5.3-Flash-Uncensored-V2-UD-Q4_K_XL-00005-of-00005.gguf | GGUF | Q4_K_XL | 5.55 GB | Download |
Model Details
| Model ID | AliceThirty/GLM-5.3-Flash-UNCENSORED-V2-GGUF |
|---|---|
| Author | AliceThirty |
| Pipeline | — |
| License | — |
| Base model | MorinoNushi/GLM-5.3-Flash-Heretic-Abliterated-LoRA-V2-GGUF,zai-org/GLM-5.3-Flash |
| Last modified | 2026-09-18T20:42:48.000Z |
Model README
---
base_model:
- MorinoNushi/GLM-5.3-Flash-Heretic-Abliterated-LoRA-V2-GGUF
- zai-org/GLM-5.3-Flash
---
A better uncensored version of GLM-5.3-Flash than my previous version.
I merged MorinoNushi's lora with the fp16 model weights of GLM-5.3-Flash. Then I quantized the result using Unsloth's pipeline.
It allows faster inference time than loading the lora on top of the model, and it fixes some precision errors.
For instance, loading the lora on top of the base UD-Q4_K_XL considerably reduced the reasoning budget for some reason. And now it doesn't happen anymore.
Run AliceThirty/GLM-5.3-Flash-UNCENSORED-V2-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models