Jerome0207/Huihui-GLM-5.2-abliterated-UD-Q2_K_MXFP4-GGUF overview
Huihui GLM 5.2 Abliterated — UD Q2 K MXFP4 GGUF This is a single quant operational mirror for RunPod cached models. It contains only the seven UD Q2 K MXFP4 GG…
Runs locally from ~9.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.2-UD-Q2_K_MXFP4-00001-of-00007.gguf | GGUF | Q2_K_MXFP4 | 9.0 MB | Download |
| GLM-5.2-UD-Q2_K_MXFP4-00002-of-00007.gguf | GGUF | Q2_K_MXFP4 | 45.71 GB | Download |
| GLM-5.2-UD-Q2_K_MXFP4-00003-of-00007.gguf | GGUF | Q2_K_MXFP4 | 45.45 GB | Download |
| GLM-5.2-UD-Q2_K_MXFP4-00004-of-00007.gguf | GGUF | Q2_K_MXFP4 | 45.45 GB | Download |
| GLM-5.2-UD-Q2_K_MXFP4-00005-of-00007.gguf | GGUF | Q2_K_MXFP4 | 45.45 GB | Download |
| GLM-5.2-UD-Q2_K_MXFP4-00006-of-00007.gguf | GGUF | Q2_K_MXFP4 | 46.29 GB | Download |
| GLM-5.2-UD-Q2_K_MXFP4-00007-of-00007.gguf | GGUF | Q2_K_MXFP4 | 6.90 GB | Download |
Model Details
| Model ID | Jerome0207/Huihui-GLM-5.2-abliterated-UD-Q2_K_MXFP4-GGUF |
|---|---|
| Author | Jerome0207 |
| Pipeline | text-generation |
| License | mit |
| Base model | zai-org/GLM-5.2 |
| Last modified | 2026-07-29T07:14:29.000Z |
Model README
---
license: mit
language:
- en
- zh
pipeline_tag: text-generation
base_model:
- zai-org/GLM-5.2
tags:
- gguf
- glm
- glm-5.2
- abliterated
- quantized
- llama.cpp
---
Huihui GLM-5.2 Abliterated — UD-Q2_K_MXFP4 GGUF
This is a single-quant operational mirror for RunPod cached models. It contains
only the seven UD-Q2_K_MXFP4 GGUF shards from the pinned huihui-ai source.
It does not claim authorship of the model or its quantization.
Destination repository: Jerome0207/Huihui-GLM-5.2-abliterated-UD-Q2_K_MXFP4-GGUF
Provenance
Source revision:
994200a058539c553f0977002e58a2f05de845c7.
The seven GGUF shards total 252,600,253,568 bytes. See
checksums.sha256 for the source SHA-256 object identifiers.
Runtime
The repository is intended for a current CUDA build of
llama.cpp. Start with the first shard;
llama.cpp discovers the remaining split files automatically:
GLM-5.2-UD-Q2_K_MXFP4-00001-of-00007.gguf
The validated deployment target uses three 96 GB GPUs, 131,072 context tokens,
one parallel slot and Q8_0 K/V caches.
License and limitations
The source model card declares the MIT license. See NOTICE for attribution
and exact provenance.
This is an “abliterated” derivative with intentionally reduced refusal
behavior. It is not a safety-tuned public service. Deploy it only behind
authentication and rate limits, and run coding/security agents in isolated
sandboxes without production credentials.
GLM-5.2 support in llama.cpp is evolving. Full DSA, Lightning Indexer,
IndexShare and MTP behavior must be validated before treating this deployment
as equivalent to the official inference implementation.
Run Jerome0207/Huihui-GLM-5.2-abliterated-UD-Q2_K_MXFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models