pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF overview
GLM 5.2 REAP50 Q3 K M GGUF A GGUF build of GLM 5.2, REAP expert pruned 50% and quantized to Q3 K M ~169 GB — sized to run on 2× 96 GB GPUs e.g. RTX PRO 6000 , …
Runs locally from ~2.88 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.2-REAP50-Q3_K_M-00001-of-00005.gguf | GGUF | Q3_K_M | 41.42 GB | Download |
| GLM-5.2-REAP50-Q3_K_M-00002-of-00005.gguf | GGUF | Q3_K_M | 41.77 GB | Download |
| GLM-5.2-REAP50-Q3_K_M-00003-of-00005.gguf | GGUF | Q3_K_M | 41.57 GB | Download |
| GLM-5.2-REAP50-Q3_K_M-00004-of-00005.gguf | GGUF | Q3_K_M | 41.69 GB | Download |
| GLM-5.2-REAP50-Q3_K_M-00005-of-00005.gguf | GGUF | Q3_K_M | 2.88 GB | Download |
Model Details
| Model ID | pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF |
|---|---|
| Author | pipenetwork |
| Pipeline | text-generation |
| License | mit |
| Base model | zai-org/GLM-5.2 |
| Last modified | 2026-06-18T17:34:32.000Z |
Model README
---
license: mit
base_model: zai-org/GLM-5.2
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- moe
- glm
- reap
- pruned
---
GLM-5.2-REAP50-Q3_K_M-GGUF
A GGUF build of GLM-5.2, REAP expert-pruned (50%) and quantized to Q3_K_M (~169 GB) — sized to run on 2× 96 GB GPUs (e.g. RTX PRO 6000), ~192 GB VRAM, with room for context.
What this is
- Base: zai-org/GLM-5.2 (
glm_moe_dsa, ~753B MoE). - REAP-50: the 128 most-salient experts per layer kept (of 256) via Cerebras REAP saliency (
gate × ‖expert_output‖), MTP layer dropped → ~394B params. - Quantized to Q3_K_M, split into 5 shards (~45 GB each).
- Runs as full MLA attention (the DSA lightning-indexer is not used at inference — same simplification as the upstream conversion).
⚠️ Requires a patched llama.cpp (for now)
Stock llama.cpp can't load any GLM-5.2 GGUF yet: its GLM-DSA loader requires the DSA indexer tensors on every layer, but GLM-5.2 only ships them on a subset ("full") of layers → missing tensor 'blk.N.indexer.k_norm.weight'. The indexer is loaded-but-unused (the graph is DeepSeek-V2 MLA), so the fix is simply to make those tensors optional.
Apply the included llama.cpp-glm-dsa-indexer-optional.patch (src/models/glm-dsa.cpp) and rebuild, or wait for the upstream GLM-DSA runtime PR. After patching it loads and runs normally.
# in a recent llama.cpp checkout:
git apply llama.cpp-glm-dsa-indexer-optional.patch
cmake -B build -DGGML_CUDA=ON && cmake --build build -j
./build/bin/llama-cli -m GLM-5.2-REAP50-Q3_K_M-00001-of-00005.gguf --jinja -ngl 99 -p "..."
Quality caveat
This is the most aggressive variant: REAP-50 (~+37.5% perplexity vs full GLM-5.2) compounded with Q3_K 3-bit quant. It generates coherently (chain-of-thought intact, correct simple code) but is not a quality champion — it's the "fits 192 GB and runs fast" option. For higher quality at a larger footprint, see the MLX REAP-25 (+2.3% PPL) or the full GLM-5.2 ladder under pipenetwork.
Smoke-tested on Apple Metal (~17 tok/s); not tested on CUDA/RTX 6000.
Run pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models