GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF overview

GLM 5.2 REAP50 Q3 K M GGUF A GGUF build of GLM 5.2, REAP expert pruned 50% and quantized to Q3 K M ~169 GB — sized to run on 2× 96 GB GPUs e.g. RTX PRO 6000 , …

ggufllama.cppmoeglmreapprunedtext-generationbase_model:zai-org/GLM-5.2base_model:quantized:zai-org/GLM-5.2license:mitendpoints_compatibleregion:usconversational

Runs locally from ~2.88 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
3
Pipeline
text-generation

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
GLM-5.2-REAP50-Q3_K_M-00001-of-00005.ggufGGUFQ3_K_M41.42 GBDownload
GLM-5.2-REAP50-Q3_K_M-00002-of-00005.ggufGGUFQ3_K_M41.77 GBDownload
GLM-5.2-REAP50-Q3_K_M-00003-of-00005.ggufGGUFQ3_K_M41.57 GBDownload
GLM-5.2-REAP50-Q3_K_M-00004-of-00005.ggufGGUFQ3_K_M41.69 GBDownload
GLM-5.2-REAP50-Q3_K_M-00005-of-00005.ggufGGUFQ3_K_M2.88 GBDownload

Model Details

Model IDpipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF
Authorpipenetwork
Pipelinetext-generation
Licensemit
Base modelzai-org/GLM-5.2
Last modified2026-06-18T17:34:32.000Z

Model README

---

license: mit

base_model: zai-org/GLM-5.2

base_model_relation: quantized

pipeline_tag: text-generation

library_name: gguf

tags:

  • gguf
  • llama.cpp
  • moe
  • glm
  • reap
  • pruned

---

GLM-5.2-REAP50-Q3_K_M-GGUF

A GGUF build of GLM-5.2, REAP expert-pruned (50%) and quantized to Q3_K_M (~169 GB) — sized to run on 2× 96 GB GPUs (e.g. RTX PRO 6000), ~192 GB VRAM, with room for context.

What this is

  • Base: zai-org/GLM-5.2 (glm_moe_dsa, ~753B MoE).
  • REAP-50: the 128 most-salient experts per layer kept (of 256) via Cerebras REAP saliency (gate × ‖expert_output‖), MTP layer dropped → ~394B params.
  • Quantized to Q3_K_M, split into 5 shards (~45 GB each).
  • Runs as full MLA attention (the DSA lightning-indexer is not used at inference — same simplification as the upstream conversion).

⚠️ Requires a patched llama.cpp (for now)

Stock llama.cpp can't load any GLM-5.2 GGUF yet: its GLM-DSA loader requires the DSA indexer tensors on every layer, but GLM-5.2 only ships them on a subset ("full") of layers → missing tensor 'blk.N.indexer.k_norm.weight'. The indexer is loaded-but-unused (the graph is DeepSeek-V2 MLA), so the fix is simply to make those tensors optional.

Apply the included llama.cpp-glm-dsa-indexer-optional.patch (src/models/glm-dsa.cpp) and rebuild, or wait for the upstream GLM-DSA runtime PR. After patching it loads and runs normally.

# in a recent llama.cpp checkout:
git apply llama.cpp-glm-dsa-indexer-optional.patch
cmake -B build -DGGML_CUDA=ON && cmake --build build -j
./build/bin/llama-cli -m GLM-5.2-REAP50-Q3_K_M-00001-of-00005.gguf --jinja -ngl 99 -p "..."

Quality caveat

This is the most aggressive variant: REAP-50 (~+37.5% perplexity vs full GLM-5.2) compounded with Q3_K 3-bit quant. It generates coherently (chain-of-thought intact, correct simple code) but is not a quality champion — it's the "fits 192 GB and runs fast" option. For higher quality at a larger footprint, see the MLX REAP-25 (+2.3% PPL) or the full GLM-5.2 ladder under pipenetwork.

Smoke-tested on Apple Metal (~17 tok/s); not tested on CUDA/RTX 6000.

Run pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models