GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF overview

Introducing GLM 4.6 REAP 252B A32B, a memory efficient compressed variant of GLM 4.6 that maintains near identical performance while being 30% lighter. The ori…

gguftext-generationbase_model:zai-org/GLM-4.6base_model:quantized:zai-org/GLM-4.6license:mitendpoints_compatibleregion:usconversational

Runs locally from ~42.43 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
9
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
GLM-4.6-REAP-252B-A32B-MXFP4_MOE-00001-of-00003.ggufGGUFGGUF45.21 GBDownload
GLM-4.6-REAP-252B-A32B-MXFP4_MOE-00002-of-00003.ggufGGUFGGUF45.50 GBDownload
GLM-4.6-REAP-252B-A32B-MXFP4_MOE-00003-of-00003.ggufGGUFGGUF42.43 GBDownload

Model Details

Model IDandyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF
Authorandyjack
Pipelinetext-generation
Licensemit
Base modelzai-org/GLM-4.6
Last modified2026-06-21T04:47:43.000Z

Model README

---

license: mit

tags:

  • text-generation

base_model:

  • zai-org/GLM-4.6

---

Introducing GLM-4.6-REAP-252B-A32B, a memory-efficient compressed variant of GLM-4.6 that maintains near-identical performance while being 30% lighter.

The original model available here: https://huggingface.co/cerebras/GLM-4.6-REAP-252B-A32B

Note: this is a MXFP4-MOE accurate downstream low-bit quantization version. It has been tested with the latest version of llama.cpp.

bitcoin:

bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah

Run andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models