andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF overview
Introducing GLM 4.6 REAP 252B A32B, a memory efficient compressed variant of GLM 4.6 that maintains near identical performance while being 30% lighter. The ori…
Runs locally from ~42.43 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF |
|---|---|
| Author | andyjack |
| Pipeline | text-generation |
| License | mit |
| Base model | zai-org/GLM-4.6 |
| Last modified | 2026-06-21T04:47:43.000Z |
Model README
---
license: mit
tags:
- text-generation
base_model:
- zai-org/GLM-4.6
---
Introducing GLM-4.6-REAP-252B-A32B, a memory-efficient compressed variant of GLM-4.6 that maintains near-identical performance while being 30% lighter.
The original model available here: https://huggingface.co/cerebras/GLM-4.6-REAP-252B-A32B
Note: this is a MXFP4-MOE accurate downstream low-bit quantization version. It has been tested with the latest version of llama.cpp.
bitcoin:
bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ahRun andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models