MaliAir/GLM-5.2-MXFP4-MOE-Q8_0-GGUF overview
GLM 5.2 BF16 → mxfp4 moe Mixed Precision Quantization + Sharded This repository contains a quantized and sharded version of the GLM 5.2 base model originally i…
Runs locally from ~3.30 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.2-MXFP4-MOE-Q8_0-00001-of-00077.gguf | GGUF | Q8_0 | 4.82 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00002-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00003-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00004-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00005-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00006-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00007-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00008-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00009-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00010-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00011-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00012-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00013-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00014-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00015-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00016-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00017-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00018-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00019-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00020-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00021-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00022-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00023-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00024-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00025-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00026-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00027-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00028-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00029-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00030-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00031-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00032-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00033-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00034-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00035-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00036-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00037-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00038-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00039-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00040-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00041-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00042-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00043-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00044-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00045-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00046-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00047-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00048-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00049-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00050-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00051-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00052-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00053-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00054-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00055-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00056-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00057-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00058-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00059-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00060-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00061-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00062-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00063-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00064-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00065-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00066-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00067-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00068-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00069-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00070-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00071-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00072-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00073-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00074-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00075-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00076-of-00077.gguf | GGUF | Q8_0 | 4.99 GB | Download |
| GLM-5.2-MXFP4-MOE-Q8_0-00077-of-00077.gguf | GGUF | Q8_0 | 3.30 GB | Download |
Model Details
Model README
---
license: apache-2.0
library_name: llama.cpp
tags:
- quantized
- moe
- glm
- mixed-precision
- sharded
- gguf-split
- epyc
- rtx-5090
---
GLM-5.2-BF16 → mxfp4_moe (Mixed-Precision Quantization + Sharded)
This repository contains a quantized and sharded version of the GLM-5.2 base model (originally in BF16), converted to the GGUF format using llama.cpp.
To make the model both efficient and practical for deployment, I applied two distinct processing steps:
- Mixed-Precision Quantization – to aggressively reduce file size while preserving critical quality.
- Sharding (Splitting) – to break the large model into smaller, filesystem-friendly chunks (e.g., to comply with Hugging Face's 6 GB file limit or to enable resumable downloads).
---
1. Quantization Strategy
To achieve the best balance between model size and generation quality, I used a specialized mixed-precision approach with the llama-quantize tool.
This model uses:
mxfp4_moefor the majority of the tensors (especially the MoE expert layers). This is a specialized 4-bit floating-point format designed to efficiently compress Mixture-of-Experts (MoE) models while preserving reasoning capabilities.Q8_0specifically for the embedding table (token_embd.weight) and the output layer (output.weight). These layers are notoriously sensitive to quantization—keeping them at higher precision (8-bit) significantly reduces degradation in output quality, especially for long generations and complex reasoning tasks.
Quantization Command
The exact command used to generate the base quantized file was:
./llama-quantize \
--allow-requantize \
--tensor-type token_embd.weight=Q8_0 \
--tensor-type output.weight=Q8_0 \
GLM-5.2-BF16.gguf \
glm-5.2-mxfp4-moe-q8.gguf \
mxfp4_moe
---
2. Performance Benchmarks
Inference speed was evaluated on the following enterprise-grade hardware configuration:
- Motherboard: Gigabyte MZ73-LM0
- CPU: AMD EPYC 9654 (96 cores / 192 threads, Genoa architecture)
- Memory: 768 GB DDR5 (12-channel configuration, fully populated)
- GPU: NVIDIA RTX 5090 (32 GB VRAM)
Under this setup, the quantized model achieves a generation speed of:
> ≈ 6.8 tokens per second (t/s)
This benchmark was recorded using llama-server with standard inference settings (batch size = 2048, context length = 98304, with layers offloaded to the GPU via -ngl). The combination of the EPYC 9654's massive memory bandwidth (over 400 GB/s thanks to 12-channel DDR5) and the RTX 5090's computational muscle allows the MoE architecture to run smoothly, fully leveraging the mixed-precision quantization. 392GB DDR5 RAM from the MZ73-LM0 and 30GB GDDR7 RAM from the RTX 5090 are used. We are using this quantized and sharded GLM-5.2 base model within our EASA-compliant HPC environment, purpose-built for Aviation Safety and Compliance at masi-ai.com
At 6.8 t/s, the model provides a fluid, interactive experience suitable for real-time chat applications, complex multi-turn dialogues, and light batch processing—delivering near-instantaneous responses without frustrating delays.
./llama-server -m /path/GLM-5.2-MXFP4-MOE-Q8_0.GGUF -ngl 99 -ot "exps=CPU" -t 80 -tb 96 --host xxx.xxx.xxx.xxx --port 8080 --ctx-size 98304 --batch-size 2048 --ubatch-size 2048 --no-mmap --embeddings --pooling mean -fa on
This quantized and sharded GLM-5.2 base model powers our EASA-compliant High-Performance Computing (HPC) cluster, specifically architected for Aviation Safety and Compliance workflows at masi-ai.com.
KHM
Run MaliAir/GLM-5.2-MXFP4-MOE-Q8_0-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models