neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF overview
GLM 5.3 Flash GSQ RCO 3.0 bit, Q4 K attention GGUF Big thanks to @csantiago78 https://github.com/csantiago78 : the expert cache here builds on their implementa…
Runs locally from ~105.78 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf | GGUF | Q4KATTN | 105.78 GB | Download |
Model Details
Model README
---
base_model: pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF
base_model_relation: quantized
license: mit
library_name: gguf
tags:
- gguf
- glm5-next
- moe
- expert-cache
---
GLM-5.3-Flash GSQ-RCO 3.0-bit, Q4_K attention (GGUF)
> Big thanks to @csantiago78: the expert cache
> here builds on their implementation in llama.cpp PR
> #27861 ("GPU-resident LRU cache
> for host-offloaded MoE expert weights"), the first to get a working hot-expert
> cache into llama.cpp. This fork takes it further: filling all free VRAM, smarter
> eviction, CPU/GPU overlap and fused kernels.
A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF
(3.0-bit) made for faster single-stream decoding of this 117 GB model on
consumer GPUs with an expert cache.
Only the non-expert Q8_0 weights changed (attention, shared experts, dense FFN,
output head: Q8_0 to Q4_K, ~8.3 GB of the file). All 126 routed-expert tensors,
token_embd and every F32/BF16 tensor are copied bit-for-bit from the original.
Those Q8_0 weights are streamed by the GPU for every token, so shrinking them
speeds up decoding; the experts are unchanged.
Requires a llama.cpp fork
Stock llama.cpp can't load GLM-5.3-Flash yet. Use
neurall/llama.cpp (GLM-5.3-Flash support
plus a VRAM-filling MoE expert cache). **On 2 GPUs it decodes ~1.56x faster than
running without the cache** (13.8 to 21.5 t/s on a 1500-token chat reply):
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf
The fork has automatic defaults that differ from stock llama.cpp, used only when a MoE
model is bigger than free VRAM: experts in RAM with a GPU expert cache in all free VRAM,
CPU repacking off, -ub 2048 on 20+ GiB GPUs (faster long prompts), 32k context (-c N
for more, at the cost of cache VRAM) and one CPU core per GPU left free. Settings you
pass win; the fork README explains each one. The file also loads on any build with
GLM-5.3-Flash support (PRs #27773 / #27917), at the no-cache speed.
Why the fork matters: stock llama.cpp splits a model across GPUs by layer, so for
a single request only one GPU works at a time: GPU1 idles while GPU2 runs its layers,
and both idle while the CPU computes the experts that didn't fit in VRAM. A second
GPU adds memory, not parallel compute. This fork turns that memory into a live cache
of the experts actually being used, so most expert work runs on the GPUs, and in
parallel with the CPU computing the rest.
Same VRAM, different use. Each token uses only 8 of the 288 experts in each layer.
- Without the cache, experts are placed statically, whole layers at a time: ~36 GB
fits all 288 experts of ~14 of the 42 MoE layers. Most of that VRAM holds experts
the current token doesn't touch, so only ~33% of each token's expert work runs on
GPU and the CPU does ~67%, one after the other.
- The fork fills the same VRAM with the ~100 most-used experts of every layer. Usage
is skewed, so those cover ~74% of what tokens actually pick in real chat output:
~74% of expert work runs on GPU and the CPU does ~26%, at the same time as the GPUs.
Results
2x RTX 3090 (48 GB VRAM; one CPU x16, one X570 chipset x4 slot) + Ryzen 7 3700X +
125 GB DDR4, single stream, temperature 0. Short prompt: 1500-token chat reply via
/v1/chat/completions. Long prompt: 12k-token code prompt (llama.cpp sources), prompt
processing then decode. Perplexity: wikitext-2 test, 40 x 512-token chunks, same build
for both files. Speeds are t/s, without cache -> with cache.
| | short: decode | long: prompt processing | long: decode | PPL |
|---|---|---|---|---|
| original 3.0-bit | 12.5 -> 18.5 | 209 -> 178 | 11.2 -> 14.7* | 3.5534 |
| this file | 13.8 -> 21.5 | 217 -> 181 | 12.3 -> 15.7* | 3.5871 (+0.95%) |
\* Model already in RAM (OS page cache), as on a server after its first request. The
first run after switching models is slower, once, while the file is read from disk.
Prompt processing with the cache is ~17% slower: without it whole layers stay in VRAM
and are never uploaded; the cache trades that VRAM for faster decode.
Note: the original quant's RCO allocation chose each tensor's precision on
purpose; overriding the Q8_0 tensors to Q4_K trades ~1% perplexity for ~7-16% speed.
If you want the RCO allocation as designed, use the original file.
How it was made
llama-quantize --allow-requantize --tensor-type-file tensor-types-q4kattn.txt \
GLM-5.3-Flash-GSQ-RCO-3.0bit.gguf GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf Q4_K
tensor-types-q4kattn.txt (in this repo) lists every tensor with its target type:
Q8_0 non-expert tensors (except token_embd) to q4_k, everything else its
original type. Tensors whose type doesn't change are copied, not requantized.
Credits and license
- Base model: zai-org/GLM-5.3-Flash, MIT.
- Original quantization: pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF,
a community reproduction of the GSQ and RCO methods by IST-DASLab
Not an IST-DASLab release.
- Expert cache: builds on llama.cpp PR #27861 (csantiago78).
- GLM-5.3-Flash support: llama.cpp PRs #27773
and #27917 (timkhronos).
MIT license, see LICENSE.
Run neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models