GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF overview

GLM 5.3 Flash GSQ RCO 3.0 bit, Q4 K attention GGUF Big thanks to @csantiago78 https://github.com/csantiago78 : the expert cache here builds on their implementa…

ggufglm5-nextmoeexpert-cachearxiv:2604.18556arxiv:2605.00649base_model:pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUFbase_model:quantized:pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUFlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~105.78 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
5,602
Likes
1
Pipeline
—
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.ggufGGUFQ4KATTN105.78 GBDownload

Model Details

Model IDneuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF
Authorneuralll
Pipeline—
Licensemit
Base modelpfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF
Last modified2026-09-26T11:42:48.000Z

Model README

---

base_model: pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF

base_model_relation: quantized

license: mit

library_name: gguf

tags:

- gguf

- glm5-next

- moe

- expert-cache

---

GLM-5.3-Flash GSQ-RCO 3.0-bit, Q4_K attention (GGUF)

> Big thanks to @csantiago78: the expert cache

> here builds on their implementation in llama.cpp PR

> #27861 ("GPU-resident LRU cache

> for host-offloaded MoE expert weights"), the first to get a working hot-expert

> cache into llama.cpp. This fork takes it further: filling all free VRAM, smarter

> eviction, CPU/GPU overlap and fused kernels.

A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF

(3.0-bit) made for faster single-stream decoding of this 117 GB model on

consumer GPUs with an expert cache.

Only the non-expert Q8_0 weights changed (attention, shared experts, dense FFN,

output head: Q8_0 to Q4_K, ~8.3 GB of the file). All 126 routed-expert tensors,

token_embd and every F32/BF16 tensor are copied bit-for-bit from the original.

Those Q8_0 weights are streamed by the GPU for every token, so shrinking them

speeds up decoding; the experts are unchanged.

Requires a llama.cpp fork

Stock llama.cpp can't load GLM-5.3-Flash yet. Use

neurall/llama.cpp (GLM-5.3-Flash support

plus a VRAM-filling MoE expert cache). **On 2 GPUs it decodes ~1.56x faster than

running without the cache** (13.8 to 21.5 t/s on a 1500-token chat reply):

llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf

The fork has automatic defaults that differ from stock llama.cpp, used only when a MoE

model is bigger than free VRAM: experts in RAM with a GPU expert cache in all free VRAM,

CPU repacking off, -ub 2048 on 20+ GiB GPUs (faster long prompts), 32k context (-c N

for more, at the cost of cache VRAM) and one CPU core per GPU left free. Settings you

pass win; the fork README explains each one. The file also loads on any build with

GLM-5.3-Flash support (PRs #27773 / #27917), at the no-cache speed.

Why the fork matters: stock llama.cpp splits a model across GPUs by layer, so for

a single request only one GPU works at a time: GPU1 idles while GPU2 runs its layers,

and both idle while the CPU computes the experts that didn't fit in VRAM. A second

GPU adds memory, not parallel compute. This fork turns that memory into a live cache

of the experts actually being used, so most expert work runs on the GPUs, and in

parallel with the CPU computing the rest.

Same VRAM, different use. Each token uses only 8 of the 288 experts in each layer.

  • Without the cache, experts are placed statically, whole layers at a time: ~36 GB

fits all 288 experts of ~14 of the 42 MoE layers. Most of that VRAM holds experts

the current token doesn't touch, so only ~33% of each token's expert work runs on

GPU and the CPU does ~67%, one after the other.

  • The fork fills the same VRAM with the ~100 most-used experts of every layer. Usage

is skewed, so those cover ~74% of what tokens actually pick in real chat output:

~74% of expert work runs on GPU and the CPU does ~26%, at the same time as the GPUs.

Results

2x RTX 3090 (48 GB VRAM; one CPU x16, one X570 chipset x4 slot) + Ryzen 7 3700X +

125 GB DDR4, single stream, temperature 0. Short prompt: 1500-token chat reply via

/v1/chat/completions. Long prompt: 12k-token code prompt (llama.cpp sources), prompt

processing then decode. Perplexity: wikitext-2 test, 40 x 512-token chunks, same build

for both files. Speeds are t/s, without cache -> with cache.

| | short: decode | long: prompt processing | long: decode | PPL |

|---|---|---|---|---|

| original 3.0-bit | 12.5 -> 18.5 | 209 -> 178 | 11.2 -> 14.7* | 3.5534 |

| this file | 13.8 -> 21.5 | 217 -> 181 | 12.3 -> 15.7* | 3.5871 (+0.95%) |

\* Model already in RAM (OS page cache), as on a server after its first request. The

first run after switching models is slower, once, while the file is read from disk.

Prompt processing with the cache is ~17% slower: without it whole layers stay in VRAM

and are never uploaded; the cache trades that VRAM for faster decode.

Note: the original quant's RCO allocation chose each tensor's precision on

purpose; overriding the Q8_0 tensors to Q4_K trades ~1% perplexity for ~7-16% speed.

If you want the RCO allocation as designed, use the original file.

How it was made

llama-quantize --allow-requantize --tensor-type-file tensor-types-q4kattn.txt \
    GLM-5.3-Flash-GSQ-RCO-3.0bit.gguf GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf Q4_K

tensor-types-q4kattn.txt (in this repo) lists every tensor with its target type:

Q8_0 non-expert tensors (except token_embd) to q4_k, everything else its

original type. Tensors whose type doesn't change are copied, not requantized.

Credits and license

a community reproduction of the GSQ and RCO methods by IST-DASLab

(GSQ, RCO).

Not an IST-DASLab release.

  • Expert cache: builds on llama.cpp PR #27861 (csantiago78).
  • GLM-5.3-Flash support: llama.cpp PRs #27773

and #27917 (timkhronos).

MIT license, see LICENSE.

Run neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models