Myric/GLM-4.7-Flash-APEX-GGUF overview
GLM 4.7 Flash — APEX GGUF IMPORTANT Requires llama.cpp b10251 or newer. These files keep the MTP nextn draft head, which older builds do not know how to create…
Runs locally from ~9.54 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | Myric/GLM-4.7-Flash-APEX-GGUF |
|---|---|
| Author | Myric |
| Pipeline | text-generation |
| License | mit |
| Base model | zai-org/GLM-4.7-Flash |
| Last modified | 2026-08-07T12:20:43.000Z |
Model README
---
license: mit
base_model: zai-org/GLM-4.7-Flash
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- moe
- apex
- quantized
- imatrix
- glm4_moe_lite
- mla
- mtp
- speculative-decoding
- agentic-coding
- llama.cpp
---
GLM-4.7-Flash — APEX GGUF
> [!IMPORTANT]
> Requires llama.cpp b10251 or newer. These files keep the MTP (nextn) draft head, which
> older builds do not know how to create, so they refuse to load the file:
>
> ```
> error loading model: done_getting_tensors: wrong number of tensors; expected 868, got 862
> ```
>
> That is a loader version problem, not a corrupt download — the file has all 868 tensors, but
> a pre-b10251 build only asks for 862 because it has no GLM-4.7-Flash MTP support
> (#24868, commit 57c0921). Update
> llama.cpp and it loads. Open an issue here if you need an MTP-stripped build for an older
> runtime.
MoE-aware, mixed-precision APEX quantization of
zai-org/GLM-4.7-Flash — 31.2B total /
~3.6B active, and the most-downloaded model in the GLM family by a wide margin.
Three tiers, all built from the bf16 checkpoint (no intermediate requantization), all
imatrix-guided, all with the MTP draft head preserved and quantized so speculative
decoding works out of the box.
Architecture, and why it quantizes well
Despite the GLM name this is a DeepSeek-lineage design — llama.cpp converts it via
Glm4MoeLiteModel(DeepseekV2Model) to the deepseek2 GGUF architecture.
| | |
|---|---|
| layers | 47 (+1 MTP), first_k_dense_replace=1 |
| hidden | 2048, moe_intermediate_size 1536 |
| experts | 64 routed, top-4, + 1 shared, noaux_tc sigmoid routing |
| attention | MLA — kv_lora_rank 512, qk_nope/rope 192/64, v_head_dim 256 |
| context | 202,752 (rope_theta 1e6) |
| MTP | 1 nextn layer — a full MoE block + draft projections, 1.28B params |
Two properties make it an unusually good APEX target:
- Routed experts are 89% of all parameters (27.8B of 31.2B). Everything whose error
compounds on every token — MLA attention, shared experts, embeddings — is only ~2.1B, so
all of it stays at ~6.5–6.9 bpw for about 1.8GB while the entire remaining budget goes to
experts.
moe_intermediate_sizeis 1536, cleanly divisible by 256, so all three expert
matrices (gate/up/down) can use the full k-quant and IQ menu. Models with a non-256
expert dimension cannot: their ffn_down_exps is restricted to block-32/64/128 types.
MLA keeps long context affordable. Instead of caching full K and V per head, it caches
one 576-dim latent per token per layer:
| KV cache @ 198K ctx, bf16 | size |
|---|---|
| MLA (512+64 latent) | 11 GB |
| the same model as plain MHA | 342 GB |
Tiers
| tier | file size | bits/weight | routed-expert bpw | non-expert bpw | wikitext PPL | vs bf16 |
|---|---|---|---|---|---|---|
| bf16 (reference) | 62.5 GB | 16.00 | 16.00 | 16.00 | 12.112 | — |
| APEX-i-quality | 16.60 GB | 4.25 | 4.03 | 6.92 | 12.746 | +5.2% |
| APEX-i-compact | 13.27 GB | 3.40 | 3.05 | 6.65 | 13.837 | +14.2% |
| APEX-i-mini | 10.24 GB | 2.62 | 2.24 | 6.39 | 15.801 | +30.5% |
Allocation is by tensor role and layer band: edge layers keep more bits than the redundant
middle, routed experts absorb the compression, and shared experts / MLA projections /
embeddings stay high. The MTP draft head is pinned to Q4_K in every tier (see below).
On the PPL column: the imatrix is deliberately code-weighted (~70% multi-language code,
30% prose) and these numbers are measured on wikitext, so they are the pessimistic direction
of that domain mismatch. The task benchmark below tracks quality much better — i-quality
gives up 5.2% PPL but only 3.6 points of task score.
Agentic coding benchmark
Five hard from-scratch and fix-existing-code challenges driven through a real agentic loop
(OpenCode), scored by partial credit over independently re-run pytest — never the agent's
own self-report. MTP speculative decoding enabled (n_max=2). One run per tier.
| challenge | bf16 | i-quality | i-compact | i-mini |
|---|---|---|---|---|
| batch_scheduler | 10/10 | 10/10 | 10/10 | 10/10 |
| buddy_allocator | 7/11 | 5/11 | 5/11 | 5/11 |
| circuit_breaker | 7/7 | 7/7 | 7/7 | 5/7 |
| expr_parser_advanced | 20/20 | 20/20 | 20/20 | 7/20 |
| inventory_fix_and_extend | 11/11 | 11/11 | 11/11 | 11/11 |
| test cases | 55/59 (92.7%) | 53/59 (89.1%) | 53/59 (89.1%) | 38/59 (70.4%) |
| challenges fully passed | 4/5 | 4/5 | 4/5 | 2/5 |
| wall clock | 21.3 min | 8.9 min | 13.1 min | 11.5 min |
This is a sweep, not a cutoff claim. Three sizes are shipped with their measured degradation
so you can pick against your own budget — one model, one workload, one run per tier is not
enough to assert a general "quantize no further than X bpw" threshold, and none is asserted
here.
Read honestly:
- i-quality and i-compact are indistinguishable (both 53/59) despite a 24% difference in
expert precision and 3.3GB of file size. Between 4.03 and 3.05 bpw the compression is
effectively free on this workload.
- i-mini is genuinely weaker — 70.4%. It holds on the shorter challenges and degrades on
the longest, most compositional one (expr_parser_advanced, 20/20 → 7/20). This is the one
tier where PPL and task score agree that something real was lost. It is published because a
10.2GB file competes against other ~10GB models, not against the larger tiers here — but
that comparison is not benchmarked, so judge it on its own numbers.
- bf16's 2-test-case lead is a single run on one challenge and is within plausible
run-to-run variance for a stochastic agentic loop. Treat it as "bf16 is not clearly better
than i-quality" rather than a measured 3.6-point gap.
- Wall clock mixes decode speed with how long each tier flailed on the hard challenge. The
cleanest same-work comparison is expr_parser_advanced, identical 20/20: bf16 404s vs
i-quality 157s (2.6×).
MTP speculative decoding
The nextn draft head is preserved and quantized (not stripped). llama-imatrix never
exercises it — a plain forward pass reports the whole MTP layer as unused — so it receives no
imatrix data, and is pinned to Q4_K rather than an IQ type, which would be fitted against
statistics that do not exist for it.
Measured on i-mini:
| | decode |
|---|---|
| without MTP | 78.7 t/s |
| with MTP (n_max=2) | 116.2 t/s (+48%) |
Draft acceptance across tiers, from the benchmark runs:
| tier | drafter | acceptance | accepted / verify step |
|---|---|---|---|
| bf16 | bf16 | 58.5% | — |
| i-quality | Q4_K | 65.2% | — |
| i-compact | Q4_K | 59.0% | — |
| i-mini | Q4_K | 64.5% | 1.29 |
Acceptance is ~59–65% for every quantized tier, from 4.03 down to 2.24 bpw experts — the
speculative speedup survives aggressive quantization intact. Notably the bf16 file, whose
draft head is also bf16, posts the lowest acceptance of the four: a low-precision drafter is
not a compromise here, which matches independent results on other MoE checkpoints.
Enable it with:
llama-server -m GLM-4.7-Flash-APEX-i-quality.gguf \
--spec-type draft-mtp --spec-draft-n-max 2 \
--jinja -ngl 999 --ctx-size 32768
Plain decode speeds without MTP: i-quality 64.2 t/s, i-compact 71.4 t/s, i-mini 78.7 t/s.
Usage
# chat / agentic serving (add --spec-type draft-mtp for the MTP speedup)
llama-server -m GLM-4.7-Flash-APEX-i-quality.gguf --jinja -ngl 999 --ctx-size 32768
# one-shot
llama-cli -m GLM-4.7-Flash-APEX-i-quality.gguf --jinja -ngl 999 --ctx-size 8192 -st \
-p "Write a Python LFU cache with O(1) get and put."
This is a reasoning model — it emits thinking before its answer, so allow a generous token
budget or responses will be truncated mid-thought. For multi-turn agentic use, ZAI recommends
Preserved Thinking mode; llama.cpp exposes --reasoning-preserve.
Requires llama.cpp b10251 or newer (see the note at the top): Glm4MoeLite conversion
support plus GLM-4.7-Flash MTP support (upstream #24868).
Reproducing
- Converted from the bf16 checkpoint with
convert_hf_to_gguf.py(no intermediate
requantization).
- imatrix:
llama-imatrix, ctx 512, over a ~70% multi-language code / 30% prose corpus,
published at
Myric/Laguna-S-2.1-imatrix-calibration-study
(corpora/laguna_calib_codeweighted.txt). Expert coverage 100% except 9 expert stacks at
98.44% (63 of 64 experts observed).
- Per-tensor allocation emitted as a
--tensor-type-fileconsumed by stockllama-quantize;
no kernel or C++ changes.
- PPL:
llama-perplexityon wikitext-2 test, ctx 512.
Attribution & licenses
- Base: zai-org/GLM-4.7-Flash (MIT)
- Engine: llama.cpp (MIT)
- APEX: localai-org/apex-quant
(@mudler) (MIT)
- Calibration corpus: permissively-licensed sources (MIT / Apache-2.0 / BSD / Unlicense /
public-domain prose)
Unofficial community quantization; not affiliated with or endorsed by ZAI.
Run Myric/GLM-4.7-Flash-APEX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models