Joakimpalm-Zen/gpt-oss-20b-keep30-MXFP4-GGUF overview
gpt oss 20b keep 30 — a structurally pruned derivative MXFP4, GGUF < BEGIN runner toolcalling Serve with xyntetik runner — tool calls that still parse wh…
Runs locally from ~10.68 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gpt-oss-20b-keep30-MXFP4.gguf | GGUF | GGUF | 10.68 GB | Download |
Model Details
| Model ID | Joakimpalm-Zen/gpt-oss-20b-keep30-MXFP4-GGUF |
|---|---|
| Author | Joakimpalm-Zen |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | openai/gpt-oss-20b |
| Last modified | 2026-08-28T09:15:10.000Z |
Model README
---
license: apache-2.0
base_model: openai/gpt-oss-20b
tags:
- gguf
- mxfp4
- moe
- expert-pruning
- structural-derivative
- xyntetik-runner
pipeline_tag: text-generation
---
gpt-oss-20b keep-30 — a structurally pruned derivative (MXFP4, GGUF)
<!-- BEGIN runner-toolcalling -->
Serve with xyntetik-runner — tool calls that still parse when the budget runs out
This GGUF is built to be served with xyntetik-runner, a single-binary, dependency-free C11 inference engine (CPU / CUDA / Metal, OpenAI-compatible server). Its one differentiator you cannot get from another engine:
> When a tool call runs past its token budget, Runner closes it to the smallest schema-legal document, so the arguments still parse.
Most OpenAI-compatible engines return finish_reason: "length" with an empty or malformed tool_calls the caller has to repair or retry. Runner does forced-truncation recovery: once a JSON document starts, it emits a legal ending when the budget expires. On local models — tight context, slow generation — that is the difference between an agent loop that finishes and one that retries from scratch.
Measured — what each engine hands the caller when the token budget cuts a tool call short (same box, same tool schema, same prompt, tool_choice:"required", temperature 0, budgets 1→64):
| engine | budget too small (1–16 tokens) | enough budget (64, control) |
|---|---|---|
| Runner | executable tool_calls, arguments parse | completes |
| vLLM 0.27.1 | no call; protocol framing leaks into content | completes |
| llama.cpp b10488 | no call; leak, then tool_calls with unparseable args | completes |
| Ollama 0.32.14 | no call; empty content, then HTTP 500 | completes |
| TensorRT-LLM 1.2.1 † | no call; <tool_call> leak, then empty content | completes |
| SGLang 0.5.17 † | no call; <tool_call> leak, then empty content | completes |
Every engine completes at 64 tokens — the failure is truncation, not misconfiguration. Below that, only Runner returns an executable call. This is the behaviour across every OpenAI-compatible engine measured, not a claim about engines not measured. † TensorRT-LLM and SGLang were measured on a Qwen3-1.7B substitute; truncation recovery is a property of the runtime, not the model.
Constrained decoding also holds the shape of a tool call under quantization: on a full quant ladder, schema conformance and tool selection stayed 100% down to Q4_0 (argument agreement decayed to 50%) — it guarantees the shape of a call at any quantization, not its contents.
Quickstart
# build the single binary (no deps), then serve this GGUF:
runner -m gpt-oss-20b-keep30-MXFP4.gguf --serve
# point any OpenAI-compatible client at http://localhost:8080/v1
Full recipe, raw responses, and the per-release regression gate: the truncation benchmark.
<!-- END runner-toolcalling -->
> **Measured status, updated 2026-08-28: this artifact does not pass the
> current quality bar, and it now fails both terms.** Re-measured at
> 43,209 scored positions on a decontaminated held-out split it scores
> mean KLD 0.07629 (bound: 0.05) and margin-qualified top-1 96.37%
> (bound: 97%). The earlier 400-position re-gate (KLD 0.062, margin-q
> 98.5%, failing on KLD only) was itself optimistic: both metrics are
> tail-sensitive and 400 positions under-samples them. A healing rescue
> was attempted on 2026-08-28 and FAILED, making the model measurably
> worse; details in the re-measurement section below. The artifact stays
> published as a measured near-miss with its numbers in the open. For a
> pruned artifact that passes the current bar, see
> for the newer artifact class, see
Run it
This file is 11.5 GB and keeps gpt-oss-20b's MXFP4 experts native, so it wants a
machine with roughly 12 GB of usable memory or more. On a smaller machine, start
with gemma-4-E2B-it-Q4_0
instead.
# Apple Silicon Mac
curl -L https://github.com/Joakimpalm-Zen/xyntetik-runner/releases/latest/download/runner-macos-arm64.tar.gz | tar xz
curl -L -o gpt-oss-20b-keep30-MXFP4.gguf \
https://huggingface.co/Joakimpalm-Zen/gpt-oss-20b-keep30-MXFP4-GGUF/resolve/main/gpt-oss-20b-keep30-MXFP4.gguf
./runner-macos-arm64 -m gpt-oss-20b-keep30-MXFP4.gguf -i
Linux and Windows builds are on the
Using a different runtime? That is the point of this artifact. The reduced
expert count is declared in the GGUF metadata, so a runtime that resolves
num_experts from the file loads it unchanged, and one that assumes
"gpt-oss ⇒ 32 experts" fails. Either result is informative — please report it
either way, including the runtime and its version.
**This is not a quantization of gpt-oss-20b. It is a structural
modification:** each MoE layer's expert roster is pruned from 32 to its
own top-30 experts, selected from measured routing data. The expert FFN
tensors remain in their native MXFP4 format — untouched, not requantized —
and top-4 routing is retained. Category-wise:
| Model | Architecture change | Experts/layer | Routing | Expert format | Category |
|---|---|---|---|---|---|
| gpt-oss-20b (OpenAI) | none | 32 | top-4 | MXFP4 | official |
| this file | 2 experts pruned per layer | 30 | top-4 | MXFP4 (unchanged) | structural / expert-pruned derivative |
- File:
gpt-oss-20b-keep30-MXFP4.gguf(11,469,796,192 bytes) - SHA256:
e6b884b3de28a78858097bbee25c9c503975992a2bbc8943e0ef2c4bbbc14312 - Base:
ggml-org/gpt-oss-20b-GGUF→gpt-oss-20b-MXFP4.gguf
(SHA256 27cd6c432c7672cb812a92f611cf3ba7bbc35928262bb1e1253ff4ee6ae35901, 12.1 GB)
The reduced expert count is declared in the GGUF metadata. A runtime that
resolves num_experts from metadata loads this file unchanged; a runtime
that assumes "gpt-oss ⇒ 32 experts" fails. That property is deliberate —
this artifact doubles as a compatibility probe for exactly that class of
assumption.
Why these two experts — the methodology
Reproducible pipeline, tooling public in the
Xyntetik Runner repo:
- Trace: run the unpruned model over a mixed evaluation corpus with
the runner's MoE trace enabled (RUNNER_MOE_TRACE → JSONL of every
routing decision: position, layer, selected experts, gate weights).
- Saliency:
scripts/moe-prune-plan.pysums gate mass per
(layer, expert) over the whole trace — popularity × confidence.
(The script also supports REAP-style gate×activation-norm saliency.)
- Plan:
--keep-n 30— every MoE layer keeps its own top-30 by
saliency; the discarded pair differs per layer. The plan is a JSON
consumed directly by the runner's --prune-experts. **The exact
plan used for this file is published here as
gptoss-keep30-plan.json** — with the
base file and the runner, the build is fully reproducible.
- Prune:
runner --prune-experts plan.json --quantize out.gguf—
drops the expert tensors and the matching router rows; MXFP4 blocks
are copied, not recoded.
- Gate: KL-divergence + top-1 agreement against the unpruned parent
(scripts/kld-compare.py, mixed corpus).
Plan determinism was verified across four platforms: the plan generated
on an M1 is byte-identical to the archived plan from the original x86
run.
Validation — and where the cliff is
The interesting result is not just that keep-30 passes; it's that the
pruning frontier was measured, and this artifact sits exactly on it
(publication bar: top-1 ≥ 97%, mean KLD ≤ 0.05 vs parent):
| Variant | Top-1 vs parent | Verdict |
|---|---|---|
| keep-30 (this file) | 97.5% (mean KLD 0.0328, 80 scored positions) | passed the 2026-08-05 bar; superseded, see the 2026-08-15 re-gate and the 2026-08-28 re-measurement below |
| keep-28 | 91.25% | fails — 0.6 GB more saved, real quality drop; not shipped |
| keep-30 + sub-4-bit expert requant ("q4ne") | 22.5% | catastrophic — MXFP4-QAT experts tolerate no grid but their own; not shipped |
Two experts per layer is what gpt-oss-20b's expert roster measurably
tolerates. The third expert starts costing accuracy, and requantizing the
survivors destroys them. As a data point about expert redundancy in
GPT-OSS: it exists, and it is small.
Measured delta vs stock gpt-oss-20b-MXFP4
| | stock (12.1 GB) | keep-30 (11.5 GB) |
|---|---|---|
| Disk / resident weights | 12.1 GB | 11.5 GB (−5%) |
| 16 GB Mac, Metal | init fails (exceeds ~10.7 GB working-set limit) | fits after sudo sysctl iogpu.wired_limit_mb=13312 |
| 16 GB Mac, CPU, desktop load | 2.33 tok/s (pages) | holds resident with ~12 GB actually free; --mlock pins |
| 8 GB Mac | 0.54 tok/s (pages) | 0.78 tok/s (still pages — not the artifact for this class) |
| Quality | reference | superseded: mean KLD 0.07629, margin-qualified top-1 96.37% at 43,209 positions (2026-08-28 re-measurement below) |
The point of the −5%: on 16 GB Apple Silicon it is exactly the difference
between exceeding the Metal working-set ceiling and fitting under a
raised one. On machines with headroom it buys little — use the stock
file there.
2026-08-15 re-gate: this artifact does not pass the current bar
The numbers above were measured on 2026-08-05 at 80 scored positions. The
project's quality bar has since been tightened twice (400 scored positions as
the standard, and a margin-qualified top-1 that forgives argmax flips the
reference itself could not separate), and every published artifact is being
re-measured against it. keep-30 was re-run on 2026-08-15. It does not pass.
Both files were SHA-256 verified against the recorded hashes before measuring,
and the corpus fixture has not been modified since it was created. The artifact
has not changed. What changed is how carefully it is being measured.
| protocol | scored positions | plain top-1 | margin-qualified top-1 | mean KLD | verdict |
|---|---|---|---|---|---|
| published 2026-08-05 | 80 | 97.5% | not measured then | 0.0328 | claimed pass |
| same protocol, re-run 2026-08-15 | 80 | 92.5% | 100.0% | 0.0164 | does not reproduce |
| current standard, 2026-08-15 | 400 | 89.0% | 98.5% | 0.0618 | FAIL on KLD |
Three things this says, plainly.
The published 97.5% does not reproduce. Re-run at its own 80-position
protocol on current code, the same two files score 92.5%. A chat-endpoint
variant of the same comparison scores 90.0%. The gap is in the measurement
environment of 2026-08-05, not in the artifact, but it means the 97.5% figure
cannot be defended and should not be cited.
The short protocol was systematically flattering. Mean KLD against the
unpruned parent grows from 0.0164 at 80 positions to 0.0618 at 400. The bar is
0.05, so the artifact passes the short measurement and fails the long one. This
is exactly why 400 positions is now the standard: 80 positions did not sample
enough of the corpus to see the divergence.
**Under the current bar it fails on distributional distance, not on token
choice.** Margin-qualified top-1 is 98.5%, comfortably over the 97% line: where
this file disagrees with its parent, it is almost always at a near-tie the
parent could barely separate. The failure is the KLD term. Forgiving near-tie
flips does not forgive distributions being further apart, and it should not.
What that means if you are using this file
The 5% footprint saving and the Metal working-set argument below are unchanged
and still measured. The engineering claim ("two experts per layer is what this
roster tolerates; the third starts costing accuracy") is also unchanged, and the
keep-28 and sub-4-bit rows still fail far worse. What is withdrawn is the
narrower claim that this file clears the project's quality bar. On the current
standard it does not, and it stays published as a measured near-miss with its
numbers in the open rather than being quietly deleted or quietly relabelled.
If you want a gpt-oss-20b that has not been pruned, use the stock file. If you
need the 0.6 GB and can accept a measured mean KLD of 0.062 against the parent,
this file is exactly that trade, now stated honestly.
Reproduction
python3 scripts/kld-compare-raw.py \
--model-a gpt-oss-20b-keep30-MXFP4.gguf \
--model-b gpt-oss-20b-MXFP4.gguf \
--runner ./runner --corpus tests/fixtures/mixed-corpus.txt \
--max-positions 400 --out keep30.json
One measurement note recorded for anyone reproducing this on a mixture-of-
experts model: top-1 agreement is stable run to run, but mean KLD moves in the
fourth decimal (0.0164 against 0.0171 across repeats of the identical command)
depending on prefix-cache state. Dense models on the same harness reproduce to
the digit. The variation is far too small to move this verdict, but quote MoE
KLD to three decimals, not five.
2026-08-28 re-measurement: both terms fail, and a healing rescue failed too
As part of an expert-healing study, this artifact's fidelity was re-measured a
third time: 43,209 scored positions on a held-out split decontaminated at the
64-token window level, against a torch reconstruction verified to reproduce the
parent exactly (harness null 0.00000000) and to match this file's published
prune plan set-for-set.
| protocol | scored positions | plain top-1 | margin-qualified top-1 | mean KLD | verdict |
|---|---|---|---|---|---|
| published 2026-08-05 | 80 | 97.5% | not measured then | 0.0328 | does not reproduce |
| re-gate 2026-08-15 | 400 | 89.0% | 98.5% | 0.0618 | FAIL on KLD |
| re-measurement 2026-08-28 | 43,209 | 88.8% | 96.37% | 0.07629 | FAIL on BOTH terms |
Plain top-1 agrees with the 2026-08-15 measurement to 0.2 points, which is
strong evidence all three rows measured the same artifact. The tail-sensitive
metrics do not agree: KLD and margin-qualified top-1 both degrade as the sample
grows, because 400 positions still under-samples the tail. The 2026-08-15
section's claim that this file fails "on distributional distance, not on token
choice" is hereby corrected: at adequate sample size it fails the agreement
term as well, 96.37% against a 97% bound. Lesson recorded for this account's
methodology: re-gating uses tens of thousands of positions from now on.
A healing rescue was attempted and failed. The obvious repair, retraining
the 30 surviving experts and the router against the frozen unpruned parent
(the same local-distillation recipe that produced this account's bar-passing
dense surgical artifacts), was run as a preregistered study on 2026-08-28:
- Router-only retraining made the model worse under every configuration tried
(3 objectives, 8 learning-rate and dtype settings). Mechanism: with top-k
routing, gradients only flow through already-selected experts, so gradient
descent cannot reach the selection flips a pruned roster needs.
- Full expert healing across all 24 layers ended 5.58% WORSE than this
un-healed file (KLD 0.08055, margin-q 95.89%), even though all 24 layers
improved on their own local training objective. Local reconstruction quality
and end-to-end fidelity pointed in opposite directions.
- Even a successful heal would have shipped a worse artifact: healed tensors
cannot stay MXFP4 without a quantization-aware pass, so the file would grow
from 10.68 GiB to roughly 2x (Q8_0 experts) or 3.8x (BF16 experts).
So this artifact is not merely unhealed pending future work; on the evidence,
this recipe class cannot heal it. What the study did find is that the parent's
own router sends 5.13% of its (token, slot) selections to the pruned experts,
with a 56x spread across layers that this file's near-uniform plan never
consulted. A utilisation-weighted prune plan is therefore expected to beat
keep-30 without any training, and if that artifact is built and passes
measurement it will supersede this one. The full study write-up is
forthcoming.
Running it
Built and certified with Xyntetik Runner
— a from-scratch C inference engine (CPU AVX2/NEON, CUDA, Metal, OpenAI-
compatible server). Its --caps reports gpu.max_working_set_bytes so a
scheduler can apply the Metal placement rule before loading:
./runner -m gpt-oss-20b-keep30-MXFP4.gguf --serve --port 8080
Any GGUF runtime with gpt-oss MXFP4 support that reads the expert count
from metadata should also load it (llama.cpp does).
Publication policy
This project publishes only artifacts that passed the quality gate;
experiments that failed it (keep-28, the sub-4-bit requant) ship as
negative-result documentation in the runner repo instead of as uploads.
An optimization doesn't pass because the benchmark got faster — it has to
preserve the model and produce a configuration worth running.
Measured envelope
Every artifact this account publishes ships a measured-envelope sidecar
(gpt-oss-20b-keep30-MXFP4.gguf.envelope.json, in this repo) — the same file
xyntetik-runner reads at
load: download it next to the GGUF and the runner reports the measured state at
load. A measured envelope is a **dated observation, not a standing quality
claim** — it records exactly what this file is, where it came from, and what
was measured, so the claims above stay traceable.
| | |
|---|---|
| artifact sha256 | e6b884b3de28a78858097bbee25c9c503975992a2bbc8943e0ef2c4bbbc14312 |
| what changed | keep-30 expert prune (32 → 30 per layer) via --prune-experts |
| reference (unpruned parent) sha256 | 27cd6c432c7672cb812a92f611cf3ba7bbc35928262bb1e1253ff4ee6ae35901 |
| fidelity vs parent | does NOT pass the bar, both terms - mean KLD 0.07629 vs the 0.05 bound, margin-qualified top-1 96.37% vs the 97% bound, at 43,209 positions (a measured near-miss, numbers in the open; healing rescue attempted and failed) |
| tool calling | native Harmony protocol (gpt-oss family; truncation-safe) |
| measured | 2026-08-28 |
| verdict | experimental — loads with a banner; never a standing certification |
Run Joakimpalm-Zen/gpt-oss-20b-keep30-MXFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models