1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF overview
Qwen3 Coder 30B A3B Instruct Q4 0 H32 Hadamard rotated for the 1bit engine Qwen/Qwen3 Coder 30B A3B Instruct https://huggingface.co/Qwen/Qwen3 Coder 30B A3B In…
Runs locally from ~16.12 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3-Coder-30B-A3B-Q4_0-H32.gguf | GGUF | Q4_0 | 16.12 GB | Download |
Model Details
| Model ID | 1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3-Coder-30B-A3B-Instruct |
| Last modified | 2026-09-28T19:53:29.000Z |
Model README
---
license: apache-2.0
base_model: Qwen/Qwen3-Coder-30B-A3B-Instruct
tags:
- gguf
- 1bit-engine
- strix-halo
- rocm
- hadamard
- w4a4
- moe
pipeline_tag: text-generation
---
Qwen3-Coder-30B-A3B-Instruct Q4_0-H32 (Hadamard-rotated) for the 1bit engine
Qwen/Qwen3-Coder-30B-A3B-Instruct as a
Q4_0 GGUF whose matmul weights, the MoE experts included, are rotated by a 32-point
Walsh-Hadamard transform. It is made for the prompt-processing route of the
1bit engine on AMD Strix Halo (Ryzen AI Max+ 395,
Radeon 8060S): 4-bit weights times 4-bit activations on the GPU's matrix units, with the rotation
keeping the 4-bit activations close to the full model.
It only runs correctly on the 1bit engine's lean ROCm build. Its weights are rotated, so the
activations must be rotated the same way, and only that build does it. Stock llama.cpp loads the
file but computes garbage. 1bit serve recognises the file (it carries
onebit.hadamard_q4_0 = 32) and picks the right route by itself.
Run it
Best for agents and coding tools, whose prompts are long: pair it with an ordinary GGUF of the
same model. Conversations that start with 2,048 prompt tokens or more go to this file, shorter
ones to the ordinary file on Vulkan (engine docs/serve.md, "Short and long prompts"):
# engine built with -DONEBIT_LEAN=ON -DONEBIT_LEAN_ROCM=ON (docs/lean.md)
1bit serve -m Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf --long-model Qwen3-Coder-30B-A3B-Q4_0-H32.gguf
or on its own: 1bit serve -m Qwen3-Coder-30B-A3B-Q4_0-H32.gguf.
Measured (Strix Halo, 2026-09-28)
What a user waits for is the first token plus the answer. Effective speed = output tokens over
that whole time (engine tools/e2e_bench.py; 256 output tokens, median of 3), tok/s:
| Prompt tokens | 512 | 2,048 | 8,192 | 16,384 |
|---|---|---|---|---|
| Vulkan, Q4_K_M | 70.7 | 52.1 | 20.3 | 8.6 |
| ROCm W4A4, this file | 66.6 | 55.0 | 27.0 | 13.4 |
| 1bit serve --long-model (both) | 71.3 | 53.1 | 25.5 | 13.1 |
At 16K prompt tokens the first token comes after 13.8 s instead of 24.9 s.
Accuracy: KLD against Unsloth's Q8_0 over wikitext-2 (40 x 512; Q8_0 PPL 8.489). Speed from
llama-bench (3 runs):
| This file on the lean ROCm build | pp512 | pp2048 | tg128 | Mean KLD | Same top token |
|---|---|---|---|---|---|
| exact int8 path (GGML_W4A4_TENSORS=) | 1,218 | 1,186 | 79.0 | 0.048 | 91.5% |
| W4A4 on attention and experts (default) | 2,156 | 2,048 | 79.2 | 0.081 | 88.7% |
| W4A4 on an unrotated Q4_0, for comparison | 2,155 | 2,058 | 79.0 | 0.137 | 85.3% |
How it was made
tools/hadamard_q4_0.py Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf Qwen3-Coder-30B-A3B-Q4_0-H32.gguf
From unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
(revision b17cb02d), its Q8_0. Rotated and Q4_0: the attention projections and the routed
experts (336 tensors); the token embedding is Q4_0 unrotated (a row lookup, not a matmul), the
output head Q6_K, norms and the router F32. 16,503 MiB (4.53 bits per weight). No imatrix went into this file; the quantize.imatrix.*
keys it carries are inherited from Unsloth's Q8_0 (on a rotated Q4_0 an imatrix has no effect,
engine docs/lean.md).
sha256 8aa9bf38cc9fd47cfd0b6e68ef40b79d69fde45ff3c1c903ffeec71de16b6b5f
Licence
Apache-2.0, as the base model (LICENSE). Qwen3-Coder by the Qwen team. Quantization
source by Unsloth.
Run 1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models