GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF overview

Qwen3 Coder 30B A3B Instruct Q4 0 H32 Hadamard rotated for the 1bit engine Qwen/Qwen3 Coder 30B A3B Instruct https://huggingface.co/Qwen/Qwen3 Coder 30B A3B In…

gguf1bit-enginestrix-halorocmhadamardw4a4moetext-generationbase_model:Qwen/Qwen3-Coder-30B-A3B-Instructbase_model:quantized:Qwen/Qwen3-Coder-30B-A3B-Instructlicense:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~16.12 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3-Coder-30B-A3B-Q4_0-H32.ggufGGUFQ4_016.12 GBDownload

Model Details

Model ID1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF
Author1bit-MONSTER
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3-Coder-30B-A3B-Instruct
Last modified2026-09-28T19:53:29.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3-Coder-30B-A3B-Instruct

tags:

  • gguf
  • 1bit-engine
  • strix-halo
  • rocm
  • hadamard
  • w4a4
  • moe

pipeline_tag: text-generation

---

Qwen3-Coder-30B-A3B-Instruct Q4_0-H32 (Hadamard-rotated) for the 1bit engine

Qwen/Qwen3-Coder-30B-A3B-Instruct as a

Q4_0 GGUF whose matmul weights, the MoE experts included, are rotated by a 32-point

Walsh-Hadamard transform. It is made for the prompt-processing route of the

1bit engine on AMD Strix Halo (Ryzen AI Max+ 395,

Radeon 8060S): 4-bit weights times 4-bit activations on the GPU's matrix units, with the rotation

keeping the 4-bit activations close to the full model.

It only runs correctly on the 1bit engine's lean ROCm build. Its weights are rotated, so the

activations must be rotated the same way, and only that build does it. Stock llama.cpp loads the

file but computes garbage. 1bit serve recognises the file (it carries

onebit.hadamard_q4_0 = 32) and picks the right route by itself.

Run it

Best for agents and coding tools, whose prompts are long: pair it with an ordinary GGUF of the

same model. Conversations that start with 2,048 prompt tokens or more go to this file, shorter

ones to the ordinary file on Vulkan (engine docs/serve.md, "Short and long prompts"):

# engine built with -DONEBIT_LEAN=ON -DONEBIT_LEAN_ROCM=ON (docs/lean.md)
1bit serve -m Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf --long-model Qwen3-Coder-30B-A3B-Q4_0-H32.gguf

or on its own: 1bit serve -m Qwen3-Coder-30B-A3B-Q4_0-H32.gguf.

Measured (Strix Halo, 2026-09-28)

What a user waits for is the first token plus the answer. Effective speed = output tokens over

that whole time (engine tools/e2e_bench.py; 256 output tokens, median of 3), tok/s:

| Prompt tokens | 512 | 2,048 | 8,192 | 16,384 |

|---|---|---|---|---|

| Vulkan, Q4_K_M | 70.7 | 52.1 | 20.3 | 8.6 |

| ROCm W4A4, this file | 66.6 | 55.0 | 27.0 | 13.4 |

| 1bit serve --long-model (both) | 71.3 | 53.1 | 25.5 | 13.1 |

At 16K prompt tokens the first token comes after 13.8 s instead of 24.9 s.

Accuracy: KLD against Unsloth's Q8_0 over wikitext-2 (40 x 512; Q8_0 PPL 8.489). Speed from

llama-bench (3 runs):

| This file on the lean ROCm build | pp512 | pp2048 | tg128 | Mean KLD | Same top token |

|---|---|---|---|---|---|

| exact int8 path (GGML_W4A4_TENSORS=) | 1,218 | 1,186 | 79.0 | 0.048 | 91.5% |

| W4A4 on attention and experts (default) | 2,156 | 2,048 | 79.2 | 0.081 | 88.7% |

| W4A4 on an unrotated Q4_0, for comparison | 2,155 | 2,058 | 79.0 | 0.137 | 85.3% |

How it was made

tools/hadamard_q4_0.py Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf Qwen3-Coder-30B-A3B-Q4_0-H32.gguf

From unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF

(revision b17cb02d), its Q8_0. Rotated and Q4_0: the attention projections and the routed

experts (336 tensors); the token embedding is Q4_0 unrotated (a row lookup, not a matmul), the

output head Q6_K, norms and the router F32. 16,503 MiB (4.53 bits per weight). No imatrix went into this file; the quantize.imatrix.*

keys it carries are inherited from Unsloth's Q8_0 (on a rotated Q4_0 an imatrix has no effect,

engine docs/lean.md).

sha256 8aa9bf38cc9fd47cfd0b6e68ef40b79d69fde45ff3c1c903ffeec71de16b6b5f

Licence

Apache-2.0, as the base model (LICENSE). Qwen3-Coder by the Qwen team. Quantization

source by Unsloth.

Run 1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models