GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-gguf overview

Niwaki niwaki.png Qwen3.6 19B A3B Niwaki v2 2bit GGUF GGUF builds of Qwen3.6 19B A3B Niwaki v2 2bit mlx https://huggingface.co/neopolita/Qwen3.6 19B A3B Niwaki…

ggufllama.cppmoepruningmixture-of-expertstext-generationbase_model:neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlxbase_model:quantized:neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlxlicense:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~11.32 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-19B-A3B-Niwaki-v2-2bit-Q4_K_M.ggufGGUFQ4_K_M13.75 GBDownload
Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.ggufGGUFQ3K11.32 GBDownload

Model Details

Model IDneopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-gguf
Authorneopolita
Pipelinetext-generation
Licenseapache-2.0
Base modelneopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlx
Last modified2026-08-23T17:51:47.000Z

Model README

---

license: apache-2.0

license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE

base_model: neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlx

library_name: gguf

pipeline_tag: text-generation

tags:

- gguf

- llama.cpp

- moe

- pruning

- mixture-of-experts

---

!Niwaki

Qwen3.6-19B-A3B-Niwaki-v2-2bit-GGUF

**GGUF builds of Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlx

Qwen3.6-35B-A3B pruned to 19B total / ~3.3B active parameters — for

llama.cpp and everything built on it. No custom code: unlike the MLX repo,

these files run on stock llama.cpp.**

Niwaki (庭木) are Japan's garden trees, sculpted by meticulous pruning so

that every branch serves the form of the whole. This model applies that

spirit to a Mixture-of-Experts. **A paper with the full method is coming

soon.**

Files

| file | size | wt2 ppl (llama.cpp, 512-ctx) |

|---|---|---|

| Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.gguf (recommended) | 12.2 GB | 11.44 ±0.08 |

| Qwen3.6-19B-A3B-Niwaki-v2-2bit-Q4_K_M.gguf | 14.8 GB | 11.49 ±0.08 |

First-generation GGUF builds under the identical protocol:

| model (UD-Q3K builds) | size | wt2 ppl |

|---|---|---|

| Qwen3.6-27B-A3B-Niwaki-2bit | 13.5 GB | 10.41 ±0.07 |

| Qwen3.6-19B-A3B-Niwaki-2bit | 9.0 GB | 13.15 ±0.10 |

| Qwen3.6-11B-A3B-Niwaki-4bit | 5.9 GB | 17.24 ±0.13 |

This build beats the same-size first-generation 19B by 13% (11.44

vs 13.15) at 3.2 GB more on disk — the format pad described below.

Reference Qwen3.6-35B-A3B at Q8_0 measures 6.95 under the identical

protocol (llama-perplexity, WikiText-2 test, 512-token windows). These

llama.cpp numbers are not directly comparable to the MLX repo's 2048-window

benchmarks; the relative standings match across both.

**Generation battery (measured on the canonical MLX weights; reference

scores 0.63 / 0.51 under the identical battery):** bigram-diversity avg/min

= 0.84 / 0.54 across an 8-prompt code/reasoning/chat/creative battery.

The recommended UD-Q3K build is quantized structure-aware (importance

matrices calibrated on a mixed corpus), mirroring the artifact's native

allocation: the always-active backbone (attention, shared experts,

embeddings) is kept at high precision (Q6_K) while the routed experts ride

a compact carrier (q3_k, imatrix-guided). It matches or beats uniform

Q4_K_M quality at ~18% fewer bytes on this model family.

Format note: GGUF requires a uniform expert count per model, so the

shared-only layers carry zero-valued expert tensors stored at ~1.6

bits/weight (~3.1 GB of the file). They contribute nothing to outputs;

this is why these files are larger than the MLX repo at equal quality.

Model dimensions

| | |

|---|---|

| total / active parameters | ~19B / ~3.3B |

| layers / routed experts / top-k | 40 / 256 / 8 (layers 10–29 are shared-expert-only) |

| expert intermediate size | 512 (unchanged) |

| context | as base model |

| conversion note | speculative-decoding (MTP) draft block not included |

Usage

llama-cli -m Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.gguf -p "your prompt" -n 256
# or serve:
llama-server -m Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.gguf

Requires a recent llama.cpp with Qwen3.6 (hybrid linear-attention) support.

Canonical benchmarks and the MLX-native artifact:

Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlx. Family:

v2-4bit ·

first-generation 27B-2bit · 19B-2bit · 11B-4bit.

Base model by the Qwen team (Apache 2.0); pruning, distillation, and GGUF

builds by the Niwaki project, 2026-08.

Run neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models