neopolita/Qwen3.6-11B-A3B-Niwaki-4bit-GGUF overview
Niwaki niwaki.png Qwen3.6 11B A3B Niwaki 4bit GGUF GGUF builds of Qwen3.6 11B A3B Niwaki 4bit mlx https://huggingface.co/neopolita/Qwen3.6 11B A3B Niwaki 4bit …
Runs locally from ~5.50 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | neopolita/Qwen3.6-11B-A3B-Niwaki-4bit-GGUF |
|---|---|
| Author | neopolita |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | neopolita/Qwen3.6-11B-A3B-Niwaki-4bit-mlx |
| Last modified | 2026-08-09T16:43:37.000Z |
Model README
---
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE
base_model: neopolita/Qwen3.6-11B-A3B-Niwaki-4bit-mlx
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- moe
- pruning
- mixture-of-experts
---
Qwen3.6-11B-A3B-Niwaki-4bit-GGUF
**GGUF builds of Qwen3.6-11B-A3B-Niwaki-4bit-mlx —
Qwen3.6-35B-A3B pruned to 19B total / 3.3B active parameters — for
llama.cpp and everything built on it.**
Niwaki (庭木): every routed expert individually width-pruned to the neurons
its own routed tokens actually use, reconstructed to compensate, then briefly distilled from the full model, and
stored at low precision. A paper with the full method is coming soon.
Files
| file | size | wt2 ppl (llama.cpp, 512-ctx) |
|---|---|---|
| Qwen3.6-11B-A3B-Niwaki-4bit-UD-Q3K.gguf (recommended) | 5.9 GB | 18.29 ±0.14 |
| Qwen3.6-11B-A3B-Niwaki-4bit-Q4_K_M.gguf | 7.0 GB | 18.23 ±0.14 |
Reference Qwen3.6-35B-A3B at Q8_0 measures 6.95 under the identical
protocol (llama-perplexity, WikiText-2 test, 512-token windows). These
llama.cpp numbers are not directly comparable to the MLX repo's 2048-window
benchmarks; the relative standings match across both.
Generation battery (measured on the canonical MLX weights; reference scores 0.89 / 0.78): bigram-diversity avg/min = 0.82 / 0.62 across an 8-prompt code/reasoning/chat/creative battery — see the limitation note below.
The recommended UD-Q3K build is quantized structure-aware (importance matrices calibrated on the same code-weighted corpus as the model itself), mirroring the
artifact's native allocation: the always-active backbone (attention, shared
experts, embeddings) is kept at high precision (Q6_K) while the pruned routed
experts ride a compact carrier (q3_k, imatrix-guided). It matches or beats
uniform Q4_K_M quality at ~20% fewer bytes on this model family.
Known limitation (measured on the MLX canonical): sustained code generation is the weakest axis of this model (battery minimum 0.62 on long-file code prompts, vs 0.75+ for its larger siblings); output tends to truncate early rather than stay coherent to the end of a long file. Chat, reasoning, and short-form writing are functional. Choose the 19B for heavy code work.
Model dimensions
| | |
|---|---|
| total / active parameters | 11B / ~3.05B |
| layers / routed experts / top-k | 40 / 256 / 8 |
| expert intermediate size | 128 (from 512) |
| context | as base model |
| conversion note | speculative-decoding (MTP) draft block not included |
Usage
llama-cli -m Qwen3.6-11B-A3B-Niwaki-4bit-UD-Q3K.gguf -p "your prompt" -n 256
# or serve:
llama-server -m Qwen3.6-11B-A3B-Niwaki-4bit-UD-Q3K.gguf
Requires a recent llama.cpp with Qwen3.6 (hybrid linear-attention) support.
Canonical benchmarks, method outline, and the MLX-native artifact:
Qwen3.6-11B-A3B-Niwaki-4bit-mlx. Family:
Base model by the Qwen team (Apache 2.0); pruning, distillation, and GGUF
builds by the Niwaki project, 2026-08.
Run neopolita/Qwen3.6-11B-A3B-Niwaki-4bit-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models