bogdan-radulescu/qwen3.6-35B-A3B-asym-GGUF overview
Qwen3.6 35B A3B — Asymmetric 2 bit Routed Expert Quant GGUF A GGUF quantization of Qwen3.6 35B A3B built for larger than RAM / SSD streaming inference . Rather…
Runs locally from ~10.93 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| qwen3.6-35b-a3b-asym.gguf | GGUF | GGUF | 10.93 GB | Download |
Model Details
| Model ID | bogdan-radulescu/qwen3.6-35B-A3B-asym-GGUF |
|---|---|
| Author | bogdan-radulescu |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.6-35B-A3B |
| Last modified | 2026-07-22T08:26:57.000Z |
Model README
---
license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- moe
- quantization
- imatrix
- asymmetric-quant
- larger-than-ram
- streaming-inference
---
Qwen3.6-35B-A3B — Asymmetric 2-bit Routed-Expert Quant (GGUF)
A GGUF quantization of Qwen3.6-35B-A3B built for **larger-than-RAM /
SSD-streaming inference**. Rather than quantizing every tensor to the same width,
it quantizes asymmetrically: the routed-expert weights (the bulk of the bytes
moved per token in a Mixture-of-Experts model) are pushed to ~2 bits, while
everything quality-sensitive (attention, router, shared experts, embeddings,
norms) stays at Q8_0.
It matches the 4-bit IQ4_XS baseline on quality (identical 9/10 coding) while
being ~33% smaller and streaming fewer cold bytes per token.
Why this exists - intelligence that fits on basic hardware
This quant is built for Fono, whose aim is to
run the most capable model that will actually fit on ordinary, memory-limited
hardware. The guiding principle is intelligence per gigabyte: spend precision only
where the model is genuinely sensitive, and compress everything else hard. That lets
a larger, smarter model drop into the same memory budget, so a modest machine gets
to punch well above its weight instead of settling for a smaller, weaker model.
What "asymmetric quant" means here
In an MoE, decode speed under memory pressure is dominated by how many
routed-expert bytes must be pulled from SSD per token. So this recipe spends
bits where they matter and saves them where they don't:
| Tensor group | Precision | Why |
|---|---|---|
| Routed experts ffn_gate_exps / ffn_up_exps | IQ2_XXS (~2.06 bpw, imatrix) | dominant streamed mass — the big lever |
| Routed expert ffn_down_exps | Q2_K (imatrix) | smart 2-bit; dims are natively ÷256, so no padding needed |
| Attention / router / shared experts (ffn_*_shexp) / embeddings / norms | Q8_0 | near-lossless where the model is sensitive |
Note the architectural contrast with the gemma-4 sibling: Qwen uses separate
gate/up/down expert tensors plus always-on shared experts, and its routed
reduction dims are already 256-aligned (down_exps = 512, gate/up = 2048).
So the smart imatrix i-quants tile directly — no zero-padding required. This
is what makes the technique general: the only arch-specific wrinkle (alignment)
is a no-op here.
Only imatrix-aware i-quants (IQ2_XXS, Q2_K) are used for the 2-bit tensors;
crude legacy 2-bit formats at the same bit budget collapse model capability. The
importance matrix was computed with llama-imatrix over Bartowski's
calibration_datav3 (generic web/wiki text, deliberately disjoint from our
benchmark tasks so scores aren't inflated), 100 chunks × 512 tokens.
Why you might use it
There are two distinct wins:
- It fits in RAM where the full-size model won't. At 11.74 GB the asymmetric
quant can sit entirely in memory on a machine where the 17.7 GB 4-bit baseline
would have to spill to SSD (or fail to load). Running fully in RAM is far faster
than streaming — so on that hardware the smaller-but-equally-smart model simply
is the faster model, and you get 35B-class capability where you otherwise
couldn't.
- When it must stream, it streams less. Past the RAM budget it moves fewer cold
bytes per token than the uniform 4-bit build, so decode stays quick.
Other reasons: a Qwen3.6 GGUF that keeps full coding capability at ~⅔ the disk of the
4-bit build, and low bytes-moved-per-token rather than just a small file.
(If the full-size model already fits comfortably in your RAM, a uniform 4-bit build
is marginally faster there — fewer dequant ops. The asymmetric quant's advantage is
exactly the common case where it fits and the bigger one doesn't.)
Measured results
Coding = run-the-code pass rate (10 objective tasks); polish = instruction-formatting
score.
| variant | size | coding | polish | note |
|---|---:|:---:|:---:|---|
| IQ4_XS (uniform 4-bit baseline) | 17.7 GB | 10/10 | 0.986 | reference |
| asym 2-bit (this model) | 11.74 GB | 9/10 | 0.981 | −33% size, fewer cold bytes/token, faster @6G cap |
The single lost coding point is the same hard Rust task the baseline's asym-class
peers also miss; general prose, polish, and the other 9 tasks are unaffected.
Files
qwen3.6-35b-a3b-asym.gguf(11.74 GB)
Usage (llama.cpp)
llama-cli -m qwen3.6-35b-a3b-asym.gguf -p "Write a prime sieve in Rust."
Limitations
- ~2-bit routed experts trade a little robustness on the hardest reasoning/parsing
tasks vs the 4-bit baseline.
Run bogdan-radulescu/qwen3.6-35B-A3B-asym-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models