Jianqiao1/Qwen3.6-27B-MTP-MoQ-GGUF overview
Qwen3.6 27B MTP MoQ GGUF This repository contains imatrix aware GGUF quantizations of the original Qwen3.6 27B MTP model using projected Mixture of Quantizatio…
Runs locally from ~10.07 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.6-27B-MTP-MoQ-3.2.gguf | GGUF | GGUF | 10.07 GB | Download |
| Qwen3.6-27B-MTP-MoQ-3.6.gguf | GGUF | GGUF | 11.31 GB | Download |
| Qwen3.6-27B-MTP-MoQ-3.8.gguf | GGUF | GGUF | 11.95 GB | Download |
| Qwen3.6-27B-MTP-MoQ-4.1.gguf | GGUF | GGUF | 13.27 GB | Download |
| Qwen3.6-27B-MTP-MoQ-4.3.gguf | GGUF | GGUF | 13.97 GB | Download |
| Qwen3.6-27B-MTP-MoQ-4.6.gguf | GGUF | GGUF | 14.20 GB | Download |
| Qwen3.6-27B-MTP-MoQ-4.8.gguf | GGUF | GGUF | 15.04 GB | Download |
| Qwen3.6-27B-MTP-MoQ-4.9.gguf | GGUF | GGUF | 15.39 GB | Download |
| Qwen3.6-27B-MTP-MoQ-5.1.gguf | GGUF | GGUF | 16.25 GB | Download |
Model Details
Model README
---
license: apache-2.0
base_model:
- Qwen/Qwen3.6-27B
base_model_relation: quantized
tags:
- gguf
- llama.cpp
- moq
- mtp
- quantized
---
Qwen3.6-27B-MTP-MoQ-GGUF
This repository contains imatrix-aware GGUF quantizations of the original Qwen3.6-27B MTP model using projected Mixture-of-Quantization (MoQ) tensor policies. These are quantizations of the base Qwen model, not a fine-tune or merge.
The complete local series was re-evaluated alongside kaitchup/Qwen3.6-27B-GGUF-MoQ under identical conditions. Lower is better in all three charts.
The interactive comparison report supports pan, wheel/mode-bar zoom, view reset, and cross-chart series toggles. The companion CSV contains all measured values and tensor-composition summaries.
Comparison Summary
All 18 measured GGUF files contain 866 tensors and 27,320,697,856 parameters. They were evaluated from scratch against the same BF16 reference logits on WikiText-2 with context length 512.
Across the 8 Jianqiao1 points inside the kaitchup size range, linearly interpolating the kaitchup curve at exactly the same file size favors Jianqiao1 on:
- Mean KLD: 8 of 8 points
- PPL: 8 of 8 points
- p999 KLD: 8 of 8 points
4 especially close pairs favor Jianqiao1 on PPL, Mean KLD, and p999 KLD while the Jianqiao1 GGUF is slightly smaller:
| Jianqiao1 | GB | kaitchup | GB | PPL (J / K) | Mean KLD (J / K) | p999 KLD (J / K) |
|---|---:|---|---:|---:|---:|---:|
| MoQ-3.8 | 12.836 | MoQ-3.5 | 12.888 | 7.097372 / 7.410869 | 0.067854 / 0.103733 | 4.466393 / 4.815393 |
| MoQ-4.1 | 14.249 | MoQ-4.0 | 14.430 | 7.030544 / 7.077410 | 0.044025 / 0.054240 | 3.270178 / 3.549793 |
| MoQ-4.6 | 15.242 | MoQ-4.25 | 15.312 | 6.974555 / 7.030861 | 0.026953 / 0.038407 | 2.159290 / 2.421159 |
| MoQ-4.8 | 16.150 | MoQ-4.5 | 16.214 | 6.934147 / 6.992102 | 0.023306 / 0.027908 | 1.785062 / 2.034667 |
After refreshing the 3.2–4.1 recipes, the earlier local exceptions disappear: every Jianqiao1 point within the comparison range is now below the linearly interpolated kaitchup curve on all three headline metrics. The improvement is especially visible at 3.6 and 4.1, while 3.8 now forms a substantially stronger direct comparison with kaitchup 3.5.
Full Quality Results
Actual BPW is computed from the complete GGUF file size, including metadata and alignment. GB is decimal. PPL and all KLD values are lower-is-better; same top-p is higher-is-better.
| Series | Recipe | Actual BPW | GB | PPL | Mean KLD | p999 KLD | p99 KLD | Max KLD | RMS delta-p | Same top-p |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| Jianqiao1 | 3.2 | 3.1649 | 10.808 | 7.446454 | 0.146377 | 6.369240 | 1.550732 | 22.504015 | 11.149% | 84.226% |
| Jianqiao1 | 3.6 | 3.5547 | 12.140 | 7.249732 | 0.109670 | 5.453876 | 1.140977 | 23.917309 | 9.526% | 86.368% |
| Jianqiao1 | 3.8 | 3.7586 | 12.836 | 7.097372 | 0.067854 | 4.466393 | 0.607818 | 26.374474 | 7.200% | 89.505% |
| Jianqiao1 | 4.1 | 4.1725 | 14.249 | 7.030544 | 0.044025 | 3.270178 | 0.382479 | 24.835104 | 5.764% | 91.354% |
| Jianqiao1 | 4.3 | 4.3929 | 15.002 | 7.019089 | 0.032922 | 2.589048 | 0.268183 | 20.440199 | 5.013% | 92.337% |
| Jianqiao1 | 4.6 | 4.4631 | 15.242 | 6.974555 | 0.026953 | 2.159290 | 0.241018 | 21.444906 | 4.444% | 93.530% |
| Jianqiao1 | 4.8 | 4.7291 | 16.150 | 6.934147 | 0.023306 | 1.785062 | 0.206606 | 19.002287 | 4.186% | 94.007% |
| Jianqiao1 | 4.9 | 4.8385 | 16.524 | 6.951682 | 0.022715 | 1.951161 | 0.200312 | 22.226179 | 4.073% | 94.172% |
| Jianqiao1 | 5.1 | 5.1102 | 17.452 | 6.921706 | 0.019116 | 1.664910 | 0.155483 | 23.198441 | 3.718% | 94.671% |
| kaitchup | 3.0 | 3.2891 | 11.232 | 7.731805 | 0.172659 | 6.413453 | 1.775455 | 24.286598 | 12.210% | 82.250% |
| kaitchup | 3.25 | 3.5421 | 12.097 | 7.553780 | 0.142644 | 5.771519 | 1.442072 | 23.057383 | 11.101% | 83.508% |
| kaitchup | 3.5 | 3.7739 | 12.888 | 7.410869 | 0.103733 | 4.815393 | 1.014503 | 24.127443 | 9.397% | 85.579% |
| kaitchup | 3.75 | 3.9913 | 13.631 | 7.150180 | 0.070493 | 4.349184 | 0.642432 | 23.034918 | 7.568% | 89.194% |
| kaitchup | 4.0 | 4.2255 | 14.430 | 7.077410 | 0.054240 | 3.549793 | 0.503634 | 23.122169 | 6.619% | 90.449% |
| kaitchup | 4.25 | 4.4838 | 15.312 | 7.030861 | 0.038407 | 2.421159 | 0.334282 | 24.549507 | 5.537% | 91.695% |
| kaitchup | 4.5 | 4.7478 | 16.214 | 6.992102 | 0.027908 | 2.034667 | 0.224925 | 22.857042 | 4.619% | 92.787% |
| kaitchup | 4.75 | 5.0489 | 17.243 | 7.032774 | 0.025216 | 1.862780 | 0.204153 | 22.090981 | 4.351% | 93.166% |
| kaitchup | 5.0 | 5.3188 | 18.164 | 7.033246 | 0.021426 | 1.817995 | 0.192830 | 20.425156 | 4.032% | 94.184% |
The kaitchup rows are comparison measurements only. kaitchup GGUF files are not redistributed in this repository.
Available Models
| File | Actual BPW | Size GB | Size GiB |
|---|---:|---:|---:|
| Qwen3.6-27B-MTP-MoQ-3.2.gguf | 3.1649 | 10.808 | 10.066 |
| Qwen3.6-27B-MTP-MoQ-3.6.gguf | 3.5547 | 12.140 | 11.306 |
| Qwen3.6-27B-MTP-MoQ-3.8.gguf | 3.7586 | 12.836 | 11.955 |
| Qwen3.6-27B-MTP-MoQ-4.1.gguf | 4.1725 | 14.249 | 13.271 |
| Qwen3.6-27B-MTP-MoQ-4.3.gguf | 4.3929 | 15.002 | 13.972 |
| Qwen3.6-27B-MTP-MoQ-4.6.gguf | 4.4631 | 15.242 | 14.195 |
| Qwen3.6-27B-MTP-MoQ-4.8.gguf | 4.7291 | 16.150 | 15.041 |
| Qwen3.6-27B-MTP-MoQ-4.9.gguf | 4.8385 | 16.524 | 15.389 |
| Qwen3.6-27B-MTP-MoQ-5.1.gguf | 5.1102 | 17.452 | 16.253 |
Quantization Approach
The recipes in this repository are projected tensor policies derived from the observable assignments in w-ahmad/Qwen3.5-9B-GGUF-MoQ-MTP, adapted to Qwen3.6-27B MTP.
The process is:
- Read tensor names, shapes, and GGML types from the source Qwen3.5-9B MoQ GGUF files.
- Split tensor names into layer id and tensor suffix.
- Map source and destination layers by normalized relative depth.
- Reuse the source tensor type for matching suffixes, with a suffix-majority fallback where needed.
- Apply an imatrix during quantization.
- Keep normalization and other non-matmul tensors at their selected high precision.
- Preserve MTP and force its eight large projection tensors in
blk.64toQ8_0.
This is a projection of observed MoQ policies, not the original unpublished optimizer.
All files in this repository were produced with the unsloth Qwen3.6-27B imatrix. A customized llama.cpp quantizer with per-tensor type assignments was used; no specific quantization command is required to use the resulting GGUF files.
Tensor Distribution
The counts below are retained from the existing master CSV. The refreshed 3.2–4.1 evaluator CSV does not include GGUF header metadata, so those tensor-composition counts were not recalculated in this update. Every model has 866 tensors in total.
| Recipe | BF16 | F32 | IQ3_XXS | IQ3_S | IQ4_XS | Q2_K | Q3_K | Q4_K | Q5_K | Q8_0 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| 3.2 | 0 | 360 | 256 | 0 | 17 | 224 | 1 | 0 | 0 | 8 |
| 3.6 | 0 | 360 | 128 | 240 | 18 | 112 | 0 | 0 | 0 | 8 |
| 3.8 | 0 | 360 | 64 | 256 | 82 | 96 | 0 | 0 | 0 | 8 |
| 4.1 | 0 | 360 | 64 | 64 | 258 | 96 | 0 | 16 | 0 | 8 |
| 4.3 | 0 | 360 | 16 | 0 | 258 | 96 | 0 | 128 | 0 | 8 |
| 4.6 | 0 | 360 | 0 | 0 | 352 | 0 | 0 | 128 | 18 | 8 |
| 4.8 | 48 | 360 | 0 | 0 | 288 | 0 | 0 | 80 | 82 | 8 |
| 4.9 | 0 | 360 | 0 | 0 | 256 | 0 | 0 | 112 | 130 | 8 |
| 5.1 | 96 | 360 | 0 | 0 | 176 | 0 | 0 | 32 | 194 | 8 |
Evaluation Conditions
The latest comparison used llama.cpp build 10122, commit d67c0b410, with CUDA on an RTX 5090. WikiText-2 wiki.test.raw was evaluated against BF16 logits saved from the original Qwen3.6-27B MTP GGUF. Context length remained fixed at 512; logical batch was 2048, ubatch was 8192, GPU layers were selected automatically with fit enabled, op offload and flash attention were enabled, and 16 CPU threads were used. All GPU evaluations were run strictly one at a time.
The BF16 reference PPL reported by the KLD evaluator was 6.902375.
Existing Performance Measurements
These throughput results were collected separately on an RTX 5090 and are retained as a practical reference. They are not part of the 2026-08-04 KLD comparison.
| Model | pp512 tok/s | tg128 tok/s | pg32768,256 tok/s | MTP p512 prefill | MTP gen128 | MTP p32768 prefill | MTP gen256 |
|---|---:|---:|---:|---:|---:|---:|---:|
| MoQ-4.8 MTP-Q8_0 | 2297.23 | 67.51 | 1793.82 | 1428.80 | 109.60 | 2270.90 | 90.10 |
| MoQ-4.9 MTP-Q8_0 | 2273.29 | 66.83 | 1762.49 | 1459.60 | 102.30 | 2247.00 | 108.80 |
| MoQ-5.1 MTP-Q8_0 | 2203.42 | 61.56 | 1716.57 | 1391.70 | 81.20 | 2209.70 | 87.50 |
| Unsloth Q4_K_M | 2217.93 | 65.52 | 1755.85 | 1265.20 | 94.80 | 2171.10 | 82.00 |
Usage
Use a recent llama.cpp build with Qwen3.6 MTP support. The files support ordinary generation and MTP speculative decoding. Choose a recipe according to available memory and the quality curves above; 4.8 is a strong balance point, while 5.1 provides the best measured overall quality in this set.
License and Acknowledgements
Released under the Apache License 2.0, following the base model license metadata.
Thanks to:
- the Qwen team for Qwen3.6-27B and its MTP architecture;
- w-ahmad for publishing the Qwen3.5-9B MoQ GGUF tensor policies used as the projection reference;
- kaitchup for publishing an independent Qwen3.6-27B MoQ series that made this controlled comparison possible;
- the unsloth team for the Qwen3.6-27B imatrix;
- the llama.cpp project and contributors for GGUF quantization and evaluation tooling.
Run Jianqiao1/Qwen3.6-27B-MTP-MoQ-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models