pearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid 99f6e4; border radius: 16px; box shadow: 0 10px 15…
Runs locally from ~8.89 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | pearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF |
|---|---|
| Author | pearsonkyle |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Jackrong/Qwopus3.6-27B-Coder |
| Last modified | 2026-06-15T05:33:25.000Z |
Model README
---
library_name: gguf
base_model:
- Jackrong/Qwopus3.6-27B-Coder
tags:
- gguf
- llama.cpp
- text-generation
- text-generation-inference
- transformers
- quantization
- quantized
- imatrix
- importance-matrix
- mtp
- multi-token-prediction
- speculative-decoding
- low-bit
- 2-bit
- iq2_xs
- iq2_m
- q2_k_s
- qwopus
- qwen3
- 27b
- coder
- tool-use
- function-calling
- long-context
license: apache-2.0
language:
- en
pipeline_tag: text-generation
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">
<div style="background: linear-gradient(135deg, #0d9488 0%, #134e4a 100%); padding: 24px; color: white;">
<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">
<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 Qwopus3.6-27B-Coder · imatrix · 2-bit · MTP · GGUF</h1>
<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix + MTP</span>
</div>
</div>
<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">
<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 8.89 / 9.74 / 9.96 GiB</span>
<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;"> IQ2_XS / IQ2_M / Q2_K_S</span>
<span style="background: #ede9fe; color: #5b21b6; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #ddd6fe;">⚡ MTP bundled (Q8) · 1.26× · 79.9% accept</span>
<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ llama.cpp 32782998</span>
<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 medKLD 0.044 · top_p 86.3%</span>
</div>
<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">
<div style="background: #f0fdfa; border-left: 5px solid #0d9488; padding: 16px; border-radius: 0 8px 8px 0;">
<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #115e59; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">Three aggressively compressed (under 3.2 bits per weight) quantizations of <b>Jackrong/Qwopus3.6-27B-Coder</b>, each calibrated with a <b>hybrid importance matrix</b> from real usage logs + wiki text, and each shipping the model's own <b>Multi-Token-Prediction (MTP) draft head bundled in at Q8_0</b> for built-in speculative decoding. The imatrix spends the 2-bit codebook's precision where the model is most sensitive; the MTP head — kept near-lossless at Q8 while the trunk goes 2-bit — drafts the next token for a <b>~1.26× decode speedup</b> at <b>79.9% acceptance</b>, no separate draft model required. Plain GGUF, no custom runtime.</p>
</div>
<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~5× smaller on disk</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">8.9–10.0 GiB on disk (incl. the bundled MTP head) vs 50.9 GiB for FP16.</span></div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">⚡ 1.26× faster decode</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Built-in MTP speculative decoding: 22.9 vs 18.1 tok/s on Metal (IQ2_M, n-max=1), 79.9% draft acceptance.</span></div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🎯 up to 86.3% top-p</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Top-token agreement with FP16: 86.3% (IQ2_M) / 85.5% (Q2_K_S) — strong at sub-3.2 bpw.</span></div>
</div>
</div>
</div>
🧰 1. Files & comparison
Three imatrix-calibrated quants, each with the MTP head bundled at Q8_0. Plain
Q2_K (no imatrix) is the no-calibration anchor.
FP16 reference: 50.90 GiB (not included; fetch from Jackrong/Qwopus3.6-27B-Coder).
All rows benched on the same external corpus.eval.txt (code+math+tools, ~90k
tokens — see §2.2), same llama.cpp build, ctx=4096, against FP16. KLD is
median. Every file includes the nextn/MTP draft layer (blk.64) at Q8_0.
| File | Quant | Technique | Size (GiB) | BPW | PPL | KLD (median) | same_top_p |
|---|---|---|---:|---:|---:|---:|---:|
| n/a | FP16 | none (reference) | 50.90 | 16.000 | 4.6585 | 0.00000 | 100.00% |
| …-Q2_K-plain-mtp.gguf | Q2_K | plain (no imatrix) | 10.40 | 3.269 | 4.2384 | 0.0935 | 81.58% |
|||||||||
| …-IQ2_XS-imatrix-mtp.gguf | IQ2_XS | hybrid imatrix | 8.89 | 2.794 | 5.6255 | 0.0783 | 82.01% |
| …-IQ2_M-imatrix-mtp.gguf | IQ2_M | hybrid imatrix | 9.74 | 3.062 | 4.6087 | 0.0442 | 86.26% |
| …-Q2_K_S-imatrix-mtp.gguf | Q2_K_S | hybrid imatrix | 9.96 | 3.133 | 4.6583 | 0.0449 | 85.53% |
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 12px; background: #f0fdfa; padding: 16px; margin: 16px 0; color: #115e59; font-size: 13px; line-height: 1.7;">
<b>Headline.</b> <b>IQ2_M</b> and <b>Q2_K_S</b> are the picks: median KLD 0.044 / 0.045 and top_p 86.3% / 85.5%, close to FP16 at ~3 bpw, with a 1.26× MTP speedup on top. <b>IQ2_XS</b> trades ~4 top_p points for the smallest file (8.89 GiB). PPL sits below FP16 on several rows — that's the usual quant-noise-vs-KLD split; read median KLD + top_p (both FP16-anchored) for fidelity.
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #fde68a; border-radius: 12px; background: #fffbeb; padding: 16px; margin: 16px 0; color: #92400e; font-size: 13px; line-height: 1.7;">
<b>⚠️ Caveat.</b> Sub-3.2-bpw quants of a 27B model. Strong for their size, but not a substitute for FP16 / Q4_K_M / Q5_K_M when you have the VRAM. Use them when memory is the binding constraint.
</div>
---
🔬 2. How they were made
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 16px; margin-bottom: 24px;">
<div style="border: 1px solid #99f6e4; border-radius: 12px; overflow: hidden; background: #ffffff;">
<div style="background: linear-gradient(135deg, #0d9488 0%, #115e59 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">🧮 2.1 Hybrid importance matrix</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">At 2-bit the quantizer must decide <i>where</i> to spend its limited precision. An <b>importance matrix</b> measures, per input channel, how much that channel drives each layer's output on a calibration corpus, and tells <code>llama-quantize</code> to preserve the high-impact channels. This release uses a <b>hybrid</b> imatrix blending activation energy <code>E[a²]</code> with weight-column energy <code>‖W[:, c]‖² · E[a²]</code>, collected at ctx=4096. Linear-attention / SSM tensors (this is a Qwen3.6 hybrid architecture) pass through with raw <code>E[a²]</code>. The output is a standard GGUF with no runtime overhead.</p>
</div>
</div>
<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff;">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #5b21b6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">⚡ 2.2 Bundled MTP (multi-token prediction)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">Qwopus3.6 ships a trained <b>MTP draft head</b> (one nextn layer, <code>blk.64</code>) that predicts the next token from the trunk's hidden state. llama.cpp runs it as built-in speculative decoding (<code>--spec-type draft-mtp</code>): the head drafts, the trunk verifies in parallel, and accepted drafts skip a full decode step.</p>
<p style="margin: 0 0 10px 0;">We <b>keep the MTP head near-lossless at Q8_0</b> while the trunk goes 2-bit — the head is tiny relative to the model, and a 2-bit draft head would draft poorly. Measured on Metal (IQ2_M, n-max=1, holdout prompts):</p>
<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">
<thead><tr style="background:#f5f3ff;"><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Config</th><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Decode tok/s</th><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Draft acceptance</th></tr></thead>
<tbody>
<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>MTP on</b> (n-max=1)</td><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>22.9 ± 0.7</b></td><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>79.9%</b></td></tr>
<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">baseline (off)</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">18.1 ± 1.7</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">—</td></tr>
</tbody>
</table>
<p style="margin: 10px 0 0 0;"><b>→ 1.26× speedup</b> on Metal. Qwen3.6 exposes <b>one</b> nextn layer, so <code>--spec-draft-n-max 1</code> is optimal (higher values don't help). GPU bandwidth matters — the upstream Qwen3.6 figure is ~1.66× on an RTX 5090. See <a href="./MTP/README.md"><code>MTP/README.md</code></a> for details.</p>
</div>
</div>
<div style="border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff;">
<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">📚 2.3 Calibration & evaluation data</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">Two disjoint corpora — the eval corpus never appears in calibration, so §1 measures generalization, not fit.</p>
<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">
<thead><tr style="background:#eff6ff;"><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Corpus</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Source</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Used for</th></tr></thead>
<tbody>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Calibration</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">~500k tokens of usage-log text + all of <code>wiki.test.raw</code></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">hybrid imatrix collection</td></tr>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Eval</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">~90k tokens from <a href="https://huggingface.co/datasets/eaddario/imatrix-calibration"><code>eaddario/imatrix-calibration</code></a> (code+math+tools)</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">all numbers in §1</td></tr>
</tbody>
</table>
</div>
</div>
</div>
Toolchain: imatrix calibration orchestrated by quant-tuner; quantization with llama-quantize from llama.cpp pinned to commit 32782998; the vision tower is stripped and the MTP head kept during extraction.
---
🚀 3. Usage
Building llama.cpp from source (GPU)
apt-get update && apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON # -DGGML_CUDA=OFF for CPU/Metal
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server
cp llama.cpp/build/bin/llama-* llama.cpp/
> MTP needs a recent llama.cpp — --spec-type draft-mtp support was merged in 2026-06. Build from current master.
Running the server with MTP speculative decoding
./llama-server \
--model Qwopus3.6-27B-Coder-IQ2_M-imatrix-mtp.gguf \
--ctx-size 16384 \
--n-gpu-layers 999 \
--spec-type draft-mtp \
--spec-draft-n-max 1 \
--flash-attn on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--host 0.0.0.0 --port 1234
Drop --spec-type draft-mtp --spec-draft-n-max 1 to run without MTP. Note:
-np > 1 (parallel slots) is not yet compatible with MTP.
Querying via the OpenAI-compatible API
import json, urllib.request
def ask(content, max_tokens=256):
body = {
"messages": [{"role": "user", "content": content}],
"max_tokens": max_tokens,
# Coder variant emits <think> reasoning. Set enable_thinking False
# (or raise max_tokens) so the answer lands in "content".
"chat_template_kwargs": {"enable_thinking": False},
}
req = urllib.request.Request("http://127.0.0.1:1234/v1/chat/completions",
json.dumps(body).encode(),
{"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"]
print(ask("Write a Python function that reverses a linked list."))
Which file to pick:
IQ2_M-imatrix-mtp(9.74 GiB) — best overall: median KLD 0.044, top_p 86.3%; benchmarked MTP arm (1.26×, 79.9% accept).Q2_K_S-imatrix-mtp(9.96 GiB) — essentially tied (KLD 0.045, top_p 85.5%).IQ2_XS-imatrix-mtp(8.89 GiB) — smallest; trades ~4 top_p points (82.0%) for ~1 GiB.Q2_K-plain-mtp(10.40 GiB) — no-calibration anchor; larger and weaker than the imatrix quants. Not recommended.
---
🪪 4. License & attribution
- Inherits its license from the base model
Jackrong/Qwopus3.6-27B-Coder. Confirm the exact terms and update the frontmatterlicense:before publishing. - Base weights:
Jackrong/Qwopus3.6-27B-Coder(full finetune of Qwen3.6-27B, ships its own MTP head). - Calibration + quantization performed locally with Quant-Tuner; vendored llama.cpp at commit
32782998. - Calibration data (usage logs) scraped using LogMiner.
Run pearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models