GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

pearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid 99f6e4; border radius: 16px; box shadow: 0 10px 15…

ggufllama.cpptext-generationtext-generation-inferencetransformersquantizationquantizedimatriximportance-matrixmtpmulti-token-predictionspeculative-decodinglow-bit2-bitiq2_xsiq2_mq2_k_sqwopusqwen327bcodertool-usefunction-callinglong-context

Runs locally from ~8.89 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwopus3.6-27B-Coder-IQ2_M-imatrix-mtp.ggufGGUFIQ2_M9.74 GBDownload
Qwopus3.6-27B-Coder-IQ2_XS-imatrix-mtp.ggufGGUFIQ2_XS8.89 GBDownload
Qwopus3.6-27B-Coder-Q2_K-plain-mtp.ggufGGUFQ2_K10.40 GBDownload
Qwopus3.6-27B-Coder-Q2_K_S-imatrix-mtp.ggufGGUFQ2_K_S9.96 GBDownload

Model Details

Model IDpearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF
Authorpearsonkyle
Pipelinetext-generation
Licenseapache-2.0
Base modelJackrong/Qwopus3.6-27B-Coder
Last modified2026-06-15T05:33:25.000Z

Model README

---

library_name: gguf

base_model:

  • Jackrong/Qwopus3.6-27B-Coder

tags:

  • gguf
  • llama.cpp
  • text-generation
  • text-generation-inference
  • transformers
  • quantization
  • quantized
  • imatrix
  • importance-matrix
  • mtp
  • multi-token-prediction
  • speculative-decoding
  • low-bit
  • 2-bit
  • iq2_xs
  • iq2_m
  • q2_k_s
  • qwopus
  • qwen3
  • 27b
  • coder
  • tool-use
  • function-calling
  • long-context

license: apache-2.0

language:

  • en

pipeline_tag: text-generation

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">

<div style="background: linear-gradient(135deg, #0d9488 0%, #134e4a 100%); padding: 24px; color: white;">

<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">

<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 Qwopus3.6-27B-Coder · imatrix · 2-bit · MTP · GGUF</h1>

<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix + MTP</span>

</div>

</div>

<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">

<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 8.89 / 9.74 / 9.96 GiB</span>

<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;"> IQ2_XS / IQ2_M / Q2_K_S</span>

<span style="background: #ede9fe; color: #5b21b6; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #ddd6fe;">⚡ MTP bundled (Q8) · 1.26× · 79.9% accept</span>

<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ llama.cpp 32782998</span>

<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 medKLD 0.044 · top_p 86.3%</span>

</div>

<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">

<div style="background: #f0fdfa; border-left: 5px solid #0d9488; padding: 16px; border-radius: 0 8px 8px 0;">

<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #115e59; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">Three aggressively compressed (under 3.2 bits per weight) quantizations of <b>Jackrong/Qwopus3.6-27B-Coder</b>, each calibrated with a <b>hybrid importance matrix</b> from real usage logs + wiki text, and each shipping the model's own <b>Multi-Token-Prediction (MTP) draft head bundled in at Q8_0</b> for built-in speculative decoding. The imatrix spends the 2-bit codebook's precision where the model is most sensitive; the MTP head — kept near-lossless at Q8 while the trunk goes 2-bit — drafts the next token for a <b>~1.26× decode speedup</b> at <b>79.9% acceptance</b>, no separate draft model required. Plain GGUF, no custom runtime.</p>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~5× smaller on disk</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">8.9–10.0 GiB on disk (incl. the bundled MTP head) vs 50.9 GiB for FP16.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">⚡ 1.26× faster decode</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Built-in MTP speculative decoding: 22.9 vs 18.1 tok/s on Metal (IQ2_M, n-max=1), 79.9% draft acceptance.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #115e59; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🎯 up to 86.3% top-p</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Top-token agreement with FP16: 86.3% (IQ2_M) / 85.5% (Q2_K_S) — strong at sub-3.2 bpw.</span></div>

</div>

</div>

</div>

🧰 1. Files & comparison

Three imatrix-calibrated quants, each with the MTP head bundled at Q8_0. Plain

Q2_K (no imatrix) is the no-calibration anchor.

FP16 reference: 50.90 GiB (not included; fetch from Jackrong/Qwopus3.6-27B-Coder).

All rows benched on the same external corpus.eval.txt (code+math+tools, ~90k

tokens — see §2.2), same llama.cpp build, ctx=4096, against FP16. KLD is

median. Every file includes the nextn/MTP draft layer (blk.64) at Q8_0.

| File | Quant | Technique | Size (GiB) | BPW | PPL | KLD (median) | same_top_p |

|---|---|---|---:|---:|---:|---:|---:|

| n/a | FP16 | none (reference) | 50.90 | 16.000 | 4.6585 | 0.00000 | 100.00% |

| …-Q2_K-plain-mtp.gguf | Q2_K | plain (no imatrix) | 10.40 | 3.269 | 4.2384 | 0.0935 | 81.58% |

|||||||||

| …-IQ2_XS-imatrix-mtp.gguf | IQ2_XS | hybrid imatrix | 8.89 | 2.794 | 5.6255 | 0.0783 | 82.01% |

| …-IQ2_M-imatrix-mtp.gguf | IQ2_M | hybrid imatrix | 9.74 | 3.062 | 4.6087 | 0.0442 | 86.26% |

| …-Q2_K_S-imatrix-mtp.gguf | Q2_K_S | hybrid imatrix | 9.96 | 3.133 | 4.6583 | 0.0449 | 85.53% |

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #99f6e4; border-radius: 12px; background: #f0fdfa; padding: 16px; margin: 16px 0; color: #115e59; font-size: 13px; line-height: 1.7;">

<b>Headline.</b> <b>IQ2_M</b> and <b>Q2_K_S</b> are the picks: median KLD 0.044 / 0.045 and top_p 86.3% / 85.5%, close to FP16 at ~3 bpw, with a 1.26× MTP speedup on top. <b>IQ2_XS</b> trades ~4 top_p points for the smallest file (8.89 GiB). PPL sits below FP16 on several rows — that's the usual quant-noise-vs-KLD split; read median KLD + top_p (both FP16-anchored) for fidelity.

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #fde68a; border-radius: 12px; background: #fffbeb; padding: 16px; margin: 16px 0; color: #92400e; font-size: 13px; line-height: 1.7;">

<b>⚠️ Caveat.</b> Sub-3.2-bpw quants of a 27B model. Strong for their size, but not a substitute for FP16 / Q4_K_M / Q5_K_M when you have the VRAM. Use them when memory is the binding constraint.

</div>

---

🔬 2. How they were made

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 16px; margin-bottom: 24px;">

<div style="border: 1px solid #99f6e4; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: linear-gradient(135deg, #0d9488 0%, #115e59 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">🧮 2.1 Hybrid importance matrix</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">At 2-bit the quantizer must decide <i>where</i> to spend its limited precision. An <b>importance matrix</b> measures, per input channel, how much that channel drives each layer's output on a calibration corpus, and tells <code>llama-quantize</code> to preserve the high-impact channels. This release uses a <b>hybrid</b> imatrix blending activation energy <code>E[a²]</code> with weight-column energy <code>‖W[:, c]‖² · E[a²]</code>, collected at ctx=4096. Linear-attention / SSM tensors (this is a Qwen3.6 hybrid architecture) pass through with raw <code>E[a²]</code>. The output is a standard GGUF with no runtime overhead.</p>

</div>

</div>

<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: linear-gradient(135deg, #7c3aed 0%, #5b21b6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">⚡ 2.2 Bundled MTP (multi-token prediction)</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">Qwopus3.6 ships a trained <b>MTP draft head</b> (one nextn layer, <code>blk.64</code>) that predicts the next token from the trunk's hidden state. llama.cpp runs it as built-in speculative decoding (<code>--spec-type draft-mtp</code>): the head drafts, the trunk verifies in parallel, and accepted drafts skip a full decode step.</p>

<p style="margin: 0 0 10px 0;">We <b>keep the MTP head near-lossless at Q8_0</b> while the trunk goes 2-bit — the head is tiny relative to the model, and a 2-bit draft head would draft poorly. Measured on Metal (IQ2_M, n-max=1, holdout prompts):</p>

<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">

<thead><tr style="background:#f5f3ff;"><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Config</th><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Decode tok/s</th><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Draft acceptance</th></tr></thead>

<tbody>

<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>MTP on</b> (n-max=1)</td><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>22.9 ± 0.7</b></td><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>79.9%</b></td></tr>

<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">baseline (off)</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">18.1 ± 1.7</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">—</td></tr>

</tbody>

</table>

<p style="margin: 10px 0 0 0;"><b>→ 1.26× speedup</b> on Metal. Qwen3.6 exposes <b>one</b> nextn layer, so <code>--spec-draft-n-max 1</code> is optimal (higher values don't help). GPU bandwidth matters — the upstream Qwen3.6 figure is ~1.66× on an RTX 5090. See <a href="./MTP/README.md"><code>MTP/README.md</code></a> for details.</p>

</div>

</div>

<div style="border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">📚 2.3 Calibration & evaluation data</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">Two disjoint corpora — the eval corpus never appears in calibration, so §1 measures generalization, not fit.</p>

<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">

<thead><tr style="background:#eff6ff;"><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Corpus</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Source</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Used for</th></tr></thead>

<tbody>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Calibration</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">~500k tokens of usage-log text + all of <code>wiki.test.raw</code></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">hybrid imatrix collection</td></tr>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Eval</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">~90k tokens from <a href="https://huggingface.co/datasets/eaddario/imatrix-calibration"><code>eaddario/imatrix-calibration</code></a> (code+math+tools)</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">all numbers in §1</td></tr>

</tbody>

</table>

</div>

</div>

</div>

Toolchain: imatrix calibration orchestrated by quant-tuner; quantization with llama-quantize from llama.cpp pinned to commit 32782998; the vision tower is stripped and the MTP head kept during extraction.

---

🚀 3. Usage

Building llama.cpp from source (GPU)

apt-get update && apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON   # -DGGML_CUDA=OFF for CPU/Metal
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server
cp llama.cpp/build/bin/llama-* llama.cpp/

> MTP needs a recent llama.cpp--spec-type draft-mtp support was merged in 2026-06. Build from current master.

Running the server with MTP speculative decoding

 ./llama-server \
    --model Qwopus3.6-27B-Coder-IQ2_M-imatrix-mtp.gguf \
    --ctx-size 16384 \
    --n-gpu-layers 999 \
    --spec-type draft-mtp \
    --spec-draft-n-max 1 \
    --flash-attn on \
    --cache-type-k q8_0 --cache-type-v q8_0 \
    --host 0.0.0.0 --port 1234

Drop --spec-type draft-mtp --spec-draft-n-max 1 to run without MTP. Note:

-np > 1 (parallel slots) is not yet compatible with MTP.

Querying via the OpenAI-compatible API

import json, urllib.request

def ask(content, max_tokens=256):
    body = {
        "messages": [{"role": "user", "content": content}],
        "max_tokens": max_tokens,
        # Coder variant emits <think> reasoning. Set enable_thinking False
        # (or raise max_tokens) so the answer lands in "content".
        "chat_template_kwargs": {"enable_thinking": False},
    }
    req = urllib.request.Request("http://127.0.0.1:1234/v1/chat/completions",
                                 json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"]

print(ask("Write a Python function that reverses a linked list."))

Which file to pick:

  • IQ2_M-imatrix-mtp (9.74 GiB) — best overall: median KLD 0.044, top_p 86.3%; benchmarked MTP arm (1.26×, 79.9% accept).
  • Q2_K_S-imatrix-mtp (9.96 GiB) — essentially tied (KLD 0.045, top_p 85.5%).
  • IQ2_XS-imatrix-mtp (8.89 GiB) — smallest; trades ~4 top_p points (82.0%) for ~1 GiB.
  • Q2_K-plain-mtp (10.40 GiB) — no-calibration anchor; larger and weaker than the imatrix quants. Not recommended.

---

🪪 4. License & attribution

  • Inherits its license from the base model Jackrong/Qwopus3.6-27B-Coder. Confirm the exact terms and update the frontmatter license: before publishing.
  • Base weights: Jackrong/Qwopus3.6-27B-Coder (full finetune of Qwen3.6-27B, ships its own MTP head).
  • Calibration + quantization performed locally with Quant-Tuner; vendored llama.cpp at commit 32782998.
  • Calibration data (usage logs) scraped using LogMiner.

Run pearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models