plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF overview
<div style="border:2px solid currentColor; font family:ui monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;" <div style="border bottom:…
Runs locally from ~1.72 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF |
|---|---|
| Author | plunderstruck |
| Pipeline | — |
| License | other |
| Base model | unsloth/Qwen3.6-27B-MTP-GGUF |
| Last modified | 2026-06-21T04:43:32.000Z |
Model README
---
base_model: unsloth/Qwen3.6-27B-MTP-GGUF
license: other
library_name: gguf
tags:
- gguf
- rocmfp4
- qwen3.6
- mtp
- speculative-decoding
- strix-halo
- amd
- rocm
- vulkan
base_model_relation: quantized
---
<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">
<div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>
<div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">
<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">
▗▇▇▇▇▇▇▇▖
▗█▘▝██████▖
▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅
▟▛ ▗█████████████████▙▖
▄▄▄▄▄▟▛ ▟████████████████████▖
▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘
▗████▖ ▜▖ ▗█▘
▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙
▜█████▙ ▝████████████▛ ▜▙
▜█████▙ ▝██████████▛ ▃ ▜▙
▀█████▙▖ ▝████████▘ ▟█▙ ▀▙
▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█
▟███████▖ ▜███▘ ▗███████████▛
▟█████████▄ ▜▛ ▗███████████▀
▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘
▜██▘ ▗▛ ▟█████▛▘
▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛
▝█▖ ▟█████▛
▝███████▀
</pre>
<div style="flex:0 1 auto; max-width:100%; text-align:center;">
<div style="font-size:23px; font-weight:800; letter-spacing:1px;">QWEN3.6-27B-MTP</div>
<div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">imatrix + f16 EMBEDDINGS</span> · <span style="white-space:nowrap;">MTP SELF-SPECULATIVE DECODE</span> · <span style="white-space:nowrap;">VISION-CAPABLE</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>
</div>
</div>
<table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
<tr>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">4.82 BPW</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">16.9 GB</div></td>
<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">262 K</div></td>
</tr>
<tr>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">MTP n-max 5</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">VISION</div><div style="font-weight:700;">QWEN3-VL</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>
<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CALIBRATION</div><div style="font-weight:700;">imatrix (CODE)</div></td>
</tr>
</table>
</div>
<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">
<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>
The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, Ollama, Jan, or koboldcpp</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>mtp-rocmfp4-strix</code>:
<br><br>
<code>git clone https://github.com/charlie12345/ROCmFPX</code><br>
<code>cd ROCmFPX</code><br>
<code>env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh</code>
</div>
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">
<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16" / 16-bit badge — its parser only knows standard GGUF quant types, can't read ROCmFP4, and "sees" only the genuinely-f16 token embeddings. These are <b>~4.8 bpw 4-bit</b> files; pick by filename in <i>Files and versions</i>.
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Output head</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-imatrix-embF16-headQ6.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">16.9 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">Q6_K</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> — best speed/quality balance: f16 embeddings + Q6 output head on the fast single-scale body</td></tr>
</tbody>
</table>
</div>
One file — the best speed/quality balance in ROCmFP4 for Strix Halo. It keeps the two quality levers that are actually felt — genuine f16 token embeddings (from BF16) and a Q6_K output head — on the fast single-scale q4_0_rocmfp4_fast body + a code-calibrated imatrix + the MTP head. Not the leanest-fastest possible (a Q5-embedding build squeezes out a few more tok/s, at a quality cost you'll notice), and not the most faithful possible (see the Unsloth fidelity link in §05) — it's the point where speed and quality meet best. Repo also bundles the mmproj-F32.gguf Qwen3-VL vision projector, chat_template.jinja (froggeric's unified Qwen3.6 template — tool calls + inline <|think_off|>/<|think_on|> + vision), and the qwen3.6-27b-code.imatrix (339 chunks) for exact reproduction.
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">token_embd</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">F16 (full precision — a lookup, ~zero decode cost)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">attention K/V (+ fused QKV)</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;"><code>q4_0_rocmfp4</code> (dual-scale)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">FFN, lm-head, rest</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;"><code>q4_0_rocmfp4_fast</code> (single-scale)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP head</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">preserved (<code>blk.64.nextn.*</code>)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">imatrix</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">code-calibrated, applied to all 496 quantizable tensors</td></tr>
</tbody>
</table>
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>
Run from the folder holding the .gguf + chat_template.jinja:
env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
-m Qwen3.6-27B-MTP-ROCmFP4-STRIX-imatrix-embF16-headQ6.gguf \
--alias qwen27b-mtp \
--host 0.0.0.0 \
--port 8080 \
-dev Vulkan0 \
-ngl 999 \
-fa on \
-c 262144 \
-b 2048 \
-ub 256 \
-t 16 \
-tb 16 \
-ctk f16 \
-ctv f16 \
-cpent 256 \
-ctxcp 32 \
--cache-reuse 256 \
--cache-ram 65536 \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--spec-type draft-mtp \
--spec-draft-device Vulkan0 \
--spec-draft-ngl all \
--spec-draft-type-k f16 \
--spec-draft-type-v f16 \
--spec-draft-n-max 5 \
--spec-draft-n-min 0 \
--spec-draft-p-min 0.0 \
--spec-draft-p-split 0.10 \
--chat-template-file chat_template.jinja \
--reasoning on \
--reasoning-format deepseek \
--chat-template-kwargs '{"preserve_thinking": true}' \
--jinja \
--parallel 1 \
--metrics \
--no-mmap \
--mmproj mmproj-F32.gguf \
--image-min-tokens 1024
The last two lines enable vision — the mmproj-F32.gguf Qwen3-VL projector is bundled in this repo (projection_dim 5120); omit them for text-only. --image-min-tokens 1024 is required whenever --mmproj is set (see §04).
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan (KHR_coopmat) — beats ROCm/HIP here, ~+1.7× prefill</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch (256 is the optimum here — bigger ubatch is <i>slower</i>) · CPU threads</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it; drop to <code>q8_0</code>/<code>q4_0</code> to use less memory</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing (every 256 tok, keep 32, reuse ≥256-tok prefix) + 64 GB resident reuse cache</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3.6 "precise coding" sampling (temp 1.0 for general tasks)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-type draft-mtp · --spec-draft-n-max 5</code></td><td style="border:1px solid currentColor; padding:6px 10px;">built-in MTP head, self-speculative; draft depth 5</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-draft-device Vulkan0 · -ngl all · type-k/v f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">draft head on Vulkan, fully offloaded, f16 draft KV</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--chat-template-file chat_template.jinja</code></td><td style="border:1px solid currentColor; padding:6px 10px;">bundled froggeric template (tool calls + think-toggle + vision)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--reasoning on --reasoning-format deepseek + kwargs {preserve_thinking:true}</code></td><td style="border:1px solid currentColor; padding:6px 10px;">clean <code>content</code> + <code>reasoning_content</code>; keep <code><think></code> across turns so cross-turn cache survives</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply template · single slot · metrics · weights in RAM</td></tr>
</tbody>
</table>
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · CODING AGENT / OPENCODE</div>
Multi-turn prompt-cache reuse is what makes this usable. Qwen3.6's recurrent (SSM) state can't be partially rewound, so multi-turn reuse needs a context checkpoint at/before the divergence point. Two defaults otherwise force a full re-prefill every turn — both fixed by the flags above:
- Checkpoint cadence. Default
-cpentis 8192, so prompts under 8K never get a usable checkpoint. Fix:-cpent 256 -ctxcp 32 --cache-reuse 256(checkpoint every 256 tokens, keep 32, reuse a matching prefix of ≥256 tokens). Verified: a shared 3,000-token prefix re-prefill dropped 12.4 s → ~0.1 s. - Thinking text breaking the prefix match.
--reasoning-formatcontrols where<think>goes.deepseek(used here) gives cleancontent+reasoning_content, auto-paired with--chat-template-kwargs '{"preserve_thinking": true}'so the template keeps<think>for all turns and reuse holds (with OpenCode the large stable leading context reuses via checkpoints regardless).noneleaves<think>inline incontentso any content-echoing client gets reuse;deepseek-legacy/autodo not reuse. - Vision +
--cache-reuse. With--mmprojloaded the server disables the--cache-reusefeature (it logs "cache_reuse is not supported by multimodal"); we haven't measured whether ordinary cross-turn caching survives with vision (see §04).
--jinja is required so the chat template (and preserve_thinking) apply.
OpenCode — point it at the server as an OpenAI-compatible provider. In single-model mode llama-server ignores the request's model field, so the client's model name is just a label (it does not have to match --alias). The provider below is named lmstudio only because it uses the generic OpenAI-compatible adapter — it points at this llama-server, not LM Studio:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"lmstudio": {
"npm": "@ai-sdk/openai-compatible",
"name": "local llama-server (ROCmFP4)",
"options": { "baseURL": "http://<host>:8080/v1", "apiKey": "sk-local" },
"models": {
"qwen3.6-27b-mtp": {
"name": "Qwen 3.6 27B",
"limit": { "context": 262144, "output": 32768 }
}
}
}
},
"model": "lmstudio/qwen3.6-27b-mtp",
"compaction": { "auto": true, "reserved": 16384 }
}
Project-local opencode.json — disable the task tool so agents don't spawn subagents, keeping the whole session on one cache-friendly context:
{
"$schema": "https://opencode.ai/config.json",
"agent": {
"build": { "tools": { "task": false } },
"plan": { "tools": { "task": false } }
}
}
The fork: PlunderStruck/opencode. compaction.auto summarizes history when the context fills — which in stock OpenCode rewrites the leading prompt and invalidates the cache, forcing a full re-prefill. This fork compacts without breaking the cached prefix (plus a few other adjustments), so cache reuse survives compaction. Paired with the checkpoint flags above, long sessions stay fast and actually usable.
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · VISION</div>
Qwen3-VL lineage — vision works via the bundled mmproj-F32.gguf projector at launch with --mmproj (no different LLM GGUF needed). It's the Qwen3-VL projector (projection_dim 5120, matches this model's hidden size), shipped in this repo.
# add to your llama-server launch:
--mmproj mmproj-F32.gguf \
--image-min-tokens 1024 # REQUIRED — Qwen-VL needs >=1024 image tokens or it misreads fine detail
Without --image-min-tokens 1024 the server feeds too few image tokens and the model describes images incorrectly (right gist, wrong detail) — the server even logs a warning at load. Verified: a code label misread at default tokens read correctly once the flag was set.
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE //</b> thinking model → for one-shot image Q&A use the bundled template's inline <code><|think_off|></code> or allow enough tokens to finish <code><think></code>, else the visible answer can come back empty. With <code>--mmproj</code> loaded the server disables the <code>--cache-reuse</code> feature (it logs <i>"cache_reuse is not supported by multimodal"</i>); whether ordinary cross-turn caching still helps with vision isn't something we've benchmarked.
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · PERFORMANCE & QUALITY</div>
This is the best speed/quality balance in ROCmFP4 — by design, not the absolute fastest. We tested the alternatives within rocmfp4: an all-dual-scale body and selective higher-precision tensors both cost decode speed for a KL improvement that sat inside the measurement noise — so the fast single-scale body + f16 embeddings + Q6 head is the right point. A leaner Q5-embedding build (what some others ship) is a few tok/s faster but degrades the one quality lever that's actually felt; we keep full f16 embeddings.
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">
<b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> Unsloth's <a href="https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF"><b>UD-Q4_K_XL</b></a> (a dynamic K-quant) runs on this same fork — measured <b>~2× lower KL divergence</b> vs BF16, at <b>~14% slower</b> decode, and <b>MTP still works</b> (it carries the nextn head). We optimize for throughput in ROCmFP4; if you want the last bit of fidelity over speed, that's the one to grab.
</div>
How we landed on this recipe — the full sweep. We measured every rocmfp4 lever against the BF16 reference by KL divergence (the right fidelity metric) plus decode speed (llama-bench), and compared to the best stock 4-bit. The frontier on this model:
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Build</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Decode t/s ↑</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Mean KLD vs BF16 ↓</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Verdict</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><b>STRIX</b> — fast body, f16 emb, Q6 head ★</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>13.66</b></td><td style="border:1px solid currentColor; padding:7px 10px;">0.126</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>this file — the balance point</b></td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">COHERENT — all-dual-scale body</td><td style="border:1px solid currentColor; padding:7px 10px;">13.19</td><td style="border:1px solid currentColor; padding:7px 10px;">0.119</td><td style="border:1px solid currentColor; padding:7px 10px;">slower; KL gain inside the noise → dropped</td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">DYN-max — Unsloth-style selective Q5/Q6 bumps on rocmfp4</td><td style="border:1px solid currentColor; padding:7px 10px;">11.59</td><td style="border:1px solid currentColor; padding:7px 10px;">0.090</td><td style="border:1px solid currentColor; padding:7px 10px;">better KL but slower <i>and</i> bigger than Q4_K_XL → dominated</td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">stock Unsloth UD-Q4_K_XL — dynamic K-quant (<i>not</i> rocmfp4)</td><td style="border:1px solid currentColor; padding:7px 10px;">11.99</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.064</b></td><td style="border:1px solid currentColor; padding:7px 10px;">most faithful — the fidelity pick (link above)</td></tr>
</tbody>
</table>
</div>
What the sweep showed:
- All-dual-scale (COHERENT) and selective higher-precision (DYN) both trade speed for a KL gain that sits inside the measurement noise — neither beats the fast body for the balance, so we dropped them.
- *Even copying Unsloth's entire high-precision allocation onto rocmfp4 (DYN-max) stalled at 0.090 KLD — still well above UD-Q4_K_XL's 0.064, and slower and bigger. That's a format limit: rocmfp4's FP4 is intrinsically less faithful than Q4_K's 4-bit, so you can't out-allocate it without simply becoming* Q4_K_XL.
- Conclusion: within rocmfp4, fast body + f16 embeddings + Q6 head is the optimal balance (this file). For maximum fidelity, the dynamic K-quant wins — so we link it rather than ship a worse copy. (Directional internal measurements — KL vs BF16 on held-out text + llama-bench decode; reproduce before citing.)
Hands-on observations from daily use on a Framework Desktop / AMD Ryzen AI Max+ 395 (gfx1151, 128 GB unified, ROCm 7.2.0) — directional internal checks, not formal benchmarks; reproduce before citing.
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE · short context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~33 t/s (Vulkan / Strix Halo)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">DECODE · ~140K context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~18 t/s</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP DRAFT ACCEPTANCE · warm, f16 KV</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~0.87–0.90</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">ARCHITECTURE</td><td style="border:1px solid currentColor; padding:8px 11px;">hybrid SSM + attention (48 SSM + 17 attention blocks) — only attention layers grow a KV cache, so it degrades gracefully at long context</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">code-calibrated imatrix + f16 embeddings (measured small win on code)</td></tr>
</tbody>
</table>
</div>
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">
<b>UPSTREAM BENCHMARK //</b> Base is Qwen's <b>Qwen3.6-27B</b> (via <a href="https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF">unsloth</a>). Qwen publishes official benchmarks as a figure on the base card — see there. <b>NOT re-measured on this ROCmFP4 quant.</b>
</div>
f16 KV is a config choice — full-precision KV is how we run it; 128 GB unified RAM affords it. On less memory drop to -ctk q8_0 -ctv q8_0.
f16 token embeddings were the single change felt the most: raising the token-embedding layer to full precision made the model follow instructions noticeably better — the embedding is the foundation every layer builds on, and the vocab is large, so a faithful embedding pays off at near-zero speed cost (it's a lookup, not a matmul). The code-calibrated imatrix is a free polish on top (same size and speed) — small, in the right direction on code:
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Test set (n_ctx=512)</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">no-imatrix</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">this (imatrix)</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">held-out code</td><td style="border:1px solid currentColor; padding:7px 10px;">1.8631</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">1.8596</td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">held-out prose</td><td style="border:1px solid currentColor; padding:7px 10px;">5.7109</td><td style="border:1px solid currentColor; padding:7px 10px;">5.7165</td></tr>
</tbody>
</table>
</div>
Tiny improvement on code (the calibration domain), neutral on prose — expected at this bit rate; at 4+ bpw the base quant is already close to the original, so imatrix is a polish, not a transformation.
The Q6-head variant — a step up (experimental). It raises the output head (output.weight) from 4-bit ROCmFP4 to standard Q6_K and leaves everything else untouched. The embedding is the input side; the output head is the output side — sharpening both beats sharpening either. Observed: a further step up in instruction-following beyond the f16 embeddings (reaching for the specific tool asked for, sticking to task rules/format more reliably). Two held-out measurements:
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Test set</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">daily (4-bit head)</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Q6 head</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">held-out code (perplexity)</td><td style="border:1px solid currentColor; padding:7px 10px;">1.8596</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">1.8550</td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">held-out prose (perplexity)</td><td style="border:1px solid currentColor; padding:7px 10px;">5.7165</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">5.6761</td></tr>
<tr><td style="border:1px solid currentColor; padding:7px 10px;">KL vs BF16 (mean, lower=more faithful)</td><td style="border:1px solid currentColor; padding:7px 10px;">≈0.0369</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">≈0.0345 (~6% nearer)</td></tr>
</tbody>
</table>
</div>
The Q6 head improved both code and prose perplexity (the imatrix alone only helped code) and was closer to BF16 on every measure. It still agrees with BF16's top word ~96% of the time either way — so the head mostly sharpens confidence on the same choice rather than flipping it. The cost: decode is ~5–7% slower at short context (the head is a fixed per-token cost, so the gap shrinks at long context); size grows ~0.4 GB. Small but consistent gains across two tests and two text types — internal checks, not formal benchmarks; reproduce before citing.
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · BUILD (REPRODUCIBLE)</div>
Calibration corpus (code_calibration.txt): a concatenation of three files from the froggeric/imatrix dataset — groups_merged.txt + code.txt + technical.txt (~646 KB total) — code-heavy but diverse enough to avoid domain overfitting. The resulting imatrix (qwen3.6-27b-code.imatrix, 339 chunks) is included in this repo.
# 1) importance matrix
llama-imatrix -m Qwen3.6-27B-BF16-00001-of-00002.gguf \
-f code_calibration.txt -o qwen3.6-27b-code.imatrix \
-dev Vulkan0 -ngl 999 -fa on -c 512
# 2) quantize: quality-biased STRIX preset + f16 embeddings + imatrix (daily driver)
llama-quantize \
--imatrix qwen3.6-27b-code.imatrix \
--token-embedding-type f16 \
Qwen3.6-27B-BF16-00001-of-00002.gguf \
Qwen3.6-27B-MTP-ROCmFP4-STRIX-imatrix-embF16.gguf \
Q4_0_ROCMFP4_STRIX
# 3) headQ6 variant — same as above + one extra flag (--output-tensor-type q6_K)
llama-quantize \
--imatrix qwen3.6-27b-code.imatrix \
--token-embedding-type f16 \
--output-tensor-type q6_K \
Qwen3.6-27B-BF16-00001-of-00002.gguf \
Qwen3.6-27B-MTP-ROCmFP4-STRIX-imatrix-embF16-headQ6.gguf \
Q4_0_ROCMFP4_STRIX
> Experimental research build for AMD Strix Halo — hardware/driver/model/prompt-sensitive, may not reproduce on other GPUs. Not native FP4 tensor-core execution. Do not treat these numbers as upstream llama.cpp claims.
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">07</span> · LINEAGE & CREDITS</div>
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;">Qwen3.6-27B (Qwen team) — dense, with built-in MTP head (<code>nextn_predict_layers=1</code>, so draft-MTP survives quantization). Derivative quant <b>inherits the base model's license</b>.</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">BF16 GGUF SOURCE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF">unsloth/Qwen3.6-27B-MTP-GGUF</a> @ <code>5cb35eb3dcbf52dbce5f87dbc64df6aaffadcace</code></td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> (based on llama.cpp, MIT)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CHAT TEMPLATE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates">froggeric/Qwen-Fixed-Chat-Templates</a></td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CALIBRATION DATA</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a></td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">AGENT FORK</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/PlunderStruck/opencode">PlunderStruck/opencode</a></td></tr>
</tbody>
</table>
</div>
Derivative quantization — verify the base model's (Qwen3.6) license before redistribution / use.
Run plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models