GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF overview

<div style="border:2px solid currentColor; font family:ui monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;" <div style="border bottom:…

ggufrocmfp4qwen3.6qwopuscodermtpspeculative-decodingvisionmultimodalstrix-haloamdrocmvulkanenlicense:apache-2.0region:us

Runs locally from ~1.72 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
7,314
Likes
6
Pipeline

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwopus3.6-27B-Coder-MTP-ROCmFP4-STRIX-embF16-headQ6.ggufGGUFGGUF15.70 GBDownload
mmproj-F32.ggufGGUFF321.72 GBDownload

Model Details

Model IDplunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF
Authorplunderstruck
Pipeline
Licenseapache-2.0
Base modelJackrong/Qwopus3.6-27B-Coder-MTP
Last modified2026-06-21T04:43:46.000Z

Model README

---

base_model: Jackrong/Qwopus3.6-27B-Coder-MTP

license: apache-2.0

library_name: gguf

tags:

  • gguf
  • rocmfp4
  • qwen3.6
  • qwopus
  • coder
  • mtp
  • speculative-decoding
  • vision
  • multimodal
  • strix-halo
  • amd
  • rocm
  • vulkan

language:

  • en

base_model_relation: quantized

---

<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">

<div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>

<div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">

<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">

▗▇▇▇▇▇▇▇▖

▗█▘▝██████▖

▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅

▟▛ ▗█████████████████▙▖

▄▄▄▄▄▟▛ ▟████████████████████▖

▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘

▗████▖ ▜▖ ▗█▘

▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙

▜█████▙ ▝████████████▛ ▜▙

▜█████▙ ▝██████████▛ ▃ ▜▙

▀█████▙▖ ▝████████▘ ▟█▙ ▀▙

▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█

▟███████▖ ▜███▘ ▗███████████▛

▟█████████▄ ▜▛ ▗███████████▀

▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘

▜██▘ ▗▛ ▟█████▛▘

▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛

▝█▖ ▟█████▛

▝███████▀

</pre>

<div style="flex:0 1 auto; max-width:100%; text-align:center;">

<div style="font-size:23px; font-weight:800; letter-spacing:1px;">QWOPUS3.6-27B-CODER-MTP</div>

<div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">MTP SELF-SPECULATIVE DECODE</span> · <span style="white-space:nowrap;">VISION-CAPABLE</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>

</div>

</div>

<table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">

<tr>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">~4.5 BPW</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">16 GB</div></td>

<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">262 K</div></td>

</tr>

<tr>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">MTP n-max 5</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">VISION</div><div style="font-weight:700;">QWEN3-VL</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>

<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">LICENSE</div><div style="font-weight:700;">APACHE-2.0</div></td>

</tr>

</table>

</div>

<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">

<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>

The custom <code>q4_0_rocmfp4</code> tensor types <b>will not load in stock llama.cpp, LM Studio, or Ollama</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>mtp-rocmfp4-strix</code>.

</div>

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">

<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16" badge — its parser can't read ROCmFP4 and mislabels by the f16 embeddings. These are <b>~4.4–4.5 bpw 4-bit</b> files; pick by filename.

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Output head</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-headQ6.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">16.9 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">Q6_K</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> — best speed/quality balance: f16 embeddings + Q6 output head on the fast single-scale body</td></tr>

</tbody>

</table>

</div>

One file — the best speed/quality balance in ROCmFP4 for the Qwopus coder. It keeps the two quality levers that are actually felt — genuine f16 token embeddings (from BF16) and a Q6_K output head — on the fast single-scale q4_0_rocmfp4_fast body + the preserved MTP head, and ships no imatrix (deliberate — imatrix worsened this coder's code-PPL, see §05). Not the leanest-fastest possible (a Q5-embedding build squeezes out a few more tok/s, at a quality cost you'll notice), and not the most faithful possible (see the Jackrong fidelity link in §05) — it's the point where speed and quality meet best. Repo also bundles chat_template.jinja — froggeric's unified Qwen3.6 template (tool calls + inline <|think_off|> + vision).

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>

Run from the folder holding the .gguf + chat_template.jinja:

env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
  -m Qwopus3.6-27B-Coder-MTP-ROCmFP4-STRIX-embF16-headQ6.gguf \
  --alias qwopus-coder \
  --host 0.0.0.0 \
  --port 8080 \
  -dev Vulkan0 \
  -ngl 999 \
  -fa on \
  -c 262144 \
  -b 2048 \
  -ub 256 \
  -t 16 \
  -tb 16 \
  -ctk f16 \
  -ctv f16 \
  -cpent 256 \
  -ctxcp 32 \
  --cache-reuse 256 \
  --cache-ram 65536 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.0 \
  --spec-type draft-mtp \
  --spec-draft-device Vulkan0 \
  --spec-draft-ngl all \
  --spec-draft-type-k f16 \
  --spec-draft-type-v f16 \
  --spec-draft-n-max 5 \
  --spec-draft-n-min 0 \
  --spec-draft-p-min 0.0 \
  --spec-draft-p-split 0.10 \
  --chat-template-file chat_template.jinja \
  --reasoning-format deepseek \
  --chat-template-kwargs '{"enable_thinking": false, "preserve_thinking": true}' \
  --jinja \
  --parallel 1 \
  --metrics \
  --no-mmap \
  --mmproj mmproj-F32.gguf \
  --image-min-tokens 1024

The last two lines enable vision (Qwen3-VL) — omit them for text-only. The mmproj-F32.gguf projector is bundled in this repo (Qwen3-VL, projection_dim 5120). --image-min-tokens 1024 is required whenever --mmproj is set — fewer image tokens and Qwen-VL misreads fine detail (see §04).

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>

<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan — fastest backend for ROCmFP4 on Strix Halo</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch · CPU threads</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it; drop to <code>q8_0</code>/<code>q4_0</code> to use less memory</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3.6 recommended sampling</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-type draft-mtp · --spec-draft-n-max 5</code></td><td style="border:1px solid currentColor; padding:6px 10px;">built-in MTP head, self-speculative; draft depth 5 (measured optimum)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-draft-device Vulkan0 · -ngl all · type-k/v f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">draft head on Vulkan, fully offloaded, f16 KV</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--chat-template-file chat_template.jinja</code></td><td style="border:1px solid currentColor; padding:6px 10px;">bundled froggeric template (tool calls + think-toggle + vision)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--reasoning-format deepseek + kwargs {enable_thinking:false, preserve_thinking:true}</code></td><td style="border:1px solid currentColor; padding:6px 10px;">thinking-off for agentic use, but keep cross-turn cache (~86% reuse)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply template · single slot · metrics · weights in RAM</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--mmproj &lt;Qwen3-VL projector&gt;</code> &nbsp;<i>(vision · optional)</i></td><td style="border:1px solid currentColor; padding:6px 10px;">enable image input — any Qwen3-VL projector with <code>projection_dim 5120</code> (the Qwen3.6-27B one, f16/f32); see §04</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--image-min-tokens 1024</code> &nbsp;<i>(vision)</i></td><td style="border:1px solid currentColor; padding:6px 10px;"><b>required whenever <code>--mmproj</code> is set</b> — fewer image tokens and Qwen-VL misreads fine detail (e.g. OCR)</td></tr>

</tbody>

</table>

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · CODING AGENT</div>

Run thinking-off — it commits straight to the tool call instead of over-planning in <think>. Naive enable_thinking:false breaks cross-turn prompt-cache reuse (measured 0% → full re-prefill); add preserve_thinking:true and the cache stays (~86% reuse measured):

--reasoning-format deepseek --chat-template-kwargs '{"enable_thinking": false, "preserve_thinking": true}'

OpenCode — via my fork PlunderStruck/opencode (compaction doesn't rewrite the leading prompt → cache survives long sessions). The model must be tool_call: true, or OpenCode won't send tools natively and the model just narrates code instead of calling tools:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "strix": {
      "npm": "@ai-sdk/openai-compatible",
      "options": { "baseURL": "http://<server-ip>:8080/v1" },
      "models": { "qwopus-coder": { "tool_call": true, "limit": { "context": 262144, "output": 65536 } } }
    }
  },
  "model": "strix/qwopus-coder"
}

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · VISION</div>

Qwen3-VL lineage — vision works by adding the bundled mmproj-F32.gguf projector with --mmproj (same LLM GGUF, no separate vision model). It's the Qwen3-VL projector (projection_dim 5120), shipped in this repo:

  --mmproj mmproj-F32.gguf \
  --image-min-tokens 1024     # REQUIRED — Qwen-VL needs >=1024 image tokens or it misreads fine detail

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">

<b>NOTE //</b> thinking model → for one-shot image Q&amp;A use <code>&lt;|think_off|&gt;</code> or allow enough tokens, else the answer can come back empty. With <code>--mmproj</code> loaded the server disables the <code>--cache-reuse</code> feature (it logs <i>"cache_reuse is not supported by multimodal"</i>); whether ordinary cross-turn caching still helps with vision isn't something we've benchmarked.

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · PERFORMANCE &amp; QUALITY</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE · thinking-off</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~34–38 t/s (Vulkan / Strix Halo)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP DRAFT ACCEPTANCE · code</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~0.8</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">BIGCODEBENCH HARD · instruct · pass@1</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">46/148 (31.1%) · thinking-off, greedy</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">non-imatrix (measured better for code)</td></tr>

</tbody>

</table>

</div>

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">

<b>UPSTREAM BENCHMARK //</b> Published by <a href="https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF">Jackrong</a> for the base Qwopus-Coder — <b>NOT re-measured on this ROCmFP4 quant:</b> SWE-bench Verified <b>335/500 = 67.0%</b>, run <b>thinking-off</b> on the Qwopus-3.6-27B-Coder <b>Q5_K_M</b> GGUF. Our quant is a different 4-bit ROCmFP4 build; we have not re-run SWE-bench.

</div>

Why no imatrix (we measured it): a code-weighted importance matrix improved fidelity-to-BF16 (median KL −15%, top-token +0.6 pp) but measurably worsened held-out-code perplexity (+2.6%, significant). For a coder, code-prediction is the task-relevant metric, so we shipped the non-imatrix quant. (On Qwen3-Coder-Next the imatrix was a clean win — it's model-dependent.)

This is the best speed/quality balance in ROCmFP4 — by design, not the absolute fastest. We swept the same rocmfp4 levers we mapped in detail on the base Qwen3.6-27B (embedding precision, output-head precision, fast single-scale vs all-dual-scale body) and the frontier landed in the same place for the coder: an all-dual-scale body trims worst-case token divergences only inside the measurement noise while costing decode speed, and top-token agreement is tied — so greedy output is effectively identical and the fast single-scale body is the right point. A leaner Q5-embedding build is a few tok/s faster but degrades the one quality lever that's actually felt; we keep full f16 embeddings.

So the recipe is the same one the base-model sweep settled on — fast single-scale body + f16 embeddings + Q6 output head — applied here over Jackrong's tuned coder, and shipped non-imatrix (above) because that's what wins on code-prediction. See the base 27B card's §05 for the full lever-by-lever sweep, KL/decode frontier table, and the format-limit discussion; we don't re-print its numbers here.

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">

<b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> Jackrong's higher-bit <a href="https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF"><b>Qwopus3.6-27B-Coder-MTP-GGUF</b></a> (standard K-quants) run on this same fork — roughly <b>~2× lower KL divergence</b> vs BF16, at slower decode, and MTP still works. We optimize for throughput in ROCmFP4; if you want the last bit of fidelity over speed, that's the one to grab.

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · BUILD (REPRODUCIBLE)</div>

# from Jackrong's BF16+MTP GGUF -> ROCmFP4, genuine f16 embeddings, no imatrix
llama-quantize --token-embedding-type f16 \
  Qwopus3.6-27B-Coder-MTP-BF16.gguf \
  Qwopus3.6-27B-Coder-MTP-ROCmFP4-STRIX-embF16.gguf  Q4_0_ROCMFP4_STRIX

# headQ6 variant adds the Q6_K output head
llama-quantize --token-embedding-type f16 --output-tensor-type q6_K \
  Qwopus3.6-27B-Coder-MTP-BF16.gguf \
  Qwopus3.6-27B-Coder-MTP-ROCmFP4-STRIX-embF16-headQ6.gguf  Q4_0_ROCMFP4_STRIX

> Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">07</span> · LINEAGE &amp; CREDITS</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF">Jackrong/Qwopus3.6-27B-Coder-MTP</a> (Apache-2.0) · from Qwen3.6-27B (Qwen)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> (llama.cpp, MIT)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">CHAT TEMPLATE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates">froggeric/Qwen-Fixed-Chat-Templates</a></td></tr>

</tbody>

</table>

</div>

Derivative quantization — verify the base model's license before redistribution / use.

Run plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models