GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF overview

<div style="border:2px solid currentColor; font family:ui monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;" <div style="border bottom:…

ggufrocmfp4qwen3.6moemtpspeculative-decodingstrix-haloamdrocmvulkanbase_model:unsloth/Qwen3.6-35B-A3B-MTP-GGUFbase_model:quantized:unsloth/Qwen3.6-35B-A3B-MTP-GGUFlicense:otherregion:us

Runs locally from ~1.66 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
4,954
Likes
4
Pipeline

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-35B-A3B-MTP-ROCmFP4-STRIX-embF16-headQ6.ggufGGUFGGUF18.50 GBDownload
mmproj-F32.ggufGGUFF321.66 GBDownload

Model Details

Model IDplunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF
Authorplunderstruck
Pipeline
Licenseother
Base modelunsloth/Qwen3.6-35B-A3B-MTP-GGUF
Last modified2026-06-21T04:43:37.000Z

Model README

---

base_model: unsloth/Qwen3.6-35B-A3B-MTP-GGUF

license: other

library_name: gguf

tags:

  • gguf
  • rocmfp4
  • qwen3.6
  • moe
  • mtp
  • speculative-decoding
  • strix-halo
  • amd
  • rocm
  • vulkan

base_model_relation: quantized

---

<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">

<div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>

<div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">

<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">

▗▇▇▇▇▇▇▇▖

▗█▘▝██████▖

▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅

▟▛ ▗█████████████████▙▖

▄▄▄▄▄▟▛ ▟████████████████████▖

▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘

▗████▖ ▜▖ ▗█▘

▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙

▜█████▙ ▝████████████▛ ▜▙

▜█████▙ ▝██████████▛ ▃ ▜▙

▀█████▙▖ ▝████████▘ ▟█▙ ▀▙

▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█

▟███████▖ ▜███▘ ▗███████████▛

▟█████████▄ ▜▛ ▗███████████▀

▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘

▜██▘ ▗▛ ▟█████▛▘

▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛

▝█▖ ▟█████▛

▝███████▀

</pre>

<div style="flex:0 1 auto; max-width:100%; text-align:center;">

<div style="font-size:23px; font-weight:800; letter-spacing:1px;">QWEN3.6-35B-A3B-MTP</div>

<div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">MIXTURE-OF-EXPERTS (A3B)</span> · <span style="white-space:nowrap;">MTP SELF-SPECULATIVE DECODE</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>

</div>

</div>

<table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">

<tr>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">4.44 BPW</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">18.8 GB</div></td>

<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">262 K</div></td>

</tr>

<tr>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">ARCH</div><div style="font-weight:700;">MoE · 256 EXPERTS</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">ACTIVE / HIDDEN</div><div style="font-weight:700;">~3B · 2048</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">MTP n-max 5</div></td>

<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>

</tr>

</table>

</div>

<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">

<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>

The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, Ollama, Jan, or koboldcpp</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>mtp-rocmfp4-strix</code>.

</div>

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">

<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16" badge — its parser can't read ROCmFP4 and mislabels by the genuinely-f16 token embeddings. These are <b>~4.44 bpw 4-bit</b> files; pick by filename.

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Output head</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-headQ6.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">18.8 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">Q6_K</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> — best speed/quality balance: f16 embeddings + Q6 output head on the fast single-scale body</td></tr>

</tbody>

</table>

</div>

One file — the best speed/quality balance in ROCmFP4 for Strix Halo. It keeps the two quality levers that are actually felt — genuine f16 token embeddings (from BF16) and a Q6_K output head — on the fast single-scale q4_0_rocmfp4_fast body, with the F32 MoE router and the MTP head preserved (no imatrix). Not the leanest-fastest possible, and not the most faithful possible (see the Unsloth fidelity link in §03) — it's the point where speed and quality meet best. Vision: this MoE is multimodal (Qwen3.6 is natively VL) — the repo bundles the mmproj-F32.gguf Qwen3-VL projector (projection_dim 2048, matched to the MoE hidden size; verified reading a test image) plus chat_template.jinja (tool calls + think-toggle + vision).

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">

<b>NOTE //</b> The <b>Q6_K output head</b> — the layer that turns the final hidden state into the next-token choice — is raised from 4-bit ROCmFP4 to standard <b>Q6_K</b> in this build. It's the output-side complement to running f16 token embeddings, sharpening both ends of the model; the head/embedding trade-off is characterized in detail on the <a href="https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF">27B card</a>.

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>

Run from the folder holding the .gguf + chat_template.jinja:

env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
  -m Qwen3.6-35B-A3B-MTP-ROCmFP4-STRIX-embF16-headQ6.gguf \
  --alias qwen35b-a3b-mtp \
  --host 0.0.0.0 \
  --port 8080 \
  -c 262144 \
  -ctk f16 \
  -ctv f16 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.0 \
  -dev Vulkan0 \
  -ngl 999 \
  -fa on \
  -b 2048 \
  -ub 256 \
  -t 16 \
  -tb 16 \
  -cpent 256 \
  -ctxcp 32 \
  --cache-reuse 256 \
  --cache-ram 65536 \
  --jinja \
  --parallel 1 \
  --metrics \
  --no-mmap \
  --spec-type draft-mtp \
  --spec-draft-device Vulkan0 \
  --spec-draft-ngl all \
  --spec-draft-type-k f16 \
  --spec-draft-type-v f16 \
  --spec-draft-n-max 5 \
  --spec-draft-n-min 0 \
  --spec-draft-p-min 0.0 \
  --spec-draft-p-split 0.10 \
  --chat-template-file chat_template.jinja \
  --reasoning on \
  --reasoning-format deepseek \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --mmproj mmproj-F32.gguf \
  --image-min-tokens 1024

The last two lines enable vision — the bundled mmproj-F32.gguf is the Qwen3-VL projector for this MoE (projection_dim 2048); omit them for text-only. --image-min-tokens 1024 is required whenever --mmproj is set.

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>

<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan (KHR_coopmat) — beats ROCm here for ROCmFP4 on Strix Halo</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K) — loads with ~92 GB free; KV footprint is modest</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch (256 = prefill optimum) · CPU threads</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it; drop to <code>q8_0</code>/<code>q4_0</code> to use less memory</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3.6 "precise coding" sampling (1.0 for general)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-type draft-mtp · --spec-draft-n-max 5</code></td><td style="border:1px solid currentColor; padding:6px 10px;">built-in MTP head, self-speculative; draft depth 5</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-draft-device Vulkan0 · -ngl all · type-k/v f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">draft head on Vulkan, fully offloaded, f16 KV</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--chat-template-file chat_template.jinja</code></td><td style="border:1px solid currentColor; padding:6px 10px;">froggeric unified Qwen3.6 template (tool calls + think-toggle)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--reasoning on --reasoning-format deepseek + kwargs {preserve_thinking:true}</code></td><td style="border:1px solid currentColor; padding:6px 10px;">keep <code>&lt;think&gt;</code> across turns with clean <code>content</code>+<code>reasoning_content</code>, so cross-turn cache survives</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply template · single slot · metrics · weights in RAM</td></tr>

</tbody>

</table>

</div>

Multi-turn prompt-cache reuse (OpenCode). Qwen3.6's recurrent state can't partial-rewind, so multi-turn reuse needs a context checkpoint. Two defaults otherwise force a full re-prefill every turn; both are fixed above:

  1. Checkpoints — default -cpent is 8192, so prompts under 8K never checkpoint. Fix: -cpent 256 -ctxcp 32 --cache-reuse 256.
  2. Thinking--reasoning-format deepseek + --chat-template-kwargs '{"preserve_thinking": true}' keeps <think> across turns with clean content+reasoning_content. (none = raw tags inline but works with any content-echoing client; deepseek-legacy/auto do not reuse.)

--jinja is required for the chat template + preserve_thinking.

OpenAI-compatible client (e.g. OpenCode). In single-model mode llama-server ignores the request's model field, so the client's model name is just a label.

  • Base URL: http://<host>:8080/v1 · API key: any non-empty string (e.g. sk-local)
  • Model id this server reports: qwen35b-a3b-mtp

A patched OpenCode that compacts conversation history without invalidating the prompt cache is at PlunderStruck/opencode — pair it with the checkpoint flags to keep long sessions fast.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · PERFORMANCE &amp; QUALITY</div>

This is the best speed/quality balance in ROCmFP4 — by design, not the absolute fastest. It keeps the two quality levers that are actually felt — genuine f16 token embeddings and a Q6_K output head — on the fast single-scale body, with the F32 MoE router untouched. We tested the alternatives within rocmfp4 (an all-dual-scale body, selective higher-precision tensors); they cost decode speed for a KL improvement that sat inside the measurement noise, so the fast single-scale body + f16 embeddings + Q6 head is the right point. A leaner build (no Q6 head, or Q5 embeddings) is a few tok/s faster but degrades a quality lever you'll notice; we keep both.

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">

<b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> Unsloth's <a href="https://huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF"><b>UD-Q4_K_XL</b></a> (a dynamic K-quant) runs on this <b>same fork</b>, and <b>MTP still works</b> (it carries the nextn head) — at roughly <b>~2× lower KL divergence</b> vs BF16, at a slower decode. We optimize for throughput in ROCmFP4; if you want the last bit of fidelity over speed, that's the one to grab.

</div>

How we landed on this recipe. We ran the full lever sweep on the 27B dense sibling — measuring every rocmfp4 build against the BF16 reference by KL divergence (the right fidelity metric) plus decode speed (llama-bench), and comparing to the best stock 4-bit. The finding generalizes here: an all-dual-scale body (COHERENT) and selective higher-precision bumps (DYN) both trade decode speed for a KL gain that sits inside the noise, while even copying Unsloth's entire high-precision allocation onto rocmfp4 still can't match a dynamic K-quant's fidelity — that's a format limit (rocmfp4's FP4 is intrinsically less faithful than Q4_K's 4-bit, a fidelity floor you can't out-allocate). So within rocmfp4 the fast body + f16 embeddings + Q6 head is the optimal balance (this file), and for maximum fidelity we link the dynamic K-quant rather than ship a worse copy. The numbered sweep — full experiments table, KLD numbers, and verdicts — is on the 27B card (those figures are 27B-specific; this 35B MoE follows the same frontier). (Directional internal measurements — reproduce before citing.)

Hands-on, on a Framework Desktop / AMD Ryzen AI Max+ 395 (gfx1151, 128 GB unified, ROCm 7.2.0):

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">78–90 t/s (Vulkan / Strix Halo)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP DRAFT ACCEPTANCE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~0.6–0.95 (content-dependent)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT @ LOAD</td><td style="border:1px solid currentColor; padding:8px 11px;">full 262144 with ~92 GB free</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">non-imatrix · F32 MoE router</td></tr>

</tbody>

</table>

</div>

MoE decode is naturally fast — only ~3B params active per token — and the F32 router keeps expert selection clean. The router stays F32 for free: the quantizer excludes expert-gating tensors (ffn_gate_inp) from quantization, so routing — which experts each token goes to, a discrete, high-sensitivity decision — keeps full precision automatically, while the experts run on the custom ROCmFP4 kernel.

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">

<b>NOTE //</b> <b>f16 KV</b> is how we run it (128 GB unified affords it; drop to q8_0/q4_0 to save memory). The <b>Q6_K output head</b> adds ~0.14 GB and a fixed per-token cost (a few % slower decode at short context, shrinking at long context) for a measurable fidelity gain — see the <a href="https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF">27B card</a>.

</div>

The companion 27B dense quant (same recipe) is at plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · BUILD (REPRODUCIBLE)</div>

Build the fork:

git clone https://github.com/charlie12345/ROCmFPX
cd ROCmFPX
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh

Quantize from the unsloth BF16+MTP GGUF — ROCmFP4 body, genuine f16 embeddings, no imatrix:

# the one build: STRIX preset + f16 embeddings + Q6_K output head
llama-quantize \
  --token-embedding-type f16 \
  --output-tensor-type q6_K \
  Qwen3.6-35B-A3B-BF16-00001-of-00002.gguf \
  Qwen3.6-35B-A3B-MTP-ROCmFP4-STRIX-embF16-headQ6.gguf \
  Q4_0_ROCMFP4_STRIX

Architecture (qwen35moe): 41 blocks, 2048 hidden, 256 experts, with the nextn_predict_layers=1 MTP head (blk.40.nextn.) — so self-speculative draft-MTP survives quantization. Format: ROCmFP4 is a 4-bit weight format for AMD using an FP4-derived value codebook plus one (FAST) or two (dual) UE4M3/FP8 scale bytes per 32-weight block; tensor-aware. This build (STRIX-embF16-headQ6): quality-biased STRIX preset + f16 token embeddings (full precision; a lookup, so ~zero decode cost) + a Q6_K output head. Experts (ffn__exps) run q4_0_rocmfp4_fast; attention K/V (+ fused QKV) run q4_0_rocmfp4 (dual-scale).

> Experimental research build for AMD Strix Halo — hardware-, driver-, model-, and prompt-sensitive, may not reproduce on other GPUs. Not native FP4 tensor-core execution. Do not treat these numbers as upstream llama.cpp claims. Base BF16 GGUF pinned at revision 5bc3e238d916f48a861bac2f8a1990a0e9b7e98d.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · LINEAGE &amp; CREDITS</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;">Qwen3.6-35B-A3B (Qwen team) — derivative quantization that <b>inherits the base model's license</b>; verify the original Qwen3.6 terms before redistribution / use</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">BF16 GGUF SOURCE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF">unsloth/Qwen3.6-35B-A3B-MTP-GGUF</a> @ <code>5bc3e238d916f48a861bac2f8a1990a0e9b7e98d</code></td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> (based on llama.cpp, MIT)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">CHAT TEMPLATE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates">froggeric/Qwen-Fixed-Chat-Templates</a></td></tr>

</tbody>

</table>

</div>

Derivative quantization — verify the base model's license before redistribution / use.

Run plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models