plunderstruck/ThinkingCap-Qwen3.6-27B-MTP-ROCmFP4-GGUF overview
<div style="border:2px solid currentColor; font family:ui monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;" <div style="border bottom:…
Runs locally from ~888.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | plunderstruck/ThinkingCap-Qwen3.6-27B-MTP-ROCmFP4-GGUF |
|---|---|
| Author | plunderstruck |
| Pipeline | — |
| License | apache-2.0 |
| Base model | bottlecapai/ThinkingCap-Qwen3.6-27B |
| Last modified | 2026-07-07T17:08:07.000Z |
Model README
---
base_model: bottlecapai/ThinkingCap-Qwen3.6-27B
license: apache-2.0
library_name: gguf
tags:
- gguf
- rocmfp4
- qwen3.6
- mtp
- speculative-decoding
- efficient-thinking
- strix-halo
- amd
- rocm
- vulkan
base_model_relation: quantized
---
<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">
<div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>
<div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">
<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:6px; line-height:1.1; letter-spacing:0;">
▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄
▟███████████████████████▙
▄▟█████████████████████████████▙▄
▟███████████████████████████████████▙
▐███████████████▛▀▀▀▀▀▜███████████████▌
▝████████████████ ████████████▛▘
▐████▛▘ ▝▜████▌
▐█▌ ▟▙ ▐█▌
▐█▌ ▟███▙ ▐█▌
▀▀ ▟█████▙ ▀▀
▟███████▙
</pre>
<div style="flex:0 1 auto; max-width:100%; text-align:center;">
<div style="font-size:23px; font-weight:800; letter-spacing:1px;">THINKINGCAP-QWEN3.6-27B-MTP</div>
<div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">EFFICIENT-THINKING</span> · <span style="white-space:nowrap;">MTP SELF-SPECULATIVE DECODE</span> · <span style="white-space:nowrap;">VISION-CAPABLE</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>
</div>
</div>
<table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
<tr>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">4.94 BPW</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">16.9 GB</div></td>
<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">262 K</div></td>
</tr>
<tr>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">MTP n-max 5</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">VISION</div><div style="font-weight:700;">QWEN3-VL</div></td>
<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>
<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CALIBRATION</div><div style="font-weight:700;">non-imatrix</div></td>
</tr>
</table>
</div>
<div style="border:1px solid currentColor; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">
<b>ThinkingCap-Qwen3.6-27B</b> is <a href="https://huggingface.co/bottlecapai">bottlecapai</a>'s efficient-reasoning finetune of <b>Qwen3.6-27B</b> — trained to cut thinking-token consumption roughly in half (<b>↓45.8%</b> macro average) while holding answer quality, across knowledge, math, code, long-context, and agentic tasks. It scores <b>GPQA-Diamond 83.8</b>, <b>MMLU-Pro 85.4</b>, <b>GSM8K 96.5</b> (bottlecapai reported). This repo is the <b>ROCmFP4</b> quantization tuned for a single AMD Strix Halo APU, with the <b>MTP head preserved</b> for self-speculative decode and vision carried through from the base model.
</div>
<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">
<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFPX FORK</b><br>
The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, Ollama, Jan, or koboldcpp</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>experimental-rocmfpx-branch</code>:
<br><br>
<code>git clone https://github.com/charlie12345/ROCmFPX</code><br>
<code>cd ROCmFPX && git checkout experimental-rocmfpx-branch</code><br>
<code>env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh</code>
</div>
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">
<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16" badge — its parser can't read ROCmFP4 and mislabels by the genuinely-f16 token embeddings. This is a <b>~4.9 bpw 4-bit</b> file; pick by filename in <i>Files and versions</i>.
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Output head</th>
<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-headQ6.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">16.9 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">Q6_K</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> — best speed/quality balance: f16 embeddings + Q6 output head on the fast single-scale body</td></tr>
</tbody>
</table>
</div>
One file — the best speed/quality balance in ROCmFP4 for Strix Halo. It keeps the two quality levers that are actually felt — genuine f16 token embeddings (from F16 source) and a Q6_K output head — on the fast single-scale q4_0_rocmfp4_fast body, plus the MTP head, with no imatrix (this recipe's daily-driver default; see §05). Repo bundles the mmproj-F32.gguf Qwen3-VL vision projector and chat_template.jinja (froggeric's unified Qwen3.6 template — tool calls + inline think-toggle + vision).
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">token_embd</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">F16 (full precision — a lookup, ~zero decode cost)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">attention K/V (+ fused QKV)</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;"><code>q4_0_rocmfp4</code> (dual-scale)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">FFN, lm-head, rest</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;"><code>q4_0_rocmfp4_fast</code> (single-scale)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP head</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">preserved (<code>nextn_predict_layers=1</code>)</td></tr>
</tbody>
</table>
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>
Run from the folder holding the .gguf + chat_template.jinja:
env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
-m ThinkingCap-Qwen3.6-27B-ROCmFP4-STRIX-embF16-headQ6.gguf \
--alias thinkingcap-27b \
--host 0.0.0.0 \
--port 8080 \
-dev Vulkan0 \
-ngl 999 \
-fa on \
-c 262144 \
-b 2048 \
-ub 256 \
-t 16 \
-tb 16 \
-ctk f16 \
-ctv f16 \
-cpent 256 \
-ctxcp 32 \
--cache-reuse 256 \
--cache-ram 65536 \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--spec-type draft-mtp \
--spec-draft-device Vulkan0 \
--spec-draft-ngl all \
--spec-draft-type-k f16 \
--spec-draft-type-v f16 \
--spec-draft-n-max 5 \
--spec-draft-n-min 2 \
--spec-draft-p-min 0.0 \
--spec-draft-p-split 0.10 \
--chat-template-file chat_template.jinja \
--reasoning on \
--reasoning-format deepseek \
--chat-template-kwargs '{"preserve_thinking": true}' \
--jinja \
--parallel 1 \
--metrics \
--no-mmap \
--mmproj mmproj-F32.gguf \
--image-min-tokens 1024
The last two lines enable vision — the mmproj-F32.gguf Qwen3-VL projector is bundled in this repo; omit them for text-only. --image-min-tokens 1024 is required whenever --mmproj is set.
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE //</b> bottlecapai's own recommended sampling for general use is <code>temp=1.0, top_p=0.95, top_k=20, min_p=0.0</code>. We serve at <code>temp 0.6</code> (Qwen3.6 "precise coding" preset) by default — raise to <code>1.0</code> for open-ended/creative tasks.
</div>
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">
<thead><tr>
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>
<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>
</tr></thead>
<tbody>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan (KHR_coopmat) — beats ROCm/HIP here for ROCmFP4 on Strix Halo</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch (256 = prefill optimum) · CPU threads</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it; drop to <code>q8_0</code>/<code>q4_0</code> to use less memory</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">precise-coding sampling (bottlecapai recommends 1.0 for general use)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-type draft-mtp · --spec-draft-n-max 5 · n-min 2</code></td><td style="border:1px solid currentColor; padding:6px 10px;">built-in MTP head, self-speculative; draft depth up to 5, at least 2</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-draft-device Vulkan0 · -ngl all · type-k/v f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">draft head on Vulkan, fully offloaded, f16 draft KV</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--chat-template-file chat_template.jinja</code></td><td style="border:1px solid currentColor; padding:6px 10px;">bundled froggeric template (tool calls + think-toggle + vision)</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--reasoning on --reasoning-format deepseek + kwargs {preserve_thinking:true}</code></td><td style="border:1px solid currentColor; padding:6px 10px;">clean <code>content</code>+<code>reasoning_content</code>; keep <code><think></code> across turns so cross-turn cache survives</td></tr>
<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply template · single slot · metrics · weights in RAM</td></tr>
</tbody>
</table>
</div>
OpenAI-compatible client (e.g. OpenCode). In single-model mode llama-server ignores the request's model field, so the client's model name is just a label.
- Base URL:
http://<host>:8080/v1· API key: any non-empty string (e.g.sk-local) - Model id this server reports:
thinkingcap-27b
A patched OpenCode that compacts conversation history without invalidating the prompt cache is at PlunderStruck/opencode — pair it with the checkpoint flags to keep long sessions fast.
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · VISION</div>
Qwen3-VL lineage — vision works via the bundled mmproj-F32.gguf projector at launch with --mmproj (no different LLM GGUF needed).
# add to your llama-server launch:
--mmproj mmproj-F32.gguf \
--image-min-tokens 1024 # REQUIRED — Qwen-VL needs >=1024 image tokens or it misreads fine detail
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE //</b> this is a thinking model — for one-shot image Q&A, use the bundled template's think-toggle or allow enough tokens to finish <code><think></code>, else the visible answer can come back empty. With <code>--mmproj</code> loaded the server disables the <code>--cache-reuse</code> feature (multimodal caching isn't supported).
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · PERFORMANCE & QUALITY</div>
This is the best speed/quality balance in ROCmFP4 — by design, not the absolute fastest. It keeps the two quality levers that are actually felt — genuine f16 token embeddings and a Q6_K output head — on the fast single-scale body, with no imatrix calibration (this recipe's default for the dense Qwen3.6-27B line — see the 27B card for the full lever sweep and rationale).
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">
<b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> bottlecapai's own GGUF repo ships standard K-quants — <a href="https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF"><b>Q4_K_M / Q8_0</b></a> — more faithful to the source at a slower decode. We optimize for throughput in ROCmFP4 on Strix Halo; grab one of those for the last bit of fidelity.
</div>
Hands-on, on a Framework Desktop / AMD Ryzen AI Max+ 395 (gfx1151, 128 GB unified):
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~30 t/s (Vulkan / Strix Halo, thinking on)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP DRAFT ACCEPTANCE</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~4.15 mean accepted length (n-max 5 / n-min 2)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CONTEXT @ LOAD</td><td style="border:1px solid currentColor; padding:8px 11px;">full 262144, f16 KV</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">non-imatrix</td></tr>
</tbody>
</table>
</div>
<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;">
<b>NOTE //</b> built on a fork snapshot carrying an upstream fix for a self-speculative-decode bug affecting M-RoPE architectures (Qwen3.6 included) — MTP was silently degrading to plain decode under certain batch shapes on some builds. This quant's measured acceptance (~4.15) reflects the fixed path; older ROCmFPX builds may show materially lower MTP throughput on this model family.
</div>
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · BUILD (REPRODUCIBLE)</div>
Build the fork:
git clone https://github.com/charlie12345/ROCmFPX
cd ROCmFPX && git checkout experimental-rocmfpx-branch
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh
Quantize from the bottlecapai F16 GGUF — ROCmFP4 body, genuine f16 embeddings, Q6_K head, no imatrix:
# the one build: STRIX preset + f16 embeddings + Q6_K output head
llama-quantize \
--token-embedding-type f16 \
--output-tensor-type q6_K \
ThinkingCap-Qwen3.6-27B-f16.gguf \
ThinkingCap-Qwen3.6-27B-ROCmFP4-STRIX-embF16-headQ6.gguf \
Q4_0_ROCMFP4_STRIX
Architecture (qwen35): 64 blocks, 5120 hidden, dense (not MoE), nextn_predict_layers=1 MTP head — self-speculative draft-MTP survives quantization. Format: ROCmFP4 is a 4-bit weight format for AMD using an FP4-derived value codebook plus one (FAST) or two (dual) UE4M3/FP8 scale bytes per 32-weight block; tensor-aware. This build (STRIX-embF16-headQ6): quality-biased STRIX preset + f16 token embeddings (full precision; a lookup, so ~zero decode cost) + a Q6_K output head. Attention K/V (+ fused QKV) run q4_0_rocmfp4 (dual-scale); FFN/rest run q4_0_rocmfp4_fast (single-scale).
> Experimental research build for AMD Strix Halo — hardware-, driver-, model-, and prompt-sensitive, may not reproduce on other GPUs. Not native FP4 tensor-core execution. Do not treat these numbers as upstream llama.cpp claims.
<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · LINEAGE & CREDITS</div>
<div style="overflow:hidden; border-radius:0;">
<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">
<tbody>
<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B">bottlecapai/ThinkingCap-Qwen3.6-27B</a> — efficient-thinking finetune of <a href="https://huggingface.co/Qwen/Qwen3.6-27B">Qwen/Qwen3.6-27B</a> (Apache 2.0). This is a derivative quantization that inherits the Apache 2.0 license.</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">F16 GGUF SOURCE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF">bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF</a> · <code>ThinkingCap-Qwen3.6-27B-f16.gguf</code></td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> (based on llama.cpp, MIT)</td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">CHAT TEMPLATE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates">froggeric/Qwen-Fixed-Chat-Templates</a></td></tr>
<tr><td style="border:1px solid currentColor; padding:8px 11px;">BENCHMARKS (BASE)</td><td style="border:1px solid currentColor; padding:8px 11px;">GPQA-Diamond 83.8±1.9 · MMLU-Pro 85.4±0.2 · GSM8K 96.5±0.3 · thinking-token reduction ↓45.8% macro avg (bottlecapai reported)</td></tr>
</tbody>
</table>
</div>
Derivative quantization — Apache 2.0, same as the base model.
Run plunderstruck/ThinkingCap-Qwen3.6-27B-MTP-ROCmFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models