GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

agentionai/Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST-GGUF overview

Qwen3.8 Flash Next MTP draft head ROCmFP4 FAST The multi token prediction head from Qwen/Qwen3.8 Flash Next https://huggingface.co/Qwen/Qwen3.8 Flash Next , 2.…

ggufmtpspeculative-decodingrocmfp4vulkanstrix-haloqwen4exptext-generationbase_model:Qwen/Qwen3.8-Flash-Nextbase_model:quantized:Qwen/Qwen3.8-Flash-Nextlicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~2.28 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
920
Likes
2
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST.ggufGGUFGGUF2.28 GBDownload

Model Details

Model IDagentionai/Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST-GGUF
Authoragentionai
Pipelinetext-generation
Licenseother
Base modelQwen/Qwen3.8-Flash-Next
Last modified2026-08-30T19:09:31.000Z

Model README

---

base_model:

  • Qwen/Qwen3.8-Flash-Next

base_model_relation: quantized

license: other

license_name: qwen-community-1.0

license_link: LICENSE

library_name: gguf

pipeline_tag: text-generation

tags:

  • gguf
  • mtp
  • speculative-decoding
  • rocmfp4
  • vulkan
  • strix-halo
  • qwen4exp

---

Qwen3.8-Flash-Next MTP draft head (ROCmFP4-FAST)

The multi-token-prediction head from

Qwen/Qwen3.8-Flash-Next, 2.28 GiB.

Qwen trains it jointly with the target model, so it drafts better than a separate small

model would.

This is a draft head, not a model. On its own it does nothing. It is used with

agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF,

which is an experimental build; see that card first.

Setup

qwen4exp and the ROCmFPx quant types are not in upstream llama.cpp yet, so build this branch:

git clone https://github.com/LaurentZuijdwijk/llama.cpp
cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Run

./build/bin/llama-server \
  -m Qwen3.8-Flash-Next-ROCmFP4-FAST-00001-of-00005.gguf \
  -md Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST.gguf \
  -ngl 99 --n-gpu-layers-draft 99 \
  --spec-type draft-mtp --spec-draft-n-max 3 -c 32768

Adds ~2.3 GiB to the ~85 GiB the target uses.

Measurements

Radeon 8060S (Ryzen AI MAX+ 395), 250 tokens at temp 0, each config warmed up first.

| draft | t/s | acceptance |

|---|---|---|

| none | 28.1 | -- |

| n-max 2 | 31.8 | 0.695 |

| n-max 3 | 32.4 | 0.612 |

Quantized to match the target, not above it. A Q8_0 draft measured worse on both

throughput and acceptance and cost 1.5 GiB more (30.3 t/s, 0.587): acceptance is the

draft agreeing with the target, and two models quantized the same way are wrong in the

same places.

Adaptive drafting also measured worse here (28.4 t/s, 0.468). It drafts longer when

acceptance looks good, and this head carries its own 512-expert MoE, so each extra

drafted token is a real forward pass.

Credits

qwen4exp support is the work of Daniel Han

(@danielhanchen), from

ggml-org/llama.cpp#27742, and the MTP

graph is from #27739 (JJJYmmm). Both

are unmerged drafts; if they land upstream, prefer upstream.

Quant formats hand-ported from ciru-ai/ROCmFPX.

The ROCmFP4 format was created by charlie12345 in

charlie12345/ROCmFPX, which ciru-ai's tree forks.

Both upstream projects are MIT-licensed.

Base model by the Qwen team.

Quantized and published by Agention.

License

Qwen Community License 1.0, included as LICENSE.

Run agentionai/Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models