GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF overview

Qwen3.8 Flash Next ROCmFP4 FAST GGUF Superseded by the imatrix build. Qwen3.8 Flash Next ROCmFP4 FAST imatrix GGUF https://huggingface.co/agentionai/Qwen3.8 Fl…

ggufrocmfp4rocmfpxvulkanstrix-haloqwen4exptext-generationbase_model:Qwen/Qwen3.8-Flash-Nextbase_model:quantized:Qwen/Qwen3.8-Flash-Nextlicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~2.28 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
5,641
Likes
15
Pipeline
text-generation

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Next-ROCmFP4-FAST-00001-of-00005.ggufGGUFGGUF17.45 GBDownload
Qwen3.8-Flash-Next-ROCmFP4-FAST-00002-of-00005.ggufGGUFGGUF18.60 GBDownload
Qwen3.8-Flash-Next-ROCmFP4-FAST-00003-of-00005.ggufGGUFGGUF18.43 GBDownload
Qwen3.8-Flash-Next-ROCmFP4-FAST-00004-of-00005.ggufGGUFGGUF18.45 GBDownload
Qwen3.8-Flash-Next-ROCmFP4-FAST-00005-of-00005.ggufGGUFGGUF10.72 GBDownload
mtp/Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST.ggufGGUFGGUF2.28 GBDownload

Model Details

Model IDagentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Authoragentionai
Pipelinetext-generation
Licenseother
Base modelQwen/Qwen3.8-Flash-Next
Last modified2026-08-30T19:09:28.000Z

Model README

---

base_model:

  • Qwen/Qwen3.8-Flash-Next

base_model_relation: quantized

license: other

license_name: qwen-community-1.0

license_link: LICENSE

library_name: gguf

pipeline_tag: text-generation

tags:

  • gguf
  • rocmfp4
  • rocmfpx
  • vulkan
  • strix-halo
  • qwen4exp

---

Qwen3.8-Flash-Next ROCmFP4-FAST GGUF

> Superseded by the imatrix build.

> Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF

> measures 4.1062 perplexity against this file's 4.6785, for 3.4 GiB more, and adds

> the vision tower. Use it instead unless you specifically need this file.

ROCmFP4 quantization of Qwen/Qwen3.8-Flash-Next,

sized for the 96 GiB VRAM carve-out of a Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S).

83.65 GiB across 5 shards, 4.06 bpw.

The quant recipe and the fork's per-head PLE sharding both exist for the same reason: keep

the entire model resident on GPU. Nothing falls back to host RAM or CPU compute - not the

experts, not the 51.2 B-parameter n-gram table, none of it.

This is an experimental build, made for speed on one machine rather than for quality.

ROCmFPx is an experimental quant family, it is carried in a fork rather than upstream, and

this file is quantized without an importance matrix. It measures 4.6785 perplexity against

4.0068 for the unquantized model, which is a wider gap than a good 4-bit quant should have.

If you want quality, use a mainstream quant; if you want ROCmFP4 kernels on RDNA3.5, this

is what it is for.

The imatrix build linked above closes about 85% of that gap, for 3.4 GiB more

(87.06 GiB against this file's 83.65). Prefer it unless you have a specific reason not to.

Setup

Qwen3.8-Flash-Next itself merged into upstream llama.cpp

(ggml-org/llama.cpp#27742) - this file

still needs this fork for two things upstream doesn't have: the ROCmFPx quant types, and

the n-gram table split per head (upstream only reads it joined, which is past what most

Vulkan devices accept as a single buffer). Build this branch:

git clone https://github.com/LaurentZuijdwijk/llama.cpp
cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Run

Point at the first shard; the rest follow automatically.

./build/bin/llama-cli \
  -m Qwen3.8-Flash-Next-ROCmFP4-FAST-00001-of-00005.gguf \
  -ngl 99 -c 32768

Runs fully on the GPU: ~85 GiB VRAM, negligible host RAM. KV is roughly 24 KiB

per token.

The n-gram table ships as 16 per-head tensors of ~1.3 GiB rather than one 20.9 GiB

tensor. A single tensor that size is past maxStorageBufferRange (4 GiB on most Vulkan

devices), so joined it can only ever live on the host - and on a machine whose VRAM

carve-out leaves less than 21 GiB for the host, that means swap.

Quantization

| tensors | type | bpw |

|---|---|---|

| MoE experts, attention, GDN, hyper-connections | Q4_0_ROCMFP4_FAST | 4.25 |

| n-gram PLE table (51.2 B params) | Q3_0_ROCMFPX | 3.50 |

| token_embd, output | Q6_K | 6.56 |

Perplexity

wikitext-2 raw, 145 chunks at -c 2048.

| build | PPL |

|---|---|

| unquantized reference (as reported in PR 27742) | 4.0068 +/- 0.02271 |

| this file | 4.6785 +/- 0.02780 |

| imatrix version | 4.1062 +/- 0.02329 |

Quantized without an importance matrix, from the FP8 release, with a flat 4.25 bpw

backbone. The imatrix version above changes all three - BF16 source, Q6_K backbone, and

calibration on 1540 chunks merged from two corpora - and closes about 85% of the remaining

gap to the unquantized reference for 3.4 GiB.

Converting other qwen4exp GGUFs

Files built the upstream way carry the table joined. This fork reads them, but the

table stays host-side. To split it per head:

python gguf-py/gguf/scripts/gguf_split_ple_heads.py in-00001-of-000NN.gguf out.gguf

Head bounds come from the file's own KV, and the quantized bytes are copied through

untouched - no dequantize, no requantize, no quality change. Works on any quant and on

split inputs. Verified on unsloth's UD-IQ4_XS, which goes from OOMing a 30 GB host to

88.6 GiB fully resident on the GPU.

MTP draft head

mtp/ holds the model's own multi-token-prediction head, 2.27 GiB, exported from the same

checkpoint. Qwen trains it jointly with the target, so it drafts better than a separate

small model would.

./build/bin/llama-server \
  -m Qwen3.8-Flash-Next-ROCmFP4-FAST-00001-of-00005.gguf \
  -md mtp/Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST.gguf \
  -ngl 99 --n-gpu-layers-draft 99 \
  --spec-type draft-mtp --spec-draft-n-max 3 -c 32768

Measured on a Radeon 8060S, 250 tokens at temp 0, each config warmed up first:

| draft | t/s | acceptance |

|---|---|---|

| none | 28.1 | -- |

| n-max 2 | 31.8 | 0.695 |

| n-max 3 | 32.4 | 0.612 |

It is quantized to match the target rather than above it. A Q8_0 draft measured worse on

both throughput and acceptance and cost 1.5 GiB more: acceptance is the draft agreeing with

the target, and two models quantized the same way are wrong in the same places.

Adds ~2.3 GiB to the ~85 GiB the target uses.

Credits

qwen4exp support is the work of Daniel Han

(@danielhanchen), from

ggml-org/llama.cpp#27742, merged

upstream 2026-08-28. This fork is only still needed for what's listed under Setup above.

Quant formats hand-ported from ciru-ai/ROCmFPX.

The ROCmFP4 format was created by charlie12345 in

charlie12345/ROCmFPX, which ciru-ai's tree forks.

Both upstream projects are MIT-licensed.

Base model by the Qwen team.

Quantized and published by Agention.

Not included

Vision tower.

License

Qwen Community License 1.0, included as LICENSE.

Run agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models