cygnal/Qwen3.8-Flash-Next-Uncensored-IQ4XS-NGQ4-GGUF overview
Qwen3.8 Flash Next Uncensored — IQ4XS NGQ4 GGUF — AMD Strix Halo gfx1151 First working GGUF build of Qwen3.8 Flash Next qwen4exp architecture with vision, runn…
Runs locally from ~865.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
Model README
---
license: other
license_name: qwen-community-1.0
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
tags:
- gguf
- llama.cpp
- strix-halo
- gfx1151
- rocm
- amd
- ryzen-ai-max
- uncensored
- vision
- qwen4exp
---
Qwen3.8-Flash-Next-Uncensored — IQ4XS-NGQ4 GGUF — AMD Strix Halo (gfx1151)
First working GGUF build of Qwen3.8-Flash-Next (qwen4exp architecture) with vision, running on stock
llama.cpp — no custom tensor formats or forked runtime required. Built from
orcarouter/Qwen3.8-Flash-Next-Uncensored,
the abliterated (uncensored) release of Qwen's newest hybrid architecture.
⚠️ Research artifact. Refusal behaviour has been removed from the source model. This does not add
capability — it removes guardrails. Use it deliberately, in a context where that is appropriate, and own
the output.
Architecture
- 125B total + 51B n-gram PLE (per-layer embedding) table + 4B MTP head = ~180B parameters, 6B active per
token (512 experts, 10 routed + 1 shared)
- Hybrid Gated DeltaNet (36 layers) + Qwen Sparse Attention (12 layers)
- Natively multimodal (Qwen3-VL-style vision tower, same family as our heretic-ara builds)
- 262K native context (this build tested to 6.5K; larger contexts untested but should work — QSA caps KV
growth the same way it does on other Qwen3.8-Flash-Next builds)
- MTP head not exported — this GGUF has no speculative-decode acceleration. All numbers below are
plain autoregressive decode.
Quant recipe — IQ4XS-NGQ4
Straightforward stock quantization, no custom AMD tensor formats:
| tensor group | type |
|---|---|
| bulk weights (attention, experts, embeddings) | IQ4_XS (falls back to IQ4_NL/F16 on ~194 odd-shaped tensors) |
| output.weight (lm head) | Q6_K |
| per_layer_token_embd.weight (the 51.2B-param n-gram PLE table) | Q4_0 |
5.61 BPW average, 98.4 GB total. The n-gram table alone is ~29 GB of that at Q4_0 — it's lookup-only
(hash → embedding), so it tolerates low-bit quantization well, but it needed a separate streaming
dequant/requant pass (chunked, ~2M rows at a time) because standard llama-quantize can't process its
unusual 160-wide row shape and would need ~200+ GB to dequantize it in one pass.
The lm head is pinned to Q6_K — every sampled token passes through it, so its quantization error lands
directly in the argmax; keeping it high-precision costs under 1 GB and removes that error class entirely.
Measured — Ryzen AI Max+ 395, gfx1151, Vulkan, stock llama.cpp + PR #27742
Full GPU offload (-ngl 99 -dev Vulkan0 -fa on), no MTP, no ROCmFPX — plain Vulkan backend on the same
llama.cpp tree unslothai/danielhanchen maintain the qwen4exp branch on.
| prompt length | prefill (tok/s) | decode (tok/s) |
|---|---:|---:|
| 3,063 tokens | 368.7 | 21.7 |
| 6,516 tokens | 345.4 | 20.5 |
Decode speed is in the same range as our 27B dense heretic-ara build (17-23 tok/s) despite Flash-Next
being a ~180B-parameter model — a reasonable result for 6B active params, though not yet a clear win, and
notably without any speculative decoding. Adding MTP (once/if we get an export path for it) or a
ROCmFP4-class quant should improve this further.
Coding accuracy — EvalPlus, greedy, 4096 max tokens
| benchmark | pass@1 |
|---|---:|
| HumanEval | 82.3% |
| HumanEval+ | 78.0% |
For reference, our other local quants on the same harness: Q6_K (27B dense) 82.9%, Q4_K_M (27B dense)
75.6%, ROCmFP6 (27B dense) 73.8%, ROCmFP4 (27B dense) 71.3%. This IQ4XS-NGQ4 Flash-Next build lands
between Q6_K and Q4_K_M on the same benchmark, despite averaging a lower bits-per-weight — likely the MoE
architecture and larger total parameter count compensating for the more aggressive quantization.
Uncensoring
Verified qualitatively: a lockpicking-mechanics prompt that stock-aligned models typically refuse or
deflect on got a direct, technical answer with no refusal framing. This is inherited from the
orcarouter abliteration — see their
model card for their methodology and safety benchmark numbers. We did not run a formal refusal-rate
benchmark ourselves; treat this as anecdotal confirmation that the abliteration survived quantization
intact, not a rigorous safety audit.
Files
| file | size |
|---|---|
| Qwen3.8-Flash-Next-Uncensored-IQ4XS-NGQ4.gguf | 98.4 GB |
| mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf | 908 MB (vision tower) |
Usage
Requires llama.cpp built from the qwen4exp architecture branch (PR #27742 —
not yet merged to master as of this writing). Stock master builds and tools like Ollama/LM Studio
cannot load this file until that PR lands. We built from
danielhanchen/llama.cpp branch qwen4exp/qwen3.8-flash-next,
plain Vulkan, no ROCm/HIP-specific patches needed:
git clone -b qwen4exp/qwen3.8-flash-next --single-branch https://github.com/danielhanchen/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON
cmake --build build -j$(nproc) --target llama-server
./build/bin/llama-server \
--model Qwen3.8-Flash-Next-Uncensored-IQ4XS-NGQ4.gguf \
--mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf \
--host 127.0.0.1 --port 8080 \
--n-gpu-layers 99 --flash-attn on \
--ctx-size 32768 --jinja
Needs ~99 GB free on the GPU/unified-memory pool for the weights alone, plus headroom for KV cache and
the vision tower. On a 128 GB Strix Halo box, load this before anything else has claimed UMA.
The community-maintained kyuz0/amd-strix-halo-toolboxes
project also ships a pre-built container (rocm-7.14-qwen-3.8-flash-next) that builds the same PR branch
on ROCm instead of Vulkan, if you'd rather not build from source.
Acknowledgements
orcarouter — published the
uncensored BF16 checkpoint this is built from. The abliteration is their engineering; we only converted
and quantized it.
danielhanchen / Unsloth — authored llama.cpp PR
#27742 adding qwen4exp architecture support
(Gated DeltaNet, Qwen Sparse Attention, the n-gram PLE table, and MTP), which this build depends on
entirely. Not yet merged to master.
llama.cpp — ggml-org and contributors — the engine, GGUF
format, and conversion tooling this is built on.
Qwen team — the original base model. See base_model; license qwen-community-1.0.
kyuz0/amd-strix-halo-toolboxes — independently
confirmed the same PR branch + build approach we used, and maintains pre-built containers for it.
Run cygnal/Qwen3.8-Flash-Next-Uncensored-IQ4XS-NGQ4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models