vcruz305/Qwen3.8-Flash-Next-GGUF overview
Qwen3.8 Flash Next GGUF vcruz305 Real K quant ladder from full HF BF16 weights with MTP heads in main blk.48 / nextn. and a shared BF16 PLE . Layout | path | c…
Runs locally from ~553.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| MTP/Qwen3.8-Flash-Next-MTP-Q4_K_M.gguf | GGUF | Q4_K_M | 2.59 GB | Download |
| MTP/Qwen3.8-Flash-Next-MTP-Q8_0.gguf | GGUF | Q8_0 | 3.85 GB | Download |
| MTP/Qwen3.8-Flash-Next-MTP.gguf | GGUF | GGUF | 7.24 GB | Download |
| PLE-BF16/Qwen3.8-Flash-Next-PLE-BF16.gguf | GGUF | BF16 | 95.37 GB | Download |
| Q2_K/Qwen3.8-Flash-Next-Q2_K-00001-of-00007.gguf | GGUF | Q2_K | 9.19 GB | Download |
| Q2_K/Qwen3.8-Flash-Next-Q2_K-00002-of-00007.gguf | GGUF | Q2_K | 9.57 GB | Download |
| Q2_K/Qwen3.8-Flash-Next-Q2_K-00003-of-00007.gguf | GGUF | Q2_K | 9.17 GB | Download |
| Q2_K/Qwen3.8-Flash-Next-Q2_K-00004-of-00007.gguf | GGUF | Q2_K | 9.15 GB | Download |
| Q2_K/Qwen3.8-Flash-Next-Q2_K-00005-of-00007.gguf | GGUF | Q2_K | 9.37 GB | Download |
| Q2_K/Qwen3.8-Flash-Next-Q2_K-00006-of-00007.gguf | GGUF | Q2_K | 3.07 GB | Download |
| Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00001-of-00007.gguf | GGUF | Q3_K_M | 11.73 GB | Download |
| Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00002-of-00007.gguf | GGUF | Q3_K_M | 12.14 GB | Download |
| Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00003-of-00007.gguf | GGUF | Q3_K_M | 11.60 GB | Download |
| Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00004-of-00007.gguf | GGUF | Q3_K_M | 11.58 GB | Download |
| Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00005-of-00007.gguf | GGUF | Q3_K_M | 11.83 GB | Download |
| Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00006-of-00007.gguf | GGUF | Q3_K_M | 3.58 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00001-of-00007.gguf | GGUF | Q4_K_M | 14.88 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00002-of-00007.gguf | GGUF | Q4_K_M | 15.59 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00003-of-00007.gguf | GGUF | Q4_K_M | 14.82 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00004-of-00007.gguf | GGUF | Q4_K_M | 14.50 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00005-of-00007.gguf | GGUF | Q4_K_M | 16.15 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00006-of-00007.gguf | GGUF | Q4_K_M | 4.30 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00001-of-00007.gguf | GGUF | Q5_K_M | 17.04 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00002-of-00007.gguf | GGUF | Q5_K_M | 17.75 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00003-of-00007.gguf | GGUF | Q5_K_M | 16.99 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00004-of-00007.gguf | GGUF | Q5_K_M | 16.72 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00005-of-00007.gguf | GGUF | Q5_K_M | 18.07 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00006-of-00007.gguf | GGUF | Q5_K_M | 4.75 GB | Download |
| Q6_K/Qwen3.8-Flash-Next-Q6_K-00001-of-00007.gguf | GGUF | Q6_K | 20.32 GB | Download |
| Q6_K/Qwen3.8-Flash-Next-Q6_K-00002-of-00007.gguf | GGUF | Q6_K | 21.02 GB | Download |
| Q6_K/Qwen3.8-Flash-Next-Q6_K-00003-of-00007.gguf | GGUF | Q6_K | 20.28 GB | Download |
| Q6_K/Qwen3.8-Flash-Next-Q6_K-00004-of-00007.gguf | GGUF | Q6_K | 20.24 GB | Download |
| Q6_K/Qwen3.8-Flash-Next-Q6_K-00005-of-00007.gguf | GGUF | Q6_K | 20.51 GB | Download |
| Q6_K/Qwen3.8-Flash-Next-Q6_K-00006-of-00007.gguf | GGUF | Q6_K | 5.42 GB | Download |
| imatrix/Qwen3.8-Flash-Next.imatrix.gguf | GGUF | GGUF | 553.2 MB | Download |
Model Details
Model README
---
license: other
base_model: Qwen/Qwen3.8-Flash-Next
tags:
- gguf
- qwen
- qwen3.8
- flash-next
- llama.cpp
- mtp
- k-quant
---
Qwen3.8-Flash-Next GGUF (vcruz305)
Real K-quant ladder from full HF BF16 weights with MTP heads in-main (blk.48 / nextn.*) and a shared BF16 PLE.
Layout
| path | content |
|---|---|
| Q2_K/ … Q6_K/ | Backbone shards 00001–00006 (MTP included in tensor set) |
| PLE-BF16/Qwen3.8-Flash-Next-PLE-BF16.gguf | Shared BF16 PLE (~95.4 GiB) — use as each quant's 00007 |
| imatrix/ | Importance matrix used for Q2–Q6 |
| scripts/link-ple.sh | ln/cp PLE → …-00007-of-00007.gguf per quant |
| MTP/ | Optional standalone MTP draft heads (if present) |
Setup
bash scripts/link-ple.sh
Approximate backbone sizes (GiB) + shared PLE
| quant | backbone | MTP tensors |
|---|---:|---:|
| Q2_K | ~49.5 | ≥20 |
| Q3_K_M | ~62.5 | ≥20 |
| Q4_K_M | ~80.2 | ≥20 |
| Q5_K_M | ~91.3 | ≥20 |
| Q6_K | ~107.8 | ≥20 |
| PLE BF16 (shared) | ~95.4 | — |
Run (DGX Spark / unified memory)
Requires qwen4exp-capable llama.cpp (PR #27742 class).
llama-cli \
-m Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00001-of-00007.gguf \
--load-mode mmap \
-ngl 99 \
-ot "per_layer_token_embd.weight=CPU" \
-c 1024 -n 64 -st --temp 0 \
-p "The capital of France is"
Do not mlock the PLE on 128G unified-memory boxes. PLE stays NVMe-backed via mmap.
MTP speculative decode
Weights include MTP (nextn_predict_layers / blk.48). Runtime --spec-type draft-mtp needs a llama.cpp build with qwen4exp graph_mtp (not all mainline builds yet).
Build notes
- Converter: llama.cpp qwen4exp with MTP export enabled
- Imatrix: AtomicChat-compatible matrix
- PLE left BF16 (not re-quantized); tok emb Q8_0; output Q6_K
Run vcruz305/Qwen3.8-Flash-Next-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models