GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

vcruz305/Qwen3.8-Flash-Next-GGUF overview

Qwen3.8 Flash Next GGUF vcruz305 Real K quant ladder from full HF BF16 weights with MTP heads in main blk.48 / nextn. and a shared BF16 PLE . Layout | path | c…

ggufqwenqwen3.8flash-nextllama.cppmtpk-quantbase_model:Qwen/Qwen3.8-Flash-Nextbase_model:quantized:Qwen/Qwen3.8-Flash-Nextlicense:otherregion:us

Runs locally from ~553.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

35 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
MTP/Qwen3.8-Flash-Next-MTP-Q4_K_M.ggufGGUFQ4_K_M2.59 GBDownload
MTP/Qwen3.8-Flash-Next-MTP-Q8_0.ggufGGUFQ8_03.85 GBDownload
MTP/Qwen3.8-Flash-Next-MTP.ggufGGUFGGUF7.24 GBDownload
PLE-BF16/Qwen3.8-Flash-Next-PLE-BF16.ggufGGUFBF1695.37 GBDownload
Q2_K/Qwen3.8-Flash-Next-Q2_K-00001-of-00007.ggufGGUFQ2_K9.19 GBDownload
Q2_K/Qwen3.8-Flash-Next-Q2_K-00002-of-00007.ggufGGUFQ2_K9.57 GBDownload
Q2_K/Qwen3.8-Flash-Next-Q2_K-00003-of-00007.ggufGGUFQ2_K9.17 GBDownload
Q2_K/Qwen3.8-Flash-Next-Q2_K-00004-of-00007.ggufGGUFQ2_K9.15 GBDownload
Q2_K/Qwen3.8-Flash-Next-Q2_K-00005-of-00007.ggufGGUFQ2_K9.37 GBDownload
Q2_K/Qwen3.8-Flash-Next-Q2_K-00006-of-00007.ggufGGUFQ2_K3.07 GBDownload
Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00001-of-00007.ggufGGUFQ3_K_M11.73 GBDownload
Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00002-of-00007.ggufGGUFQ3_K_M12.14 GBDownload
Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00003-of-00007.ggufGGUFQ3_K_M11.60 GBDownload
Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00004-of-00007.ggufGGUFQ3_K_M11.58 GBDownload
Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00005-of-00007.ggufGGUFQ3_K_M11.83 GBDownload
Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00006-of-00007.ggufGGUFQ3_K_M3.58 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00001-of-00007.ggufGGUFQ4_K_M14.88 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00002-of-00007.ggufGGUFQ4_K_M15.59 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00003-of-00007.ggufGGUFQ4_K_M14.82 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00004-of-00007.ggufGGUFQ4_K_M14.50 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00005-of-00007.ggufGGUFQ4_K_M16.15 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00006-of-00007.ggufGGUFQ4_K_M4.30 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00001-of-00007.ggufGGUFQ5_K_M17.04 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00002-of-00007.ggufGGUFQ5_K_M17.75 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00003-of-00007.ggufGGUFQ5_K_M16.99 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00004-of-00007.ggufGGUFQ5_K_M16.72 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00005-of-00007.ggufGGUFQ5_K_M18.07 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00006-of-00007.ggufGGUFQ5_K_M4.75 GBDownload
Q6_K/Qwen3.8-Flash-Next-Q6_K-00001-of-00007.ggufGGUFQ6_K20.32 GBDownload
Q6_K/Qwen3.8-Flash-Next-Q6_K-00002-of-00007.ggufGGUFQ6_K21.02 GBDownload
Q6_K/Qwen3.8-Flash-Next-Q6_K-00003-of-00007.ggufGGUFQ6_K20.28 GBDownload
Q6_K/Qwen3.8-Flash-Next-Q6_K-00004-of-00007.ggufGGUFQ6_K20.24 GBDownload
Q6_K/Qwen3.8-Flash-Next-Q6_K-00005-of-00007.ggufGGUFQ6_K20.51 GBDownload
Q6_K/Qwen3.8-Flash-Next-Q6_K-00006-of-00007.ggufGGUFQ6_K5.42 GBDownload
imatrix/Qwen3.8-Flash-Next.imatrix.ggufGGUFGGUF553.2 MBDownload

Model Details

Model IDvcruz305/Qwen3.8-Flash-Next-GGUF
Authorvcruz305
Pipeline
Licenseother
Base modelQwen/Qwen3.8-Flash-Next
Last modified2026-08-27T18:50:33.000Z

Model README

---

license: other

base_model: Qwen/Qwen3.8-Flash-Next

tags:

- gguf

- qwen

- qwen3.8

- flash-next

- llama.cpp

- mtp

- k-quant

---

Qwen3.8-Flash-Next GGUF (vcruz305)

Real K-quant ladder from full HF BF16 weights with MTP heads in-main (blk.48 / nextn.*) and a shared BF16 PLE.

Layout

| path | content |

|---|---|

| Q2_K/Q6_K/ | Backbone shards 00001–00006 (MTP included in tensor set) |

| PLE-BF16/Qwen3.8-Flash-Next-PLE-BF16.gguf | Shared BF16 PLE (~95.4 GiB) — use as each quant's 00007 |

| imatrix/ | Importance matrix used for Q2–Q6 |

| scripts/link-ple.sh | ln/cp PLE → …-00007-of-00007.gguf per quant |

| MTP/ | Optional standalone MTP draft heads (if present) |

Setup

bash scripts/link-ple.sh

Approximate backbone sizes (GiB) + shared PLE

| quant | backbone | MTP tensors |

|---|---:|---:|

| Q2_K | ~49.5 | ≥20 |

| Q3_K_M | ~62.5 | ≥20 |

| Q4_K_M | ~80.2 | ≥20 |

| Q5_K_M | ~91.3 | ≥20 |

| Q6_K | ~107.8 | ≥20 |

| PLE BF16 (shared) | ~95.4 | — |

Run (DGX Spark / unified memory)

Requires qwen4exp-capable llama.cpp (PR #27742 class).

llama-cli \
  -m Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00001-of-00007.gguf \
  --load-mode mmap \
  -ngl 99 \
  -ot "per_layer_token_embd.weight=CPU" \
  -c 1024 -n 64 -st --temp 0 \
  -p "The capital of France is"

Do not mlock the PLE on 128G unified-memory boxes. PLE stays NVMe-backed via mmap.

MTP speculative decode

Weights include MTP (nextn_predict_layers / blk.48). Runtime --spec-type draft-mtp needs a llama.cpp build with qwen4exp graph_mtp (not all mainline builds yet).

Build notes

  • Converter: llama.cpp qwen4exp with MTP export enabled
  • Imatrix: AtomicChat-compatible matrix
  • PLE left BF16 (not re-quantized); tok emb Q8_0; output Q6_K

Run vcruz305/Qwen3.8-Flash-Next-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models