GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AesSedai/Qwen3.8-Flash-Next-GGUF overview

Updates / Notes 09/02/2027: Added the PLEQ4 0 for IQ2 S and IQ3 S 09/01/2027: The original quants all have Q8 0 for the engram embedding tensors, and I've uplo…

ggufbase_model:Qwen/Qwen3.8-Flash-Nextbase_model:quantized:Qwen/Qwen3.8-Flash-Nextendpoints_compatibleregion:usimatrixconversational

Runs locally from ~10.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2,318
Likes
15
Pipeline
Author

Repository Files & Downloads

47 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
IQ2_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ2_S-PLEQ4_0-00001-of-00003.ggufGGUFIQ2_S10.4 MBDownload
IQ2_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ2_S-PLEQ4_0-00002-of-00003.ggufGGUFIQ2_S46.41 GBDownload
IQ2_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ2_S-PLEQ4_0-00003-of-00003.ggufGGUFIQ2_S30.79 GBDownload
IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00001-of-00004.ggufGGUFIQ2_S10.4 MBDownload
IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00002-of-00004.ggufGGUFIQ2_S503.2 MBDownload
IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00003-of-00004.ggufGGUFIQ2_S50.66 GBDownload
IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00004-of-00004.ggufGGUFIQ2_S46.51 GBDownload
IQ3_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ3_S-PLEQ4_0-00001-of-00003.ggufGGUFIQ3_S10.4 MBDownload
IQ3_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ3_S-PLEQ4_0-00002-of-00003.ggufGGUFIQ3_S46.20 GBDownload
IQ3_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ3_S-PLEQ4_0-00003-of-00003.ggufGGUFIQ3_S33.35 GBDownload
IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00001-of-00005.ggufGGUFIQ3_S10.4 MBDownload
IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00002-of-00005.ggufGGUFIQ3_S503.2 MBDownload
IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00003-of-00005.ggufGGUFIQ3_S50.66 GBDownload
IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00004-of-00005.ggufGGUFIQ3_S46.56 GBDownload
IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00005-of-00005.ggufGGUFIQ3_S2.29 GBDownload
IQ4_XS-PLEQ4_0/Qwen3.8-Flash-Next-IQ4_XS-PLEQ4_0-00001-of-00003.ggufGGUFIQ4_XS10.4 MBDownload
IQ4_XS-PLEQ4_0/Qwen3.8-Flash-Next-IQ4_XS-PLEQ4_0-00002-of-00003.ggufGGUFIQ4_XS46.27 GBDownload
IQ4_XS-PLEQ4_0/Qwen3.8-Flash-Next-IQ4_XS-PLEQ4_0-00003-of-00003.ggufGGUFIQ4_XS42.29 GBDownload
IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00005.ggufGGUFIQ4_XS10.4 MBDownload
IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00002-of-00005.ggufGGUFIQ4_XS650.8 MBDownload
IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00003-of-00005.ggufGGUFIQ4_XS50.66 GBDownload
IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00004-of-00005.ggufGGUFIQ4_XS46.37 GBDownload
IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00005-of-00005.ggufGGUFIQ4_XS11.42 GBDownload
Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00001-of-00004.ggufGGUFQ4_K_M10.4 MBDownload
Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00002-of-00004.ggufGGUFQ4_K_M46.55 GBDownload
Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00003-of-00004.ggufGGUFQ4_K_M46.14 GBDownload
Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00004-of-00004.ggufGGUFQ4_K_M12.86 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00001-of-00005.ggufGGUFQ4_K_M10.4 MBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00002-of-00005.ggufGGUFQ4_K_M650.8 MBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00003-of-00005.ggufGGUFQ4_K_M50.66 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00004-of-00005.ggufGGUFQ4_K_M46.52 GBDownload
Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00005-of-00005.ggufGGUFQ4_K_M28.26 GBDownload
Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00001-of-00004.ggufGGUFQ5_K_M10.4 MBDownload
Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00002-of-00004.ggufGGUFQ5_K_M46.56 GBDownload
Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00003-of-00004.ggufGGUFQ5_K_M46.11 GBDownload
Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00004-of-00004.ggufGGUFQ5_K_M33.97 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00001-of-00006.ggufGGUFQ5_K_M10.4 MBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00002-of-00006.ggufGGUFQ5_K_M650.8 MBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00003-of-00006.ggufGGUFQ5_K_M50.66 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00004-of-00006.ggufGGUFQ5_K_M46.34 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00005-of-00006.ggufGGUFQ5_K_M46.45 GBDownload
Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00006-of-00006.ggufGGUFQ5_K_M3.09 GBDownload
imatrix.ggufGGUFGGUF553.2 MBDownload
mmproj-Qwen3.8-Flash-Next-BF16.ggufGGUFBF16865.5 MBDownload
mmproj-Qwen3.8-Flash-Next-F16.ggufGGUFF16862.1 MBDownload
mmproj-Qwen3.8-Flash-Next-F32.ggufGGUFF321.67 GBDownload
mmproj-Qwen3.8-Flash-Next-Q8_0.ggufGGUFQ8_0588.1 MBDownload

Model Details

Model IDAesSedai/Qwen3.8-Flash-Next-GGUF
AuthorAesSedai
Pipeline
License
Base modelQwen/Qwen3.8-Flash-Next
Last modified2026-09-02T18:15:39.000Z

Model README

---

base_model:

  • Qwen/Qwen3.8-Flash-Next

---

Updates / Notes

  • 09/02/2027: Added the PLEQ4_0 for IQ2_S and IQ3_S
  • 09/01/2027: The original quants all have Q8_0 for the engram embedding tensors, and I've uploaded three variants with Q4_0 for the embedded tensors. The only difference is the Q8_0 vs Q4_0 for those tensors. I'm leaving the original quants up for those who want to use the Q8_0 PLE's.

This repo contains specialized MoE-quants for Qwen/Qwen3.8-Flash-Next. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |

| :----- | :-------------------- | :------------------------------- | :------------------ | :------------------------ | :------------------ |

| Q5_K_M | 147.17 GiB (7.14 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 4.600285 ± 0.027992 | +0.2179% | 0.030432 ± 0.000210 |

| Q5_K_M (Q4_0 PLE) | 126.64 GiB (6.15 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 4.616165 ± 0.028059 | +0.5638% | 0.030814 ± 0.000217 |

| Q4_K_M | 126.08 GiB (6.12 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 4.618153 ± 0.028121 | +0.6071% | 0.041479 ± 0.000274 |

| IQ4_XS | 109.09 GiB (5.30 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 4.733298 ± 0.029068 | +3.1156% | 0.079730 ± 0.000510 |

| Q4_K_M (Q4_0 PLE) | 105.55 GiB (5.12 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 4.633212 ± 0.028203 | +0.9352% | 0.043324 ± 0.000287 |

| IQ3_S | 100.01 GiB (4.85 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 4.870320 ± 0.030101 | +6.1006% | 0.162764 ± 0.000918 |

| IQ2_S | 97.66 GiB (4.74 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 5.037301 ± 0.031450 | +9.7383% | 0.205950 ± 0.001125 |

| IQ4_XS (Q4_0 PLE) | 88.56 GiB (4.30 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 4.750635 ± 0.029142 | +3.4933% | 0.084118 ± 0.000515 |

| IQ3_S (Q4_0 PLE) | 79.55 GiB (3.86 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 4.887335 ± 0.030175 | +6.4713% | 0.167344 ± 0.000940 |

| IQ2_S (Q4_0 PLE) | 77.21 GiB (3.75 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 5.076957 ± 0.031749 | +10.6023% | 0.213340 ± 0.001164 |

!kld_graph

!ppl_graph

Run AesSedai/Qwen3.8-Flash-Next-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models