AesSedai/Qwen3.8-Flash-Next-GGUF overview
Updates / Notes 09/02/2027: Added the PLEQ4 0 for IQ2 S and IQ3 S 09/01/2027: The original quants all have Q8 0 for the engram embedding tensors, and I've uplo…
Runs locally from ~10.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| IQ2_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ2_S-PLEQ4_0-00001-of-00003.gguf | GGUF | IQ2_S | 10.4 MB | Download |
| IQ2_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ2_S-PLEQ4_0-00002-of-00003.gguf | GGUF | IQ2_S | 46.41 GB | Download |
| IQ2_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ2_S-PLEQ4_0-00003-of-00003.gguf | GGUF | IQ2_S | 30.79 GB | Download |
| IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00001-of-00004.gguf | GGUF | IQ2_S | 10.4 MB | Download |
| IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00002-of-00004.gguf | GGUF | IQ2_S | 503.2 MB | Download |
| IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00003-of-00004.gguf | GGUF | IQ2_S | 50.66 GB | Download |
| IQ2_S/Qwen3.8-Flash-Next-IQ2_S-00004-of-00004.gguf | GGUF | IQ2_S | 46.51 GB | Download |
| IQ3_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ3_S-PLEQ4_0-00001-of-00003.gguf | GGUF | IQ3_S | 10.4 MB | Download |
| IQ3_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ3_S-PLEQ4_0-00002-of-00003.gguf | GGUF | IQ3_S | 46.20 GB | Download |
| IQ3_S-PLEQ4_0/Qwen3.8-Flash-Next-IQ3_S-PLEQ4_0-00003-of-00003.gguf | GGUF | IQ3_S | 33.35 GB | Download |
| IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00001-of-00005.gguf | GGUF | IQ3_S | 10.4 MB | Download |
| IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00002-of-00005.gguf | GGUF | IQ3_S | 503.2 MB | Download |
| IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00003-of-00005.gguf | GGUF | IQ3_S | 50.66 GB | Download |
| IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00004-of-00005.gguf | GGUF | IQ3_S | 46.56 GB | Download |
| IQ3_S/Qwen3.8-Flash-Next-IQ3_S-00005-of-00005.gguf | GGUF | IQ3_S | 2.29 GB | Download |
| IQ4_XS-PLEQ4_0/Qwen3.8-Flash-Next-IQ4_XS-PLEQ4_0-00001-of-00003.gguf | GGUF | IQ4_XS | 10.4 MB | Download |
| IQ4_XS-PLEQ4_0/Qwen3.8-Flash-Next-IQ4_XS-PLEQ4_0-00002-of-00003.gguf | GGUF | IQ4_XS | 46.27 GB | Download |
| IQ4_XS-PLEQ4_0/Qwen3.8-Flash-Next-IQ4_XS-PLEQ4_0-00003-of-00003.gguf | GGUF | IQ4_XS | 42.29 GB | Download |
| IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00005.gguf | GGUF | IQ4_XS | 10.4 MB | Download |
| IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00002-of-00005.gguf | GGUF | IQ4_XS | 650.8 MB | Download |
| IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00003-of-00005.gguf | GGUF | IQ4_XS | 50.66 GB | Download |
| IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00004-of-00005.gguf | GGUF | IQ4_XS | 46.37 GB | Download |
| IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00005-of-00005.gguf | GGUF | IQ4_XS | 11.42 GB | Download |
| Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00001-of-00004.gguf | GGUF | Q4_K_M | 10.4 MB | Download |
| Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00002-of-00004.gguf | GGUF | Q4_K_M | 46.55 GB | Download |
| Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00003-of-00004.gguf | GGUF | Q4_K_M | 46.14 GB | Download |
| Q4_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q4_K_M-PLEQ4_0-00004-of-00004.gguf | GGUF | Q4_K_M | 12.86 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00001-of-00005.gguf | GGUF | Q4_K_M | 10.4 MB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00002-of-00005.gguf | GGUF | Q4_K_M | 650.8 MB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00003-of-00005.gguf | GGUF | Q4_K_M | 50.66 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00004-of-00005.gguf | GGUF | Q4_K_M | 46.52 GB | Download |
| Q4_K_M/Qwen3.8-Flash-Next-Q4_K_M-00005-of-00005.gguf | GGUF | Q4_K_M | 28.26 GB | Download |
| Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00001-of-00004.gguf | GGUF | Q5_K_M | 10.4 MB | Download |
| Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00002-of-00004.gguf | GGUF | Q5_K_M | 46.56 GB | Download |
| Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00003-of-00004.gguf | GGUF | Q5_K_M | 46.11 GB | Download |
| Q5_K_M-PLEQ4_0/Qwen3.8-Flash-Next-Q5_K_M-PLEQ4_0-00004-of-00004.gguf | GGUF | Q5_K_M | 33.97 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00001-of-00006.gguf | GGUF | Q5_K_M | 10.4 MB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00002-of-00006.gguf | GGUF | Q5_K_M | 650.8 MB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00003-of-00006.gguf | GGUF | Q5_K_M | 50.66 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00004-of-00006.gguf | GGUF | Q5_K_M | 46.34 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00005-of-00006.gguf | GGUF | Q5_K_M | 46.45 GB | Download |
| Q5_K_M/Qwen3.8-Flash-Next-Q5_K_M-00006-of-00006.gguf | GGUF | Q5_K_M | 3.09 GB | Download |
| imatrix.gguf | GGUF | GGUF | 553.2 MB | Download |
| mmproj-Qwen3.8-Flash-Next-BF16.gguf | GGUF | BF16 | 865.5 MB | Download |
| mmproj-Qwen3.8-Flash-Next-F16.gguf | GGUF | F16 | 862.1 MB | Download |
| mmproj-Qwen3.8-Flash-Next-F32.gguf | GGUF | F32 | 1.67 GB | Download |
| mmproj-Qwen3.8-Flash-Next-Q8_0.gguf | GGUF | Q8_0 | 588.1 MB | Download |
Model Details
Model README
---
base_model:
- Qwen/Qwen3.8-Flash-Next
---
Updates / Notes
- 09/02/2027: Added the PLEQ4_0 for IQ2_S and IQ3_S
- 09/01/2027: The original quants all have Q8_0 for the engram embedding tensors, and I've uploaded three variants with Q4_0 for the embedded tensors. The only difference is the Q8_0 vs Q4_0 for those tensors. I'm leaving the original quants up for those who want to use the Q8_0 PLE's.
This repo contains specialized MoE-quants for Qwen/Qwen3.8-Flash-Next. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.
| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |
| :----- | :-------------------- | :------------------------------- | :------------------ | :------------------------ | :------------------ |
| Q5_K_M | 147.17 GiB (7.14 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 4.600285 ± 0.027992 | +0.2179% | 0.030432 ± 0.000210 |
| Q5_K_M (Q4_0 PLE) | 126.64 GiB (6.15 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 4.616165 ± 0.028059 | +0.5638% | 0.030814 ± 0.000217 |
| Q4_K_M | 126.08 GiB (6.12 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 4.618153 ± 0.028121 | +0.6071% | 0.041479 ± 0.000274 |
| IQ4_XS | 109.09 GiB (5.30 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 4.733298 ± 0.029068 | +3.1156% | 0.079730 ± 0.000510 |
| Q4_K_M (Q4_0 PLE) | 105.55 GiB (5.12 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 4.633212 ± 0.028203 | +0.9352% | 0.043324 ± 0.000287 |
| IQ3_S | 100.01 GiB (4.85 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 4.870320 ± 0.030101 | +6.1006% | 0.162764 ± 0.000918 |
| IQ2_S | 97.66 GiB (4.74 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 5.037301 ± 0.031450 | +9.7383% | 0.205950 ± 0.001125 |
| IQ4_XS (Q4_0 PLE) | 88.56 GiB (4.30 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 4.750635 ± 0.029142 | +3.4933% | 0.084118 ± 0.000515 |
| IQ3_S (Q4_0 PLE) | 79.55 GiB (3.86 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 4.887335 ± 0.030175 | +6.4713% | 0.167344 ± 0.000940 |
| IQ2_S (Q4_0 PLE) | 77.21 GiB (3.75 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 5.076957 ± 0.031749 | +10.6023% | 0.213340 ± 0.001164 |
Run AesSedai/Qwen3.8-Flash-Next-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models