0xKitkat/Qwen3.8-Flash-Next-GGUF overview
Qwen3.8 Flash Next GGUF GGUF quantizations of Qwen/Qwen3.8 Flash Next https://huggingface.co/Qwen/Qwen3.8 Flash Next will be published here after the official …
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Browse files on Hugging Face | ||||
Model Details
| Model ID | 0xKitkat/Qwen3.8-Flash-Next-GGUF |
|---|---|
| Author | 0xKitkat |
| Pipeline | image-text-to-text |
| License | — |
| Base model | Qwen/Qwen3.8-Flash-Next |
| Last modified | 2026-08-26T04:13:49.000Z |
Model README
---
base_model: Qwen/Qwen3.8-Flash-Next
language:
- en
- zh
tags:
- qwen
- qwen3.8
- gguf
- quantization
- multimodal
- moe
pipeline_tag: image-text-to-text
---
Qwen3.8-Flash-Next-GGUF
GGUF quantizations of Qwen/Qwen3.8-Flash-Next will be published here after the official weights and license are released and compatible conversion support is available.
> 🔔 Like and follow this repository for the GGUF release.
> 🐦 Follow @procrastiness on X for quantization updates.
Release status
The upstream model is currently listed by Qwen as an upcoming release scheduled for August 26, 2026. This repository does not contain model weights yet. It is being prepared for a real quantized release—not a renamed or unrelated checkpoint.
Qwen describes Qwen3.8-Flash-Next as:
- A preview of the next-generation Qwen4 architecture
- A multimodal Mixture-of-Experts (MoE) model
- An upcoming open release under the official model ID
Qwen/Qwen3.8-Flash-Next
Final architecture details, parameter counts, context length, license, chat template, multimodal projector requirements, and runtime compatibility will be copied from the official release—not inferred from rumors.
Planned GGUF files
The exact set will depend on the released architecture and practical file sizes. Intended variants include:
Q4_K_MQ5_K_MQ6_KQ8_0
If the model requires a separate multimodal projector, compatible mmproj files will also be provided when the conversion toolchain supports them.
Release checklist
- Verify the official source revision and license
- Convert directly from
Qwen/Qwen3.8-Flash-Next - Record the converter and
llama.cpprevisions - Validate the tokenizer and chat template
- Test text and multimodal inference where supported
- Publish checksums, file sizes, and memory guidance
- Credit Qwen and link the upstream model card
Expected usage
Usage commands will be added after compatibility is verified against an actual llama.cpp release. Commands will not be published before they can be tested with the released architecture.
Sources
Disclaimer
This is an independent community quantization project and is not affiliated with or endorsed by Qwen, Alibaba, or Hugging Face. The upstream model's license will govern redistribution and use of derived GGUF files.
Changelog
- 2026-08-26: Repository prepared ahead of the official upstream release; no weights uploaded yet.
Run 0xKitkat/Qwen3.8-Flash-Next-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models