axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF overview
Qwen3.8 Flash Next W4A16 NVFP4 GGUF GGUF conversion of axiomofmind/Qwen3.8 Flash Next W4A16 NVFP4 https://huggingface.co/axiomofmind/Qwen3.8 Flash Next W4A16 N…
Runs locally from ~865.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF |
|---|---|
| Author | axiomofmind |
| Pipeline | image-text-to-text |
| License | other |
| Base model | axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4 |
| Last modified | 2026-08-27T08:29:44.000Z |
Model README
---
base_model: axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4
pipeline_tag: image-text-to-text
license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE
tags:
- gguf
- qwen
- qwen4exp
- nvfp4
- modelopt
- moe
- vision-language
---
Qwen3.8-Flash-Next W4A16 NVFP4 GGUF
GGUF conversion of
axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4,
based on Qwen/Qwen3.8-Flash-Next.
The routed-expert weights use W4A16 NVFP4. Attention, shared experts,
routers, embeddings, PLE, and other retained tensors remain in BF16 or F32.
Files
| File | Description | Size |
| --- | --- | ---: |
| Qwen3.8-Flash-Next-W4A16-NVFP4-BF16attn-PLE-noMTP.gguf | Main text model | 180.4 GB |
| mmproj-Qwen3.8-Flash-Next-BF16.gguf | BF16 vision projector | 907.5 MB |
The vision projector is required for image inputs. It is not required for
text-only use.
This GGUF does not include MTP weights.
Requirements
A llama.cpp build with Qwen4Exp and NVFP4 GGUF support is required.
For multimodal use, load the main model together with the included mmproj
file.
License
This model is distributed under the
Refer to the official model card
for architecture details, usage guidance, and limitations.
Run axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models