GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF overview

Qwen3.8 Flash Next W4A16 NVFP4 GGUF GGUF conversion of axiomofmind/Qwen3.8 Flash Next W4A16 NVFP4 https://huggingface.co/axiomofmind/Qwen3.8 Flash Next W4A16 N…

ggufqwenqwen4expnvfp4modeloptmoevision-languageimage-text-to-textbase_model:axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4base_model:quantized:axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4license:otherendpoints_compatibleregion:usconversational

Runs locally from ~865.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
image-text-to-text

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Next-W4A16-NVFP4-BF16attn-PLE-noMTP.ggufGGUFBF16168.00 GBDownload
mmproj-Qwen3.8-Flash-Next-BF16.ggufGGUFBF16865.5 MBDownload

Model Details

Model IDaxiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF
Authoraxiomofmind
Pipelineimage-text-to-text
Licenseother
Base modelaxiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4
Last modified2026-08-27T08:29:44.000Z

Model README

---

base_model: axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4

pipeline_tag: image-text-to-text

license: other

license_name: qwen-community-1.0

license_link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE

tags:

- gguf

- qwen

- qwen4exp

- nvfp4

- modelopt

- moe

- vision-language

---

Qwen3.8-Flash-Next W4A16 NVFP4 GGUF

GGUF conversion of

axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4,

based on Qwen/Qwen3.8-Flash-Next.

The routed-expert weights use W4A16 NVFP4. Attention, shared experts,

routers, embeddings, PLE, and other retained tensors remain in BF16 or F32.

Files

| File | Description | Size |

| --- | --- | ---: |

| Qwen3.8-Flash-Next-W4A16-NVFP4-BF16attn-PLE-noMTP.gguf | Main text model | 180.4 GB |

| mmproj-Qwen3.8-Flash-Next-BF16.gguf | BF16 vision projector | 907.5 MB |

The vision projector is required for image inputs. It is not required for

text-only use.

This GGUF does not include MTP weights.

Requirements

A llama.cpp build with Qwen4Exp and NVFP4 GGUF support is required.

For multimodal use, load the main model together with the included mmproj

file.

License

This model is distributed under the

Qwen Community License 1.0.

Refer to the official model card

for architecture details, usage guidance, and limitations.

Run axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models