GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

grapeV-ai/Qwen3.8-Flash-Next-GGUF overview

What is this? Qwen4のアーキテクチャを先取り! Qwen3.8 Flash Next https://huggingface.co/Qwen/Qwen3.8 Flash Next をGGUFフォーマットに変換したものです。 imatrix dataset 日本語能力を重視し、日本語が多量に含まれる …

gguflicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~553.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
55
Likes
0
Pipeline
Author

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Next-512x56B-BF16.ggufGGUFBF16329.72 GBDownload
Qwen3.8-Flash-Next-IQ4_XS.ggufGGUFIQ4_XS90.78 GBDownload
Qwen3.8-Flash-Next-MXFP4_MOE.ggufGGUFGGUF115.53 GBDownload
Qwen3.8-Flash-Next-Q4_K_M.ggufGGUFQ4_K_M110.97 GBDownload
Qwen3.8-Flash-Next-Q5_K_M.ggufGGUFQ5_K_M124.90 GBDownload
Qwen3.8-Flash-Next-imatrix.ggufGGUFGGUF553.2 MBDownload
mmproj-Qwen3.8-Flash-Next-BF16.ggufGGUFBF16865.5 MBDownload
mmproj-Qwen3.8-Flash-Next-Q8_0.ggufGGUFQ8_0588.1 MBDownload

Model Details

Model IDgrapeV-ai/Qwen3.8-Flash-Next-GGUF
AuthorgrapeV-ai
Pipeline
Licenseother
Base model
Last modified2026-08-31T13:54:32.000Z

Model README

---

license: other

license_name: qwen-community-1.0

license_link: LICENSE

---

What is this?

Qwen4のアーキテクチャを先取り!Qwen3.8-Flash-NextをGGUFフォーマットに変換したものです。

imatrix dataset

日本語能力を重視し、日本語が多量に含まれるTFMC/imatrix-dataset-for-japanese-llmデータセットを使用しました。<br>

なお、計算リソースの関係上imatrixの算出にはQ6_K量子化モデルを使用しました。

Quants

各クオンツ・推論努力とそのベンチマークスコア(API版Gemma4 31B採点によるElyza_tasks 100)をまとめておきます。

|クオンツ|スコア|コメント|

|---|---|---|

|Q5_K_M(No Think)|4.53||

|Q4_K_M(No Think)|4.52||

|IQ4_XS(No Think)|4.55||

|MXFP4(No Think)|4.43||

||||

|reasoning_strength: low|4.5||

|reasoning_strength: medium|4.515||

|reasoning_strength: xhigh|4.54||

Note

-mm mmproj-Qwen3.8-Flash-Next-BF16.ggufでビジョンエンコーダーをロードし、Vision対応モデルとして使用することができます。

推論努力(Reasoning effort)はlow / medium / xhighから選択可能で、デフォルトはxhighです。<br>

llama.cppのserverの場合は起動時に以下の引数を追加することで変更可能です。

--chat-template-kwargs '{\"reasoning_effort\":\"xhigh\"}'

License

qwen-community-1.0

Developer

Alibaba Cloud

Run grapeV-ai/Qwen3.8-Flash-Next-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models