GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

RemySkye/Qwen3.8-Flash-Next-REAM-60Pct-GGUF-56GiB overview

Qwen3.8 Flash Next REAM 60Pct GGUF 55GB This is a static GGUF quant of Akicou/Qwen3.8 Flash Next REAM 60Pct https://huggingface.co/Akicou/Qwen3.8 Flash Next RE…

ggufqwenqwen3.8base_model:Akicou/Qwen3.8-Flash-Next-REAM-60Pctbase_model:quantized:Akicou/Qwen3.8-Flash-Next-REAM-60Pctendpoints_compatibleregion:usconversational

Runs locally from ~6.42 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
Author

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
BF16/Qwen3.8-Flash-Next-REAM-60Pct-BF16-00001-of-00006.ggufGGUFBF166.42 GBDownload
BF16/Qwen3.8-Flash-Next-REAM-60Pct-BF16-00002-of-00006.ggufGGUFBF1695.37 GBDownload
BF16/Qwen3.8-Flash-Next-REAM-60Pct-BF16-00003-of-00006.ggufGGUFBF1642.36 GBDownload
BF16/Qwen3.8-Flash-Next-REAM-60Pct-BF16-00004-of-00006.ggufGGUFBF1642.65 GBDownload
BF16/Qwen3.8-Flash-Next-REAM-60Pct-BF16-00005-of-00006.ggufGGUFBF1642.40 GBDownload
BF16/Qwen3.8-Flash-Next-REAM-60Pct-BF16-00006-of-00006.ggufGGUFBF1610.79 GBDownload
Qwen3.8-Flash-Next-REAM-60Pct-56GiB-Q3_K_L.ggufGGUFQ3_K_L56.67 GBDownload
Qwen3.8-Flash-Next-REAM-60Pct-PLE2of16-Q3_K_L.gguf.ggufGGUFQ3_K_L33.20 GBDownload

Model Details

Model IDRemySkye/Qwen3.8-Flash-Next-REAM-60Pct-GGUF-56GiB
AuthorRemySkye
Pipeline
License
Base modelAkicou/Qwen3.8-Flash-Next-REAM-60Pct
Last modified2026-08-30T17:56:50.000Z

Model README

---

base_model:

  • Akicou/Qwen3.8-Flash-Next-REAM-60Pct

library_name: gguf

tags:

  • gguf
  • qwen
  • qwen3.8

---

Qwen3.8-Flash-Next-REAM-60Pct-GGUF-55GB

This is a static GGUF quant of Akicou/Qwen3.8-Flash-Next-REAM-60Pct.

I made this quant to keep the model weights below 64 GB so it has a better chance of running on systems with about 64 GB of combined memory.

The final quantized GGUF is:

  • 60.84 GB / 56.67 GiB
  • 3.779 effective BPW

Extra memory is still needed for the KV cache, compute buffers, the operating system, and other programs.

This model is heavily compressed. Quality may not be great compared with BF16 because there are deliberate precision cuts to make the model smaller.

No imatrix or calibration dataset was used. This is a direct static quantization.

Size comparison

| Model | Size | BPW |

|---|---:|---:|

| Qwen3.8-Flash-Next BF16 GGUF | 354.03 GB / 329.72 GiB | 16-bit weights |

| Qwen3.8-Flash-Next-REAM-60Pct BF16 GGUF | 257.67 GB / 239.97 GiB | 16-bit weights |

| This static quant | 60.84 GB / 56.67 GiB | 3.779 effective BPW |

The repo name uses 55GB as the target label. The exact finished size is shown above.

Quantization

  • PLE n-gram table: Q4_0
  • Routed expert gate/up: Q2_K
  • Routed expert down: Q4_0
  • Token embedding: Q4_K
  • Output head: Q4_K
  • Shared expert gate/up: Q3_K
  • Shared expert down: Q4_0
  • Most other compatible matrices: Q3_K
  • Small or unsupported tensors stay in BF16/F32 when required by llama.cpp

Q2_0 and IQ quants are not used.

The quant was made directly from the BF16 GGUF in the BF16 folder using llama.cpp commit cc231cb0da565440cf6a3e5b55dfeba477972cb6.

The current llama.cpp Qwen3.8 converter does not export the separate MTP draft head, so the BF16 GGUF numbers here refer to the llama.cpp text-model export.

Run RemySkye/Qwen3.8-Flash-Next-REAM-60Pct-GGUF-56GiB with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models