GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sokann/DeepSeek-V4-Flash-GGUF overview

DeepSeek V4 Flash GGUF This is a GGUF model for DeepSeek V4 Flash, created with the convert hf to gguf.py script from the merged commit of https://github.com/g…

ggufdeepseek_v4conversationalik_llama.cppllama.cppbase_model:deepseek-ai/DeepSeek-V4-Flashbase_model:quantized:deepseek-ai/DeepSeek-V4-Flashlicense:mitendpoints_compatibleregion:us

Runs locally from ~145.64 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
1,937
Likes
3
Pipeline
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash.ggufGGUFGGUF145.64 GBDownload

Model Details

Model IDsokann/DeepSeek-V4-Flash-GGUF
Authorsokann
Pipeline
Licensemit
Base modeldeepseek-ai/DeepSeek-V4-Flash
Last modified2026-07-03T23:03:24.000Z

Model README

---

base_model: deepseek-ai/DeepSeek-V4-Flash

base_model_relation: quantized

license: mit

license_link: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/LICENSE

tags:

  • deepseek_v4
  • conversational
  • ik_llama.cpp
  • llama.cpp

---

DeepSeek-V4-Flash-GGUF

This is a GGUF model for DeepSeek-V4-Flash, created with the convert_hf_to_gguf.py script from the merged commit of https://github.com/ggml-org/llama.cpp/pull/24162 (8c146a83, b9840), thanks to the amazing works from u/fairydreaming and u/am17an.

(The weights are identical with the one uploaded previously; the only difference is that the chat template is now included in the GGUF)

The script does 2 things:

  • Repacks routed experts tensors in FP4 to MXFP4.
  • Quantizes other tensors in FP8 to Q8_0.

The GGUF model can be considered as lossless when compared to the original weights in safetensors.

The size of the model is about 146 GiB. It fits comfortably on a machine with 160 GiB of RAM and 48 GiB of VRAM, and runs reasonably well:

$ build/bin/llama-server -m ~/models/DeepSeek-V4-Flash.gguf -cram 0 --jinja --no-mmap -dev CUDA0,CUDA1 --fit on -c 32768 -b 2048 -ub 2048
...
3.23.716.821 I slot print_timing: id  3 | task 0 | prompt eval time =   15977.50 ms /  5407 tokens (    2.95 ms per token,   338.41 tokens per second)
3.23.716.827 I slot print_timing: id  3 | task 0 |        eval time =   43043.40 ms /   586 tokens (   73.45 ms per token,    13.61 tokens per second)
3.23.716.828 I slot print_timing: id  3 | task 0 |       total time =   59020.89 ms /  5993 tokens

1M context

And after applying the lightning indexer change from https://github.com/ggml-org/llama.cpp/pull/24231, we can have 1M context with just 6GiB of VRAM!

To join the One million tokens prompt club, ~~apply lcpp-dsv4-lid-combo.diff~~ check out from https://github.com/fairydreaming/llama.cpp/tree/dsv4 directly!

The speed is decent on my machine with 2 x RTX PRO 4000:

$ ~/repo/llama.cpp/build/bin/llama-batched-bench -m ~/models/DeepSeek-V4-Flash-r3.gguf -b 2048 -ub 2048 -npl 1 -npp 8192,16384,32768,65536,131072,262144,524288,1048064 -ntg 128 -fa 1 -cmoe --no-repack --no-mmap

llama_batched_bench: n_kv_max = 1048576, n_batch = 2048, n_ubatch = 2048, flash_attn = 1, is_pp_shared = 0, is_tg_separate = 0, n_gpu_layers = -1, n_threads = 112, n_threads_batch = 112

|    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
|-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
|  8192 |    128 |    1 |   8320 |   18.262 |   448.57 |    6.668 |    19.20 |   24.931 |   333.73 |
| 16384 |    128 |    1 |  16512 |   37.376 |   438.36 |    6.671 |    19.19 |   44.047 |   374.87 |
| 32768 |    128 |    1 |  32896 |   81.523 |   401.95 |    6.959 |    18.39 |   88.482 |   371.78 |
| 65536 |    128 |    1 |  65664 |  187.222 |   350.04 |    7.347 |    17.42 |  194.570 |   337.48 |
|131072 |    128 |    1 | 131200 |  468.701 |   279.65 |    8.153 |    15.70 |  476.854 |   275.14 |
|262144 |    128 |    1 | 262272 | 1334.002 |   196.51 |    9.837 |    13.01 | 1343.839 |   195.17 |
|524288 |    128 |    1 | 524416 | 4348.049 |   120.58 |   13.055 |     9.80 | 4361.104 |   120.25 |
|1048064 |    128 |    1 | 1048192 | 15886.781 |    65.97 |   19.503 |     6.56 | 15906.284 |    65.90 |

At long context, PP and TG started to converge. But we made it to 1M!

Previously, without the hc-ops, the speed wasn't as good, especially at long context, and I had to stop it after 500k:

|    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
|-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
|  8192 |    128 |    1 |   8320 |   25.401 |   322.50 |    8.754 |    14.62 |   34.155 |   243.60 |
| 16384 |    128 |    1 |  16512 |   58.092 |   282.03 |    8.827 |    14.50 |   66.920 |   246.74 |
| 32768 |    128 |    1 |  32896 |  148.657 |   220.43 |    9.158 |    13.98 |  157.815 |   208.45 |
| 65536 |    128 |    1 |  65664 |  426.688 |   153.59 |   10.113 |    12.66 |  436.801 |   150.33 |
|131072 |    128 |    1 | 131200 | 1373.003 |    95.46 |   11.775 |    10.87 | 1384.778 |    94.74 |
|262144 |    128 |    1 | 262272 | 4838.849 |    54.17 |   15.295 |     8.37 | 4854.144 |    54.03 |
|524288 |    128 |    1 | 524416 | 18152.150 |    28.88 |   22.333 |     5.73 | 18174.482 |    28.85 |

KLD comparison

Using wiki.test.raw with -c 512, llama-perplexity was used to first generate the base logits using this near-lossless quant, and then used to do measure the KLD of the 2 quants from antirez:

These 2 quants work amazingly well. However, their KLD scores are quite bad.

Summary for the Q2 quant:

Mean PPL(Q)                   :   6.155042 ±   0.042084
Mean PPL(base)                :   4.738434 ±   0.030638
Cor(ln(PPL(Q)), ln(PPL(base))):  87.64%
...
Mean    KLD:   0.422277 ±   0.002287
Maximum KLD:  15.166180
99.9%   KLD:   8.191362
...
RMS Δp    : 21.606 ± 0.086 %
Same top p: 77.586 ± 0.110 %

Summary for the Q4 quant:

Mean PPL(Q)                   :   4.570599 ±   0.029570
Mean PPL(base)                :   4.738434 ±   0.030638
Cor(ln(PPL(Q)), ln(PPL(base))):  96.77%
...
Mean    KLD:   0.115089 ±   0.000817
Maximum KLD:   9.996851
99.9%   KLD:   3.656791
...
RMS Δp    : 11.098 ± 0.065 %
Same top p: 88.794 ± 0.083 %

Not sure I did something wrong in the measurement.

Run sokann/DeepSeek-V4-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models