GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ddh0/DeepSeek-V4-Flash-GGUF overview

GGUF quantizations of DeepSeek V4 Flash. Using MTP requires am17an/llama.cpp:dsv4 mtp https://github.com/am17an/llama.cpp/tree/dsv4 mtp until llama.cpp 25784 h…

ggufbase_model:deepseek-ai/DeepSeek-V4-Flashbase_model:quantized:deepseek-ai/DeepSeek-V4-Flashendpoints_compatibleregion:usimatrixconversational

Runs locally from ~3.92 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
93
Likes
2
Pipeline
Author

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-3.86bpw.ggufGGUFGGUF127.82 GBDownload
DeepSeek-V4-Flash-MTP-Q4_0.ggufGGUFQ4_03.92 GBDownload
DeepSeek-V4-Flash-MTP-Q8_0.ggufGGUFQ8_04.41 GBDownload
DeepSeek-V4-Flash-MXFP4.ggufGGUFGGUF145.29 GBDownload

Model Details

Model IDddh0/DeepSeek-V4-Flash-GGUF
Authorddh0
Pipeline
License
Base modeldeepseek-ai/DeepSeek-V4-Flash
Last modified2026-07-16T21:25:00.000Z

Model README

---

base_model:

  • deepseek-ai/DeepSeek-V4-Flash

---

GGUF quantizations of DeepSeek-V4-Flash.

Using MTP requires am17an/llama.cpp:dsv4-mtp until llama.cpp#25784 is merged.

Run ddh0/DeepSeek-V4-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models