GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ddh0/DeepSeek-V4-Flash-GGUF overview

GGUF quantizations of DeepSeek V4 Flash. Using MTP requires am17an/llama.cpp:dsv4 mtp https://github.com/am17an/llama.cpp/tree/dsv4 mtp until llama.cpp 25784 h…

ggufbase_model:deepseek-ai/DeepSeek-V4-Flashbase_model:quantized:deepseek-ai/DeepSeek-V4-Flashendpoints_compatibleregion:usimatrixconversational

Runs locally from ~3.51 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6,168
Likes
7
Pipeline
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-3.86bpw-IQ3_S.ggufGGUFIQ3_S127.82 GBDownload
DeepSeek-V4-Flash-3.86bpw-Q3_K.ggufGGUFQ3_K127.82 GBDownload
DeepSeek-V4-Flash-MTP-3.93bpw.ggufGGUFGGUF3.51 GBDownload
DeepSeek-V4-Flash-MTP-4.93bpw.ggufGGUFGGUF4.41 GBDownload
DeepSeek-V4-Flash-MXFP4.ggufGGUFGGUF145.29 GBDownload

Model Details

Model IDddh0/DeepSeek-V4-Flash-GGUF
Authorddh0
Pipeline
License
Base modeldeepseek-ai/DeepSeek-V4-Flash
Last modified2026-08-07T01:34:35.000Z

Model README

---

base_model:

  • deepseek-ai/DeepSeek-V4-Flash

---

GGUF quantizations of DeepSeek-V4-Flash.

Using MTP requires am17an/llama.cpp:dsv4-mtp until llama.cpp#25784 is merged.

Run ddh0/DeepSeek-V4-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models