GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ggml-org/DeepSeek-V4-Flash-GGUF overview

DeepSeek V4 Flash Run with https://llama.app bash llama serve hf ggml org/DeepSeek V4 Flash GGUF Source models https://huggingface.co/deepseek ai/DeepSeek V4 F…

ggufquantizedtext-generationbase_model:deepseek-ai/DeepSeek-V4-Flashbase_model:quantized:deepseek-ai/DeepSeek-V4-Flashlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~5.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2,641
Likes
9
Pipeline
text-generation
Author

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-MXFP4-00001-of-00002.ggufGGUFGGUF5.0 MBDownload
DeepSeek-V4-Flash-MXFP4-00002-of-00002.ggufGGUFGGUF144.34 GBDownload
DeepSeek-V4-Flash-Q2_K-00001-of-00002.ggufGGUFQ2_K5.0 MBDownload
DeepSeek-V4-Flash-Q2_K-00002-of-00002.ggufGGUFQ2_K109.29 GBDownload
DeepSeek-V4-Flash-Q2_K_S-00001-of-00002.ggufGGUFQ2_K_S5.0 MBDownload
DeepSeek-V4-Flash-Q2_K_S-00002-of-00002.ggufGGUFQ2_K_S91.82 GBDownload
mtp-DeepSeek-V4-Flash-BF16.ggufGGUFBF165.33 GBDownload
mtp-DeepSeek-V4-Flash-MXFP4.ggufGGUFGGUF4.41 GBDownload

Model Details

Model IDggml-org/DeepSeek-V4-Flash-GGUF
Authorggml-org
Pipelinetext-generation
Licensemit
Base modeldeepseek-ai/DeepSeek-V4-Flash
Last modified2026-08-06T11:41:36.000Z

Model README

---

license: mit

pipeline_tag: text-generation

tags:

  • gguf
  • quantized

base_model:

  • deepseek-ai/DeepSeek-V4-Flash

---

DeepSeek-V4-Flash

Run with https://llama.app

llama serve -hf ggml-org/DeepSeek-V4-Flash-GGUF

Source models

  • https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash

Notes

  • Currently, the Q2 models do not use an imatrix calibration due to lack of one.

TODOs

  • add info

> [!IMPORTANT]

> This model is automatically converted using https://github.com/ggml-org/convert

Run ggml-org/DeepSeek-V4-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models