GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

bartowski/DeepSeek-V4-Flash-GGUF overview

Llamacpp Quantizations of DeepSeek V4 Flash by deepseek ai Using <a href="https://github.com/ggml org/llama.cpp/" llama.cpp</a release <a href="https://github.…

gguftext-generationbase_model:deepseek-ai/DeepSeek-V4-Flashbase_model:quantized:deepseek-ai/DeepSeek-V4-Flashlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~34.64 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-MXFP4/DeepSeek-V4-Flash-MXFP4-00001-of-00004.ggufGGUFGGUF36.83 GBDownload
DeepSeek-V4-Flash-MXFP4/DeepSeek-V4-Flash-MXFP4-00002-of-00004.ggufGGUFGGUF36.91 GBDownload
DeepSeek-V4-Flash-MXFP4/DeepSeek-V4-Flash-MXFP4-00003-of-00004.ggufGGUFGGUF36.91 GBDownload
DeepSeek-V4-Flash-MXFP4/DeepSeek-V4-Flash-MXFP4-00004-of-00004.ggufGGUFGGUF34.64 GBDownload

Model Details

Model IDbartowski/DeepSeek-V4-Flash-GGUF
Authorbartowski
Pipelinetext-generation
Licensemit
Base modeldeepseek-ai/DeepSeek-V4-Flash
Last modified2026-06-30T03:17:40.000Z

Model README

---

quantized_by: bartowski

pipeline_tag: text-generation

license: mit

base_model_relation: quantized

base_model: deepseek-ai/DeepSeek-V4-Flash

---

Llamacpp Quantizations of DeepSeek-V4-Flash by deepseek-ai

Using <a href="https://github.com/ggml-org/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggml-org/llama.cpp/releases/tag/b9843">b9843</a> for quantization.

Original model: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash

This model is in MXFP4 and as such has only been provided in MXFP4 format!

No other sizes can be provided unfortunately as MXFP4 does not quantize properly.

Run in your choice of tools:

Note: since it's a newly supported model, you may need to wait for an update from the developers.

Prompt format

No prompt format found

Download the MXFP4 files:

| Filename | Quant type | File Size | Split | Description |

| -------- | ---------- | --------- | ----- | ----------- |

| DeepSeek-V4-Flash-MXFP4.gguf | MXFP4 | 156.00GB | true | Original quality. |

Downloading using huggingface-cli

<details>

<summary>Click to view download instructions</summary>

First, make sure you have hugginface-cli installed:

pip install -U "huggingface_hub[cli]"
huggingface-cli download bartowski/DeepSeek-V4-Flash-GGUF --include "DeepSeek-V4-Flash-MXFP4*" --local-dir ./

You can either specify a new local-dir (DeepSeek-V4-Flash-MXFP4) or download them all in place (./)

</details>

Credits

Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.

Thank you ZeroWw for the inspiration to experiment with embed/output.

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

Run bartowski/DeepSeek-V4-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models