GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

6block/DeepSeek-V4-Pro-0813-GGUF overview

DeepSeek V4 Pro 0813 GGUF Sub 4 bit GGUF quantizations of deepseek ai/DeepSeek V4 Pro 0813 https://huggingface.co/deepseek ai/DeepSeek V4 Pro 0813 , produced b…

ggufdeepseekdeepseek_v4imatrixmoe6blocktext-generationenzhbase_model:deepseek-ai/DeepSeek-V4-Pro-0813base_model:quantized:deepseek-ai/DeepSeek-V4-Pro-0813license:mitendpoints_compatibleregion:usconversational

Runs locally from ~1.57 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
3,410
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

50 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Pro-0813-IQ1_M.ggufGGUFIQ1_M346.58 GBDownload
DeepSeek-V4-Pro-0813-IQ1_S.ggufGGUFIQ1_S314.10 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00001-of-00015.ggufGGUFIQ3_XXS39.46 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00002-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00003-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00004-of-00015.ggufGGUFIQ3_XXS41.10 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00005-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00006-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00007-of-00015.ggufGGUFIQ3_XXS41.10 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00008-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00009-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00010-of-00015.ggufGGUFIQ3_XXS41.10 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00011-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00012-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00013-of-00015.ggufGGUFIQ3_XXS41.10 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00014-of-00015.ggufGGUFIQ3_XXS40.78 GBDownload
DeepSeek-V4-Pro-0813-IQ3_XXS-00015-of-00015.ggufGGUFIQ3_XXS6.10 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00001-of-00014.ggufGGUFQ2_K40.90 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00002-of-00014.ggufGGUFQ2_K41.32 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00003-of-00014.ggufGGUFQ2_K41.79 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00004-of-00014.ggufGGUFQ2_K38.70 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00005-of-00014.ggufGGUFQ2_K41.79 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00006-of-00014.ggufGGUFQ2_K38.71 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00007-of-00014.ggufGGUFQ2_K41.79 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00008-of-00014.ggufGGUFQ2_K38.70 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00009-of-00014.ggufGGUFQ2_K41.79 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00010-of-00014.ggufGGUFQ2_K38.71 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00011-of-00014.ggufGGUFQ2_K41.79 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00012-of-00014.ggufGGUFQ2_K38.70 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00013-of-00014.ggufGGUFQ2_K41.79 GBDownload
DeepSeek-V4-Pro-0813-Q2_K-00014-of-00014.ggufGGUFQ2_K20.51 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00001-of-00018.ggufGGUFQ3_K_M39.44 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00002-of-00018.ggufGGUFQ3_K_M39.21 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00003-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00004-of-00018.ggufGGUFQ3_K_M39.22 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00005-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00006-of-00018.ggufGGUFQ3_K_M39.21 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00007-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00008-of-00018.ggufGGUFQ3_K_M39.22 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00009-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00010-of-00018.ggufGGUFQ3_K_M39.21 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00011-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00012-of-00018.ggufGGUFQ3_K_M39.22 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00013-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00014-of-00018.ggufGGUFQ3_K_M39.21 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00015-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00016-of-00018.ggufGGUFQ3_K_M39.22 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00017-of-00018.ggufGGUFQ3_K_M41.90 GBDownload
DeepSeek-V4-Pro-0813-Q3_K_M-00018-of-00018.ggufGGUFQ3_K_M22.87 GBDownload
imatrix.ggufGGUFGGUF1.57 GBDownload

Model Details

Model ID6block/DeepSeek-V4-Pro-0813-GGUF
Author6block
Pipelinetext-generation
Licensemit
Base modeldeepseek-ai/DeepSeek-V4-Pro-0813
Last modified2026-08-20T10:47:02.000Z

Model README

---

license: mit

language:

  • en
  • zh

pipeline_tag: text-generation

library_name: gguf

base_model: deepseek-ai/DeepSeek-V4-Pro-0813

base_model_relation: quantized

tags:

  • deepseek
  • deepseek_v4
  • gguf
  • imatrix
  • moe
  • 6block

---

DeepSeek-V4-Pro-0813 GGUF

Sub-4-bit GGUF quantizations of deepseek-ai/DeepSeek-V4-Pro-0813,

produced by the 6block team with importance-matrix (imatrix) calibration.

1.57T parameters, 48B active per token. 61 layers, 384 routed experts (top-6) + 1 shared expert.

Why only sub-4-bit tiers

The upstream weights ship in FP4 (expert_dtype: fp4 in config.json). Routed-expert tensors

are stored pre-packed, so convert_hf_to_gguf.py writes them straight to GGUF's MXFP4 type without

ever materialising BF16. The resulting F16 GGUF is 812.7 GiB at 4.33 bpw, and expert tensors are

96.4% of it.

That means the usual "higher tier = better" ladder does not apply. Measured with llama-quantize --dry-run

against this exact master:

| Tier | Size | vs master | Verdict |

|---|---|---|---|

| Q8_0 | 1556.7 GiB | +96% | inflates, no quality gained |

| Q6_K | 1201.9 GiB | +48% | inflates |

| Q5_K_M | 1038.9 GiB | +28% | inflates |

| Q4_K_M | 885.6 GiB | +12% | inflates |

| IQ4_XS | 787.4 GiB | −3% | not worth publishing |

Quantizing a 4.25-bpw tensor up to 8 bpw only doubles the file; it cannot recover information the

factory FP4 step already discarded. Everything published here is below the master's 4.33 bpw.

For 4-bit and 8-bit builds of this model, see unsloth/DeepSeek-V4-Pro-0813-GGUF

(UD-Q4_K_XL 850 GB, UD-Q8_K_XL 873 GB). This repository covers the range below that.

Available quantizations

| Tier | Size | bpw | PPL (12 chunks) | Layout | Notes |

|---|---|---|---|---|---|

| Q3_K_M | 711.3 GiB | 3.88 | 1.6217 ± 0.0528 | sharded | highest quality here |

| IQ3_XXS | 577.0 GiB | 3.15 | 1.6708 ± 0.0547 | sharded | best size/quality balance |

| Q2_K | 547.0 GiB | 2.99 | 1.7621 ± 0.0594 | sharded | |

| IQ1_M | 346.6 GiB | 1.89 | 3.6966 ± 0.1640 | single file | quality drops sharply |

| IQ1_S | 314.1 GiB | 1.72 | 4.1095 ± 0.1799 | single file | smallest |

Sharded tiers

The top three tiers exceed HuggingFace's 500 GB per-file limit, so they ship as

…-NNNNN-of-NNNNN.gguf shards of roughly 42 GiB each. Download every shard of a tier into one

directory and point -m at the first one — llama.cpp reads split.count from shard 00001 and pulls

in the rest automatically. Do not try to concatenate them; they are individually valid GGUF files,

not split-style byte chunks.

Master F16 GGUF baseline: PPL 4.0795 ± 0.0458 (measured during imatrix, 220 chunks — a different

chunk count than the table above, so it is not directly comparable; see caveats).

The quality cliff sits between Q2_K and IQ1_M: 200 GiB of savings costs +1.93 PPL, whereas the entire

Q3_K_M → Q2_K range costs only +0.14.

Tiers deliberately not published

IQ2_XS (445 GiB, PPL 4.4705) and IQ2_XXS (401 GiB, PPL 21.6798) were built and then rejected.

Both are beaten outright by smaller files — IQ1_S is 315 GiB at PPL 4.11 — so they occupy a size

bracket while delivering worse output. The IQ2 expert-quantization path appears to break down on this

sparse-routing MoE; the same failure mode showed up on DeepSeek-V4-Flash's IQ2_M. Sizes and tensor

counts looked completely normal, which is why every tier here was PPL-tested before release.

Quantization details

  • Tool: llama.cpp @ 4ed2b13 (needs LLM_ARCH_DEEPSEEK4; older builds reject deepseek4)
  • imatrix: 220 chunks over a 476 KB multilingual corpus (EN/ZH), final PPL 4.0795, published as imatrix.gguf
  • Requantization: --allow-requantize is mandatory. Expert tensors arrive already quantized as

MXFP4, and llama.cpp refuses to requantize by default

(requantizing from type mxfp4 is disabled). Note this makes every tier here a second

quantization pass on top of the factory FP4 step.

  • Non-expert tensors are protected explicitly, because a global low-bit setting would otherwise

crush the sparse-attention indexer and the per-layer control tensors:

```

hc_* → F32 (per-layer control)

attn_q/k/v/output → Q8_0

indexer, compressor → Q8_0 (sparse-attention index path)

ffn_gate_inp → F32 (router)

shexp → Q8_0 (shared expert)

token_embd, output → Q6_K

```

--tensor-type matches substrings and first match wins, so attn_ alone would also swallow

hc_attn_fn. The four attention projections are listed separately on purpose.

  • Metadata: general.quantized_by=6block, no absolute paths in any KV field.

Usage

llama.cpp

Sharded tier (Q3_K_M / IQ3_XXS / Q2_K) — fetch all shards, then load the first:

hf download 6block/DeepSeek-V4-Pro-0813-GGUF \
  --include "DeepSeek-V4-Pro-0813-IQ3_XXS-*.gguf" --local-dir .

llama-server -m DeepSeek-V4-Pro-0813-IQ3_XXS-00001-of-00015.gguf -c 8192 --jinja

Single-file tier (IQ1_M / IQ1_S):

hf download 6block/DeepSeek-V4-Pro-0813-GGUF \
  DeepSeek-V4-Pro-0813-IQ1_S.gguf --local-dir .

llama-server -m DeepSeek-V4-Pro-0813-IQ1_S.gguf -c 8192 --jinja

Do not pass -ngl or --n-cpu-moe manually. Setting either makes llama.cpp abandon automatic

VRAM fitting and split by layer count instead, which overflows individual cards on a model this

size (common_fit_params: n_gpu_layers already set by user to 99, abort, then cudaMalloc failed).

Let it fit the model itself.

Ollama

FROM takes one file, so for a sharded tier merge the shards first (needs free space for both the

shards and the merged result):

llama-gguf-split --merge \
  DeepSeek-V4-Pro-0813-IQ3_XXS-00001-of-00015.gguf \
  DeepSeek-V4-Pro-0813-IQ3_XXS.gguf
cat > Modelfile <<'EOF'
FROM ./DeepSeek-V4-Pro-0813-IQ3_XXS.gguf
PARAMETER temperature 0.6
PARAMETER top_p 0.95
EOF

ollama create deepseek-v4-pro -f Modelfile
ollama run deepseek-v4-pro

Caveats

Read these before comparing numbers with any other repository.

  1. PPL is wikitext-2, n_ctx=512, 12 chunks. Cross-tier comparisons in the table are valid;

comparisons against other models or other repos' published figures are not. Perplexity's running

average climbs monotonically as more corpus is covered, so a 12-chunk number and a 568-chunk

number are different measurements even for the same file.

  1. The 4.0795 master baseline was measured at 220 chunks, during the imatrix pass — not at 12.

It indicates the master's general range, not a like-for-like delta against the table.

  1. Every tier is a double quantization (factory FP4 → MXFP4 → target). Losses appear smaller than

they would from a BF16 master, because the first pass already removed most of the information.

That is a property of this master, not evidence of a better recipe.

  1. PPL is not generation quality. It measures language-modelling loss on one English corpus.

The 1-bit tiers pass the numeric gate but have not been evaluated for instruction following,

long-context behaviour, or agentic use. Test before deploying.

  1. No benchmark suite was run. No MMLU, GSM8K, or coding evaluations — only perplexity.

License

MIT, inherited from the upstream model. See the

original repository for terms.

---

Quantized by the 6block team.

Run 6block/DeepSeek-V4-Pro-0813-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models