GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sleepyeldrazi/deepseek-v4-flash-reap-k128-Q2-GGUF overview

DeepSeek V4 Flash — REAP K128 Uniform REAP pruned DeepSeek V4 Flash at K128 128 routed experts per MoE layer . Prunes 50% of routed experts via Cerebras REAP R…

ggufdeepseekdeepseek-v4deepseek-v4-flashmixture-of-expertsreapexpert-pruningds4experimentaldgx-sparktext-generationarxiv:2510.13999license:mitendpoints_compatibleregion:usconversational

Runs locally from ~46.98 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-REAP-K128-uniform.ggufGGUFGGUF46.98 GBDownload

Model Details

Model IDsleepyeldrazi/deepseek-v4-flash-reap-k128-Q2-GGUF
Authorsleepyeldrazi
Pipelinetext-generation
Licensemit
Base model
Last modified2026-06-19T15:56:40.000Z

Model README

---

license: mit

tags:

  • deepseek
  • deepseek-v4
  • deepseek-v4-flash
  • mixture-of-experts
  • reap
  • expert-pruning
  • gguf
  • ds4
  • experimental
  • dgx-spark

pipeline_tag: text-generation

---

DeepSeek V4 Flash — REAP K128 (Uniform)

REAP-pruned DeepSeek V4 Flash at K128 (128 routed experts per MoE layer).

Prunes 50% of routed experts via Cerebras REAP (Router-weighted Expert Activation Pruning),

preserving all attention, embeddings, shared experts, router, and MTP components.

Drop-in compatible with the standard ds4-engine runtime — uniform IQ2_XXS/Q2_K expert

quantization throughout all layers. No per-layer quant dispatch required.

At a Glance

| | |

|---|---|

| Base model | DeepSeek V4 Flash |

| Donor GGUF | antirez IQ2XXS-w2Q2K-AProjQ8 (80.8 GiB) |

| Pruning method | REAP (Cerebras Research) |

| Routed experts | 128 per layer (down from 256) |

| Kept slots | 5,888 / 11,008 |

| Hash-preserved | Layers 0-2 (256 experts each) |

| Pruned | Layers 3-42 (128 experts each) |

| Format | ds4-compact-v1 GGUF |

| File size | 46.98 GiB |

| Quantization | Uniform IQ2_XXS / Q2_K experts in all layers |

> This is the recommended variant for most users. Uniform expert quantization means

> it works out of the box with any ds4-engine build. The mixed-precision variant

> (52 GiB) has Q4_K in layers 37-42 but requires runtime per-layer quant dispatch.

Domain Split (Calibration)

8,000 prompts · 5.0M tokens · 1.3B routed expert observations

| Domain | Share |

|---|---|

| Coding & development | 35-40% |

| Agentic tool-calling | 16% |

| Research & knowledge | 15-20% |

| Math & science | 10-15% |

| Design & planning | 5-10% |

| Trivia & general QA | 3-5% |

Calibration used the REAP activation_energy_sum2 score metric with 4,096 token context per prompt.

Top-to-bottom expert score gap in layer 3: 2,200x (strong pruning signal).

How to Run

Requires eouya2/ds4-for-reaped (ds4 engine with compact GGUF support):

git clone https://github.com/eouya2/ds4-for-reaped
cd ds4-for-reaped
make cuda-spark -j$(nproc)  # DGX Spark / CUDA
# or: make                    # Metal / macOS

./ds4 --cuda -m DeepSeek-V4-Flash-REAP-K128-uniform.gguf --ctx 131072

API server mode:

./ds4-server --cuda -m DeepSeek-V4-Flash-REAP-K128-uniform.gguf \
  --host 0.0.0.0 --port 17777 --ctx 131072

How It Was Built

  1. Donor GGUF: Downloaded antirez IQ2XXS-w2Q2K-AProjQ8 variant (80.8 GiB) — uniform IQ2_XXS/Q2_K experts throughout
  2. Calibration: 8,000 prompts collected and run through ds4's imatrix collector on a DGX Spark (NVIDIA GB10) at 4,096 token context
  3. REAP scoring: Imatrix activation data converted to per-expert REAP scores using activation_energy_sum2 (same calibration used for the mixed-precision variant — REAP scores are quantization-independent)
  4. Pruning: 50% expert removal via ds4_prune_gguf.py from eouya2/reap-for-ds4. Layers 0-2 (hash-routed) preserved. Expert tensors copied byte-for-byte — no dequant/requant.
  5. Output: ds4-compact-v1 GGUF

No fine-tuning. Purely structural expert removal. Weights are unmodified — a subset of the original experts.

Comparison with Mixed-Precision Variant

| | Uniform (this) | Mixed Precision |

|---|---|---|

| File size | 46.98 GiB | 52.04 GiB |

| Expert quants | IQ2_XXS/Q2_K all layers | Q4_K in layers 37-42 |

| Compatibility | Drop-in | Needs quant-aware runtime |

| Quality | Baseline | Slightly higher in deep layers |

Acknowledgments

Run sleepyeldrazi/deepseek-v4-flash-reap-k128-Q2-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models