GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-DS4-GGUF overview

DeepSeek V4 Flash — REAM96 111B · DS4 Q2 A 2 bit build of DeepSeek V4 Flash 0731 REAM96 111B https://huggingface.co/WaveCut/DeepSeek V4 Flash 0731 REAM96 111B …

ggufdeepseek-v4mixture-of-expertsreamds4text-generationenruzhesbase_model:WaveCut/DeepSeek-V4-Flash-0731-REAM96-111Bbase_model:quantized:WaveCut/DeepSeek-V4-Flash-0731-REAM96-111Blicense:mitendpoints_compatibleregion:usimatrixconversational

Runs locally from ~5.58 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
dspark.ggufGGUFGGUF5.58 GBDownload
ream96.ggufGGUFGGUF35.21 GBDownload

Model Details

Model IDWaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-DS4-GGUF
AuthorWaveCut
Pipelinetext-generation
Licensemit
Base modelWaveCut/DeepSeek-V4-Flash-0731-REAM96-111B
Last modified2026-08-14T05:23:24.000Z

Model README

---

base_model: WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B

license: mit

pipeline_tag: text-generation

language:

  • en
  • ru
  • zh
  • es

tags:

  • deepseek-v4
  • mixture-of-experts
  • ream
  • gguf
  • ds4

---

DeepSeek V4 Flash — REAM96 (111B) · DS4 Q2

A 2-bit build of DeepSeek-V4-Flash-0731-REAM96-111B

DeepSeek-V4-Flash with 96 of the original 256 experts per layer, sized for 48 GiB configurations.

Quantized with the standard DS4 recipe (2-bit experts, 8-bit attention) and a fresh

importance matrix.

> [!IMPORTANT]

> This is a DS4-specific GGUF. Run it with the DS4 fork

> the 96-expert topology needs its variable expert count support. Generic llama.cpp

> will not load this file.

> [!WARNING]

> Live smoke testing passed 5/10 scenarios on the first run. Independent reruns show the failures (Russian wordplay, multi-turn, English → Russian code-switching, Tool calling (DSML), Long-dialog focus (drift check), Tool call → code chain) are intermittent, not absolute — see the Stability column below for per-scenario pass rates. Multilingual chat, reasoning and long dialogs are consistently healthy. The full-precision native checkpoint may behave better — 2-bit quantization hits agentic behavior hardest.

Files

| File | Size | What it is |

| --- | ---: | --- |

| ream96.gguf | 36.4 GB | The model |

| dspark.gguf | 5.7 GB | Optional speculative decoding (DSpark) |

| imatrix.dat | 0.2 GB | Importance matrix used for this quant |

| SMOKE_REPORT.json | — | Raw smoke-test evidence |

Run

Plain:

./ds4 -m ream96.gguf -c 8192

With DSpark speculative decoding (faster generation, more memory):

./ds4 -m ream96.gguf --mtp dspark.gguf --dspark -c 8192

DSpark fits alongside on 64 GiB hosts, but in our Apple Silicon runs it gave no measurable speedup (12.9 vs 13.0 tok/s) — treat it as experimental.

How it was made

One pruning step, straight from the original — no cascading. Expert importance was

measured by running deepseek-ai/DeepSeek-V4-Flash-0731

over a ~5-million-token calibration mix (multi-turn dialogs, thinking and direct modes,

rendered with the model's own chat encoder). The strongest experts of every domain were

protected from pruning, the survivors were carried over byte-identical, and the

router was re-balanced to keep the original selection behavior.

| Calibration domain | Share |

| --- | ---: |

| Code | 35% |

| Agentic / tool use | 19% |

| Multilingual chat | 16% |

| Math | 8% |

| General chat | 6% |

| Roleplay | 6% |

| Russian | 5% |

| Long docs | 4% |

This line replaces the earlier cascaded REAM builds (now archived under -exp names),

which degraded badly in multi-turn use.

Smoke results

Every scenario is a live multi-turn conversation run end-to-end on the DS4 runtime (raw evidence ships in SMOKE_REPORT.json).

| Scenario | First run | Stability (reruns) |

| --- | :---: | :---: |

| Russian wordplay, multi-turn | ❌ | 2/10 |

| English → Russian code-switching | ❌ | 0/10 |

| Code Q&A over a 4k-token file | ✅ | — |

| Tool calling (DSML) | ❌ | 0/10 |

| Russian multi-turn reasoning | ✅ | — |

| Spanish creative writing | ✅ | — |

| Code refactoring | ✅ | — |

| Chinese summarization | ✅ | — |

| Long-dialog focus (drift check) | ❌ | 2/10 |

| Tool call → code chain | ❌ | 0/10 |

Stability = pass rate over independent reruns of the scenarios that failed the first run; passing scenarios were not re-run.

Limitations

  • Needs the DS4 fork; not a generic llama.cpp file.
  • 2-bit quantization is aggressive: expect the native checkpoint to be smarter than

this build, especially on agentic tool use.

  • Memory use grows with context length and DSpark; the sizes above are the files alone.

Run WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-DS4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models