GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Jerome0207/Huihui-DeepSeek-V4-Flash-0731-abliterated-Q4-MXFP4-GGUF overview

Huihui DeepSeek V4 Flash 0731 Abliterated — Q4 MXFP4 This is a single file mirror of the Q4 MXFP4 GGUF from huihui ai/Huihui DeepSeek V4 Flash 0731 abliterated…

ggufdeepseek-v4text-generationabliteratedenzhbase_model:deepseek-ai/DeepSeek-V4-Flash-0731base_model:quantized:deepseek-ai/DeepSeek-V4-Flash-0731license:mitendpoints_compatibleregion:usconversational

Runs locally from ~145.26 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-Q4-mxfp4-0731.ggufGGUFQ4145.26 GBDownload

Model Details

Model IDJerome0207/Huihui-DeepSeek-V4-Flash-0731-abliterated-Q4-MXFP4-GGUF
AuthorJerome0207
Pipelinetext-generation
Licensemit
Base modeldeepseek-ai/DeepSeek-V4-Flash-0731
Last modified2026-08-04T11:24:24.000Z

Model README

---

license: mit

base_model:

- deepseek-ai/DeepSeek-V4-Flash-0731

language:

- en

- zh

tags:

- gguf

- deepseek-v4

- text-generation

- abliterated

---

Huihui DeepSeek V4 Flash 0731 Abliterated — Q4-MXFP4

This is a single-file mirror of the Q4-MXFP4 GGUF from

huihui-ai/Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF.

It exists so RunPod's cached-model feature downloads only the selected quantization instead of the complete multi-quant repository.

The model and quantization were created by their respective upstream authors. This repository does not alter the GGUF payload.

File

| File | Bytes | SHA-256 |

| --- | ---: | --- |

| DeepSeek-V4-Flash-Q4-mxfp4-0731.gguf | 155,976,458,848 | 897c9d82ec412e1d983dcda579c96bd256454414b5b80d2570452477abe7beed |

The GGUF metadata declares the deepseek4 architecture, a maximum context length of 1,048,576 tokens, and a DSML-capable chat template. The associated RunPod deployment intentionally uses 262,144 tokens to retain generous KV-cache and runtime headroom on two 96 GB Blackwell GPUs.

Provenance

Review the upstream model cards, limitations, license, and acceptable-use requirements before deployment.

Run Jerome0207/Huihui-DeepSeek-V4-Flash-0731-abliterated-Q4-MXFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models