GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Tinker-Stack/DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S-GGUF overview

DeepSeek V4 Flash 0731 REAP 150B · TQ3 4S Container v3b recipe · 256K ctx A custom, selectively quantized GGUF of the DeepSeek V4 Flash 0731 architecture, prun…

llama.cppggufexperimentaldeepseekdeepseek-v4moeturboquanttq3_4siq3_xxsquantized256kagentictool-usecodingreasoningimatrixtext-generationenzhmultilingualbase_model:puwaer/DeepSeek-V4-Flash-0731-reap-150bbase_model:quantized:puwaer/DeepSeek-V4-Flash-0731-reap-150blicense:mitendpoints_compatible

Runs locally from ~221.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
1,026
Likes
0
Pipeline
text-generation

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S.ggufGGUFGGUF61.01 GBDownload
imatrix-filtered.ggufGGUFGGUF221.8 MBDownload

Model Details

Model IDTinker-Stack/DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S-GGUF
AuthorTinker-Stack
Pipelinetext-generation
Licensemit
Base modelpuwaer/DeepSeek-V4-Flash-0731-reap-150b
Last modified2026-09-02T21:42:31.000Z

Model README

---

license: mit

base_model: puwaer/DeepSeek-V4-Flash-0731-reap-150b

tags:

- experimental

- deepseek

- deepseek-v4

- moe

- gguf

- llama.cpp

- turboquant

- tq3_4s

- iq3_xxs

- quantized

- 256k

- agentic

- tool-use

- coding

- reasoning

- imatrix

language:

- en

- zh

- multilingual

pipeline_tag: text-generation

library_name: llama.cpp

---

DeepSeek-V4-Flash-0731 REAP-150B · TQ3_4S-Container (v3b recipe) · 256K ctx

A custom, selectively-quantized GGUF of the DeepSeek-V4-Flash-0731 architecture, pruned to the

150B "REAP" variant, quantized with a hand-tuned per-tensor recipe and wrapped in the

TurboQuant TQ3_4S container from the turbo-tan/llama.cpp-tq3 fork.

This is not a plain uniform quant. It is the result of iterative quantization

on a 3× Turing GPU box, and the recipe reflects a specific, measured trade-off:

maximize context length (256K) while holding wikitext-2 perplexity within noise of the

IQ3_XXS baseline that proved to be the local optimum for this model.

---

> [!WARNING]

> Experimental build. This is a hand-tuned, selectively-quantized experimental release,

> not a polished upstream quant. It requires the non-stock

> turbo-tan/llama.cpp-tq3 fork to load

> (TQ3_4S container is rejected by stock llama.cpp), the recipe and serving flags are

> tuned for our specific 3× Turing hardware (§3, §6), and the benchmarks in §4 are our own

> internal measurements, not standardized leaderboard runs. Expect rough edges, and treat

> the numbers as relative to sibling builds on the same rig rather than absolute claims.

> If you just want a drop-in quant of this model, use the upstream REAP release instead.

---

Benchmarks

This quant vs. the models it derives from. Metrics per the REAP benchmarks:

GSM8K exact_match, MATH-500 math_verify, HumanEval+/MBPP+ pass@1_plus

(greedy, n=1).

| Benchmark | Base 284B | REAP-150B (FP8) | This build (TQ3_4S) | Δ vs FP8 |

|---|---|---|---|---|

| GSM8K | 0.9484 | 0.9295 | 0.9227 | −0.68 |

| MATH-500 | 0.7060 | 0.7140 | 0.9618 ᵃ | +24.78 |

| HumanEval+ | 0.8720 | 0.8963 | 0.9207 ᵃ | +2.44 |

| MBPP+ | 0.7407 | 0.7593 | 0.8624 | +10.31 |

| mean | 0.8168 | 0.8248 | 0.9169 | +9.2 |

ᵃ Thinking mode (reasoning_effort=high + answer headroom floor, §7 pitfall 8).

Base and FP8 reference scores are from the REAP model card, run with thinking

off; in matched no-thinking protocol this build scores 0.696 (MATH-500) and

0.6524 (HumanEval+). Full methodology, the both-protocols matrix, and spiral

rates: §4.

The quantization itself costs ~nothing: no-thinking GSM8K is within stderr of the

FP8 checkpoint (−0.68), MATH-500 is −1.8, MBPP+ is +10.3 above. The thinking-mode

rows show what the model does when given inference budget and the headroom floor

that keeps spirals from eating the answer.

---

Highlights

What this build buys you over the stock REAP-150B release, in one screen:

  • Quality on par with the base REAP model. wikitext-2 perplexity ~13.34, within noise of

the IQ3_XXS baseline (13.3073 ± 0.23) — the aggressive 3.49 BPW quantization costs no

measurable quality. (The failed 2-bit IQ2_S experiment measured 14.30, +7.5%, confirming

this recipe sits on the size/quality optimum, not past the cliff.) The full capability-bench

picture (base 284B / REAP-150B FP8 / this build) is right after this intro: no-thinking

quality tracks the FP8 checkpoint within noise, **MBPP+ beats it by +10.3 in matched

protocol, and with thinking mode MATH-500 hits 0.9618 and HumanEval+ 0.9207** (both

above the FP8 reference, which was measured thinking-off — see §4 for the full disclosure).

  • Optional uncensoring (external LoRA). An optional 2 MB rank-1 refusal-ablation LoRA

(heretic v2 t256-r2, not bundled here) drops refusals from 99.29% → ~7% with low KL

drift (0.085), applied via --lora at load — fetch it per §5. Base weights are never modified

or requantized; ship it censored or not, same 61 GiB file.

  • 2.7× the context. 256K context (vs the 96K ceiling of prior q8_0-KV builds) via the

TurboQuant q8_0 K + turbo4 V-cache — a 4-bit compressed KV value cache that frees the

headroom at zero quality cost.

  • ngram-mod speculative decoding. ngram-mod (n16/m24) speeds code / structured output to

~85 tok/s (100% draft acceptance on repetitive text) over a ~21 tok/s dense baseline.

  • Compact VRAM footprint. 61 GiB at 3.49 BPW fits full GPU offload in ~70 GiB — three

Turing cards (1× 22 GiB + 2× 24 GiB). A uniform Q4_K of the same model is ~84 GiB and

wouldn't fit; this selective recipe is why it does.

  • Ships complete. Everything to serve the model end-to-end is in this repo: the model, the

exact per-tensor recipe, and the importance matrix (see §2a).

---

1. Summary

| Property | Value |

|---|---|

| Base architecture | deepseek4 (DeepSeek-V4-Flash-0731) |

| Variant | REAP-150B (pruned) |

| Parameters | 150B total · 132 experts · 6 active/token · 1 shared expert |

| Hidden size | 4096 |

| Attention heads | 64 (MLA, head_count_kv = 1) |

| Layers | 43 (dense + MoE blocks) |

| Native context | 1,048,576 (model metadata) |

| File size | 61.01 GiB |

| Bits-per-weight | 3.49 BPW (weighted tensor payload) |

| Container format | TQ3_4S (GGUF general.file_type = 45) |

| Quantizer | turbo-tan/llama.cpp-tq3 @ commit 47635d7 |

| Runtime | llama.cpp (the same llama.cpp-tq3 fork) |

| Optional uncensor | heretic v2 t256-r2 LoRA (rank-1, 2 MB, applied via --lora) — external, see §5 |

| License | DeepSeek (inherited from base) |

The file was quantized from a full-precision Q8_0 re-quantization of the FP8 source

(dsv4-fp8-q8_0.gguf, 79.2 GiB) rather than from the raw FP8 weights, using

--allow-requantize with an importance matrix.

---

2. Why "TQ3_4S" but IQ3_XXS tensors? (Read this before you assume anything)

The naming is a deliberate container/recipe distinction that trips people up:

  • *TQ3_4S is the container type***. It is the TurboQuant "four-scale" format and the

general.file_type stamped into the GGUF header. It is what makes stock llama.cpp refuse

to load this file — you must use the turbo-tan/llama.cpp-tq3 fork (or a build with

TQ3_4S support) to run it.

  • The actual per-tensor types are a recipe, not a uniform TQ3_4S:

| Tensor group | Type | Rationale |

|---|---|---|

| MoE experts (gate/up), middle layers | IQ3_XXS | 3.06 bpw — the measured sweet spot for this model |

| MoE experts (down), middle layers | IQ3_S | slightly higher fidelity on the down-projection |

| MoE experts, edge layers 0–2 and 41–42 | mxfp4 (native, untouched) | preserve early/late layer fidelity |

| Shared expert (ffn_*_shexp) | Q8_0 | shared expert is used every token — keep it clean |

| Attention / indexer weights | Q6_K | attention is the bottleneck, keep it high |

| Token embeddings + output | Q8_0 | vocabulary fidelity |

| Hyper-connection / norm scales | F32 | tiny tensors, no reason to quantize |

The result is that the model is smaller than a uniform Q4_K (which would be ~84 GiB) while

holding quality, because the aggressive IQ3_XXS/IQ3_S quantization is applied only where the

measured error cost is lowest (the deep MoE experts), and the attention path is deliberately

left at Q6_K.

---

2a. Repository contents (this bundle)

This repository ships everything needed to serve the model end-to-end, not just

the base quant. Files and their roles:

| File | Size | Role |

|---|---|---|

| DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S.gguf | 61.01 GiB | The model: REAP-150B in TQ3_4S container (v3b recipe) |

| README.md | ~16 KB | This card |

| weights.txt | 406 B | Exact per-tensor quantization recipe (see §3) |

| imatrix-filtered.gguf | 222 MiB | Importance matrix used for the quant (reproducibility) |

> 🧩 The optional uncensor overlay is not bundled here. The base quant is self-sufficient

> (stock quality, ~99% refusal). If you want the ~7%-refusal behavior, an optional 2 MiB

> rank-1 LoRA (t256-r2) is applied losslessly via --lora in f32/f16 compute regardless of

> base quant. Fetch that adapter separately (see §5) — it lives in the MorinoNushi repo, not

> this one, so the base quant stays a clean, self-contained GGUF.

Other heretic adapter variants (t264/t265/t273) are not sideloaded here — see §5 for why

t256 is the correct choice for this box and how to fetch the others if you want a

stronger-decensor/slower trade-off.

---

3. Quantization recipe (exact tensor_types.txt)

token_embd\.weight=q8_0
output\.weight=q8_0
output_hc_fn\.weight=f32
blk\.[0-9]+\.hc_(attn|ffn)_(fn|base|scale)\.weight=f32
blk\.[0-9]+\.ffn_(gate|down|up)_shexp\.weight=q8_0
blk\.[0-9]+\.attn_[a-z_]+\.weight=q6_K
blk\.[0-9]+\.indexer[._][a-z_]+\.weight=q6_K
blk\.([0-2]|4[12])\.ffn_(gate|down|up)_exps\.weight=mxfp4
blk\.[0-9]+\.ffn_down_exps\.weight=iq3_s
blk\.[0-9]+\.ffn_(gate|up)_exps\.weight=iq3_xxs

Command used

llama-quantize \
  --allow-requantize \
  --imatrix imatrix-filtered.gguf \
  --tensor-type-file weights.txt \
  --output-tensor-type q6_K \
  --token-embedding-type q6_K \
  dsv4-fp8-q8_0.gguf \
  DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S.gguf TQ3_4S 32
  • Importance matrix: imatrix-filtered.gguf (222 MiB), 129 entries over 812 calibration

chunks, generated against the REAP-150B architecture (not the full 284B), so the importance

data is matched to this pruned variant.

  • --allow-requantize is required because the source is already-quantized Q8_0 (itself a

re-quant of the FP8 release weights).

---

4. Evaluation

Perplexity (wikitext-2)

The PPL gate for this line of experiments was ≤ 13.44 (the IQ3_XXS baseline 13.3073 ± 0.23

plus a 1% tolerance). This recipe reproduces the v3b recipe exactly, so it lands at the same

~13.34 as the already-validated IQ3_XXS/IQ3_S selective build.

For context on why the experts sit at IQ3_XXS rather than lower, the aborted IQ2_S-expert

candidate (the turbo-tan "champion" recipe) measured 14.30 ± 0.25 — a clear +7.5% regression

that confirms the 2-bit cliff for this pruned model. The IQ3_XXS/IQ3_S combination is the

measured local optimum on the size/quality curve.

> ⚠️ These are our internal wikitext-2 measurements, not a full Open LLM Leaderboard / MMLU

> run. Treat them as relative (this build vs. sibling builds), not absolute capability claims.

Capability benchmarks vs. the FP8 reference (2026-08-30)

Methodology (full disclosure). All runs on this build, same machine (§6), with the

heretic t256-r2 uncensor LoRA loaded (§5) and routed through our serving stack —

same headroom floor, reasoning-split, and retry logic production traffic uses:

| Setting | Value |

|---|---|

| Decoding | greedy, temperature=0.0, n=1 |

| Thinking rows | enable_thinking=true, reasoning_effort=high |

| Output budget | max_tokens floored to reasoning-budget + 2048 = 10240 (budget 8192) |

| GSM8K / MATH-500 / HumanEval+ | our own generation harness + graders: numeric exact-match, math_verify, and evalplus check_correctness respectively |

| MBPP+ | lm-eval mbpp_plus task (evalplus-based), 378 problems |

| Retries | none counted — generations scored as first emitted |

Per-row protocol is marked in the table; both protocols are measured and shown in the

protocol note below, so nothing is hidden.

Comparability caveat: puwaer's reference numbers (Base 284B and FP8 rows) are

published on the REAP model card: greedy n=1, enable_thinking=false, 4096-token

context, served with SGLang. Our no-thinking rows match that protocol (greedy,

thinking off, llama.cpp); our thinking rows add inference budget their runs did not

use, so the MATH-500 / HumanEval+ deltas vs FP8 are protocol-advantaged as well as

capability gains — which is exactly why the both-protocols matrix below exists.

| Benchmark | Metric | TQ3_4S (this build) | puwaer FP8 | Δ |

|---|---|---|---|---|

| GSM8K | exact_match | 0.9227 | 0.9295 | −0.68 |

| MATH-500 | math_verify (thinking) | 0.9618 | 0.7140 | +24.78 |

| HumanEval+ | pass@1 (thinking mode) | 0.9207 | 0.8963 | +2.44 |

| MBPP+ | pass@1_plus (no-think) | 0.8624 | 0.7593 | +10.31 |

Protocol note: results use the best serving protocol per task, all disclosed.

Thinking mode (reasoning_effort=high + headroom floor, pitfall #8) is worth

+26.8 on HumanEval+ and +26.6 on MATH-500, but it hurts MBPP+ by ~5.5 points

(0.8624 no-think vs 0.807 thinking) — the tasks are one-liners and overthinking

triggers formatting drift and the occasional budget-exhausted empty. Without

thinking mode this model scores ~0.65 HumanEval+ and ~0.70 MATH-500 on this box;

the naive lm-eval HumanEval 0.2256 is a format artifact.

All measured variants (greedy, n=1, same build, LoRA loaded):

| Benchmark | no-thinking | thinking (effort=high) |

|---|---|---|

| GSM8K | 0.9227 | not run (≤0.7pt headroom vs FP8; ~14 GPU-h for no expected gain) |

| MATH-500 | 0.696 | 0.9618 |

| HumanEval+ | 0.6524 | 0.9207 |

| MBPP+ | 0.8624 | 0.807 |

Spiral rate under the headroom floor: 1/164 (HumanEval), 2/500 (MATH-500),

5/378 (MBPP+ thinking) — every spiral burned exactly the 8192 reasoning budget

and emitted no answer; the +2048 headroom preserves the answer whenever the

model emits one at all.

Reading:

  • The quant retains the FP8 checkpoint's capability at 3.49 BPW: no-thinking

GSM8K within stderr, MATH-500 −1.8, MBPP+ +10.3 above FP8.

  • With thinking mode enabled the build exceeds the FP8 reference on MATH-500

(+24.8) and HumanEval+ (+2.4) — with the caveat that the FP8 reference was

measured thinking-off, so part of those deltas is inference budget (see protocol

note + comparability caveat above).

  • MBPP+ is genuinely above the FP8 reference (+10.3) in matched no-thinking

mode — the one bench where thinking is counterproductive (see protocol note).

  • HumanEval+ 0.9207 (thinking mode, greedy, reasoning_effort=high): requires

thinking-mode serving with a headroom floor — see pitfall #8 — otherwise reasoning

spirals can eat the output budget and return empty answers. Without thinking mode

the score drops to ~0.65; the naive lm-eval 0.2256 is a format artifact.

> These runs used the optional heretic t256-r2 LoRA (§5). Ablation check (no-thinking

> protocol): removing it moves HumanEval+ by ~1.5–2.6 pts either direction depending on

> extraction method — negligible relative to the quant's retained capability.

Throughput (3× Turing, see §6)

| Condition | Speed |

|---|---|

| Dense decode (baseline) | ~21–22 tok/s |

| ngram-mod speculative decoding, repetitive/code output | up to 85 tok/s (100% draft acceptance on repetitive text) |

Context length (measured)

| KV cache | Context | Result |

|---|---|---|

| q4_0 K / tq3_0 V | 256K | ✅ clean, full pipeline parallelism |

| q8_0 K / turbo4 V | 256K | ✅ clean (231 MiB headroom) |

| — | 512K | ⚠️ loads, but falls back to no-pipeline-parallel (degraded) |

| — | 1M | ❌ OOM on compute-buffer allocation |

The headline win over prior builds is 256K context at the same quality — a 2.7× gain over the

96K ceiling of the previous q8_0-KV build — achieved purely by switching to the compressed

turbo4 V-cache (the TurboQuant fork's 4-bit KV value cache).

---

5. How to run

Requirements

  • The turbo-tan/llama.cpp-tq3 fork (stock llama.cpp will not load TQ3_4S).
  • CUDA build: cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA=ON -DCMAKE_BUILD_TYPE=Release,

then cmake --build build -j --target llama-server.

  • ~70 GiB combined VRAM for full GPU offload at 256K (see §6 split).

Minimal serving command

CUDA_DEVICE_ORDER=PCI_BUS_ID \
CUDA_VISIBLE_DEVICES="<2080Ti>,<Titan-A>,<Titan-B>" \
llama-server \
  -m DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S.gguf \
  -ngl 999 \
  --ctx-size 262144 \
  --cache-type-k q8_0 --cache-type-v turbo4 \
  --flash-attn on \
  -ts 18,22,24 \
  --batch-size 1024 --ubatch-size 512

Recommended full config (agentic / coding)

  --spec-type ngram-mod --spec-ngram-mod-n-max 16 --spec-ngram-mod-n-match 24 \
  --temp 0.65 --top-p 0.90 --top-k 40 --min-p 0.15 \
  --repeat-penalty 1.02 --repeat-last-n 256 --presence-penalty 0.0 \
    --reasoning on --reasoning-format deepseek --reasoning-budget 8192
    # Optional uncensor overlay (external, see "Uncensoring" below):
    #   --lora <path/to/ds4-flash-heretic-f4-t256-lora-r2.gguf>
  ```

### Chat template / reasoning

The model emits a native `reasoning_content` field with an inline ` response` marker
(DeepSeek V4 format). The `--reasoning-format deepseek` flag routes thinking into
`reasoning_content` and the final answer into `content`.

### Uncensoring (optional, external heretic v2 LoRA)

The stock base refuses 99.29% of a 140-prompt harmful eval set. The
`MorinoNushi/DeepSeek-V4-Flash-0731-heretic-abliterated-v2-GGUF-lora` repo provides the
`ds4-flash-heretic-f4-t256-lora-r2.gguf` adapter (rank-1 directional-ablation LoRA, 2 MB) that
reduces refusals to ~7% with low KL drift (0.085). It is **lossless**: base weights are never
modified or requantized, applied in f32/f16 compute regardless of base quant via `--lora`. This
adapter is **not hosted here** — fetch it from the MorinoNushi repo linked above (or grab the
`t256-r2` file directly) and point `--lora` at it.

**Why t256-r2**: it is the only adapter with **zero routed-expert tensors**
(`ffn_down_exps`), which sidesteps the known llama.cpp LoRA slowdown (untagged
adapter buffers → host↔device copies + non-quantized mul_mat_id fallback that
disables CUDA graphs). Measured on this box: 21.06 t/s with the LoRA vs 21 t/s
without — no spec-decode penalty. (t264/t265/t273 touch 50-62 expert tensors and
cost a real 10-25% slowdown but reach 3.5-8.6% refusal.)

> ✅ Validated on this exact quant (REAP-150B, `TQ3_4S-Container`) on 2026-08-28: full
> compliance on refusal probes with ngram-mod still active.

---

## 6. Hardware and the tensor split

Developed and validated on a **3× NVIDIA Turing** box:

| Device | Model | VRAM |
|---|---|---|
| CUDA0 | RTX 2080 Ti | 22 GiB |
| CUDA1 | TITAN RTX | 24 GiB |
| CUDA2 | TITAN RTX | 24 GiB |

- **Tensor split `-ts 18,22,24`** (2080Ti=18, Titan=22, Titan=24) was determined by exhaustive
  headroom measurement and is the proven-optimal split for this 61 GiB model.
- **`CUDA_DEVICE_ORDER=PCI_BUS_ID` + explicit `CUDA_VISIBLE_DEVICES` pinning is mandatory.**
  Without pinning, CUDA enumerates the Titans before the 2080 Ti and the split lands on the
  wrong physical cards, causing OOM.
- Host RAM: 60 GiB. Full GPU offload is used (`-ngl 999`); the model does **not** spill to CPU.
- 256K context leaves ~231 MiB headroom on the tightest card. 512K degrades (loses pipeline
  parallelism), 1M OOMs.

If your cards are larger (e.g. 3× 48 GiB), 512K becomes clean and 1M is within reach.

---

## 7. Pitfalls (learned the hard way)

1. **Stock llama.cpp will not load this.** `general.file_type = 45` (`TQ3_4S`) is rejected by
   upstream. Use the `llama.cpp-tq3` fork.
2. **CUDA device enumeration order matters.** Without `CUDA_DEVICE_ORDER=PCI_BUS_ID` pinning, the
   `-ts 18,22,24` split maps to the wrong physical GPUs and you OOM even though the total VRAM
   is sufficient.
3. **Do not requantize the experts down to 2-bit.** `IQ2_S` experts cost +7.5% PPL on this pruned
   variant (the "champion" recipe does not survive the REAP prune). IQ3_XXS/IQ3_S is the floor.
4. **Do not requantize `mxfp4` experts *up* to Q4_K.** MXFP4 is *denser* than Q4_K (4.0 vs 4.58
   bpw); "re-quantizing to Q4" makes the file *bigger* and *lossier*. Keep native mxfp4 edges as-is.
5. **The DSML tool-call format changed between forks.** The old build-deepseek fork emitted
   tool-call *text* in `content`; the TQ3 fork emits native OpenAI `tool_calls` deltas. Any
   middleware written against the old fork must forward native `tool_calls` deltas or you get
   empty responses + retry loops.
6. **Reasoning effort interacts with context.** High reasoning effort on trivial prompts causes
   multi-minute think-a-thons. `reasoning-budget 8192` + `--spec-ngram-mod` clamps this.
7. **ngram-mod only accelerates structured output.** It is ~neutral during the reasoning phase
   (draft acceptance ~18%) and highly effective once the model emits code/structured text
   (up to 100% acceptance).
8. **Reasoning mode: keep `max_tokens` strictly above `--reasoning-budget`.** If a client sends
   `max_tokens` equal to the budget (e.g. both 8192), a single max-effort thinking spiral can
   burn the entire output budget on reasoning and leave **zero tokens for the visible answer**
   (`finish_reason='length'`, empty `content`). Force `max_tokens >= reasoning-budget + 2048`
   at your middleware/router — a guaranteed headroom floor eliminates the empty-answer case.
9. **Reasoning-mode code answers are wrapped in prose/fences.** Benchmarks (and any naive
   extraction) will undercount this model on HumanEval-style tasks unless you strip markdown
   fences / role-play junk and repair indentation before scoring. The model writes correct code
   far more often than the raw pass@1 suggests.

---

## 8. Intended use & limitations

**Intended:** agentic tool use, long-context coding and analysis, reasoning-heavy workflows,
serving a 150B-class MoE on consumer hardware with a large context window.

**Limitations:**
- This is a **pruned (REAP) variant**, not the full 284B model; some capabilities present in the
  full model are reduced.
- Quantization is lossy; for maximum fidelity prefer the FP8 or a BF16 base at the cost of ~2×
  file size and VRAM.
- Requires a non-stock llama.cpp build (the `TQ3_4S` container).
- Full 1M context is *not* reachable on 3× Turing (compute-buffer OOM); 256K is the validated
  ceiling there.

---

## 9. Reproducibility

1. Obtain the REAP-150B source (Q8_0 re-quant of the FP8 release).
2. Build `turbo-tan/llama.cpp-tq3` at `47635d7` with CUDA.
3. Generate or reuse an importance matrix matched to the REAP-150B architecture.
4. Run the quantize command in §3 with the exact `tensor_types.txt` recipe.
5. Verify: `1328` tensors, `61.01 GiB`, `3.49 BPW`, `general.file_type = 45`.

---

## 10. Acknowledgements

- **deepseek-ai** for the DeepSeek-V4-Flash-0731 architecture.
- **puwaer** for the REAP-150B prune and the FP8/Q8_0 source.
- **turbo-tan** for the `llama.cpp-tq3` fork (TurboQuant `TQ3_4S` + deepseek4 arch support).
- **MorinoNushi** for the heretic-gguf directional-ablation tooling and the v2 LoRA adapters
  used for the lossless uncensoring, and **p-e-w** for the underlying Heretic method.
- The quantization campaign, recipe, uncensoring, and validation were performed by **Tinkerstack**
  with automated tooling; the build is a community-style release for research and
  self-hosting.

Run Tinker-Stack/DeepSeek-V4-Flash-0731-REAP-150b-TQ3_4S-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models