GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF overview

KAT Coder V2.5 Dev — MTP GGUFs KAT ships mtp num hidden layers: 0 — no draft head. These builds graft Qwen3.6 35B A3B's original MTP head onto KAT's trunk, qua…

ggufmoecodeagentic-codingmtpspeculative-decodingimatrixenzhbase_model:Kwaipilot/KAT-Coder-V2.5-Devbase_model:quantized:Kwaipilot/KAT-Coder-V2.5-Devlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~13.38 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
Author

Repository Files & Downloads

11 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
BF16/Kwaipilot_KAT-Coder-V2.5-Dev-BF16-MTP-00001-of-00002.ggufGGUFBF1642.46 GBDownload
BF16/Kwaipilot_KAT-Coder-V2.5-Dev-BF16-MTP-00002-of-00002.ggufGGUFBF1623.73 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Balanced.ggufGGUFGGUF24.43 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Compact-v2D-lite.ggufGGUFGGUF16.08 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Compact.ggufGGUFGGUF16.24 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Mini.ggufGGUFGGUF13.38 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Quality.ggufGGUFGGUF22.09 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-UD-IQ4_XS.ggufGGUFIQ4_XS16.96 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-UD-Q4_K_XL.ggufGGUFQ4_K_XL21.29 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-UD-Q5_K_S.ggufGGUFQ5_K_S23.79 GBDownload
Kwaipilot_KAT-Coder-V2.5-Dev-MTP-UD-Q6_K.ggufGGUFQ6_K27.95 GBDownload

Model Details

Model IDgbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF
Authorgbuzhf
Pipeline
Licenseapache-2.0
Base modelKwaipilot/KAT-Coder-V2.5-Dev
Last modified2026-08-02T20:04:17.000Z

Model README

---

license: apache-2.0

language:

  • en
  • zh

base_model:

  • Kwaipilot/KAT-Coder-V2.5-Dev

tags:

  • gguf
  • moe
  • code
  • agentic-coding
  • mtp
  • speculative-decoding
  • imatrix

---

KAT-Coder-V2.5-Dev — MTP GGUFs

KAT ships mtp_num_hidden_layers: 0 — no draft head. These builds graft

Qwen3.6-35B-A3B's original MTP head onto KAT's trunk, quantized with an

imatrix calibrated on KAT's own output.

Includes the bf16 master so you can build any tier yourself without a 69 GB

safetensors pull or a conversion.

---

Which head is in here, and why it matters

We fine-tuned this head twice on KAT's own rollouts. **Both fine-tunes made it

worse.** Measured live on 79 configs, same tier, same flags, only the head

differing:

| MTP head | COPY | NOVEL | AGENTIC |

|---|---|---|---|

| Qwen donor (shipped here) | 76% | 48% | 73% |

| our fine-tune, 450 steps | 50% | 24% | 46% |

| our fine-tune, 80 steps | 47% | 37% | 45% |

| (reference) Qwen head on Qwen's own trunk | 89% | 53% | 76% |

Draft acceptance, --spec-type draft-mtp, DraftMax 2, temp 1.0 / top_k 20 /

top_p 0.95 / presence_penalty 1.5.

**The donor head on KAT is within 3 points of Qwen's own co-trained head on its

own trunk.** There is essentially no trunk-swap penalty. Every file here carries

that head, verified byte-identical to the donor at build time:

donor-head sha256  faac91f15cbe54475faa2578bedc46a7c29a947b8a3e7ef3ecd376ae079826ab
blk.40.nextn.hnorm.weight  sha256  6dda2c53989ed9a8   <- fingerprint, verify yours

---

Files

| tier | recipe |

|---|---|

| UD-IQ4_XS | Unsloth Dynamic 2.0 |

| UD-Q4_K_XL | Unsloth Dynamic 2.0 |

| UD-Q5_K_S | Unsloth Dynamic 2.0 |

| UD-Q6_K | Unsloth Dynamic 2.0 |

| APEX-I-Mini | mudler APEX |

| APEX-I-Compact | mudler APEX |

| APEX-I-Quality | mudler APEX |

| APEX-I-Balanced | mudler APEX |

| APEX-I-Compact-v2D-lite | mudler APEX + v2D-lite |

| BF16/*-00001..2-of-00002.gguf | bf16 master, MTP embedded |

| original-mtp-head.safetensors | the head alone, for re-grafts |

Every map was read from that tier's own published GGUF header — none assumed,

none shared between tiers.

v2D-lite is applied to I-Compact only. It raises attn_k/attn_v on the

10 full-attention layers and output.weight, funded by token_embd. Unsloth's

maps already sit at Q8_0 on all of those, so applying it there would only lower

token_embd — measurably worse, so we didn't.

---

Serving

llama-server -m <model>.gguf -c 65536 -fa on --jinja \
  --spec-type draft-mtp,ngram-mod \
  --spec-draft-n-max 1 --spec-draft-n-min 0 --spec-draft-p-min 0.75 \
  --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 24 --spec-ngram-mod-n-match 48

Found by coordinate ascent over 79 live configs. Measured, RTX 3070 Ti Laptop

8 GB, 35 of 40 MoE layers on CPU:

| workload | t/s | draft acceptance |

|---|---|---|

| copy-heavy | 71.0 | 97% |

| agentic | 33.0 | 64% |

| novel prose | 33.9 | 81% |

Two knobs carry most of it:

  • --spec-draft-p-min 0.75 — the highest-leverage setting found. Drafting

only when confident turns a mediocre head into a useful one.

  • draft-mtp + ngram-mod together. Either alone is far worse: on this

hardware MTP alone is a net loss versus no speculation. With ngram, every

head reaches 96-97% on copy — ngram covers the repeats, and the head covers

the rest.

--spec-draft-n-max 1 beat 2 and 3: a longer MTP chain starves ngram-mod's

dispatch opportunities.

---

Building your own tier

No conversion, no graft, no 69 GB pull:

hf download gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF --include "BF16/*" --local-dir .
llama-gguf-split --merge BF16/Kwaipilot_KAT-Coder-V2.5-Dev-BF16-MTP-00001-of-00002.gguf master.gguf
llama-quantize --imatrix imatrix.gguf --tensor-type-file your_map.txt master.gguf out.gguf Q4_K_M

---

Known limitation

No imatrix contains statistics for blk.40llama-imatrix never executes

the MTP head during a forward pass. That block is quantized unguided in every

build, ours and everyone else's.

---

Credits

Kwaipilot — KAT-Coder-V2.5-Dev ·

Qwen — Qwen3.6-35B-A3B base and the MTP head ·

Unsloth — Dynamic 2.0 maps ·

mudler — APEX maps ·

bartowski — calibration corpus ·

llama.cpp

License: apache-2.0, inherited from the base model.

Run gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models