GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

cHunter789/Qwen3.6-27B-i1-IQ4_KS_KT-GGUF overview

Qwen3.6 27B i1 IQ4 KS KT GGUF EXPERIMENTAL UPDATE I have uploaded what I believe to be the final versions of the quantized models, created with a strict maximu…

gguftext-generationqwenik_llama.cppnvidiaimatrix4-bitiq4_ksbase_model:Qwen/Qwen3.6-27Bbase_model:quantized:Qwen/Qwen3.6-27Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~13.09 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6,858
Likes
0
Pipeline
text-generation

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-27B.CPU.i1-IQ4_KT-attn_qkv-IQ4_KS-i1_MTP.ggufGGUFIQ4_KT13.46 GBDownload
Qwen3.6-27B.i1-IQ4_KS_KT-attn_qkv-IQ4_KS-MTP.ggufGGUFIQ4_KS_KT13.51 GBDownload
Qwen3.6-27B.i1-IQ4_KS_KT-attn_qkv-IQ4_KS-i1_MTP.ggufGGUFIQ4_KS_KT13.30 GBDownload
Qwen3.6-27B.i1-IQ4_KS_KT-attn_qkv-IQ4_KS.ggufGGUFIQ4_KS_KT13.09 GBDownload
Qwen3.6-27B.i1-IQ4_KT-attn_qkv-IQ4_KS-i1_MTP.ggufGGUFIQ4_KT13.33 GBDownload
Qwen3.6-27B.i1-IQ4_KT-attn_qkv-IQ4_KS.ggufGGUFIQ4_KT13.12 GBDownload

Model Details

Model IDcHunter789/Qwen3.6-27B-i1-IQ4_KS_KT-GGUF
AuthorcHunter789
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3.6-27B
Last modified2026-08-08T19:47:21.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3.6-27B

tags:

  • gguf
  • text-generation
  • qwen
  • ik_llama.cpp
  • nvidia
  • imatrix
  • 4-bit
  • iq4_ks

pipeline_tag: text-generation

quantized_by: cHunter789

---

Qwen3.6-27B-i1-IQ4_KS_KT-GGUF !!!EXPERIMENTAL!!!

UPDATE

I have uploaded what I believe to be the final versions of the quantized models, created with a strict maximum size target of 14.1 GB (without MTP).

Updated Quantizations:

  • Qwen3.6-27B.i1-IQ4_KT-attn_qkv-IQ4_KS.gguf (no MTP, for GPU Only)
  • Qwen3.6-27B.i1-IQ4_KT-attn_qkv-IQ4_KS-i1_MTP.gguf (with MTP, for GPU Only)
  • Qwen3.6-27B.CPU.i1-IQ4_KT-attn_qkv-IQ4_KS-i1_MTP.gguf (with MTP, for GPU+CPU)

The updated quants achieve lower perplexity (PPL).

./llama-perplexity     -m Qwen3.6-27B.i1-IQ4_KS-attn_qkv-IQ4_KS_test36n.gguf  -f /mnt/Samsung4TB/models/pg19.txt     -c 65536     --chunks 32     -ngl 99       -ctk q4_0     -ctv q4_0     -khad     -vhad     --no-mmap     --merge-qkv    -fa 1     -b 512     -ub 512     -t 8 
perplexity: tokenization took 957.272 ms
perplexity: calculating perplexity over 12 chunks, n_ctx=65536, batch_size=512, n_seq=1
perplexity: 70.16 seconds per pass - ETA 14.02 minutes
[1]6.6908,[2]6.9999,[3]7.1898,[4]7.3280,[5]7.4747,[6]7.3713,[7]7.4271,[8]7.4306,[9]7.4697,[10]7.5156,[11]7.5640,[12]7.3997,
Final estimate: PPL over 12 chunks for n_ctx=65536 = 7.3997 +/- 0.02772

The most interesting build here is Qwen3.6-27B.CPU.i1-IQ4_KT-attn_qkv-IQ4_KS-i1_MTP.gguf, which uses iq4_ks quantization for the first 16 FFN blocks.

This was done to test out the concept discussed in this Reddit thread.

GPU/CPU setup will give you 140k ctx (q5_0/q4_0) with decode speed starting at 35-30t/s and fall to 15-20t/s at the end (RTX 5070Ti + 8845HS dual channel DDR 5600)

llama-server" \
        -m "$MODEL_PATH" \
        -a Qwen3.6-27B \
        --ctx-size 140000 \
        --chat-template-file chat_template.jinja \
        --n-gpu-layers 99 \
        --cache-type-k q5_0 \
        --cache-type-v q4_0 \
        --spec-type ngram-mod:n_max=2 \
        --spec-type mtp:n_max=3 \
        --batch-size 512 \
        --ubatch-size 512 \
        -ot "blk\.([0-9]|1[0-5])\.ffn_.*=CPU" \
        --flash-attn on \
        --host 0.0.0.0 \
        --port 8081 \
        --reasoning on \
        --reasoning-format none \
        --reasoning-budget 32000 \
        -t 8 \
        -tb 8 \
        --parallel 1 \
        --metrics \
        --merge-qkv \
        -khad \
        -vhad \
        --chat-template-kwargs '{"preserve_thinking": true}' \
        --defrag-thold 0.4 \
        --jinja \
        --cont-batching \
        --temp 0.6 \
        --top-k 20 \
        --min-p 0.05 \
        --top-p 0.95 \
        --presence-penalty 0.0 \
        --repeat-last-n 512 \
        --repeat-penalty 1.05

IMPORTAN: Somehow the ubatch-size is important for the quality of this model. The recommended value is --ubatch-size 192. This is clearly visible in that test: https://sql-benchmark.nicklothian.com/#run-your-own-benchmark

This repository contains GGUF format weights for the Qwen3.6-27B model, quantized using the ik_llama.cpp project.

This model was specifically created to run on consumer GPUs with 16GB VRAM. By utilizing q4_0 KV cache quantization, it allows pushing the context length up to 105k tokens.

> Note: This model is designed exclusively for nVidia GPUs and is based on the advanced KS and KT quants developed by ikawrakow from the ik_llama.cpp repository.

---

Quantization Details & Imatrix File

  • Quantization Base: IQ_KS,IQ_KT Quants (ik_llama.cpp).
  • Imatrix File Used: Provided by mradermacher.
  • Other Tested Imatrix Files:

bartowski* – yielded significantly worse results.

ubergarm* – yielded comparable results.

* _If you find or generate a better Imatrix file, please let me know in the Community tab!_

<details>

<summary><b>secret recipe</b></summary>

# 1. Base Tensors & Embeddings
^output\.weight$ -> iq5_k
^output_norm\.weight$ -> f32
^token_embd\.weight$ -> iq4_ks

# 2. Attention Gate Exceptions (per-block)
^blk\.0\.attn_gate\.weight$ -> q8_0
^blk\.([1-2]|60|61)\.attn_gate\.weight$ -> iq6_k
^blk\.62\.attn_gate\.weight$ -> q8_0
.*attn_gate\.weight$ -> iq4_ks

# 3. Feed-Forward Networks (FFN) — Sandwich Pattern
# Outer blocks 0-3 and 60-63: standard iq4_ks
# Middle blocks 4-59: aggressive iq4_kss for down/gate
^blk\.([4-9]|[1-5][0-9])\.ffn_(down|gate)\.weight$ -> iq4_kss
.*ffn_(down|gate)\.weight$ -> iq4_ks
# All blocks: ffn_up uses iq4_kt
.*ffn_up\.weight$ -> iq4_kt

# 4. Attention Layers
.*attn_qkv\.weight$ -> iq4_ks
.*attn_(q|k)\.weight$ -> iq4_ks
.*attn_v\.weight$ -> iq5_ks
.*attn_output\.weight$ -> iq4_kt

# 5. Mamba / SSM Layers
.*ssm_out\.weight$ -> iq4_kt
.*ssm_(alpha|beta)\.weight$ -> iq4_ks
.*ssm_conv1d\.weight$ -> f32
.*ssm_a$ -> f32
.*ssm_dt\.bias$ -> f32

# 6. All Normalization Layers (Kept in f32 for stability)
.*norm\.weight$ -> f32

</details>

Benchmark Results & Real-World Impressions

This quantization is experimental. It uses i4_kt (trellis) quantization on a portion of the tensors. Normally, applying such quantization destroys outlier values, which impairs the model's reasoning capabilities. Here, I only applied it to areas where it should theoretically perform well. However, it remains an experiment. The model itself shows remarkably high performance in the https://sql-benchmark.nicklothian.com/ benchmark, scoring 24/25 points - just set ubatch-size to 192, yes really. In the SVG chess test (https://qwen3-6-27b-benchmark.vercel.app/), the results are mixed.

Perplexity (PPL) Testing

Perplexity evaluations were conducted focusing exclusively on the KV Cache quantization setup (q4_0), as this is the primary target use case:

wget [https://www.gutenberg.org/files/2600/2600-0.txt](https://www.gutenberg.org/files/2600/2600-0.txt) -O pg19.txt \

./llama-perplexity -m Qwen3.6-27B.i1-IQ4_KS_KT-attn_qkv-IQ4_KS.gguf -f /mnt/Samsung4TB/models/pg19.txt -c 65536 --chunks 32 -ngl 99 -khad -vhad -ctk q4_0 -ctv q4_0 -fa 1 -b 512 -ub 256

Test Log Output:

[1]6.6915,[2]7.0030,[3]7.1945,[4]7.3323,[5]7.4815,[6]7.3783,[7]7.4367,[8]7.4409,[9]7.4804,[10]7.5251,[11]7.5728,[12]7.4091,

Final estimate: PPL over 12 chunks for n_ctx=65536 = 7.4091 +/- 0.02777

---

Production Deployment Configuration (ik_llama.cpp)

Below is the verified and stable execution configuration used for llama-server. It utilizes the corrected chat template (chat.template) by ex-arman68, sourced from froggeric/Qwen-Fixed-Chat-Templates.

llama-server \
        -m "$MODEL_PATH" \
        -a Qwen3.6-27B \
        --ctx-size 105000 \
        --chat-template-file chat_template.jinja \
        --n-gpu-layers 99 \
        --cache-type-k q4_0 \
        --cache-type-v q4_0 \
        --batch-size 512 \
        --ubatch-size 192 \
        --flash-attn on \
        --no-mmap \
        --host 0.0.0.0 \
        --port 8081 \
        --reasoning on \
        --reasoning-format deepseek \
        -t 8 \
        --parallel 1 \
        -khad \
        -vhad \
        --chat-template-kwargs '{"preserve_thinking": true}' \
        --defrag-thold 0.3 \
        --jinja \
        --cont-batching \
        --temp 0.15 \
        --top-k 1 \
        --min-p 0.1 \
        --keep -1 \
        --repeat-last-n 512 \
        --repeat-penalty 1.05

Versions of the model with MTP have also been prepared. i1_MTP with iq4_ks quants and MPT with q8_0. Here is an example command to run the model with MTP:

llama-server \
        -m "$MODEL_PATH" \
        -a Qwen3.6-27B \
        --ctx-size 60000 \
        --chat-template-file chat_template.jinja \
        --n-gpu-layers 99 \
        -mtp \
        --spec-type mtp \
        --draft-n 3 \
        --cache-type-k q4_0 \
        --cache-type-v q4_0 \
        --batch-size 512 \
        --ubatch-size 192 \
        --flash-attn on \
        --no-mmap \
        --host 0.0.0.0 \
        --port 8081 \
        --reasoning on \
        --reasoning-format deepseek \
        -t 8 \
        --parallel 1 \
        -khad \
        -vhad \
        --chat-template-kwargs '{"preserve_thinking": true}' \
        --defrag-thold 0.3 \
        --jinja \
        --cont-batching \
        --temp 0.15 \
        --top-k 1 \
        --min-p 0.1 \
        --keep -1 \
        --repeat-last-n 512 \
        --repeat-penalty 1.05

Run cHunter789/Qwen3.6-27B-i1-IQ4_KS_KT-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models