GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

deucebucket/Qwen3.6-27B-Cerebellum-GGUF overview

<p align="center" <img src="cerebellum banner.png" alt="Cerebellum" width="640" </p Qwen 3.6 27B Cerebellum GGUF Ablation informed mixed precision quantization…

ggufquantizedcerebellumqwen3.6ablation-informedtext-generationbase_model:Qwen/Qwen3.6-27Bbase_model:quantized:Qwen/Qwen3.6-27Blicense:apache-2.0model-indexendpoints_compatibleregion:usimatrixconversational

Runs locally from ~11.98 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
650
Likes
11
Pipeline
text-generation

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-27B-Cerebellum-v4-Q2_K_Mixed.ggufGGUFQ2_K_MIXED11.98 GBDownload
Qwen3.6-27B-Cerebellum-v5-Q2_K_Mixed.ggufGGUFQ2_K_MIXED12.42 GBDownload

Model Details

Model IDdeucebucket/Qwen3.6-27B-Cerebellum-GGUF
Authordeucebucket
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3.6-27B
Last modified2026-06-16T07:24:43.000Z

Model README

---

license: apache-2.0

tags:

  • gguf
  • quantized
  • cerebellum
  • qwen3.6
  • ablation-informed

base_model: Qwen/Qwen3.6-27B

model_type: qwen3

quantized_by: deucebucket

pipeline_tag: text-generation

model-index:

  • name: Qwen3.6-27B-Cerebellum-GGUF

results:

- task:

name: Text Generation

type: text-generation

dataset:

name: AI2 Reasoning Challenge

type: ai2_arc

config: ARC-Challenge

split: test

metrics:

- name: normalized accuracy

type: acc_norm

value: 0.969

source:

name: Local benchmark run (RTX 3090, llama.cpp)

url: https://huggingface.co/deucebucket/Qwen3.6-27B-Cerebellum-GGUF/tree/main/benchmarks

- task:

name: Text Generation

type: text-generation

dataset:

name: HellaSwag

type: hellaswag

split: validation

metrics:

- name: accuracy

type: acc

value: 0.933

source:

name: Local benchmark run (RTX 3090, llama.cpp)

url: https://huggingface.co/deucebucket/Qwen3.6-27B-Cerebellum-GGUF/tree/main/benchmarks

- task:

name: Text Generation

type: text-generation

dataset:

name: MMLU-Redux

type: cais/mmlu

config: all

split: test

metrics:

- name: accuracy

type: acc

value: 0.783

source:

name: Local benchmark run (RTX 3090, llama.cpp)

url: https://huggingface.co/deucebucket/Qwen3.6-27B-Cerebellum-GGUF/tree/main/benchmarks

- task:

name: Text Generation

type: text-generation

dataset:

name: HumanEval (pass@1)

type: openai_humaneval

split: test

metrics:

- name: pass@1

type: pass@1

value: 0.909

source:

name: EvalPlus chat-mode (llama-server), samples + eval JSON in benchmark_results/

url: https://huggingface.co/deucebucket/Qwen3.6-27B-Cerebellum-GGUF/tree/main/benchmark_results

- task:

name: Text Generation

type: text-generation

dataset:

name: HumanEval+ (pass@1)

type: evalplus/humanevalplus

split: test

metrics:

- name: pass@1

type: pass@1

value: 0.890

source:

name: EvalPlus chat-mode (llama-server), samples + eval JSON in benchmark_results/

url: https://huggingface.co/deucebucket/Qwen3.6-27B-Cerebellum-GGUF/tree/main/benchmark_results

- task:

name: Text Generation

type: text-generation

dataset:

name: WikiText-2 Perplexity

type: wikitext

config: wikitext-2-raw-v1

split: test

metrics:

- name: perplexity

type: perplexity

value: 7.034

source:

name: Local benchmark run (RTX 3090, llama.cpp)

url: https://huggingface.co/deucebucket/Qwen3.6-27B-Cerebellum-GGUF/tree/main/benchmarks

---

<p align="center">

<img src="cerebellum_banner.png" alt="Cerebellum" width="640">

</p>

Qwen 3.6 27B Cerebellum GGUF

Ablation-informed mixed-precision quantization of Qwen 3.6 27B. Two builds in this repo:

  • v5, 13 GB: attention QKV projections (q/k/v/qkv) at Q5_K. 8-pack 105/150. File: Qwen3.6-27B-Cerebellum-v5-Q2_K_Mixed.gguf.
  • v4, 12 GB: 0.4 GB lighter, so more KV-cache room (longer max context on the same card). 8-pack 96/150. File: Qwen3.6-27B-Cerebellum-v4-Q2_K_Mixed.gguf.

Standard GGUFs; stock llama.cpp consumes them (same method, base, imatrix). v5's measured context ceiling and serving numbers are in the Serving section below.

What changed in v5

v5 is the v4 override map with the attention QKV projections (q/k/v/qkv across the 64 blocks) promoted from Q2/Q3 to Q5_K under the same ablation-informed budget. +1 GB (12 to 13). No other tensor changed, nothing re-quantized.

8-pack think-off: 105/150 vs v4 96/150. Per-pack deltas: dataextract +1, reasonmath +2, hermesagent +2, cli-40 +4; the other four equal. Wikitext PPL: 7.22 vs v4 7.03.

Behavioral quality (club-3090 8-pack, think-off /150)

Standardized behavioral suite, greedy/think-off, measured on each GGUF. This is a much broader read than perplexity or HumanEval alone.

| pack | v5 (13 GB) | v4 (12 GB) |

|------|:---:|:---:|

| toolcall-15 | 14 | 14 |

| instructfollow-15 | 13 | 13 |

| structoutput-15 | 14 | 14 |

| dataextract-15 | 8 | 7 |

| reasonmath-15 | 11 | 9 |

| bugfind-15 | 12 | 12 |

| hermesagent-20 | 15 | 13 |

| cli-40 | 18 | 14 |

| total | 105/150 | 96/150 |

Tool-calling, instruction-following, structured-output, and bug-finding hold at 80-93% on both builds. Data-extraction and multi-step reasoning are the lower packs; the v5 QKV promotion moves dataextract +1, reasonmath +2, hermesagent +2, cli-40 +4.

Benchmarks

Measured on each GGUF, RTX 3090, llama.cpp, temperature 0, thinking off. HumanEval via the upstream EvalPlus chat pipeline (greedy); v4 and v5 both re-measured on it here for a like-for-like row.

| Benchmark | v5 | v4 | Questions |

|-----------|:---:|:---:|-----------|

| Perplexity (WikiText-2, 2048 ctx) | 7.22 | 7.03 | |

| HumanEval pass@1 | 90.9% | 89.6% | 164 |

| HumanEval+ pass@1 | 86.0% | 86.0% | 164 |

| ARC-Challenge | 96.9% | 96.8% | 1,172 |

| HellaSwag | 93.3% | 92.2% | 10,042 |

| MMLU-Redux | 78.3% | 76.6% | 2,400 |

Samples JSONL, EvalPlus eval JSON, and 8-pack per-scenario JSONL are in benchmark_results/.

vs Q2_K imatrix (10 GB), the same-size-class baseline

| Benchmark | Cerebellum v5 (13 GB) | Q2_K imatrix (10 GB) |

|-----------|:---:|:---:|

| Perplexity | 7.22 | 7.500 |

| ARC-Challenge | 96.9% | 95.0% |

| HellaSwag | 93.3% | 90.8% |

| MMLU-Redux | 78.3% | 74.3% |

Short-answer benchmarks (ARC, HellaSwag) barely move at 2-bit either way. The separation is on code, multi-step reasoning, and agentic use.

How Cerebellum Works

Standard quantization applies the same precision level uniformly across every tensor. Cerebellum measures the actual sensitivity of each tensor and allocates bits where they matter.

Step 1: Ablation Sweep

Each tensor is individually crushed to Q2_K while keeping all other tensors at their baseline quant. The perplexity impact of each crush is measured. This produces a sensitivity map of the entire model.

Example measurements from this model (baseline PPL 8.256):

| Tensor | PPL when crushed | Delta | Verdict |

|--------|-----------------|-------|---------|

| blk.63.attn_q | 8.418 | +0.162 | Sacred, needs max precision |

| blk.63.ffn_down | 8.393 | +0.138 | Sacred |

| blk.1.ffn_gate | 8.294 | +0.039 | Sensitive |

| blk.50.ffn_down | 8.246 | -0.010 | Safe to crush |

| blk.34.ffn_down | 8.161 | -0.095 | Demotable, improves when crushed |

| blk.2.ffn_gate | 8.109 | -0.147 | Demotable, actively helps |

Step 2: Budget Allocation

Given a target file size, the allocator promotes sacred tensors to higher quant levels (Q3_K, Q4_K, Q5_K, Q6_K, Q8_0) in multiple passes, spending the size budget on tensors with the highest measured sensitivity. Demotable tensors are explicitly kept at Q2_K. v5 adds the attention QKV path to the promote list on top of the v4 map.

Step 3: Build

The final GGUF is built with llama-quantize --tensor-type @tensor_types.txt, which applies per-tensor quant overrides.

What the ablation found

Key findings on this model:

  • Layer 63 is the most sensitive: q_proj (+0.162 PPL) and ffn_down (+0.138 PPL) need maximum precision.
  • 7 tensors actively improve at Q2_K: crushing them reduces perplexity (negative delta).
  • Same-layer interactions are destructive: crushing two FFN tensors in the same layer at once regresses worse than expected (interaction ratio 0.13).
  • Cross-layer effects are about 86% additive: single-tensor deltas predict multi-tensor outcomes with roughly 14% attenuation.
  • Attention QKV carries behavior that perplexity underweights: protecting it (v5) lifts the 8-pack +9 while perplexity barely moves. Trust the behavioral bench, not perplexity alone.

Serving (measured, 1x RTX 3090 24 GB, mainline llama.cpp, q8_0 KV)

| metric | v5 |

|---|---|

| decode | 31.2 tok/s (CV 0.5%) |

| context | NIAH-clean to 120K (91% of a 131K window); needles recalled at 94K and 120K, all boundary checks pass |

| VRAM | ~12.4 GB weights; ~20.1 GB at 94K context (3.9 GB free on 24 GB) |

llama-server -m Qwen3.6-27B-Cerebellum-v5-Q2_K_Mixed.gguf \
  -ngl 99 -c 131072 -fa on --cache-type-k q8_0 --cache-type-v q8_0 --jinja

On a 24 GB card v5 runs large context comfortably; the +0.4 GB over v4 only narrows headroom near a 16 GB card. Decode is ~31 tok/s vs v4's ~34 (the Q5 attention projections read more per token). The continuous-soak finding from v4 (fill-scaled flash-attn scratch plateau, no leak, full retention) is architecture-level and carries to v5; it was not re-run here (the soak harness needs a club-3090 container, not a bare server).

Usage

Recommended: thinking OFF (reliable for chat and code):

llama-server -m Qwen3.6-27B-Cerebellum-v5-Q2_K_Mixed.gguf \
  -ngl 99 -c 16384 --jinja --reasoning-budget 0

Per request: temperature 0 for code and exact tasks.

Thinking / reasoning mode (works, but give it room):

llama-server -m Qwen3.6-27B-Cerebellum-v5-Q2_K_Mixed.gguf \
  -ngl 99 -c 32768 --jinja

Per request: temperature 0.6, top_p 0.95, top_k 20, and max_tokens >= 8192.

> Thinking mode uses 3,000 to 4,000+ tokens of reasoning before it answers. With a small

> context or max_tokens (e.g. 4096) it spends the whole budget thinking and returns an empty or

> half-finished reply (and temperature 0 can make the reasoning degenerate). This is a known

> llama.cpp serving behavior for Qwen3.6-27B

> (#22255,

> #22398) and reproduces on BF16/FP8, so it is

> not specific to this quant. For reliable single-shot output, use the thinking-OFF default

> above. Use a current llama.cpp build; avoid CUDA 13.2 (it produces gibberish).

Ollama

echo 'FROM ./Qwen3.6-27B-Cerebellum-v5-Q2_K_Mixed.gguf' > Modelfile
ollama create qwen36-cerebellum -f Modelfile
ollama run qwen36-cerebellum

Reproducing This Quant

The full ablation data, tensor type allocations, and tools are in the Cerebellum repo. The v5 override map is the v4 map plus the attention QKV promotions; both tensor-type files ship in benchmark_results/.

Run deucebucket/Qwen3.6-27B-Cerebellum-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models