GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

schneewolflabs/Wichtelchen-Qwen3.5-9B-GGUF overview

Wichtelchen Qwen3.5 9B — GGUF image/png https://huggingface.co/nbeerbower/Wichtel Qwen3.6 27B/resolve/main/wichtel banner.png?download=true GGUF quants of schn…

llama.cppggufagentstool-usecodehemlockqwen3.5image-text-to-textenbase_model:schneewolflabs/Wichtelchen-Qwen3.5-9Bbase_model:quantized:schneewolflabs/Wichtelchen-Qwen3.5-9Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~879.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
image-text-to-text

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Wichtelchen-Qwen3.5-9B-Q4_K_M.ggufGGUFQ4_K_M5.38 GBDownload
Wichtelchen-Qwen3.5-9B-Q5_K_M.ggufGGUFQ5_K_M6.19 GBDownload
Wichtelchen-Qwen3.5-9B-Q6_K.ggufGGUFQ6_K7.04 GBDownload
Wichtelchen-Qwen3.5-9B-Q8_0.ggufGGUFQ8_09.11 GBDownload
Wichtelchen-Qwen3.5-9B-mmproj-f16.ggufGGUFF16879.0 MBDownload

Model Details

Model IDschneewolflabs/Wichtelchen-Qwen3.5-9B-GGUF
Authorschneewolflabs
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelschneewolflabs/Wichtelchen-Qwen3.5-9B
Last modified2026-08-29T00:30:37.000Z

Model README

---

base_model:

  • schneewolflabs/Wichtelchen-Qwen3.5-9B

library_name: llama.cpp

license: apache-2.0

pipeline_tag: image-text-to-text

language:

  • en

tags:

  • gguf
  • agents
  • tool-use
  • code
  • hemlock
  • qwen3.5

---

Wichtelchen-Qwen3.5-9B — GGUF

!image/png

GGUF quants of schneewolflabs/Wichtelchen-Qwen3.5-9B

the Wichtel operator recipe on Qwen3.5-9B: delegation 10/10, Hemlock 56.1% on hembench.

See the main card for the full ladder and limitations.

| file | quant | use |

|---|---|---|

| Wichtelchen-Qwen3.5-9B-Q8_0.gguf | Q8_0 | reference quality — all benchmark numbers were measured on this |

| Wichtelchen-Qwen3.5-9B-Q6_K.gguf | Q6_K | near-Q8 quality, smaller |

| Wichtelchen-Qwen3.5-9B-Q5_K_M.gguf | Q5_K_M | balanced |

| Wichtelchen-Qwen3.5-9B-Q4_K_M.gguf | Q4_K_M | smallest recommended |

Serving

llama-server -m Wichtelchen-Qwen3.5-9B-Q8_0.gguf -ngl 99 -c 8192 --jinja -fa on -np 1 \
    --spec-type draft-mtp --spec-draft-n-max 4

The checkpoint carries the MTP head, so --spec-type draft-mtp speculative decoding works.

For vision, add the mmproj: --mmproj Wichtelchen-Qwen3.5-9B-mmproj-f16.gguf.

Run schneewolflabs/Wichtelchen-Qwen3.5-9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models