GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ManniX-ITA/Qwen3.5-4B-M8-GGUF overview

Qwen3.5 4B M8 — GGUF GGUF quantizations of ManniX ITA/Qwen3.5 4B M8 https://huggingface.co/ManniX ITA/Qwen3.5 4B M8 . See the BF16 model card for the recipe an…

ggufqwen3.5mergellama.cppbase_model:ManniX-ITA/Qwen3.5-4B-M8base_model:quantized:ManniX-ITA/Qwen3.5-4B-M8license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~3.23 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
39
Likes
0
Pipeline

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
m8-Q6_K.ggufGGUFQ6_K3.23 GBDownload

Model Details

Model IDManniX-ITA/Qwen3.5-4B-M8-GGUF
AuthorManniX-ITA
Pipeline
Licenseapache-2.0
Base modelManniX-ITA/Qwen3.5-4B-M8
Last modified2026-07-12T19:01:40.000Z

Model README

---

base_model: ManniX-ITA/Qwen3.5-4B-M8

tags:

- qwen3.5

- merge

- gguf

- llama.cpp

license: apache-2.0

---

Qwen3.5-4B-M8 — GGUF

GGUF quantizations of ManniX-ITA/Qwen3.5-4B-M8.

See the BF16 model card for the recipe and full ablation matrix.

Quants

| File | Quant | Size |

|---|---|---:|

| m8-Q6_K.gguf | Q6_K | 3.46 GB |

Usage

llama-server -m m8-Q6_K.gguf \
    --port 8099 -c 32768 -ngl 99 --no-warmup \
    --reasoning-format deepseek --reasoning-budget 8192 \
    --parallel 2 --cache-type-k q8_0 --cache-type-v q8_0

Eval (Q6_K, lm-evaluation-harness)

| Task | Score |

|---|---:|

| HumanEval | 54.27 |

| MBPP | 51.40 |

Run ManniX-ITA/Qwen3.5-4B-M8-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models