Model Intelligence Sheet
nuofang/Qwen3.5-4B-MiniFantasy-GGUF overview
Auto Quantized GGUF Model This repository contains automated GGUF quantization files for Nubinu/Qwen3.5 4B MiniFantasy https://huggingface.co/Nubinu/Qwen3.5 4B…
Runs locally from ~3.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
5 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.5-4B-MiniFantasy-IQ4_XS.gguf | GGUF | IQ4_XS | 2.34 GB | Download |
| Qwen3.5-4B-MiniFantasy-Q4_K_M.gguf | GGUF | Q4_K_M | 2.52 GB | Download |
| Qwen3.5-4B-MiniFantasy-Q5_K_M.gguf | GGUF | Q5_K_M | 2.86 GB | Download |
| imatrix.gguf | GGUF | GGUF | 3.5 MB | Download |
| mmproj-Qwen3.5-4B-MiniFantasy-f16.gguf | GGUF | F16 | 644.3 MB | Download |
Model Details
Model README
---
base_model: Nubinu/Qwen3.5-4B-MiniFantasy
tags:
- llama.cpp
- quantized
- imatrix
---
Auto-Quantized GGUF Model
This repository contains automated GGUF quantization files for Nubinu/Qwen3.5-4B-MiniFantasy.
The calibration data for the imatrix is targeted at Chinese novels and role-playing (RP), while preserving logic and common sense.
imatrix 的校准数据以中文的小说、角色扮演为目标,同时保留逻辑和常识。
📊 Perplexity Evaluation
(Tested against the provided calibration dataset)
- Base (F16/BF16): PPL = 16.5300 +/- 0.13828
- IQ4_XS: PPL = 14.1688 +/- 0.11587
- Q4_K_M: PPL = 14.0570 +/- 0.11458
- Q5_K_M: PPL = 13.9765 +/- 0.11426
Run nuofang/Qwen3.5-4B-MiniFantasy-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models