GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

nuofang/Qwen3.5-4B-MiniFantasy-GGUF overview

Auto Quantized GGUF Model This repository contains automated GGUF quantization files for Nubinu/Qwen3.5 4B MiniFantasy https://huggingface.co/Nubinu/Qwen3.5 4B…

ggufllama.cppquantizedimatrixbase_model:Nubinu/Qwen3.5-4B-MiniFantasybase_model:quantized:Nubinu/Qwen3.5-4B-MiniFantasyendpoints_compatibleregion:usconversational

Runs locally from ~3.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.5-4B-MiniFantasy-IQ4_XS.ggufGGUFIQ4_XS2.34 GBDownload
Qwen3.5-4B-MiniFantasy-Q4_K_M.ggufGGUFQ4_K_M2.52 GBDownload
Qwen3.5-4B-MiniFantasy-Q5_K_M.ggufGGUFQ5_K_M2.86 GBDownload
imatrix.ggufGGUFGGUF3.5 MBDownload
mmproj-Qwen3.5-4B-MiniFantasy-f16.ggufGGUFF16644.3 MBDownload

Model Details

Model IDnuofang/Qwen3.5-4B-MiniFantasy-GGUF
Authornuofang
Pipeline
License
Base modelNubinu/Qwen3.5-4B-MiniFantasy
Last modified2026-07-07T11:37:37.000Z

Model README

---

base_model: Nubinu/Qwen3.5-4B-MiniFantasy

tags:

  • llama.cpp
  • quantized
  • imatrix

---

Auto-Quantized GGUF Model

This repository contains automated GGUF quantization files for Nubinu/Qwen3.5-4B-MiniFantasy.

The calibration data for the imatrix is targeted at Chinese novels and role-playing (RP), while preserving logic and common sense.

imatrix 的校准数据以中文的小说、角色扮演为目标,同时保留逻辑和常识。

📊 Perplexity Evaluation

(Tested against the provided calibration dataset)

  • Base (F16/BF16): PPL = 16.5300 +/- 0.13828
  • IQ4_XS: PPL = 14.1688 +/- 0.11587
  • Q4_K_M: PPL = 14.0570 +/- 0.11458
  • Q5_K_M: PPL = 13.9765 +/- 0.11426

Run nuofang/Qwen3.5-4B-MiniFantasy-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models