Iloqt/Versi-StyleTune-31B-GGUF overview
license: apache 2.0 base model: Iloqt/Versi StyleTune 31B base model relation: quantized tags: gguf gemma 4 31B merge mergekit reasoning creative writing rolep…
Runs locally from ~11.53 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Versi-StyleTune-31B-Q2_K.gguf | GGUF | Q2_K | 11.53 GB | Download |
| Versi-StyleTune-31B-Q3_K_M.gguf | GGUF | Q3_K_M | 14.80 GB | Download |
| Versi-StyleTune-31B-Q4_K_M-attn8-HB.gguf | GGUF | Q4_K_M | 25.39 GB | Download |
| Versi-StyleTune-31B-Q4_K_M-hb16.gguf | GGUF | Q4_K_M | 21.58 GB | Download |
| Versi-StyleTune-31B-Q4_K_M.gguf | GGUF | Q4_K_M | 18.14 GB | Download |
| Versi-StyleTune-31B-Q5_K_M-attn8-HB.gguf | GGUF | Q5_K_M | 27.41 GB | Download |
| Versi-StyleTune-31B-Q5_K_M-hb16.gguf | GGUF | Q5_K_M | 24.52 GB | Download |
| Versi-StyleTune-31B-Q5_K_M.gguf | GGUF | Q5_K_M | 21.25 GB | Download |
| Versi-StyleTune-31B-Q6_K-attn8-HB.gguf | GGUF | Q6_K | 29.56 GB | Download |
| Versi-StyleTune-31B-Q6_K-hb16.gguf | GGUF | Q6_K | 27.64 GB | Download |
| Versi-StyleTune-31B-Q6_K.gguf | GGUF | Q6_K | 24.55 GB | Download |
| Versi-StyleTune-31B-Q8_0.gguf | GGUF | Q8_0 | 31.79 GB | Download |
| Versi-StyleTune-31B.i1-Q4_K_M-hb16.gguf | GGUF | Q4_K_M | 21.58 GB | Download |
| Versi-StyleTune-31B.i1-Q6_K-hb16.gguf | GGUF | Q6_K | 27.64 GB | Download |
| Versi-StyleTune-BF16.gguf | GGUF | BF16 | 59.82 GB | Download |
Model Details
Model README
---
license: apache-2.0
base_model:
- Iloqt/Versi-StyleTune-31B
base_model_relation: quantized
tags:
- gguf
- gemma-4
- 31B
- merge
- mergekit
- reasoning
- creative writing
- roleplay
- conversational
not-for-all-audiences: true
---
Versi-StyleTune-31B (GGUF)
GGUF quants of Iloqt/Versi-StyleTune-head: a text-only Gemma 4 31B with Nimbz/Versipellis-31B as the base and the lm_head (output projection) grafted from Gryphe/Gemma-4-31B-StyleTune.
All credits go to Nimbz and Gryphe for the original models, I only committed the merge.
Variants
Standard quants:
- Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0 — body-only quantization.
hb16 variants (head + embeddings kept at BF16):
- Q4_K_M-hb16, Q5_K_M-hb16, Q6_K-hb16
- Preserves the grafted StyleTune
lm_headandembed_tokensat full precision while quantizing the rest of the body to K-quant.
attn8-HB variants (Q8_0 attention + BF16 head + embeddings):
- Q4_K_M-attn8-HB, Q5_K_M-attn8-HB, Q6_K-attn8-HB
- Adds Q8_0 attention layers on top of the hb16 protection.
Notes
- Run with the Gemma 4 chat template; thinking off by default.
- Like all Gemma 4 models, benefits from repetition penalty or DRY to avoid token loops.
Run Iloqt/Versi-StyleTune-31B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models