GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF overview

spoomplesmaxx whiskeyjack 12B i1 GGUF Weighted imatrix GGUF quants of aimeri/spoomplesmaxx whiskeyjack 12B https://huggingface.co/aimeri/spoomplesmaxx whiskeyj…

ggufimatrixquantizedgemma-4roleplaybase_model:aimeri/spoomplesmaxx-whiskeyjack-12Bbase_model:quantized:aimeri/spoomplesmaxx-whiskeyjack-12Bendpoints_compatibleregion:usconversational

Runs locally from ~4.47 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
spoomplesmaxx-whiskeyjack-12B.i1-IQ2_M.ggufGGUFIQ2_M4.47 GBDownload
spoomplesmaxx-whiskeyjack-12B.i1-IQ3_M.ggufGGUFIQ3_M5.74 GBDownload
spoomplesmaxx-whiskeyjack-12B.i1-IQ4_XS.ggufGGUFIQ4_XS6.68 GBDownload
spoomplesmaxx-whiskeyjack-12B.i1-Q4_K_M.ggufGGUFQ4_K_M7.40 GBDownload
spoomplesmaxx-whiskeyjack-12B.i1-Q5_K_M.ggufGGUFQ5_K_M8.60 GBDownload
spoomplesmaxx-whiskeyjack-12B.i1-Q6_K.ggufGGUFQ6_K9.88 GBDownload

Model Details

Model IDaimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF
Authoraimeri
Pipeline
License
Base modelaimeri/spoomplesmaxx-whiskeyjack-12B
Last modified2026-07-31T08:48:44.000Z

Model README

---

base_model: aimeri/spoomplesmaxx-whiskeyjack-12B

library_name: gguf

tags: [gguf, imatrix, quantized, gemma-4, roleplay]

---

spoomplesmaxx-whiskeyjack-12B-i1-GGUF

Weighted (imatrix) GGUF quants of aimeri/spoomplesmaxx-whiskeyjack-12B,

quantized on the box that trained it.

Measurements

KL divergence against this model's own bf16, on held-out text that was

excluded from the calibration corpus by construction. Not against the base

model, not against a benchmark — the number answers one question: how much did

quantization change this model.

| file | GB | KLD mean | KLD median | KLD p99 | RMS Δp % | same top-1 % |

|---|---|---|---|---|---|---|

| spoomplesmaxx-whiskeyjack-12B.i1-IQ2_M.gguf | 4.80 | 0.2262 | 0.1187 | 1.8559 | 14.50 | 81.03 |

| spoomplesmaxx-whiskeyjack-12B.i1-IQ3_M.gguf | 6.17 | 0.0545 | 0.0252 | 0.4786 | 7.14 | 90.44 |

| spoomplesmaxx-whiskeyjack-12B.i1-IQ4_XS.gguf | 7.17 | 0.0287 | 0.0114 | 0.2847 | 5.52 | 93.46 |

| spoomplesmaxx-whiskeyjack-12B.i1-Q4_K_M.gguf | 7.95 | 0.0221 | 0.0088 | 0.1982 | 5.01 | 94.21 |

| spoomplesmaxx-whiskeyjack-12B.i1-Q5_K_M.gguf | 9.24 | 0.0102 | 0.0034 | 0.1006 | 3.63 | 96.42 |

| spoomplesmaxx-whiskeyjack-12B.i1-Q6_K.gguf | 10.61 | 0.0038 | 0.0012 | 0.0366 | 2.16 | 97.93 |

Calibration

| | |

|---|---|

| context | 4096 tokens, document-aligned |

| corpus | 11.8M tokens, 2093 documents |

| sources | 100% in-domain (the model's own training corpora) |

| separator | <eos> (a special token — see below) |

| reference imatrix | merged from unsloth/gemma-4-12b-it-GGUF |

Every document is truncated to an exact multiple of the calibration context, so

llama-imatrix's non-overlapping windows land on document boundaries rather than

straddling two unrelated scenes. The separator is a special token because

whitespace separators are BPE-mergeable: a document ending in a newline and the

next beginning with one can fuse into a single token and shift every subsequent

window.

CALIB_CTX is 4096 rather than the 8192 used for larger models in this

family, and that is measured rather than inherited: at 8192 only 0.1% of the

aviary corpus clears a single window, which would have made calibration ~95% one

sub-corpus with no tool-calling coverage at all. Gemma 4 also runs 40 of its 48

layers as sliding-window attention at 1024, so context beyond a few thousand

tokens sharpens statistics for only the 8 global layers.

Serving

The stop token is <turn|> (id 106), not <eos>. generation_config.json

carries a list and GGUF stores a single u32; this build is verified to have

picked <turn|>. Anything that waits for <eos> will run past the turn.

Thinking is selected by <|think|> at the top of the system turn. With it,

every model turn opens a <|channel>thought ... <channel|> block; without it,

there is no channel at all.

For tool use, a whole episode lives inside ONE <|turn>model: serve with

stop=["<tool_call|>"], inject <|tool_response>response:NAME{...}<tool_response|>,

and continue the same turn. A harness waiting for <turn|> after a tool call

will hang.

Run aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models