Abiray/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF overview
๐ป Gemma4 12B Coder GGUF โ Composer 2.5 ร Fable 5 โจ ๐ฃ Tiny footprint, big brain โ a local coding model for everyone No matter your GPU. No matter your RAM. Ifโฆ
Runs locally from ~5.15 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gemma-4-12B-coder-fable5-composer2.5-v1-IQ4_XS.gguf | GGUF | IQ4_XS | 6.18 GB | Download |
| gemma-4-12B-coder-fable5-composer2.5-v1-Q3_K_M.gguf | GGUF | Q3_K_M | 5.67 GB | Download |
| gemma-4-12B-coder-fable5-composer2.5-v1-Q3_K_S.gguf | GGUF | Q3_K_S | 5.15 GB | Download |
| gemma-4-12B-coder-fable5-composer2.5-v1-Q4_K_M.gguf | GGUF | Q4_K_M | 6.87 GB | Download |
| gemma-4-12B-coder-fable5-composer2.5-v1-Q5_K_M.gguf | GGUF | Q5_K_M | 7.96 GB | Download |
| gemma-4-12B-coder-fable5-composer2.5-v1-Q6_K.gguf | GGUF | Q6_K | 9.11 GB | Download |
| gemma-4-12B-coder-fable5-composer2.5-v1-Q8_0.gguf | GGUF | Q8_0 | 11.80 GB | Download |
Model Details
| Model ID | Abiray/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF |
|---|---|
| Author | Abiray |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF |
| Last modified | 2026-06-18T16:59:18.000Z |
Model README
---
license: apache-2.0
base_model: yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
library_name: gguf
pipeline_tag: text-generation
tags: [gemma4, coding, code, reasoning, thinking, gguf, llama.cpp, local-llm]
---
๐ป Gemma4-12B-Coder (GGUF) โ Composer 2.5 ร Fable 5 โจ
๐ฃ Tiny footprint, big brain โ a local coding model for everyone
> No matter your GPU. No matter your RAM. If you've got ~5.5 GB of VRAM or unified memory free,
> you can run your own private, offline coding assistant right now. ๐
> This is the v1 / code edition โ distilled from real chain-of-thought so it thinks through a problem
> before writing the solution. ๐ง ๐ป All local, all yours, no API, no cloud.
๐ฏ What it is
A focused fine-tune of Gemma 4 12B on verifiable Python coding data โ every training example's reasoning leads to
code that actually passed its tests. The result reasons in the open (edge cases, complexity, approach) and then
emits a clean, runnable solution. ๐
---
๐ฃ Context length fixed: now 256K (was 131K) โ thanks, community! ๐
A community member spotted that this model was reporting only a 131K context window. That turned out to be
the well-known upstream Gemma 4 metadata bug โ Google's initial config.json shipped with
max_position_embeddings: 131072 instead of the real 262144 (256K), and that value got baked into a lot of
downstream finetunes and quants (including this one) before it was fixed upstream.
The weights were always fine โ it was purely a metadata field. **All GGUF quants in this repository have been fully re-patched to the
full 256K context** (gemma4.context_length = 262144). Just re-download if you grabbed an earlier copy. ๐
---
๐ Training data (the interesting part ๐ณ)
This is a distillation of two complementary chain-of-thought sources, both over verifiable Python coding tasks
(algorithmic / function-level problems that come with deterministic tests):
- *๐ฅ Main set โ Composer 2.5 real CoT.* Genuine, model-authored reasoning traces. The teacher solved each problem,
its code was run against the task's tests, and only the passing solutions were kept. So the reasoning you're
learning from leads to code that actually works.
- ๐ฅ Aux set โ Fable 5 (released today! ๐). A clever twist: we took the problems where Composer 2.5 got it wrong
and handed them to Fable 5 to redo โ re-deriving a fresh, self-consistent chain-of-thought and a correct
solution, again gated on passing the tests. This recovers the hard cases the main teacher missed. These traces
are synthetic (rationalized CoT), and are tagged separately so the two sources stay distinguishable.
The recipe: real CoT for the bulk of solid coverage, plus synthetic "second-attempt" CoT to patch the failures โ
both verified by execution before anything entered training. โ
---
๐ฆ Pick your size (GGUF quants)
All files follow the standardized gemma-4-12B-coder-fable5-composer2.5-v1-(quant_name).gguf naming convention.
| Quant filename | Size | Vibe |
|------|------|------|
| ๐ข gemma-4-12B-coder-fable5-composer2.5-v1-Q3_K_S.gguf | 5.53 GB | Smallest available โ solid for tight memory profiles |
| ๐ก gemma-4-12B-coder-fable5-composer2.5-v1-Q3_K_M.gguf | 6.09 GB | Good balance for 8 GB total system/VRAM configurations |
| ๐ต gemma-4-12B-coder-fable5-composer2.5-v1-IQ4_XS.gguf | 6.64 GB | Sub-4-bit performance optimization via I-quantization |
| ๐ต gemma-4-12B-coder-fable5-composer2.5-v1-Q4_K_M.gguf | 7.38 GB | The ultimate sweet spot ๐ (highly recommended) |
| ๐ฃ gemma-4-12B-coder-fable5-composer2.5-v1-Q5_K_M.gguf | 8.55 GB | Excellent trade-off between preservation and weight size |
| ๐ฃ gemma-4-12B-coder-fable5-composer2.5-v1-Q6_K.gguf | 9.79 GB | Near-lossless precision representation |
| โช gemma-4-12B-coder-fable5-composer2.5-v1-Q8_0.gguf | 12.70 GB | Maximum precision, basically full teacher quality |
---
๐งฎ "Will it fit?" โ context length cheat-sheet
Rough estimates ๐ค (assumes q8_0 KV cache + ~1.5 GB overhead; use q4_0 KV cache for โ2ร more context!).
Max context is 256K. "โ" = won't fit, pick a smaller quant. โ๏ธ
| Your VRAM / unified mem | ๐ข Q3_K_S / M (~5.5-6G) | ๐ต IQ4_XS / Q4_K_M (~6.6-7.4G) | ๐ฃ Q5_K_M / Q6_K (~8.5-9.8G) | โช Q8_0 (12.7G) |
|---|---|---|---|---|
| 8 GB | ~12K ctx | tight (~2โ4K) | โ | โ |
| 12 GB | ~40K | ~30K | ~12K | โ |
| 16 GB | ~75K | ~64K | ~44K | ~22K |
| 24 GB | ~180K | ~128K | ~110K | ~88K |
| 32 GB | 256K (max) ๐ | 256K | ~230K | ~190K |
> ๐ก Apple Silicon / integrated GPUs with unified memory count too โ same numbers, just slower than a dGPU.
> ๐ก Low on room? Drop a quant or switch KV cache to q4_0 and your context roughly doubles.
---
๐ How to run it (super easy)
Option A โ llama.cpp (recommended) ๐ฆ
- Grab a quant above (e.g.
gemma-4-12B-coder-fable5-composer2.5-v1-Q4_K_M.gguf) andllama-serverfrom llama.cpp.
> โ ๏ธ Needs a recent llama.cpp (this uses the updated gemma4_unified architecture โ older builds won't load it).
- Run a server (Windows
.batshown โ tweak--port,--ctx-sizeto taste):
@echo off
cd /d C:\llama.cpp
llama-server.exe ^
-m C:\models\gemma-4-12B-coder-fable5-composer2.5-v1-Q4_K_M.gguf ^
--ctx-size 16384 ^
--n-gpu-layers 99 ^
--no-mmap ^
-fa on ^
--cache-type-k q8_0 --cache-type-v q8_0 ^
--temp 1.0 --top-p 0.95 --top-k 64 ^
--host 0.0.0.0 --port 18080
pause
- Open
http://localhost:18080and chat. ๐ (Tip: bump--ctx-sizeper the table; useq4_0KV for more.)
Option B โ one-click apps ๐ฑ๏ธ
Works in LM Studio, Jan, Ollama, etc. โ just import the local GGUF file manually, select your quant level, and launch. ๐พ
๐ง Thinking mode
This model thinks in Gemma's native thought channel before answering โ exactly how it was trained. Keep
enable_thinking=true (the default chat template handles it). Recommended sampling: temp 1.0, top_p 0.95, top_k 64.
For coding tasks you can also go greedy (temp 0) for more deterministic, structured solutions.
---
โ ๏ธ Good to know
- Reduced refusals: the training data is task-focused with no safety hedging, so this refuses less than the base
model. It is not safety-aligned โ add your own guardrails for production environments. Use responsibly. ๐
- Specialized for Python / algorithmic coding. Reasoning quality is strongest in that domain; general-knowledge
facts/numbers should still be double-checked.
- English-centric.
---
๐ Base & License
- License: Apache 2.0. Gemma 4 is released by Google under
Apache 2.0 (unlike the older Gemma 1/2/3 terms), so this fine-tune is
Apache 2.0 too โ free to use, modify, and redistribute. ๐
- Base model:
google/gemma-4-12B-it. - Personal/hobby project โ shared as-is, no warranty. Have fun, and happy hacking! ๐พโจ
Run Abiray/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF with guIDE
Download guIDE โ the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face ยท Compare models