GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF overview

Gemma 4 12B Coder — SFT v5 GGUF gemma 4 12B coder for local, agentic tool use — GGUF quantizations for llama.cpp / Ollama. Run it: llama server hf tpls/gemma 4…

ggufgemma4codetext-generationfunction-callingtool-useagenticllama.cppimatrixendataset:Agent-Ark/Toucan-1.5Mdataset:Nanbeige/ToolMinddataset:NousResearch/hermes-function-calling-v1dataset:Salesforce/xlam-function-calling-60kbase_model:tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5base_model:quantized:tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5license:gemmamodel-indexendpoints_compatibleregion:usconversational

Runs locally from ~7.98 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
130
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-12B-coder-fable5-composer2.5-v1-sft-Q4_K_M.ggufGGUFQ4_K_M7.98 GBDownload

Model Details

Model IDtpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF
Authortpls
Pipelinetext-generation
Licensegemma
Base modeltpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5
Last modified2026-06-26T13:59:04.000Z

Model README

---

base_model: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5

base_model_relation: quantized

library_name: gguf

pipeline_tag: text-generation

language:

- en

license: gemma

quantized_by: triplezrobotics

tags:

- gemma4

- code

- text-generation

- function-calling

- tool-use

- agentic

- gguf

- llama.cpp

- imatrix

datasets:

- Agent-Ark/Toucan-1.5M

- Nanbeige/ToolMind

- NousResearch/hermes-function-calling-v1

- Salesforce/xlam-function-calling-60k

model-index:

- name: Gemma-4 12B Coder — SFT v5 (GGUF)

results:

- task:

type: text-generation

name: Function calling (tool use)

dataset:

name: gemma4-coder-tool-eval

type: tpls/gemma4-coder-tool-eval

metrics:

- type: pass_rate

value: 1.0

name: Tool-call pass rate (shim, prod path)

- type: pass_rate

value: 0.125

name: Tool-call pass rate (raw llama.cpp --jinja)

---

Gemma-4 12B Coder — SFT v5 (GGUF)

gemma-4 12B coder for local, agentic tool use — GGUF quantizations for llama.cpp / Ollama.

Run it: llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja (full commands below).

> ⚠️ Tool-calling needs the recovery shim. The model emits gemma-4's native tool markup, which llama.cpp --jinja under-parses — wrap your endpoint with the tool-shim (see Tool-calling below) to get standard tool_calls.

> 💡 Pick this for the best tool-calling (our gate winner). For an uncensored model, use SFT v5 + abliterated GGUF.

At a glance

| | |

|--|--|

| Type | GGUF quantizations · llama.cpp / Ollama |

| Techniques | sft-qloraimatrix-quanttool-shim |

| Tool-calling | ✅ 100% gate pass (recovery-shim path) |

| Status | ✅ Active / supported |

| Use | llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja |

Use it

# llama.cpp (server) — tool-calling needs the recovery shim, see below
llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja --ctx-size 16384

# Ollama
ollama run hf.co/tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M

Files

Sizes and a one-click loader are in the file browser / Quantizations widget above;

the note says which quant to reach for.

| Quant | Notes |

|-------|-------|

| Q4_K_M | good default — fits 12 GB VRAM, best size/quality balance |

Tool-calling

Tool-calling works — but llama.cpp --jinja doesn't recognise gemma-4's native

tool-call markup, so the bare parser under-reports calls. The model is fine; the

parser is blind to the format. Recover standard tool_calls with a small serve-side

post-processor (no weight change, no latency beyond a regex scan).

Ready-to-use → tpls/gemma4-tool-shim — a drop-in

callback for OpenAI-compatible proxies, a standalone (dependency-free) example, and the pure

parser, all Apache-2.0, with the full recovery algorithm documented. Point your

OpenAI-compatible endpoint through it.

You send tools the usual OpenAI way (tools=[…]); the model emits native markup; the

shim turns it into a standard tool_calls object:

# model completion (raw):
<|tool_call>get_weather{"city": "Paris", "units": "celsius"}
// after the shim:
{"finish_reason": "tool_calls",
 "message": {"role": "assistant", "content": null,
   "tool_calls": [{"id": "call_0", "type": "function",
     "function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\", \"units\": \"celsius\"}"}}]}}

Tool-calling gate

Served the GGUF on llama.cpp (llama-server --jinja), prompted 7 tool-use cases + 1 no-tool abstain, scored whether a structured tool call was emitted. raw = llama.cpp native parse; shim = same outputs re-parsed for gemma-4 native markup. Tools folded into the prompt at eval time, matching training.

The rows are this model under two parse paths (raw and shim); the shim path is how it's served in production.

| Measured on | Pass rate |

|-------------|-----------|

| this model — raw (--jinja) | 0.125 |

| this model — shim (prod path) | 1.000 |

Intended use & limitations

Built for code generation and agentic tool use; serve locally via llama.cpp /

Ollama, or use as a base to fine-tune / merge / quantize. Outputs can be wrong or

fabricated — validate tool arguments before executing, and keep a human in the loop

for anything consequential.

Where this sits in the family

- SFT v5 (weights)

- SFT v5 + abliterated (weights)

- SFT v5 + abliterated (GGUF)

- SFT v5 (GGUF)you are here

---

Provenance & reproduction

How this model was built — technique chain, training mix, and the exact knobs/pins,

so the result is reproducible without any of our tooling.

Mechanics applied

| Step | Technique | What it does | Provenance |

|------|-----------|--------------|------------|

| 1 | sft-qlora | QLoRA supervised fine-tune to keep + improve native tool-calling | — |

| 2 | imatrix-quant | llama.cpp quantization with an importance matrix (imatrix) | — |

| 3 | tool-shim | serve-side recovery of structured tool_calls from the model's native markup | — |

1. sft-qlora

  • tools_mode: mixed

2. imatrix-quant

  • calibration: code + tool-call markup
  • embed/output: kept at f16 (protects tool-call logits)
  • eog_patch: tokens 105/106 → EOG (bounds the <|turn> runaway)

3. tool-shim

  • where: a thin pre/post wrapper on the OpenAI-compatible endpoint
  • format: re-parse <|tool_call>NAME{json-args} (and leaked <|tool>…) into tool_calls

> serve-side only — does not modify the weights; recommended for abliterated variants.

Training data & mix

Public sources; weights/row-caps are the exact balance.

| Dataset | Subset | Role | Weight | Max rows |

|---------|--------|------|-------:|---------:|

| Agent-Ark/Toucan-1.5M | Kimi-K2 | tool-dense multiturn | 2.0 | 1000 |

| Nanbeige/ToolMind | graph_syn_datasets/graphsyn.jsonl | reasoning-heavy (down-weighted vs v4 — the v4 culprit) | 1.0 | 500 |

| NousResearch/hermes-function-calling-v1 | func-calling.json | multiturn function calling | 1.5 | 600 |

| NousResearch/hermes-function-calling-v1 | func-calling-singleturn.json | terse single-call (up vs v4) | 2.0 | 600 |

| Salesforce/xlam-function-calling-60k | tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts) | terse, verifiable single-call (up vs v4) | 2.0 | 600 |

Pinned revisions (byte-exact reproduction):

  • Agent-Ark/Toucan-1.5M (Kimi-K2) @ 0df3cf37f2abefb380370cfb02eabea2a35ae782
  • Nanbeige/ToolMind (graph_syn_datasets/graphsyn.jsonl) @ 8020ed1c03c367e4eb720ac3828ab4b0b95d8baf
  • NousResearch/hermes-function-calling-v1 (func-calling.json) @ dae3e1d28cfbcf4b915c04ea1e072030529b4bda
  • NousResearch/hermes-function-calling-v1 (func-calling-singleturn.json) @ dae3e1d28cfbcf4b915c04ea1e072030529b4bda
  • Salesforce/xlam-function-calling-60k (tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts)) @ 26d14ebfe18b1f7b524bd39b404b50af5dc97866

Training hyperparameters

| Knob | Value |

|------|-------|

| method | QLoRA (4-bit NF4 base, bf16 compute) |

| lora_r / lora_alpha / dropout | 32 / 32 / 0.05 |

| target_modules | all-linear |

| objective | train on assistant turns only (responses-only masking) |

| optimizer | adamw_8bit |

| lr / scheduler / warmup | 2e-4 / cosine / 0.03 |

| epochs | 1 |

| seq_len | 4096 |

| effective_batch | 16 |

| chat_template | gemma (native turn boundaries 105/106) |

Training environment

Exact pins the run trained against (the base arch needs a recent transformers).

| Package | Version |

|---------|---------|

| torch | 2.11.0 |

| transformers | 5.13.0.dev0 @ c21da1b (git pin) |

| peft | 0.19.1 |

| trl | 1.6.0 |

| datasets | 5.0.0 |

| bitsandbytes | 0.49.2 |

| accelerate | 1.14.0 |

| liger-kernel | 0.8.0 |

| attention | eager (no flash-attn) |

Quantization environment

The GGUF bytes depend on the quantizer build, not just the weights — a different

llama.cpp release rounds tensors differently and can change the convert mapping. Pins

the toolchain these quants were produced with:

| Step | Tool / setting |

|------|----------------|

| quantizer | llama.cpp tools image ghcr.io/ggml-org/llama.cpp:full |

| convert | convert_hf_to_gguf.py → f16 GGUF |

| imatrix | llama-imatrix over the calibration set (CPU forward pass) |

| quantize | llama-quantize --imatrix, token-embeddings + output tensor kept at f16 |

> The image is the rolling :full tag, not a digest — for byte-exact reproduction pin

> the image digest you build with. The imatrix-quant step above lists the calibration

> set and the EOG patch this build applied.

---

Part of the Gemma-4 12B Coder — active collection.

Something not right, or a request? Open a discussion — happy to help.

Run tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models