GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

iromu/Gemma3-1B-tools-GGUF overview

Gemma3 1B Tools GGUF The Gemma 3 1B tool calling model in GGUF format, fine tuned with LoRA for tool calling and agent style interactions. Base model This mode…

ggufgemma3tool-callingfunction-callingagentsloraendataset:r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillationbase_model:google/gemma-3-1b-itbase_model:adapter:google/gemma-3-1b-itlicense:gemmaendpoints_compatibleregion:usconversational

Runs locally from ~776.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Gemma3-1B-tools-BF16.ggufGGUFBF161.88 GBDownload
Gemma3-1B-tools-Q4_K_M.ggufGGUFQ4_K_M776.5 MBDownload
Gemma3-1B-tools-Q5_K_M.ggufGGUFQ5_K_M819.7 MBDownload
Gemma3-1B-tools-Q8_0.ggufGGUFQ8_01.00 GBDownload

Model Details

Model IDiromu/Gemma3-1B-tools-GGUF
Authoriromu
Pipeline
Licensegemma
Base modelgoogle/gemma-3-1b-it
Last modified2026-08-29T12:33:20.000Z

Model README

---

language: en

license: gemma

base_model: google/gemma-3-1b-it

datasets:

  • r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation

tags:

  • gemma3
  • tool-calling
  • function-calling
  • agents
  • lora
  • gguf

library_name: gguf

---

Gemma3 1B Tools GGUF

The Gemma 3 1B tool-calling model in GGUF format, fine-tuned with

LoRA for tool calling and agent-style interactions.

Base model

This model was fine-tuned from:

google/gemma-3-1b-it

GGUF files

The model is provided in GGUF format at the following precisions:

| Precision | File |

|---|---|

| BF16 | Gemma3-1B-tools-BF16.gguf (original precision) |

| Q4_K_M | Gemma3-1B-tools-Q4_K_M.gguf |

| Q5_K_M | Gemma3-1B-tools-Q5_K_M.gguf |

| Q8_0 | Gemma3-1B-tools-Q8_0.gguf |

Training

Training was performed using NVIDIA NeMo AutoModel with LoRA/PEFT.

LoRA configuration

  • LoRA dimension: 32
  • LoRA alpha: 32
  • Dropout: 0.05
  • Target modules: .proj (all _proj linear layers)

Training configuration

  • Max sequence length: 4096
  • Learning rate: 5e-5 (cosine decay, 15 warmup steps, min 1e-6)
  • Weight decay: 0.01
  • Global batch size: 64 (micro batch 2 x 32 accumulation)
  • Training steps: 336 (4 epochs)
  • Mixed precision: bf16
  • Validation loss: 0.5790.4715 (final epoch)

Dataset

Training used the sft_tools split of the

r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation dataset.

Tool-calling format

This model was trained with a custom chat template (embedded in the

GGUF metadata). It renders the tool schemas into a developer turn and

emits tool calls as:

<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call>

Serving stacks must render prompts with this template for tool

calling to work.

Intended use

  • Structured tool/function calling
  • Agent-style multi-step interactions
  • Small-footprint on-device or edge deployment

It is not intended to be a general replacement for larger Gemma models.

GGUF versions

The model is available in GGUF format at:

  • BF16
  • Q4_K_M
  • Q5_K_M
  • Q8_0

Usage

Run the model with llama.cpp:

llama-cli -hf iromu/Gemma3-1B-tools-GGUF:Q4_K_M

The BF16 GGUF file can be quantized locally to other GGUF

precisions with llama-quantize if needed.

<!-- VALIDATION:BEGIN (auto-generated, do not edit) -->

Validation matrix

Tool-calling validation on the sft_tools validation split (greedy decoding, 384 max new tokens). Throughput is single-stream greedy decode, not serving throughput.

Pretrained base (google/gemma-3-1b-it): 2.0% exact-args match (1/50).

Fine-tuned (BF16): 66.0% exact-args match (33/50) (+64pp vs base).

  • GGUF-BF16: 20/50 (40.0%) exact, 66.0 tok/s — 61% of BF16.
  • GGUF-Q4_K_M: 10/50 (20.0%) exact, 89.5 tok/s — 30% of BF16.
  • GGUF-Q5_K_M: 24/50 (48.0%) exact, 60.7 tok/s — 73% of BF16.
  • GGUF-Q8_0: 22/50 (44.0%) exact, 50.2 tok/s — 67% of BF16.

| Model | Quant | n | Tool call emitted | Names match | Exact args match | Δ exact vs BASE | tok/s |

|---|---|---|---|---|---|---|---|

| Gemma3-1B-tools | BASE (google/gemma-3-1b-it) | 50 | 6/50 (12.0%) | 1/50 (2.0%) | 1/50 (2.0%) | — | 68.5 |

| Gemma3-1B-tools | BF16 | 50 | 50/50 (100.0%) | 41/50 (82.0%) | 33/50 (66.0%) | +64pp | 47.1 |

| Gemma3-1B-tools | GGUF-BF16 | 50 | 50/50 (100.0%) | 36/50 (72.0%) | 20/50 (40.0%) | +38pp | 66.0 |

| Gemma3-1B-tools | GGUF-Q4_K_M | 50 | 50/50 (100.0%) | 19/50 (38.0%) | 10/50 (20.0%) | +18pp | 89.5 |

| Gemma3-1B-tools | GGUF-Q5_K_M | 50 | 50/50 (100.0%) | 34/50 (68.0%) | 24/50 (48.0%) | +46pp | 60.7 |

| Gemma3-1B-tools | GGUF-Q8_0 | 50 | 50/50 (100.0%) | 37/50 (74.0%) | 22/50 (44.0%) | +42pp | 50.2 |

<!-- VALIDATION:END -->

Run iromu/Gemma3-1B-tools-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models