GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sigmanih/gemma-4-12B-it-GGUF-Q4_K_M overview

<div align="center" ⚡ gemma 4 12B it GGUF Q4 K M High Performance Model Published via Σ SIGMA Studio https://github.com/Sigmanih/SigmaStudio SigmaStudio GitHub…

transformersgguftext-generationsigma-studiosigmanihconversationalcustom-modelsafetensorspytorchenitlicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~6.87 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
google--gemma-4-12B-it.Q4_K_M.ggufGGUFGGUF6.87 GBDownload

Model Details

Model IDsigmanih/gemma-4-12B-it-GGUF-Q4_K_M
Authorsigmanih
Pipelinetext-generation
Licenseapache-2.0
Base model
Last modified2026-08-27T21:48:36.000Z

Model README

---

language:

  • en
  • it

license: apache-2.0

tags:

  • text-generation
  • sigma-studio
  • sigmanih
  • conversational
  • custom-model
  • safetensors
  • transformers
  • pytorch

pipeline_tag: text-generation

---

<div align="center">

⚡ gemma-4-12B-it-GGUF-Q4_K_M

High-Performance Model Published via Σ-SIGMA Studio

![SigmaStudio GitHub](https://github.com/Sigmanih/SigmaStudio) ![HuggingFace Hub](https://huggingface.co/sigmanih/gemma-4-12B-it-GGUF-Q4_K_M) ![Engine](https://github.com/Sigmanih/SigmaStudio) ![License: Apache-2.0](https://opensource.org/licenses/Apache-2.0)

</div>

> ❤️ Support & Community: If you find this model helpful, please give this repository a Like on Hugging Face and a ⭐ Star on our SigmaStudio GitHub!

🌐 English Overview

gemma-4-12B-it-GGUF-Q4_K_M is a production-ready model optimized and published using the Model Hub module of Sigma Studio.

⚙️ Technical Specifications & Architecture

| Specification | Value |

| :--- | :--- |

| Model Repository | sigmanih/gemma-4-12B-it-GGUF-Q4_K_M |

| Weight Format | Safetensors (BF16 / FP16) |

| Base Architecture | CausalLM |

| Active Parameters | 12B |

| Context Window | 32,768 tokens |

| Total Disk Footprint | 6.87 GB |

| Recommended Usage | High-speed coding, Everyday assistants, Autonomous agentic loops & reasoning. |

🏆 Official Benchmark Performance

Evaluated directly on GPU via Sigma Studio Training Lab (Deterministic seed 42, Temp 0.0):

| Benchmark Suite | Score / Accuracy | Pass Rate | Test Date | Execution Engine |

| :--- | :---: | :---: | :---: | :---: |

| Tutti i Benchmark Ufficiali | 72.0% | 72/100 quesiti superati | 2026-08-27 | ⚡ SigmaEngine Direct GPU |

<details>

<summary>Per-suite breakdown</summary>

| Suite | Passed | Total | % |

| :--- | :---: | :---: | :---: |

| ARC-Challenge | 8 | 9 | 89% |

| BIG-Bench Hard | 7 | 7 | 100% |

| GPQA | 5 | 9 | 56% |

| GSM8K | 9 | 9 | 100% |

| HellaSwag | 5 | 9 | 56% |

| HumanEval | 5 | 7 | 71% |

| MATH | 8 | 9 | 89% |

| MBPP | 0 | 9 | 0% |

| MMLU | 11 | 14 | 79% |

| MMLU-Pro | 5 | 9 | 56% |

| TruthfulQA | 9 | 9 | 100% |

</details>

Protocol: code_execution, continuation_logprob, cot_generation, letter_logprob · temp 0 · seed 42

Reproducibility hash: SHA256-C38959F4D3D7998E

> ⚠️ Measured on a slice of the dataset, not the full suite: this score is not comparable with a full-suite run.

⚡ Real Measured Speed & Hardware Performance Matrix

  • Local Host Verified Speed: 87.8 tok/s (measured on NVIDIA GeForce RTX 5070 Ti15.9 GB VRAM).

| Hardware Tier | Typical Devices | Estimated Speed | Recommended Workload |

| :--- | :--- | :---: | :--- |

| 🚀 Tier 1 (Flagship Ultra) | NVIDIA RTX 4090, RTX 3090, A100, H100 | ~158 - 219 tok/s | Production APIs, Heavy Coding & Autonomous Agents |

| ⚡ Tier 2 (High Performance) | NVIDIA RTX 4070 Ti, RTX 4070, RTX 3080 12GB | ~96 - 131 tok/s | Interactive Chat, Dev Workstations & SLM Studio |

| 💻 Tier 3 (Mainstream / Mac) | RTX 4060 Ti 16GB, RTX 3060 12GB, Apple M2/M3/M4 | ~61 - 87 tok/s | Personal Assistant, Summarization & Edge Dev |

| 🧩 Tier 4 (CPU Offloading) | Multi-core CPU (Intel i7/i9, AMD Ryzen, 32GB RAM) | ~17 - 35 tok/s | Verification, Batch & Offline Processing |

🚀 Quick Start Guide

1. Running with Sigma Studio (Recommended)

Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:

# Clone and run Sigma Studio
git clone https://github.com/Sigmanih/SigmaStudio.git
cd SigmaStudio
.\sigma_studio.bat

2. Running with Transformers / PyTorch

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "sigmanih/gemma-4-12B-it-GGUF-Q4_K_M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto")

---

🇮🇹 Documentazione in Italiano

gemma-4-12B-it-GGUF-Q4_K_M è un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso Σ-SIGMA Studio.

📋 Specifiche e Configurazione

  • Architettura Base: CausalLM (12B parametri)
  • Formato Pesi: Safetensors (BF16 / FP16)
  • Spazio su Disco: 6.87 GB
  • Finestra di Contesto: 32,768 token
  • Profilo d'Uso Consigliato: Coding ad alta velocità, assistenti quotidiani, loop di agenti autonomi e ragionamento.

📊 Risultati Benchmark Ufficiali

  • Suite di Valutazione: Tutti i Benchmark Ufficiali
  • Punteggio Ufficiale: 72.0% (72/100 quesiti superati)

<details>

<summary>Dettaglio per suite</summary>

| Suite | Pass | Totale | % |

| :--- | :---: | :---: | :---: |

| ARC-Challenge | 8 | 9 | 89% |

| BIG-Bench Hard | 7 | 7 | 100% |

| GPQA | 5 | 9 | 56% |

| GSM8K | 9 | 9 | 100% |

| HellaSwag | 5 | 9 | 56% |

| HumanEval | 5 | 7 | 71% |

| MATH | 8 | 9 | 89% |

| MBPP | 0 | 9 | 0% |

| MMLU | 11 | 14 | 79% |

| MMLU-Pro | 5 | 9 | 56% |

| TruthfulQA | 9 | 9 | 100% |

</details>

Protocollo: code_execution, continuation_logprob, cot_generation, letter_logprob · temp 0 · seed 42

Impronta di riproducibilità: SHA256-C38959F4D3D7998E

> ⚠️ Misurato su una porzione del dataset, non sulla suite intera: il punteggio non è confrontabile con uno ottenuto sull'intero.

  • Data Test: 2026-08-27 su motore deterministico SigmaEngine

⏱️ Throughput Hardware e Fasce Consigliate

  • Velocità Verificata in Locale: 87.8 tok/s su NVIDIA GeForce RTX 5070 Ti.
  • Fascia Top GPU (RTX 4090/3090): ~158 - 219 tok/s (Ideale per produzione)
  • Fascia Media (RTX 4070/3080): ~96 - 131 tok/s (Ideale per sviluppo e studio)
  • Fascia Entry / Apple Silicon (RTX 3060/Mac): ~61 - 87 tok/s (Ideale per uso personale)
  • CPU Offload: ~17 - 35 tok/s

⭐ Supporta il Progetto Open Source

Se questo modello ti è utile o vuoi esplorare l'ecosistema completo:

  • 🌟 Metti una Stella al repository GitHub: Sigmanih/SigmaStudio
  • ❤️ Lascia un Like a questa scheda su Hugging Face

---

Creato e distribuito con il Model Hub di Σ-SIGMA Studio (27/08/2026 23:43)

Run sigmanih/gemma-4-12B-it-GGUF-Q4_K_M with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models