GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Alittlehammmer/gemma-4-26B-A4B-it-DFlash-GGUF-llama.cpp overview

Gemma 4 26B A4B it DFlash GGUF quantizations of z lab/gemma 4 26B A4B it DFlash https://huggingface.co/z lab/gemma 4 26B A4B it DFlash . Converted to BF16 usin…

ggufquantizedllama.cpptext-generationagentic-codingdflashbase_model:z-lab/gemma-4-26B-A4B-it-DFlashbase_model:quantized:z-lab/gemma-4-26B-A4B-it-DFlashlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~253.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-26B-A4B-it-DFlash-BF16.ggufGGUFBF16833.7 MBDownload
gemma-4-26B-A4B-it-DFlash-Q4_K_M.ggufGGUFQ4_K_M253.9 MBDownload
gemma-4-26B-A4B-it-DFlash-Q5_K.ggufGGUFQ5_K300.6 MBDownload
gemma-4-26B-A4B-it-DFlash-Q6_K.ggufGGUFQ6_K350.3 MBDownload
gemma-4-26B-A4B-it-DFlash-Q8_0.ggufGGUFQ8_0449.5 MBDownload

Model Details

Model IDAlittlehammmer/gemma-4-26B-A4B-it-DFlash-GGUF-llama.cpp
AuthorAlittlehammmer
Pipelinetext-generation
Licenseapache-2.0
Base modelz-lab/gemma-4-26B-A4B-it-DFlash
Last modified2026-06-29T18:40:04.000Z

Model README

---

base_model:

- z-lab/gemma-4-26B-A4B-it-DFlash

base_model_relation: quantized

quantized_by: Alittlehammmer

license: apache-2.0

license_link: https://huggingface.co/z-lab/gemma-4-26B-A4B-it-DFlash/blob/main/LICENSE

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- quantized

- llama.cpp

- text-generation

- agentic-coding

- dflash

---

Gemma-4-26B-A4B-it-DFlash

GGUF quantizations of z-lab/gemma-4-26B-A4B-it-DFlash.

Converted to BF16 using convert_hf_to_gguf.py, then quantized using llama-quantize from llama.cpp.

Available quants

| Quant | Bits | Size | Notes |

| ------ | ----- | ------- | -------------------------------- |

| Q4_K_M | 4 | 226 MB | Average quality |

| Q5_K | 5 | 315 MB | High quality |

| Q6_K | 6 | 367 MB | Very high quality |

| Q8_0 | 8 | 471 MB | Highest quality, near lossless, Recommended |

| BF16 | 16 | 874 MB | Full precision, reference file |

Usage

Use in conjunction with existing Gemma 4 Quants, example config if using llama-server:

[Gemma-4-26B-A4B-it-DFlash]
sm = layer
model = /mnt/gguf/Gemma-4-26B-A4B-it/Gemma-4-26B-A4B-it-Q8_0.gguf
model-draft = /mnt/gguf/Gemma-4-26B-A4B-it-DFlash/Gemma-4-26B-A4B-it-DFlash-Q8_0.gguf
spec-type = draft-dflash
spec-draft-n-max = 6 

(Note: For some reason I cannot get sm = tensor to work, it crashes on launch, pretty sure this is an issue in llama.cpp)

Original model

See the original model card

for details on capabilities, benchmarks, and license.

Run Alittlehammmer/gemma-4-26B-A4B-it-DFlash-GGUF-llama.cpp with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models