GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Alittlehammmer/gemma-4-31B-it-DFlash-GGUF-llama.cpp overview

Gemma 4 31B it DFlash GGUF quantizations of z lab/gemma 4 31B it DFlash https://huggingface.co/z lab/gemma 4 31B it DFlash . Converted to BF16 using convert hf…

ggufquantizedllama.cpptext-generationagentic-codingdflashbase_model:z-lab/gemma-4-31B-it-DFlashbase_model:quantized:z-lab/gemma-4-31B-it-DFlashlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~869.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-4-31B-it-DFlash-BF16.ggufGGUFBF162.87 GBDownload
gemma-4-31B-it-DFlash-Q4_K_M.ggufGGUFQ4_K_M869.4 MBDownload
gemma-4-31B-it-DFlash-Q5_K.ggufGGUFQ5_K1.01 GBDownload
gemma-4-31B-it-DFlash-Q6_K.ggufGGUFQ6_K1.19 GBDownload
gemma-4-31B-it-DFlash-Q8_0.ggufGGUFQ8_01.53 GBDownload

Model Details

Model IDAlittlehammmer/gemma-4-31B-it-DFlash-GGUF-llama.cpp
AuthorAlittlehammmer
Pipelinetext-generation
Licenseapache-2.0
Base modelz-lab/gemma-4-31B-it-DFlash
Last modified2026-06-29T18:40:20.000Z

Model README

---

base_model:

- z-lab/gemma-4-31B-it-DFlash

base_model_relation: quantized

quantized_by: Alittlehammmer

license: apache-2.0

license_link: https://huggingface.co/z-lab/gemma-4-31B-it-DFlash/blob/main/LICENSE

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- quantized

- llama.cpp

- text-generation

- agentic-coding

- dflash

---

Gemma-4-31B-it-DFlash

GGUF quantizations of z-lab/gemma-4-31B-it-DFlash.

Converted to BF16 using convert_hf_to_gguf.py, then quantized using llama-quantize from llama.cpp.

Available quants

| Quant | Bits | Size | Notes |

| ------ | ----- | ------- | -------------------------------- |

| Q4_K_M | 4 | 912 MB | Average quality |

| Q5_K | 5 | 1.09 GB | High quality |

| Q6_K | 6 | 1.27 GB | Very high quality |

| Q8_0 | 8 | 1.65 GB | Highest quality, near lossless, Recommended |

| BF16 | 16 | 3.07 GB | Full precision, reference file |

Usage

Use in conjunction with existing Gemma 4 Quants, example config if using llama-server:

[Gemma-4-31B-it-DFlash]
sm = layer
model = /mnt/gguf/Gemma-4-31B-it/Gemma-4-31B-it-Q8_0.gguf
model-draft = /mnt/gguf/Gemma-4-31B-it-DFlash/Gemma-4-31B-it-DFlash-Q8_0.gguf
spec-type = draft-dflash
spec-draft-n-max = 6 

(Note: For some reason I cannot get sm = tensor to work, it crashes on launch, pretty sure this is an issue in llama.cpp)

Original model

See the original model card

for details on capabilities, benchmarks, and license.

Run Alittlehammmer/gemma-4-31B-it-DFlash-GGUF-llama.cpp with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models