GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

UlukaDev/saturn-31b-Q4_K_M-GGUF overview

Saturn 31B — Q4 K M GGUF A fine tune of Gemma 4 31B dense, instruct , trained on Fireworks AI and merged + quantized to Q4 K M GGUF for local inference with ll…

ggufgemma4q4_k_mcodehtmltext-generationbase_model:google/gemma-4-31B-itbase_model:quantized:google/gemma-4-31B-itlicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~17.40 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
557
Likes
4
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
saturn-31b-Q4_K_M.ggufGGUFQ4_K_M17.40 GBDownload

Model Details

Model IDUlukaDev/saturn-31b-Q4_K_M-GGUF
AuthorUlukaDev
Pipelinetext-generation
Licenseother
Base modelgoogle/gemma-4-31B-it
Last modified2026-07-26T23:47:00.000Z

Model README

---

base_model: google/gemma-4-31B-it

library_name: gguf

pipeline_tag: text-generation

license: other

license_name: saturn-community-license

license_link: LICENSE

tags:

- gemma4

- gguf

- q4_k_m

- code

- html

---

Saturn 31B — Q4_K_M GGUF

A fine-tune of Gemma 4 31B (dense, instruct), trained on Fireworks AI and

merged + quantized to Q4_K_M GGUF for local inference with llama.cpp.

Built for app and UI generation — HTML, frontend components, and small self-contained web applications.

License summary

You may use this model to build and ship applications, including commercial

ones. You may not resell or redistribute the model weights themselves as a

product, or present the model as your own. This model is a derivative of

Google Gemma 4 and remains subject to the Gemma Terms of Use, which apply in

addition to the terms below. See the LICENSE file for full text.

Specs

  • Base: google/gemma-4-31B-it (text-only export)
  • Quant: Q4_K_M (~4.87 bpw, ~17.4 GB)
  • Context: up to 262,144 tokens

Recommended hardware

Total memory (VRAM + RAM) should exceed the file size (~17.4 GB).

  • ~24 GB+ VRAM (RTX 3090/4090): full GPU offload, fast
  • 16 GB VRAM: partial offload, usable
  • CPU-only: works, but very slow

Run with llama.cpp

\\\`

llama-server -m saturn-31b-Q4_K_M.gguf -ngl 99 -c 32768 --jinja

\\\`

Gemma 4 has a thinking mode; disable with

\--chat-template-kwargs '{"enable_thinking":false}'\ for direct answers

Run UlukaDev/saturn-31b-Q4_K_M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models