GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

satgeze/Gemma4-26B-A4B-Uncensored-HauhauCS-1M-GGUF overview

<img src="banner.jpeg" width="720"/ Gemma4 26B A4B Uncensored: 1M Context + MTP + Vision HauhauCS/Gemma4 26B A4B QAT Uncensored HauhauCS Balanced MTP https://h…

gguflong-contextyarngemma4uncensoredmtpspeculative-decodingvisionllama.cppollamatext-generationbase_model:HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTPbase_model:quantized:HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTPlicense:gemmaendpoints_compatibleregion:usconversational

Runs locally from ~240.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2,130
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma4-26b-a4b-uncensored-1M-Q4_K_M.ggufGGUFQ4_K_M15.64 GBDownload
mmproj-gemma26b-hauhau.ggufGGUFGGUF1.11 GBDownload
mtp-gemma-4-26B-A4B-it.ggufGGUFGGUF240.3 MBDownload

Model Details

Model IDsatgeze/Gemma4-26B-A4B-Uncensored-HauhauCS-1M-GGUF
Authorsatgeze
Pipelinetext-generation
Licensegemma
Base modelHauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
Last modified2026-07-10T22:59:09.000Z

Model README

---

license: gemma

pipeline_tag: text-generation

base_model: HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP

tags:

  • gguf
  • long-context
  • yarn
  • gemma4
  • uncensored
  • mtp
  • speculative-decoding
  • vision
  • llama.cpp
  • ollama

---

<img src="banner.jpeg" width="720"/>

Gemma4-26B-A4B Uncensored: 1M Context + MTP + Vision

HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP (26B MoE, 4B active, Google QAT checkpoint) with a 1,048,576-token context baked in (4x the native 262,144), shipping with its MTP speculative-decoding draft head and vision tower. All numbers below were measured on these exact files.

<table>

<tr>

<th style="background:#1a73e8;color:#fff;padding:8px 14px;">Capability</th>

<th style="background:#1a73e8;color:#fff;padding:8px 14px;">Status</th>

</tr>

<tr><td><b>1M context</b></td><td>~91% mean recall across two seed sets (honesty note below)</td></tr>

<tr><td><b>MTP speculative decoding</b></td><td>249.1 to 369.2 tok/s (<b>+48%</b>), acceptance 0.679 (measured on this trunk, RTX 5090)</td></tr>

<tr><td><b>Vision</b></td><td>Verified July 6, 2026: reads image text and identifies objects</td></tr>

<tr><td><b>Uncensored</b></td><td>HauhauCS Balanced abliteration; trunk weights bit-identical to the source release</td></tr>

</table>

Needle-in-a-haystack

<img src="niah_heatmap.png" width="640"/>

Full transparency: across two complete seed sets this trunk averages ~91 percent needle recall, dropping roughly one needle per rung at random depths, including inside the native 262K range. The official censored trunk scored 10/10 at 393K under the same harness, so this is a small flat abliteration tax, not a context-length failure. If a rare retrieval miss is unacceptable for your workload, use the 12B, which is certified clean. Every run, including the imperfect ones, is in results.jsonl. Rungs above 524K exceed 32 GB VRAM and have not been run yet; they are queued for larger hardware and will publish here as measured, pass or fail.

MTP speculative decoding

<img src="mtp_speedup.png" width="480"/>

The draft head predicts ahead and the trunk verifies every token, so output is identical to standard decoding, only faster. Measured speedup on this uncensored trunk beats the ~35 percent claimed upstream.

Files

| File | Size | Role |

|---|---|---|

| gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf | 16.8 GB | Trunk, 1M baked, QAT 4-bit |

| mtp-gemma-4-26B-A4B-it.gguf | 252 MB | MTP draft head, pair with -md |

| mmproj-gemma26b-hauhau.gguf | 1.2 GB | Vision tower, pair with --mmproj |

| niah_heatmap.png, mtp_speedup.png, results.jsonl | small | Verification evidence |

Every file, every mirror

Nothing was discontinued: every quant is one click away. Hugging Face carries the curated picks, ModelScope always carries everything, and Ollama serves ready-to-run tags.

On Ollama every tag ships with the vision tower bundled and the 1M rope metadata baked in.

| File | Size | Hugging Face | ModelScope | Ollama |

|---|---|---|---|---|

| gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf | 16.8 GB | download | download | ollama run satgeze/gemma4-26b-uncensored-1m |

| mmproj-gemma26b-hauhau.gguf | 1.2 GB | download | download | bundled in every tag |

| mtp-gemma-4-26B-A4B-it.gguf | 252 MB | download | download | - |

Run it

llama.cpp, everything on:

llama-server -m gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf \
  -c 1048576 -np 1 --jinja \
  -md mtp-gemma-4-26B-A4B-it.gguf --spec-type draft-mtp --spec-draft-n-max 3 \
  --mmproj mmproj-gemma26b-hauhau.gguf

Ollama (1M and vision work; Ollama has no speculative decoding yet, so the MTP head adds no speed there):

FROM ./gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf
RENDERER gemma4
PARSER gemma4
PARAMETER num_ctx 262144

The RENDERER and PARSER lines avoid imported-GGUF template bugs under tool-heavy use. Raise num_ctx as memory allows.

How this was built

YaRN rope-scaling metadata (factor 4.0 over native 262,144) baked into the GGUF header with gguf-py; weights are bit-identical to the HauhauCS release, no fine-tuning. Gemma 4's dual-rope design takes YaRN on its global-attention layers. Certification harness: 10 needles per rung at depths 5 to 95 percent, temperature 0, seeded prompts, f16 KV only. Method and tooling: github.com/satindergrewal/aviary-1m.

For base capability benchmarks see Google's official Gemma 4 cards; uncensoring quality versus the official trunk has not been independently benchmarked here.

How to actually use a 1M-context model

Habits that measurably help, from our RULER, hop and adherence testing across this fleet:

  1. Re-state standing instructions near the end of long prompts; recency beats depth.
  2. One big reference dump beats a long accumulated conversation. Fresh session per task.
  3. After any compaction or summarization, repeat your active rules yourself.
  4. Prefill at 500K+ takes real time on any hardware; stage your questions accordingly.
  5. Know your quant: the results tables on this card show what each quant actually holds at depth; pick the strongest one your memory allows.

Credits

Base model and QAT: Google (Gemma license; its terms flow down to these files). Uncensoring and packaging: HauhauCS. MTP head: Unsloth (via the HauhauCS repo). 1M YaRN extension, benchmarking, and certification: SatGeze.

Sister repos: 12B | 26B-A4B | 31B | Qwen3.6-35B

Mirrors: Hugging Face | ModelScope

Run satgeze/Gemma4-26B-A4B-Uncensored-HauhauCS-1M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models