williamliao/gemma-4-31B-it-EAGLE3-Speculator-GGUF overview
license: apache 2.0 base model: RedHatAI/gemma 4 31B it speculator.eagle3 gemma 4 31B it base model relation: quantized library name: llama.cpp tags: gguf llam…
Runs locally from ~1.26 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | williamliao/gemma-4-31B-it-EAGLE3-Speculator-GGUF |
|---|---|
| Author | williamliao |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | RedHatAI/gemma-4-31B-it-speculator.eagle3,gemma-4-31B-it |
| Last modified | 2026-06-14T04:04:11.000Z |
Model README
---
license: apache-2.0
base_model:
- RedHatAI/gemma-4-31B-it-speculator.eagle3
- gemma-4-31B-it
base_model_relation: quantized
library_name: llama.cpp
tags:
- gguf
- llama.cpp
- eagle3
- speculative-decoding
- speculator
- draft-model
- gemma-4
- gemma
- redhatai
pipeline_tag: text-generation
---
Gemma 4 31B IT EAGLE3 Speculator GGUF
This repository contains GGUF conversions and quantizations of RedHatAI/gemma-4-31B-it-speculator.eagle3 for use with llama.cpp EAGLE3 speculative decoding.
> [!IMPORTANT]
> This is not a standalone chat model. It is an EAGLE3 draft/speculator model and must be used together with the matching target/verifier model.
- Target model:
gemma-4-31B-it - Speculator source:
RedHatAI/gemma-4-31B-it-speculator.eagle3 - Runtime: llama.cpp with
--spec-type draft-eagle3
Files
| File | Type | Notes |
| ---------------------------------------------- | -------------------------------- | ------------------------------------------------------------- |
| gemma-4-31B-it-speculator.eagle3-F16.gguf | EAGLE3 speculator GGUF | Converted from the original RedHatAI safetensors checkpoint |
| gemma-4-31B-it-speculator.eagle3-Q8_0.gguf | Quantized EAGLE3 speculator GGUF | Quantized from the F16 GGUF |
| gemma-4-31B-it-speculator.eagle3-Q4_K_M.gguf | Quantized EAGLE3 speculator GGUF | Quantized from the F16 GGUF; may be faster for draft decoding |
Usage with llama.cpp
Example:
llama-server \
-m gemma-4-31B-it-Q4_K_M.gguf \
-md gemma-4-31B-it-speculator.eagle3-Q4_K_M.gguf \
--spec-type draft-eagle3 \
--spec-draft-n-max 4 \
--spec-draft-p-min 0.5 \
-c 32768 \
-ngl 99 \
-fa on
Windows CMD example:
llama-server.exe ^
-m gemma-4-31B-it-Q4_K_M.gguf ^
-md gemma-4-31B-it-speculator.eagle3-Q4_K_M.gguf ^
--spec-type draft-eagle3 ^
--spec-draft-n-max 4 ^
--spec-draft-p-min 0.5 ^
-c 32768 ^
-ngl 99 ^
-fa on
PowerShell example:
.\llama-server.exe `
-m "gemma-4-31B-it-Q4_K_M.gguf" `
-md "gemma-4-31B-it-speculator.eagle3-Q4_K_M.gguf" `
--spec-type draft-eagle3 `
--spec-draft-n-max 4 `
--spec-draft-p-min 0.5 `
-c 32768 `
-ngl 99 `
-fa on
Important Notes
This GGUF file is only the draft/speculator model. You still need a compatible GGUF of the target model, such as a GGUF conversion or quantization of gemma-4-31B-it.
Do not use this speculator with unrelated models such as Gemma 4 12B, Gemma 3, Qwen, Llama, Mistral, or other non-matching models. EAGLE3 speculators are target-specific.
Even small differences in the target model, prompt format, quantization, or runtime settings may affect draft acceptance rate and overall speed.
Tested Configuration
Tested with:
- Runtime: llama.cpp with EAGLE3 support
- Target model: Gemma 4 31B IT GGUF
- Draft model: this EAGLE3 GGUF
- Example settings:
* --spec-type draft-eagle3
* --spec-draft-n-max 4
* --spec-draft-p-min 0.5
Local benchmark observations may vary depending on GPU, quantization, context length, batch size, sampling settings, and prompt type.
Benchmark Notes
In local testing, EAGLE3 showed stronger gains on structured or predictable outputs such as:
- code generation
- JSON output
- repeated pattern generation
- summarization
- code completion
It was less effective on some tasks such as translation, creative writing, and some open-ended prompts, where draft acceptance may be lower.
For this speculator, lower-bit draft quantization may perform better than F16 in some llama.cpp setups because the draft model becomes cheaper to run. In local testing, Q4_K_M was a practical candidate for draft decoding, while results may vary across hardware and runtime versions.
On smaller VRAM setups, the extra draft/speculator model may reduce the practical benefit of EAGLE3. In those cases, native MTP models may be more efficient.
Conversion
Converted with llama.cpp convert_hf_to_gguf.py using the original speculator repository and the matching target model directory.
Example conversion command:
python convert_hf_to_gguf.py \
RedHatAI/gemma-4-31B-it-speculator.eagle3 \
--outtype f16 \
--target-model-dir gemma-4-31B-it \
--outfile gemma-4-31B-it-speculator.eagle3-F16.gguf
PowerShell example:
python .\convert_hf_to_gguf.py `
"E:\OLLAMA_MODELS\gemma-4-31B-it-speculator.eagle3" `
--outtype f16 `
--target-model-dir "E:\OLLAMA_MODELS\gemma-4-31B-it" `
--outfile "E:\OLLAMA_MODELS\gemma-4-31B-it-speculator.eagle3-F16.gguf"
Quantization
The F16 GGUF can be quantized with llama-quantize.
Q8_0 example:
llama-quantize \
gemma-4-31B-it-speculator.eagle3-F16.gguf \
gemma-4-31B-it-speculator.eagle3-Q8_0.gguf \
Q8_0
Q4_K_M example:
llama-quantize \
gemma-4-31B-it-speculator.eagle3-F16.gguf \
gemma-4-31B-it-speculator.eagle3-Q4_K_M.gguf \
Q4_K_M
PowerShell example:
.\llama-quantize.exe `
"E:\OLLAMA_MODELS\gemma-4-31B-it-speculator.eagle3-F16.gguf" `
"E:\OLLAMA_MODELS\gemma-4-31B-it-speculator.eagle3-Q4_K_M.gguf" `
Q4_K_M
Credits
Original EAGLE3 speculator model by RedHatAI:
RedHatAI/gemma-4-31B-it-speculator.eagle3
Target model:
gemma-4-31B-it
GGUF support and runtime:
ggml-org/llama.cpp
License
This repository is a converted GGUF version of the original speculator model. The original model license and usage terms apply. Please refer to the upstream repositories for full license details.
Run williamliao/gemma-4-31B-it-EAGLE3-Speculator-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models