GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Mediform/gemma4-e4b-v13-assistant-rollout-gguf overview

gemma4 e4b v13 assistant rollout — GGUF BF16 + Q8 0 llama.cpp GGUF of Scribion's MTP draft assistant , rollout distilled against the finetuned v13 target E4B p…

ggufllama.cppspeculative-decodingmtpgemma4base_model:google/gemma-4-E4B-it-assistantbase_model:quantized:google/gemma-4-E4B-it-assistantlicense:gemmaendpoints_compatibleregion:usconversational

Runs locally from ~94.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
assistant-rollout-bf16.ggufGGUFBF16163.8 MBDownload
assistant-rollout-q8_0.ggufGGUFQ8_094.1 MBDownload

Model Details

Model IDMediform/gemma4-e4b-v13-assistant-rollout-gguf
AuthorMediform
Pipeline
Licensegemma
Base modelgoogle/gemma-4-E4B-it-assistant
Last modified2026-06-24T07:52:39.000Z

Model README

---

license: gemma

base_model: google/gemma-4-E4B-it-assistant

tags:

  • gguf
  • llama.cpp
  • speculative-decoding
  • mtp
  • gemma4

---

gemma4-e4b-v13-assistant-rollout — GGUF (BF16 + Q8_0)

llama.cpp GGUF of Scribion's MTP draft assistant, rollout-distilled against the finetuned

v13 target (E4B plain-LoRA r16). EAGLE-style multi-step rollout distillation on in-domain German

medical extraction data lifts deep-draft acceptance on long dialogues (froehlich +13.6% accept/step

at draft length 7) — a pure decode-speed win (speculative decoding is exact, output unchanged).

Files

| file | precision | size |

|------|-----------|------|

| assistant-rollout-bf16.gguf | bf16 | 172 MB |

| assistant-rollout-q8_0.gguf | Q8_0 | 99 MB |

78.5M-param 4-layer EAGLE-style draft model. Pair it as the draft model for speculative decoding

with the target:

Requires a llama.cpp build with Gemma-4 MTP-assistant / speculative support.

Run Mediform/gemma4-e4b-v13-assistant-rollout-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models