GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

1bit-MONSTER/Zamba-7B-v1-GGUF overview

Zamba 7B v1 — GGUF Our own GGUF conversion of Zyphra's Zamba 7B v1 https://huggingface.co/Zyphra/Zamba 7B v1 a Mamba/attention hybrid with a single shared atte…

ggufmambabase_model:Zyphra/Zamba-7B-v1base_model:quantized:Zyphra/Zamba-7B-v1license:apache-2.0endpoints_compatibleregion:us

Runs locally from ~7.31 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
—

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Zamba-7B-v1-Q8_0.ggufGGUFQ8_07.31 GBDownload

Model Details

Model ID1bit-MONSTER/Zamba-7B-v1-GGUF
Author1bit-MONSTER
Pipeline—
Licenseapache-2.0
Base modelZyphra/Zamba-7B-v1
Last modified2026-09-26T10:30:31.000Z

Model README

---

license: apache-2.0

base_model:

  • Zyphra/Zamba-7B-v1

tags:

  • gguf
  • mamba

---

Zamba-7B-v1 — GGUF

Our own GGUF conversion of Zyphra's Zamba-7B-v1

(a Mamba/attention hybrid with a single shared attention block), added to the

1bit engine's architecture showcase.

From-scratch conversion and a new src/models/zamba.cpp in our llama.cpp fork.

Contents

  • Zamba-7B-v1-Q8_0.gguf

Validation

96/96 teacher-forced token agreement against transformers running the model in FP32

on CPU.

Getting the conversion right also meant fixing an upstream tokenizer bug we found

along the way: LlamaHfVocab was scoring all merges as -1000, so byte-pair merges

for tokenizer.json-only SentencePiece BPE models (like this one) were resolved in

arbitrary order — before the fix, only 318/1500 wikitext lines round-tripped correctly

against HF's own tokenizer; after (deriving scores from the actual merge ranks), all

1500/1500 do. GGUFs converted before this fix landed upstream need reconverting.

Other traps specific to this architecture: the shared attention block's weights are

stored once in the checkpoint, not duplicated per layer (naively duplicating them

inflates the parameter count by ~4.6B); in_proj interleaves [x, gate] rows; the

Mamba state-space ssm_n_group must stay at 0 or the conv state gets inflated; and the

tokenizer config's [PAD] token adds a 32001st entry that needs capping at the real

vocab size.

Measured performance (Strix Halo, Vulkan, Q8_0)

pp512: 266 tok/s · tg128: 8.9 tok/s

Decode is slower than a comparably-sized dense transformer here — Zamba v1's SSM scan

hasn't had the same Vulkan-side optimization pass its sibling Zamba2 has (see the

Zamba2 model card in this collection), so there's real headroom left on this number.

Running it

1bit serve -m Zamba-7B-v1-Q8_0.gguf --device vulkan

Attribution

  • Base model: Zyphra/Zamba-7B-v1, Apache 2.0.
  • GGUF conversion, llama.cpp architecture support, and the upstream LlamaHfVocab fix: this project's engine team.
  • License: Apache 2.0, inherited from the base model.

Run 1bit-MONSTER/Zamba-7B-v1-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models