1bit-MONSTER/Zamba-7B-v1-GGUF overview
Zamba 7B v1 — GGUF Our own GGUF conversion of Zyphra's Zamba 7B v1 https://huggingface.co/Zyphra/Zamba 7B v1 a Mamba/attention hybrid with a single shared atte…
Runs locally from ~7.31 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Zamba-7B-v1-Q8_0.gguf | GGUF | Q8_0 | 7.31 GB | Download |
Model Details
| Model ID | 1bit-MONSTER/Zamba-7B-v1-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/Zamba-7B-v1 |
| Last modified | 2026-09-26T10:30:31.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/Zamba-7B-v1
tags:
- gguf
- mamba
---
Zamba-7B-v1 — GGUF
Our own GGUF conversion of Zyphra's Zamba-7B-v1
(a Mamba/attention hybrid with a single shared attention block), added to the
1bit engine's architecture showcase.
From-scratch conversion and a new src/models/zamba.cpp in our llama.cpp fork.
Contents
Zamba-7B-v1-Q8_0.gguf
Validation
96/96 teacher-forced token agreement against transformers running the model in FP32
on CPU.
Getting the conversion right also meant fixing an upstream tokenizer bug we found
along the way: LlamaHfVocab was scoring all merges as -1000, so byte-pair merges
for tokenizer.json-only SentencePiece BPE models (like this one) were resolved in
arbitrary order — before the fix, only 318/1500 wikitext lines round-tripped correctly
against HF's own tokenizer; after (deriving scores from the actual merge ranks), all
1500/1500 do. GGUFs converted before this fix landed upstream need reconverting.
Other traps specific to this architecture: the shared attention block's weights are
stored once in the checkpoint, not duplicated per layer (naively duplicating them
inflates the parameter count by ~4.6B); in_proj interleaves [x, gate] rows; the
Mamba state-space ssm_n_group must stay at 0 or the conv state gets inflated; and the
tokenizer config's [PAD] token adds a 32001st entry that needs capping at the real
vocab size.
Measured performance (Strix Halo, Vulkan, Q8_0)
pp512: 266 tok/s · tg128: 8.9 tok/s
Decode is slower than a comparably-sized dense transformer here — Zamba v1's SSM scan
hasn't had the same Vulkan-side optimization pass its sibling Zamba2 has (see the
Zamba2 model card in this collection), so there's real headroom left on this number.
Running it
1bit serve -m Zamba-7B-v1-Q8_0.gguf --device vulkan
Attribution
- Base model: Zyphra/Zamba-7B-v1, Apache 2.0.
- GGUF conversion, llama.cpp architecture support, and the upstream
LlamaHfVocabfix: this project's engine team. - License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/Zamba-7B-v1-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models