1bit-MONSTER/Zamba2-1.2B-instruct-GGUF overview
Zamba2 1.2B instruct — GGUF GGUF conversion of Zyphra's Zamba2 1.2B instruct https://huggingface.co/Zyphra/Zamba2 1.2B instruct a Mamba2/attention hybrid , ser…
Runs locally from ~1.71 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Zamba2-1.2B-instruct-Q8_0.gguf | GGUF | Q8_0 | 1.71 GB | Download |
Model Details
| Model ID | 1bit-MONSTER/Zamba2-1.2B-instruct-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/Zamba2-1.2B-instruct |
| Last modified | 2026-09-26T10:15:32.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/Zamba2-1.2B-instruct
tags:
- gguf
- mamba
---
Zamba2-1.2B-instruct — GGUF
GGUF conversion of Zyphra's Zamba2-1.2B-instruct
(a Mamba2/attention hybrid), served through the 1bit engine.
Zamba2 architecture support was carried from upstream llama.cpp (ggml-org#21412); we
added the Vulkan-side fix that makes it actually fast on this hardware.
Contents
Zamba2-1.2B-instruct-Q8_0.gguf
The Vulkan fix
Zamba2's Vulkan SSM_SCAN kernel didn't handle d_state = 64 correctly, which capped
decode speed badly across the whole model family — we measured 4.3 tok/s at 2.7B and
2.1 tok/s at 7B before fixing it. After the fix, those became 45 and 17 tok/s
respectively (same hardware). This 1.2B checkpoint benefits from the same fix.
Measured performance (Strix Halo, Vulkan, Q8_0, this checkpoint)
pp512: 3302 tok/s · tg128: 83.4 tok/s
Running it
1bit serve -m Zamba2-1.2B-instruct-Q8_0.gguf --device vulkan
Attribution
- Base model: Zyphra/Zamba2-1.2B-instruct, Apache 2.0.
- Zamba2 llama.cpp support: carried from upstream
ggml-org#21412. - Vulkan
SSM_SCANd_state=64 fix: this project's engine team. - License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/Zamba2-1.2B-instruct-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models