1bit-MONSTER/BlackMamba-1.5B-GGUF overview
BlackMamba 1.5B — GGUF Our own GGUF conversion of Zyphra's BlackMamba 1.5B https://huggingface.co/Zyphra/BlackMamba 1.5B a hybrid Mamba/MoE architecture , adde…
Runs locally from ~1.45 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| BlackMamba-1.5B-Q8_0.gguf | GGUF | Q8_0 | 1.45 GB | Download |
Model Details
| Model ID | 1bit-MONSTER/BlackMamba-1.5B-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/BlackMamba-1.5B |
| Last modified | 2026-09-26T10:14:23.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/BlackMamba-1.5B
tags:
- gguf
- mamba
- moe
---
BlackMamba-1.5B — GGUF
Our own GGUF conversion of Zyphra's BlackMamba-1.5B
(a hybrid Mamba/MoE architecture), added to the 1bit engine's
architecture showcase. From-scratch conversion and a new src/models/blackmamba.cpp in
our llama.cpp fork — BlackMamba wasn't previously supported by llama.cpp.
Contents
BlackMamba-1.5B-Q8_0.gguf
Validation
96/96 teacher-forced token agreement against Zyphra's own PyTorch reference implementation
(run on CPU with Triton/CUDA kernels stubbed out).
A few non-obvious things the config doesn't tell you, that our conversion had to get
right: the config says swiglu but the activation is actually gated erf-GELU
(F.gelu); the router's Linear has a bias; the final norm is a biased LayerNorm,
not RMSNorm; the checkpoint's embedding key is embedding.word_embeddings.weight; and
config.json doesn't declare an architectures field at all.
Measured performance (Strix Halo, Vulkan, Q8_0)
pp512: 12890 tok/s · tg128: 378 tok/s
Running it
1bit serve -m BlackMamba-1.5B-Q8_0.gguf --device vulkan
Attribution
- Base model: Zyphra/BlackMamba-1.5B, Apache 2.0.
- GGUF conversion and llama.cpp architecture support: this project's engine team.
- License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/BlackMamba-1.5B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models