1bit-MONSTER/BlackMamba-2.8B-GGUF overview
BlackMamba 2.8B — GGUF Our own GGUF conversion of Zyphra's BlackMamba 2.8B https://huggingface.co/Zyphra/BlackMamba 2.8B : Mamba 1 layers alternating with a Sw…
Runs locally from ~2.76 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| BlackMamba-2.8B-Q8_0.gguf | GGUF | Q8_0 | 2.76 GB | Download |
Model Details
| Model ID | 1bit-MONSTER/BlackMamba-2.8B-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/BlackMamba-2.8B |
| Last modified | 2026-09-26T21:48:58.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/BlackMamba-2.8B
tags:
- gguf
- mamba
- moe
---
BlackMamba-2.8B — GGUF
Our own GGUF conversion of Zyphra's BlackMamba-2.8B: Mamba-1 layers alternating with a Switch MoE (sigmoid top-1 router with bias, gated GELU experts). The model code is ours (src/models/blackmamba.cpp in our llama.cpp fork); upstream llama.cpp has no BlackMamba.
Contents
BlackMamba-2.8B-Q8_0.gguf
Validation
- The same port matches Zyphra's own PyTorch code (FP32, CPU) at 96/96 teacher-forced positions on BlackMamba-1.5B. The 2.8B was not separately teacher-forced.
- Wikitext-2 test perplexity, 60 chunks of 512 tokens, Q8_0 on Vulkan: 15.28 ± 0.34 (BlackMamba-1.5B: 17.84 on the same test).
- Decode on Strix Halo (Radeon 8060S, Vulkan), Q8_0: 229 tok/s.
- It is a base model: it continues text rather than chatting.
Running it
With the 1bit engine:
1bit serve -m BlackMamba-2.8B-Q8_0.gguf --device vulkan
BlackMamba runs from our llama.cpp fork (branch 1bit/vulkan-upstream), which also has a Vulkan Mamba-1 scan; upstream llama.cpp has neither.
Attribution
- Base model: Zyphra/BlackMamba-2.8B, Apache 2.0.
- GGUF conversion and validation: the 1bit engine project, with the converter and model code in our llama.cpp fork.
- License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/BlackMamba-2.8B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models