GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

1bit-MONSTER/BlackMamba-2.8B-GGUF overview

BlackMamba 2.8B — GGUF Our own GGUF conversion of Zyphra's BlackMamba 2.8B https://huggingface.co/Zyphra/BlackMamba 2.8B : Mamba 1 layers alternating with a Sw…

ggufmambamoebase_model:Zyphra/BlackMamba-2.8Bbase_model:quantized:Zyphra/BlackMamba-2.8Blicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~2.76 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
—

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
BlackMamba-2.8B-Q8_0.ggufGGUFQ8_02.76 GBDownload

Model Details

Model ID1bit-MONSTER/BlackMamba-2.8B-GGUF
Author1bit-MONSTER
Pipeline—
Licenseapache-2.0
Base modelZyphra/BlackMamba-2.8B
Last modified2026-09-26T21:48:58.000Z

Model README

---

license: apache-2.0

base_model:

  • Zyphra/BlackMamba-2.8B

tags:

  • gguf
  • mamba
  • moe

---

BlackMamba-2.8B — GGUF

Our own GGUF conversion of Zyphra's BlackMamba-2.8B: Mamba-1 layers alternating with a Switch MoE (sigmoid top-1 router with bias, gated GELU experts). The model code is ours (src/models/blackmamba.cpp in our llama.cpp fork); upstream llama.cpp has no BlackMamba.

Contents

  • BlackMamba-2.8B-Q8_0.gguf

Validation

  • The same port matches Zyphra's own PyTorch code (FP32, CPU) at 96/96 teacher-forced positions on BlackMamba-1.5B. The 2.8B was not separately teacher-forced.
  • Wikitext-2 test perplexity, 60 chunks of 512 tokens, Q8_0 on Vulkan: 15.28 ± 0.34 (BlackMamba-1.5B: 17.84 on the same test).
  • Decode on Strix Halo (Radeon 8060S, Vulkan), Q8_0: 229 tok/s.
  • It is a base model: it continues text rather than chatting.

Running it

With the 1bit engine:

1bit serve -m BlackMamba-2.8B-Q8_0.gguf --device vulkan

BlackMamba runs from our llama.cpp fork (branch 1bit/vulkan-upstream), which also has a Vulkan Mamba-1 scan; upstream llama.cpp has neither.

Attribution

Run 1bit-MONSTER/BlackMamba-2.8B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models