1bit-MONSTER/Mamba-370M-GGUF overview
Mamba 370M — GGUF GGUF conversion of Zyphra's Mamba 370M https://huggingface.co/Zyphra/Mamba 370M pure Mamba 1, no attention/MoE , included in the 1bit engine …
Runs locally from ~714.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| mamba-370m-f16.gguf | GGUF | F16 | 714.7 MB | Download |
Model Details
| Model ID | 1bit-MONSTER/Mamba-370M-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/Mamba-370M |
| Last modified | 2026-09-26T11:45:21.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/Mamba-370M
tags:
- gguf
- mamba
---
Mamba-370M — GGUF
GGUF conversion of Zyphra's Mamba-370M
(pure Mamba-1, no attention/MoE), included in the
1bit engine's architecture showcase as the
reference case for our Vulkan Mamba-1 SSM_SCAN fix.
Contents
mamba-370m-f16.gguf
The Vulkan fix
Vulkan's Mamba-1 SSM_SCAN kernel had a correctness/performance gap that capped decode
speed hard. Fixing it took this model from 25.8 tok/s to 173.5 tok/s on Strix Halo —
this repo exists mainly to document and reproduce that number on real hardware.
Measured performance (Strix Halo, Vulkan, F16, this checkpoint)
pp512: 6530 tok/s · tg128: 163 tok/s
Running it
1bit serve -m mamba-370m-f16.gguf --device vulkan
Attribution
- Base model: Zyphra/Mamba-370M, Apache 2.0.
- GGUF conversion and the Vulkan
SSM_SCANfix: this project's engine team. - License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/Mamba-370M-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models