rpatel622/mamba2-130m-hf-Q8_0-GGUF overview
Mamba2 130M Q8 0 GGUF Quantized GGUF version of benchang1110/mamba2 130m hf https://huggingface.co/benchang1110/mamba2 130m hf for use with llama.cpp https://g…
Runs locally from ~172.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
Model README
---
tags:
- mamba2
- ssms
- quantized
- q8_0
- gguf
- llama.cpp
license: apache-2.0
---
Mamba2-130M Q8_0 GGUF
Quantized GGUF version of benchang1110/mamba2-130m-hf for use with llama.cpp.
Details
| Attribute | Value |
|-----------|-------|
| Base model | benchang1110/mamba2-130m-hf |
| Original size | 335 MB (bfloat16 safetensors) |
| Quantized size | 173 MB |
| Format | GGUF V3 |
| Quantization | Q8_0 (50 tensors quantized, 169 kept F32) |
| Architecture | mamba2 |
| Layers | 24 |
| Vocab size | 50,288 |
| Context length | 1,048,576 |
Usage (llama.cpp / llama-cpp-python)
CLI
License
Apache-2.0 (same as base model)
Run rpatel622/mamba2-130m-hf-Q8_0-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models