Model Intelligence Sheet
1bit-MONSTER/Zamba2-VL-1.2B-GGUF overview
Zamba2 VL 1.2B — GGUF Our own GGUF conversion of Zyphra's Zamba2 VL 1.2B https://huggingface.co/Zyphra/Zamba2 VL 1.2B : the Zamba2 language model with a Qwen2.…
Runs locally from ~1.25 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | 1bit-MONSTER/Zamba2-VL-1.2B-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/Zamba2-VL-1.2B |
| Last modified | 2026-09-26T21:57:05.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/Zamba2-VL-1.2B
tags:
- gguf
- zamba2
- vision
- multimodal
---
Zamba2-VL-1.2B — GGUF
Our own GGUF conversion of Zyphra's Zamba2-VL-1.2B: the Zamba2 language model with a Qwen2.5-VL vision tower. The converter is ours (llama.cpp #16, #17), including a chat template that places images the way Zyphra's does.
Contents
Zamba2-VL-1.2B-Q8_0.gguf: the language model.mmproj-Zamba2-VL-1.2B-F16.gguf: the vision tower.
Validation
- Vision embeddings against transformers: per-token cosine 0.99989 on the CPU and 0.99944 on Vulkan.
- On a synthetic image (a red square, a blue circle and "HELLO 42") the model gets the shapes, colours and layout right, and reads the text as "HElo 42".
Running it
With the 1bit engine:
1bit serve -m Zamba2-VL-1.2B-Q8_0.gguf --mmproj mmproj-Zamba2-VL-1.2B-F16.gguf --device vulkan
With llama.cpp (our fork), pass --jinja so the embedded chat template is used.
Attribution
- Base model: Zyphra/Zamba2-VL-1.2B, Apache 2.0.
- GGUF conversion and validation: the 1bit engine project, with the converter and model code in our llama.cpp fork.
- License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/Zamba2-VL-1.2B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models