1bit-MONSTER/ZAYA1-VL-8B-GGUF overview
ZAYA1 VL 8B — GGUF Our own GGUF conversion of Zyphra's ZAYA1 VL 8B https://huggingface.co/Zyphra/ZAYA1 VL 8B : the ZAYA1 language model with a Qwen2.5 VL visio…
Runs locally from ~1.25 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | 1bit-MONSTER/ZAYA1-VL-8B-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/ZAYA1-VL-8B |
| Last modified | 2026-09-27T14:40:00.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/ZAYA1-VL-8B
tags:
- gguf
- zaya
- vision
- multimodal
---
ZAYA1-VL-8B — GGUF
Our own GGUF conversion of Zyphra's ZAYA1-VL-8B: the ZAYA1 language model with a Qwen2.5-VL vision tower, a vision-only LoRA applied to image tokens, and bidirectional attention within each image. Ported in our llama.cpp fork (PR #18); it was not available in GGUF form before.
Contents
ZAYA1-VL-8B-F16.gguf: the language model in F16, for further quantization.ZAYA1-VL-8B-Q4_K_M.gguf: the language model in Q4_K_M, for serving.mmproj-ZAYA1-VL-8B-F16.gguf: the vision tower (Qwen2.5-VL).
Validation
- F16 against Zyphra's own code (their transformers branch
zaya1-vl, FP32 on the CPU), teacher-forced on three image questions (101 tokens): top-1 agreement 100/101, both with the image attended causally and bidirectionally. - Q4_K_M on a synthetic image (a red square, a blue circle and the text "HELLO 42"): "two shapes: a red square and a blue circle ... black text that reads "HELLO 42"".
- Adding the vision path leaves text-only ZAYA1-8B unchanged (same wikitext perplexity, 95/96 against FP32).
- Decode on Strix Halo (Radeon 8060S, Vulkan), F16: 38-51 tok/s.
Running it
With the 1bit engine:
1bit serve -m ZAYA1-VL-8B-Q4_K_M.gguf --mmproj mmproj-ZAYA1-VL-8B-F16.gguf --device vulkan
With llama.cpp (our fork), decode each image in one ubatch: -b 4096 -ub 4096 covers Qwen2.5-VL's 4,096-token cap.
ZAYA runs from our llama.cpp fork (branch 1bit/hrx-vulkan-patched); upstream llama.cpp has no ZAYA
model. The GGUFs must come from our converter: it writes the grouped convolution's weights
tap-major, which the graph expects.
Attribution
- Base model: Zyphra/ZAYA1-VL-8B, Apache 2.0.
- GGUF conversion and validation: the 1bit engine project, with the converter and model code in our llama.cpp fork.
- License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/ZAYA1-VL-8B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models