GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

1bit-MONSTER/ZAYA1-VL-8B-GGUF overview

ZAYA1 VL 8B — GGUF Our own GGUF conversion of Zyphra's ZAYA1 VL 8B https://huggingface.co/Zyphra/ZAYA1 VL 8B : the ZAYA1 language model with a Qwen2.5 VL visio…

ggufzayavisionmultimodalbase_model:Zyphra/ZAYA1-VL-8Bbase_model:quantized:Zyphra/ZAYA1-VL-8Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.25 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6
Likes
0
Pipeline
—

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
ZAYA1-VL-8B-F16.ggufGGUFF1616.92 GBDownload
ZAYA1-VL-8B-Q4_K_M.ggufGGUFQ4_K_M5.31 GBDownload
mmproj-ZAYA1-VL-8B-F16.ggufGGUFF161.25 GBDownload

Model Details

Model ID1bit-MONSTER/ZAYA1-VL-8B-GGUF
Author1bit-MONSTER
Pipeline—
Licenseapache-2.0
Base modelZyphra/ZAYA1-VL-8B
Last modified2026-09-27T14:40:00.000Z

Model README

---

license: apache-2.0

base_model:

  • Zyphra/ZAYA1-VL-8B

tags:

  • gguf
  • zaya
  • vision
  • multimodal

---

ZAYA1-VL-8B — GGUF

Our own GGUF conversion of Zyphra's ZAYA1-VL-8B: the ZAYA1 language model with a Qwen2.5-VL vision tower, a vision-only LoRA applied to image tokens, and bidirectional attention within each image. Ported in our llama.cpp fork (PR #18); it was not available in GGUF form before.

Contents

  • ZAYA1-VL-8B-F16.gguf: the language model in F16, for further quantization.
  • ZAYA1-VL-8B-Q4_K_M.gguf: the language model in Q4_K_M, for serving.
  • mmproj-ZAYA1-VL-8B-F16.gguf: the vision tower (Qwen2.5-VL).

Validation

  • F16 against Zyphra's own code (their transformers branch zaya1-vl, FP32 on the CPU), teacher-forced on three image questions (101 tokens): top-1 agreement 100/101, both with the image attended causally and bidirectionally.
  • Q4_K_M on a synthetic image (a red square, a blue circle and the text "HELLO 42"): "two shapes: a red square and a blue circle ... black text that reads "HELLO 42"".
  • Adding the vision path leaves text-only ZAYA1-8B unchanged (same wikitext perplexity, 95/96 against FP32).
  • Decode on Strix Halo (Radeon 8060S, Vulkan), F16: 38-51 tok/s.

Running it

With the 1bit engine:

1bit serve -m ZAYA1-VL-8B-Q4_K_M.gguf --mmproj mmproj-ZAYA1-VL-8B-F16.gguf --device vulkan

With llama.cpp (our fork), decode each image in one ubatch: -b 4096 -ub 4096 covers Qwen2.5-VL's 4,096-token cap.

ZAYA runs from our llama.cpp fork (branch 1bit/hrx-vulkan-patched); upstream llama.cpp has no ZAYA

model. The GGUFs must come from our converter: it writes the grouped convolution's weights

tap-major, which the graph expects.

Attribution

Run 1bit-MONSTER/ZAYA1-VL-8B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models