GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

1bit-MONSTER/ZAYA1-74B-preview-GGUF overview

ZAYA1 74B preview — GGUF Our own GGUF conversion of Zyphra's ZAYA1 74B preview https://huggingface.co/Zyphra/ZAYA1 74B preview : 74.8B parameters, 60 layers al…

ggufzayamoebase_model:Zyphra/ZAYA1-74B-previewbase_model:quantized:Zyphra/ZAYA1-74B-previewlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~42.58 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
—

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
ZAYA1-74B-preview-Q4_K_M.ggufGGUFQ4_K_M42.58 GBDownload

Model Details

Model ID1bit-MONSTER/ZAYA1-74B-preview-GGUF
Author1bit-MONSTER
Pipeline—
Licenseapache-2.0
Base modelZyphra/ZAYA1-74B-preview
Last modified2026-09-27T14:56:24.000Z

Model README

---

license: apache-2.0

base_model:

  • Zyphra/ZAYA1-74B-preview

tags:

  • gguf
  • zaya
  • moe

---

ZAYA1-74B-preview — GGUF

Our own GGUF conversion of Zyphra's ZAYA1-74B-preview: 74.8B parameters, 60 layers alternating sliding-window attention (a 4,097-token window, rope theta 1e4) and full attention (rope theta 1e7). Sliding-window support was added in our llama.cpp fork (PR #11).

Contents

  • ZAYA1-74B-preview-Q4_K_M.gguf (45.7 GB), converted with --remote to Q8_0 and requantized to Q4_K_M. Includes the model's chat template.

Validation

  • The engine's end-to-end serve test on Strix Halo (--device vulkan --ctx-size 8192): pass. Chat with the model's own template gives coherent answers (code, prose, a one-line fact).
  • Measured with llama-bench on Vulkan (Radeon 8060S): prefill 505.8 tok/s (pp512), decode 35.4 tok/s (tg128).
  • There is no reference check against transformers: a 74B FP32 run needs about 300 GB of memory.

Running it

With the 1bit engine:

1bit serve -m ZAYA1-74B-preview-Q4_K_M.gguf --device vulkan --ctx-size 8192

It needs about 46 GB of GPU-visible memory.

ZAYA runs from our llama.cpp fork (branch 1bit/hrx-vulkan-patched); upstream llama.cpp has no ZAYA

model. The GGUFs must come from our converter: it writes the grouped convolution's weights

tap-major, which the graph expects.

Attribution

Run 1bit-MONSTER/ZAYA1-74B-preview-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models