1bit-MONSTER/ZAYA1-74B-preview-GGUF overview
ZAYA1 74B preview — GGUF Our own GGUF conversion of Zyphra's ZAYA1 74B preview https://huggingface.co/Zyphra/ZAYA1 74B preview : 74.8B parameters, 60 layers al…
Runs locally from ~42.58 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| ZAYA1-74B-preview-Q4_K_M.gguf | GGUF | Q4_K_M | 42.58 GB | Download |
Model Details
| Model ID | 1bit-MONSTER/ZAYA1-74B-preview-GGUF |
|---|---|
| Author | 1bit-MONSTER |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Zyphra/ZAYA1-74B-preview |
| Last modified | 2026-09-27T14:56:24.000Z |
Model README
---
license: apache-2.0
base_model:
- Zyphra/ZAYA1-74B-preview
tags:
- gguf
- zaya
- moe
---
ZAYA1-74B-preview — GGUF
Our own GGUF conversion of Zyphra's ZAYA1-74B-preview: 74.8B parameters, 60 layers alternating sliding-window attention (a 4,097-token window, rope theta 1e4) and full attention (rope theta 1e7). Sliding-window support was added in our llama.cpp fork (PR #11).
Contents
ZAYA1-74B-preview-Q4_K_M.gguf(45.7 GB), converted with--remoteto Q8_0 and requantized to Q4_K_M. Includes the model's chat template.
Validation
- The engine's end-to-end serve test on Strix Halo (
--device vulkan --ctx-size 8192): pass. Chat with the model's own template gives coherent answers (code, prose, a one-line fact). - Measured with
llama-benchon Vulkan (Radeon 8060S): prefill 505.8 tok/s (pp512), decode 35.4 tok/s (tg128). - There is no reference check against transformers: a 74B FP32 run needs about 300 GB of memory.
Running it
With the 1bit engine:
1bit serve -m ZAYA1-74B-preview-Q4_K_M.gguf --device vulkan --ctx-size 8192
It needs about 46 GB of GPU-visible memory.
ZAYA runs from our llama.cpp fork (branch 1bit/hrx-vulkan-patched); upstream llama.cpp has no ZAYA
model. The GGUFs must come from our converter: it writes the grouped convolution's weights
tap-major, which the graph expects.
Attribution
- Base model: Zyphra/ZAYA1-74B-preview, Apache 2.0.
- GGUF conversion and validation: the 1bit engine project, with the converter and model code in our llama.cpp fork.
- License: Apache 2.0, inherited from the base model.
Run 1bit-MONSTER/ZAYA1-74B-preview-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models