GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

1bit-MONSTER/ZAYA1-reasoning-base-GGUF overview

ZAYA1 reasoning base — GGUF Our own GGUF conversion of Zyphra's ZAYA1 reasoning base https://huggingface.co/Zyphra/ZAYA1 reasoning base . It keeps Zyphra's ori…

ggufzayamoebase-modelbase_model:Zyphra/ZAYA1-reasoning-basebase_model:quantized:Zyphra/ZAYA1-reasoning-baselicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~5.17 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
26
Likes
0
Pipeline
—

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
ZAYA1-reasoning-base-Q4_K_M.ggufGGUFQ4_K_M5.17 GBDownload

Model Details

Model ID1bit-MONSTER/ZAYA1-reasoning-base-GGUF
Author1bit-MONSTER
Pipeline—
Licenseapache-2.0
Base modelZyphra/ZAYA1-reasoning-base
Last modified2026-09-27T14:44:42.000Z

Model README

---

license: apache-2.0

base_model:

  • Zyphra/ZAYA1-reasoning-base

tags:

  • gguf
  • zaya
  • moe
  • base-model

---

ZAYA1-reasoning-base — GGUF

Our own GGUF conversion of Zyphra's ZAYA1-reasoning-base. It keeps Zyphra's original Megatron-style checkpoint layout, which transformers cannot load; our converter maps it to ZAYA's standard layout (llama.cpp #14). Converting Zyphra/ZAYA1-8B-legacy this way gives tensors byte-identical to the ones from Zyphra/ZAYA1-8B.

Contents

  • ZAYA1-reasoning-base-Q4_K_M.gguf (8B shape: 40 layers, 16 experts; rope theta 1e6; 32,768-token context).

Validation

  • Wikitext-2 test perplexity, 60 chunks of 512 tokens, Q4_K_M on Vulkan (Radeon 8060S): 10.15 ± 0.22. Decode 89-91 tok/s (chat, Vulkan). It answers chat prompts with its own template, thinking first.
  • Raw-text perplexity favours base models; the post-trained ZAYA1-8B scores 32.13 on the same test.

Running it

With the 1bit engine:

1bit serve -m ZAYA1-reasoning-base-Q4_K_M.gguf --device vulkan

ZAYA runs from our llama.cpp fork (branch 1bit/hrx-vulkan-patched); upstream llama.cpp has no ZAYA

model. The GGUFs must come from our converter: it writes the grouped convolution's weights

tap-major, which the graph expects.

Attribution

Run 1bit-MONSTER/ZAYA1-reasoning-base-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models