GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

kingjones777/ZAYA1-8B-ROCmFPX-Q8_0-GGUF overview

πŸ”§ Runtime: build the ROCmFPX fork below Stock llama.cpp will not load this file. You need both the zaya architecture and the ROCmFP4 tensor types in one tree.…

ggufllama.cpprocmgfx1151strix-haloamdryzen-ai-max-395rocmfpxzayazyphraccatext-generationenbase_model:Zyphra/ZAYA1-basebase_model:quantized:Zyphra/ZAYA1-baselicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~8.59 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
115
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
ZAYA1-8B-Q8_0_ROCMFPX.ggufGGUFQ8_0_ROCMFPX8.59 GBDownload

Model Details

Model IDkingjones777/ZAYA1-8B-ROCmFPX-Q8_0-GGUF
Authorkingjones777
Pipelinetext-generation
Licenseapache-2.0
Base modelZyphra/ZAYA1-base
Last modified2026-08-28T04:07:50.000Z

Model README

---

license: apache-2.0

base_model: Zyphra/ZAYA1-base

base_model_relation: quantized

tags: [gguf, llama.cpp, rocm, gfx1151, strix-halo, amd, ryzen-ai-max-395, rocmfpx, zaya, zyphra, cca]

language: [en]

pipeline_tag: text-generation

---

> ### πŸ”§ Runtime: build the ROCmFPX fork below

> Stock llama.cpp will not load this file. You need both the zaya architecture

> and the ROCmFP4 tensor types in one tree. Upstream

> charlie12345/ROCmFPX has the ROCmFP4 types but

> not zaya. Our fork has both:

>

> kingjones30/ROCmFPX β€” a fork of charlie12345/ROCmFPX, branch main.

>

> ```bash

> git clone https://github.com/kingjones30/ROCmFPX.git

> cd ROCmFPX

> cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release

> cmake --build build --target llama-server llama-quantize -j$(nproc)

> ```

>

> Verified 2026-08-27 on gfx1151: clean clone β†’ 0 build errors β†’ llama-server loads a

> zaya ROCmFP4 GGUF from this family and generates coherent text.

> ### ⚠️ STOCK llama.cpp WILL NOT LOAD THIS MODEL

>

> 8.59 GiB Β· 21.08 tok/s on a Ryzen AI MAX+ 395.

ZAYA1-8B β€” ROCmFPX 8-bit GGUF

An 8-bit ROCmFPX quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo),

quantized from BF16 GGUF β€” a lossless source, not a requantization of a lower-bit build.

| | |

|---|---|

| File | ZAYA1-8B-Q8_0_ROCMFPX.gguf |

| Size | 8.59 GiB |

| BPW | 8.32 |

| ftype | Q8_0_ROCMFPX (111) |

β›” tie_word_embeddings is TRUE, so output.weight does not exist β€” --output-tensor-type is a silent no-op here and --token-embedding-type is the flag that lands (262K vocab).

β›” Requires a llama.cpp with the ROCmFPX quant types

Q8_0_ROCMFPX (ftype 111) and Q8_0_ROCMFPX_AGENT (ftype 115) exist only in

charlie12345/ROCmFPX, not upstream llama.cpp.

Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated

"Use this model" commands above.

---

All quant variants

Three builds of this model, all measured in one session on one box with one binary

(Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, ROCmFPX-2809dc5) β€” so these rows are directly

comparable. Median of 3, warm-up discarded, otherwise-idle box.

| variant | ftype | size | bpw | decode (median) | range | repo |

|---|---|---|---|---|---|---|

| 4-bit COHERENT | 102 | 4.86 GiB | 4.71 | 23.04 | 22.83 – 23.70 | ZAYA1-8B-ROCmFP4-GGUF |

| 8-bit AGENT | 115 | 8.72 GiB | 8.45 | 21.02 | 20.95 – 21.47 | ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF |

| 8-bit plain | 111 | 8.59 GiB | 8.32 | 21.08 | 20.99 – 21.20 | ZAYA1-8B-ROCmFPX-Q8_0-GGUF |

⚠️ Decode is ~88% weight-independent on this architecture (the CCA grouped conv is ~55% of decode). All three builds land within ~10% of each other; the 4-bit is smallest and marginally fastest. No 8-bit or 4-bit format will make this model meaningfully faster.

What AGENT actually changes: it keeps far more tensors at true Q8_0 instead of the

packed 8-bit type β€” measured in these files, 154 tensors vs 1 tensor. On models with an

MTP draft head that raises draft acceptance and wins ~6%; these two models have no MTP head,

and here the two 8-bit builds are within noise of each other.

Correctness: All three builds answer correctly. On some prompts content is empty with finish_reason=length and the correct answer sits in reasoning_content β€” this model is verbose, give it β‰₯1024 tokens.

Per-tensor types (audited in this finished file)

token_embd Q8_0 Β· 400 packed TYPE_103 Β· 842 F32 Β· 40 BF16 (CCA conv kept at BF16) Β· 1 Q8_0

---

What was NOT measured

  • No perplexity run, and no quality A/B against the source. The checks above are

memorized-fact prompts β€” necessary but not sufficient; a damaged model can pass them.

  • No long-context testing. Β· No tool-calling evaluation.

Base model licence inherited; credit for the model goes to its authors.

Run kingjones777/ZAYA1-8B-ROCmFPX-Q8_0-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models