kingjones777/ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF overview
⚠️ STOCK llama.cpp WILL NOT LOAD THIS MODEL 8.72 GiB · 21.02 tok/s on a Ryzen AI MAX+ 395. ZAYA1 8B — ROCmFPX 8 bit AGENT GGUF An 8 bit ROCmFPX quantization fo…
Runs locally from ~8.72 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| ZAYA1-8B-Q8_0_ROCMFPX_AGENT.gguf | GGUF | Q8_0_ROCMFPX_AGENT | 8.72 GB | Download |
Model Details
| Model ID | kingjones777/ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF |
|---|---|
| Author | kingjones777 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Zyphra/ZAYA1-base |
| Last modified | 2026-08-16T15:17:58.000Z |
Model README
---
license: apache-2.0
base_model: Zyphra/ZAYA1-base
base_model_relation: quantized
tags: [gguf, llama.cpp, rocm, gfx1151, strix-halo, amd, ryzen-ai-max-395, rocmfpx, zaya, zyphra, cca]
language: [en]
pipeline_tag: text-generation
---
> ### ⚠️ STOCK llama.cpp WILL NOT LOAD THIS MODEL
>
> 8.72 GiB · 21.02 tok/s on a Ryzen AI MAX+ 395.
ZAYA1-8B — ROCmFPX 8-bit AGENT GGUF
An 8-bit ROCmFPX quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo),
quantized from BF16 GGUF — a lossless source, not a requantization of a lower-bit build.
| | |
|---|---|
| File | ZAYA1-8B-Q8_0_ROCMFPX_AGENT.gguf |
| Size | 8.72 GiB |
| BPW | 8.45 |
| ftype | Q8_0_ROCMFPX_AGENT (115) |
⛔ tie_word_embeddings is TRUE, so output.weight does not exist — --output-tensor-type is a silent no-op here and --token-embedding-type is the flag that lands (262K vocab).
⛔ Requires a llama.cpp with the ROCmFPX quant types
Q8_0_ROCMFPX (ftype 111) and Q8_0_ROCMFPX_AGENT (ftype 115) exist only in
charlie12345/ROCmFPX, not upstream llama.cpp.
Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated
"Use this model" commands above.
---
All quant variants
Three builds of this model, all measured in one session on one box with one binary
(Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, ROCmFPX-2809dc5) — so these rows are directly
comparable. Median of 3, warm-up discarded, otherwise-idle box.
| variant | ftype | size | bpw | decode (median) | range | repo |
|---|---|---|---|---|---|---|
| 4-bit COHERENT | 102 | 4.86 GiB | 4.71 | 23.04 | 22.83 – 23.70 | ZAYA1-8B-ROCmFP4-GGUF |
| 8-bit AGENT | 115 | 8.72 GiB | 8.45 | 21.02 | 20.95 – 21.47 | ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF |
| 8-bit plain | 111 | 8.59 GiB | 8.32 | 21.08 | 20.99 – 21.20 | ZAYA1-8B-ROCmFPX-Q8_0-GGUF |
⚠️ Decode is ~88% weight-independent on this architecture (the CCA grouped conv is ~55% of decode). All three builds land within ~10% of each other; the 4-bit is smallest and marginally fastest. No 8-bit or 4-bit format will make this model meaningfully faster.
What AGENT actually changes: it keeps far more tensors at true Q8_0 instead of the
packed 8-bit type — measured in these files, 154 tensors vs 1 tensor. On models with an
MTP draft head that raises draft acceptance and wins ~6%; these two models have no MTP head,
and here the two 8-bit builds are within noise of each other.
Correctness: All three builds answer correctly. On some prompts content is empty with finish_reason=length and the correct answer sits in reasoning_content — this model is verbose, give it ≥1024 tokens.
Per-tensor types (audited in this finished file)
token_embd Q8_0 · 247 packed TYPE_103 · 842 F32 · 40 BF16 · 154 Q8_0
---
What was NOT measured
- No perplexity run, and no quality A/B against the source. The checks above are
memorized-fact prompts — necessary but not sufficient; a damaged model can pass them.
- No long-context testing. · No tool-calling evaluation.
Base model licence inherited; credit for the model goes to its authors.
Run kingjones777/ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models