GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

kingjones777/Xing4.0-29B-A4B-ROCmFP4-GGUF overview

Xing4.0 29B A4B — ROCmFP4 GGUF Strix Halo / gfx1151 4 bit ROCmFP4 quantizations of XingChen AGI/Xing4.0 29B A4B https://huggingface.co/XingChen AGI/Xing4.0 29B…

ggufrocmrocmfp4strix-halogfx1151amdmoemlatext-generationbase_model:XingChen-AGI/Xing4.0-29B-A4Bbase_model:quantized:XingChen-AGI/Xing4.0-29B-A4Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~15.52 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
175
Likes
2
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Xing4.0-29B-A4B-Q4_0-ROCmFP4-COHERENT-noimatrix.ggufGGUFQ4_016.62 GBDownload
Xing4.0-29B-A4B-Q4_0-ROCmFP4-FAST-noimatrix.ggufGGUFQ4_015.52 GBDownload
Xing4.0-29B-A4B-Q4_0-ROCmFP4-STRIX_LEAN-noimatrix.ggufGGUFQ4_015.88 GBDownload

Model Details

Model IDkingjones777/Xing4.0-29B-A4B-ROCmFP4-GGUF
Authorkingjones777
Pipelinetext-generation
Licenseapache-2.0
Base modelXingChen-AGI/Xing4.0-29B-A4B
Last modified2026-09-20T01:03:40.000Z

Model README

---

library_name: gguf

license: apache-2.0

base_model: XingChen-AGI/Xing4.0-29B-A4B

base_model_relation: quantized

pipeline_tag: text-generation

tags:

- gguf

- rocm

- rocmfp4

- strix-halo

- gfx1151

- amd

- moe

- mla

---

Xing4.0-29B-A4B — ROCmFP4 GGUF (Strix Halo / gfx1151)

4-bit ROCmFP4 quantizations of XingChen-AGI/Xing4.0-29B-A4B,

China Telecom AI's 29B-total / 4B-active MoE (formerly the TeleChat series), built for AMD

Ryzen AI Max+ "Strix Halo" (gfx1151).

The single imatrix build lives in its own repo:

Xing4.0-29B-A4B-ROCmFP4-FAST-imatrix-GGUF.

⛔ These files need the ROCmFPX fork — stock llama.cpp cannot load them

Two independent reasons:

  1. ROCmFP4 tensor types are ≥ 100, past stock llama.cpp's GGML_TYPE_COUNT. The file is

rejected outright — no metadata edit fixes this.

  1. xing4 is a new architecture. Xing uses mHC (hyper-connections) + MLA with

n_embd_head_k=192 / n_embd_head_v=128, which no existing llama.cpp arch implements.

Build it from our fork — the arch lives on a branch, not main:

git clone -b xing4-port https://github.com/kingjones30/ROCmFPX.git
cd ROCmFPX && cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 \
  -DGGML_HIP_GRAPHS=ON -DGGML_HIP_NO_VMM=ON -DLLAMA_CURL=OFF -DLLAMA_BUILD_WEBUI=OFF
cmake --build build -j --target llama-server llama-quantize llama-perplexity

Pinned at commit cbc2cd5 (the port plus the hyper-connection correctness fix).

⚠️ main does not carry xing4 — you must use the xing4-port branch.

Files

| file | tier | ftype | size | PPL | tg128 | pp512 |

|---|---|---|---|---|---|---|

| …-COHERENT-noimatrix.gguf | COHERENT | 102 | 16.62 GiB | 9.0639 ± 0.274 | 29.04 t/s | 829.8 t/s |

| …-STRIX_LEAN-noimatrix.gguf | STRIX_LEAN (q6_K head) | 106 | 15.88 GiB | 9.1095 ± 0.277 | 28.98 t/s | 838.3 t/s |

| …-FAST-noimatrix.gguf | FAST (q4_K head) | 103 | 15.51 GiB | 9.2258 ± 0.282 | 29.67 t/s | 832.0 t/s |

BF16 reference: 8.3608 ± 0.249. Upstream HF transformers scores 8.4252 on the

identical tokens (see Verification).

Which one

  • COHERENT — best quality of the three. The default pick.
  • FAST — fastest decode and smallest. +1.8% PPL over COHERENT.
  • STRIX_LEAN — not recommended for this model. It is measurably the slowest at decode

and worse quality than COHERENT here. STRIX_LEAN's advantage comes from its PLE

upgrade, and Xing has no PLE. Published for completeness, not because it wins.

What we serve. Our own seat runs the COHERENT file at -c 32768 with

--spec-type draft-mtp --spec-draft-n-max 1 (see the MTP section).

Honest caveat: at ±0.27 the tier-to-tier PPL gaps are inside the individual error bars.

They are paired measurements on identical data so the ordering is meaningful, but do not

read the quality differences as large — the whole spread is 1.8%.

Usage

llama-server -m Xing4.0-29B-A4B-Q4_0-ROCmFP4-COHERENT-noimatrix.gguf \
  -c 32768 -ngl 99 -fit off --jinja

--jinja is required for tool calling. -fit off with an explicit -c is the house rule on

gfx1151 — auto-fit reads MemAvailable and can silently cut context or offload layers.

⚠️ This is a reasoning model — give it tokens

Thinking is on by default. A two-sentence answer consumed 569 completion tokens, and

max_tokens: 200 returns empty content. Budget ≥ 600, or you will conclude the model is

broken when it is only mid-thought. llama-server separates the trace into

reasoning_content correctly.

Reasoning is often emitted in Chinese even when the answer is English — expected for this

model, not a defect.

Tool calling

Verified working:

finish_reason: tool_calls
get_weather({"city":"Tokyo","unit":"celsius"})

Verification

Every number above is measured on the hardware named in the reproduction block. Nothing is

estimated or carried over from a similar model.

Reference check. The port was validated against upstream transformers running the same

model on the same tokens, not merely "looks coherent":

| | PPL |

|---|---|

| this port, BF16 | 8.3608 |

| HF transformers bf16, identical tokens | 8.4252 |

The port lands 0.7% below HF because llama.cpp computes the hyper-connection projection in

F32 where HF casts to bf16 — the direction and magnitude were predicted in advance by an

independent numpy reference before being measured.

⚠️ llama-perplexity scores only the second half of each window. A naive script scoring

every token from position 1 gives 11.91 for this same model and data. If you compare against

these numbers, match the scoring window or the comparison is meaningless.

Not tested

  • Contexts beyond 32K. The architecture declares 262144 but that context cannot be

allocated on 128 GiB hardware (see MTP section); 32768 is what was served and measured.

  • Vulkan. All numbers are the ROCm backend.

MTP: measured, and it is the biggest lever on this model

Layer 40 is the model's MTP block and its weights are in these files (block_count = 41,

nextn_predict_layers = 1). Measured 2026-09-19 on the STRIX_LEAN file, same flags, 589-token

prompt generating 160 tokens, 32K context, quiet box, ROCm0:

| --spec-type | prefill | decode | vs off | mean accepted length |

| --- | ---: | ---: | ---: | ---: |

| off | 454.8 t/s | 12.67 t/s | — | — |

| draft-mtp --spec-draft-n-max 1 | 388.9 t/s | 18.87 t/s | +48.9% | 1.65 |

| draft-mtp --spec-draft-n-max 2 | 411.9 t/s | 14.45 t/s | +14.0% | 1.83 |

Use --spec-draft-n-max 1. n-max 2 accepts slightly longer drafts (1.83 vs 1.65) and is still

much slower, because the extra draft work costs more than it returns. Prefill pays 14.5% for the

gain, so MTP favours generation-heavy traffic over long-prompt batch work.

⛔ MTP costs you half your context

The MTP draft context allocates its own full-size KV cache. Two consequences measured on a

128 GiB Strix Halo:

  • The declared 262144 context is not reachable on this hardware at all: the KV allocation asks

for 209,920 MiB (205 GiB) and fails. That is ~0.80 MiB per token, uncompressed.

  • With MTP enabled, 131072 still fails (failed to create MTP context) and 32768 is what fits.

⚠️ -ctk q8_0 -ctv q8_0 is silently ignored on this arch. At 131072 with both flags set the

allocation was 104,960 MiB — exactly half the 262144 figure, i.e. pure context scaling with no

quantisation benefit. Do not plan context around quantised KV here.

Reproduction

box     AMD Ryzen AI Max+ 395 "Strix Halo", gfx1151, ROCm 7.2.4, 123 GiB unified
build   github.com/kingjones30/ROCmFPX, arch xing4
source  XingChen-AGI/Xing4.0-29B-A4B (bf16 safetensors, 8307 tensors)
convert per-arch modular converter -> 936-tensor BF16 GGUF (62.4 GB)
quant   llama-quantize <src> <out> <ftype> 32
        COHERENT 102 | STRIX_LEAN 106 --output-tensor-type q6_K | FAST 103 --output-tensor-type q4_K

ppl     llama-perplexity -m <f> -f wiki.test.raw -c 512 --chunks 32 -ub 64 -ngl 99 -fit off
        (wikitext-2-raw test split)
speed   llama-bench -m <f> -p 512 -n 128 -ngl 99 -r 3      (quiet box, no co-resident seats)

⚠️ *-ub must be ≤ 128 when running an unquantized MoE on gfx1151* (BF16/F16 — e.g. when

computing your own imatrix). The ROCm mul_mat_id sorted fallback returns wrong values there:

test-backend-ops -o MUL_MAT_ID reports 1311/1320, with all 9 failures on f16/bf16. **All 920

quantized MUL_MAT_ID cases pass, so the files in this repo are unaffected** — this only bites

if you re-derive artifacts from the BF16 yourself.

Notes on the port

Getting this arch correct required fixing two defects in the hyper-connection ops that took

wikitext PPL from 80.05 to 8.3608:

  1. Sinkhorn eps placement. eps was added to the value rather than the denominator. At

1e-6 that sounds harmless, but after exp(logits − rowmax) the legitimate entries in a row

are ~2.5e-7 and ~6e-8, so it acts as a 5–17× floor that 20 Sinkhorn iterations amplify.

  1. comb orientation. hc_expand contracts ne1 and indexes the destination stream by ne0,

so the doubly-stochastic mixing matrix arrived transposed — rows and columns both sum to

~1, so nothing crashed and every structural check passed while the streams mixed wrongly.

Credits

Base model © China Telecom Artificial Intelligence Technology Co., Ltd.

(XingChen-AGI ·

GitHub), Apache-2.0. Quantized and

arch-ported by kingjones777; these GGUFs inherit Apache-2.0.

Run kingjones777/Xing4.0-29B-A4B-ROCmFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models