GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

IsValorum/Xing4.0-29B-A4B-APEX-I-MiniPlus-V2.1-Abliterated-GGUF overview

<a id="quick navigation" </a Quick Navigation Index 1. Xing4.0 Architecture & Abliterated Source toc 01 2. Measured Fidelity & MiniPlus / NanoPlus Comparison t…

ggufllama.cppquantizedquantizationapexapex-quantapex-i-miniplusv2.1custom-quantizationimatrixabliterateduncensoredmoemlahyper-connectionsmtpreasoningcodingagenticlong-contextxing4_0text-generationconversationalen

Runs locally from ~12.82 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
578
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.ggufGGUFGGUF12.82 GBDownload

Model Details

Model IDIsValorum/Xing4.0-29B-A4B-APEX-I-MiniPlus-V2.1-Abliterated-GGUF
AuthorIsValorum
Pipelinetext-generation
Licenseapache-2.0
Base modelhuihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
Last modified2026-10-03T17:42:39.000Z

Model README

---

base_model: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated

base_model_relation: quantized

quantized_by: IsValorum

library_name: gguf

license: apache-2.0

language:

- en

- zh

pipeline_tag: text-generation

tags:

- gguf

- llama.cpp

- quantized

- quantization

- apex

- apex-quant

- apex-i-miniplus

- v2.1

- custom-quantization

- imatrix

- abliterated

- uncensored

- moe

- mla

- hyper-connections

- mtp

- reasoning

- coding

- agentic

- long-context

- xing4_0

---

<a id="quick-navigation"></a>Quick Navigation Index

  1. Xing4.0 Architecture & Abliterated Source
  2. Measured Fidelity & MiniPlus / NanoPlus Comparison
  3. Model Files & Technical Specifications
  4. Surgical Tensor Quantization Map (Audited from GGUF)
  5. Hardware Throughput: GPU Projections & Tested RAM Offload
  6. 256K Native Context & Deployment Notes
  7. llama.cpp Quickstart
  8. Official Xing4.0 Generation Parameters
  9. Optional Support

Xing4.0-29B-A4B APEX-I-MiniPlus-V2.1 Abliterated GGUF

29B / 4B Active MoE · mHC + MLA + MTP · 256K Native Context · 13.76 GB

> [!IMPORTANT]

> ### APEX-I-MINIPLUS V2.1 — HIGH-FIDELITY 13–14 GB RELEASE

> This APEX-I-MiniPlus-V2.1 build uses a tensor-by-tensor mixed-precision layout designed for Xing4.0 rather than a flat quantization. The final GGUF is 13.76 GB (12.82 GiB / 3.42 BPW) and preserves the measured language-modeling quality of the source at a Q5_K_M-class quality target while remaining compact enough for consumer GPU and hybrid RAM/GPU deployments.

> [!NOTE]

> ### ABLITERATED / UNCENSORED SOURCE

> This GGUF quantizes huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated. Huihui describes that checkpoint as an uncensored version of the official Xing4.0 model created with an abliteration workflow and references Sumandora/remove-refusals-with-transformers for the method. No additional refusal-removal procedure is performed by this GGUF quantization.

> [!TIP]

> ### MEASURED FIDELITY

>

> | Quantization | File Size | Memory Footprint | Average BPW | WikiText-2 Perplexity | Delta vs. BF16 | Quality Target |

> | :--- | :---: | :---: | :---: | :---: | :---: | :---: |

> | BF16 source | 62.40 GB | 58.11 GiB | 16.00 | 7.3060 +/- 0.18633 | Baseline | Full-precision reference |

> | APEX-I-MiniPlus V2.1 (CURRENT) | 13.76 GB | 12.82 GiB | 3.42 | 7.7238 +/- 0.19671 | +0.4178 (+5.71%) | Q5_K_M-class quality |

> | APEX-I-NanoPlus | 11.74 GB | 10.94 GiB | 2.92 | 8.3197 +/- 0.21528 | +1.0137 (+13.87%) | Q4_K_M / Q4_K_S-class quality |

>

> The PPL values above are direct measurements from the final GGUFs. The MiniPlus/NanoPlus quality labels describe the intended practical fidelity tier; they are not derived from generic flat-quant PPL tables.

---

<a id="toc-01"></a>

Xing4.0 Architecture & Abliterated Source

The official XingChen-AGI/Xing4.0-29B-A4B is a 29B-parameter Mixture-of-Experts model with approximately 4B parameters activated per token. The upstream architecture combines mHC (manifold-constrained Hyper-Connections), Multi-head Latent Attention (MLA), and one Multi-Token Prediction (MTP) layer, and is positioned for reasoning, coding, agentic planning, and tool use. XingChen reports a 256K native context window, extensible to 512K, and training on the Ascend NPU / MindSpore stack.

The architecture is not a uniform 40-layer MoE:

  • 40 main transformer layers in total.
  • Layers 0–1 are dense FFN layers (first_k_dense_replace: 2).
  • Layers 2–39 are MoE layers with 64 routed experts, 4 selected per token, plus one always-active shared expert.
  • One additional MTP block is stored as blk.40 in the GGUF.
  • Native context: 262,144 tokens (256K); upstream documents extension to 512K.

The MiniPlus allocation preserves the two dense entry layers at high precision, keeps all main MoE routers in F32, protects shared experts in Q5_K, and uses calibrated 3-bit precision for the routed expert backbone.

---

<a id="toc-02"></a>

Measured Fidelity & MiniPlus / NanoPlus Comparison

Both APEX-I editions were evaluated directly from their compiled GGUF binaries with llama-perplexity on WikiText-2 using the same 2048-context / 512-batch / 10-chunk setup:

| Metric | BF16 Source | MiniPlus V2.1 | NanoPlus |

| :--- | :---: | :---: | :---: |

| WikiText-2 PPL | 7.3060 +/- 0.18633 | 7.7238 +/- 0.19671 | 8.3197 +/- 0.21528 |

| Delta PPL | 0 | +0.4178 (+5.71%) | +1.0137 (+13.87%) |

| GGUF size | 62.40 GB BF16 source | 13.76 GB | 11.74 GB |

| Average BPW | 16.00 | 3.42 | 2.92 |

| Practical quality target | Full precision | Q5_K_M class | Q4_K_M / Q4_K_S class |

Choose MiniPlus V2.1 when fidelity is the priority and the extra ~2 GB is acceptable. Choose NanoPlus when the smaller footprint materially improves full GPU offload or RAM-streaming behavior.

---

<a id="toc-03"></a>

Model Files & Technical Specifications

| File | Size | Memory Footprint | BPW | Description |

| :--- | :---: | :---: | :---: | :--- |

| Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf | 13.76 GB | 12.82 GiB | 3.42 | High-fidelity mixed-precision Xing4.0 Abliterated GGUF |

  • Quantized source: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
  • Official model: XingChen-AGI/Xing4.0-29B-A4B
  • Parameters: 29B-class total / approximately 4B active per token (official specification)
  • Main stack: 40 layers = 2 dense + 38 MoE
  • MoE: 64 routed experts, 4 active per token, plus 1 shared expert on MoE layers
  • Attention / residual architecture: MLA + mHC
  • MTP: 1 additional prediction block (blk.40 in GGUF)
  • Context: 256K native; upstream supports extension to 512K

---

<a id="toc-04"></a>

Surgical Tensor Quantization Map (Audited from GGUF)

The table below was checked against the tensor directory of the published 13.76 GB GGUF, including the two dense layers and the MTP block:

| Component | Tensor / Layer Range | Precision | Purpose |

| :--- | :--- | :---: | :--- |

| Output head | output.weight | Q6_K | High-precision final token projection |

| Token embeddings | token_embd.weight | Q4_K | Vocabulary representation |

| Normalizations | RMS / attention / FFN norms | F32 | Numerical stability |

| Dense FFN entry layers | blk.0–1.ffn_down | Q6_K | Protects the two upstream-defined dense layers |

| Dense FFN entry layers | blk.0–1.ffn_gate/up | Q4_K | High-fidelity dense feature projections |

| Main MoE routers | blk.2–39.ffn_gate_inp | F32 | Preserves routed expert selection |

| Shared experts | blk.2–39.ffn_*_shexp | Q5_K | Always-active MoE foundation path |

| Boundary routed experts | blk.2–9 and blk.30–39 ffn_*_exps | Q3_K | Linear K-quant boundary protection and fast RAM/GPU execution |

| Core routed experts | blk.10–29.ffn_*_exps | IQ3_XXS | IMatrix-calibrated compact core |

| MLA latent projections | attn_q_a, attn_q_b, attn_kv_a_mqa, attn_v_b | Q4_K (31 layers) / Q5_K (10 anchor layers) | Protects MLA reconstruction geometry with higher precision at layers 3, 7, 11, ..., 39 |

| MLA K projection | attn_k_b | Q5_0 (31 layers) / Q5_1 (10 anchor layers) | Higher precision at the 10 periodic attention anchors |

| Attention output | attn_output | Q6_K | High-precision attention output projection |

| mHC transform matrices | hc_attn_fn, hc_ffn_fn | Q8_0 | Preserves Hyper-Connection transforms |

| mHC base / scale tensors | hc__base, hc__scale | F32 | Full-precision residual mixing controls |

| MTP experts / next-token weights | blk.40.ffn__{exps,shexp}, nextn. weight tensors | Q4_K | Protects the integrated prediction head; MTP normalization tensors remain F32 |

| MTP attention | blk.40 MLA projections / output | Q4_K / Q5_0 / Q6_K | Retains the same protected MLA precision pattern |

| MTP router / norms | blk.40.ffn_gate_inp and normalization tensors | F32 | Full-precision routing and normalization |

---

<a id="toc-05"></a>

Hardware Throughput: GPU Projections & Tested RAM Offload

Dedicated-GPU throughput projections

The following figures are estimates, not measured Xing4.0 benchmarks. They are deployment projections for the 13.76 GB GGUF after the weights fit in the target configuration; context length, KV format, batch size, backend, clocks, and llama.cpp build can materially change throughput.

| Hardware Target | Mode | Estimated Generation | Notes |

| :--- | :--- | :---: | :--- |

| RTX 5080 / 5090 | Full GPU where memory permits | 120–160+ tok/s | Blackwell projection; available VRAM determines context headroom |

| RTX 4090 / 3090 | Full GPU | 75–110+ tok/s | 24GB cards leave substantially more room for long context |

| RTX 4080 / 4070 Ti Super | Full GPU | 55–80+ tok/s | 16GB class; context headroom is configuration-dependent |

Tested hybrid / system-RAM behavior

Hybrid and system-RAM offload were tested during development. Depending on CPU, memory bandwidth, DDR4/DDR5 configuration, GPU offload level, and active context, observed generation behavior falls around 20–45 tok/s. Treat this as a hardware-dependent author-tested range, not a guarantee for every system.

---

<a id="toc-06"></a>

256K Native Context & Deployment Notes

Xing4.0's official configuration declares max_position_embeddings: 262144; the upstream model card documents 256K native context with extension to 512K.

The GGUF weights occupy 12.82 GiB, but maximum fully resident context is not determined by weight size alone. KV-cache type, Flash Attention, batch size, concurrent slots, compute buffers, display-driver reservation, and llama.cpp build all affect the final VRAM requirement. On 16GB cards, shorter contexts offer the safest full-offload margin; 24GB+ cards provide much more room for the native long-context window. The 512K mode should be treated as an upstream-supported extended context rather than an automatic fit guarantee on a single GPU.

---

<a id="toc-07"></a>

llama.cpp Quickstart

Xing4.0's upstream Jinja template uses Xing-specific role tokens such as <_system>, <_user>, <_bot>, and <_end>. Let llama.cpp apply the chat template stored in the model metadata instead of manually constructing role tokens.

Interactive llama-cli

llama-cli \
  -m Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf \
  -ngl 99 \
  -c 32768 \
  -cnv \
  --jinja \
  --temp 1.0 \
  --top-p 0.95 \
  --repeat-penalty 1.05

OpenAI-compatible llama-server

llama-server \
  -m Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf \
  --port 8080 \
  --host 0.0.0.0 \
  -ngl 99 \
  -c 65536 \
  --flash-attn on \
  --jinja \
  --cache-type-k q8_0 \
  --cache-type-v q8_0

llama.cpp documents that --jinja enables Jinja chat templating and that the template can be taken from the model metadata. Increase context only after accounting for your actual KV-cache and buffer usage.

---

<a id="toc-08"></a>

Official Xing4.0 Generation Parameters

The official XingChen model card recommends the following sampler settings:

| Scenario | Temperature | Top-P | Repetition Penalty |

| :--- | :---: | :---: | :---: |

| Complex reasoning / general tasks | 1.0 | 0.95 | 1.05 |

| Coding / agent tasks | 0.8 | 0.95 | 1.05 |

XingChen also reports using long contexts in its evaluations, including 210K for SWE-bench-style agent runs and 256K for several long-context evaluations.

---

<a id="toc-09"></a>

Optional Support

<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>

If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

Run IsValorum/Xing4.0-29B-A4B-APEX-I-MiniPlus-V2.1-Abliterated-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models