IsValorum/Xing4.0-29B-A4B-APEX-I-NanoPlus-Abliterated-GGUF overview
<a id="quick navigation" </a Quick Navigation Index 1. Xing4.0 Architecture & Abliterated Source toc 01 2. Measured Fidelity & NanoPlus / MiniPlus Comparison t…
Runs locally from ~10.94 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf | GGUF | GGUF | 10.94 GB | Download |
Model Details
| Model ID | IsValorum/Xing4.0-29B-A4B-APEX-I-NanoPlus-Abliterated-GGUF |
|---|---|
| Author | IsValorum |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated |
| Last modified | 2026-10-03T17:42:40.000Z |
Model README
---
base_model: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:
- en
- zh
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- quantized
- quantization
- apex
- apex-quant
- apex-i-nanoplus
- nanoplus
- custom-quantization
- imatrix
- abliterated
- uncensored
- moe
- mla
- hyper-connections
- mtp
- reasoning
- coding
- agentic
- long-context
- xing4_0
---
<a id="quick-navigation"></a>Quick Navigation Index
- Xing4.0 Architecture & Abliterated Source
- Measured Fidelity & NanoPlus / MiniPlus Comparison
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Hardware Throughput: GPU Projections & Tested RAM Offload
- 256K Native Context & Deployment Notes
- llama.cpp Quickstart
- Official Xing4.0 Generation Parameters
- Optional Support
Xing4.0-29B-A4B APEX-I-NanoPlus Abliterated GGUF
29B / 4B Active MoE · mHC + MLA + MTP · 256K Native Context · 11.74 GB
> [!IMPORTANT]
> ### APEX-I-NANOPLUS — ULTRA-COMPACT 11–12 GB RELEASE
> This APEX-I-NanoPlus build uses a Xing4.0-specific tensor-by-tensor allocation to compress the model to 11.74 GB (10.94 GiB / 2.92 BPW) while retaining measured Q4_K_M / Q4_K_S-class practical quality. It is intended for configurations where an extra ~2 GB of headroom materially improves GPU offload, context capacity, or hybrid RAM/GPU execution.
> [!NOTE]
> ### ABLITERATED / UNCENSORED SOURCE
> This GGUF quantizes huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated. Huihui describes that checkpoint as an uncensored version of the official Xing4.0 model created with an abliteration workflow and references Sumandora/remove-refusals-with-transformers for the method. No additional refusal-removal procedure is performed by this GGUF quantization.
> [!TIP]
> ### MEASURED FIDELITY
>
> | Quantization | File Size | Memory Footprint | Average BPW | WikiText-2 Perplexity | Delta vs. BF16 | Quality Target |
> | :--- | :---: | :---: | :---: | :---: | :---: | :---: |
> | BF16 source | 62.40 GB | 58.11 GiB | 16.00 | 7.3060 +/- 0.18633 | Baseline | Full-precision reference |
> | APEX-I-MiniPlus V2.1 | 13.76 GB | 12.82 GiB | 3.42 | 7.7238 +/- 0.19671 | +0.4178 (+5.71%) | Q5_K_M-class quality |
> | APEX-I-NanoPlus (CURRENT) | 11.74 GB | 10.94 GiB | 2.92 | 8.3197 +/- 0.21528 | +1.0137 (+13.87%) | Q4_K_M / Q4_K_S-class quality |
>
> The PPL values above are direct measurements from the final GGUFs. The MiniPlus/NanoPlus quality labels describe the intended practical fidelity tier; they are not derived from generic flat-quant PPL tables.
---
<a id="toc-01"></a>
Xing4.0 Architecture & Abliterated Source
The official XingChen-AGI/Xing4.0-29B-A4B is a 29B-parameter Mixture-of-Experts model with approximately 4B parameters activated per token. The upstream architecture combines mHC (manifold-constrained Hyper-Connections), Multi-head Latent Attention (MLA), and one Multi-Token Prediction (MTP) layer, and is positioned for reasoning, coding, agentic planning, and tool use. XingChen reports a 256K native context window, extensible to 512K, and training on the Ascend NPU / MindSpore stack.
The architecture is not a uniform 40-layer MoE:
- 40 main transformer layers in total.
- Layers 0–1 are dense FFN layers (
first_k_dense_replace: 2). - Layers 2–39 are MoE layers with 64 routed experts, 4 selected per token, plus one always-active shared expert.
- One additional MTP block is stored as
blk.40in the GGUF. - Native context: 262,144 tokens (256K); upstream documents extension to 512K.
NanoPlus concentrates precision on the residual down-projections, the two dense entry layers, MLA, mHC, routers, and output path while pushing selected routed gate/up projections lower to reach the 11.74 GB target.
---
<a id="toc-02"></a>
Measured Fidelity & NanoPlus / MiniPlus Comparison
Both APEX-I editions were evaluated directly from their compiled GGUF binaries with llama-perplexity on WikiText-2 using the same 2048-context / 512-batch / 10-chunk setup:
| Metric | BF16 Source | NanoPlus | MiniPlus V2.1 |
| :--- | :---: | :---: | :---: |
| WikiText-2 PPL | 7.3060 +/- 0.18633 | 8.3197 +/- 0.21528 | 7.7238 +/- 0.19671 |
| Delta PPL | 0 | +1.0137 (+13.87%) | +0.4178 (+5.71%) |
| GGUF size | 62.40 GB BF16 source | 11.74 GB | 13.76 GB |
| Average BPW | 16.00 | 2.92 | 3.42 |
| Practical quality target | Full precision | Q4_K_M / Q4_K_S class | Q5_K_M class |
Choose NanoPlus when memory efficiency is the priority. Choose MiniPlus V2.1 when the additional ~2 GB is available and maximum retained fidelity is preferred.
---
<a id="toc-03"></a>
Model Files & Technical Specifications
| File | Size | Memory Footprint | BPW | Description |
| :--- | :---: | :---: | :---: | :--- |
| Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf | 11.74 GB | 10.94 GiB | 2.92 | Ultra-compact mixed-precision Xing4.0 Abliterated GGUF |
- Quantized source: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
- Official model: XingChen-AGI/Xing4.0-29B-A4B
- Parameters: 29B-class total / approximately 4B active per token (official specification)
- Main stack: 40 layers = 2 dense + 38 MoE
- MoE: 64 routed experts, 4 active per token, plus 1 shared expert on MoE layers
- Attention / residual architecture: MLA + mHC
- MTP: 1 additional prediction block (
blk.40in GGUF) - Context: 256K native; upstream supports extension to 512K
---
<a id="toc-04"></a>
Surgical Tensor Quantization Map (Audited from GGUF)
The table below was checked against the tensor directory of the published 11.74 GB GGUF. It includes the two dense layers and the layer-range changes that are specific to this NanoPlus build:
| Component | Tensor / Layer Range | Precision | Purpose |
| :--- | :--- | :---: | :--- |
| Output head | output.weight | Q6_K | High-precision final token projection |
| Token embeddings | token_embd.weight | Q4_K | Vocabulary representation |
| Normalizations | RMS / attention / FFN norms | F32 | Numerical stability |
| Dense FFN entry layers | blk.0–1.ffn_down | Q6_K | Protects the two upstream-defined dense layers |
| Dense FFN entry layers | blk.0–1.ffn_gate/up | Q4_K | High-fidelity dense feature projections |
| Main MoE routers | blk.2–39.ffn_gate_inp | F32 | Preserves routed expert selection |
| Shared experts | blk.2–39.ffn_*_shexp | Q4_K | Always-active MoE foundation path |
| Routed down-projections | blk.2–37.ffn_down_exps | IQ3_XXS | Protects the residual stream at calibrated 3-bit precision |
| Routed gate projections | blk.2–37.ffn_gate_exps | IQ2_S | IMatrix-guided compact gating path |
| Early/intermediate up-projections | blk.2–15.ffn_up_exps | IQ2_S | Extra protection in earlier MoE layers |
| Deep up-projections | blk.16–37.ffn_up_exps | IQ2_XXS | Aggressive compression concentrated in the deep up path |
| Final MoE down-projections | blk.38–39.ffn_down_exps | Q3_K | Linear boundary protection |
| Final MoE gate/up projections | blk.38–39.ffn_gate/up_exps | IQ3_XXS | Higher-precision output-side expert path |
| MLA latent projections | attn_q_a, attn_q_b, attn_kv_a_mqa, attn_v_b | Q4_K | Protects MLA reconstruction geometry |
| MLA K projection | attn_k_b | Q5_0 | Additional precision for latent-key reconstruction |
| Attention output | attn_output | Q6_K | High-precision attention output projection |
| mHC transform matrices | hc_attn_fn, hc_ffn_fn | Q8_0 | Preserves Hyper-Connection transforms |
| mHC base / scale tensors | hc__base, hc__scale | F32 | Full-precision residual mixing controls |
| MTP experts / next-token weights | blk.40.ffn__{exps,shexp}, nextn. weight tensors | Q4_K | Protects the integrated prediction head; MTP normalization tensors remain F32 |
| MTP attention | blk.40 MLA projections / output | Q4_K / Q5_0 / Q6_K | Retains the same protected MLA precision pattern |
| MTP router / norms | blk.40.ffn_gate_inp and normalization tensors | F32 | Full-precision routing and normalization |
---
<a id="toc-05"></a>
Hardware Throughput: GPU Projections & Tested RAM Offload
Dedicated-GPU throughput projections
The following figures are estimates, not measured Xing4.0 benchmarks. They are deployment projections for the 11.74 GB GGUF after the weights fit in the target configuration; context length, KV format, batch size, backend, clocks, and llama.cpp build can materially change throughput.
| Hardware Target | Mode | Estimated Generation | Notes |
| :--- | :--- | :---: | :--- |
| RTX 5080 / 5090 | Full GPU where memory permits | 130–175+ tok/s | Blackwell projection; available VRAM determines context headroom |
| RTX 4090 / 3090 | Full GPU | 85–120+ tok/s | 24GB cards leave substantial room for long context |
| RTX 4080 / 4070 Ti Super | Full GPU | 65–95+ tok/s | 16GB class benefits directly from the smaller weight footprint |
Tested hybrid / system-RAM behavior
Hybrid and system-RAM offload were tested during development. Depending on CPU, memory bandwidth, DDR4/DDR5 configuration, GPU offload level, and active context, observed generation behavior falls around 25–50 tok/s. Treat this as a hardware-dependent author-tested range, not a guarantee for every system.
---
<a id="toc-06"></a>
256K Native Context & Deployment Notes
Xing4.0's official configuration declares max_position_embeddings: 262144; the upstream model card documents 256K native context with extension to 512K.
NanoPlus weights occupy 10.94 GiB, leaving roughly 1.88 GiB more headroom than MiniPlus before KV cache and compute buffers are allocated. Actual fully resident context still depends on KV-cache type, Flash Attention, batch size, concurrent slots, compute buffers, display-driver reservation, and llama.cpp build. The 512K mode should be treated as an upstream-supported extended context rather than an automatic fit guarantee on a single GPU.
---
<a id="toc-07"></a>
llama.cpp Quickstart
Xing4.0's upstream Jinja template uses Xing-specific role tokens such as <_system>, <_user>, <_bot>, and <_end>. Let llama.cpp apply the chat template stored in the model metadata instead of manually constructing role tokens.
Interactive llama-cli
llama-cli \
-m Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf \
-ngl 99 \
-c 32768 \
-cnv \
--jinja \
--temp 1.0 \
--top-p 0.95 \
--repeat-penalty 1.05
OpenAI-compatible llama-server
llama-server \
-m Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf \
--port 8080 \
--host 0.0.0.0 \
-ngl 99 \
-c 65536 \
--flash-attn on \
--jinja \
--cache-type-k q8_0 \
--cache-type-v q8_0
llama.cpp documents that --jinja enables Jinja chat templating and that the template can be taken from the model metadata. Increase context only after accounting for your actual KV-cache and buffer usage.
---
<a id="toc-08"></a>
Official Xing4.0 Generation Parameters
The official XingChen model card recommends the following sampler settings:
| Scenario | Temperature | Top-P | Repetition Penalty |
| :--- | :---: | :---: | :---: |
| Complex reasoning / general tasks | 1.0 | 0.95 | 1.05 |
| Coding / agent tasks | 0.8 | 0.95 | 1.05 |
XingChen also reports using long contexts in its evaluations, including 210K for SWE-bench-style agent runs and 256K for several long-context evaluations.
---
<a id="toc-09"></a>
Optional Support
<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>
If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.
Run IsValorum/Xing4.0-29B-A4B-APEX-I-NanoPlus-Abliterated-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models