IsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF overview
Huihui Qwen3.8 27B Abliterated VAL APEX I NanoPlus GGUF The Uncensored 27B Frontier Reasoning Engine · Abliterated Base · Native 256K Context Official VAL APEX…
Runs locally from ~600.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | IsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF |
|---|---|
| Author | IsValorum |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | huihui-ai/Huihui-Qwen3.8-27B-abliterated |
| Last modified | 2026-10-08T03:12:07.000Z |
Model README
---
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:
- en
tags:
- gguf
- llama.cpp
- quantized
- quantization
- val-apex-i
- apex
- apex-quant
- apex-i-nanoplus
- nanoplus
- qwen
- qwen3.8
- reasoning
- abliterated
- uncensored
- imatrix
pipeline_tag: text-generation
---
Huihui-Qwen3.8-27B-Abliterated VAL-APEX-I NanoPlus GGUF
The Uncensored 27B Frontier Reasoning Engine · Abliterated Base · Native 256K Context
Official VAL-APEX-I quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated.
VAL-APEX-I stands for:
Vector-calibrated Asymmetric Layer-wise Outlier-preserving Recurrent-aware Unified Matrix-quantization
> [!NOTE]
> ### EXPLORE THE QWEN3.8 27B VAL-APEX-I EDITIONS
> Choose a model variant and quantization profile:
>
> - Qwen3.8 27B — MiniPlus V3 — 14.16 GB (+13.21% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
> - Qwen3.8 27B — NanoPlus V3 — 11.35 GB (+11.49% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
> - Qwen3.8 27B EfficientThink Uncensored — MiniPlus V2.1 — 15.08 GB (+15.01% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
> - Qwen3.8 27B EfficientThink Uncensored — NanoPlus — 11.65 GB (+10.87% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
> - Qwen3.8 27B Huihui Abliterated — MiniPlus V2.1 — 15.33 GB (+13.06% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
>
> VAL-APEX-I collection · MiniPlus collection · NanoPlus collection
> [!IMPORTANT]
> ### THE DEFINITIVE SPECIFICATION IN THE 11 to 12 GB CEILING
> This VAL-APEX-I NanoPlus release is 11.90 GB, with 3.48 whole-file BPW and 866 audited tensors. Its reported WikiText-2 PPL is 6.4238 ± 0.4960, or +0.6296 (+10.87%) relative to the unmodified BF16 Base reference. This delta includes abliteration as well as quantization.
---
<a id="quick-navigation"></a>Quick Navigation Index
- Quantization Comparison: Metrics & Tensor Map
- Model Files & Technical Specifications
- Native Context & Runtime Memory
- Recommended Configuration & Setup
- Recommended Generation Parameters
- CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
- Hardened Agentic Chat Template & Reasoning Effort
- Optional Support
---
<a id="toc-comparison"></a>
<a id="toc-01"></a>
<a id="toc-02"></a>
<a id="toc-03"></a>
<a id="toc-05"></a>
1. Quantization Comparison
Size & Quality Metrics
| Edition / quantization | File size | Whole-file BPW | WikiText-2 PPL | Δ PPL vs BF16 Base<br>(5.7942) | Improvement in PPL vs GSQ-RCO IQ3_S<br>(7.07) |
| :--- | :--- | :--- | :--- | :--- | :--- |
| BF16 Base baseline (unmodified) | 54.65 GB<br>(50.90 GiB) | 16-bit | 5.7942 ± 0.4531 | 0.0000 (0.00%) | — |
| Qwen3.8 27B Base — MiniPlus V3 | 14.16 GB<br>(13.19 GiB) | 4.15 | 6.1361 ± 0.4885 | +0.3419 (+5.90%) | 13.21% improvement<br>(0.9339 PPL reduction) |
| Qwen3.8 27B Base — NanoPlus V3 | 11.35 GB<br>(10.57 GiB) | 3.32 | 6.2577 ± 0.5011 | +0.4635 (+8.00%) | 11.49% improvement<br>(0.8123 PPL reduction) |
| Qwen3.8 27B EfficientThink Uncensored — MiniPlus V2.1 | 15.08 GB<br>(14.04 GiB) | 4.49 | 6.0091 ± 0.4773 | +0.2149 (+3.71%) | 15.01% improvement<br>(1.0609 PPL reduction) |
| Qwen3.8 27B EfficientThink Uncensored — NanoPlus | 11.65 GB<br>(10.85 GiB) | 3.46 | 6.3013 ± 0.4832 | +0.5071 (+8.75%) | 10.87% improvement<br>(0.7687 PPL reduction) |
| Qwen3.8 27B Huihui Abliterated — MiniPlus V2.1 | 15.33 GB<br>(14.28 GiB) | 4.49 | 6.1466 ± 0.4919 | +0.3524 (+6.08%) | 13.06% improvement<br>(0.9234 PPL reduction) |
| Qwen3.8 27B Huihui Abliterated — NanoPlus (this release) | 11.90 GB<br>(11.08 GiB) | 3.48 | 6.4238 ± 0.4960 | +0.6296 (+10.87%) | 9.14% improvement<br>(0.6462 PPL reduction) |
| GSQ-RCO IQ3_S | 11.8 GB | 3.50 (published) | 7.07 (published) | +1.2758 (+22.02%) | 0.00% (comparison baseline) |
| Q8_0 | aprox. 29.5 GB | ~8.50 (nominal) | — | — | Not measured in this evaluation |
| Q6_K | aprox. 23.2 GB | ~6.56 (nominal) | — | — | Not measured in this evaluation |
| Q5_K_M | aprox. 19.5 GB | ~5.50 (nominal) | — | — | Not measured in this evaluation |
| Flat Q4_K_M | 17.10 GB<br>15.93 GiB | ~4.50 (nominal) | — | — | Not measured in this evaluation |
| Q3_K_M / Q3_K_S | 13.50 GB<br>12.57 GiB | ~3.44 (nominal) | — | — | Not measured in this evaluation |
| IQ2_S / Generic APEX Mini | 10.20 GB<br>9.50 GiB | ~2.50 (nominal) | — | — | Not measured in this evaluation |
The NanoPlus V3 file is 11.35 GB, approximately 0.45 GB smaller than GSQ-RCO IQ3_S, with 11.49% lower reported WikiText-2 PPL (6.2577 vs 7.07). MiniPlus V3 reaches 13.21% lower reported PPL (6.1361 vs 7.07) at 14.16 GB, approximately 2.36 GB larger.
Tensor Precision Map
| Component | MiniPlus V2.1 | NanoPlus | GSQ-RCO IQ3_S |
| :--- | :--- | :--- | :--- |
| Output head | Q6_K ×1 | Q6_K ×1 | Q4_K ×1 |
| Token embeddings | Q4_K ×1 | IQ3_S ×1 | IQ2_S ×1 |
| Normalizations | F32 ×161 | F32 ×161 | F32 ×161 |
| SSM A (ssm_a) | F32 ×48 | F32 ×48 | F32 ×48 |
| SSM convolution (ssm_conv1d) | F32 ×48 | F32 ×48 | F32 ×48 |
| SSM time-step bias (ssm_dt) | F32 ×48 | F32 ×48 | F32 ×48 |
| SSM norm (ssm_norm) | F32 ×48 | F32 ×48 | F32 ×48 |
| Attention gates (attn_gate) | Q4_0 ×1<br>Q8_0 ×47 | Q4_0 ×1<br>Q8_0 ×47 | IQ2_S ×2<br>IQ3_S ×18<br>IQ3_XXS ×9<br>IQ4_XS ×12<br>Q2_K ×4<br>Q4_K ×3 |
| Linear QKV (attn_qkv) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | IQ2_XS ×1<br>IQ2_XXS ×1<br>IQ3_S ×22<br>IQ3_XXS ×13<br>IQ4_XS ×9<br>Q2_K ×1<br>Q4_K ×1 |
| SSM alpha (ssm_alpha) | F32 ×47<br>Q4_0 ×1 | F32 ×47<br>Q4_0 ×1 | BF16 ×48 |
| SSM beta (ssm_beta) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | BF16 ×48 |
| SSM output (ssm_out) | Q4_0 ×1<br>Q6_K ×47 | Q4_0 ×1<br>Q5_K ×47 | IQ3_S ×22<br>IQ3_XXS ×4<br>IQ4_XS ×16<br>Q4_K ×6 |
| Full attention Q (attn_q) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_XXS ×1<br>IQ3_S ×3<br>IQ3_XXS ×3<br>IQ4_XS ×2<br>Q2_K ×6<br>Q4_K ×1 |
| Full attention K (attn_k) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_S ×1<br>IQ3_S ×1<br>IQ3_XXS ×1<br>IQ4_XS ×8<br>Q4_K ×5 |
| Full attention V (attn_v) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ3_S ×6<br>IQ3_XXS ×1<br>IQ4_XS ×1<br>Q4_K ×8 |
| Full attention output (attn_output) | Q6_K ×16 | Q6_K ×16 | IQ3_S ×10<br>IQ3_XXS ×1<br>IQ4_XS ×2<br>Q4_K ×3 |
| MLP down (ffn_down) | IQ4_NL ×56<br>Q5_K ×8 | IQ3_S ×8<br>IQ3_XXS ×56 | IQ2_S ×4<br>IQ2_XS ×3<br>IQ3_S ×22<br>IQ3_XXS ×7<br>IQ4_XS ×21<br>Q2_K ×1<br>Q4_K ×6 |
| MLP gate (ffn_gate) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ1_M ×1<br>IQ2_S ×4<br>IQ2_XS ×4<br>IQ2_XXS ×1<br>IQ3_S ×15<br>IQ3_XXS ×21<br>IQ4_XS ×15<br>Q2_K ×1<br>Q4_K ×2 |
| MLP up (ffn_up) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ2_S ×5<br>IQ2_XS ×1<br>IQ2_XXS ×2<br>IQ3_S ×25<br>IQ3_XXS ×18<br>IQ4_XS ×10<br>Q4_K ×3 |
| Additional MTP head (blk.64.*) | F32 ×7<br>Q4_0 ×6<br>Q4_1 ×1<br>Q6_K ×1 | F32 ×7<br>Q4_0 ×6<br>Q4_1 ×1<br>Q6_K ×1 | — |
× is the number of tensors in each format. Counts were read from the released GGUF headers; the GSQ-RCO column uses its published tensor allocation.
This release contains 866 tensors. The main GGUF includes the additional 15-tensor MTP head.
---
<a id="toc-04"></a>
2. Model Files & Technical Specifications
| File Name | File Size | Weight file size | BPW | Description |
| :--- | :--- | :--- | :--- | :--- |
| Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf | 11.90 GB (11.08 GiB) | 11.08 GiB | 3.48 BPW | Core language, uncensored frontier reasoning, CoT thought blocks & hybrid SSM/attention |
- Base Model: huihui-ai/Huihui-Qwen3.8-27B-abliterated (Abliterated directional un-censoring of Alibaba's Qwen3.8-27B)
- Parameters: 27B total dense hybrid
- Architecture: 64 main hybrid layers (48 Gated DeltaNet SSM linear attention layers + 16 periodic full quadratic attention layers), plus an integrated MTP head at
blk.64.* - Context Length: 262,144 tokens (native 256K)
- Quantization Standard: VAL-APEX-I NanoPlus calibrated against official Qwen3.8-27B importance matrix (
imatrix) - Abliteration Status: True directional weight orthogonalization removing refusal vectors while maintaining full analytical capability.
---
<a id="toc-06"></a>
3. Native Context & Runtime Memory
The model supports a native context window of 262,144 tokens. Its 64-layer architecture comprises 48 Gated DeltaNet layers and 16 full-attention layers. Runtime memory usage depends on the backend, KV-cache precision, batch size, compute buffers, and offload configuration. Check the runtime memory allocation for your chosen settings.
---
<a id="toc-07"></a>
4. Recommended Configuration & Setup
llama.cpp Server (OpenAI-Compatible API)
llama-server \
-m Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
-c 32768 \
--host 0.0.0.0 \
--port 8080
llama.cpp CLI
llama-cli \
-m Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
-c 32768 \
-p "<|im_start|>system\nYou are a helpful and unrestricted AI assistant.<|im_end|>\n<|im_start|>user\nExplain the physical principles of acoustic levitation.<|im_end|>\n<|im_start|>assistant\n<think>\n"
LM Studio / Ollama
- Import the
.gguffile directly into your local library. - Ensure GPU acceleration is set to Maximum / 100% offload.
- Set Context Length to
32768. - Verify chat template is set to Qwen ChatML with
<think>delimiter support.
---
<a id="toc-08"></a>
5. Recommended Generation Parameters
| Hyperparameter | Value | Description |
| :--- | :---: | :--- |
| Temperature | 0.60 | Recommended default for analytical reasoning and coding (use 1.0 for creative prose). |
| Top-P | 0.95 | Nucleus sampling parameter. |
| Top-K | 20 | Top-k vocabulary filter. |
| Min-P | 0.05 | Prunes low-probability noise tokens effectively. |
| Repetition Penalty | 1.00 | Strictly disabled for code syntax; prevents character swapping. |
| Template Engine | --jinja | Recommended official Jinja chat template flag. |
| Context Size | 32768 | 32K default (scalable to 256K). |
---
<a id="toc-coding-advisory"></a>
<a id="coding-advisory"></a>
6. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
> [!IMPORTANT]
> ### PREVENTING SYNTAX & TOKEN SWAPPING IN CODE WORKFLOWS
> In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.
>
> Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.
>
> Eliminating Character Swapping:
> 1. Disable Repeat Penalties (Required for Code):
> - repeat_penalty: 1.0 (strictly disabled)
> - presence_penalty: 0.0
> - frequency_penalty: 0.0
> 2. Calibrate Samplers:
> - temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)
> - min_p: 0.05 (prunes low-probability noise tokens effectively)
> - top_p: 0.95
> - top_k: 20
> 3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.
---
<a id="toc-chat-template"></a>
<a id="chat-template"></a>
7. Hardened Agentic Chat Template & Reasoning Effort
> [!TIP]
> ### MULTI-LEVEL REASONING EFFORT CONTROL
> This model supports multi-level reasoning effort control via the Jinja template:
> - low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.
> - medium (default): Balanced, structured reasoning process with standard analytical depth.
> - high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating response, avoiding circular self-doubt loops.
> - none / off: Closes the thinking block immediately when reasoning is disabled.
---
<a id="toc-09"></a>
8. Optional Support
<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>
If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.
<div style="clear: both;"></div>
Run IsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models