GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

IsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF overview

Huihui Qwen3.8 27B Abliterated VAL APEX I NanoPlus GGUF The Uncensored 27B Frontier Reasoning Engine · Abliterated Base · Native 256K Context Official VAL APEX…

ggufllama.cppquantizedquantizationval-apex-iapexapex-quantapex-i-nanoplusnanoplusqwenqwen3.8reasoningabliterateduncensoredimatrixtext-generationconversationalenbase_model:huihui-ai/Huihui-Qwen3.8-27B-abliteratedbase_model:quantized:huihui-ai/Huihui-Qwen3.8-27B-abliteratedlicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~600.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
473
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.ggufGGUFGGUF11.08 GBDownload
mmproj-Qwen3.8-27B-Q8_0.ggufGGUFQ8_0600.1 MBDownload

Model Details

Model IDIsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF
AuthorIsValorum
Pipelinetext-generation
Licenseapache-2.0
Base modelhuihui-ai/Huihui-Qwen3.8-27B-abliterated
Last modified2026-10-08T03:12:07.000Z

Model README

---

base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated

base_model_relation: quantized

quantized_by: IsValorum

library_name: gguf

license: apache-2.0

language:

  • en

tags:

  • gguf
  • llama.cpp
  • quantized
  • quantization
  • val-apex-i
  • apex
  • apex-quant
  • apex-i-nanoplus
  • nanoplus
  • qwen
  • qwen3.8
  • reasoning
  • abliterated
  • uncensored
  • imatrix

pipeline_tag: text-generation

---

Huihui-Qwen3.8-27B-Abliterated VAL-APEX-I NanoPlus GGUF

The Uncensored 27B Frontier Reasoning Engine · Abliterated Base · Native 256K Context

Official VAL-APEX-I quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated.

VAL-APEX-I stands for:

Vector-calibrated Asymmetric Layer-wise Outlier-preserving Recurrent-aware Unified Matrix-quantization

> [!NOTE]

> ### EXPLORE THE QWEN3.8 27B VAL-APEX-I EDITIONS

> Choose a model variant and quantization profile:

>

> - Qwen3.8 27B — MiniPlus V3 — 14.16 GB (+13.21% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).

> - Qwen3.8 27B — NanoPlus V3 — 11.35 GB (+11.49% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).

> - Qwen3.8 27B EfficientThink Uncensored — MiniPlus V2.1 — 15.08 GB (+15.01% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).

> - Qwen3.8 27B EfficientThink Uncensored — NanoPlus — 11.65 GB (+10.87% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).

> - Qwen3.8 27B Huihui Abliterated — MiniPlus V2.1 — 15.33 GB (+13.06% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).

>

> VAL-APEX-I collection · MiniPlus collection · NanoPlus collection

> [!IMPORTANT]

> ### THE DEFINITIVE SPECIFICATION IN THE 11 to 12 GB CEILING

> This VAL-APEX-I NanoPlus release is 11.90 GB, with 3.48 whole-file BPW and 866 audited tensors. Its reported WikiText-2 PPL is 6.4238 ± 0.4960, or +0.6296 (+10.87%) relative to the unmodified BF16 Base reference. This delta includes abliteration as well as quantization.

---

<a id="quick-navigation"></a>Quick Navigation Index

  1. Quantization Comparison: Metrics & Tensor Map
  2. Model Files & Technical Specifications
  3. Native Context & Runtime Memory
  4. Recommended Configuration & Setup
  5. Recommended Generation Parameters
  6. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  7. Hardened Agentic Chat Template & Reasoning Effort
  8. Optional Support

---

<a id="toc-comparison"></a>

<a id="toc-01"></a>

<a id="toc-02"></a>

<a id="toc-03"></a>

<a id="toc-05"></a>

1. Quantization Comparison

Size & Quality Metrics

| Edition / quantization | File size | Whole-file BPW | WikiText-2 PPL | Δ PPL vs BF16 Base<br>(5.7942) | Improvement in PPL vs GSQ-RCO IQ3_S<br>(7.07) |

| :--- | :--- | :--- | :--- | :--- | :--- |

| BF16 Base baseline (unmodified) | 54.65 GB<br>(50.90 GiB) | 16-bit | 5.7942 ± 0.4531 | 0.0000 (0.00%) | — |

| Qwen3.8 27B Base — MiniPlus V3 | 14.16 GB<br>(13.19 GiB) | 4.15 | 6.1361 ± 0.4885 | +0.3419 (+5.90%) | 13.21% improvement<br>(0.9339 PPL reduction) |

| Qwen3.8 27B Base — NanoPlus V3 | 11.35 GB<br>(10.57 GiB) | 3.32 | 6.2577 ± 0.5011 | +0.4635 (+8.00%) | 11.49% improvement<br>(0.8123 PPL reduction) |

| Qwen3.8 27B EfficientThink Uncensored — MiniPlus V2.1 | 15.08 GB<br>(14.04 GiB) | 4.49 | 6.0091 ± 0.4773 | +0.2149 (+3.71%) | 15.01% improvement<br>(1.0609 PPL reduction) |

| Qwen3.8 27B EfficientThink Uncensored — NanoPlus | 11.65 GB<br>(10.85 GiB) | 3.46 | 6.3013 ± 0.4832 | +0.5071 (+8.75%) | 10.87% improvement<br>(0.7687 PPL reduction) |

| Qwen3.8 27B Huihui Abliterated — MiniPlus V2.1 | 15.33 GB<br>(14.28 GiB) | 4.49 | 6.1466 ± 0.4919 | +0.3524 (+6.08%) | 13.06% improvement<br>(0.9234 PPL reduction) |

| Qwen3.8 27B Huihui Abliterated — NanoPlus (this release) | 11.90 GB<br>(11.08 GiB) | 3.48 | 6.4238 ± 0.4960 | +0.6296 (+10.87%) | 9.14% improvement<br>(0.6462 PPL reduction) |

| GSQ-RCO IQ3_S | 11.8 GB | 3.50 (published) | 7.07 (published) | +1.2758 (+22.02%) | 0.00% (comparison baseline) |

| Q8_0 | aprox. 29.5 GB | ~8.50 (nominal) | — | — | Not measured in this evaluation |

| Q6_K | aprox. 23.2 GB | ~6.56 (nominal) | — | — | Not measured in this evaluation |

| Q5_K_M | aprox. 19.5 GB | ~5.50 (nominal) | — | — | Not measured in this evaluation |

| Flat Q4_K_M | 17.10 GB<br>15.93 GiB | ~4.50 (nominal) | — | — | Not measured in this evaluation |

| Q3_K_M / Q3_K_S | 13.50 GB<br>12.57 GiB | ~3.44 (nominal) | — | — | Not measured in this evaluation |

| IQ2_S / Generic APEX Mini | 10.20 GB<br>9.50 GiB | ~2.50 (nominal) | — | — | Not measured in this evaluation |

The NanoPlus V3 file is 11.35 GB, approximately 0.45 GB smaller than GSQ-RCO IQ3_S, with 11.49% lower reported WikiText-2 PPL (6.2577 vs 7.07). MiniPlus V3 reaches 13.21% lower reported PPL (6.1361 vs 7.07) at 14.16 GB, approximately 2.36 GB larger.

Tensor Precision Map

| Component | MiniPlus V2.1 | NanoPlus | GSQ-RCO IQ3_S |

| :--- | :--- | :--- | :--- |

| Output head | Q6_K ×1 | Q6_K ×1 | Q4_K ×1 |

| Token embeddings | Q4_K ×1 | IQ3_S ×1 | IQ2_S ×1 |

| Normalizations | F32 ×161 | F32 ×161 | F32 ×161 |

| SSM A (ssm_a) | F32 ×48 | F32 ×48 | F32 ×48 |

| SSM convolution (ssm_conv1d) | F32 ×48 | F32 ×48 | F32 ×48 |

| SSM time-step bias (ssm_dt) | F32 ×48 | F32 ×48 | F32 ×48 |

| SSM norm (ssm_norm) | F32 ×48 | F32 ×48 | F32 ×48 |

| Attention gates (attn_gate) | Q4_0 ×1<br>Q8_0 ×47 | Q4_0 ×1<br>Q8_0 ×47 | IQ2_S ×2<br>IQ3_S ×18<br>IQ3_XXS ×9<br>IQ4_XS ×12<br>Q2_K ×4<br>Q4_K ×3 |

| Linear QKV (attn_qkv) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | IQ2_XS ×1<br>IQ2_XXS ×1<br>IQ3_S ×22<br>IQ3_XXS ×13<br>IQ4_XS ×9<br>Q2_K ×1<br>Q4_K ×1 |

| SSM alpha (ssm_alpha) | F32 ×47<br>Q4_0 ×1 | F32 ×47<br>Q4_0 ×1 | BF16 ×48 |

| SSM beta (ssm_beta) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | BF16 ×48 |

| SSM output (ssm_out) | Q4_0 ×1<br>Q6_K ×47 | Q4_0 ×1<br>Q5_K ×47 | IQ3_S ×22<br>IQ3_XXS ×4<br>IQ4_XS ×16<br>Q4_K ×6 |

| Full attention Q (attn_q) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_XXS ×1<br>IQ3_S ×3<br>IQ3_XXS ×3<br>IQ4_XS ×2<br>Q2_K ×6<br>Q4_K ×1 |

| Full attention K (attn_k) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_S ×1<br>IQ3_S ×1<br>IQ3_XXS ×1<br>IQ4_XS ×8<br>Q4_K ×5 |

| Full attention V (attn_v) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ3_S ×6<br>IQ3_XXS ×1<br>IQ4_XS ×1<br>Q4_K ×8 |

| Full attention output (attn_output) | Q6_K ×16 | Q6_K ×16 | IQ3_S ×10<br>IQ3_XXS ×1<br>IQ4_XS ×2<br>Q4_K ×3 |

| MLP down (ffn_down) | IQ4_NL ×56<br>Q5_K ×8 | IQ3_S ×8<br>IQ3_XXS ×56 | IQ2_S ×4<br>IQ2_XS ×3<br>IQ3_S ×22<br>IQ3_XXS ×7<br>IQ4_XS ×21<br>Q2_K ×1<br>Q4_K ×6 |

| MLP gate (ffn_gate) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ1_M ×1<br>IQ2_S ×4<br>IQ2_XS ×4<br>IQ2_XXS ×1<br>IQ3_S ×15<br>IQ3_XXS ×21<br>IQ4_XS ×15<br>Q2_K ×1<br>Q4_K ×2 |

| MLP up (ffn_up) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ2_S ×5<br>IQ2_XS ×1<br>IQ2_XXS ×2<br>IQ3_S ×25<br>IQ3_XXS ×18<br>IQ4_XS ×10<br>Q4_K ×3 |

| Additional MTP head (blk.64.*) | F32 ×7<br>Q4_0 ×6<br>Q4_1 ×1<br>Q6_K ×1 | F32 ×7<br>Q4_0 ×6<br>Q4_1 ×1<br>Q6_K ×1 | — |

× is the number of tensors in each format. Counts were read from the released GGUF headers; the GSQ-RCO column uses its published tensor allocation.

This release contains 866 tensors. The main GGUF includes the additional 15-tensor MTP head.

---

<a id="toc-04"></a>

2. Model Files & Technical Specifications

| File Name | File Size | Weight file size | BPW | Description |

| :--- | :--- | :--- | :--- | :--- |

| Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf | 11.90 GB (11.08 GiB) | 11.08 GiB | 3.48 BPW | Core language, uncensored frontier reasoning, CoT thought blocks & hybrid SSM/attention |

  • Base Model: huihui-ai/Huihui-Qwen3.8-27B-abliterated (Abliterated directional un-censoring of Alibaba's Qwen3.8-27B)
  • Parameters: 27B total dense hybrid
  • Architecture: 64 main hybrid layers (48 Gated DeltaNet SSM linear attention layers + 16 periodic full quadratic attention layers), plus an integrated MTP head at blk.64.*
  • Context Length: 262,144 tokens (native 256K)
  • Quantization Standard: VAL-APEX-I NanoPlus calibrated against official Qwen3.8-27B importance matrix (imatrix)
  • Abliteration Status: True directional weight orthogonalization removing refusal vectors while maintaining full analytical capability.

---

<a id="toc-06"></a>

3. Native Context & Runtime Memory

The model supports a native context window of 262,144 tokens. Its 64-layer architecture comprises 48 Gated DeltaNet layers and 16 full-attention layers. Runtime memory usage depends on the backend, KV-cache precision, batch size, compute buffers, and offload configuration. Check the runtime memory allocation for your chosen settings.

---

<a id="toc-07"></a>

4. Recommended Configuration & Setup

llama.cpp Server (OpenAI-Compatible API)

llama-server \
  -m Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  --host 0.0.0.0 \
  --port 8080

llama.cpp CLI

llama-cli \
  -m Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  -p "<|im_start|>system\nYou are a helpful and unrestricted AI assistant.<|im_end|>\n<|im_start|>user\nExplain the physical principles of acoustic levitation.<|im_end|>\n<|im_start|>assistant\n<think>\n"

LM Studio / Ollama

  1. Import the .gguf file directly into your local library.
  2. Ensure GPU acceleration is set to Maximum / 100% offload.
  3. Set Context Length to 32768.
  4. Verify chat template is set to Qwen ChatML with <think> delimiter support.

---

<a id="toc-08"></a>

5. Recommended Generation Parameters

| Hyperparameter | Value | Description |

| :--- | :---: | :--- |

| Temperature | 0.60 | Recommended default for analytical reasoning and coding (use 1.0 for creative prose). |

| Top-P | 0.95 | Nucleus sampling parameter. |

| Top-K | 20 | Top-k vocabulary filter. |

| Min-P | 0.05 | Prunes low-probability noise tokens effectively. |

| Repetition Penalty | 1.00 | Strictly disabled for code syntax; prevents character swapping. |

| Template Engine | --jinja | Recommended official Jinja chat template flag. |

| Context Size | 32768 | 32K default (scalable to 256K). |

---

<a id="toc-coding-advisory"></a>

<a id="coding-advisory"></a>

6. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)

> [!IMPORTANT]

> ### PREVENTING SYNTAX & TOKEN SWAPPING IN CODE WORKFLOWS

> In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

>

> Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.

>

> Eliminating Character Swapping:

> 1. Disable Repeat Penalties (Required for Code):

> - repeat_penalty: 1.0 (strictly disabled)

> - presence_penalty: 0.0

> - frequency_penalty: 0.0

> 2. Calibrate Samplers:

> - temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)

> - min_p: 0.05 (prunes low-probability noise tokens effectively)

> - top_p: 0.95

> - top_k: 20

> 3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.

---

<a id="toc-chat-template"></a>

<a id="chat-template"></a>

7. Hardened Agentic Chat Template & Reasoning Effort

> [!TIP]

> ### MULTI-LEVEL REASONING EFFORT CONTROL

> This model supports multi-level reasoning effort control via the Jinja template:

> - low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.

> - medium (default): Balanced, structured reasoning process with standard analytical depth.

> - high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating response, avoiding circular self-doubt loops.

> - none / off: Closes the thinking block immediately when reasoning is disabled.

---

<a id="toc-09"></a>

8. Optional Support

<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>

If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

<div style="clear: both;"></div>

Run IsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models