IsValorum/Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus-GGUF overview
Qwen3.8 27B EfficientThink Uncensored VAL APEX I NanoPlus GGUF The SimPO + DFlash2 Frontier Reasoning Specialist · Synthetic SFT Opus 5 / Grok 4.6 / GPT 5.6 So…
Runs locally from ~498.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | IsValorum/Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus-GGUF |
|---|---|
| Author | IsValorum |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2 |
| Last modified | 2026-10-07T02:58:14.000Z |
Model README
---
base_model: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:
- en
tags:
- gguf
- llama.cpp
- quantized
- quantization
- val-apex-i
- apex
- apex-quant
- apex-i-nanoplus
- nanoplus
- qwen
- qwen3.8
- reasoning
- chain-of-thought
- dflash2
- speculative-decoding
- simpo
- uncensored
- imatrix
pipeline_tag: text-generation
---
Qwen3.8-27B-EfficientThink-Uncensored VAL-APEX-I NanoPlus GGUF
The SimPO + DFlash2 Frontier Reasoning Specialist · Synthetic SFT (Opus 5 / Grok 4.6 / GPT 5.6 Sol) · Native 256K Context
Official VAL-APEX-I quantization of nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2.
VAL-APEX-I stands for:
Vector-calibrated Asymmetric Layer-wise Outlier-preserving Recurrent-aware Unified Matrix-quantization
> [!NOTE]
> ### EXPLORE THE OTHER QWEN3.8 27B VAL-APEX-I EDITIONS
> Choose another base variant or fidelity profile:
>
> - Qwen3.8 27B Huihui Abliterated — VAL-APEX-I MiniPlus V2.1 — 15.33 GB.
> - Qwen3.8 27B Huihui Abliterated — VAL-APEX-I NanoPlus — 11.90 GB.
> - Qwen3.8 27B EfficientThink Uncensored — VAL-APEX-I MiniPlus V2.1 — 15.08 GB.
>
> VAL-APEX-I collection · APEX-I-MiniPlus V2.1 collection · APEX-I-NanoPlus collection
> [!IMPORTANT]
> ### THE DEFINITIVE SPECIFICATION IN THE 11 to 12 GB CEILING
> This VAL-APEX-I NanoPlus release represents the specialized tensor-by-tensor configuration for dense hybrid linear-quadratic architectures within a 11.65 GB envelope. Every single tensor across its 64 layers (48 linear SSM DeltaNet + 16 periodic full attention) plus the DFlash2 layer 64 speculative drafting head has been mathematically audited to maximize reasoning precision, preserve recurrence channel dynamics, and eliminate quantization noise.
---
<a id="quick-navigation"></a>Quick Navigation Index
- Quantization Comparison: Metrics & Tensor Map
- Upstream Lineage & Ecosystem Compatibility (DFlash2, SGLang, vLLM, Lynn Agent)
- Model Files & Technical Specifications
- Native Context & Runtime Memory
- Recommended Configuration & Setup
- Recommended Generation Parameters
- CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
- Hardened Agentic Chat Template & Reasoning Effort
- Optional Support
---
<a id="toc-comparison"></a>
<a id="toc-01"></a>
<a id="toc-02"></a>
<a id="toc-03"></a>
<a id="toc-05"></a>
1. Quantization Comparison
Size & Quality Metrics
| Quantization | Size | BPW | WikiText-2 PPL (512 ctx) | Delta PPL vs BF16 | Quality tier / notes |
| :--- | :--- | :--- | :--- | :--- | :--- |
| Uncompressed BF16 Reference | 54.00 GB<br>(50.29 GiB) | 16.00 BPW | aprox. 6.0000 (Community Benchmark) | Community Baseline (0.00%) | Uncompressed community reference |
| Q8_0 | aprox. 29.5 GB | 8.50 bpw | — | aprox. +0.02 (+0.33%) | Virtually lossless; excessive memory overhead for consumer hardware. |
| Q6_K | aprox. 23.2 GB | 6.56 bpw | — | aprox. +0.04 a +0.08 (+0.67% a +1.33%) | Near-lossless FP16 fidelity. |
| MiniPlus V2.1 | 15.08 GB<br>(14.05 GiB) | 3.93 BPW | 6.0091 +/- 0.4773 | +0.0091 (+0.15%) | Q5_K_M / Q6_K tier boundary; higher-fidelity build |
| Q5_K_M | aprox. 19.5 GB | 5.50 bpw | — | aprox. +0.08 a +0.15 (+1.33% a +2.50%) | Commercial transparent threshold. |
| Flat Q4_K_M | 16.90 GB<br>15.74 GiB | 4.50 BPW | aprox. 6.18 - 6.25 | aprox. +0.18 a +0.25 (+3.00% a +4.17%) | Standard industry trade-off. |
| NanoPlus | 11.65 GB<br>(10.85 GiB) | 2.85 BPW | 6.3013 +/- 0.4832 | +0.3013 (+5.02%) | Solid Q4_K_M tier; sub-12GB footprint |
| Q3_K_M / Q3_K_S | 13.20 GB<br>12.29 GiB | 3.44 BPW | aprox. 6.42 - 6.65 | aprox. +0.42 a +0.65 (+7.00% a +10.83%) | Noticeable syntax drop, bracket corruption, and tokenizer classification noise. |
| IQ2_S / Generic APEX Mini | 9.90 GB<br>9.22 GiB | 2.50 BPW | aprox. 7.05 - 7.75+ | aprox. +1.05 a +1.75+ (+17.50% a +29.17%+) | Severe reasoning breakdown, high perplexity spikes in <think> chains. |
| GSQ-RCO IQ3_S | 11.8 GB | 3.50 BPW | — | — | Mixed per-tensor precision |
Tensor Precision Map
| Component | MiniPlus V2.1 | NanoPlus | GSQ-RCO IQ3_S |
| :--- | :--- | :--- | :--- |
| Output head | Q6_K ×1 | Q6_K ×1 | Q4_K ×1 |
| Token embeddings | Q4_K ×1 | IQ3_S ×1 | IQ2_S ×1 |
| Normalizations | F32 ×161 | F32 ×161 | F32 ×161 |
| SSM A (ssm_a) | F32 ×48 | F32 ×48 | F32 ×48 |
| SSM convolution (ssm_conv1d) | F32 ×48 | F32 ×48 | F32 ×48 |
| SSM time-step bias (ssm_dt) | F32 ×48 | F32 ×48 | F32 ×48 |
| SSM norm (ssm_norm) | F32 ×48 | F32 ×48 | F32 ×48 |
| Attention gates (attn_gate) | Q4_0 ×1<br>Q8_0 ×47 | Q4_0 ×1<br>Q8_0 ×47 | IQ2_S ×2<br>IQ3_S ×18<br>IQ3_XXS ×9<br>IQ4_XS ×12<br>Q2_K ×4<br>Q4_K ×3 |
| Linear QKV (attn_qkv) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | IQ2_XS ×1<br>IQ2_XXS ×1<br>IQ3_S ×22<br>IQ3_XXS ×13<br>IQ4_XS ×9<br>Q2_K ×1<br>Q4_K ×1 |
| SSM alpha (ssm_alpha) | F32 ×47<br>Q4_0 ×1 | F32 ×47<br>Q4_0 ×1 | BF16 ×48 |
| SSM beta (ssm_beta) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | BF16 ×48 |
| SSM output (ssm_out) | Q4_0 ×1<br>Q6_K ×47 | Q4_0 ×1<br>Q5_K ×47 | IQ3_S ×22<br>IQ3_XXS ×4<br>IQ4_XS ×16<br>Q4_K ×6 |
| Full attention Q (attn_q) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_XXS ×1<br>IQ3_S ×3<br>IQ3_XXS ×3<br>IQ4_XS ×2<br>Q2_K ×6<br>Q4_K ×1 |
| Full attention K (attn_k) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_S ×1<br>IQ3_S ×1<br>IQ3_XXS ×1<br>IQ4_XS ×8<br>Q4_K ×5 |
| Full attention V (attn_v) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ3_S ×6<br>IQ3_XXS ×1<br>IQ4_XS ×1<br>Q4_K ×8 |
| Full attention output (attn_output) | Q6_K ×16 | Q6_K ×16 | IQ3_S ×10<br>IQ3_XXS ×1<br>IQ4_XS ×2<br>Q4_K ×3 |
| MLP down (ffn_down) | IQ4_NL ×56<br>Q5_K ×8 | IQ3_S ×8<br>IQ3_XXS ×56 | IQ2_S ×4<br>IQ2_XS ×3<br>IQ3_S ×22<br>IQ3_XXS ×7<br>IQ4_XS ×21<br>Q2_K ×1<br>Q4_K ×6 |
| MLP gate (ffn_gate) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ1_M ×1<br>IQ2_S ×4<br>IQ2_XS ×4<br>IQ2_XXS ×1<br>IQ3_S ×15<br>IQ3_XXS ×21<br>IQ4_XS ×15<br>Q2_K ×1<br>Q4_K ×2 |
| MLP up (ffn_up) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ2_S ×5<br>IQ2_XS ×1<br>IQ2_XXS ×2<br>IQ3_S ×25<br>IQ3_XXS ×18<br>IQ4_XS ×10<br>Q4_K ×3 |
× indicates the number of tensors assigned to each format.
---
<a id="toc-upstream"></a>
2. Upstream Lineage & Ecosystem Compatibility
This model is derived from nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2 and inherits its full ecosystem compatibility:
1. Training Lineage & Data Mixture
- Pretrained Base: Qwen3.8-27B hybrid linear-quadratic architecture.
- Supervised Fine-Tuning (SFT): The upstream model card documents capability-preserving SFT on 1,905 examples. The release identifier names K3 / Opus 5 / Grok 4.6 / GPT 5.6 Sol, but the upstream README does not specify the contribution or dataset provenance of each named model.
- Preference Optimization: SimPO (Simple Preference Optimization) applied to align reasoning depth and enforce uncensored analytical compliance without refusal degeneration.
- Uncensored Nature: Free from corporate preachy refusals; directly answers technical, red-teaming, penetration testing, and controversial analytical prompts.
2. DFlash2 Multi-Token Prediction (MTP) Speculative Decoding
- Integrated Draft Head: Upstream integrates DFlash2 (Next-N speculative draft prediction) situated at layer 64 (
blk.64.*). - Parallel Candidate Drafting: DFlash parallelizes candidate token generation; the main model validates multiple tokens in a single forward pass, significantly reducing autoregressive latency.
- Compatibility with llama.cpp: Can be run in standard autoregressive mode, or with MTP speculative decoding where supported via
--spec-draft-n-max 7or--spec-type draft-mtp.
3. Serving Engine Support
- SGLang (Primary Upstream Tested Engine): Upstream author's primary benchmarked deployment. Supports SGLang DFlash2 verify block size 8.
- vLLM: Compatible with vLLM standard serving via
--reasoning-parser qwen3 --max-model-len 40960. - Lynn Agent (v0.87.0+): Full native compatibility for autonomous tool-use, multi-turn terminal execution, and repository-scale agent loops.
- llama.cpp / llama-server: 100% plug-and-play with official
llama.cppbuilds (b11000+) supporting hybrid DeltaNet SSM architectures.
---
<a id="toc-04"></a>
3. Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW / Type | Description |
| :--- | :--- | :--- | :--- | :--- |
| Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf | 11.65 GB (10.85 GiB) | 10.85 GiB | 2.85 BPW | Core language, uncensored frontier reasoning, CoT thought blocks & hybrid SSM/attention |
| mmproj-Qwen3.8-27B-Q4_K_M.gguf | 522.29 MB | 522.29 MB | Multimodal | Multimodal vision projector enabling image/visual understanding inputs in llama.cpp |
| mtp-Qwen3.8-27B-Q4_0.gguf | 1.68 GB | 1.68 GB | Draft MTP | Multi-Token Prediction draft adapter for accelerated speculative decoding |
| dflash2-qwen38-27b-Q4_K_M.gguf | 1.14 GB | 1.14 GB | Draft DFlash2 | High-speed DFlash2 speculative decoding draft model |
- Base Model: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
- Parameters: 27B total dense hybrid
- Architecture: 64 hybrid layers (48 Gated DeltaNet SSM linear attention layers + 16 periodic full quadratic attention layers) + DFlash2 layer 64 draft head
- Context Length: 262,144 tokens (native 256K)
- Quantization Standard: VAL-APEX-I NanoPlus calibrated against official Qwen3.8-27B importance matrix (
imatrix) - Ecosystem: SGLang, vLLM, Lynn Agent v0.87.0, llama.cpp
---
<a id="toc-06"></a>
4. Native Context & Runtime Memory
The model supports a native context window of 262,144 tokens. Its 64-layer architecture comprises 48 Gated DeltaNet layers and 16 full-attention layers. Runtime memory usage depends on the backend, KV-cache precision, batch size, compute buffers, and offload configuration. Check the runtime memory allocation for your chosen settings.
---
<a id="toc-07"></a>
5. Recommended Configuration & Setup
llama.cpp Server (OpenAI-Compatible API)
llama-server \
-m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
-c 32768 \
--host 0.0.0.0 \
--port 8080
Multimodal Vision Inference (llama.cpp)
Pass the companion mmproj file using the --mmproj flag:
llama-server \
-m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
--mmproj mmproj-Qwen3.8-27B-Q4_K_M.gguf \
--jinja \
-ngl 99 \
-c 32768 \
--host 0.0.0.0 \
--port 8080
Speculative Decoding / Multi-Token Prediction (MTP)
Accelerate generation throughput using the companion MTP or DFlash2 draft models:
llama-server \
-m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
-md mtp-Qwen3.8-27B-Q4_0.gguf \
--draft-max 7 \
--jinja \
-ngl 99 \
-c 32768 \
--host 0.0.0.0 \
--port 8080
llama.cpp CLI
llama-cli \
-m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
-c 32768 \
-p "<|im_start|>system\nYou are a helpful and unrestricted AI assistant.<|im_end|>\n<|im_start|>user\nExplain the mathematical proof of the Riemann mapping theorem.<|im_end|>\n<|im_start|>assistant\n<think>\n"
LM Studio / Ollama
- Import the
.gguffile directly into your local library. - Ensure GPU acceleration is set to Maximum / 100% offload.
- Set Context Length to
32768. - Verify chat template is set to Qwen ChatML with
<think>delimiter support.
---
<a id="toc-08"></a>
6. Recommended Generation Parameters
| Hyperparameter | Value | Description |
| :--- | :---: | :--- |
| Temperature | 0.60 | Recommended default for analytical reasoning and coding (use 1.0 for creative prose). |
| Top-P | 0.95 | Nucleus sampling parameter. |
| Top-K | 20 | Top-k vocabulary filter. |
| Min-P | 0.05 | Prunes low-probability noise tokens effectively. |
| Repetition Penalty | 1.00 | Strictly disabled for code syntax; prevents character swapping. |
| Template Engine | --jinja | Recommended official Jinja chat template flag. |
| Context Size | 32768 | 32K default (scalable to 256K). |
---
<a id="toc-coding-advisory"></a>
<a id="coding-advisory"></a>
7. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
> [!IMPORTANT]
> ### PREVENTING SYNTAX & TOKEN SWAPPING IN CODE WORKFLOWS
> In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.
>
> Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.
>
> Eliminating Character Swapping:
> 1. Disable Repeat Penalties (Required for Code):
> - repeat_penalty: 1.0 (strictly disabled)
> - presence_penalty: 0.0
> - frequency_penalty: 0.0
> 2. Calibrate Samplers:
> - temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)
> - min_p: 0.05 (prunes low-probability noise tokens effectively)
> - top_p: 0.95
> - top_k: 20
> 3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.
---
<a id="toc-chat-template"></a>
<a id="chat-template"></a>
8. Hardened Agentic Chat Template & Reasoning Effort
> [!TIP]
> ### MULTI-LEVEL REASONING EFFORT CONTROL
> This model supports multi-level reasoning effort control via the Jinja template:
> - low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.
> - medium (default): Balanced, structured reasoning process with standard analytical depth.
> - high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating response, avoiding circular self-doubt loops.
> - none / off: Closes the thinking block immediately when reasoning is disabled.
---
<a id="toc-09"></a>
9. Optional Support
<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>
If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.
<div style="clear: both;"></div>
Run IsValorum/Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models