GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

IsValorum/Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus-GGUF overview

Qwen3.8 27B EfficientThink Uncensored VAL APEX I NanoPlus GGUF The SimPO + DFlash2 Frontier Reasoning Specialist · Synthetic SFT Opus 5 / Grok 4.6 / GPT 5.6 So…

ggufllama.cppquantizedquantizationval-apex-iapexapex-quantapex-i-nanoplusnanoplusqwenqwen3.8reasoningchain-of-thoughtdflash2speculative-decodingsimpouncensoredimatrixtext-generationconversationalenbase_model:nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2base_model:quantized:nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2license:apache-2.0

Runs locally from ~498.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.ggufGGUFGGUF10.85 GBDownload
dflash2-qwen38-27b-Q4_K_M.ggufGGUFQ4_K_M1.06 GBDownload
mmproj-Qwen3.8-27B-Q4_K_M.ggufGGUFQ4_K_M498.1 MBDownload
mtp-Qwen3.8-27B-Q4_0.ggufGGUFQ4_01.56 GBDownload

Model Details

Model IDIsValorum/Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus-GGUF
AuthorIsValorum
Pipelinetext-generation
Licenseapache-2.0
Base modelnerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
Last modified2026-10-07T02:58:14.000Z

Model README

---

base_model: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2

base_model_relation: quantized

quantized_by: IsValorum

library_name: gguf

license: apache-2.0

language:

  • en

tags:

  • gguf
  • llama.cpp
  • quantized
  • quantization
  • val-apex-i
  • apex
  • apex-quant
  • apex-i-nanoplus
  • nanoplus
  • qwen
  • qwen3.8
  • reasoning
  • chain-of-thought
  • dflash2
  • speculative-decoding
  • simpo
  • uncensored
  • imatrix

pipeline_tag: text-generation

---

Qwen3.8-27B-EfficientThink-Uncensored VAL-APEX-I NanoPlus GGUF

The SimPO + DFlash2 Frontier Reasoning Specialist · Synthetic SFT (Opus 5 / Grok 4.6 / GPT 5.6 Sol) · Native 256K Context

Official VAL-APEX-I quantization of nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2.

VAL-APEX-I stands for:

Vector-calibrated Asymmetric Layer-wise Outlier-preserving Recurrent-aware Unified Matrix-quantization

> [!NOTE]

> ### EXPLORE THE OTHER QWEN3.8 27B VAL-APEX-I EDITIONS

> Choose another base variant or fidelity profile:

>

> - Qwen3.8 27B Huihui Abliterated — VAL-APEX-I MiniPlus V2.1 — 15.33 GB.

> - Qwen3.8 27B Huihui Abliterated — VAL-APEX-I NanoPlus — 11.90 GB.

> - Qwen3.8 27B EfficientThink Uncensored — VAL-APEX-I MiniPlus V2.1 — 15.08 GB.

>

> VAL-APEX-I collection · APEX-I-MiniPlus V2.1 collection · APEX-I-NanoPlus collection

> [!IMPORTANT]

> ### THE DEFINITIVE SPECIFICATION IN THE 11 to 12 GB CEILING

> This VAL-APEX-I NanoPlus release represents the specialized tensor-by-tensor configuration for dense hybrid linear-quadratic architectures within a 11.65 GB envelope. Every single tensor across its 64 layers (48 linear SSM DeltaNet + 16 periodic full attention) plus the DFlash2 layer 64 speculative drafting head has been mathematically audited to maximize reasoning precision, preserve recurrence channel dynamics, and eliminate quantization noise.

---

<a id="quick-navigation"></a>Quick Navigation Index

  1. Quantization Comparison: Metrics & Tensor Map
  2. Upstream Lineage & Ecosystem Compatibility (DFlash2, SGLang, vLLM, Lynn Agent)
  3. Model Files & Technical Specifications
  4. Native Context & Runtime Memory
  5. Recommended Configuration & Setup
  6. Recommended Generation Parameters
  7. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  8. Hardened Agentic Chat Template & Reasoning Effort
  9. Optional Support

---

<a id="toc-comparison"></a>

<a id="toc-01"></a>

<a id="toc-02"></a>

<a id="toc-03"></a>

<a id="toc-05"></a>

1. Quantization Comparison

Size & Quality Metrics

| Quantization | Size | BPW | WikiText-2 PPL (512 ctx) | Delta PPL vs BF16 | Quality tier / notes |

| :--- | :--- | :--- | :--- | :--- | :--- |

| Uncompressed BF16 Reference | 54.00 GB<br>(50.29 GiB) | 16.00 BPW | aprox. 6.0000 (Community Benchmark) | Community Baseline (0.00%) | Uncompressed community reference |

| Q8_0 | aprox. 29.5 GB | 8.50 bpw | — | aprox. +0.02 (+0.33%) | Virtually lossless; excessive memory overhead for consumer hardware. |

| Q6_K | aprox. 23.2 GB | 6.56 bpw | — | aprox. +0.04 a +0.08 (+0.67% a +1.33%) | Near-lossless FP16 fidelity. |

| MiniPlus V2.1 | 15.08 GB<br>(14.05 GiB) | 3.93 BPW | 6.0091 +/- 0.4773 | +0.0091 (+0.15%) | Q5_K_M / Q6_K tier boundary; higher-fidelity build |

| Q5_K_M | aprox. 19.5 GB | 5.50 bpw | — | aprox. +0.08 a +0.15 (+1.33% a +2.50%) | Commercial transparent threshold. |

| Flat Q4_K_M | 16.90 GB<br>15.74 GiB | 4.50 BPW | aprox. 6.18 - 6.25 | aprox. +0.18 a +0.25 (+3.00% a +4.17%) | Standard industry trade-off. |

| NanoPlus | 11.65 GB<br>(10.85 GiB) | 2.85 BPW | 6.3013 +/- 0.4832 | +0.3013 (+5.02%) | Solid Q4_K_M tier; sub-12GB footprint |

| Q3_K_M / Q3_K_S | 13.20 GB<br>12.29 GiB | 3.44 BPW | aprox. 6.42 - 6.65 | aprox. +0.42 a +0.65 (+7.00% a +10.83%) | Noticeable syntax drop, bracket corruption, and tokenizer classification noise. |

| IQ2_S / Generic APEX Mini | 9.90 GB<br>9.22 GiB | 2.50 BPW | aprox. 7.05 - 7.75+ | aprox. +1.05 a +1.75+ (+17.50% a +29.17%+) | Severe reasoning breakdown, high perplexity spikes in <think> chains. |

| GSQ-RCO IQ3_S | 11.8 GB | 3.50 BPW | — | — | Mixed per-tensor precision |

Tensor Precision Map

| Component | MiniPlus V2.1 | NanoPlus | GSQ-RCO IQ3_S |

| :--- | :--- | :--- | :--- |

| Output head | Q6_K ×1 | Q6_K ×1 | Q4_K ×1 |

| Token embeddings | Q4_K ×1 | IQ3_S ×1 | IQ2_S ×1 |

| Normalizations | F32 ×161 | F32 ×161 | F32 ×161 |

| SSM A (ssm_a) | F32 ×48 | F32 ×48 | F32 ×48 |

| SSM convolution (ssm_conv1d) | F32 ×48 | F32 ×48 | F32 ×48 |

| SSM time-step bias (ssm_dt) | F32 ×48 | F32 ×48 | F32 ×48 |

| SSM norm (ssm_norm) | F32 ×48 | F32 ×48 | F32 ×48 |

| Attention gates (attn_gate) | Q4_0 ×1<br>Q8_0 ×47 | Q4_0 ×1<br>Q8_0 ×47 | IQ2_S ×2<br>IQ3_S ×18<br>IQ3_XXS ×9<br>IQ4_XS ×12<br>Q2_K ×4<br>Q4_K ×3 |

| Linear QKV (attn_qkv) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | IQ2_XS ×1<br>IQ2_XXS ×1<br>IQ3_S ×22<br>IQ3_XXS ×13<br>IQ4_XS ×9<br>Q2_K ×1<br>Q4_K ×1 |

| SSM alpha (ssm_alpha) | F32 ×47<br>Q4_0 ×1 | F32 ×47<br>Q4_0 ×1 | BF16 ×48 |

| SSM beta (ssm_beta) | Q4_0 ×1<br>Q4_K ×47 | IQ3_S ×47<br>Q4_0 ×1 | BF16 ×48 |

| SSM output (ssm_out) | Q4_0 ×1<br>Q6_K ×47 | Q4_0 ×1<br>Q5_K ×47 | IQ3_S ×22<br>IQ3_XXS ×4<br>IQ4_XS ×16<br>Q4_K ×6 |

| Full attention Q (attn_q) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_XXS ×1<br>IQ3_S ×3<br>IQ3_XXS ×3<br>IQ4_XS ×2<br>Q2_K ×6<br>Q4_K ×1 |

| Full attention K (attn_k) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ2_S ×1<br>IQ3_S ×1<br>IQ3_XXS ×1<br>IQ4_XS ×8<br>Q4_K ×5 |

| Full attention V (attn_v) | Q4_K ×14<br>Q5_K ×2 | IQ3_S ×14<br>Q4_K ×2 | IQ3_S ×6<br>IQ3_XXS ×1<br>IQ4_XS ×1<br>Q4_K ×8 |

| Full attention output (attn_output) | Q6_K ×16 | Q6_K ×16 | IQ3_S ×10<br>IQ3_XXS ×1<br>IQ4_XS ×2<br>Q4_K ×3 |

| MLP down (ffn_down) | IQ4_NL ×56<br>Q5_K ×8 | IQ3_S ×8<br>IQ3_XXS ×56 | IQ2_S ×4<br>IQ2_XS ×3<br>IQ3_S ×22<br>IQ3_XXS ×7<br>IQ4_XS ×21<br>Q2_K ×1<br>Q4_K ×6 |

| MLP gate (ffn_gate) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ1_M ×1<br>IQ2_S ×4<br>IQ2_XS ×4<br>IQ2_XXS ×1<br>IQ3_S ×15<br>IQ3_XXS ×21<br>IQ4_XS ×15<br>Q2_K ×1<br>Q4_K ×2 |

| MLP up (ffn_up) | IQ3_XXS ×56<br>Q4_K ×8 | IQ2_XXS ×56<br>IQ3_XXS ×8 | IQ2_S ×5<br>IQ2_XS ×1<br>IQ2_XXS ×2<br>IQ3_S ×25<br>IQ3_XXS ×18<br>IQ4_XS ×10<br>Q4_K ×3 |

× indicates the number of tensors assigned to each format.

---

<a id="toc-upstream"></a>

2. Upstream Lineage & Ecosystem Compatibility

This model is derived from nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2 and inherits its full ecosystem compatibility:

1. Training Lineage & Data Mixture

  • Pretrained Base: Qwen3.8-27B hybrid linear-quadratic architecture.
  • Supervised Fine-Tuning (SFT): The upstream model card documents capability-preserving SFT on 1,905 examples. The release identifier names K3 / Opus 5 / Grok 4.6 / GPT 5.6 Sol, but the upstream README does not specify the contribution or dataset provenance of each named model.
  • Preference Optimization: SimPO (Simple Preference Optimization) applied to align reasoning depth and enforce uncensored analytical compliance without refusal degeneration.
  • Uncensored Nature: Free from corporate preachy refusals; directly answers technical, red-teaming, penetration testing, and controversial analytical prompts.

2. DFlash2 Multi-Token Prediction (MTP) Speculative Decoding

  • Integrated Draft Head: Upstream integrates DFlash2 (Next-N speculative draft prediction) situated at layer 64 (blk.64.*).
  • Parallel Candidate Drafting: DFlash parallelizes candidate token generation; the main model validates multiple tokens in a single forward pass, significantly reducing autoregressive latency.
  • Compatibility with llama.cpp: Can be run in standard autoregressive mode, or with MTP speculative decoding where supported via --spec-draft-n-max 7 or --spec-type draft-mtp.

3. Serving Engine Support

  • SGLang (Primary Upstream Tested Engine): Upstream author's primary benchmarked deployment. Supports SGLang DFlash2 verify block size 8.
  • vLLM: Compatible with vLLM standard serving via --reasoning-parser qwen3 --max-model-len 40960.
  • Lynn Agent (v0.87.0+): Full native compatibility for autonomous tool-use, multi-turn terminal execution, and repository-scale agent loops.
  • llama.cpp / llama-server: 100% plug-and-play with official llama.cpp builds (b11000+) supporting hybrid DeltaNet SSM architectures.

---

<a id="toc-04"></a>

3. Model Files & Technical Specifications

| File Name | File Size | Memory Footprint | BPW / Type | Description |

| :--- | :--- | :--- | :--- | :--- |

| Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf | 11.65 GB (10.85 GiB) | 10.85 GiB | 2.85 BPW | Core language, uncensored frontier reasoning, CoT thought blocks & hybrid SSM/attention |

| mmproj-Qwen3.8-27B-Q4_K_M.gguf | 522.29 MB | 522.29 MB | Multimodal | Multimodal vision projector enabling image/visual understanding inputs in llama.cpp |

| mtp-Qwen3.8-27B-Q4_0.gguf | 1.68 GB | 1.68 GB | Draft MTP | Multi-Token Prediction draft adapter for accelerated speculative decoding |

| dflash2-qwen38-27b-Q4_K_M.gguf | 1.14 GB | 1.14 GB | Draft DFlash2 | High-speed DFlash2 speculative decoding draft model |

  • Base Model: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
  • Parameters: 27B total dense hybrid
  • Architecture: 64 hybrid layers (48 Gated DeltaNet SSM linear attention layers + 16 periodic full quadratic attention layers) + DFlash2 layer 64 draft head
  • Context Length: 262,144 tokens (native 256K)
  • Quantization Standard: VAL-APEX-I NanoPlus calibrated against official Qwen3.8-27B importance matrix (imatrix)
  • Ecosystem: SGLang, vLLM, Lynn Agent v0.87.0, llama.cpp

---

<a id="toc-06"></a>

4. Native Context & Runtime Memory

The model supports a native context window of 262,144 tokens. Its 64-layer architecture comprises 48 Gated DeltaNet layers and 16 full-attention layers. Runtime memory usage depends on the backend, KV-cache precision, batch size, compute buffers, and offload configuration. Check the runtime memory allocation for your chosen settings.

---

<a id="toc-07"></a>

5. Recommended Configuration & Setup

llama.cpp Server (OpenAI-Compatible API)

llama-server \
  -m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  --host 0.0.0.0 \
  --port 8080

Multimodal Vision Inference (llama.cpp)

Pass the companion mmproj file using the --mmproj flag:

llama-server \
  -m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
  --mmproj mmproj-Qwen3.8-27B-Q4_K_M.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  --host 0.0.0.0 \
  --port 8080

Speculative Decoding / Multi-Token Prediction (MTP)

Accelerate generation throughput using the companion MTP or DFlash2 draft models:

llama-server \
  -m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
  -md mtp-Qwen3.8-27B-Q4_0.gguf \
  --draft-max 7 \
  --jinja \
  -ngl 99 \
  -c 32768 \
  --host 0.0.0.0 \
  --port 8080

llama.cpp CLI

llama-cli \
  -m Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  -p "<|im_start|>system\nYou are a helpful and unrestricted AI assistant.<|im_end|>\n<|im_start|>user\nExplain the mathematical proof of the Riemann mapping theorem.<|im_end|>\n<|im_start|>assistant\n<think>\n"

LM Studio / Ollama

  1. Import the .gguf file directly into your local library.
  2. Ensure GPU acceleration is set to Maximum / 100% offload.
  3. Set Context Length to 32768.
  4. Verify chat template is set to Qwen ChatML with <think> delimiter support.

---

<a id="toc-08"></a>

6. Recommended Generation Parameters

| Hyperparameter | Value | Description |

| :--- | :---: | :--- |

| Temperature | 0.60 | Recommended default for analytical reasoning and coding (use 1.0 for creative prose). |

| Top-P | 0.95 | Nucleus sampling parameter. |

| Top-K | 20 | Top-k vocabulary filter. |

| Min-P | 0.05 | Prunes low-probability noise tokens effectively. |

| Repetition Penalty | 1.00 | Strictly disabled for code syntax; prevents character swapping. |

| Template Engine | --jinja | Recommended official Jinja chat template flag. |

| Context Size | 32768 | 32K default (scalable to 256K). |

---

<a id="toc-coding-advisory"></a>

<a id="coding-advisory"></a>

7. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)

> [!IMPORTANT]

> ### PREVENTING SYNTAX & TOKEN SWAPPING IN CODE WORKFLOWS

> In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

>

> Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.

>

> Eliminating Character Swapping:

> 1. Disable Repeat Penalties (Required for Code):

> - repeat_penalty: 1.0 (strictly disabled)

> - presence_penalty: 0.0

> - frequency_penalty: 0.0

> 2. Calibrate Samplers:

> - temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)

> - min_p: 0.05 (prunes low-probability noise tokens effectively)

> - top_p: 0.95

> - top_k: 20

> 3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.

---

<a id="toc-chat-template"></a>

<a id="chat-template"></a>

8. Hardened Agentic Chat Template & Reasoning Effort

> [!TIP]

> ### MULTI-LEVEL REASONING EFFORT CONTROL

> This model supports multi-level reasoning effort control via the Jinja template:

> - low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.

> - medium (default): Balanced, structured reasoning process with standard analytical depth.

> - high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating response, avoiding circular self-doubt loops.

> - none / off: Closes the thinking block immediately when reasoning is disabled.

---

<a id="toc-09"></a>

9. Optional Support

<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>

If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

<div style="clear: both;"></div>

Run IsValorum/Qwen3.8-27B-EfficientThink-Uncensored-VAL-APEX-I-NanoPlus-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models