GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF overview

base model: Jab1718/qwen3.8 flash coder 85gb bf16 base model relation: quantized quantized by: IsValorum library name: gguf license: apache 2.0 language: en ta…

ggufllama.cppquantizedquantizationapexapex-quantapex-i-nanoplusnanoplusqwenqwen4qwen4-expmoereasoningagentic-codingcodingswe-benchimatrixenbase_model:Jab1718/qwen3.8-flash-coder-85gb-bf16base_model:quantized:Jab1718/qwen3.8-flash-coder-85gb-bf16license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~17.08 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
617
Likes
1
Pipeline
—
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.ggufGGUFGGUF17.08 GBDownload

Model Details

Model IDIsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
AuthorIsValorum
Pipeline—
Licenseapache-2.0
Base modelJab1718/qwen3.8-flash-coder-85gb-bf16
Last modified2026-10-07T16:14:41.000Z

Model README

---

base_model: Jab1718/qwen3.8-flash-coder-85gb-bf16

base_model_relation: quantized

quantized_by: IsValorum

library_name: gguf

license: apache-2.0

language:

  • en

tags:

  • gguf
  • llama.cpp
  • quantized
  • quantization
  • apex
  • apex-quant
  • apex-i-nanoplus
  • nanoplus
  • qwen
  • qwen4
  • qwen4-exp
  • moe
  • reasoning
  • agentic-coding
  • coding
  • swe-bench
  • imatrix

---

<a id="quick-navigation"></a>Quick Navigation Index

  1. Optimization History & Transparency Notice
  2. Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations
  3. Model Files & Technical Specifications
  4. Surgical Tensor Quantization Map (Audited from GGUF)
  5. Inference Quickstart
  6. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  7. Hardened Agentic Chat Template & Reasoning Effort
  8. Optional Support

> [!WARNING]

> ### EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST

> This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using moe-slice (352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).

>

> - English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.

> - Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.

> - Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock llama.cpp (llama-server) or LM Studio.

Qwen3.8-Flash-Coder APEX-I-NanoPlus GGUF

The Next-Generation Frontier MoE · Extreme 18GB Footprint · Fast System RAM Streaming & Massive Context on 16GB–24GB VRAM

> [!NOTE]

> ### EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP

> These are complementary APEX-I releases, not alternate downloads of the same model:

>

> - Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 — specialist agentic coding MoE tuned for Q5–Q6 quality (21.77 GB / 3.45 BPW).

> - Qwen3.8-Flash-Coder APEX-I-NanoPlus — ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).

> - Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF — uncompressed reference baseline (85.30 GB / 16.00 BPW).

> [!TIP]

> ### EMPIRICAL BENCHMARK & QUALITY COMPARISON

>

> | Quantization Specification | File Size (Disk) | Memory Footprint (RAM/VRAM) | Average BPW | WikiText-2 Perplexity | Delta PPL vs BF16 (%) | Quality Tier Equivalent |

> | :--- | :---: | :---: | :---: | :---: | :---: | :---: |

> | Uncompressed BF16 Reference | 85.30 GB (79.44 GiB) | 79.44 GiB | 16.00 BPW | 30.0975 +/- 0.1200 | Baseline (0.00%) | Lossless Reference Baseline |

> | APEX-I-MiniPlus V2.1 | 21.77 GB (20.27 GiB) | 20.27 GiB | 3.45 BPW | 30.1495 +/- 1.0089 | +0.0520 (+0.17%) | Q5_K_L / Q6_K Tier Boundary |

> | APEX-I-NanoPlus (CURRENT) | 18.34 GB (17.08 GiB) | 17.08 GiB | 2.90 BPW | 34.4199 +/- 1.1591 | +4.3224 (+14.36%) | Solid Q4_K_M Tier |

> | Standard Flat Q3_K_S | 20.41 GB | 19.01 GiB | 3.10 BPW | aprox. 30.75 - 31.20 | +0.65 a +1.10 (+2.9%) | High syntax degradation |

> | Generic APEX Mini (IQ2_S) | 17.73 GB | 16.51 GiB | 2.50 BPW | aprox. 31.60 - 33.10+ | +1.50 a +3.00+ (+7.5%) | Severe reasoning breakdown |

---

<a id="toc-03"></a>

Model Files & Technical Specifications

| File Name | File Size | Memory Footprint | BPW | Description |

| :--- | :--- | :--- | :--- | :--- |

| Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf | 18.34 GB (17.08 GiB) | 17.08 GiB | 2.90 BPW | Ultra-compact agentic coding MoE achieving solid Q4 quality with massive context capability |

  • Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16
  • Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
  • Architecture: Qwen4ExpForCausalLM (48 hybrid layers, 160 MoE experts with 10 active)
  • Context Length: 262,144 tokens (native 256K)

---

<a id="toc-04"></a>

Surgical Tensor Quantization Map (Audited from GGUF)

| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |

| :--- | :--- | :---: | :---: | :--- |

| Global Output Head | output.weight | 1 | Q6_K | Armored in high-precision Q6_K to preserve token classification. |

| Global Embeddings | token_embd.weight | 1 | Q3_K | Compact embedding representation across 248k vocabulary. |

| All Normalizations | output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm | 146 | F32 | 100% uncompressed numerical stability across all 48 layers. |

| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp | 96 | F32 | 100% uncompressed routing fidelity across 160 experts. |

| Attention Gates | blk.*.attn_gate.weight (36 SSM Layers) | 36 | Q8_0 | High-precision attention gating across hybrid DeltaNet recurrence layers. |

| Linear Attention Projections | blk.*.attn_qkv.weight (36 SSM Layers) | 36 | Q3_K | Efficient 3-bit quantization for SSM attention state inputs. |

| SSM Linear Output | blk.*.ssm_out.weight (36 SSM Layers) | 36 | Q5_K | 5-bit precision for linear state-space recurrence output. |

| Recurrent SSM Parameters | blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} | 216 | F32 | Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |

| Periodic Sparse Attention | blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) | 36 | Q4_K | Quadratic attention checkpoints for deep retrieval. |

| Periodic Sparse Attention Output | blk.{3,7,...}.attn_output.weight (12 Layers) | 12 | Q6_K | Armored attention output projection over deep context. |

| QSA Sparse Attention Indexer | blk.{3,7,...}.indexer.{q,k}_proj.weight | 24 | BF16 | High-fidelity sparse indexer projections for query-stream attention routing. |

| QSA Indexer Norms | blk.{3,7,...}.indexer.{q,k}_norm.weight | 24 | F32 | Uncompressed indexer layer normalizations. |

| Shared Foundation Experts | blk.*.ffn_down_shexp.weight (All 48 Layers) | 48 | Q5_0 | ne0=640 in standard block-32; zero divisibility crashes. |

| Shared Foundation Experts | blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) | 96 | Q5_K | Preserves core coding knowledge active on 100% of tokens. |

| Highway Connections (Down/Inject) | blk.*.hc_{attn,ffn}_{down,inject}.weight | 192 | Q6_K | High-fidelity residual highway bypass. |

| Highway Connections (Up) | blk.*.hc_{attn,ffn}_up.weight | 96 | Q5_0 | ne0=320 in standard block-32 format. |

| Routed MoE Down-Projections | blk.*.ffn_down_exps.weight (All 48 Layers) | 48 | IQ4_NL (36) / Q4_0 (12) | Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers. |

| Routed MoE Gate/Up Projections | blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) | 96 | IQ2_S (36) / IQ2_XXS (36) / IQ3_XXS (24) | Scaled expert density: sub-2.5 BPW on deep layers, IQ3_XXS on anchor layers. |

| Residual Output Highway | output_hc_{down,up}.weight | 2 | Q4_0 | Low-rank residual highway projections at model termination. |

| Residual Output Highway Norm | output_hc_norm.weight | 1 | F32 | Final residual normalization anchor. |

---

<a id="toc-05"></a>

Inference Quickstart

llama-server \
  -m Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  --ctx-size 65536 \
  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
  --port 8080

<a id="toc-coding-advisory"></a>

> [!IMPORTANT]

> ### CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

> Disable repeat penalties (repeat_penalty: 1.0, presence_penalty: 0.0, frequency_penalty: 0.0) and use --jinja to avoid syntax bracket substitutions.

<a id="toc-06"></a>

Optional Support

<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>

If these releases are helpful, voluntary support is welcome at https://ko-fi.com/isvalorum.

Run IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models