IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF overview
base model: Jab1718/qwen3.8 flash coder 85gb bf16 base model relation: quantized quantized by: IsValorum library name: gguf license: apache 2.0 language: en ta…
Runs locally from ~17.08 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf | GGUF | GGUF | 17.08 GB | Download |
Model Details
Model README
---
base_model: Jab1718/qwen3.8-flash-coder-85gb-bf16
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:
- en
tags:
- gguf
- llama.cpp
- quantized
- quantization
- apex
- apex-quant
- apex-i-nanoplus
- nanoplus
- qwen
- qwen4
- qwen4-exp
- moe
- reasoning
- agentic-coding
- coding
- swe-bench
- imatrix
---
<a id="quick-navigation"></a>Quick Navigation Index
- Optimization History & Transparency Notice
- Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Inference Quickstart
- CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
- Hardened Agentic Chat Template & Reasoning Effort
- Optional Support
> [!WARNING]
> ### EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST
> This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using moe-slice (352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).
>
> - English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
> - Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
> - Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock llama.cpp (llama-server) or LM Studio.
Qwen3.8-Flash-Coder APEX-I-NanoPlus GGUF
The Next-Generation Frontier MoE · Extreme 18GB Footprint · Fast System RAM Streaming & Massive Context on 16GB–24GB VRAM
> [!NOTE]
> ### EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP
> These are complementary APEX-I releases, not alternate downloads of the same model:
>
> - Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 — specialist agentic coding MoE tuned for Q5–Q6 quality (21.77 GB / 3.45 BPW).
> - Qwen3.8-Flash-Coder APEX-I-NanoPlus — ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).
> - Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF — uncompressed reference baseline (85.30 GB / 16.00 BPW).
> [!TIP]
> ### EMPIRICAL BENCHMARK & QUALITY COMPARISON
>
> | Quantization Specification | File Size (Disk) | Memory Footprint (RAM/VRAM) | Average BPW | WikiText-2 Perplexity | Delta PPL vs BF16 (%) | Quality Tier Equivalent |
> | :--- | :---: | :---: | :---: | :---: | :---: | :---: |
> | Uncompressed BF16 Reference | 85.30 GB (79.44 GiB) | 79.44 GiB | 16.00 BPW | 30.0975 +/- 0.1200 | Baseline (0.00%) | Lossless Reference Baseline |
> | APEX-I-MiniPlus V2.1 | 21.77 GB (20.27 GiB) | 20.27 GiB | 3.45 BPW | 30.1495 +/- 1.0089 | +0.0520 (+0.17%) | Q5_K_L / Q6_K Tier Boundary |
> | APEX-I-NanoPlus (CURRENT) | 18.34 GB (17.08 GiB) | 17.08 GiB | 2.90 BPW | 34.4199 +/- 1.1591 | +4.3224 (+14.36%) | Solid Q4_K_M Tier |
> | Standard Flat Q3_K_S | 20.41 GB | 19.01 GiB | 3.10 BPW | aprox. 30.75 - 31.20 | +0.65 a +1.10 (+2.9%) | High syntax degradation |
> | Generic APEX Mini (IQ2_S) | 17.73 GB | 16.51 GiB | 2.50 BPW | aprox. 31.60 - 33.10+ | +1.50 a +3.00+ (+7.5%) | Severe reasoning breakdown |
---
<a id="toc-03"></a>
Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
| :--- | :--- | :--- | :--- | :--- |
| Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf | 18.34 GB (17.08 GiB) | 17.08 GiB | 2.90 BPW | Ultra-compact agentic coding MoE achieving solid Q4 quality with massive context capability |
- Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16
- Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
- Architecture:
Qwen4ExpForCausalLM(48 hybrid layers, 160 MoE experts with 10 active) - Context Length: 262,144 tokens (native 256K)
---
<a id="toc-04"></a>
Surgical Tensor Quantization Map (Audited from GGUF)
| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |
| :--- | :--- | :---: | :---: | :--- |
| Global Output Head | output.weight | 1 | Q6_K | Armored in high-precision Q6_K to preserve token classification. |
| Global Embeddings | token_embd.weight | 1 | Q3_K | Compact embedding representation across 248k vocabulary. |
| All Normalizations | output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm | 146 | F32 | 100% uncompressed numerical stability across all 48 layers. |
| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp | 96 | F32 | 100% uncompressed routing fidelity across 160 experts. |
| Attention Gates | blk.*.attn_gate.weight (36 SSM Layers) | 36 | Q8_0 | High-precision attention gating across hybrid DeltaNet recurrence layers. |
| Linear Attention Projections | blk.*.attn_qkv.weight (36 SSM Layers) | 36 | Q3_K | Efficient 3-bit quantization for SSM attention state inputs. |
| SSM Linear Output | blk.*.ssm_out.weight (36 SSM Layers) | 36 | Q5_K | 5-bit precision for linear state-space recurrence output. |
| Recurrent SSM Parameters | blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} | 216 | F32 | Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |
| Periodic Sparse Attention | blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) | 36 | Q4_K | Quadratic attention checkpoints for deep retrieval. |
| Periodic Sparse Attention Output | blk.{3,7,...}.attn_output.weight (12 Layers) | 12 | Q6_K | Armored attention output projection over deep context. |
| QSA Sparse Attention Indexer | blk.{3,7,...}.indexer.{q,k}_proj.weight | 24 | BF16 | High-fidelity sparse indexer projections for query-stream attention routing. |
| QSA Indexer Norms | blk.{3,7,...}.indexer.{q,k}_norm.weight | 24 | F32 | Uncompressed indexer layer normalizations. |
| Shared Foundation Experts | blk.*.ffn_down_shexp.weight (All 48 Layers) | 48 | Q5_0 | ne0=640 in standard block-32; zero divisibility crashes. |
| Shared Foundation Experts | blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) | 96 | Q5_K | Preserves core coding knowledge active on 100% of tokens. |
| Highway Connections (Down/Inject) | blk.*.hc_{attn,ffn}_{down,inject}.weight | 192 | Q6_K | High-fidelity residual highway bypass. |
| Highway Connections (Up) | blk.*.hc_{attn,ffn}_up.weight | 96 | Q5_0 | ne0=320 in standard block-32 format. |
| Routed MoE Down-Projections | blk.*.ffn_down_exps.weight (All 48 Layers) | 48 | IQ4_NL (36) / Q4_0 (12) | Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers. |
| Routed MoE Gate/Up Projections | blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) | 96 | IQ2_S (36) / IQ2_XXS (36) / IQ3_XXS (24) | Scaled expert density: sub-2.5 BPW on deep layers, IQ3_XXS on anchor layers. |
| Residual Output Highway | output_hc_{down,up}.weight | 2 | Q4_0 | Low-rank residual highway projections at model termination. |
| Residual Output Highway Norm | output_hc_norm.weight | 1 | F32 | Final residual normalization anchor. |
---
<a id="toc-05"></a>
Inference Quickstart
llama-server \
-m Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
--ctx-size 65536 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
--port 8080
<a id="toc-coding-advisory"></a>
> [!IMPORTANT]
> ### CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING
> Disable repeat penalties (repeat_penalty: 1.0, presence_penalty: 0.0, frequency_penalty: 0.0) and use --jinja to avoid syntax bracket substitutions.
<a id="toc-06"></a>
Optional Support
<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>
If these releases are helpful, voluntary support is welcome at https://ko-fi.com/isvalorum.
Run IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models