GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF overview

base model: Jab1718/qwen3.8 flash coder 85gb bf16 base model relation: quantized quantized by: IsValorum library name: gguf license: apache 2.0 language: en ta…

ggufllama.cppquantizedquantizationapexapex-quantapex-i-miniplusv2.1qwenqwen4qwen4-expmoereasoningagentic-codingcodingswe-benchimatrixenbase_model:Jab1718/qwen3.8-flash-coder-85gb-bf16base_model:quantized:Jab1718/qwen3.8-flash-coder-85gb-bf16license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~20.27 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
790
Likes
1
Pipeline
—
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.ggufGGUFGGUF20.27 GBDownload

Model Details

Model IDIsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
AuthorIsValorum
Pipeline—
Licenseapache-2.0
Base modelJab1718/qwen3.8-flash-coder-85gb-bf16
Last modified2026-10-07T16:14:44.000Z

Model README

---

base_model: Jab1718/qwen3.8-flash-coder-85gb-bf16

base_model_relation: quantized

quantized_by: IsValorum

library_name: gguf

license: apache-2.0

language:

  • en

tags:

  • gguf
  • llama.cpp
  • quantized
  • quantization
  • apex
  • apex-quant
  • apex-i-miniplus
  • v2.1
  • qwen
  • qwen4
  • qwen4-exp
  • moe
  • reasoning
  • agentic-coding
  • coding
  • swe-bench
  • imatrix

---

<a id="quick-navigation"></a>Quick Navigation Index

  1. Optimization History & Transparency Notice
  2. Empirical Benchmarks & Fidelity Verification
  3. Quality Spectrum: APEX-I-MiniPlus V2.1 vs. Standard Flat Quantizations
  4. Model Files & Technical Specifications
  5. Surgical Tensor Quantization Map (Audited from GGUF)
  6. Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
  7. The 24GB Miracle: Full 256K Context Runs In VRAM!
  8. Recommended Configuration & Setup
  9. Recommended Generation Parameters
  10. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  11. Hardened Agentic Chat Template & Reasoning Effort
  12. Optional Support

> [!WARNING]

> ### EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST

> This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using moe-slice (352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).

>

> - English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.

> - Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.

> - Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock llama.cpp (llama-server) or LM Studio.

Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 GGUF

The Definitive Frontier MoE · Efficient System RAM Offload · Full 256K Context on 24GB Workstations

> [!NOTE]

> ### EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP

> These are complementary APEX-I releases, not alternate downloads of the same model:

>

> - Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 - specialist agentic coding MoE tuned for Q5-Q6 quality (21.77 GB / 3.45 BPW).

> - Qwen3.8-Flash-Coder APEX-I-NanoPlus - ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).

> - Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF - uncompressed reference baseline (85.30 GB / 16.00 BPW).

> [!IMPORTANT]

> ### THE DEFINITIVE SPECIFICATION IN THE 21–22 GB CEILING

> This APEX-I-MiniPlus-V2.1 release represents the specialized tensor-by-tensor configuration for sparse Mixture-of-Experts quantization within a 21–22 GB envelope. Every tensor across its 48 hybrid layers (36 linear SSM DeltaNet + 12 sparse full attention) and 160 MoE experts has been mathematically allocated to maximize reasoning precision, preserve routing behavior, and prevent avoidable CPU dequantization stalls.

> [!TIP]

> ### EMPIRICAL BENCHMARK & QUALITY COMPARISON

>

> | Quantization Specification | File Size (Disk) | Memory Footprint (RAM/VRAM) | Average BPW | WikiText-2 Perplexity | Delta PPL vs BF16 (%) | Quality Tier Equivalent |

> | :--- | :---: | :---: | :---: | :---: | :---: | :---: |

> | Uncompressed BF16 Reference | 85.30 GB (79.44 GiB) | 79.44 GiB | 16.00 BPW | 30.0975 +/- 0.1200 | Baseline (0.00%) | Lossless Reference Baseline |

> | APEX-I-MiniPlus V2.1 (CURRENT) | 21.77 GB (20.27 GiB) | 20.27 GiB | 3.45 BPW | 30.1495 +/- 1.0089 | +0.0520 (+0.17%) | Q5_K_L / Q6_K Tier Boundary |

> | APEX-I-NanoPlus | 18.34 GB (17.08 GiB) | 17.08 GiB | 2.90 BPW | 34.4199 +/- 1.1591 | +4.3224 (+14.36%) | Solid Q4_K_M Tier |

> | Standard Flat Q4_K_M | 28.39 GB | 26.44 GiB | 4.50 BPW | aprox. 30.28 - 30.38 | +0.18 a +0.28 (+0.7%) | Standard industry trade-off |

> | Standard Flat Q3_K_M | 22.22 GB | 20.69 GiB | 3.44 BPW | aprox. 30.55 - 30.95 | +0.45 a +0.85 (+2.1%) | Noticeable syntax drop & bracket noise |

> | Generic APEX Mini (IQ2_S) | 17.73 GB | 16.51 GiB | 2.50 BPW | aprox. 31.60 - 33.10+ | +1.50 a +3.00+ (+7.5%) | Severe reasoning breakdown |

>

> Routing Fidelity: All router gates (ffn_gate_inp, ffn_gate_inp_shexp) remain in uncompressed F32, preserving zero routing drift across 160 experts.

>

> - Q6_K-Bordering Tier in Language Fidelity: WikiText-2 perplexity delta is exceptionally low (+0.0520 / +0.17% vs. BF16 baseline), placing overall code representation at the boundary of a 38 GB Q6_K build within an agile 21.77 GB footprint.

> - High-Precision Foundation Knowledge: Shared foundation experts run in uncompressed Q8_0 (down, block-32) + Q6_K (gate/up) across all 48 layers, keeping coding syntax intact across 100% of tokens.

> - Bypassed PLE Overhead: The 51B parameter N-gram table was cleanly removed (ple_layer_ids: []), enabling 100% GPU VRAM execution with zero host RAM overhead.

> [!WARNING]

> ### DO NOT CONFUSE APEX-I-MINIPLUS WITH GENERIC COMMUNITY APEX-I-MINI!

> - Generic Community APEX-I-Mini: Uniformly compresses all core MoE experts down to aggressive 2-bit IQ2_S, leaves the sensitive token output head unarmored at 3-bit Q3_K_M, and compresses attention projections down to Q3_K. In deep reasoning models, this triggers severe perplexity spikes, syntax errors, and broken code brackets.

> - Handcrafted APEX-I-MiniPlus V2.1: Preserves specified router gates in uncompressed F32, armors the token output head in high-precision Q6_K, safeguards attention gates in Q8_0, and keeps core reasoning experts at calibrated 3-bit to 4-bit non-linear quantization.

---

<a id="toc-01"></a>

Optimization History & Transparency Notice

We maintain our releases publicly as a transparent engineering record of continuous optimization. Below is the exact evolutionary roadmap of our MiniPlus architectures:

| Specification | Core Experts (10 Active/160) | Shared Expert (shexp) | Sparse Attention (12 Layers) | Highway Connections (4-Way) | Attention Gates (36 Layers) | Output Head (output.weight) | Routers (gate_inp) | Real-World Impact |

| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :--- |

| Generic APEX Mini | IQ2_S (2.50 bpw) | Q4_K / Q3_K | Q3_K | Compressed | Compressed | Q3_K_M | Compressed | High perplexity spikes, syntax errors in code generation. |

| MiniPlus V2.1 (CURRENT) | IQ3_XXS & IQ4_NL | Q8_0 + Q6_K (All 48 layers) | Q4_K (q/k/v) + Q6_K (output) | Q8_0 | Q8_0 | Q6_K | F32 | Definitive Build (21.77 GB). Zero AVX2 CPU stalls, zero divisibility crashes, and full 256K context support. |

---

<a id="toc-04"></a>

Model Files & Technical Specifications

| File Name | File Size | Memory Footprint | BPW | Description |

| :--- | :--- | :--- | :--- | :--- |

| Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.gguf | 21.77 GB (20.27 GiB) | 20.27 GiB | 3.45 BPW | Frontier agentic coding, SWE-bench vulnerability analysis, reasoning & hybrid SSM/sparse MoE |

  • Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16 (derived from Qwen/Qwen3.8-Flash-Next sliced with moe-slice and DoRA calibration)
  • Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
  • Architecture: Qwen4ExpForCausalLM (48 hybrid layers: 36 linear SSM DeltaNet + 12 sparse full attention, 160 MoE experts per layer with 10 active)
  • Context Length: 262,144 tokens (native 256K)
  • PLE / N-Gram Table: Decoupled and bypassed (ple_layer_ids: []) for 100% GPU VRAM execution.

---

<a id="toc-05"></a>

Surgical Tensor Quantization Map (Audited from GGUF)

The exact tensor breakdown below has been verified directly from the compiled binary weights:

| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |

| :--- | :--- | :---: | :---: | :--- |

| Global Output Head | output.weight | 1 | Q6_K | Preserves near-FP16 token classification; eliminates syntax errors, bracket drops, and hallucinations. |

| Global Embeddings | token_embd.weight | 1 | Q4_K | High-fidelity vocabulary embedding representation across 248k vocabulary. |

| All Normalizations | output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm | 146 | F32 | 100% uncompressed numerical stability across all 48 layers. |

| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp | 96 | F32 | 100% uncompressed routing fidelity across 160 experts; zero router drift. |

| Attention Gates | blk.*.attn_gate.weight (36 SSM Layers) | 36 | Q8_0 | High-precision attention gating across hybrid DeltaNet recurrence layers; eliminates crosstalk. |

| Linear Attention Projections | blk.*.attn_qkv.weight (36 SSM Layers) | 36 | Q4_K | Balanced precision for input state-space projections. |

| SSM Linear Output | blk.*.ssm_out.weight (36 SSM Layers) | 36 | Q6_K | Armored state-space output projection; preserves recurrence channel mixing. |

| Recurrent SSM Parameters | blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} | 216 | F32 | Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |

| Periodic Sparse Attention | blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) | 36 | Q4_K | Quadratic full-attention anchor checkpoints for deep needle-in-a-haystack retrieval. |

| Periodic Sparse Attention Output | blk.{3,7,...}.attn_output.weight (12 Layers) | 12 | Q6_K | High-precision attention output projection over deep context. |

| QSA Sparse Attention Indexer | blk.{3,7,...}.indexer.{q,k}_proj.weight | 24 | BF16 | High-fidelity sparse indexer projections for query-stream attention routing. |

| QSA Indexer Norms | blk.{3,7,...}.indexer.{q,k}_norm.weight | 24 | F32 | Uncompressed indexer layer normalizations. |

| Shared Foundation Experts | blk.*.ffn_down_shexp.weight (All 48 Layers) | 48 | Q8_0 | ne0=640 in standard block-32; eliminates divisibility validation aborts while preserving foundation knowledge. |

| Shared Foundation Experts | blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) | 96 | Q6_K | High-precision shared backbone active on 100% of tokens. |

| Highway Connections (4-Way) | blk.*.hc_{attn,ffn}_{down,inject,up}.weight | 288 | Q8_0 | High-speed residual highway bypass; ne0=320 up-projections mapped to block-32 Q8_0. |

| Routed MoE Down-Projections | blk.*.ffn_down_exps.weight (All 48 Layers) | 48 | IQ4_NL (36) / Q4_0 (12) | Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers. |

| Routed MoE Gate/Up Projections | blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) | 96 | IQ3_XXS (72) / Q4_K (24) | Calibrated with imatrix for maximum compactness, with Q4_K protection on anchor layers. |

| Residual Output Highway | output_hc_{down,up}.weight | 2 | Q4_0 | Low-rank residual highway projections at model termination. |

| Residual Output Highway Norm | output_hc_norm.weight | 1 | F32 | Final residual normalization anchor. |

---

<a id="toc-06"></a>

Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)

| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |

| :--- | :--- | :---: | :---: | :--- |

| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) | approx. 210 – 240 tok/s | 2,600 – 3,500+ tok/s | Blistering throughput on Hopper / Blackwell architectures |

| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) | 80 – 105+ tok/s | 1,800 – 2,400+ tok/s | Hybrid linear attention slashes prefill latency |

| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) | 65 – 82+ tok/s | 1,300 – 1,800+ tok/s | Full 256k native window in VRAM |

| Workstation / Laptop (DDR4 / DDR5 RAM) | Hybrid Offload (Few layers in VRAM) | Hardware-dependent | Hardware-dependent | Zero AVX2 CPU stalls; efficient streaming from system RAM |

---

<a id="toc-07"></a>

The 24GB Miracle: Full 256K Context Runs In VRAM!

Qwen3.8-Flash-Coder-85GB APEX-I-MiniPlus-V2.1 fits deep context windows within standard 24GB and 32GB GPUs:

| Context Length | Model Weights | KV Cache (q8_0) | Compute Buffers | Total GPU VRAM (Est.) | Feasibility |

| :--- | :--- | :--- | :--- | :--- | :--- |

| 32,768 (32k) | 20.27 GiB | 0.45 GiB | 1.50 GiB | 22.22 GiB | Effortless fit on 24GB GPUs (RTX 3090 / 4090) |

| 65,536 (64k) | 20.27 GiB | 0.85 GiB | 1.65 GiB | 22.77 GiB | Full offload on 24GB GPUs |

| 131,072 (128k)| 20.27 GiB | 1.45 GiB | 1.90 GiB | 23.62 GiB | Full offload on 24GB GPUs |

| 262,144 (256k)| 20.27 GiB | 2.60 GiB | 2.40 GiB | 25.27 GiB | Full offload on 32GB (RTX 5090) or partial RAM offload |

---

<a id="toc-08"></a>

Recommended Configuration & Setup

llama-server \
  -m Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.gguf \
  --jinja \
  -ngl 99 \
  --ctx-size 65536 \
  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
  --port 8080

<a id="toc-09"></a>

Recommended Generation Parameters

| Hyperparameter | Value | Description / Creator Notice |

| :--- | :---: | :--- |

| Temperature | 0.60 | Recommended default for code generation & reasoning (--temp 0.6). |

| Top-P | 0.95 | Source sampling parameter (--top-p 0.95). |

| Top-K | 20 | Source sampling parameter (--top-k 20). |

| Min-P | 0.00 | Source sampling parameter (--min-p 0.00). |

| Template Engine | --jinja | Recommended official chat template flag. |

| Context Size | 65536 | High-throughput 64K context (up to native 256K). |

<a id="toc-coding-advisory"></a>

> [!IMPORTANT]

> ### CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

> In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

>

> Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.

>

> Eliminating Character Swapping:

> 1. Disable Repeat Penalties (Required for Code):

> - repeat_penalty: 1.0 (strictly disabled)

> - presence_penalty: 0.0

> - frequency_penalty: 0.0

> 2. Calibrate Samplers:

> - temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)

> - min_p: 0.05 (prunes low-probability noise tokens effectively)

> - top_p: 0.95

> - top_k: 20

> 3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.

<a id="toc-chat-template"></a>

> [!TIP]

> ### HARDENED AGENTIC CHAT TEMPLATE (JINJA)

> The repository supports full Jinja chat templating with multi-level reasoning effort control:

> - low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.

> - medium (default): Balanced, structured reasoning process with standard analytical depth.

> - high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating code.

<a id="toc-10"></a>

Optional Support

<a href="https://ko-fi.com/isvalorum"><img src="https://huggingface.co/spaces/IsValorum/MiniPlus-NanoPlus-Requests/resolve/main/assets/dance-gold-ship.gif" alt="Gold Ship dancing" width="128" align="right"></a>

If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

Run IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models