IsValorum/Qwen3.8-Flash-Coder-85GB-BF16-GGUF overview
base model: Jab1718/qwen3.8 flash coder 85gb bf16 base model relation: quantized quantized by: IsValorum library name: gguf license: apache 2.0 language: en ta…
Runs locally from ~79.44 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.8-Flash-Coder-85GB-BF16.gguf | GGUF | BF16 | 79.44 GB | Download |
Model Details
Model README
---
base_model: Jab1718/qwen3.8-flash-coder-85gb-bf16
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:
- en
tags:
- gguf
- llama.cpp
- bf16
- lossless
- reference
- qwen
- qwen4
- qwen4-exp
- moe
- reasoning
- agentic-coding
- coding
- swe-bench
---
> [!WARNING]
> ### EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST
> This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using moe-slice (352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).
>
> - English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
> - Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
> - Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock llama.cpp (llama-server) or LM Studio.
Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF Reference (IsValorum)
The Official Uncompressed BF16 Reference GGUF · Golden Baseline for llama.cpp
> [!NOTE]
> ### EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP
> These are complementary APEX-I releases, not alternate downloads of the same model:
>
> - Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 - specialist agentic coding MoE tuned for Q5-Q6 quality (21.77 GB / 3.45 BPW).
> - Qwen3.8-Flash-Coder APEX-I-NanoPlus - ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).
> - Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF - uncompressed reference baseline (85.30 GB / 16.00 BPW).
This repository provides the official uncompressed BF16 GGUF reference format converted directly from Jab1718/qwen3.8-flash-coder-85gb-bf16.
This model serves as the lossless reference baseline for benchmarking, local testing, and high-fidelity inference with llama.cpp without any intermediate requantization noise.
Model Family & Verification Matrix
| Variant | File Size (Disk) | Memory Footprint (RAM/VRAM) | Average BPW | WikiText-2 Perplexity | Delta PPL vs BF16 (%) | Target Quality Tier | Repository Link |
| :--- | :---: | :---: | :---: | :---: | :---: | :--- | :--- |
| BF16 (Reference) | 85.30 GB (79.44 GiB) | 79.44 GiB | 16.00 BPW | 30.0975 +/- 0.1200 | Baseline (0.00%) | Uncompressed Baseline | IsValorum/Qwen3.8-Flash-Coder-85GB-BF16-GGUF |
| APEX-I-MiniPlus V2.1 | 21.77 GB (20.27 GiB) | 20.27 GiB | aprox. 3.45 BPW | 30.1495 +/- 1.0089 | +0.0520 (+0.17%) | Q5_K_L / Q6_K Tier | IsValorum/Qwen3.8-Flash-Coder-85GB-APEX-I-MiniPlus-V2.1-GGUF |
| APEX-I-NanoPlus | 18.34 GB (17.08 GiB) | 17.08 GiB | aprox. 2.90 BPW | 34.4199 +/- 1.1591 | +4.3224 (+14.36%) | Solid Q4_K_M Tier | IsValorum/Qwen3.8-Flash-Coder-85GB-APEX-I-NanoPlus-GGUF |
Architecture Details
- Base Architecture:
Qwen4ExpForCausalLM(48 hybrid layers: 36 linear attention SSM + 12 sparse attention, 4-way hyper-connections) - Active Parameters: approx. 3.7B active per token (10 active routed MoE experts out of 160 per layer + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings).
- Context Window: Native 262,144 tokens (256K).
- PLE / N-gram Table: Bypassed (
ple_layer_ids: []) for 100% GPU VRAM execution with zero host RAM offload.
Usage with llama.cpp
./llama-server \
-m ./Qwen3.8-Flash-Coder-85GB-BF16.gguf \
-c 65536 \
-ngl 999 \
--host 0.0.0.0 \
--port 8080Run IsValorum/Qwen3.8-Flash-Coder-85GB-BF16-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models