jakeatx/slimder-qwen38-reap384-depth40-mb4-5-GGUF overview
SLIMDER Qwen3.8 REAP 384 depth 40 GGUF Portable llama.cpp export of the structurally materialized SLIMDER candidate remove macroblocks 4 5 . Provenance HF sour…
Runs locally from ~60.06 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
Model README
---
base_model: sjakek/slimder-qwen38-reap384-depth40-mb4-5
library_name: llama.cpp
license: apache-2.0
tags:
- gguf
- qwen4-exp
- slimder
---
SLIMDER Qwen3.8 REAP-384 depth-40 GGUF
Portable llama.cpp export of the structurally materialized SLIMDER candidate
remove-macroblocks-4-5.
Provenance
- HF source:
sjakek/slimder-qwen38-reap384-depth40-mb4-5 - Verified HF source revision:
2835be64b0429eddef35c001b818929223476c5e - Original portable S0 source:
sjakek/slimder-qwen38-reap384-s0 - Original S0 revision:
39b2218d23fa6d37f05ed1351c1884d1d43dc859 - llama.cpp revision:
9723942adc518b43c4b95dc4dce6906903eb5e09 - Source depth: 48 layers
- Export depth: 40 layers
- Removed source layers: 16 through 23 (macroblocks 4 and 5)
- Materialized parameter count: 131,026,159,680
The source HF checkpoint passed a deterministic Transformers runtime smoke:
finite logits, no meta tensors or disk offload, non-empty generation, and an
identical repeated greedy generation. GGUF conversion, quantization, hashes,
and pinned llama.cpp runtime results are recorded alongside the artifact.
Status
Two artifacts are promoted and accompanied by exact SHA-256 records:
Q4_K_M: higher-quality reference quantization, 88,475,026,336 bytes.IQ3_XXS: 64-GB-class candidate, 64,484,008,416 bytes (60.05 GiB).
The calibrated IQ3_XXS candidate passed deterministic llama.cpp loads at
2K, 8K, and 32K contexts and a bounded WikiText-2 perplexity gate: 4.9148
versus 4.3216 for Q4_K_M, a 1.1373x ratio under the 1.25x limit. Its single
uncalibrated output_hc_down.weight tensor was retained at Q8_0; the remaining
low-bit tensors use a 64-chunk importance matrix.
Rejected Q2_0, Q2_K, and IQ2_M experiments are represented only by small
reports and manifests; their large model files were not published.
<!-- qwen38-perian-lineage:start -->
Qwen3.8 Perian project lineage
This repository is retained in the
Qwen3.8 Perian checkpoints collection.
Its exact position in the lineage is: Depth-pruning precursor at 40 layers. It retains 384 experts per layer and the full PLE table and predates the Perian QLoRA.
The final Qwen3.8 Perian GGUF release
combines three reductions and one post-training stage:
- depth: 48 to 32 transformer layers;
- routed-expert width: 384 to 288 experts per layer;
- PLE n-gram capacity: 320,001,446 to 160,000,768 rows (50%, about
25.60B parameters removed), using activation-aware bigram and
frequency-ranked trigram selections validated on a document-disjoint
5M-token holdout;
- rank-32 QLoRA on 12,558 normalized traces spanning math/STEM
reasoning, coding/debugging, agentic tool use, retrieval, and general
multi-step reasoning. The trace mixture draws from several frontier-model
families, including Fable 5, GLM 5.2, Kimi K3, Claude Opus 4.7,
Qwen3.8-Max, and GPT-5.6-Sol. The final merged milestone was trained through
9,336,692 supervised assistant tokens.
Earlier checkpoints in this collection do not inherit later stages merely by
being listed beside them; the stage statement above is authoritative for this
artifact.
<!-- qwen38-perian-lineage:end -->
Run jakeatx/slimder-qwen38-reap384-depth40-mb4-5-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models