GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

jakeatx/slimder-qwen38-reap384-depth40-mb4-5-GGUF overview

SLIMDER Qwen3.8 REAP 384 depth 40 GGUF Portable llama.cpp export of the structurally materialized SLIMDER candidate remove macroblocks 4 5 . Provenance HF sour…

llama.cppggufqwen4-expslimderbase_model:jakeatx/slimder-qwen38-reap384-depth40-mb4-5base_model:quantized:jakeatx/slimder-qwen38-reap384-depth40-mb4-5license:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~60.06 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
270
Likes
0
Pipeline
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
slimder-qwen38-reap384-depth40-mb4-5-IQ3_XXS.ggufGGUFIQ3_XXS60.06 GBDownload
slimder-qwen38-reap384-depth40-mb4-5-Q4_K_M.ggufGGUFQ4_K_M82.40 GBDownload

Model Details

Model IDjakeatx/slimder-qwen38-reap384-depth40-mb4-5-GGUF
Authorjakeatx
Pipeline
Licenseapache-2.0
Base modelsjakek/slimder-qwen38-reap384-depth40-mb4-5
Last modified2026-09-10T23:43:25.000Z

Model README

---

base_model: sjakek/slimder-qwen38-reap384-depth40-mb4-5

library_name: llama.cpp

license: apache-2.0

tags:

- gguf

- qwen4-exp

- slimder

---

SLIMDER Qwen3.8 REAP-384 depth-40 GGUF

Portable llama.cpp export of the structurally materialized SLIMDER candidate

remove-macroblocks-4-5.

Provenance

  • HF source: sjakek/slimder-qwen38-reap384-depth40-mb4-5
  • Verified HF source revision: 2835be64b0429eddef35c001b818929223476c5e
  • Original portable S0 source: sjakek/slimder-qwen38-reap384-s0
  • Original S0 revision: 39b2218d23fa6d37f05ed1351c1884d1d43dc859
  • llama.cpp revision: 9723942adc518b43c4b95dc4dce6906903eb5e09
  • Source depth: 48 layers
  • Export depth: 40 layers
  • Removed source layers: 16 through 23 (macroblocks 4 and 5)
  • Materialized parameter count: 131,026,159,680

The source HF checkpoint passed a deterministic Transformers runtime smoke:

finite logits, no meta tensors or disk offload, non-empty generation, and an

identical repeated greedy generation. GGUF conversion, quantization, hashes,

and pinned llama.cpp runtime results are recorded alongside the artifact.

Status

Two artifacts are promoted and accompanied by exact SHA-256 records:

  • Q4_K_M: higher-quality reference quantization, 88,475,026,336 bytes.
  • IQ3_XXS: 64-GB-class candidate, 64,484,008,416 bytes (60.05 GiB).

The calibrated IQ3_XXS candidate passed deterministic llama.cpp loads at

2K, 8K, and 32K contexts and a bounded WikiText-2 perplexity gate: 4.9148

versus 4.3216 for Q4_K_M, a 1.1373x ratio under the 1.25x limit. Its single

uncalibrated output_hc_down.weight tensor was retained at Q8_0; the remaining

low-bit tensors use a 64-chunk importance matrix.

Rejected Q2_0, Q2_K, and IQ2_M experiments are represented only by small

reports and manifests; their large model files were not published.

<!-- qwen38-perian-lineage:start -->

Qwen3.8 Perian project lineage

This repository is retained in the

Qwen3.8 Perian checkpoints collection.

Its exact position in the lineage is: Depth-pruning precursor at 40 layers. It retains 384 experts per layer and the full PLE table and predates the Perian QLoRA.

The final Qwen3.8 Perian GGUF release

combines three reductions and one post-training stage:

  • depth: 48 to 32 transformer layers;
  • routed-expert width: 384 to 288 experts per layer;
  • PLE n-gram capacity: 320,001,446 to 160,000,768 rows (50%, about

25.60B parameters removed), using activation-aware bigram and

frequency-ranked trigram selections validated on a document-disjoint

5M-token holdout;

  • rank-32 QLoRA on 12,558 normalized traces spanning math/STEM

reasoning, coding/debugging, agentic tool use, retrieval, and general

multi-step reasoning. The trace mixture draws from several frontier-model

families, including Fable 5, GLM 5.2, Kimi K3, Claude Opus 4.7,

Qwen3.8-Max, and GPT-5.6-Sol. The final merged milestone was trained through

9,336,692 supervised assistant tokens.

Earlier checkpoints in this collection do not inherit later stages merely by

being listed beside them; the stage statement above is authoritative for this

artifact.

<!-- qwen38-perian-lineage:end -->

Run jakeatx/slimder-qwen38-reap384-depth40-mb4-5-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models