bloomer010/Ling-3.0-flash-REAP384-97B-A5B-GGUF overview
This is an experimental REAP. Ling 3.0 flash REAP384 97B total / 5.1B active GGUF 384 of 512 routed experts kept per layer 25% of experts pruned from inclusion…
Runs locally from ~32.85 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | bloomer010/Ling-3.0-flash-REAP384-97B-A5B-GGUF |
|---|---|
| Author | bloomer010 |
| Pipeline | — |
| License | — |
| Base model | inclusionAI/Ling-3.0-flash |
| Last modified | 2026-08-21T17:25:23.000Z |
Model README
---
base_model: inclusionAI/Ling-3.0-flash
tags: [reap, expert-pruning, moe, bailingmoe, gguf]
---
This is an experimental REAP.
Ling-3.0-flash REAP384 (97B total / 5.1B active) - GGUF
[384 of 512 routed experts kept per layer - 25% of experts pruned]
from inclusionAI/Ling-3.0-flash
(124B total / 5.1B active).
Method: one-shot REAP (Router-weighted Expert Activation Pruning) -
experts scored by router-gate-value × output-L2-norm over calibration data, lowest-scoring deleted.
No fine-tuning, no recovery training.
Calibration: 1M tokens of ultrachat (chat-only calibration)
🎉 bailingmoe3 is supported in stock llama.cpp since
PR #26608 (merged 2026-08-17, commit
3733366720). Any build from that commit onward loads these files directly.
🔔 2026-08-21: added reasoning_effort support (low = thinking off, high = on, default same).
If you want reasoning_effort," re-download or override with chat_template.jinja.
Serving with experts in CPU RAM (attention on GPU, experts streamed from RAM):
llama-server -m Ling-3.0-flash-REAP384-94B-A5B-MXFP4.gguf -ngl 99 -ot "ffn_.*_exps\.weight=CPU" --no-mmap -c 65536 --flash-attn on --jinja
⚠️ Thinking model occasionally stops after thinking with empty content
(reasoning lands in reasoning_content);
serve with --reasoning-format none if your client only reads content.
Quants in this repo (all cut from a full-precision master): MXFP4, Q4_K_M, Q2_K
- MXFP4 (experts MXFP4 / rest Q8_0) is the pick for CPU-offload serving.
Run bloomer010/Ling-3.0-flash-REAP384-97B-A5B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models