GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

SC117/Ling-3.0-tiny-abliterated-APEX-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: f5f5f7; border radius:…

llama.cppggufbailing-moelingabliterateduncensoreddecensoredapexquantizationimatrixtext-generationbase_model:inclusionAI/Ling-3.0-tinybase_model:quantized:inclusionAI/Ling-3.0-tinylicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~3.17 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6,436
Likes
11
Pipeline
text-generation
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-tiny-abliterated-APEX-I-Balanced.ggufGGUFGGUF5.55 GBDownload
Ling-3.0-tiny-abliterated-APEX-I-Compact.ggufGGUFGGUF3.71 GBDownload
Ling-3.0-tiny-abliterated-APEX-I-Quality.ggufGGUFGGUF5.37 GBDownload
Ling-3.0-tiny-abliterated-APEX-Mini.ggufGGUFGGUF3.17 GBDownload
Ling-3.0-tiny-abliterated-bf16.ggufGGUFBF1614.72 GBDownload

Model Details

Model IDSC117/Ling-3.0-tiny-abliterated-APEX-GGUF
AuthorSC117
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-tiny
Last modified2026-08-19T12:19:41.000Z

Model README

---

library_name: llama.cpp

license: mit

license_link: https://huggingface.co/inclusionAI/Ling-3.0-tiny/blob/main/LICENSE

pipeline_tag: text-generation

tags:

  • bailing-moe
  • ling
  • abliterated
  • uncensored
  • decensored
  • apex
  • quantization
  • imatrix
  • gguf

base_model:

  • inclusionAI/Ling-3.0-tiny

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">

<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">

<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #dbeafe; border-radius: 50%;"></div>

<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #bfdbfe; border-radius: 50%;"></div>

<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #dbeafe; border-radius: 50%;"></div>

<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1;">

<span style="background: #007aff; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">APEX</span>

<span style="background: #ff3b30; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">Abliterated</span>

<span style="background: #af52de; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">7.9B MoE</span>

</div>

<h1 style="margin: 0 0 8px 0; font-size: 32px; font-weight: 700; color: #1d1d1f; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Ling-3.0-tiny</h1>

<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;"><span style="color: #86868b;">English</span> | <a href="https://huggingface.co/SC117/Ling-3.0-tiny-abliterated-APEX-GGUF/blob/main/README_zh.md" style="color: #007aff; text-decoration: none;">📖 中文文档</a></p>

<p style="margin: 0; font-size: 15px; color: #86868b;">Abliterated APEX GGUF quants of <a href="https://huggingface.co/inclusionAI/Ling-3.0-tiny" style="color: #007aff; text-decoration: none;">inclusionAI/Ling-3.0-tiny</a> — 7.9B total / 1.3B active hybrid-reasoning MoE</p>

</div>

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff3b30 0%, #ff6b35 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔓</span> Abliteration</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 12px 0;">Refusal direction removed with <a href="https://github.com/wuwangzhang1216/abliterix" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">abliterix</a> using LoRA steering search on a bnb-4bit loaded model. Over 130 candidate steering configurations were explored; the Pareto-selected recipe was merged into the BF16 weights.</p>

<p style="margin: 0;"><b>85% fewer refusals</b> (15/100 vs 98/100 baseline) at 0.0677 KL divergence. Post-abliteration capability spot-checks (math, logic, coding, translation, knowledge) all pass with no degradation observed.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>💡</span> What is APEX?</div>

<div style="padding: 16px;">

<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7;">These GGUF files are quantized using <a href="https://github.com/mudler/apex-quant" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">APEX</a>, a novel MoE-aware mixed-precision quantization technique that outperforms standard quantization methods while being significantly smaller.</p>

<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7; font-weight: 700;">APEX beats Q8_0 perplexity at half the size — and even beats F16.</p>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">APEX classifies every tensor by its role — routed expert, shared expert, or attention — and applies a layer-wise precision gradient, giving the most sensitive edge layers higher precision and compressing the redundant middle layers more aggressively. Ling-3.0-tiny's 128 routed experts (only 8 active per token) make it an ideal candidate.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> APEX Quantization Tiers</div>

<div style="padding: 16px;">

<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(255,107,53,0.05);"><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 40%;">File</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 15%;">Size</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 15%;">BPW</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 30%;">Best For</th></tr></thead><tbody>

<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;"><code>*-APEX-I-Quality.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">5.77 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">5.84</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">Highest quality, best accuracy</td></tr>

<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;"><code>*-APEX-I-Balanced.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">5.96 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">6.03</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">Best all-rounder, recommended</td></tr>

<tr><td style="padding: 8px 10px; color: #334155;"><code>*-APEX-I-Compact.gguf</code></td><td style="padding: 8px 10px; color: #334155;">3.99 GB</td><td style="padding: 8px 10px; color: #334155;">4.10</td><td style="padding: 8px 10px; color: #334155;">Best quality/size ratio, 8 GB GPUs</td></tr>

<tr><td style="padding: 8px 10px; color: #334155;"><code>*-APEX-Mini.gguf</code></td><td style="padding: 8px 10px; color: #334155;">3.41 GB</td><td style="padding: 8px 10px; color: #334155;">3.45</td><td style="padding: 8px 10px; color: #334155;">Smallest viable, 6 GB GPUs, long context</td></tr>

</tbody></table>

<p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; font-style: italic;">All tiers quantized from a single BF16 source (15.07 GB) with the same diverse imatrix. Expert tensor layout: edge layers (L0–4, L19–23) keep higher precision than middle layers (L10–13); shared experts stay at Q8_0; the router is never quantized.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📐</span> I-Variant: Diverse Imatrix Calibration</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0;">All tiers use a diverse calibration dataset spanning chat, code, reasoning, and tool-calling — no Wikipedia (500 chunks). This produces higher accuracy on real-world benchmarks, lower KL divergence, and only a tiny perplexity increase on wikitext.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧠</span> Model Details</div>

<div style="padding: 16px;">

<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Architecture</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">BailingMoeV3 — hybrid KDA/MLA linear-attention MoE</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Parameters</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">7.9B total, 1.3B active per token</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Experts</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">128 routed experts + 1 shared expert, 8 routed active per token</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Layers</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">24 layers, 3:1 KDA–MLA stacking</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Context</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">131,072 tokens native</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Reasoning</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">Native hybrid reasoning (thinking mode on by default)</td></tr>

<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Abliteration</td><td style="padding: 6px 10px; color: #4b5563; background: white;">abliterix LoRA steering (85% fewer refusals, 0.0677 KL)</td></tr>

</tbody></table>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Usage</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">llama.cpp</p>

<p style="margin: 0 0 8px 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">hf download SC117/Ling-3.0-tiny-abliterated-APEX-GGUF --include "*.gguf" --local-dir ./models

./llama-server -m ./models/Ling-3.0-tiny-abliterated-APEX-I-Balanced.gguf -ngl 99 -c 32768 --flash-attn on --jinja</p>

<p style="margin: 0 0 12px 0; padding: 10px 14px; background: #fef3c7; border: 1px solid #fcd34d; border-radius: 6px; font-size: 12px; color: #92400e;">⚠️ <b>bailingmoe3 architecture support:</b> BailingMoE3 (<a href="https://github.com/ggml-org/llama.cpp/pull/26608" style="color: #c2410c;">PR #26608</a>) was merged into llama.cpp master on 2026-08-17 — the first release containing it is <b>b10470</b>. Use llama.cpp <b>b10470 or newer</b>. If you see <code>unknown model architecture: 'bailingmoe3'</code>, your build is too old — update and it will load.</p>

<p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">Ollama</p>

<p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">echo 'FROM ./Ling-3.0-tiny-abliterated-APEX-I-Balanced.gguf' &gt; Modelfile

ollama create ling-tiny-abliterated -f Modelfile &amp;&amp; ollama run ling-tiny-abliterated</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🎛️</span> Recommended Settings</div>

<div style="padding: 16px;">

<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(255,107,53,0.05);"><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 30%;">Mode</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 70%;">Parameters</th></tr></thead><tbody>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Thinking (default)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white; font-family: monospace;">temp=1.0, top_p=0.95, top_k=20</td></tr>

<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Fast / instruct</td><td style="padding: 6px 10px; color: #4b5563; font-family: monospace; background: white;">temp=0.7, top_p=0.8, top_k=20 (disable thinking via chat template)</td></tr>

</tbody></table>

</div>

</div>

</div>

Links

  • Original Model: https://huggingface.co/inclusionAI/Ling-3.0-tiny
  • APEX Quantization: https://github.com/mudler/apex-quant
  • abliterix: https://github.com/wuwangzhang1216/abliterix
  • BailingMoE3 GGUF support (llama.cpp): https://github.com/ggml-org/llama.cpp/pull/26608 (merged 2026-08-17, first release b10470)

Citation

@misc{ling3tiny,
title = {{Ling-3.0-tiny}: A Lightweight Hybrid Reasoning MoE Model},
url = {https://huggingface.co/inclusionAI/Ling-3.0-tiny},
author = {{inclusionAI}},
year = {2026}
}

Run SC117/Ling-3.0-tiny-abliterated-APEX-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models