SC117/Ling-3.0-tiny-abliterated-APEX-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: f5f5f7; border radius:…
Runs locally from ~3.17 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Ling-3.0-tiny-abliterated-APEX-I-Balanced.gguf | GGUF | GGUF | 5.55 GB | Download |
| Ling-3.0-tiny-abliterated-APEX-I-Compact.gguf | GGUF | GGUF | 3.71 GB | Download |
| Ling-3.0-tiny-abliterated-APEX-I-Quality.gguf | GGUF | GGUF | 5.37 GB | Download |
| Ling-3.0-tiny-abliterated-APEX-Mini.gguf | GGUF | GGUF | 3.17 GB | Download |
| Ling-3.0-tiny-abliterated-bf16.gguf | GGUF | BF16 | 14.72 GB | Download |
Model Details
| Model ID | SC117/Ling-3.0-tiny-abliterated-APEX-GGUF |
|---|---|
| Author | SC117 |
| Pipeline | text-generation |
| License | mit |
| Base model | inclusionAI/Ling-3.0-tiny |
| Last modified | 2026-08-19T12:19:41.000Z |
Model README
---
library_name: llama.cpp
license: mit
license_link: https://huggingface.co/inclusionAI/Ling-3.0-tiny/blob/main/LICENSE
pipeline_tag: text-generation
tags:
- bailing-moe
- ling
- abliterated
- uncensored
- decensored
- apex
- quantization
- imatrix
- gguf
base_model:
- inclusionAI/Ling-3.0-tiny
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">
<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #dbeafe; border-radius: 50%;"></div>
<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #bfdbfe; border-radius: 50%;"></div>
<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #dbeafe; border-radius: 50%;"></div>
<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1;">
<span style="background: #007aff; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">APEX</span>
<span style="background: #ff3b30; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">Abliterated</span>
<span style="background: #af52de; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">7.9B MoE</span>
</div>
<h1 style="margin: 0 0 8px 0; font-size: 32px; font-weight: 700; color: #1d1d1f; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Ling-3.0-tiny</h1>
<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;"><span style="color: #86868b;">English</span> | <a href="https://huggingface.co/SC117/Ling-3.0-tiny-abliterated-APEX-GGUF/blob/main/README_zh.md" style="color: #007aff; text-decoration: none;">📖 中文文档</a></p>
<p style="margin: 0; font-size: 15px; color: #86868b;">Abliterated APEX GGUF quants of <a href="https://huggingface.co/inclusionAI/Ling-3.0-tiny" style="color: #007aff; text-decoration: none;">inclusionAI/Ling-3.0-tiny</a> — 7.9B total / 1.3B active hybrid-reasoning MoE</p>
</div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff3b30 0%, #ff6b35 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔓</span> Abliteration</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;">Refusal direction removed with <a href="https://github.com/wuwangzhang1216/abliterix" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">abliterix</a> using LoRA steering search on a bnb-4bit loaded model. Over 130 candidate steering configurations were explored; the Pareto-selected recipe was merged into the BF16 weights.</p>
<p style="margin: 0;"><b>85% fewer refusals</b> (15/100 vs 98/100 baseline) at 0.0677 KL divergence. Post-abliteration capability spot-checks (math, logic, coding, translation, knowledge) all pass with no degradation observed.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>💡</span> What is APEX?</div>
<div style="padding: 16px;">
<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7;">These GGUF files are quantized using <a href="https://github.com/mudler/apex-quant" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">APEX</a>, a novel MoE-aware mixed-precision quantization technique that outperforms standard quantization methods while being significantly smaller.</p>
<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7; font-weight: 700;">APEX beats Q8_0 perplexity at half the size — and even beats F16.</p>
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">APEX classifies every tensor by its role — routed expert, shared expert, or attention — and applies a layer-wise precision gradient, giving the most sensitive edge layers higher precision and compressing the redundant middle layers more aggressively. Ling-3.0-tiny's 128 routed experts (only 8 active per token) make it an ideal candidate.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> APEX Quantization Tiers</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(255,107,53,0.05);"><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 40%;">File</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 15%;">Size</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 15%;">BPW</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 30%;">Best For</th></tr></thead><tbody>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;"><code>*-APEX-I-Quality.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">5.77 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">5.84</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">Highest quality, best accuracy</td></tr>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;"><code>*-APEX-I-Balanced.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">5.96 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">6.03</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #334155; background: white;">Best all-rounder, recommended</td></tr>
<tr><td style="padding: 8px 10px; color: #334155;"><code>*-APEX-I-Compact.gguf</code></td><td style="padding: 8px 10px; color: #334155;">3.99 GB</td><td style="padding: 8px 10px; color: #334155;">4.10</td><td style="padding: 8px 10px; color: #334155;">Best quality/size ratio, 8 GB GPUs</td></tr>
<tr><td style="padding: 8px 10px; color: #334155;"><code>*-APEX-Mini.gguf</code></td><td style="padding: 8px 10px; color: #334155;">3.41 GB</td><td style="padding: 8px 10px; color: #334155;">3.45</td><td style="padding: 8px 10px; color: #334155;">Smallest viable, 6 GB GPUs, long context</td></tr>
</tbody></table>
<p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; font-style: italic;">All tiers quantized from a single BF16 source (15.07 GB) with the same diverse imatrix. Expert tensor layout: edge layers (L0–4, L19–23) keep higher precision than middle layers (L10–13); shared experts stay at Q8_0; the router is never quantized.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📐</span> I-Variant: Diverse Imatrix Calibration</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0;">All tiers use a diverse calibration dataset spanning chat, code, reasoning, and tool-calling — no Wikipedia (500 chunks). This produces higher accuracy on real-world benchmarks, lower KL divergence, and only a tiny perplexity increase on wikitext.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧠</span> Model Details</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Architecture</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">BailingMoeV3 — hybrid KDA/MLA linear-attention MoE</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Parameters</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">7.9B total, 1.3B active per token</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Experts</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">128 routed experts + 1 shared expert, 8 routed active per token</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Layers</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">24 layers, 3:1 KDA–MLA stacking</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Context</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">131,072 tokens native</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Reasoning</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">Native hybrid reasoning (thinking mode on by default)</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Abliteration</td><td style="padding: 6px 10px; color: #4b5563; background: white;">abliterix LoRA steering (85% fewer refusals, 0.0677 KL)</td></tr>
</tbody></table>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Usage</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">llama.cpp</p>
<p style="margin: 0 0 8px 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">hf download SC117/Ling-3.0-tiny-abliterated-APEX-GGUF --include "*.gguf" --local-dir ./models
./llama-server -m ./models/Ling-3.0-tiny-abliterated-APEX-I-Balanced.gguf -ngl 99 -c 32768 --flash-attn on --jinja</p>
<p style="margin: 0 0 12px 0; padding: 10px 14px; background: #fef3c7; border: 1px solid #fcd34d; border-radius: 6px; font-size: 12px; color: #92400e;">⚠️ <b>bailingmoe3 architecture support:</b> BailingMoE3 (<a href="https://github.com/ggml-org/llama.cpp/pull/26608" style="color: #c2410c;">PR #26608</a>) was merged into llama.cpp master on 2026-08-17 — the first release containing it is <b>b10470</b>. Use llama.cpp <b>b10470 or newer</b>. If you see <code>unknown model architecture: 'bailingmoe3'</code>, your build is too old — update and it will load.</p>
<p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">Ollama</p>
<p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">echo 'FROM ./Ling-3.0-tiny-abliterated-APEX-I-Balanced.gguf' > Modelfile
ollama create ling-tiny-abliterated -f Modelfile && ollama run ling-tiny-abliterated</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🎛️</span> Recommended Settings</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(255,107,53,0.05);"><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 30%;">Mode</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c; font-weight: bold; width: 70%;">Parameters</th></tr></thead><tbody>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Thinking (default)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white; font-family: monospace;">temp=1.0, top_p=0.95, top_k=20</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Fast / instruct</td><td style="padding: 6px 10px; color: #4b5563; font-family: monospace; background: white;">temp=0.7, top_p=0.8, top_k=20 (disable thinking via chat template)</td></tr>
</tbody></table>
</div>
</div>
</div>
Links
- Original Model: https://huggingface.co/inclusionAI/Ling-3.0-tiny
- APEX Quantization: https://github.com/mudler/apex-quant
- abliterix: https://github.com/wuwangzhang1216/abliterix
- BailingMoE3 GGUF support (llama.cpp): https://github.com/ggml-org/llama.cpp/pull/26608 (merged 2026-08-17, first release b10470)
Citation
@misc{ling3tiny,
title = {{Ling-3.0-tiny}: A Lightweight Hybrid Reasoning MoE Model},
url = {https://huggingface.co/inclusionAI/Ling-3.0-tiny},
author = {{inclusionAI}},
year = {2026}
}Run SC117/Ling-3.0-tiny-abliterated-APEX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models