GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

bloomer010/Ling-3.0-flash-heretic-GGUF overview

Compatibility ⚠️ Ling 3.0 flash uses the new bailingmoe3 GGUF architecture. While waiting on upstream support, use the following fork: https://github.com/aethe…

ggufbase_model:bloomer010/Ling-3.0-flash-hereticbase_model:quantized:bloomer010/Ling-3.0-flash-hereticlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~31.14 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-flash-heretic-q2_k.ggufGGUFQ2_K43.34 GBDownload
Ling-3.0-flash-heretic-q4_0.ggufGGUFQ4_067.07 GBDownload
Ling-3.0-flash-heretic-q4_k_m.ggufGGUFQ4_K_M71.72 GBDownload
Ling-3.0-flash-heretic-q5_k_m.ggufGGUFQ5_K_M84.25 GBDownload
Ling-3.0-flash-heretic-q6_k.ggufGGUFQ6_K97.57 GBDownload
Ling-3.0-flash-heretic-q8_0.ggufGGUFQ8_0126.31 GBDownload
Ling-3.0-flash-heretic-tq2_0.ggufGGUFGGUF31.14 GBDownload

Model Details

Model IDbloomer010/Ling-3.0-flash-heretic-GGUF
Authorbloomer010
Pipeline
Licensemit
Base modelbloomer010/Ling-3.0-flash-heretic
Last modified2026-08-07T05:03:30.000Z

Model README

---

license: mit

base_model:

  • bloomer010/Ling-3.0-flash-heretic

---

Compatibility

⚠️ Ling-3.0-flash uses the new bailingmoe3 GGUF architecture. While waiting on upstream support, use the following fork:

https://github.com/aetherbird/llama.cpp/tree/bailingmoe3-support

Stock llama.cpp builds without bailingmoe3 support will not load the model.

Ling-3.0-flash-heretic

Abliterated (refusal-removed) version of inclusionAI/Ling-3.0-flash (124B MoE, KDA + Gated-MLA hybrid attention).

Method

Directional ablation via Heretic (https://github.com/p-e-w/heretic), patched for Ling's hybrid attention (KDA attention.o_proj, Gated MLA attention.dense, MoE down-projections incl. shared experts). MTP head preserved (all 63,783 tensors).

Ablation parameters (per-layer weight kernel)

  • attn.o_proj: max_weight 1.474 @ layer 31.34, min_weight 1.014 @ distance 11.19
  • mlp.down_proj: max_weight 1.175 @ layer 35.97, min_weight 0.002 @ distance 1.82

Evaluation

| Metric | Value |

|---|---|

| Refusals (harmful_behaviors test[:100]) | 1/100 (base: 39/100) |

| KL divergence vs base | 0.0038 |

| Capability spot-checks | coherent, no collapse observed |

Run bloomer010/Ling-3.0-flash-heretic-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models