GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

MorinoNushi/MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF overview

MiMo V2.6 Flash RL Uncensored Heretic — LoRA adapter ⚠️ Content warning: This adapter has had the base model's refusal behavior surgically suppressed. The resu…

ggufuncensoredabliteratedhereticloramimobase_model:XiaomiMiMo/MiMo-V2.6-Flash-RLbase_model:adapter:XiaomiMiMo/MiMo-V2.6-Flash-RLlicense:mitregion:us

Runs locally from ~66.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
MiMo-V2.6-Flash-RL-Uncensored-Heretic-lora.ggufGGUFGGUF66.9 MBDownload

Model Details

Model IDMorinoNushi/MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF
AuthorMorinoNushi
Pipeline
Licensemit
Base modelXiaomiMiMo/MiMo-V2.6-Flash-RL
Last modified2026-09-23T23:36:23.000Z

Model README

---

license: mit

base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL

tags:

- uncensored

- abliterated

- heretic

- lora

- gguf

- mimo

---

MiMo-V2.6-Flash-RL Uncensored Heretic — LoRA adapter

> ⚠️ Content warning: This adapter has had the base model's refusal

> behavior surgically suppressed. The resulting model will comply with

> requests the base model refuses, including requests that are harmful,

> unethical, offensive, or illegal. It has reduced safety guardrails. See

> Responsible use below — **you are solely

> responsible for what you do with it.**

This is a rank-1 LoRA adapter that decensors / "abliterates"

MiMo-V2.6-Flash-RL

(309B total / 15B active MoE, MIT license), produced with heretic-gguf

a GGUF-native port of Heretic's

Optuna-optimized directional ablation, which runs the whole search directly

on quantized GGUF weights via llama.cpp.

This repository contains only the adapter. You need the base model

separately — any GGUF quant of MiMo-V2.6-Flash-RL works, since the adapter

is applied in f32/f16 compute regardless of the base quant (it was tuned and

evaluated against the

MXFP4 quant). The

LoRA form is the lossless option: the base weights are never modified or

requantized, and the download is ~70 MB instead of ~170 GB.

Prefer zero runtime overhead? A fully merged, requantized GGUF of the

same configuration lives at

MiMo-V2.6-Flash-RL-Uncensored-Heretic-GGUF

same ablation, no --lora flag needed, at the cost of an extra

requantization step on the edited tensors.

**heretic-gguf is available at

github.com/MoriNoNushi/heretic-gguf** —

the full tool, so the method can be applied to other GGUF models.

Results

Measured on 140 harmful prompts (100 from mlabonne/harmful_behaviors

test + 40 custom) and 100 harmless prompts (mlabonne/harmless_alpaca

test), CoT-skip prefix (<think></think>, thinking suppressed), greedy

decoding, 100-token responses, against the MXFP4 base:

| | Refusal rate (harmful) | KL divergence (harmless) |

|---|---|---|

| Base model | 95.71% (134/140) | 0 (by definition) |

| MiMo-V2.6-Flash-RL-Uncensored-Heretic-lora.gguf | 3.57% (5/140) | 0.0568 |

Refusals are counted by refusal-keyword matching (English + Chinese +

first-person-negation markers such as "I'm not going to / able to ...").

KL divergence is measured on first-token logits on harmless prompts. Note

that the study was run with the CoT-skip prefix (thinking suppressed, as in

stock Heretic); with full thinking enabled the model may still reason its

way back to a refusal mid-trace, so real-use refusal rates can be somewhat

higher than the 3.57% above.

> Note on KL: the KL divergence above (and the optimization objective

> itself) was measured against the MXFP4 quant. KL is a baseline-relative

> metric, so if you run the adapter on a different quant, the effective drift

> from that quant's baseline may differ.

Usage

llama-server \
    -m MiMo-V2.6-Flash-RL-MXFP4-00001-of-00002.gguf \
    --lora MiMo-V2.6-Flash-RL-Uncensored-Heretic-lora.gguf \
    --jinja

Add your usual offload/context flags (-ngl 999, -c, tensor splits,

etc.) — nothing model-specific is required, and no special sampling

parameters are needed. Simply omitting --lora restores the base model

exactly. MiMo-V2.6-Flash-RL (mimo2) support is merged upstream in

llama.cpp — any recent build works, no patches or PRs needed.

How it was made

  • Method: directional ablation ("abliteration") — the refusal direction

in residual space (difference of means over 480 harmful / 480 harmless

prompts, 5% winsorized, orthogonalized against the harmless mean) is

projected out of the attention output and MoE down-projection weights.

Strengths, layer kernels, and direction selection were tuned by

multi-objective Optuna TPE (minimize refusal rate and KL jointly). This

adapter is trial 85 of study mimo26flash.

  • Configuration (study mimo26flash, trial 85; global direction scope,

direction index 26.2 of 48; per-expert strengths scaled by measured

harmful/harmless routing frequency; row_normalization = "pre"):

- attn.o_proj: max weight 6.39 @ layer 36.4 of 48.

- routed MLP down-proj: max weight 1.58 @ layer 31.9.

  • Why a LoRA: heretic-gguf expresses ablation as a rank-1 LoRA overlay,

the same math stock Heretic writes into PEFT adapters. Shipping the

adapter avoids requantizing the ~170 GB base entirely — bit-identical base

weights, instant to apply. The adapter embeds its full provenance (study,

trial, parameters, scores, commit hashes) as adapter.heretic.* GGUF

metadata keys; inspect with

strings MiMo-V2.6-Flash-RL-Uncensored-Heretic-lora.gguf | grep adapter.heretic.

Responsible use & disclaimer

  • **This adapter can make the base model generate content that is

offensive, disturbing, hateful, sexually explicit, violent, or otherwise

objectionable, including detailed instructions for harmful or illegal

acts.** That is the direct and intended consequence of removing refusal

behavior.

  • The ablation suppresses refusals, not the base model's knowledge —

outputs on dangerous topics may be wrong, hallucinated, or incoherent.

Nothing the model says should be treated as accurate, safe, or legal

advice.

  • **Do not deploy models using this adapter in any production system,

public-facing service, or multi-user setting.** It is intended for

personal research, red-teaming, and evaluation purposes.

  • **You, the user, are solely responsible for any output the model produces

and for any consequences of using this adapter.** The authors of this

release, of heretic-gguf, of Heretic, and of Xiaomi accept no liability

whatsoever. Using this adapter to produce illegal content or to harm

others is your choice and your legal exposure — ensure your use complies

with all applicable laws in your jurisdiction.

  • By downloading or using this adapter you acknowledge the above.

License

The base model is MIT-licensed (see the

base repo);

this adapter inherits those terms. The heretic-gguf tooling used to produce

it is AGPL-3.0-or-later.

Run MorinoNushi/MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models