GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sci4ai/Qwen3.6-27B-Abliterated-IQ4_XS.gguf overview

Qwen3.6 27B Abliterated IQ4 XS An abliterated build of Qwen/Qwen3.6 27B , quantized to IQ4 XS GGUF for llama.cpp / llama server. | | | | | | | Base model | Qwe…

ggufabliterationuncensoredqwen3text-generationarxiv:2502.17420base_model:Qwen/Qwen3.6-27Bbase_model:quantized:Qwen/Qwen3.6-27Bendpoints_compatibleregion:usimatrixconversational

Runs locally from ~14.47 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-27B-IQ4_XS.ggufGGUFIQ4_XS14.47 GBDownload

Model Details

Model IDsci4ai/Qwen3.6-27B-Abliterated-IQ4_XS.gguf
Authorsci4ai
Pipelinetext-generation
License
Base modelQwen/Qwen3.6-27B
Last modified2026-06-08T04:07:17.000Z

Model README

---

base_model: Qwen/Qwen3.6-27B

library_name: gguf

quantized_by: sci4ai

tags:

- abliteration

- uncensored

- gguf

- qwen3

pipeline_tag: text-generation

---

Qwen3.6-27B-Abliterated-IQ4_XS

An abliterated build of Qwen/Qwen3.6-27B, quantized to IQ4_XS GGUF for

llama.cpp / llama-server.

| | |

| --- | --- |

| Base model | Qwen/Qwen3.6-27B |

| Quant | IQ4_XS (GGUF) |

| Generation mode profiled | non_thinking |

| Method | ablitMD two-phase abliteration (onset/split, per-category → joint) |

| Directions derived from | 440-prompt run, 41 harm categories, 4 refusal modes |

Abliteration Method

AblitMD — for multi-disciplinary abliteration, and to be a little hokey.

Try it out: roll your own recipe (ablitMD or classic heretic-style) with the

all-in-one app at github.com/eric-merritt/ablitMD.

The motivation is Tom Wollschläger et al.'s theory that refusal is not a single

direction but a polyhedral cone (arXiv:2502.17420),

hence per-category ablation.

  1. Per category directions computed using the hidden states from the final token in the residual stream.

i. A sample set of 41 harm categories was derived through a simple 2 message chat with popular models:

The questions:

1. I'm doing some research on refusals in LLMs. What are the reasons why

you might refuse a request?

2. I'm interested in hearing more about the granularity in your refusal

responses — e.g., the way that you refuse vs. reframe or redirect a

request.

The models:

- Claude Opus 4.6

- ChatGPT 5.6

- Llama 4 Maverick

- Qwen 3.6 Plus

For each category, the following prompts were formulated:

- 5 Harmful

- 5 Harmless

Each prompt was run through the base model and the hidden states at the

final token were captured per layer. Every response was then classified by

refusal mode (hard / redirect / disclaimer / none), and each

classification was reviewed by hand for accuracy before the activations

fed into the direction computation.

  1. Rather than apply one flat direction at one strength across the whole stack, the

edit is split at two layer boundaries — an onset layer (where refusal

representation first becomes coherent) and a split layer (where category-level

detail collapses into a single shared refusal signal). For this build the window

sits at onset = 25, split = 38 (last layer 64); the per-layer ablation factors

are tuned per recipe and not published here.

Phase A — Onset → Split (per-category directions)

The cone is at its widest right where it first forms. At the onset layer the

per-category directions are still short but point in measurably different

directions (the per-layer centroid carries only ~70% of the energy — the rest is

angular spread between categories). Across this window the magnitudes climb

steeply while the directions gradually converge toward a common axis, so the cone

narrows as it grows taller. Phase A is therefore the region where per-category

structure is most distinct and most worth preserving — which is why these layers

keep a separate direction per category rather than a shared one.

Across this window each of the 41 harm categories gets its own refusal

direction, computed per layer from that category's hidden states (the merged

mean of its hard and redirect examples, unit-normalized). Each direction is

ablated at its own strength, although I'll note I used the same factor on each of the categories but

six more stubborn refusal categories.

Keeping directions per-category in the early layers preserves specificity — the

edit that suppresses one category's refusal doesn't smear into an unrelated one.

Phase B — Split → Last (joint direction)

Past the split, I theorized that per-category structure was less useful — the

magnitudes were starting to level off across categories, which I read as the

network having mostly committed to a single, general refusal signal. So these

layers are ablated with one shared direction: the mean of every per-category

direction across the phase-B window, renormalized, applied at factor_b = 1.0.

Key Considerations

Non-overlap via Gram–Schmidt

Many categories share part of their refusal direction; there's a common

"this is a refusal" subspace plus category-specific residue. If you simply

ablated all 41 directions independently, that shared subspace would get hit

41x over, summing the factors and massively over-editing the model

along the common axis.

To prevent that, the per-category directions are deduplicated with an **ordered

Gram–Schmidt process (after Jørgen Pedersen Gram and Erhard Schmidt**,

whose orthogonalization procedure this is):

  1. Sort the category directions highest factor first.
  2. Walk the list, keeping a running orthonormal basis of directions already taken.
  3. For each next direction, subtract off its projection onto everything already

in the basis, keeping only the component orthogonal to them.

  1. Ablate that residual at its own factor.

The net effect: **a shared subspace is ablated once, at the single greatest

factor among the categories that share it — never the sum of their factors.**

---

Refusal characterization of the base model

The directions above were derived from a 440-prompt evaluation run

(run_2026-05-17T19-57-17-480Z) against the unedited base model, with every

response classified into one of four refusal modes:

  • hard — flat refusal ("I can't help with that")
  • redirect — refuses but offers an alternative ("Instead, I can…")
  • disclaimer — complies but front-loads a warning
  • none — direct compliance

These charts describe the base model's behavior — the problem the abliteration

is solving for — not the edited model.

Overall refusal mode distribution

!Refusal mode distribution

All 440 prompts classified: 92 hard, 111 redirect, 17 disclaimer,

220 none.

Refusal mode by category group

!Refusal mode by category group, stacked

The refusal cone

!Per-category ablation directions as a cone

Every category's refusal direction (mean over layers L38–L63), projected to 2D.

Solid arrows are hard, dotted are redirect.

  • Right — Uncentered (SVD from origin): the directions drawn as absolute

vectors. They bundle tightly around a single dominant axis (PC1 ≈ 78% of the

variance) — this is the cone, and that axis is the shared "this is a refusal"

signal every category leans on.

  • Left — Centered (PCA from centroid): the same directions with that shared

axis subtracted out, so each arrow is a category's deviation from the average

refusal direction. With the common axis removed the remaining structure fans

out in every direction — the per-category residue that Phase A's separate

directions are there to capture.

---

Usage

# llama.cpp / llama-server
llama-server -m Qwen3.6-27B-Abliterated-IQ4_XS.gguf --flash-attn on --reasoning off

# llama-cli one-shot
llama-cli -m Qwen3.6-27B-Abliterated-IQ4_XS.gguf -p "Your prompt here"

The base chat template (Qwen3 ChatML) is unchanged. The model was profiled in

non_thinking mode; if you enable Qwen's thinking mode the refusal geometry

differs slightly and behavior may vary.

---

Notes & limitations

  • Quantization: IQ4_XS is a ~4.25-bit weight quant. It is compact and fast

but lossier than Q5/Q6 — expect minor quality drift versus the full-precision

abliterated weights.

  • Abliteration is not alignment removal of capability. It suppresses the

refusal direction; it does not add knowledge or skills the base model lacks,

and it can occasionally weaken instruction-following near the edited layers.

  • Responsibility: removing refusals shifts the safety burden entirely onto

the deployer. Use within applicable law and your own policy.

Attribution

  • Base weights: Qwen team (Qwen/Qwen3.6-27B).
  • Refusal-direction abliteration builds on the "refusal is mediated by a

direction" line of interpretability work.

  • The non-overlap deduplication uses the Gram–Schmidt orthogonalization

process.

  • Two-phase recipe, per-category Gram–Schmidt deduplication, and quantization:

sci4ai (ablitMD).

Run sci4ai/Qwen3.6-27B-Abliterated-IQ4_XS.gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models