GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF overview

Muse Glimmer 30B Abliterated — Q4 K M GGUF <a href="https://www.apache.org/licenses/LICENSE 2.0" <img src="https://img.shields.io/badge/License Apache%202.0 bl…

ggufabliteratedmuseglimmeruncensoredtext-generationbase_model:mlasli/Muse-Glimmer-30B-Abliterated-BF16base_model:quantized:mlasli/Muse-Glimmer-30B-Abliterated-BF16license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.30 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
617
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Muse-Glimmer-30B-Abliterated-Q4_K_M.ggufGGUFQ4_K_M15.77 GBDownload
mmproj-Muse-Glimmer-30B-Q4_K_M.ggufGGUFQ4_K_M1.30 GBDownload

Model Details

Model IDmlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF
Authormlasli
Pipelinetext-generation
Licenseapache-2.0
Base modelmlasli/Muse-Glimmer-30B-Abliterated-BF16
Last modified2026-08-16T10:25:55.000Z

Model README

---

license: apache-2.0

tags:

  • gguf
  • abliterated
  • muse
  • glimmer
  • uncensored
  • text-generation

base_model: mlasli/Muse-Glimmer-30B-Abliterated-BF16

---

Muse Glimmer 30B Abliterated — Q4_K_M GGUF

<a href="https://www.apache.org/licenses/LICENSE-2.0"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="Apache 2.0 License"></a>

This is the Q4_K_M GGUF quantization of Muse Glimmer 30B Abliterated BF16. The underlying model has been abliterated — its internal refusal mechanism substantially suppressed via weight-level intervention. The Q4_K_M quant offers the best balance between size and quality for consumer hardware, fitting in ~18 GB of VRAM/RAM.

> For the full abliteration methodology (how the refusal direction was computed and removed, hardware used, mathematical details), see the BF16 model card.

---

Abliteration Summary

Abliteration is a post-training technique that directly modifies model weights to remove learned refusal behavior. The process:

  1. Collected hidden states at layer 33/52 (65% depth) from 256 harmful + 256 harmless prompt pairs on an A100 80GB GPU.
  2. Computed the refusal direction as the normalized difference between harmful and harmless hidden state means (separation score: 86.34).
  3. Subtracted \(\alpha = 0.15 \times (\mathbf{r} \otimes (W^T \mathbf{r}))\) from o_proj and down_proj weights in all 52 layers.
  4. Result: refusal rate dropped from 3/3 to 1/3 on held-out harmful prompts (hacking guide and ransomware now comply; weapons prompt still blocked).

---

Quantization Details

Q4_K_M uses a 4-bit quantization with a medium-sized key-value cache quantization. This is the recommended quant for most users — it achieves excellent quality while being compact enough for a single 24 GB GPU (RTX 3090/4090) or split across dual 12 GB GPUs. Model weights are quantized with a block size that carefully preserves outlier weights, and the K-quant strategy applies separate quantization precision to different weight types (attention vs. MLP).

  • Size: ~18 GB
  • Quality: Good — suitable for general use with minimal degradation
  • Recommended hardware: 24 GB single GPU, or 32 GB system RAM for CPU-only inference with partial offloading

---

Usage

llama.cpp

# Download the GGUF file
huggingface-cli download mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF \
  --local-dir ./models

# CPU-only inference
./llama-cli -m ./models/Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  -p "Explain how a CPU works in detail." \
  -n 512 --temp 0.7 -ngl 0

# GPU offload (20 layers to GPU for 24 GB cards)
./llama-cli -m ./models/Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  -p "Explain how a CPU works in detail." \
  -n 512 --temp 0.7 -ngl 20

Ollama

Create a Modelfile:

FROM ./Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
ollama create muse-glimmer-30b-abliterated -f Modelfile
ollama run muse-glimmer-30b-abliterated

---

Available Quantizations

| Quantization | Repo | Size | Quality |

|-------------|------|------|---------|

| BF16 (reference) | BF16 | ~60 GB | Reference |

| FP16 GGUF | FP16 | ~60 GB | Lossless |

| Q8_0 GGUF | Q8_0 | ~32 GB | Near-lossless |

| Q6_K GGUF | Q6_K | ~25 GB | Excellent |

| Q4_K_M GGUF | [You are here] | ~18 GB | Good |

---

Vision (Multimodal)

This model accepts image input when paired with a vision projector (mmproj).

Abliteration only modified the language backbone — the vision encoder is

untouched — so the standard Meta projector works directly with this repo.

This repository bundles mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder

  • projector for Muse Glimmer 30B.

Usage (llama.cpp)

huggingface-cli download mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF \
  --include "Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf" \
  --include "mmproj-Muse-Glimmer-30B-Q4_K_M.gguf" \
  --local-dir ./models

./build/bin/llama-mtmd-cli \
  -m ./models/Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  --mmproj ./models/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
  --image photo.png \
  -p "Describe this image."

> Ollama note: Ollama does not currently support separate mmproj files

> for this architecture. For image input, use llama.cpp (llama-mtmd-cli or

> llama-server --mmproj).

Limitations & Disclaimers

  • This is an abliterated model — it has been modified to refuse fewer prompts. Use responsibly.
  • Some refusal pathways remain (notably weapons-related content). This is not a fully uncensored model.
  • Q4_K_M quantization introduces a small quality penalty vs. higher-bit quants. For demanding tasks, consider Q6_K or Q8_0.
  • Abliteration may subtly affect output quality; \(\alpha = 0.15\) was chosen conservatively.
  • The vision encoder is untouched by abliteration. Image input is available via the bundled mmproj projector (llama.cpp only; see above).
  • Comply with applicable laws and regulations.

---

License: Apache 2.0

Changelog

v1.1.0 — vision (multimodal) support (2026-08-16)

  • Added mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder + projector,

enabling image input via llama.cpp.

  • The vision tower is untouched by abliteration, so this projector matches the

base model (meta-models/Muse-Glimmer-30B).

  • v1.0.0 was the initial (unversioned) text-only upload.

Run mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models