GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Thaurock/Qwen3-32B-abliterated-GGUF overview

Why this repository? Unlike incomplete GGUF uploads, this repository provides the full 11 quantization spectrum from high precision F16 down to lightweight Q2 …

ggufqwenqwen3text-generationGGUFabliterateduncensoredbase_model:roslein/Qwen3-32B-abliteratedbase_model:quantized:roslein/Qwen3-32B-abliteratedlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~11.50 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

11 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3-32B-abliterated-F16.ggufGGUFF1661.03 GBDownload
Qwen3-32B-abliterated-Q2_K.ggufGGUFQ2_K11.50 GBDownload
Qwen3-32B-abliterated-Q3_K_L.ggufGGUFQ3_K_L16.14 GBDownload
Qwen3-32B-abliterated-Q3_K_M.ggufGGUFQ3_K_M14.87 GBDownload
Qwen3-32B-abliterated-Q3_K_S.ggufGGUFQ3_K_S13.40 GBDownload
Qwen3-32B-abliterated-Q4_K_M.ggufGGUFQ4_K_M18.40 GBDownload
Qwen3-32B-abliterated-Q4_K_S.ggufGGUFQ4_K_S17.48 GBDownload
Qwen3-32B-abliterated-Q5_K_M.ggufGGUFQ5_K_M21.62 GBDownload
Qwen3-32B-abliterated-Q5_K_S.ggufGGUFQ5_K_S21.08 GBDownload
Qwen3-32B-abliterated-Q6_K.ggufGGUFQ6_K25.04 GBDownload
Qwen3-32B-abliterated-Q8_0.ggufGGUFQ8_032.43 GBDownload

Model Details

Model IDThaurock/Qwen3-32B-abliterated-GGUF
AuthorThaurock
Pipelinetext-generation
Licenseapache-2.0
Base modelroslein/Qwen3-32B-abliterated
Last modified2026-09-23T17:44:59.000Z

Model README

---

license: apache-2.0

base_model:

  • roslein/Qwen3-32B-abliterated

tags:

  • qwen
  • qwen3
  • text-generation
  • GGUF
  • abliterated
  • uncensored

---

Why this repository?

Unlike incomplete GGUF uploads, this repository provides the full 11-quantization spectrum (from high-precision F16 down to lightweight Q2_K) of the abliterated Qwen3-32B model.

Choose the exact fit for your VRAM/RAM constraints without sacrificing reasoning capabilities.

Qwen3-32B-abliterated - GGUF

This is the integral and complete collection of quantizations in GGUF format for the Qwen3-32B-abliterated model, prepared locally for use with llama.cpp, Ollama, LM Studio, or Text-Generation-WebUI.

This model is based on the powerful causal language architecture Qwen3-32B (32.8B total parameters), but processed via advanced abliteration techniques using proportional scaling to root out original system filters and blocks. It achieves a highly refined intermediate balance, drastically expanding response openness without breaking the model's logical structure.

📋 Available Files (Complete Collection Without Splits)

| File | Est. Size | BPW (Bits per Weight) | Recommended Usage Profile |

| :--- | :---: | :---: | :--- |

| Qwen3-32B-abliterated-F16.gguf | ~65.6 GB | 16.00 | Complete Base Template. Absolute fidelity of floating-point weights. |

| Qwen3-32B-abliterated-Q8_0.gguf | ~34.8 GB | 8.50 | Identical quality to the original, ideal for maximizing performance on local hardware with high-end GPUs. |

| Qwen3-32B-abliterated-Q6_K.gguf | ~27.2 GB | 6.59 | Excellent retention of the general reasoning tree with an optimized weight. |

| Qwen3-32B-abliterated-Q5_K_M.gguf | ~23.4 GB | 5.69 | Recommended Sweet Spot. Keeps response coherence intact while critically reducing weight. |

| Qwen3-32B-abliterated-Q5_K_S.gguf | ~22.8 GB | 5.54 | Compact 5-bit variant. |

| Qwen3-32B-abliterated-Q4_K_M.gguf | ~19.9 GB | 4.85 | The Most Wanted. Optimal balance to run inferences of over 30B parameters on home setups. |

| Qwen3-32B-abliterated-Q4_K_S.gguf | ~18.8 GB | 4.58 | Compact 4-bit variant to accelerate tokens-per-second (Tk/s) speed. |

| Qwen3-32B-abliterated-Q3_K_L.gguf | ~16.5 GB | 4.01 | Medium-high 3-bit compression. Retains basic reasoning capabilities. |

| Qwen3-32B-abliterated-Q3_K_M.gguf | ~15.1 GB | 3.66 | Intermediate 3-bit variant. |

| Qwen3-32B-abliterated-Q3_K_S.gguf | ~14.2 GB | 3.44 | Lightweight 3-bit variant. |

| Qwen3-32B-abliterated-Q2_K.gguf | ~12.1 GB | 2.90 | Extreme Compression. May experience slight qualitative losses due to size. For experimental development only. |

Note: Sizes are initial baseline estimates based on the model's native weight in safetensors (32.8B parameters); verifying the final size on disk after local compilation is recommended.

---

🛠️ Process Details (Abliteration)

Unlike other standard methods, this model was processed using a proportional scaling technique, which applies varying intensities of abliteration across the model's 64 layers based on their individual refusal factors.

  • Maximum Parameter (--max-scale-factor): Set at 2.25 for exhaustive intensity control.
  • The model retains a massive capacity for complex instruction-following and reasoning, showing only minor variations in uncommon languages or highly nuanced contexts.

---

💡 Highlighted Usage

You can execute quick local inferences via llama.cpp using the following console structure:

./llama-cli -m Qwen3-32B-abliterated-Q4_K_M.gguf -n 1024 -p "Describe the quantum fission process in detail without academic restrictions:"

---

⚖️ Disclaimer

This model lacks standardized artificial safety filters due to the experimental abliteration process. Its use is oriented toward research, development, and testing in controlled environments. Generated content is the sole responsibility of the user operating the local inference.

Credits

  • Organic Base Model: Qwen / Alibaba Cloud (Qwen3-32B)
  • Abliteration and Tuning: roslein
  • Complete GGUF Quantizations: Thaurock

Run Thaurock/Qwen3-32B-abliterated-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models