Thaurock/Qwen3-32B-abliterated-GGUF overview
Why this repository? Unlike incomplete GGUF uploads, this repository provides the full 11 quantization spectrum from high precision F16 down to lightweight Q2 …
Runs locally from ~11.50 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3-32B-abliterated-F16.gguf | GGUF | F16 | 61.03 GB | Download |
| Qwen3-32B-abliterated-Q2_K.gguf | GGUF | Q2_K | 11.50 GB | Download |
| Qwen3-32B-abliterated-Q3_K_L.gguf | GGUF | Q3_K_L | 16.14 GB | Download |
| Qwen3-32B-abliterated-Q3_K_M.gguf | GGUF | Q3_K_M | 14.87 GB | Download |
| Qwen3-32B-abliterated-Q3_K_S.gguf | GGUF | Q3_K_S | 13.40 GB | Download |
| Qwen3-32B-abliterated-Q4_K_M.gguf | GGUF | Q4_K_M | 18.40 GB | Download |
| Qwen3-32B-abliterated-Q4_K_S.gguf | GGUF | Q4_K_S | 17.48 GB | Download |
| Qwen3-32B-abliterated-Q5_K_M.gguf | GGUF | Q5_K_M | 21.62 GB | Download |
| Qwen3-32B-abliterated-Q5_K_S.gguf | GGUF | Q5_K_S | 21.08 GB | Download |
| Qwen3-32B-abliterated-Q6_K.gguf | GGUF | Q6_K | 25.04 GB | Download |
| Qwen3-32B-abliterated-Q8_0.gguf | GGUF | Q8_0 | 32.43 GB | Download |
Model Details
| Model ID | Thaurock/Qwen3-32B-abliterated-GGUF |
|---|---|
| Author | Thaurock |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | roslein/Qwen3-32B-abliterated |
| Last modified | 2026-09-23T17:44:59.000Z |
Model README
---
license: apache-2.0
base_model:
- roslein/Qwen3-32B-abliterated
tags:
- qwen
- qwen3
- text-generation
- GGUF
- abliterated
- uncensored
---
Why this repository?
Unlike incomplete GGUF uploads, this repository provides the full 11-quantization spectrum (from high-precision F16 down to lightweight Q2_K) of the abliterated Qwen3-32B model.
Choose the exact fit for your VRAM/RAM constraints without sacrificing reasoning capabilities.
Qwen3-32B-abliterated - GGUF
This is the integral and complete collection of quantizations in GGUF format for the Qwen3-32B-abliterated model, prepared locally for use with llama.cpp, Ollama, LM Studio, or Text-Generation-WebUI.
This model is based on the powerful causal language architecture Qwen3-32B (32.8B total parameters), but processed via advanced abliteration techniques using proportional scaling to root out original system filters and blocks. It achieves a highly refined intermediate balance, drastically expanding response openness without breaking the model's logical structure.
📋 Available Files (Complete Collection Without Splits)
| File | Est. Size | BPW (Bits per Weight) | Recommended Usage Profile |
| :--- | :---: | :---: | :--- |
| Qwen3-32B-abliterated-F16.gguf | ~65.6 GB | 16.00 | Complete Base Template. Absolute fidelity of floating-point weights. |
| Qwen3-32B-abliterated-Q8_0.gguf | ~34.8 GB | 8.50 | Identical quality to the original, ideal for maximizing performance on local hardware with high-end GPUs. |
| Qwen3-32B-abliterated-Q6_K.gguf | ~27.2 GB | 6.59 | Excellent retention of the general reasoning tree with an optimized weight. |
| Qwen3-32B-abliterated-Q5_K_M.gguf | ~23.4 GB | 5.69 | Recommended Sweet Spot. Keeps response coherence intact while critically reducing weight. |
| Qwen3-32B-abliterated-Q5_K_S.gguf | ~22.8 GB | 5.54 | Compact 5-bit variant. |
| Qwen3-32B-abliterated-Q4_K_M.gguf | ~19.9 GB | 4.85 | The Most Wanted. Optimal balance to run inferences of over 30B parameters on home setups. |
| Qwen3-32B-abliterated-Q4_K_S.gguf | ~18.8 GB | 4.58 | Compact 4-bit variant to accelerate tokens-per-second (Tk/s) speed. |
| Qwen3-32B-abliterated-Q3_K_L.gguf | ~16.5 GB | 4.01 | Medium-high 3-bit compression. Retains basic reasoning capabilities. |
| Qwen3-32B-abliterated-Q3_K_M.gguf | ~15.1 GB | 3.66 | Intermediate 3-bit variant. |
| Qwen3-32B-abliterated-Q3_K_S.gguf | ~14.2 GB | 3.44 | Lightweight 3-bit variant. |
| Qwen3-32B-abliterated-Q2_K.gguf | ~12.1 GB | 2.90 | Extreme Compression. May experience slight qualitative losses due to size. For experimental development only. |
Note: Sizes are initial baseline estimates based on the model's native weight in safetensors (32.8B parameters); verifying the final size on disk after local compilation is recommended.
---
🛠️ Process Details (Abliteration)
Unlike other standard methods, this model was processed using a proportional scaling technique, which applies varying intensities of abliteration across the model's 64 layers based on their individual refusal factors.
- Maximum Parameter (
--max-scale-factor): Set at2.25for exhaustive intensity control. - The model retains a massive capacity for complex instruction-following and reasoning, showing only minor variations in uncommon languages or highly nuanced contexts.
---
💡 Highlighted Usage
You can execute quick local inferences via llama.cpp using the following console structure:
./llama-cli -m Qwen3-32B-abliterated-Q4_K_M.gguf -n 1024 -p "Describe the quantum fission process in detail without academic restrictions:"
---
⚖️ Disclaimer
This model lacks standardized artificial safety filters due to the experimental abliteration process. Its use is oriented toward research, development, and testing in controlled environments. Generated content is the sole responsibility of the user operating the local inference.
Credits
Run Thaurock/Qwen3-32B-abliterated-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models