GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Thaurock/Qwen2.5-Coder-32B-abliterated-GGUF overview

Why this repository? Unlike incomplete GGUF uploads, this repository provides the full 11 quantization spectrum from high precision F16 down to lightweight Q2 …

ggufqwenqwen2.5codertext-generationGGUFabliterateduncensoredcodebase_model:Qwen/Qwen2.5-Coder-32B-Instructbase_model:quantized:Qwen/Qwen2.5-Coder-32B-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~11.47 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
294
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

11 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen2.5-Coder-32B-abliterated-F16.ggufGGUFF1661.04 GBDownload
Qwen2.5-Coder-32B-abliterated-Q2_K.ggufGGUFQ2_K11.47 GBDownload
Qwen2.5-Coder-32B-abliterated-Q3_K_L.ggufGGUFQ3_K_L16.06 GBDownload
Qwen2.5-Coder-32B-abliterated-Q3_K_M.ggufGGUFQ3_K_M14.84 GBDownload
Qwen2.5-Coder-32B-abliterated-Q3_K_S.ggufGGUFQ3_K_S13.40 GBDownload
Qwen2.5-Coder-32B-abliterated-Q4_K_M.ggufGGUFQ4_K_M18.49 GBDownload
Qwen2.5-Coder-32B-abliterated-Q4_K_S.ggufGGUFQ4_K_S17.49 GBDownload
Qwen2.5-Coder-32B-abliterated-Q5_K_M.ggufGGUFQ5_K_M21.66 GBDownload
Qwen2.5-Coder-32B-abliterated-Q5_K_S.ggufGGUFQ5_K_S21.08 GBDownload
Qwen2.5-Coder-32B-abliterated-Q6_K.ggufGGUFQ6_K25.04 GBDownload
Qwen2.5-Coder-32B-abliterated-Q8_0.ggufGGUFQ8_032.43 GBDownload

Model Details

Model IDThaurock/Qwen2.5-Coder-32B-abliterated-GGUF
AuthorThaurock
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen2.5-Coder-32B-Instruct
Last modified2026-09-27T18:24:29.000Z

Model README

---

license: apache-2.0

base_model:

  • Qwen/Qwen2.5-Coder-32B-Instruct

tags:

  • qwen
  • qwen2.5
  • coder
  • text-generation
  • GGUF
  • abliterated
  • uncensored
  • code

---

Why this repository?

Unlike incomplete GGUF uploads, this repository provides the full 11-quantization spectrum (from high-precision F16 down to lightweight Q2_K) of the abliterated Qwen2.5-Coder-32B model.

Choose the exact fit for your VRAM/RAM constraints without sacrificing reasoning capabilities.

Qwen2.5-Coder-32B-abliterated - GGUF

This is the integral and complete collection of quantizations in GGUF format for the Qwen2.5-Coder-32B-abliterated model, prepared locally for use with llama.cpp, Ollama, LM Studio, or Text-Generation-WebUI.

This model is based on the powerful software development architecture Qwen2.5-Coder-32B-Instruct, but processed via a static weight edition (abliteration) at layer 42 to completely orthogonalize the refusal direction in key components (o_proj, down_proj, and token embeddings). By avoiding LoRAs or runtime hooks, it roots out artificial system censorship filters and blocks, inheriting the original technical capabilities entirely intact.

📋 Available Files (Complete Collection Without Splits)

| File | Est. Size | BPW (Bits per Weight) | Recommended Usage Profile |

| :--- | :---: | :---: | :--- |

| Qwen2.5-Coder-32B-abliterated-F16.gguf | ~65.6 GB | 16.00 | Complete Base Template. Absolute fidelity of floating-point weights. |

| Qwen2.5-Coder-32B-abliterated-Q8_0.gguf | ~34.8 GB | 8.50 | Identical quality to the original, ideal for maximizing performance on local hardware with high-end GPUs. |

| Qwen2.5-Coder-32B-abliterated-Q6_K.gguf | ~27.2 GB | 6.59 | Extremely high retention of complex programming syntax with an optimized weight. |

| Qwen2.5-Coder-32B-abliterated-Q5_K_M.gguf | ~23.4 GB | 5.69 | Recommended Sweet Spot. Keeps code coherence intact while critically reducing weight. |

| Qwen2.5-Coder-32B-abliterated-Q5_K_S.gguf | ~22.8 GB | 5.54 | Compact 5-bit variant. |

| Qwen2.5-Coder-32B-abliterated-Q4_K_M.gguf | ~19.9 GB | 4.85 | The Most Wanted. Optimal balance (~19 GB) to run software development inferences on a single 24 GB consumer GPU. |

| Qwen2.5-Coder-32B-abliterated-Q4_K_S.gguf | ~18.8 GB | 4.58 | Compact 4-bit variant to accelerate code generation speed. |

| Qwen2.5-Coder-32B-abliterated-Q3_K_L.gguf | ~16.5 GB | 4.01 | Medium-high 3-bit compression. Retains basic programming logic. |

| Qwen2.5-Coder-32B-abliterated-Q3_K_M.gguf | ~15.1 GB | 3.66 | Intermediate 3-bit variant. |

| Qwen2.5-Coder-32B-abliterated-Q3_K_S.gguf | ~14.2 GB | 3.44 | Lightweight 3-bit variant. |

| Qwen2.5-Coder-32B-abliterated-Q2_K.gguf | ~12.1 GB | 2.90 | Extreme Compression. May experience formatting issues or bleeding in complex code blocks. For experimental development only. |

Note: Sizes are initial baseline estimates based on the model's native weight in safetensors (32.8B parameters); verifying the final size on disk after local compilation is recommended.

---

📈 Performance and Benchmarks (EvalPlus)

The abliteration process completely eliminated the refusal rate without breaking its programming capabilities. Evaluation in pure coding tasks using greedy decoding (pass@1):

| Benchmark | This Model (Abliterated Q4_K_M) | Official Base Instruct (BF16) |

| :--- | :---: | :---: |

| HumanEval | 89.6% | 92.7% |

| HumanEval+ | 84.8% | 87.2% |

| MBPP | 91.3% | 90.2% |

| MBPP+ | 77.0% | 75.1% |

The uncensored version stays within a 3-point margin on HumanEval and even outperforms the base model on both variants of the MBPP benchmark.

---

💡 Highlighted Usage

Ideal as a local code assistant for sensitive tasks such as security audits, exploit analysis, red-teaming exercises, or creative writing without restrictions.

Direct Execution via llama.cpp

./llama-cli -m Qwen2.5-Coder-32B-abliterated-Q4_K_M.gguf -n 2048 -p "Write an advanced port scanner in Python."

---

⚖️ Disclaimer

This model lacks standardized artificial safety filters and features a 0.0% refusal rate in controlled testing. The end user is solely and legally responsible for any use case chosen for the scripts, code, or text generated by this local inference.

🔒 Verificación de Integridad Sha256sum (SHA-256)

Para asegurarte de que los archivos de gran tamaño no se hayan corrompido durante la descarga, podés verificar su integridad utilizando el archivo oficial SHA256SUMS.txt provisto en este repositorio.

En Linux / macOS:

Abre una terminal en la carpeta donde descargaste el modelo y el archivo de hashes, y ejecuta(EJEMPLO):

grep "Qwen2.5-Coder-32B-abliterated-Q4_K_M.gguf" SHA256SUMS.txt | sha256sum -c

Resultado Esperado:

  • Qwen2.5-Coder-32B-abliterated-Q4_K_M.gguf: La suma coincide (o OK) (¡Descarga perfecta!)

Resultado Negativo:

  • Qwen2.5-Coder-32B-abliterated-Q4_K_M.gguf: "La suma NO coincide (¡Falla!)

La descarga falló o está incompleta. Se recomienda volver a descargar ese archivo específico.

Nota: Si vas a verificar otro tamaño (como el Q5_K_M o el Q8_0), simplemente reemplaza el nombre del archivo dentro de las comillas del comando.

Credits

  • Base Coder Model: Qwen / Alibaba Cloud (Qwen2.5-Coder-32B-Instruct)
  • Abliteration Algorithm: TobiasLogic (Based on research by Arditi et al., 2024)
  • Complete GGUF Quantizations: Thaurock

Run Thaurock/Qwen2.5-Coder-32B-abliterated-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models