GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

andyjack/Huihui-Qwen3.5-4B-abliterated-GGUF overview

This is an uncensored version of Qwen/Qwen3.5 4B https://huggingface.co/Qwen/Qwen3.5 4B created with abliteration see remove refusals with transformers https:/…

transformersggufabliterateduncensoredimage-text-to-textbase_model:Qwen/Qwen3.6-27Bbase_model:quantized:Qwen/Qwen3.6-27Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~641.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
237
Likes
0
Pipeline
image-text-to-text
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Huihui-Qwen3.5-4B-Q4_K_M.ggufGGUFQ4_K_M2.52 GBDownload
Huihui-Qwen3.5-4B-mmproj-F16.ggufGGUFF16641.3 MBDownload

Model Details

Model IDandyjack/Huihui-Qwen3.5-4B-abliterated-GGUF
Authorandyjack
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelQwen/Qwen3.6-27B
Last modified2026-07-06T18:39:09.000Z

Model README

---

library_name: transformers

license: apache-2.0

license_link: https://huggingface.co/Qwen/Qwen3.6-27B/blob/main/LICENSE

pipeline_tag: image-text-to-text

base_model:

  • Qwen/Qwen3.6-27B

tags:

  • abliterated
  • uncensored

---

This is an uncensored version of Qwen/Qwen3.5-4B created with abliteration (see remove-refusals-with-transformers to know more about it).

This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens.

I needed an uncensored multi-modal GGUF in Q4_K_M to fit on a single Radeon R9 M295X / M390X with 4GB of VRAM.

This model was tested with llama.cpp compiled with VULKAN driver

cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release

The model quant was tested with Huihui-Qwen3.5-4B-mmproj-F16.gguf for multi-modal image use.

It generate 12 tokens/sec with the mmproj loaded and 25 tokens/sec without mmproj loaded.

Testing performed on a graphics card released in 2014.

llama.cpp

llama.cpp/build/bin/llama-server \
  -m Huihui-Qwen3.5-4B-Q4_K_M.gguf \
  --mmproj Huihui-Qwen3.5-4B-mmproj-F16.gguf \
  --host 0.0.0.0 \
  --port 11434 \
  --flash-attn on \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --n-gpu-layers 99 \
  --temp 0.7 \
  --top-p 0.95 \
  --top-k 20 \
  --ctx-size = 57344

Usage Warnings

- Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.

- Not Suitable for All Audiences: Due to limited content filtering, the model’s outputs may be inappropriate for public settings, underage users, or applications requiring high security.

- Legal and Ethical Responsibilities: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.

- Research and Experimental Use: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.

- Monitoring and Review Recommendations: Users are strongly advised to monitor model outputs in real-time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.

- No Default Safety Guarantees: Unlike standard models, this model has not undergone rigorous safety optimization. huihui.ai nor I bear any responsibility for any consequences arising from its use.

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah

Run andyjack/Huihui-Qwen3.5-4B-abliterated-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models