GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF overview

Qwen3.8 27B AEON Ultimate Uncensored GGUF llama.cpp K quants of AEON 7/Qwen3.8 27B AEON ULTIMATE UNCENSORED BF16 https://huggingface.co/AEON 7/Qwen3.8 27B AEON…

ggufqwenqwen3.8llama.cppuncensoredabliteratedaeonmtptext-generationconversationalenzhbase_model:AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16base_model:quantized:AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16license:apache-2.0endpoints_compatibleregion:us

Runs locally from ~1.87 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
15
Likes
24
Pipeline
text-generation
Author

Repository Files & Downloads

9 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q2_K.ggufGGUFQ2_K10.12 GBDownload
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q3_K_M.ggufGGUFQ3_K_M12.57 GBDownload
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.ggufGGUFQ4_K_M15.66 GBDownload
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.ggufGGUFQ5_K_M18.19 GBDownload
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.ggufGGUFQ6_K20.89 GBDownload
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.ggufGGUFQ8_027.05 GBDownload
mtp-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-F16.ggufGGUFF165.54 GBDownload
mtp-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_0.ggufGGUFQ4_01.87 GBDownload
mtp-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.ggufGGUFQ8_02.95 GBDownload

Model Details

Model IDvcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF
Authorvcruz305
Pipelinetext-generation
Licenseapache-2.0
Base modelAEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
Last modified2026-08-16T23:22:26.000Z

Model README

---

language:

- en

- zh

license: apache-2.0

library_name: gguf

pipeline_tag: text-generation

base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16

base_model_relation: quantized

quantized_by: vcruz305

tags:

- gguf

- qwen

- qwen3.8

- llama.cpp

- uncensored

- abliterated

- aeon

- mtp

---

Qwen3.8-27B AEON Ultimate Uncensored GGUF

llama.cpp K-quants of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16, an abliterated BF16 of Qwen/Qwen3.8-27B.

This is not official Qwen. It is also not vcruz305/Qwen3.8-27B-GGUF (official trunk) or vcruz305/Qwen3.8-27B-Uncensored-GGUF (orcarouter FP8).

MTP is baked into every Q2–Q8 file (866 tensors, qwen35.nextn_predict_layers=1, +~0.24 GiB vs the old trunk-only files). You do not need a second GGUF for draft-mtp.

What is in these files

27B dense hybrid-attention (qwen35). 64 language-trunk blocks plus 1 nextn/MTP block. Converted from the AEON BF16 master (no --no-mtp). Pair a separate mmproj if you need vision.

The source removed the refusal direction. These GGUFs inherit that.

Chat template

Official 3.8 jinja wraps empty <think> blocks and breaks multi-turn agents. These GGUFs bake a fixed template. Use --jinja.

llama-server -m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf --jinja --reasoning-format deepseek

Files

| File | Quant | Bytes | GiB | Notes |

|---|---|---:|---:|---|

| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q2_K.gguf | Q2_K | 10864592928 | 10.12 | live; MTP baked in |

| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q3_K_M.gguf | Q3_K_M | 13500737568 | 12.57 | live; MTP baked in |

| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf | Q4_K_M | 16810715168 | 15.66 | live; MTP baked in |

| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf | Q5_K_M | 19535702048 | 18.20 | live; MTP baked in |

| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf | Q6_K | 22431000608 | 20.89 | live; MTP baked in; largest full-GPU on 24GB Turing |

| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf | Q8_0 | 29047085088 | 27.05 | live; MTP baked in; will not -ngl 99 on 24GB |

How to run MTP

Needs a llama.cpp build that understands Qwen3.5 nextn / draft-mtp. --parallel 1 is required.

llama-server \
  -m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.7 \
  --parallel 1 --jinja --reasoning-format deepseek \
  -a qwen38-27b-aeon --host 127.0.0.1 --port 8085 \
  -ngl 99 -fa on -b 512 -ub 512 -c 32768

Optional extra mtp-* sidecars are still in the repo if you want a separate -md draft. They are not required for the files above.

Download

Use hf_xet. Do not git clone. Grab only the file that exists.

export HF_XET_HIGH_PERFORMANCE=1
hf download vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
  --local-dir Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
  --include "Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf"

Q6_K is the largest file that still full-offloads a 24GB Turing card. Q8_0 does not.

Intended use

Local llama.cpp of the AEON uncensored 27B trunk + native MTP. Out of scope: treating this as official Qwen, or as a drop-in for the official-trunk GGUF pack.

Source

  • BF16 master: https://huggingface.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
  • Official base: https://huggingface.co/Qwen/Qwen3.8-27B
  • Convert: convert_hf_to_gguf.py --outtype f16 (MTP mixin on, no --no-mtp) → llama-quantize
  • License: Apache-2.0

Credits

Abliteration / BF16: AEON-7. Base: Qwen / Alibaba. GGUF pack: Victor Cruz (vcruz305).

Run vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models