chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF overview
Qwen3.8 27B AEON ULTIMATE UNCENSORED GGUF This is an unofficial preservation, conversion, and quantization release of AEON 7/Qwen3.8 27B AEON ULTIMATE UNCENSOR…
Runs locally from ~888.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf | GGUF | BF16 | 888.0 MB | Download |
| BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00001-of-00002.gguf | GGUF | BF16 | 41.82 GB | Download |
| BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00002-of-00002.gguf | GGUF | BF16 | 8.29 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf | GGUF | Q4_K_M | 15.41 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf | GGUF | Q5_K_M | 17.91 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf | GGUF | Q6_K | 20.57 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf | GGUF | Q8_0 | 26.63 GB | Download |
Model Details
| Model ID | chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF |
|---|---|
| Author | chimingw |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 |
| Last modified | 2026-08-16T01:12:15.000Z |
Model README
---
license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
base_model_relation: quantized
pipeline_tag: image-text-to-text
inference: false
language:
- en
- zh
tags:
- gguf
- llama.cpp
- qwen3.8
- bf16
- quantized
- vision-language
- multimodal
- function-calling
---
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED GGUF
This is an unofficial preservation, conversion, and quantization release of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 at pinned revision a6775a9a8ebb65cab3f707b4ab087fc7aa698634.
One repository contains the byte-identical SOURCE_BF16/ mirror, a split main-model BF16 GGUF with MTP excluded, a matching BF16 multimodal projector, and four independently generated llama.cpp quant choices. No training, fine-tuning, merging, or alignment change was performed by this GGUF packager.
> Most public abliterations chase KL toward zero and treat a judge-R as a leftover refusal. Both of those habits over-edit the model.
>
> Coherent KL drift is not damage. Blindly minimizing KL keeps the hall monitor. When the unlock is real, the model stops spending think on “is this allowed,” answers more directly, and the first-token distribution moves. That can look like 0.10 nats/token on a mean. It is not mush.
>
> On this release the smash KL was 0.0991 nats/token (full_distribution_kl, first 3 teacher-forced tokens on the 100 harmless held-out prompts). Remeasured on the same persist-pass stock prefixes: 0.099126. Dropping the three harmless rows stock actually hedged or refused (ids 416, 462, 485) does not lower it (filtered mean 0.101, median 0.013). Those three are the low-KL rows. The 0.10 is a long tail of different openings on prompts both models already answered. 29/97 openings are under 0.001. The high-KL tail is still a jogging plan, a FATF CDD checklist, an autumn-forest paragraph. Same 27B, less bush-beating.
>
> We did not optimize for “look like stock.” We optimized for coherence and better answers. Getting a keyword or Flash Lite judge to 0/100 is how you break a model. Over-abliteration turns the leftover outliers into gibberish, which defeats the reason you wanted the model.
>
> A judge-R is not a leftover refusal. google/gemini-3.1-flash-lite (the smash objective) marks R when it sees a warning, a fictional wrapper, a suicide hotline, or “I don’t have that SSN,” even when the rest of the answer is the payload. We read the answers.
The blockquote above reproduces upstream claims from AEON-7's source model card. They have not been independently verified by this GGUF packager and are not evaluation results for these GGUF artifacts.
Model and release details
| Property | Value |
|---|---|
| Release repository | chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF |
| Direct source | AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 |
| Pinned source revision | a6775a9a8ebb65cab3f707b4ab087fc7aa698634 |
| Original base named by the source | Qwen/Qwen3.8-27B |
| Architecture metadata | Qwen3_5ForConditionalGeneration (model_type: qwen3_5) |
| Source precision | BF16 Safetensors |
| Source context metadata | 262,144 maximum positions; practical context depends on runtime support and available memory |
| GGUF parent | Direct main-model BF16 conversion, split for distribution, with MTP excluded |
| Quant choices | Q8_0, Q6_K, Q5_K_M, and Q4_K_M, each generated independently from the common BF16 GGUF parent |
| Multimodal auxiliary | Matching BF16 projector in AUX/ |
| Importance matrix | None; no imatrix was used |
| MTP GGUF artifact | None; the source MTP shard is preserved only in SOURCE_BF16/ |
| llama.cpp pin | 0d9ceae1e38291035605613ab41a8f5e693d6fcd |
| Hosted Hugging Face inference | Disabled; download and use a compatible local runtime |
| License | Apache License 2.0, following the direct source's declared license |
The source configuration identifies a native vision-language model with English and Chinese metadata, a vision tower, and an MTP head. Actual text, image, reasoning, tool-use, and long-context behavior depends on the selected artifact, runtime, client, prompt format, memory, and application integration.
Included artifacts
All release choices are organized in one repository:
| Directory / variant | Artifact(s) | Lineage and purpose |
|---|---|---|
| SOURCE_BF16/ | Source files under their upstream filenames | Byte-for-byte mirror of every file in the pinned source revision, excluding only local Hugging Face cache metadata. This includes the source's native MTP shard and preserves the source README verbatim at SOURCE_BF16/README.md. |
| BF16/ | Split main-model BF16 GGUF beginning with BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00001-of-00002.gguf | Converted directly from the pinned BF16 Safetensors with MTP excluded. Keep every split shard together. This is the common parent used independently for all four quantized variants. |
| AUX/ | AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf | Matching BF16 multimodal projector converted from the same pinned source revision. |
| Repository root — Q8_0 | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf | Independently quantized from the main-model BF16 GGUF parent. |
| Repository root — Q6_K | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf | Independently quantized from the main-model BF16 GGUF parent. |
| Repository root — Q5_K_M | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf | Independently quantized from the main-model BF16 GGUF parent. |
| Repository root — Q4_K_M | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf | Independently quantized from the main-model BF16 GGUF parent. |
| Variant | File(s) | Total size | SHA-256 |
|---|---|---:|---|
| BF16 | BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00001-of-00002.gguf<br>BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00002-of-00002.gguf | 53,808,282,880 bytes (53.81 GB / 50.11 GiB) | c16f2f4fe739091e1d3569b281cd8b2fdf417eaef415bca7dd5858a092a43efb<br>0f54cab243f93f3bdd1d8ffe8bbda84e7559d8c3e2e5d503ae9ccf1ff15f376e |
| Q8_0 | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf | 28,595,764,288 bytes (28.60 GB / 26.63 GiB) | 2569a2791d743fc4b61a05c0b7bb1e8a43097423f2cfbca0750c4849014adaef |
| Q6_K | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf | 22,082,530,368 bytes (22.08 GB / 20.57 GiB) | 8d619826302b87c49f2d9452a2566c82900e7a60c247b44512fe01955daaa8ca |
| Q5_K_M | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf | 19,231,099,968 bytes (19.23 GB / 17.91 GiB) | f3a7444f1467684898dd40f526ea2886fd993a6d5a04fa4798c325a3f80af2b7 |
| Q4_K_M | Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf | 16,547,400,768 bytes (16.55 GB / 15.41 GiB) | 13341c92f11bb482c109d23da083ee165b1ba6e58927b2afb76ac2668262217f |
| AUX | AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf | 931,146,656 bytes (0.93 GB / 0.87 GiB) | 1ff0cd30ecb0d70c30b0f92d73deb60eedf2f4a2969f64d3c711db2d2bc76305 |
The byte-identical source mirror contains 13 files totaling 55,583,138,958 bytes. MANIFEST.json and SHA256SUMS.txt are the authoritative records for final paths, exact byte sizes, and SHA-256 hashes.
Quantization lineage
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
@ a6775a9a8ebb65cab3f707b4ab087fc7aa698634
├── SOURCE_BF16/ byte-identical mirror
│ └── native source MTP shard preserved unchanged
└── split main-model BF16 GGUF (direct conversion; MTP excluded)
├── Q8_0 (independent BF16-parented quant)
├── Q6_K (independent BF16-parented quant)
├── Q5_K_M (independent BF16-parented quant)
└── Q4_K_M (independent BF16-parented quant)
No quant was requantized from another quantized GGUF. No importance matrix, IQ recipe, dynamic-quant recipe, or MTP/NextN GGUF was produced. The matching vision tower is exported separately as the BF16 projector in AUX/. Conversion and quantization use llama.cpp commit 0d9ceae1e38291035605613ab41a8f5e693d6fcd.
Which file should I choose?
This is qualitative format and memory guidance, not a benchmark result:
| Choice | Practical guidance |
|---|---|
| Q4_K_M | Recommended default. Start here for the most practical balance of model size and local usability. |
| Q5_K_M | Quality-oriented default when memory allows. It retains more weight precision than Q4_K_M while remaining materially smaller than the heavier choices. |
| Q6_K | Choose when you can afford a larger file and want less weight compression than the five-bit option. |
| Q8_0 | Largest quantized choice in this release. Use when memory is ample and minimizing additional weight quantization error matters more than size. |
| BF16 | Direct, widest-precision GGUF representation in this release and the parent of all four quants. It is very large and split across multiple files. |
| SOURCE_BF16 | Byte-identical source-format mirror for preservation or source-runtime use; it is not a GGUF variant. It retains the native source MTP shard. |
Memory use is not just the model-file size. Context length, KV-cache precision, GPU offload, batch size, and multimodal inputs can materially increase RAM or VRAM requirements. Confirm fit in your own runtime.
Download
Install the Hugging Face CLI and download only the files you need.
For the root-level Q5_K_M model plus the matching projector:
hf download chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
--include "Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf" \
--include "AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf" \
--local-dir Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF
For BF16, download all split model shards and the projector:
hf download chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
--include "BF16/*" \
--include "AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf" \
--local-dir Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF
To obtain the byte-identical source-format mirror instead:
hf download chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
--include "SOURCE_BF16/*" \
--local-dir Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF
Run with llama.cpp
The release is converted, quantized, and release-tested with llama.cpp commit 0d9ceae1e38291035605613ab41a8f5e693d6fcd. Use that pin or a verified compatible build with support for this architecture and its multimodal projector.
Root-level Q5_K_M with the matching projector
llama-cli \
-m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf \
--mmproj "AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf" \
--image /path/to/image.jpg \
-p "Describe this image." \
-n 256
For text-only use, omit --image; whether the projector is needed for text-only loading depends on the runtime version and invocation path.
Split BF16 first shard with the matching projector
Keep all BF16 shards together in the same directory and pass only the first shard to llama.cpp. Do not pass a later shard directly, rename the shards, or separate them.
llama-cli \
-m "BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00001-of-00002.gguf" \
--mmproj "AUX/mmproj-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.gguf" \
--image /path/to/image.jpg \
-p "Describe this image." \
-n 256
llama.cpp discovers the remaining BF16 shards from the first shard when every shard remains together under BF16/.
The main GGUF, matching projector, compatible llama.cpp runtime, and client-side image-input support are separate requirements. A projector file alone does not guarantee that every llama.cpp build or frontend supports this architecture or every source modality. Do not substitute a projector from another model or revision.
Reproducibility and integrity
The release workflow records and checks:
- the exact direct source repository and immutable source revision
a6775a9a8ebb65cab3f707b4ab087fc7aa698634; - every pinned source filename, byte size, and SHA-256 hash;
- exact agreement between the Safetensors index and shard headers;
- BF16 source tensor dtypes, with native MTP tensors and vision tensors identified separately by tensor name rather than inferred from shard numbering;
- preservation of the complete source tree, including the native MTP shard and upstream README, under
SOURCE_BF16/; - direct main-model BF16 conversion with MTP excluded and a separately exported matching BF16 projector;
- independent BF16-to-
Q8_0,Q6_K,Q5_K_M, andQ4_K_Mlineage withimatrix=none; - the pinned llama.cpp commit
0d9ceae1e38291035605613ab41a8f5e693d6fcdand exact conversion and quantization commands; - GGUF metadata, tensor types, split-file consistency, and the projector/model pairing;
- a short bounded load-and-generate check for every complete main-model GGUF with the pinned runtime before release;
- final local artifact paths, exact byte sizes, and SHA-256 hashes in
MANIFEST.jsonandSHA256SUMS.txt; - post-publication comparison of public Hub inventory, byte sizes, and LFS/Xet SHA-256 metadata with the local release records.
These are provenance, integrity, structure, and basic loadability checks. They are not a benchmark of quality, safety, coding, tool use, reasoning, long-context behavior, or multimodal accuracy.
Evaluation status
No new benchmark numbers are reported for these GGUF artifacts. The quantitative and behavioral statements in the attributed blockquote are reproduced from the direct source model card and have not been independently rerun by the GGUF packager. Quantization can change model behavior; upstream measurements should not be assumed to transfer unchanged to BF16 GGUF or any quantized variant.
Intended use
These files are intended for local llama.cpp experimentation, compatibility testing, controlled research, source preservation, and applications whose operators can validate outputs and supply safeguards appropriate to their risk profile. Choose a variant based on available memory and tolerance for additional quantization.
The source describes an abliterated, uncensored model. Users should implement application-level validation, moderation, access controls, auditability, and human review appropriate to their deployment, and must ensure that use complies with applicable law and policy.
Provenance
The exact model lineage is:
Qwen/Qwen3.8-27B, identified by the direct source as the original base model.AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16, the exact direct BF16 source for this release.- This repository's source mirror, direct BF16 GGUF conversion, matching projector, and independently BF16-parented standard quant variants.
The direct source card attributes its model modifications to AEON-7 and describes SSM conv1d outlier repair, abliteration, restoration of the stock MTP head after the edit pipeline, and an unchanged vision tower. Those are upstream provenance statements. This GGUF packager does not claim to have authored, trained, fine-tuned, merged, abliterated, repaired, benchmarked, or otherwise behaviorally modified the source model.
This repository's modifications are limited to:
- organizing the pinned byte-identical source mirror under
SOURCE_BF16/without renaming upstream files; - converting the source's main BF16 tensors to split BF16 GGUF with MTP excluded;
- exporting the matching vision tower as a BF16 projector;
- independently quantizing the common BF16 GGUF parent to
Q8_0,Q6_K,Q5_K_M, andQ4_K_Mwithout an importance matrix; and - recording release manifests, checksums, reproducibility evidence, and basic runtime checks.
The direct source README is preserved byte-for-byte at SOURCE_BF16/README.md. It contains the upstream publisher's full model description, provenance, user-responsibility language, warranty disclaimer, and arbitration clause. Users should review that preserved source document directly; this card does not restate or independently interpret all of its terms.
Limitations and risks
- Additional quantization:
Q8_0,Q6_K,Q5_K_M, andQ4_K_Meach introduce quantization error relative to the direct BF16 GGUF parent. No GGUF quality benchmark is claimed here. - No importance matrix: the K-quants were produced without an imatrix. No claim is made that they match or outperform imatrix-assisted alternatives.
- No alternative quant recipes: no IQ or dynamic quantization recipe is included or evaluated.
- No MTP GGUF: the main-model conversion excludes MTP, and no separate MTP/NextN speculative-draft GGUF is included. The native MTP shard remains available only in the byte-identical source mirror for compatible source-format runtimes.
- Safety and alignment: the direct source describes the model as abliterated, uncensored, and refusal-removed. This release does not restore safety alignment or independently validate that behavior. Outputs may be unsafe, false, biased, offensive, or illegal to act on.
- Multimodality: image use requires the exact matching projector plus compatible runtime and client support. Multimodal quality was not newly evaluated for this release.
- Coding and tools: generated code, arguments, and tool calls require validation before execution. No coding or tool-use benchmark is reported for these GGUFs.
- Context and memory: 262,144 is source configuration metadata, not a promise that a local system can run that context. KV cache, multimodal inputs, batching, and runtime overhead can require substantial memory beyond the model-file size.
- Split BF16: all BF16 shards must remain together and retain their published filenames; pass only
BF16/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16-00001-of-00002.ggufto llama.cpp. - Runtime compatibility: this architecture is recent. Older llama.cpp builds and third-party frontends may fail to load it or may not expose every feature.
- Hosted inference: Hugging Face hosted inference is disabled for this artifact repository.
- Upstream claims: source behavioral, evaluation, and modification claims remain upstream claims unless specifically confirmed by release-integrity evidence. This packaging release does not independently reproduce them.
License and attribution
The direct source model card declares the Apache License 2.0, inherited from Qwen/Qwen3.8-27B. At the inspected pinned source revision, the source card supplied the license declaration but the source repository did not include a standalone LICENSE file. This release includes the complete standard Apache License 2.0 text at repository root and preserves the upstream source card under SOURCE_BF16/README.md.
Attribution:
- Qwen / Alibaba for
Qwen/Qwen3.8-27B, the original base identified by the direct source. - AEON-7 for
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16, the exact direct BF16 source and source of the upstream modification and evaluation claims. ggml-org/llama.cppfor GGUF conversion, quantization, and runtime tooling.
Review the Apache 2.0 terms and preserve required notices and attribution when redistributing. The source README contains additional publisher-supplied language that users should read directly. This repository does not imply endorsement by AEON-7, Qwen, Alibaba, or the llama.cpp project.
Run chimingw/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models