GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

lambsea/Qwen3.8-27B-AEON-Ultimate-Uncensored-UD-GGUF overview

Qwen3.8 27B AEON Ultimate Uncensored — GGUF UD Quants Unsloth Dynamic style UD GGUF quantizations of AEON 7/Qwen3.8 27B AEON ULTIMATE UNCENSORED BF16 https://h…

ggufquantizedqwen3qwen3.8hybridssmgated-delta-netunsloth-dynamicimatrixmtpvisionawqtext-generationenzhbase_model:AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16base_model:quantized:AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~884.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
1,333
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

9 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-27B-AEON-AWQ-UD-IQ4_XS.ggufGGUFIQ4_XS25.94 GBDownload
Qwen3.8-27B-AEON-AWQ-UD-Q5_K_M.ggufGGUFQ5_K_M28.69 GBDownload
Qwen3.8-27B-AEON-AWQ-UD-Q6_K.ggufGGUFQ6_K30.58 GBDownload
Qwen3.8-27B-AEON-AWQ-UD-Q8_0.ggufGGUFQ8_034.75 GBDownload
Qwen3.8-27B-AEON-UD-IQ4_XS.ggufGGUFIQ4_XS25.94 GBDownload
Qwen3.8-27B-AEON-UD-Q5_K_M.ggufGGUFQ5_K_M28.69 GBDownload
Qwen3.8-27B-AEON-UD-Q6_K.ggufGGUFQ6_K30.58 GBDownload
Qwen3.8-27B-AEON-UD-Q8_0.ggufGGUFQ8_034.75 GBDownload
Qwen3.8-27B-AEON-mmproj-F16.ggufGGUFF16884.6 MBDownload

Model Details

Model IDlambsea/Qwen3.8-27B-AEON-Ultimate-Uncensored-UD-GGUF
Authorlambsea
Pipelinetext-generation
Licenseapache-2.0
Base modelAEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
Last modified2026-09-04T06:28:02.000Z

Model README

---

license: apache-2.0

language:

- en

- zh

base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16

tags:

- gguf

- quantized

- qwen3

- qwen3.8

- hybrid

- ssm

- gated-delta-net

- unsloth-dynamic

- imatrix

- mtp

- vision

- awq

model_name: Qwen3.8-27B-AEON-Ultimate-Uncensored-UD-GGUF

pipeline_tag: text-generation

---

Qwen3.8-27B-AEON-Ultimate-Uncensored — GGUF (UD Quants)

Unsloth Dynamic-style (UD) GGUF quantizations of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16

Two quant families are available:

  • AWQ-UD (recommended): AWQ channel pre-scaling applied before quantization. Lower perplexity than baseline at every bit width. All AWQ-UD quants beat F16 on perplexity.
  • Baseline UD: Standard quantization without pre-scaling.

Every quant uses per-tensor overrides (sensitivity-driven) + importance matrix (multi-domain calibration). All SSM recurrence tensors are preserved at source precision. MTP speculative decoding and vision (mmproj) are preserved.

---

Quant Comparison

AWQ-UD (recommended)

AWQ pre-scaling (256 samples x 512 tokens, W4A16_ASYM target) redistributes weight magnitudes before quantization. All AWQ-UD quants beat F16 reference perplexity.

| File | Quant | Size | PPL | KL mean | tg t/s |

|------|-------|------|-----|---------|--------|

| F16 (reference) | F16 | 50.9 GB | 2.8950 | — | 30.3 |

| AWQ-UD-Q8_0 | Q8_0 | 34.7 GB | 2.8807 | 0.00384 | 42.4 |

| AWQ-UD-Q6_K | Q6_K | 30.6 GB | 2.8795 | 0.00376 | 47.2 |

| AWQ-UD-Q5_K_M | Q5_K_M | 28.7 GB | 2.8672 | 0.00945 | 49.0 |

| AWQ-UD-IQ4_XS | IQ4_XS | 25.9 GB | 2.8887 | 0.01922 | 53.2 |

Baseline UD

| File | Quant | Size | PPL | KL mean | tg t/s |

|------|-------|------|-----|---------|--------|

| UD-Q8_0 | Q8_0 | 34.7 GB | 2.8918 | 0.00181 | 42.4 |

| UD-Q6_K | Q6_K | 30.6 GB | 2.8891 | 0.00153 | 47.2 |

| UD-Q5_K_M | Q5_K_M | 28.7 GB | 2.8856 | 0.00822 | 49.0 |

| UD-IQ4_XS | IQ4_XS | 25.9 GB | 2.8990 | 0.01791 | 53.2 |

Benchmarked on NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), llama.cpp (a4501150/llama.cpp), pp=512, tg=128. Throughput is identical between AWQ-UD and baseline UD at the same bit width because AWQ does not change tensor sizes.

> AWQ-UD vs baseline: AWQ wins on perplexity (0.010-0.018 lower PPL). Baseline wins on KL divergence (closer to F16 output distribution). AWQ changes the channel basis, which shifts the output distribution away from F16, but absolute quality improves.

---

SGLang Throughput (NVFP4 + DFlash2)

For maximum throughput, serve the NVFP4 checkpoint via SGLang with DFlash2 speculative decoding:

| Config | Single-user t/s | 3 users (agg) | 6 users (agg) |

|--------|----------------|---------------|---------------|

| llama.cpp NVFP4 | 64 | — | — |

| llama.cpp UD-Q6_K + DSpark | 72 | — | 178 (5 users) |

| SGLang NVFP4 + DFlash2 | 146-178 | 391 | 493 |

SGLang is 2.3-2.8x faster single-user and 2.8x faster concurrent vs llama.cpp.

---

What Makes These Different

AWQ Pre-Scaling (AWQ-UD only)

AWQ (Activation-Aware Weight Quantization) applies per-channel scaling to redistribute weight magnitudes before quantization. This makes outlier channels less damaging when quantized. The scaling is applied at BF16 precision and is lossless — the model produces identical output before quantization. After scaling, the full GGUF pipeline runs: convert, importance matrix, sensitivity analysis, quantize with per-tensor overrides.

SSM Recurrence Preservation

Qwen3.8 is a hybrid GatedDeltaNet + attention model. 48 of 64 layers use a recurrent SSM where quantization error compounds across token positions. All SSM recurrence tensors are preserved at source precision (F16) — never quantized.

| Tensor | Count | Precision | Rationale |

|--------|-------|-----------|-----------|

| ssm_alpha, ssm_beta | 96 | F16 | State update projections — error accumulates in recurrence |

| ssm_out | 48 | F16 | Output projection feeds directly into residual stream |

| ssm_a, ssm_conv1d, ssm_dt, ssm_norm | 192 | F32 | Small state tensors (llama-quantize keeps 1D/small tensors at F32) |

| attn_qkv (SSM input projection) | 48 | F16 | Highest measured KL sensitivity |

| attn_gate (SSM gate projection) | 48 | F16 | Second-highest measured KL sensitivity |

Per-Tensor Sensitivity Analysis

Each tensor group was probed by quantizing only that group to Q4_0 while keeping the rest at F16, then measuring KL divergence. The override generator assigns precision based on measured sensitivity:

| Precision | Tensor Groups | Override Count |

|-----------|--------------|----------------|

| F16 | SSM recurrence, norms, biases, MTP layer | 512 |

| F16 | All attention tensors (attn_qkv, attn_gate, attn_v, attn_q, attn_k, attn_output), ffn_down edge | 173 |

| Base quant | FFN middle layers, FFN edge gate/up, embeddings | ~181 |

685 total overrides — every non-FFN tensor has an explicit precision assignment. No dependence on llama-quantize's internal promotion rules.

Multi-Domain Calibration + GPU Imatrix

Calibrated on a balanced mix across 4 domains from 13 HF datasets:

| Domain | Token Budget | Sources |

|--------|-------------|---------|

| General | 1M | ultrachat, OpenHermes, COIG-CQIA, LongAlpaca, pg19, froggeric/imatrix |

| Code | 750K | Magicoder-Evol-Instruct-110K |

| Reasoning | 750K | OpenMathInstruct-2, OpenR1-Math-220k |

| Agentic | 500K | glaive-function-calling-v2, xlam-function-calling-60k, hermes-function-calling-v1 |

Special tokens from source datasets are stripped automatically. Samples are kept whole — never truncated mid-conversation.

The importance matrix is generated with a PyTorch GPU-native generator at 32,768 context — uses forward hooks to accumulate squared activations on GPU with zero PCIe D2H copies. Supports multi-GPU via device_map="auto".

Per-domain imatrices are merged with equal weights (DI-MATRIX approach).

MTP + Vision Preserved

  • MTP (Multi-Token Prediction): Draft head (blk.64) pinned at F16. Use --spec-type draft-mtp --spec-draft-n-max 3 for ~1.5-2x faster generation.
  • Vision: mmproj file contains the full vision encoder. Use --mmproj flag with llama-server for image/video understanding.

---

Files

| File | Description | Size |

|------|-------------|------|

| Qwen3.8-27B-AEON-AWQ-UD-Q8_0.gguf | AWQ — Highest quality | 34.7 GB |

| Qwen3.8-27B-AEON-AWQ-UD-Q6_K.gguf | AWQ — Recommended — best quality/size | 30.6 GB |

| Qwen3.8-27B-AEON-AWQ-UD-Q5_K_M.gguf | AWQ — Balanced | 28.7 GB |

| Qwen3.8-27B-AEON-AWQ-UD-IQ4_XS.gguf | AWQ — Smallest | 25.9 GB |

| Qwen3.8-27B-AEON-UD-Q8_0.gguf | Baseline — Highest quality | 34.7 GB |

| Qwen3.8-27B-AEON-UD-Q6_K.gguf | Baseline — Best quality/size | 30.6 GB |

| Qwen3.8-27B-AEON-UD-Q5_K_M.gguf | Baseline — Balanced | 28.7 GB |

| Qwen3.8-27B-AEON-UD-IQ4_XS.gguf | Baseline — Smallest | 25.9 GB |

| Qwen3.8-27B-AEON-mmproj-F16.gguf | Vision encoder (use with --mmproj) | 885 MB |

| Qwen3.8-27B-sharp.jinja | Enhanced chat template (terse output, reasoning effort, tool error detection) | 18 KB |

| imatrix_merged.dat | Importance matrix for requantization | 13 MB |

---

Usage

llama-server (recommended)

# AWQ-UD-Q6_K with sharp template, MTP + vision
llama-server \
    -m Qwen3.8-27B-AEON-AWQ-UD-Q6_K.gguf \
    --mmproj Qwen3.8-27B-AEON-mmproj-F16.gguf \
    -ngl 99 \
    --flash-attn \
    -c 262144 \
    --parallel 3 \
    -kvu \
    --jinja \
    --chat-template-file Qwen3.8-27B-sharp.jinja \
    --reasoning-format deepseek \
    --reasoning-preserve \
    --spec-type draft-mtp \
    --spec-draft-n-max 3 \
    --host 0.0.0.0 --port 8080

> Note: --spec-type draft-mtp requires llama.cpp b9375+. --reasoning-format deepseek extracts thinking into message.reasoning_content in API responses. The sharp template (by froggeric) enables thinking by default with terse output, reasoning effort control, and tool call error detection.

llama-cli

llama-cli \
    -m Qwen3.8-27B-AEON-AWQ-UD-Q6_K.gguf \
    -ngl 99 \
    --flash-attn \
    -c 262144 \
    --jinja \
    --chat-template-file Qwen3.8-27B-sharp.jinja \
    --reasoning on \
    --reasoning-preserve

---

Sampling Parameters

From the official Qwen3.8-27B model card:

| Mode | temperature | top_p | top_k | min_p | presence_penalty | repetition_penalty |

|------|-------------|-------|-------|-------|------------------|--------------------|

| Thinking (default) | 1.0 | 0.95 | 20 | 0.0 | 0.0 | 1.0 |

| Non-thinking | 0.7 | 0.80 | 20 | 0.0 | 1.5 | 1.0 |

Do not use greedy decoding (temperature=0). reasoning_effort controls thinking depth independently: xhigh (default), medium, low.

---

Architecture

Qwen3.8-27B is a hybrid SSM-attention model:

  • 64 transformer layers + 1 MTP layer (blk.0-64)
  • 48 SSM layers (GatedDeltaNet, no KV cache) + 16 full attention layers (every 4th layer)
  • 27B parameters, 24 attention heads, 4 KV heads, head dim 256
  • Vocab: 248,320 tokens, native context: 262,144 tokens

---

Quantization Pipeline

Built with super-quant:

  1. AWQ pre-scaling (AWQ-UD only): Apply per-channel weight scaling via llm-compressor AWQModifier (256 calibration samples, W4A16_ASYM target). Strip all quantization state. Save as plain BF16.
  2. Convert HF to F16 GGUF (with MTP tensors) + mmproj GGUF (vision)
  3. Multi-domain calibration data from 13 HF datasets, special tokens stripped
  4. GPU-native importance matrix generation (PyTorch, 32k context) + weighted merge
  5. Per-tensor sensitivity analysis (KL divergence probing against F16 logits)
  6. Hybrid override generation — SSM at source precision, sensitivity-driven for the rest
  7. Quantize with per-tensor overrides + imatrix
  8. Benchmark: throughput + perplexity + KL divergence vs F16

---

Links

Credits

---

License: Apache-2.0

Run lambsea/Qwen3.8-27B-AEON-Ultimate-Uncensored-UD-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models