vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF overview
Qwen3.8 27B AEON ULTIMATE — ROCmFPX iMatrix GGUF ROCmFPX iMatrix quantizations of the full BF16 checkpoint published by Aeon / AEON 7 https://huggingface.co/AE…
Runs locally from ~13.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf | GGUF | GGUF | 16.52 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP6-iMatrix.gguf | GGUF | GGUF | 20.98 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP8-iMatrix.gguf | GGUF | GGUF | 26.26 GB | Download |
| Qwen3.8-27B-AEON-ULTIMATE-iMatrix.imatrix.gguf | GGUF | GGUF | 13.0 MB | Download |
Model Details
| Model ID | vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF |
|---|---|
| Author | vmlinux |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 |
| Last modified | 2026-08-17T22:54:28.000Z |
Model README
---
license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- rocm
- rocmfpx
- amd
- strix-halo
- qwen3.8
- qwen35
- imatrix
- thinking
- uncensored
---
Qwen3.8-27B AEON ULTIMATE — ROCmFPX iMatrix GGUF
ROCmFPX iMatrix quantizations of the full BF16 checkpoint published by
Attribution
The source weights are Aeon's Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16,
pinned at Hub revision
8f76e82ed7ef4de7735f5d4148fce7b643b00fae. Aeon deserves attribution for the
BF16 model and its model work. This repository contains derived GGUF
quantizations produced by vmlinux with the ROCmFPX toolchain; it is not a new
training run or a claim of ownership of the source model.
The source model declares Apache 2.0 licensing. Review the source model card
and applicable terms before redistribution or deployment.
Quantizations
| File | Preset | Size | SHA-256 |
| --- | --- | ---: | --- |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf | Q4_0_ROCMFP4 | 17,735,469,440 bytes | 34c04aaec2399fab8179a921d0ea1e9d3b93da213201ab0ec2930ee1420307b0 |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP6-iMatrix.gguf | Q6_0_ROCMFPX | 22,528,383,360 bytes | d096fdd0ece150cefa793a8fdae55076b63d61d9e135571de6c73b016a641ba1 |
| Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP8-iMatrix.gguf | Q8_0_ROCMFPX | 28,193,397,120 bytes | 10eff5653c51e23a8a816761feea69b235c9003e80821759ecef41814843a761 |
All three use the same model-specific importance matrix:
Qwen3.8-27B-AEON-ULTIMATE-iMatrix.imatrix.gguf. The matrix was generated
from 339 chunks of 512 tokens using the shared calibration corpus, and each
quantizer consumed 496 entries.
Runtime
These are experimental ROCmFPX tensor types and require a compatible
ROCmFPX-enabled llama.cpp build. Stock upstream llama.cpp will not load them.
Example ROCm0 invocation:
hf download vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF \
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
--local-dir ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX
./llama-completion \
-m ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
-dev ROCm0 -ngl all -c 8192 -n 256 -p "Hello."
The GGUFs retain the native one-layer MTP head. Thinking is enabled by the
embedded Qwen template by default; pass the appropriate chat-template kwargs
when an application needs thinking disabled.
Validation and provenance
All three files loaded and generated a short completion on ROCm0 with all
layers offloaded. Detailed public build information is in
BUILD_RESULTS.md, with exact hashes in
SHA256SUMS and source/toolchain details in
Run vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models