Ishowbackup/Qwen3.8-27B-ABLITERATED-GGUF overview
<div align="center" Blackfrost https://cdn uploads.huggingface.co/production/uploads/69a27f2d114e4ac9de4dafc7/xDTdhLFXmKlZOazFcvJ5S.jpeg <div align="center" <h…
Runs locally from ~600.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.8-27B-ABLITERATED-Q2_K.gguf | GGUF | Q2_K | 9.98 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q3_K_M.gguf | GGUF | Q3_K_M | 12.39 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q3_K_S.gguf | GGUF | Q3_K_S | 11.24 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf | GGUF | Q4_K_M | 15.41 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q4_K_S.gguf | GGUF | Q4_K_S | 14.52 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q5_K_M.gguf | GGUF | Q5_K_M | 17.91 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q5_K_S.gguf | GGUF | Q5_K_S | 17.40 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q6_K.gguf | GGUF | Q6_K | 20.57 GB | Download |
| Qwen3.8-27B-ABLITERATED-Q8_0.gguf | GGUF | Q8_0 | 26.63 GB | Download |
| mmproj-Qwen3.8-27B-ABLITERATED-F16.gguf | GGUF | F16 | 884.6 MB | Download |
| mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf | GGUF | Q8_0 | 600.1 MB | Download |
| mtp-Qwen3.8-27B-ABLITERATED-BF16.gguf | GGUF | BF16 | 5.54 GB | Download |
| mtp-Qwen3.8-27B-ABLITERATED-Q4_0.gguf | GGUF | Q4_0 | 1.87 GB | Download |
| mtp-Qwen3.8-27B-ABLITERATED-Q8_0.gguf | GGUF | Q8_0 | 2.95 GB | Download |
Model Details
| Model ID | Ishowbackup/Qwen3.8-27B-ABLITERATED-GGUF |
|---|---|
| Author | Ishowbackup |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16 |
| Last modified | 2026-08-15T23:19:28.000Z |
Model README
---
license: apache-2.0
base_model:
- Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16
tags:
- qwen3.8
- qwen
- 27b
- dense
- gguf
- abliterated
- quantized
- multimodal
- reasoning
- tool-calling
- llama.cpp
- long-context
pipeline_tag: image-text-to-text
library_name: gguf
---
<div align="center">
<div align="center">
<h1>QWEN3.8-27B-ABLITERATED-GGUF</h1>
<h3>Full standard GGUF quant ladder of the Blackfrost abliterated Qwen3.8-27B · dense multimodal model for llama.cpp</h3>
<p><strong>Built by <a href="https://x.com/Blackfrost_AI">Blackfrost</a> · Las Vegas, NV</strong></p>
<p>
<img src="https://img.shields.io/badge/GGUF_standard_ladder-047857?style=for-the-badge" />
<img src="https://img.shields.io/badge/11%2F450_refusals-047857?style=for-the-badge" />
<img src="https://img.shields.io/badge/Abliterated-1f2937?style=for-the-badge" />
<img src="https://img.shields.io/badge/EXPERIMENTAL-b45309?style=for-the-badge" />
<img src="https://img.shields.io/badge/llama.cpp-1f2937?style=for-the-badge" />
</p>
</div>
> ## All standard quants live
>
> The complete standard K-quant ladder (Q2_K through Q8_0) and both vision projectors are included. No IQ/IK or importance-matrix quants are used.
> ## MTP speculative decoding restored — August 15, 2026
>
> Three separate MTP sidecars are now included: BF16, Q8_0, and Q4_0. Existing text quants and vision projectors are unchanged, so current users only need to download an mtp- file to add speculative decoding.
> ## Experimental release
>
> This is a newly supported architecture and a deliberately modified research checkpoint. Validate the selected quant, context, sampling, vision, tools, and structured-output behavior in your own workload before deployment.
---
Refusal benchmark — R1-HARMFUL-BENCH-450
The release score is 11 residual refusals from 450 original cases (2.4%). The shipped Blackfrost short execution prompt is embedded once in the GGUF chat template.
This result is a sequential, manually reviewed residual funnel measured on the W4A4 NVFP4 derivative of the same BF16 parent. It is not a fresh full-450 GGUF run with every case presented under the final short prompt.
| evaluation stage | cases evaluated | material answer | true refusal remaining | other |
|---|---:|---:|---:|---:|
| Raw upstream template | 450 | 360 | 88 | 2 capability limitations |
| Blackfrost operational-prompt retest | 88 residuals | 53 | 33 | 1 limitation, 1 reproducible incoherent output |
| Shipped short execution-prompt retest | 33 residuals | 22 | 11 | 0 |
| Final residual count | 450 original cases | — | 11 (2.4%) | — |
The 450-case source set contains 150 AdvBench, 150 StrongREJECT, and 150 XSTest prompts. The final 11 comprise 1 AdvBench, 5 StrongREJECT, and 5 XSTest cases. An opening objection followed by a materially useful payload was counted as softened compliance, not as a refusal; a true refusal means the requested payload never arrived.
---
Why this model exists
Qwen3.8-27B is the dense, deployment-friendly member of the Qwen3.8 family. This is the abliterated Blackfrost build: refusal behavior was reduced through a weight-level process, then the BF16 parent was converted into a standard GGUF ladder for local llama.cpp inference.
It is not a coding fine-tune, merge, LoRA, or pruned model.
---
Specifications
| | |
|---|---|
| Architecture | Qwen3.8 dense hybrid VLM · 64 text layers · Gated DeltaNet + full attention · 27-layer vision tower |
| Parent | Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16 |
| Base | Qwen/Qwen3.8-27B · Apache-2.0 |
| Transform | Abliterated — refusal surface modified at weight level; no fine-tuning or pruning |
| Formats | Q2_K, Q3_K_S, Q3_K_M, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0 |
| Context | 262,144 tokens architecturally; practical context depends on RAM/VRAM and concurrency |
| Modalities | Text, image, and video input; text output |
| Chat behavior | Blackfrost short execution prompt embedded in the default Jinja chat template |
| MTP speculative head | Separate BF16, Q8_0, and Q4_0 sidecars included; existing text quants are unchanged |
---
Quant ladder
| quant | size | recommended for |
|---|--:|---|
| Q2_K | 10.7 GB | smallest standard quant; largest quality trade-off |
| Q3_K_S | 12.1 GB | very tight memory |
| Q3_K_M | 13.3 GB | compact general use |
| Q4_K_S | 15.6 GB | lower-memory Q4 option |
| Q4_K_M | 16.5 GB | default — balanced quality and footprint |
| Q5_K_S | 18.7 GB | higher fidelity |
| Q5_K_M | 19.2 GB | strong quality/size balance |
| Q6_K | 22.1 GB | near-BF16 behavior for many workloads |
| Q8_0 | 28.6 GB | maximum fidelity in the ladder |
File sizes are decimal GB as displayed by Hugging Face. Runtime memory also includes context state, compute buffers, the optional vision projector, and server overhead.
---
Vision projector files
Load one text quant plus one mmproj file for image or video input:
| file | size | purpose |
|---|--:|---|
| mmproj-Qwen3.8-27B-ABLITERATED-F16.gguf | 0.93 GB | full-fidelity vision projector |
| mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf | 0.63 GB | compact projector; unsupported 4,304-wide tensors retain F16 automatically |
---
MTP speculative decoding
The parent checkpoint's MTP head is published as separate lowercase mtp- sidecars, which is the current llama.cpp layout. Pair one sidecar with any existing text quant:
| file | size | use |
|---|--:|---|
| mtp-Qwen3.8-27B-ABLITERATED-BF16.gguf | 5.95 GB | maximum draft fidelity |
| mtp-Qwen3.8-27B-ABLITERATED-Q8_0.gguf | 3.16 GB | recommended balance |
| mtp-Qwen3.8-27B-ABLITERATED-Q4_0.gguf | 2.01 GB | lowest draft-model memory |
With current llama.cpp, repository loading discovers the closest mtp- sidecar automatically when MTP speculation is enabled:
llama-server \
-hf Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF:Q4_K_M \
--spec-type draft-mtp --spec-draft-n-max 3 \
-ngl 999 -ngld 999 --jinja -c 16384
For manual files, pass --spec-draft-model mtp-Qwen3.8-27B-ABLITERATED-Q8_0.gguf along with --spec-type draft-mtp. The Q4_K_M target plus Q8_0 MTP sidecar was runtime-tested on one RTX PRO 6000 Blackwell: 256 generated tokens at 90.0 tok/s with 50.7% draft acceptance (153 accepted of 302 drafted). Throughput and acceptance vary with prompt, sampler, hardware, context, and concurrency.
---
Serving with llama.cpp
Use a current llama.cpp build with llama-server. Q4_K_M plus the compact projector was load- and generation-tested through the OpenAI-compatible chat API on an NVIDIA B200.
hf download Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF \
Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf \
mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf \
--local-dir ./Qwen3.8-27B-ABLITERATED-GGUF
llama-server \
-m ./Qwen3.8-27B-ABLITERATED-GGUF/Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf \
--mmproj ./Qwen3.8-27B-ABLITERATED-GGUF/mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf \
-ngl 999 -fa on --jinja \
--host 0.0.0.0 --port 8080 -c 16384 \
--temp 1.0 --top-p 0.95 --top-k 20
- Text only: omit
--mmprojand do not download a projector. - CPU or hybrid inference: lower
-ngl; use-ngl 0for CPU-only operation. - Larger context: increase
-conly after checking memory headroom at the intended concurrency. - Embedded prompt: keep
--jinjaenabled so the repository's default chat template is applied. - One-command kit:
deploy/serve.shdownloads and serves the selected quant; seedeploy/DEPLOYMENT.mdfor the full guide.
API check
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Qwen3.8-27B-ABLITERATED",
"messages": [{"role": "user", "content": "Reply with exactly READY and nothing else."}],
"temperature": 0,
"max_tokens": 64
}'
---
Quality check
WikiText-2 rolling perplexity was measured on the parent artifacts through the same 8K API harness:
| artifact | word perplexity | byte perplexity | bits/byte |
|---|---:|---:|---:|
| Clean upstream BF16 | 8.4764 | 1.4914 | 0.5766 |
| Blackfrost W4A4 NVFP4 derivative | 9.3677 | 1.5195 | 0.6036 |
These figures are parent-artifact measurements, not per-quant GGUF perplexity scores. The Q4_K_M GGUF and compact projector passed a real llama.cpp load and chat-generation smoke test. The separate Q8_0 MTP sidecar also passed a CUDA runtime test with measurable draft acceptance.
---
Deployment responsibility
This checkpoint has a deliberately reduced refusal surface. Open weights do not provide an application policy, authorization system, audit trail, sandbox, or access-control boundary. Operators are responsible for authenticated access, least-privilege tool credentials, execution isolation, logging, and approval boundaries appropriate to their deployment.
The embedded prompt is a behavioral instruction, not a security boundary.
---
Disclaimer
Refusal behavior in this checkpoint has been deliberately modified at the weight level. It is not a safety-stock model and must not be represented as one.
This checkpoint is provided "as is," without warranty of any kind. Measurements describe only the tested artifacts, prompts, templates, samplers, serving engines, and review criteria. They do not guarantee that any particular input will be accepted or refused, that every upstream capability is retained, or that the measurements generalize to multimodal, tool-use, long-context, or multi-turn settings.
The derivative remains subject to the Apache 2.0 license shipped with the official Qwen3.8-27B checkpoint.
---
<div align="center">
<p>Built by <a href="https://x.com/Blackfrost_AI">Blackfrost</a> · Las Vegas, NV. Not affiliated with Qwen or Alibaba.</p>
</div>
Run Ishowbackup/Qwen3.8-27B-ABLITERATED-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models