esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF overview
Qwen3.8 27B TURBO Fable Cold Fusion 735 882 Heretic Uncensored NM DAU NVFP4 GGUF HF repos: esatapedico/Qwen3.8 27B TURBO Fable Cold Fusion 735 882 Heretic Unce…
Runs locally from ~14.11 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-COMPACT-LOW.gguf | GGUF | GGUF | 14.12 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGH.gguf | GGUF | GGUF | 16.60 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGHEST.gguf | GGUF | GGUF | 19.82 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-LOW.gguf | GGUF | GGUF | 14.67 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MEDIUM.gguf | GGUF | GGUF | 15.46 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MID-HIGH.gguf | GGUF | GGUF | 15.75 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-HIGH.gguf | GGUF | GGUF | 18.00 GB | Download |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-LOW.gguf | GGUF | GGUF | 14.11 GB | Download |
Model Details
| Model ID | esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF |
|---|---|
| Author | esatapedico |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU |
| Last modified | 2026-09-03T22:03:57.000Z |
Model README
---
license: apache-2.0
base_model: DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
pipeline_tag: text-generation
library_name: gguf
description: "Eight GGUFs of DavidAU's Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (TURBO Fable Cold Fusion Heretic Uncensored, 735/882): NVFP4 family with a per-tier quality ladder for lm_head / token_embd / MTP head. Seven tiers share a byte-identical 448-tensor native-NVFP4 backbone; HIGHEST is the fidelity-max tier (256 NVFP4 backbone plus Q8_0 extras, BF16 token_embd/MTP). MTP speculative decoding baked into every file. Pairs with the original Qwen3.8 vision projector. Blackwell sm_120."
tags:
- gguf
- nvfp4
- qwen3.8
- qwen3.5
- blackwell
- mtp
- speculative-decoding
- turbo
- fable
- cold-fusion
- llama.cpp
language:
- en
- multilingual
---
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF
HF repos: esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4 (safetensors, compressed-tensors NVFP4) and esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF (this GGUF family, 8 tiers). Budget no-MTP variants (BUDGET 14.72, STARVED 14.59 GB, Q3_K/Q2_K heads, 1,107 tensors, block_count 64) live in the companion repo esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-BUDGET-GGUF.
A family of eight GGUF files of DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU, DavidAU's TURBO 735-882 tune (Fable plus Cold Fusion plus Heretic/Uncensored, 735/882 Heretic merge and DPO layers) on Qwen3.8-27B: a 27B dense hybrid model (Gated DeltaNet plus Gated Attention every fourth layer, 262K native context, embedded MTP speculative head, native vision tower). NVFP4 here means W4A16 with FP8 scales, group size 16, weight only. Linear layers are NVFP4, vision tower, linear attention path, lm_head, embeddings and MTP head are kept in BF16 at the quantization source. The MTP head is baked into every file in this repo, no separate drafter is needed (--spec-type draft-mtp).
My part here is only the numerics: I converted the NVFP4 checkpoint to GGUF and built a size and precision ladder for the tensors that most affect output quality and decode speed. All credit for the model itself belongs upstream (full chain below).
Follow along & support
I post updates on new conversions, benchmarks, and what I'm working on over on Ko-fi. Follow along there to keep up with new releases and the work in progress. If you'd like to support more of it, a coffee is always welcome. I do this on consumer hardware and like seeing how far it goes. More is on the way.
☕ ko-fi.com/esatapedico. Updates, work-in-progress, and an optional coffee.
The eight files
Eight tiers in this repo share the MTP head (block_count 65, nextn_predict_layers 1, 1,122 tensors). Seven tiers share a byte-identical 448-tensor native NVFP4 backbone and differ only in lm_head / token_embd / MTP precision; HIGHEST is the fidelity-max tier that preserves more of the source precision (256 NVFP4 plus Q8_0 attention/SSM, BF16 token_embd and MTP) and is closest to ORIG:
| File | Size (decimal GB) | lm_head (output.weight) | token_embd | MTP head (blk.64) | Backbone |
|---|---|---|---|---|---|
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-LOW.gguf | 15.15 GB | Q3_K | Q2_K | Q2_K | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-COMPACT-LOW.gguf | 15.16 GB | Q4_K | Q3_K | Q2_K | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-LOW.gguf | 15.75 GB | Q5_0 | IQ4_XS | IQ4_XS | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MEDIUM.gguf | 16.60 GB | Q8_0 | Q6_K | IQ4_XS | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MID-HIGH.gguf | 16.91 GB | Q8_0 | Q8_0 | Q8_0 | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGH.gguf | 17.83 GB | BF16 | Q6_K | IQ4_XS | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-HIGH.gguf | 19.33 GB | BF16 | BF16 | BF16 | NVFP4 448 |
| Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGHEST.gguf | 21.28 GB | Q8_0 | BF16 | BF16 | 256 NVFP4 + Q8_0 |
Sizes are decimal GB as shown in the HF file browser. Picking a tier: MID-HIGH is the highest precision compact option among the 448-backbone tiers (all three head groups at Q8_0) and our expected fastest compact decode on dual GPU split; LOW and VERY-LOW trade some head precision for about 2 GB less VRAM; HIGH and VERY-HIGH restore BF16 heads where VRAM allows. VERY-HIGH needs about 18 GiB VRAM plus KV, so single 16 GB cards will fall back to CPU offload. HIGHEST is the fidelity-max tier (21.28 GB, 19.82 GiB reported, Q8_0 attention/SSM plus BF16 token_embd/MTP, closest to ORIG) and needs about 19.8 GiB VRAM plus KV; use dual-GPU split or 24 GB cards for full GPU.
Budget companions in the sibling repo: Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-BUDGET.gguf (14.72 GB, Q3_K head, Q2_K emb) and Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-STARVED.gguf (14.59 GB, Q2_K head and emb), both 1,107 tensors, block_count 64, no MTP, same 448 NVFP4 backbone with the draft block removed. Both fit a single 16 GB card.
Tensor layout
At a high level, the GGUFs in this repo contain 1,122 tensors. Seven tiers contribute a uniform 448-tensor native NVFP4 block shared byte identical across those seven tiers; HIGHEST contributes 256 NVFP4 plus Q8_0 attention/SSM extras (193 Q8_0, 9 BF16) and is not byte-identical to the 448 block but is the closest to ORIG. The NVFP4 format is W4A16 with group size 16 and FP8 E4M3 scales. Remaining tensors are F32 norms and scales, with lm_head, token embeddings and MTP head varying per tier as in the table above. Budget no-MTP variants in the companion repo have 1,107 tensors (block_count 64, nextn_predict_layers 0) and share the same 448-tensor NVFP4 backbone as the seven MTP tiers with the MTP block removed.
| Field | Value |
|---|---|
| Tensors per file | 1,122 |
| NVFP4 per tier | 448 (HIGHEST 256) |
| qwen35.block_count | 65 (64 plus 1 MTP blk.64, 15 tensors) |
| qwen35.context_length | 262144 |
| qwen35.nextn_predict_layers | 1 |
| tokenizer.chat_template | 8953 chars, intact |
| general.file_type | advisory only, per tensor type is authoritative |
| general.quantization_version | 2 |
| general.name | Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-<TIER> |
| general.description | NVFP4 <TIER> backbone. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU. |
| general.license | apache-2.0 in every file |
Vision
The TURBO 735-882 tune leaves the original Qwen3.8 vision tower untouched. Pair any tier with the matching vision projector mmproj-BF16.gguf via --mmproj. No mmproj is bundled in this repo, use the projector from the base model.
How this was made
At a high level, the steps were:
- Converted the BF16 base model to compressed tensors NVFP4 (W4A16, vision, linear attention, lm_head and MTP kept in BF16, no calibration).
- Verified NVFP4 metadata and ran a short vLLM smoke check.
- Converted the NVFP4 checkpoint to GGUF.
- Built each tier over a shared 448 tensor NVFP4 backbone, varying only lm_head, token embedding and MTP head precision per tier.
- Verified per tier NVFP4 tensor count and backbone byte identity across all tiers, and patched GGUF KV for name, description and license without changing tensor data.
Benchmarks (naive, single run, not comparable across setups)
These are rough sanity checks to confirm the files load and generate, not a formal benchmark. Method and hardware are noted so you can interpret them in context. Results will vary with hardware, sampling and context length.
All runs used direct GPU via llama.cpp tools, with a diverse synthetic English payload (about 75 kB, 28 chunks times 512 context). LocalAI gateway numbers are end to end and include gateway overhead, don't compare them to llama-bench decode.
Smoke: 8 of 8 coherent
Prompt What is a black hole? via single turn sampling, each tier produces a natural coherent completion with no repetition or truncation:
| Tier | Result |
|---|---|
| VERY-LOW | ok, coherent, 710 chars |
| COMPACT-LOW | ok, coherent, 892 chars |
| LOW | ok, coherent, 981 chars |
| MEDIUM | ok, coherent, 838 chars |
| MID-HIGH | ok, coherent, 978 chars |
| HIGH | ok, coherent, 829 chars |
| VERY-HIGH | ok, coherent, 1.1k chars |
| HIGHEST | ok, coherent, 1.0k chars |
Budget companions also smoke 2 of 2 pass (BUDGET and STARVED coherent on the same prompt).
Perplexity (PPL)
Direct llama-perplexity on the 75 kB diverse payload, single GPU, same chunks for all tiers:
| Tier | PPL (Final estimate) |
|---|---|
| VERY-LOW | 3.3410 +/- 0.07288 |
| COMPACT-LOW | 3.2808 +/- 0.07074 |
| LOW | 3.2761 +/- 0.07056 |
| MEDIUM | 3.2858 +/- 0.07109 |
| MID-HIGH | 3.2903 +/- 0.07107 |
| HIGH | 3.2851 +/- 0.07107 |
| VERY-HIGH | 3.2835 +/- 0.07089 |
| HIGHEST | 3.2367 +/- 0.07020 |
HIGHEST is best of the whole family at 3.2367, about 0.10 better than VERY-LOW and about 0.04 better than LOW. The six 448-backbone tiers span only 0.065 from best to worst, so the quantization costs almost nothing even at the smallest tier. VERY-LOW is highest as expected (Q3_K and Q2_K heads least precise). For reference, the sibling BUDGET repo measures BUDGET 3.3410 (identical to VERY-LOW, both Q3_K/Q2_K heads) and STARVED 3.4560 (+0.11, Q2_K everywhere).
Speed (llama-bench, pp512 prompt processing, tg128 generation, tok/s)
Single 5070 Ti (16 GB class) with full CUDA offload, three runs per tier:
| Tier | VRAM (llama-bench) | pp512 tok/s | tg128 tok/s | Notes |
|---|---|---|---|---|
| VERY-LOW | 14.10 GiB | 2748.05 +/- 263.70 | 47.30 +/- 0.04 | fits, full GPU |
| COMPACT-LOW | 14.11 GiB | 2437.61 +/- 343.88 | 47.49 +/- 0.08 | fits, full GPU |
| LOW | 14.66 GiB | 2728.03 +/- 264.88 | 46.77 +/- 0.08 | fits |
| MEDIUM | 15.45 GiB | 2704.70 +/- 300.87 | 45.55 +/- 0.04 | fits |
| MID-HIGH | 15.74 GiB | 2714.21 +/- 277.94 | 45.55 +/- 0.10 | fits |
| HIGH | 16.59 GiB | 34.89 +/- 0.05 | 6.32 +/- 0.01 | exceeds single 16 GB, CPU fallback, about 30 times slower on this bench |
| VERY-HIGH | 17.99 GiB | 34.44 +/- 0.05 | 6.12 +/- 0.12 | exceeds, CPU fallback |
| HIGHEST | 19.81 GiB | 35.40 +/- 0.03 | 4.73 +/- 0.00 | exceeds single 16.3 GiB, CPU fallback on single GPU, needs dual split or 24 GB |
The four compact tiers show flat prefill around 2700 tok/s and decode around 46 tok/s with less than 4 percent variance, consistent with an identical backbone and heads that are decode bound. HIGH, VERY-HIGH and HIGHEST exceed single 16 GB VRAM and fall back to partial CPU offload on this single GPU bench, on dual GPU split or 24 GB cards they run at full GPU speed. For reference, sibling BUDGET 13.70 GiB at 2753.10 +/- 249.27 pp512 and 47.51 +/- 0.06 tg128 fits single 5070 Ti, STARVED 13.58 GiB at 2762.68 +/- 235.81 and 48.05 +/- 0.13 also fits.
vLLM NVFP4 smoke
A short vLLM smoke check with tensor parallel 2 and 2048 context on the NVFP4 checkpoint passed with coherent output, confirming the checkpoint loads and generates before GGUF tier building.
LocalAI gateway check (naive, one-at-a-time)
Each tier was sent one large request with max_tokens=20000 and finish_reason=stop for all tiers. These are end to end gateway timings, not pure decode, so treat as sanity and responsiveness, not a formal benchmark.
| Tier | request_seconds | tokens per sec (20000 / s) | finish_reason | reasoning chars | content chars |
|---|---|---|---|---|---|
| VERY-LOW | 661.94 | 30.21 | stop | 1325 | 4941 |
| COMPACT-LOW | 392.98 | 50.89 | stop | 1392 | 9382 |
| LOW | 370.89 | 53.92 | stop | 795 | 2806 |
| MEDIUM | 374.34 | 53.43 | stop | 1464 | 2474 |
| MID-HIGH | 422.28 | 47.36 | stop | 3320 | 8399 |
| HIGH | 507.82 | 39.38 | stop | 2115 | 8685 |
| VERY-HIGH | 505.75 | 39.55 | stop | 1993 | 7551 |
| HIGHEST | 416.45 | 48.02 | stop | 3557 | 5056 |
Budget companions in the sibling repo on the same gateway method: BUDGET 178.84s, 111.83 tok/s, stop; STARVED 303.68s, 65.86 tok/s, stop. Both finished with finish_reason stop, no repetition collapse. Gateway overhead included, single run only, don't compare to llama-bench decode.
Tokens per sec is 20000 divided by request_seconds. It reflects the whole gateway round trip for this specific large prompt, not llama-bench decode.
All tiers had distinct 5-gram ratios near 1.0 and no adjacent duplication, so no repetition collapse.
What to take away, naive framing
- Single run only, one large synthetic prompt, gateway overhead included, no warmup average. Don't compare gateway tok/s to direct
llama-benchnumbers and don't compare across hardware. - LOW and MEDIUM were the fastest gateway round trips for this prompt at about 54 tok/s, MID-HIGH at 47 tok/s, HIGHEST at 48 tok/s, HIGH and VERY-HIGH slower at about 39 tok/s. VERY-LOW was the outlier at 30 tok/s with the longest wall time but still finished cleanly with stop. All produced coherent output with no truncation. BUDGET was fastest at 111 tok/s on its cycle for the same gateway prompt, STARVED at 65 tok/s.
SHA-256
c58f0a932957825cea5e6b73966ddf451f33b287668f670e1293a8ef132bf561 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-LOW.gguf
56f8d6f5c656ac96da20086c4e9e546c1b30d723e185386d50219b08b5e47f83 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-COMPACT-LOW.gguf
6b2d5f6d5795ceef5a9dcc18a444ef9d03dd47b1d3bffaf80dd392f5a6cc425b Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-LOW.gguf
1c9e7f3e77c6938fb0f1218ec0da9de93dcd7a83357e244d76e4ebff99e36058 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MEDIUM.gguf
4e1b87edfc2b7f58a78c657afe965344dea2415ae71264f93ecea1f94d4ac24f Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MID-HIGH.gguf
10f60fd0f0597e9d97641e6c77b637fc247e343fb39b84b145f17254dd588bc8 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGH.gguf
b08f4dcd1c5c48463c185746dd6ba594cae466c24c5e2a26628351dbadabc702 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-HIGH.gguf
c18d844039ccbd5cf5195c12937bf7043c26324a03de3c5b6ff3a821a6c68473 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGHEST.gguf
Attribution & provenance
This is a derivative work built entirely from existing Apache 2.0 artifacts. Nothing here was trained or fine tuned. Credit belongs to:
- Alibaba and Qwen team for the base model, Qwen/Qwen3.8-27B (Apache 2.0): 27B dense, 64 blocks, Gated DeltaNet plus Gated Attention hybrid, native vision language, 262,144 token context, MTP head.
- DavidAU for the tune itself, Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (Apache 2.0): the TURBO 735-882 Fable plus Cold Fusion plus Heretic/Uncensored DPO stack that this family converts.
- Unsloth, whose trainers and systems power the underlying training methods.
- This repo's author for the GGUF conversion and the tier ladder only.
Repository contents
- Eight tier GGUFs (table above, 15.15 to 21.28 GB) plus 8 override maps (
overrides-*.txt, 1,122 lines each: per-tensor target types, includesoutput.weight/token_embd.weightper tier). HIGHEST override pins attention/SSM toQ8_0and keeps token_embd/MTP atBF16. - Budget no-MTP companion repo:
esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-BUDGET-GGUF(BUDGET 14.72 GBQ3_K/Q2_K, STARVED 14.59 GBQ2_K/Q2_K, 1,107 tensors, no MTP) - Corresponding safetensors NVFP4 checkpoint (single
model.safetensors, compressed tensorsnvfp4-pack-quantized):esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4
License
apache-2.0 (inherits from Qwen base and DavidAU tune). general.license = apache-2.0 is set inside every GGUF. general.name equals filename without extension, general.description mentions full base Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU, qwen35.block_count 65, qwen35.context_length 262144.
Card written by AI assistance at the request of the repository author, who did the engineering. As with any AI-generated text, there may be errors; please verify anything important (hashes, sizes, commands) against the file itself before relying on it.
Run esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models