GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

jan1k/Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-GGUF overview

Qwen3.8 27B Uncensored Genesis V1 — NVFP4 imatrix GGUF Absolutely experimental NVFP4 quantisations of LuffyTheFox/Qwen3.8 27B Uncensored Genesis V1 MTP GGUF ht…

ggufnvfp4fp4imatrixblackwellllama.cppqwen3.8uncensoredfastmtptext-generationenzhbase_model:LuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUFbase_model:quantized:LuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUFlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~861.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v5-noMTP.ggufGGUFGGUF14.46 GBDownload
Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v5.ggufGGUFGGUF14.68 GBDownload
Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v6-froggeric-noMTP.ggufGGUFGGUF14.46 GBDownload
Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v6-froggeric.ggufGGUFGGUF14.68 GBDownload
Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2.ggufGGUFGGUF14.68 GBDownload
Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-FastMTP-32K.ggufGGUFGGUF861.6 MBDownload
mmproj-Qwen3.8-27B-Uncensored-Genesis-BF16.ggufGGUFBF16888.0 MBDownload

Model Details

Model IDjan1k/Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-GGUF
Authorjan1k
Pipelinetext-generation
Licenseapache-2.0
Base modelLuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUF
Last modified2026-09-10T01:58:38.000Z

Model README

---

base_model: LuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUF

library_name: gguf

pipeline_tag: text-generation

license: apache-2.0

tags:

- gguf

- nvfp4

- fp4

- imatrix

- blackwell

- llama.cpp

- qwen3.8

- uncensored

- fastmtp

language:

- en

- zh

---

Qwen3.8-27B-Uncensored-Genesis-V1 — NVFP4 imatrix GGUF

Absolutely experimental NVFP4 quantisations of

LuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUF,

calibrated with two different imatrix corpora and built with the

advanced-gguf-quantizer

fork of llama.cpp.

Two calibration variants are provided, each with an MTP and a noMTP build:

| File | Calibration | MTP | Size | Tensors |

|---|---|---|---|---|

| ...NVFP4-v2-imx-v5.gguf | v5 — plain text (Bartowski) | yes (blk.64 preserved) | 14.68 GB | 1866 |

| ...NVFP4-v2-imx-v5-noMTP.gguf | v5 — plain text | no (blk.64 stripped) | 14.46 GB | 1835 |

| ...NVFP4-v2-imx-v6-froggeric.gguf | v6 — chat-template rendered | yes | 14.68 GB | 1866 |

| ...NVFP4-v2-imx-v6-froggeric-noMTP.gguf | v6 — chat-template rendered | no | 14.46 GB | 1835 |

All four share the same source, the same per-tensor protection policy, and the

same architecture profile (qwen35dense). They differ only in the calibration

data used to compute the NVFP4 input scales.

Source and architecture

| | |

|---|---|

| Source GGUF | Qwen3.8-27B-Uncensored-Genesis-V1-Q8_K_P.gguf (29990 MiB, 9.21 BPW) |

| Architecture | qwen35 (dense Qwen3.8), 27.32 B params, 64 layers + 1 MTP/NextN head |

| Context | 262144 |

| Quantiser | advanced-gguf-quantizer llama-quantize |

| general.file_type | 39 (LLAMA_FTYPE_MOSTLY_NVFP4) |

The source is from Q8_K_P, not from BF16 — no BF16 source was

available. As with any lossy-on-lossy requant, tensors that were genuinely q8 in

the source go through one extra rounding step; the substantial number of bf16

source tensors requantise cleanly.

Why imatrix, and why two variants

NVFP4 is a 4-bit float format with per-block input scales. Without an

importance matrix those scales fall back to identity (all 1.0), which works on

Qwen3.x but leaves quality on the table. An imatrix pass measures how much each

weight influences activations on real text, and the quantiser spends the input

scale budget where it matters.

Two calibration corpora were used:

v5 — plain text

Bartowski's v5 calibration set, ~1.7 MB of plain prose. Broad language coverage,

no chat-template structure. Useful as a general-purpose baseline.

v6 — chat-template rendered

Bartowski prose plus 173 conversations (137 base + 36 tool-calling) rendered

through the target model's chat template using

froggeric/Qwen-Fixed-Chat-Templates

v22.5. Rendering preserves the special/control tokens (<|im_start|>,

<|im_end|>, <|tool_call|>, etc.) and role structure the model actually sees

at inference time, so the activation statistics better match the chat/tool-use

distribution. llama-imatrix was run with --parse-special so the rendered

special-token markup is treated as real tokens, not literal text.

If the model is used primarily for chat, reasoning, or tool calling, v6 is the

recommended variant. If it is used for plain completion or broad text work, v5

is a reasonable choice and is slightly cheaper to have produced.

Per-tensor protection policy

Luffy's protection policy was written for the MoE Qwen 3.6 35B (which has

ffn_down_exps). This model is dense, so the equivalent tensor is

ffn_down. Applied that way, every line resolves.

The policy is implemented with anchored --tensor-type regex overrides. The

advanced quantiser matches overrides with std::regex_search and applies them

in reverse order (last match wins), so patterns are anchored with ^...$ to

avoid accidental partial matches (e.g. ssm_a would otherwise also match

ssm_alpha).

Singular-collapse protection (F16)

Four tensors are kept at F16 because they collapse badly under further

quantisation:

| tensor | type | note |

|---|---|---|

| blk.0.attn_gate.weight | F16 | as specified |

| blk.0.attn_qkv.weight | F16 | as specified |

| blk.0.ffn_down.weight | F16 | dense equivalent of ffn_down_exps |

| blk.13.ffn_down.weight | F16 | " |

A caveat: blk.0.attn_gate / attn_qkv were q8_0 in the source, so F16 here

stores dequantised q8 values. It does the intended job (those tensors keep

q8-level fidelity instead of dropping to 4 bits) but it is not true F16

precision, and it costs more bytes than the source did.

Architecture-specific F32 protection

Additional tensors are kept at F32 for FastMTP loader and Ampere CUDA kernel

compatibility. These protections were not in the original Luffy policy;

they were derived from runtime failures during testing.

| tensor | type | reason |

|---|---|---|

| blk.*.attn_norm.weight | F32 | 1D norm — avoids emitting .scale/.input_scale the FastMTP loader does not declare |

| blk.*.post_attention_norm.weight | F32 | " |

| blk.*.attn_q_norm.weight | F32 | " |

| blk.*.attn_k_norm.weight | F32 | " |

| blk.*.ssm_norm.weight | F32 | " |

| output_norm.weight | F32 | " |

| blk.*.ssm_conv1d.weight | F32 | patched CUDA SSM conv kernel requires F32 input (ssm-conv.cu asserts src1->nb[0] == sizeof(float)) |

| blk.*.ssm_dt.bias | F32 | SSM scalar, kept F32 for kernel compatibility |

| blk.*.ssm_a | F32 | SSM scalar, kept F32 for kernel compatibility |

Without the 1D-norm F32 protection, the quantiser emits 198 extra one-dimensional

.scale / .input_scale tensors (99 each) that the FastMTP loader at the

investigated revision does not declare, causing a tensor-count mismatch

(expected 1866, got 2064). Keeping the 1D norms at F32 avoids the auxiliary scale

tensors entirely.

Forced NVFP4 (no further collapse)

| tensor | type | note |

|---|---|---|

| blk.0.ssm_out.weight | NVFP4 | as specified — do not push lower, collapses |

| blk.1.attn_gate.weight | NVFP4 | as specified |

| blk.1.attn_qkv.weight | NVFP4 | as specified |

Catch-all

Everything else eligible takes NVFP4.

Tensor mix (verified in the finished files)

| type | count |

|---|---|

| F32 | 1357 |

| NVFP4 | 505 |

| F16 | 4 |

| total | 1866 |

  • general.file_type = 39 (NVFP4)
  • qwen35.block_count = 65 (64 + 1 MTP)
  • qwen35.nextn_predict_layers = 1
  • No unexpected one-dimensional norm .scale tensors
  • All four F16 protections and all three forced-NVFP4 tensors verified by name

The noMTP derivatives strip the 31 blk.64.* tensors and the

qwen35.nextn_predict_layers metadata key, decrement block_count 65 → 64,

and end at 1835 tensors / 14.46 GB.

FastMTP compatibility

The MTP variants were tested with

HauhauCS FastMTP

using the prebuilt draft sidecar

Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-FastMTP-32K.gguf and a patched

llama.cpp build. With the correct local model path (not a Hugging Face cache

variable accidentally pointing at a different quant), the model loaded, the

draft sidecar initialised, generation completed, and speculative decoding

produced measurable acceptance (~96 accepted / 234 generated, acceptance rate

~0.41, mean draft length ~2.23 on the tested prompt).

Important: FastMTP compatibility was verified by actually running the model

in the patched runtime, not by inspecting tensor counts alone. A model that

loads under stock llama.cpp is not automatically FastMTP-compatible.

Running it

llama-cli -m Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v6-froggeric.gguf \
  -ngl 99 -c 4096 --temp 0.6 --top-k 20 --top-p 1.0 --min-p 0.0

Sampling above follows the Genesis guidance (temp 0.6, top_k 20, top_p and

min_p disabled). For code, also set --repeat-penalty 1.0: code repeats

identifiers by nature, and any penalty pushes the model off the identifier it

just chose, which is a reliable way to get generated code that won't run.

Give it room to think

Genesis reasons at length before answering. In testing it can spend 900,

even 2600 tokens entirely inside the thinking block without reaching its final

answer. That is the model, not the quantisation — but with a small -n it

looks like the model is broken or rambling. Budget generously (-n 4096 or

more) for anything non-trivial, or you will only ever see it thinking.

FastMTP

llama-server \
  -m Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v6-froggeric.gguf \
  --spec-draft-model Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-FastMTP-32K.gguf \
  --spec-draft-ngl all \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --spec-draft-p-min 0 \
  --ctx-size 32768 \
  --parallel 1 \
  --batch-size 2048 \
  --ubatch-size 512 \
  --n-gpu-layers all \
  --split-mode none \
  --flash-attn on \
  --jinja

Use the MTP file (...v6-froggeric.gguf) for FastMTP, not the noMTP

derivative. The noMTP file is for ordinary inference where the MTP/NextN head is

not wanted.

Hardware notes

  • Blackwell (RTX 50xx): native FP4 path, fastest. Build llama.cpp with

BLACKWELL_NATIVE_FP4 = 1.

  • Ampere (RTX 30xx): NVFP4 inference works via fallback kernels. The

F32 SSM protections in these files prevent a CUDA assert in the patched SSM

convolution kernel on this architecture.

  • The imatrix pass itself runs on Ampere (it is ordinary source-model

inference); only native NVFP4 CUDA encoding is inefficient on Ampere, so

quantisation was done CPU-only.

How it was made

v5 imatrix

llama-imatrix \
  -m Qwen3.8-27B-Uncensored-Genesis-V1-Q8_K_P.gguf \
  -f calibration_datav5.txt \
  -o imatrix_v5.dat \
  --output-format dat \
  -ngl all -ts 3,1 -b 2048 -ub 512 -t 6

802 chunks

v6 imatrix

Conversations were rendered through the froggeric v22.5 chat template with

`transformers.AutoTokenizer.apply_chat_template(..., tools=tools,

add_generation_prompt=False, tokenize=False)`, concatenated with the prose, and

fed to llama-imatrix with --parse-special:

llama-imatrix \
  -m Qwen3.8-27B-Uncensored-Genesis-V1-Q8_K_P.gguf \
  -f calibration_v6_froggeric_rendered.txt \
  -o imatrix_v6_froggeric.dat \
  --output-format dat \
  -ngl all -ts 3,1 -b 2048 -ub 512 -t 6 \
  --parse-special

708 chunks

Quantisation

Both variants use the same protection policy and the same CPU-only quantise

command, differing only in the imatrix file:

llama-quantize \
  --allow-requantize --mode fast \
  --tensor-type .*=nvfp4 \
  --tensor-type '^blk\..*\.attn_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.post_attention_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.attn_q_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.attn_k_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.ssm_norm\.weight$=f32' \
  --tensor-type '^output_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.ssm_conv1d\.weight$=f32' \
  --tensor-type '^blk\..*\.ssm_dt\.bias$=f32' \
  --tensor-type '^blk\..*\.ssm_a$=f32' \
  --tensor-type '^blk.0.ssm_out.weight$=nvfp4' \
  --tensor-type '^blk.1.attn_gate.weight$=nvfp4' \
  --tensor-type '^blk.1.attn_qkv.weight$=nvfp4' \
  --tensor-type '^blk.0.attn_gate.weight$=f16' \
  --tensor-type '^blk.0.attn_qkv.weight$=f16' \
  --tensor-type '^blk.0.ffn_down.weight$=f16' \
  --tensor-type '^blk.13.ffn_down.weight$=f16' \
  --imatrix imatrix_v6_froggeric.dat \
  Qwen3.8-27B-Uncensored-Genesis-V1-Q8_K_P.gguf \
  Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-v2-imx-v6-froggeric.gguf \
  Q8_0 6

general.file_type is patched to 39 (LLAMA_FTYPE_MOSTLY_NVFP4) after

quantisation. The catch-all .*=nvfp4 is listed first; more-specific rules

after it. Patterns are anchored to avoid partial matches.

noMTP derivatives

strip_mtp.py removes the 31 blk.64.* tensors and the

qwen35.nextn_predict_layers metadata key, and decrements

qwen35.block_count 65 → 64.

Honest notes

  • Speed: these files use the advanced-gguf-quantizer's separate-scale

NVFP4 format (1866 tensors with explicit .scale / .input_scale tensors),

not the inline-scale format stock llama.cpp produces (866 tensors). The

separate-scale format was chosen for FastMTP loader compatibility at the

investigated llama.cpp revision. It is slightly slower at decode than the

inline-scale format on otherwise identical hardware, because dequant reads

scales from separate tensors rather than inline. If FastMTP is not needed,

an inline-scale build may be faster at the same quality.

  • imatrix does not change speed. It only changes the input_scale values

used during quantisation; the tensor layout and count are the same with and

without imatrix.

  • **Quality evidence here is limited to tensor inspection, metadata

verification, ordinary generation, and FastMTP acceptance measurement.** PPL

and KLD comparisons against the Q8_K_P source and against a data-free NVFP4

baseline are planned but not yet recorded. Do not treat the imatrix variants

as proven-better on the basis of this card alone; the calibration rationale is

sound, but the quality delta should be measured before being claimed.

  • **The F32 SSM and 1D-norm protections are a runtime compatibility choice,

not a quality choice.** They trade a small amount of size (and a small decode

cost) for FastMTP loader and Ampere CUDA kernel compatibility. On Blackwell

with a stock loader they may be unnecessary.

Credits

  • Base model and per-tensor protection policy:

LuffyTheFox

  • Genesis V1 finetune: HauhauCS
  • Chat template used for v6 calibration:

froggeric/Qwen-Fixed-Chat-Templates

v22.5

  • FastMTP draft sidecar and patched llama.cpp:

HauhauCS

(llama.cpp fork)

  • Calibration corpus: Bartowski v5 plain text; v6 prose + conversations rendered

through the target chat template

Run jan1k/Qwen3.8-27B-Uncensored-Genesis-V1-NVFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models