GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

vcruz305/SuperHY3-abliterated-GGUF overview

SuperHY3 abliterated GGUF GGUF conversions of Jiunsong/SuperHY3 abliterated NVFP4 https://huggingface.co/Jiunsong/SuperHY3 abliterated NVFP4 295B A32B MoE, hy …

ggufhy_v3moeabliterateduncensoredmtpbase_model:Jiunsong/SuperHY3-abliterated-NVFP4base_model:quantized:Jiunsong/SuperHY3-abliterated-NVFP4license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~2.22 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
188
Likes
0
Pipeline
Author

Repository Files & Downloads

27 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
BF16/SuperHY3-abliterated-BF16-00001-of-00014.ggufGGUFBF1641.50 GBDownload
BF16/SuperHY3-abliterated-BF16-00002-of-00014.ggufGGUFBF1640.25 GBDownload
BF16/SuperHY3-abliterated-BF16-00003-of-00014.ggufGGUFBF1641.51 GBDownload
BF16/SuperHY3-abliterated-BF16-00004-of-00014.ggufGGUFBF1641.88 GBDownload
BF16/SuperHY3-abliterated-BF16-00005-of-00014.ggufGGUFBF1639.87 GBDownload
BF16/SuperHY3-abliterated-BF16-00006-of-00014.ggufGGUFBF1641.48 GBDownload
BF16/SuperHY3-abliterated-BF16-00007-of-00014.ggufGGUFBF1641.50 GBDownload
BF16/SuperHY3-abliterated-BF16-00008-of-00014.ggufGGUFBF1641.89 GBDownload
BF16/SuperHY3-abliterated-BF16-00009-of-00014.ggufGGUFBF1641.64 GBDownload
BF16/SuperHY3-abliterated-BF16-00010-of-00014.ggufGGUFBF1641.59 GBDownload
BF16/SuperHY3-abliterated-BF16-00011-of-00014.ggufGGUFBF1641.59 GBDownload
BF16/SuperHY3-abliterated-BF16-00012-of-00014.ggufGGUFBF1641.74 GBDownload
BF16/SuperHY3-abliterated-BF16-00013-of-00014.ggufGGUFBF1639.94 GBDownload
BF16/SuperHY3-abliterated-BF16-00014-of-00014.ggufGGUFBF1620.25 GBDownload
Q4_K_M/SuperHY3-abliterated-Q4_K_M-00001-of-00005.ggufGGUFQ4_K_M41.40 GBDownload
Q4_K_M/SuperHY3-abliterated-Q4_K_M-00002-of-00005.ggufGGUFQ4_K_M41.70 GBDownload
Q4_K_M/SuperHY3-abliterated-Q4_K_M-00003-of-00005.ggufGGUFQ4_K_M41.70 GBDownload
Q4_K_M/SuperHY3-abliterated-Q4_K_M-00004-of-00005.ggufGGUFQ4_K_M41.55 GBDownload
Q4_K_M/SuperHY3-abliterated-Q4_K_M-00005-of-00005.ggufGGUFQ4_K_M2.22 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00001-of-00008.ggufGGUFQ8_041.80 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00002-of-00008.ggufGGUFQ8_041.71 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00003-of-00008.ggufGGUFQ8_041.71 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00004-of-00008.ggufGGUFQ8_041.78 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00005-of-00008.ggufGGUFQ8_041.71 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00006-of-00008.ggufGGUFQ8_041.71 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00007-of-00008.ggufGGUFQ8_041.78 GBDownload
Q8_0/SuperHY3-abliterated-Q8_0-00008-of-00008.ggufGGUFQ8_03.64 GBDownload

Model Details

Model IDvcruz305/SuperHY3-abliterated-GGUF
Authorvcruz305
Pipeline
Licenseapache-2.0
Base modelJiunsong/SuperHY3-abliterated-NVFP4
Last modified2026-07-23T00:54:11.000Z

Model README

---

license: apache-2.0

base_model: Jiunsong/SuperHY3-abliterated-NVFP4

base_model_relation: quantized

library_name: gguf

tags:

- gguf

- hy_v3

- moe

- abliterated

- uncensored

- mtp

---

SuperHY3-abliterated-GGUF

GGUF conversions of Jiunsong/SuperHY3-abliterated-NVFP4

(295B-A32B MoE, hy_v3 architecture, 80 layers + MTP layer, 192 experts).

⚠️ Provenance: dequantized from NVFP4 — read this first

No BF16 source of this abliteration exists anywhere. The upstream model is

published ONLY as an NVFP4 (compressed-tensors, FP4-E2M1 group-16) checkpoint.

These GGUFs were produced by exactly dequantizing that NVFP4 checkpoint to

BF16, then converting/quantizing as normal:

  • The dequantization step is lossless with respect to the NVFP4 checkpoint

(bit-exact against the reference compressed-tensors implementation;

round-trip verified on the actual shards).

  • BUT the expert FFN weights (the bulk of the parameters) carry the **FP4

ceiling (~4.25 bpw effective information)** of the source into every tier

below. Attention, dense layers, embeddings, and router weights were never

quantized upstream and are true BF16.

  • Practical consequence: tiers up to ~Q4 lose essentially nothing vs a

hypothetical BF16 source. Q5/Q6/Q8/BF16 buy fidelity only on the

attention/dense tensors. The BF16 tier is provided as a conversion-faithful

archival source, not because it contains BF16-grade expert weights.

Dequantization tool: llm-dequant

(streaming NVFP4 → BF16 safetensors, byte-exact round-trip verification).

Requirements

hy_v3 support is not yet in llama.cpp master. Build from

PR #25395

(this repo was produced at commit cecbf5fb0):

git fetch origin pull/25395/head:hy3 && git checkout hy3
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

Files

| Tier | Size | BPW | Notes |

|---|---|---|---|

| BF16 | 597.7 GB | 16 | conversion-faithful source (expert weights carry FP4 ceiling — see above) |

Quantized tiers (IQ1_S … Q8_0, imatrix) are being produced with the same

verified pipeline and will appear here as they upload.

MTP / speculative decoding

The MTP layer (blk.80, eh_proj/enorm) is present in the source and bundled

in these GGUFs. PR #25395 supports --spec-type draft-mtp. On the same

architecture (Hy3 IQ2_M, one GB10) we measured +27% decode speed with:

--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.75 --parallel 1

--spec-draft-p-min 0.75 matters: the head is trained single-depth and the

p_min=0 default makes speculation a net loss. MTP acceptance on this

fine-tune has not been separately measured yet; numbers will be added with the

quant tiers.

Known quirks (inherited from the hy_v3 family)

  • Native OpenAI-style tools API fails on llama-server (unsupported

<tool_calls:opensource> markup) — use prompt-injected tools; the model

tool-calls well, the server-side parser is what's missing.

  • EOG metadata warning at load (special_eos_id is not in special_eog_ids)

— if you see looping, add an explicit stop on the EOS token.

  • Chat template requires --jinja.

Provenance chain

tencent/Hy3 (BF16)
  └─ Jiunsong abliteration + fine-tune, released as NVFP4 only
       └─ llm-dequant: exact NVFP4 → BF16 dequantization
            └─ convert_hf_to_gguf.py (llama.cpp PR #25395 @ cecbf5fb0) → BF16 GGUF
                 └─ llama-quantize (+imatrix) → quant tiers

Run vcruz305/SuperHY3-abliterated-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models