vcruz305/SuperHY3-abliterated-GGUF overview
SuperHY3 abliterated GGUF GGUF conversions of Jiunsong/SuperHY3 abliterated NVFP4 https://huggingface.co/Jiunsong/SuperHY3 abliterated NVFP4 295B A32B MoE, hy …
Runs locally from ~2.22 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| BF16/SuperHY3-abliterated-BF16-00001-of-00014.gguf | GGUF | BF16 | 41.50 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00002-of-00014.gguf | GGUF | BF16 | 40.25 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00003-of-00014.gguf | GGUF | BF16 | 41.51 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00004-of-00014.gguf | GGUF | BF16 | 41.88 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00005-of-00014.gguf | GGUF | BF16 | 39.87 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00006-of-00014.gguf | GGUF | BF16 | 41.48 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00007-of-00014.gguf | GGUF | BF16 | 41.50 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00008-of-00014.gguf | GGUF | BF16 | 41.89 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00009-of-00014.gguf | GGUF | BF16 | 41.64 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00010-of-00014.gguf | GGUF | BF16 | 41.59 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00011-of-00014.gguf | GGUF | BF16 | 41.59 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00012-of-00014.gguf | GGUF | BF16 | 41.74 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00013-of-00014.gguf | GGUF | BF16 | 39.94 GB | Download |
| BF16/SuperHY3-abliterated-BF16-00014-of-00014.gguf | GGUF | BF16 | 20.25 GB | Download |
| Q4_K_M/SuperHY3-abliterated-Q4_K_M-00001-of-00005.gguf | GGUF | Q4_K_M | 41.40 GB | Download |
| Q4_K_M/SuperHY3-abliterated-Q4_K_M-00002-of-00005.gguf | GGUF | Q4_K_M | 41.70 GB | Download |
| Q4_K_M/SuperHY3-abliterated-Q4_K_M-00003-of-00005.gguf | GGUF | Q4_K_M | 41.70 GB | Download |
| Q4_K_M/SuperHY3-abliterated-Q4_K_M-00004-of-00005.gguf | GGUF | Q4_K_M | 41.55 GB | Download |
| Q4_K_M/SuperHY3-abliterated-Q4_K_M-00005-of-00005.gguf | GGUF | Q4_K_M | 2.22 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00001-of-00008.gguf | GGUF | Q8_0 | 41.80 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00002-of-00008.gguf | GGUF | Q8_0 | 41.71 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00003-of-00008.gguf | GGUF | Q8_0 | 41.71 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00004-of-00008.gguf | GGUF | Q8_0 | 41.78 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00005-of-00008.gguf | GGUF | Q8_0 | 41.71 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00006-of-00008.gguf | GGUF | Q8_0 | 41.71 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00007-of-00008.gguf | GGUF | Q8_0 | 41.78 GB | Download |
| Q8_0/SuperHY3-abliterated-Q8_0-00008-of-00008.gguf | GGUF | Q8_0 | 3.64 GB | Download |
Model Details
Model README
---
license: apache-2.0
base_model: Jiunsong/SuperHY3-abliterated-NVFP4
base_model_relation: quantized
library_name: gguf
tags:
- gguf
- hy_v3
- moe
- abliterated
- uncensored
- mtp
---
SuperHY3-abliterated-GGUF
GGUF conversions of Jiunsong/SuperHY3-abliterated-NVFP4
(295B-A32B MoE, hy_v3 architecture, 80 layers + MTP layer, 192 experts).
⚠️ Provenance: dequantized from NVFP4 — read this first
No BF16 source of this abliteration exists anywhere. The upstream model is
published ONLY as an NVFP4 (compressed-tensors, FP4-E2M1 group-16) checkpoint.
These GGUFs were produced by exactly dequantizing that NVFP4 checkpoint to
BF16, then converting/quantizing as normal:
- The dequantization step is lossless with respect to the NVFP4 checkpoint
(bit-exact against the reference compressed-tensors implementation;
round-trip verified on the actual shards).
- BUT the expert FFN weights (the bulk of the parameters) carry the **FP4
ceiling (~4.25 bpw effective information)** of the source into every tier
below. Attention, dense layers, embeddings, and router weights were never
quantized upstream and are true BF16.
- Practical consequence: tiers up to ~Q4 lose essentially nothing vs a
hypothetical BF16 source. Q5/Q6/Q8/BF16 buy fidelity only on the
attention/dense tensors. The BF16 tier is provided as a conversion-faithful
archival source, not because it contains BF16-grade expert weights.
Dequantization tool: llm-dequant
(streaming NVFP4 → BF16 safetensors, byte-exact round-trip verification).
Requirements
hy_v3 support is not yet in llama.cpp master. Build from
(this repo was produced at commit cecbf5fb0):
git fetch origin pull/25395/head:hy3 && git checkout hy3
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build -j
Files
| Tier | Size | BPW | Notes |
|---|---|---|---|
| BF16 | 597.7 GB | 16 | conversion-faithful source (expert weights carry FP4 ceiling — see above) |
Quantized tiers (IQ1_S … Q8_0, imatrix) are being produced with the same
verified pipeline and will appear here as they upload.
MTP / speculative decoding
The MTP layer (blk.80, eh_proj/enorm) is present in the source and bundled
in these GGUFs. PR #25395 supports --spec-type draft-mtp. On the same
architecture (Hy3 IQ2_M, one GB10) we measured +27% decode speed with:
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.75 --parallel 1
--spec-draft-p-min 0.75 matters: the head is trained single-depth and the
p_min=0 default makes speculation a net loss. MTP acceptance on this
fine-tune has not been separately measured yet; numbers will be added with the
quant tiers.
Known quirks (inherited from the hy_v3 family)
- Native OpenAI-style
toolsAPI fails on llama-server (unsupported
<tool_calls:opensource> markup) — use prompt-injected tools; the model
tool-calls well, the server-side parser is what's missing.
- EOG metadata warning at load (
special_eos_id is not in special_eog_ids)
— if you see looping, add an explicit stop on the EOS token.
- Chat template requires
--jinja.
Provenance chain
tencent/Hy3 (BF16)
└─ Jiunsong abliteration + fine-tune, released as NVFP4 only
└─ llm-dequant: exact NVFP4 → BF16 dequantization
└─ convert_hf_to_gguf.py (llama.cpp PR #25395 @ cecbf5fb0) → BF16 GGUF
└─ llama-quantize (+imatrix) → quant tiersRun vcruz305/SuperHY3-abliterated-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models