GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

mrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF overview

Ornith 1.0 35B AEON Ultimate Uncensored MTP GGUF GGUF conversions of AEON 7/Ornith 1.0 35B AEON Ultimate Uncensored https://huggingface.co/AEON 7/Ornith 1.0 35…

ggufmtpmulti-token-predictionspeculative-decodingbf16nvfp4q8_0q4_k_mqwen3_5_moeabliterateduncensoredrefusal-removedmoemixture-of-expertsreasoningthinkingcodingagentictool-callingconversational35btext-generationbase_model:AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16base_model:quantized:AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16

Runs locally from ~20.98 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-BF16.ggufGGUFBF1666.19 GBDownload
Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-NVFP4.ggufGGUFGGUF21.80 GBDownload
Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q4_K_M.ggufGGUFQ4_K_M20.98 GBDownload
Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q8_0.ggufGGUFQ8_035.21 GBDownload

Model Details

Model IDmrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF
Authormrexodia
Pipelinetext-generation
Licensemit
Base modelAEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16,AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4,unsloth/Qwen3.6-35B-A3B-MTP-GGUF
Last modified2026-06-28T22:24:15.000Z

Model README

---

model_name: Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF

license: mit

base_model:

- AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16

- AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4

- unsloth/Qwen3.6-35B-A3B-MTP-GGUF

base_model_relation: quantized

base_model_sources:

- name: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16

organization: AEON-7

description: Abliterated BF16 safetensors source used for the BF16 conversion and Q8_0 quantization.

repo_url: https://huggingface.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16

- name: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4

organization: AEON-7

description: Abliterated NVFP4 safetensors source used for the NVFP4 GGUF conversion.

repo_url: https://huggingface.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4

- name: unsloth/Qwen3.6-35B-A3B-MTP-GGUF

organization: unsloth

description: Qwen3.6 MTP tensor donor used for speculative decoding support.

repo_url: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF

library_name: gguf

pipeline_tag: text-generation

tags:

- gguf

- mtp

- multi-token-prediction

- speculative-decoding

- bf16

- nvfp4

- q8_0

- q4_k_m

- qwen3_5_moe

- abliterated

- uncensored

- refusal-removed

- moe

- mixture-of-experts

- reasoning

- thinking

- coding

- agentic

- tool-calling

- conversational

- 35b

quantized_by: mrexodia

---

Ornith-1.0-35B-AEON-Ultimate-Uncensored MTP GGUF

GGUF conversions of AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored with MTP (Multi-Token Prediction) tensors grafted in for speculative decoding support.

The base model is the AEON abliterated variant of deepreinforce-ai/Ornith-1.0-35B. The MTP tensors are from unsloth/Qwen3.6-35B-A3B-MTP-GGUF.

NOTE: I do not recommend enabling MTP for this model, it performs worse (BF16: 56 tok/s, BF16-MTP: 27 tok/s).

Available Quantizations

| File | Quant | Size | Source Weights | MTP Source |

|---|---:|---:|---|---|

| Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-BF16.gguf | BF16 | 71.07 GB (66.19 GiB) | AEON-7/...-BF16 | unsloth/Qwen3.6-35B-A3B-MTP-GGUF |

| Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-NVFP4.gguf | NVFP4 | 23.40 GB (21.80 GiB) | AEON-7/...-NVFP4 | unsloth/Qwen3.6-35B-A3B-MTP-GGUF via s-batman |

| Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q8_0.gguf | Q8_0 | 37.80 GB (35.21 GiB) | Quantized from the BF16 GGUF | unsloth/Qwen3.6-35B-A3B-MTP-GGUF |

| Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q4_K_M.gguf | Q4_K_M | 22.53 GB (20.98 GiB) | Quantized from the BF16 MTP GGUF | MTP preserved and quantized in-place |

Checksums

a13df4cce8a32b2065d8aea51dcc80d7056fea6c3277266d9b040923a2641840  Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-BF16.gguf
d78f62f6c112de9721390ce8f75b22cf753b3766a640257cc13ca85f16030292  Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-NVFP4.gguf
16696fb2e19b5b3faa316b198524be3dff3652555c67c3f3ea11e811147b219a  Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q8_0.gguf
3ed574a680d4f5d3db57d54271d31d6102180e746c5b70bf64864c13c3179f85  Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q4_K_M.gguf

Provenance

Base model

deepreinforce-ai/Ornith-1.0-35B
  -> AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
      -> AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4

MTP tensors

All MTP prediction heads originate from unsloth/Qwen3.6-35B-A3B-MTP-GGUF. These are compatible at the tensor-shape level because Ornith-1.0-35B uses the Qwen3.5 MoE architecture and tokenizer family.

  • BF16: blk.40.* MTP tensors grafted directly from Unsloth's BF16 split GGUF.
  • NVFP4: blk.40.* MTP tensors grafted via s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF. Byte-level verification confirms this block is identical to Unsloth's Qwen3.6-35B-A3B-MXFP4_MOE.gguf MTP block: 20 tensors, 512,079,872 tensor payload bytes, combined tensor-name-plus-payload SHA-256 8b8ba06cf776d2cdbf4d4db6714cf69b8a455105fc848bc02c4e5acb62f585f1.
  • Q8_0: blk.40.* MTP tensors grafted directly from Unsloth's Q8_0 GGUF.

Credit for the Qwen3.6 MTP tensors goes to Unsloth and the original Qwen release. s-batman is acknowledged as the intermediary who performed the NVFP4 graft used here as the practical donor for the NVFP4 file.

Usage

MTP requires a llama.cpp build with draft-mtp speculative decoding support.

llama-cli

llama-cli \
  -m Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q8_0.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  -p "Explain gradient descent in 3 sentences."

llama-server

llama-server \
  -m Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-NVFP4.gguf \
  --host 0.0.0.0 --port 8080 \
  -ngl all \
  -c 65536 \
  --spec-type draft-mtp \
  --spec-draft-n-max 3

Some frontends expose this as "MTP" or "speculative decoding" rather than the raw llama.cpp --spec-type draft-mtp flag.

LM Studio

Download the desired quant file. The model should appear as:

mrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF:BF16
mrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF:NVFP4
mrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF:Q8_0
mrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF:Q4_K_M

Enable MTP/speculative decoding in advanced settings if your frontend supports it.

Notes on correctness

MTP is used as a speculative draft. The target model verifies proposed tokens, so a poorly matched MTP head should reduce acceptance rate or speedup rather than change the final verified output distribution. The graft is still experimental and should be benchmarked for your workload.

NVFP4 is intended for hardware and software stacks with NVFP4 support. On unsupported hardware, use the BF16, Q8_0, or Q4_K_M files.

Exact production steps

All commands below were run from a llama.cpp checkout with a CUDA build available. Local cache paths are omitted for readability and shown as HuggingFace repo names.

1. Convert AEON BF16 safetensors to body-only BF16 GGUF

The source config advertises MTP, but the AEON BF16 safetensors snapshot does not contain MTP tensors. The body conversion was therefore done with --no-mtp.

python convert_hf_to_gguf.py \
  AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 \
  --outtype bf16 \
  --no-mtp \
  --outfile Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16.gguf

2. Graft BF16 MTP tensors

Copied all 20 donor tensors with prefix blk.40. from:

unsloth/Qwen3.6-35B-A3B-MTP-GGUF/BF16/Qwen3.6-35B-A3B-BF16-00002-of-00002.gguf

into the BF16 body GGUF, then updated:

qwen35moe.block_count = 41
qwen35moe.nextn_predict_layers = 1

The final published filename is:

Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-BF16.gguf

3. Convert AEON NVFP4 safetensors to body-only NVFP4 GGUF

python convert_hf_to_gguf.py \
  AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 \
  --outtype bf16 \
  --no-mtp \
  --outfile Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4.gguf

The converter detected and preserved the source NVFP4 quantization. --outtype bf16 only affects non-NVFP4 tensors that remain floating point.

4. Graft NVFP4/MXFP4_MOE MTP tensors

Copied all 20 donor tensors with prefix blk.40. from:

s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF/ornith-1.0-35b-NVFP4_MOE-MTP.gguf

This block was verified byte-for-byte identical to the MTP block in:

unsloth/Qwen3.6-35B-A3B-MTP-GGUF/Qwen3.6-35B-A3B-MXFP4_MOE.gguf

Then updated:

qwen35moe.block_count = 41
qwen35moe.nextn_predict_layers = 1

The final published filename is:

Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-NVFP4.gguf

5. Quantize BF16 body to Q8_0

The Q8_0 trunk was quantized from the body-only BF16 GGUF:

llama-quantize \
  Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16.gguf \
  Ornith-1.0-35B-AEON-Ultimate-Uncensored-Q8_0-body.gguf \
  q8_0

6. Graft Q8_0 MTP tensors

Copied all 20 donor tensors with prefix blk.40. from:

unsloth/Qwen3.6-35B-A3B-MTP-GGUF/Qwen3.6-35B-A3B-Q8_0.gguf

Then updated:

qwen35moe.block_count = 41
qwen35moe.nextn_predict_layers = 1

The final published filename is:

Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q8_0.gguf

7. Quantize BF16 MTP GGUF to Q4_K_M

The Q4_K_M file was quantized from the final BF16 MTP GGUF with token embeddings kept at Q8_0, output.weight set to Q6_K, and MTP matrix tensors set to Q8_0:

llama-quantize \
  --token-embedding-type Q8_0 \
  --output-tensor-type Q6_K \
  --mtp-tensor-type Q8_0 \
  Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-BF16.gguf \
  Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-Q4_K_M.gguf \
  Q4_K_M

8. Metadata updates

Each final GGUF was rewritten with file-specific metadata:

  • general.name
  • general.author = mrexodia
  • general.quantized_by = mrexodia
  • general.license = mit
  • general.license.name = MIT License
  • general.license.link = https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B/blob/main/LICENSE
  • general.source.huggingface.repository
  • general.description
  • general.base_model.count = 2
  • general.base_model.0.* for the AEON source
  • general.base_model.1.* for the Unsloth MTP donor
  • general.tags

The legacy custom key general.base_model was removed in favor of the interoperable general.base_model.{id}.name mapping used by HuggingFace GGUF metadata.

9. Verification

For each final GGUF:

  • qwen35moe.block_count = 41
  • qwen35moe.nextn_predict_layers = 1
  • 20 tensors with prefix blk.40. are present
  • 4 tensors under blk.40.nextn.* are present
  • llama.cpp loaded the model with --spec-type draft-mtp
  • Smoke test prompt What is 2+2? Answer with just the number. produced 4

License

MIT, inherited from the base model.

Run mrexodia/Ornith-1.0-35B-AEON-Ultimate-Uncensored-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models