GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

singulared/Ornith-1.0-35B-MTP-GGUF overview

Ornith 1.0 35B MTP GGUF ๐Ÿ’ก Strix Halo / gfx1151? A fork only ROCmFP4 build ROCmFPX, ~19 GiB, Vulkan 86.7 t/s MTP is available separately: โ†’ singulared/Ornith 1โ€ฆ

ggufllama.cppmtpspeculative-decodingqwen35moecodeagentictext-generationbase_model:Qwen/Qwen3.6-35B-A3Bbase_model:merge:Qwen/Qwen3.6-35B-A3Bbase_model:deepreinforce-ai/Ornith-1.0-35Bbase_model:merge:deepreinforce-ai/Ornith-1.0-35Blicense:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~20.55 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
747
Likes
3
Pipeline
text-generation

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
ornith-1.0-35b-MTP-Q4_K_M.ggufGGUFQ4_K_M20.55 GBDownload
ornith-1.0-35b-MTP-Q8_0.ggufGGUFQ8_035.21 GBDownload

Model Details

Model IDsingulared/Ornith-1.0-35B-MTP-GGUF
Authorsingulared
Pipelinetext-generation
Licenseapache-2.0
Base modeldeepreinforce-ai/Ornith-1.0-35B,Qwen/Qwen3.6-35B-A3B
Last modified2026-07-23T00:43:45.000Z

Model README

---

license: apache-2.0

base_model:

- deepreinforce-ai/Ornith-1.0-35B

- Qwen/Qwen3.6-35B-A3B

base_model_relation: merge

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- llama.cpp

- mtp

- speculative-decoding

- qwen35moe

- code

- agentic

---

Ornith-1.0-35B-MTP (GGUF)

> ๐Ÿ’ก Strix Halo / gfx1151? A fork-only ROCmFP4 build (ROCmFPX, ~19 GiB, Vulkan 86.7 t/s MTP) is available separately: โ†’ singulared/Ornith-1.0-35B-MTP-ROCmFP4-GGUF

Ornith-1.0-35B (DeepReinforce) with an embedded MTP (Multi-Token Prediction / nextn) head grafted in, enabling self-speculative decoding in llama.cpp at identical output quality.

Which file?

| file | size | decode t/s | acceptance | notes |

|---|---|---|---|---|

| ornith-1.0-35b-MTP-Q4_K_M.gguf | 20.6 GiB | ~80 | 0.847 | recommended |

| ornith-1.0-35b-MTP-Q8_0.gguf | 35.2 GiB | ~63โ€“66 | 0.859 | near-lossless weights |

Sizes are weights only โ€” budget additional headroom for the KV cache and compute buffers, which grow with context length. At long context (128K) plan well above the file size.

Both carry the MTP head at Q8_0 precision. That matters: the nextn.eh_proj tensor is only ~9 MB but it largely determines draft acceptance โ€” quantizing it down costs a substantial slice of the speedup, so it is kept at Q8 even in the Q4_K_M build.

Why

Ornith-1.0-35B is an agentic coder fine-tuned from Qwen3.6-35B-A3B (same qwen35moe architecture, same tokenizer). The base Qwen ships with an embedded MTP head; Ornith's release does not โ€” so it decodes without self-speculation (~50 t/s at Q8_0). Because the fine-tune barely shifts the relevant hidden states, the base model's MTP head transfers almost perfectly when grafted in.

Notably the head also transfers down the quant ladder: moving from a Q8_0 body to a Q4_K_M body costs only ~1.2 points of acceptance (0.859 โ†’ 0.847), so the smaller build keeps essentially all of the benefit.

Measured

Radeon 8060S / Strix Halo (gfx1151), llama.cpp build 387 (571d0d5), -c 8192, temp 0.

| build | backend | decode t/s | draft acceptance | mean accepted len |

|---|---|---|---|---|

| Q4_K_M + MTP | Vulkan | 80.2 | 0.847 | 3.39 |

| Q4_K_M + MTP | ROCm/HIP | 61.0 | 0.870 | 3.28 |

| Q8_0 + MTP | Vulkan | 63โ€“66 | 0.859 | 3.05 |

| Q8_0, no MTP (reference) | Vulkan | ~50 | โ€” | โ€” |

Prefer Vulkan for this model โ€” it is ~31% faster than ROCm/HIP on the same build, while acceptance is statistically identical across backends (acceptance is a property of the math, not the kernel), so the gap is pure kernel throughput.

> โš ๏ธ ROCm/HIP on gfx1151: throughput measured, output not validated at long context. The ROCm row above is a speed measurement from short prompts only. A different K-quant model on this hardware produced corrupted output on ROCm at long context while being clean on Vulkan, and the ROCm figure here was never checked for correctness at depth. Treat ROCm/HIP as unverified for this model and use Vulkan.

The weights are unchanged โ†’ output quality is identical; the speedup is pure self-speculation.

Usage (llama.cpp)

llama-server -m ornith-1.0-35b-MTP-Q4_K_M.gguf \
  --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.6 \
  -fa on -ngl 99 -c 131072 --jinja --alias ornith

The MTP head is embedded โ€” no separate draft model (-md) required.

--spec-draft-n-max is worth a quick sweep on your hardware: this MoE peaks around 4, and drafting deeper eventually costs more than it wins because acceptance falls with draft depth.

How it was made (reproducible)

The donor MTP block is blk.40 โ€” a full nextn layer: attention + MoE experts + nextn.{eh_proj, enorm, hnorm, shared_head_norm}, 20 tensors โ€” originating from Qwen3.6-35B-A3B-Q8_0.gguf. Splice with gguf-py:

  1. Copy all of the target's tensors + metadata (raw quantized round-trip โ€” no dequant/requant). Skip the reader's virtual GGUF.* fields, which GGUFWriter emits itself.
  2. Append the donor's 20 blk.40.* tensors.
  3. Set qwen35moe.block_count = 41 and add qwen35moe.nextn_predict_layers = 1.

The Q4_K_M build was produced the same way, using the Q8_0 build above as the donor so the head keeps its Q8 precision on top of a Q4_K_M body.

Both bases share the architecture + tokenizer, so the head plugs in directly with no retraining.

Licensing & attribution

A derivative of two permissively-licensed models; both are credited and their licenses apply to their respective parts:

  • Ornith-1.0-35B โ€” ยฉ DeepReinforce โ€” MIT โ€” the base weights (733 of 753 tensors).
  • Qwen3.6-35B-A3B โ€” ยฉ Alibaba Cloud / Qwen โ€” Apache-2.0 โ€” the grafted MTP (blk.40) head.

Quantizations: Q4_K_M (Q8_0 MTP head) and Q8_0. Not affiliated with or endorsed by DeepReinforce or the Qwen team.

Run singulared/Ornith-1.0-35B-MTP-GGUF with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models