singulared/Ornith-1.0-35B-MTP-GGUF overview
Ornith 1.0 35B MTP GGUF ๐ก Strix Halo / gfx1151? A fork only ROCmFP4 build ROCmFPX, ~19 GiB, Vulkan 86.7 t/s MTP is available separately: โ singulared/Ornith 1โฆ
Runs locally from ~20.55 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | singulared/Ornith-1.0-35B-MTP-GGUF |
|---|---|
| Author | singulared |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | deepreinforce-ai/Ornith-1.0-35B,Qwen/Qwen3.6-35B-A3B |
| Last modified | 2026-07-23T00:43:45.000Z |
Model README
---
license: apache-2.0
base_model:
- deepreinforce-ai/Ornith-1.0-35B
- Qwen/Qwen3.6-35B-A3B
base_model_relation: merge
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- mtp
- speculative-decoding
- qwen35moe
- code
- agentic
---
Ornith-1.0-35B-MTP (GGUF)
> ๐ก Strix Halo / gfx1151? A fork-only ROCmFP4 build (ROCmFPX, ~19 GiB, Vulkan 86.7 t/s MTP) is available separately: โ singulared/Ornith-1.0-35B-MTP-ROCmFP4-GGUF
Ornith-1.0-35B (DeepReinforce) with an embedded MTP (Multi-Token Prediction / nextn) head grafted in, enabling self-speculative decoding in llama.cpp at identical output quality.
Which file?
| file | size | decode t/s | acceptance | notes |
|---|---|---|---|---|
| ornith-1.0-35b-MTP-Q4_K_M.gguf | 20.6 GiB | ~80 | 0.847 | recommended |
| ornith-1.0-35b-MTP-Q8_0.gguf | 35.2 GiB | ~63โ66 | 0.859 | near-lossless weights |
Sizes are weights only โ budget additional headroom for the KV cache and compute buffers, which grow with context length. At long context (128K) plan well above the file size.
Both carry the MTP head at Q8_0 precision. That matters: the nextn.eh_proj tensor is only ~9 MB but it largely determines draft acceptance โ quantizing it down costs a substantial slice of the speedup, so it is kept at Q8 even in the Q4_K_M build.
Why
Ornith-1.0-35B is an agentic coder fine-tuned from Qwen3.6-35B-A3B (same qwen35moe architecture, same tokenizer). The base Qwen ships with an embedded MTP head; Ornith's release does not โ so it decodes without self-speculation (~50 t/s at Q8_0). Because the fine-tune barely shifts the relevant hidden states, the base model's MTP head transfers almost perfectly when grafted in.
Notably the head also transfers down the quant ladder: moving from a Q8_0 body to a Q4_K_M body costs only ~1.2 points of acceptance (0.859 โ 0.847), so the smaller build keeps essentially all of the benefit.
Measured
Radeon 8060S / Strix Halo (gfx1151), llama.cpp build 387 (571d0d5), -c 8192, temp 0.
| build | backend | decode t/s | draft acceptance | mean accepted len |
|---|---|---|---|---|
| Q4_K_M + MTP | Vulkan | 80.2 | 0.847 | 3.39 |
| Q4_K_M + MTP | ROCm/HIP | 61.0 | 0.870 | 3.28 |
| Q8_0 + MTP | Vulkan | 63โ66 | 0.859 | 3.05 |
| Q8_0, no MTP (reference) | Vulkan | ~50 | โ | โ |
Prefer Vulkan for this model โ it is ~31% faster than ROCm/HIP on the same build, while acceptance is statistically identical across backends (acceptance is a property of the math, not the kernel), so the gap is pure kernel throughput.
> โ ๏ธ ROCm/HIP on gfx1151: throughput measured, output not validated at long context. The ROCm row above is a speed measurement from short prompts only. A different K-quant model on this hardware produced corrupted output on ROCm at long context while being clean on Vulkan, and the ROCm figure here was never checked for correctness at depth. Treat ROCm/HIP as unverified for this model and use Vulkan.
The weights are unchanged โ output quality is identical; the speedup is pure self-speculation.
Usage (llama.cpp)
llama-server -m ornith-1.0-35b-MTP-Q4_K_M.gguf \
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.6 \
-fa on -ngl 99 -c 131072 --jinja --alias ornith
The MTP head is embedded โ no separate draft model (-md) required.
--spec-draft-n-max is worth a quick sweep on your hardware: this MoE peaks around 4, and drafting deeper eventually costs more than it wins because acceptance falls with draft depth.
How it was made (reproducible)
The donor MTP block is blk.40 โ a full nextn layer: attention + MoE experts + nextn.{eh_proj, enorm, hnorm, shared_head_norm}, 20 tensors โ originating from Qwen3.6-35B-A3B-Q8_0.gguf. Splice with gguf-py:
- Copy all of the target's tensors + metadata (raw quantized round-trip โ no dequant/requant). Skip the reader's virtual
GGUF.*fields, whichGGUFWriteremits itself. - Append the donor's 20
blk.40.*tensors. - Set
qwen35moe.block_count = 41and addqwen35moe.nextn_predict_layers = 1.
The Q4_K_M build was produced the same way, using the Q8_0 build above as the donor so the head keeps its Q8 precision on top of a Q4_K_M body.
Both bases share the architecture + tokenizer, so the head plugs in directly with no retraining.
Licensing & attribution
A derivative of two permissively-licensed models; both are credited and their licenses apply to their respective parts:
- Ornith-1.0-35B โ ยฉ DeepReinforce โ MIT โ the base weights (733 of 753 tensors).
- Qwen3.6-35B-A3B โ ยฉ Alibaba Cloud / Qwen โ Apache-2.0 โ the grafted MTP (
blk.40) head.
Quantizations: Q4_K_M (Q8_0 MTP head) and Q8_0. Not affiliated with or endorsed by DeepReinforce or the Qwen team.
Run singulared/Ornith-1.0-35B-MTP-GGUF with guIDE
Download guIDE โ the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face ยท Compare models