wepiqx/Ornith-1.5-9B-MTP-ASHQ1-GGUF overview
Ornith 1.5 9B MTP — ASHQ1 Quantization ASHQ1 quants of Ornith 1.5 9B MTP GDN hybrid, 32 layers + MTP head . MTP head pinned at Q8 0 outside the budget, output/…
Runs locally from ~4.41 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | wepiqx/Ornith-1.5-9B-MTP-ASHQ1-GGUF |
|---|---|
| Author | wepiqx |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | flowerdreaming/Ornith-1.5-9B-MTP |
| Last modified | 2026-09-10T14:54:36.000Z |
Model README
---
license: apache-2.0
language:
- en
base_model: flowerdreaming/Ornith-1.5-9B-MTP
pipeline_tag: text-generation
library_name: gguf
tags:
- quantization
- gguf
- llama-cpp
- imatrix
- hybrid-quantization
- ASHQ1
- gdn-hybrid
---
Ornith-1.5-9B-MTP — ASHQ1 Quantization
ASHQ1 quants of Ornith-1.5-9B-MTP (GDN hybrid, 32 layers + MTP head).
MTP head pinned at Q8_0 outside the budget, output/token_embd pinned at Q5_K.
> 🧪 KLD-validated. Our PPL-only wins are cross-checked on Soulfate24's
> KLD battery (KLD vs BF16 reference — the rank column; PPL is a canary, not
> a rank). PPL below the BF16 base flags over-confidence, not quality.
Quants
| File | Size | PPL (ctx 1024) | KLD | Verdict |
|:-----|:----:|:--------------:|:---:|:--------|
| Ornith-1.5-9B-MTP-BF16-ASHQ1-6500.gguf | 6511 MiB | 8.6341 | 0.0763 (solid) | ✅ beats Remix Quality-36 (8.9895) at +123 MiB, KLD-clean |
| Ornith-1.5-9B-MTP-BF16-ASHQ1-4500.gguf | 4523 MiB | 10.1077 | pending A/B | standard build for KLD A/B vs TOX |
| Ornith-1.5-9B-MTP-BF16-ASHQ1-4500-TOX.gguf | 4523 MiB | 9.9629 | 0.2894 (degraded) | ⚠️ cautionary: PPL win is a sharpening artifact, do not use |
PPL: wikitext-2-raw wiki.test.raw, -ngl 99 -c 1024 -b 512 --seed 7.
KLD measured by Soulfate24 (ctx 512×64, Flash-Attention, vs BF16 reference).
Run wepiqx/Ornith-1.5-9B-MTP-ASHQ1-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models