GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ManniX-ITA/Ornith-1.5-27B-A3B-CoderX-MTP-GGUF overview

Ornith 1.5 27B A3B CoderX — GGUF imatrix, MTP GGUF quantizations of ManniX ITA/Ornith 1.5 27B A3B CoderX https://huggingface.co/ManniX ITA/Ornith 1.5 27B A3B C…

ggufllama.cppmoeexpert-pruningcodeimatrixmtpornithtext-generationbase_model:ManniX-ITA/Ornith-1.5-27B-A3B-CoderXbase_model:quantized:ManniX-ITA/Ornith-1.5-27B-A3B-CoderXlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~132.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

22 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ornith-1.5-27B-A3B-CoderX-F16.ggufGGUFF1648.87 GBDownload
Ornith-1.5-27B-A3B-CoderX-IQ2_M.ggufGGUFIQ2_M8.44 GBDownload
Ornith-1.5-27B-A3B-CoderX-IQ2_XS.ggufGGUFIQ2_XS7.64 GBDownload
Ornith-1.5-27B-A3B-CoderX-IQ3_M.ggufGGUFIQ3_M10.94 GBDownload
Ornith-1.5-27B-A3B-CoderX-IQ4_NL.ggufGGUFIQ4_NL13.98 GBDownload
Ornith-1.5-27B-A3B-CoderX-IQ4_XS.ggufGGUFIQ4_XS13.25 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q2_K_L.ggufGGUFQ2_K_L9.65 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q3_K_L.ggufGGUFQ3_K_L12.80 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q3_K_M.ggufGGUFQ3_K_M11.85 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q3_K_S.ggufGGUFQ3_K_S10.74 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q3_K_XL.ggufGGUFQ3_K_XL12.26 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q4_K_L.ggufGGUFQ4_K_L15.31 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q4_K_M.ggufGGUFQ4_K_M14.95 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q4_K_S.ggufGGUFQ4_K_S14.04 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q5_K_L.ggufGGUFQ5_K_L17.73 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q5_K_M.ggufGGUFQ5_K_M17.44 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q5_K_S.ggufGGUFQ5_K_S16.91 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q6_K.ggufGGUFQ6_K20.09 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q6_K_L.ggufGGUFQ6_K_L20.32 GBDownload
Ornith-1.5-27B-A3B-CoderX-Q8_0.ggufGGUFQ8_025.99 GBDownload
Ornith-1.5-27B-A3B-CoderX.imatrix.ggufGGUFGGUF132.7 MBDownload
mmproj-Ornith-1.5-27B-A3B-CoderX-F16.ggufGGUFF16857.6 MBDownload

Model Details

Model IDManniX-ITA/Ornith-1.5-27B-A3B-CoderX-MTP-GGUF
AuthorManniX-ITA
Pipelinetext-generation
Licenseapache-2.0
Base modelManniX-ITA/Ornith-1.5-27B-A3B-CoderX
Last modified2026-09-09T14:06:15.000Z

Model README

---

base_model: ManniX-ITA/Ornith-1.5-27B-A3B-CoderX

base_model_relation: quantized

license: apache-2.0

pipeline_tag: text-generation

tags:

- gguf

- llama.cpp

- moe

- expert-pruning

- code

- imatrix

- mtp

- ornith

---

Ornith-1.5-27B-A3B-CoderX — GGUF (imatrix, MTP)

GGUF quantizations of

ManniX-ITA/Ornith-1.5-27B-A3B-CoderX

— the REAP-stability-floor 184-expert cut of Ornith-1.5-35B-A3B.

**→ Recipe, drop maps, the full 11-bench evaluation, the cross-family comparison and the

length/deliberation analysis all live on the

weights card.** This card

covers the quantized files only.

> Its sibling Coder

> is the same 184-expert budget and the same competence map without the stability floor —

> the pair differs by exactly 6 experts per layer. This arm measures better on almost every

> code and reasoning axis at the same file size.

The -MTP- claim, verified from this file's own metadata

qwen35moe.block_count          = 41     (40 trunk + 1 MTP)
qwen35moe.nextn_predict_layers = 1
distinct blk indices           = 41, max = 40      <- block 40 is actually SHIPPED
MTP tensors                    = blk.40.nextn.{eh_proj,enorm,hnorm,shared_head_norm}
qwen35moe.expert_count         = 184
qwen35moe.expert_used_count    = 8

Worth stating explicitly for this arm: the REAM save path *drops the mtp. block**,

producing a GGUF that declares 41 blocks and ships only 40 — quantization passes, then

llama-imatrix dies on blk.40.attn_norm.weight not found. The head was grafted back from

the anchor, and the counts above are the read-back proof that it stuck.

Quants

Built with an importance matrix at every tier, including Q6_K, from this arm's own

imatrix. An imatrix cannot be shared between two different expert cuts, so CoderX has a

separate one from Coder. Sizes are GB (bytes ÷ 1e9).

| file | size | notes |

|---|---|---|

| Ornith-1.5-27B-A3B-CoderX-F16.gguf | 52.47 GB | unquantised source for re-quantising |

| Ornith-1.5-27B-A3B-CoderX-Q8_0.gguf | 27.91 GB | near-lossless reference |

| Ornith-1.5-27B-A3B-CoderX-Q6_K_L.gguf | 21.81 GB | Q6_K with embed/output at Q8_0 |

| Ornith-1.5-27B-A3B-CoderX-Q6_K.gguf | 21.57 GB | recommended; this is the eval tier |

| Ornith-1.5-27B-A3B-CoderX-Q5_K_L.gguf | 19.04 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q5_K_M.gguf | 18.73 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q5_K_S.gguf | 18.16 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q4_K_L.gguf | 16.43 GB | Q4_K_M with embed/output at Q8_0 |

| Ornith-1.5-27B-A3B-CoderX-Q4_K_M.gguf | 16.06 GB | default :latest on Ollama; fits 24 GB VRAM with context |

| Ornith-1.5-27B-A3B-CoderX-Q4_K_S.gguf | 15.07 GB | |

| Ornith-1.5-27B-A3B-CoderX-IQ4_NL.gguf | 15.01 GB | |

| Ornith-1.5-27B-A3B-CoderX-IQ4_XS.gguf | 14.22 GB | best size/quality under 15 GB |

| Ornith-1.5-27B-A3B-CoderX-Q3_K_XL.gguf | 13.17 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q3_K_L.gguf | 13.74 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q3_K_M.gguf | 12.72 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q3_K_S.gguf | 11.53 GB | |

| Ornith-1.5-27B-A3B-CoderX-IQ3_M.gguf | 11.74 GB | |

| Ornith-1.5-27B-A3B-CoderX-Q2_K_L.gguf | 10.36 GB | |

| Ornith-1.5-27B-A3B-CoderX-IQ2_M.gguf | 9.06 GB | |

| Ornith-1.5-27B-A3B-CoderX-IQ2_XS.gguf | 8.20 GB | smallest; expect real degradation |

| Ornith-1.5-27B-A3B-CoderX.imatrix.gguf | 0.14 GB | this arm's own imatrix, for reproduction/audit |

| mmproj-Ornith-1.5-27B-A3B-CoderX-F16.gguf | 0.90 GB | vision projector — see Vision below |

A Q4_K_M build of these exact weights (16.06 GB) also lives in

Ornith-1.5-27B-A3B-P6-MTP-GGUF.

Calibration (imatrix)

The importance matrix for every tier in this repo was computed on

AtomicChat/calib-corpora,

build builds/ornith-1.5-9b — a published, manifest-verified calibration corpus

rather than an ad-hoc text dump, so the quantisation basis is auditable and reproducible.

| | |

|---|---|

| corpus | AtomicChat/calib-corpora builds/ornith-1.5-9b |

| calib_train.txt | 3,079 documents · 4,958,490 tokens |

| chunks used | 9,686 (full file — no --chunks cap) |

| special tokens | --parse-special on (the corpus is chat-rendered; without it the chat control tokens never reach the statistics) |

| mixture | agentic 24.7% · code 17.8% · reasoning 14.8% · multilingual 13.8% · longctx 11.9% · vocab-sweep 9.9% · structured 3.9% · graphics 3.0% |

Corpus/model match was verified, not assumed. The build is keyed to a tokenizer, and

this arm's tokenizer.json hashes to

5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42 — byte-identical to the

tokenizer.sha256 recorded in that build's manifest.json. A calibration corpus rendered

for a different tokenizer would silently mis-tokenise, so this check gates the reuse.

The imatrix itself ships in this repo (*.imatrix.gguf) so the quants can be reproduced

or audited. Note it is this arm's own imatrix: an imatrix cannot be shared between two

different expert cuts.

Serving

llama-server -m Ornith-1.5-27B-A3B-CoderX-Q6_K.gguf --port 8080 \
    -c 262144 -ngl 99 --no-warmup \
    --reasoning-format deepseek --reasoning-budget 12288

--reasoning-budget is mandatory for this family. Use the vendor sampler

(temp 0.6 / top_p 0.95 / top_k 20); greedy is not viable here.

Budget the tail, not the median. On LiveCodeBench the median generation is ~12.7k

tokens but 5 of 77 problems run to a 32,768-token ceiling. If you care about worst-case

latency, bound it at serve time rather than hoping for a short draw — the weights card has

the full distribution.

Vision

These weights are text-only on their own; the vision tower ships separately as

mmproj-Ornith-1.5-27B-A3B-CoderX-F16.gguf in this repo. Pass it alongside any

quant:

llama-server -m Ornith-1.5-27B-A3B-CoderX-Q6_K.gguf \
    --mmproj mmproj-Ornith-1.5-27B-A3B-CoderX-F16.gguf \
    --port 8080 -c 262144 -ngl 99

Then post an image to /v1/chat/completions as an image_url data URI, exactly as with

any other multimodal llama.cpp model.

On Ollama the vision-capable tags are published under vision-<tier>:

ollama run mannix/ornith-1.5-27b-a3b-coderx:vision-Q4_K_M

Verified, not assumed. The projector was loaded against this arm in llama.cpp and

shown a synthetic 5-band image in a deliberately non-obvious colour order

(orange, purple, yellow, teal, brown). It named all five correctly and gave the right

band count; an identical prompt with no image attached correctly answered that it

could not see one — so the description comes from the pixels, not the prompt.

Evaluation (Q6_K + imatrix, llama.cpp, sampler recommended)

| Benchmark | base 35B (256e) | Coder (184e) | CoderX (184e) |

|---|---|---|---|

| GPQA-Diamond (198) | 0.8283 | 0.7677 | 0.8131 |

| AIME (30) | 0.9000 | 0.9000 | 0.9667 |

| LiveCodeBench v6 (77 hard) | 0.6623 | 0.7273 | 0.7662 |

| HumanEval (164) | 0.8720 | 0.8902 | 0.8963 |

| MultiPL-E (300) | 0.8533 | 0.8300 | 0.8367 |

| Mean (11 benches) | 0.8252 | 0.8292 | 0.8386 |

Full 11-bench table, per-language MultiPL-E, the Qwen3.6-27B-A3B cross-family comparison

and every basis caveat: weights card.

Apache-2.0 · research checkpoint.

Tool-calling benchmark — tool-eval-bench hardmode (88 scenarios, 176 pts)

Benchmarked file: Ornith-1.5-27B-A3B-CoderX-IQ4_XS.gguf (this repo), served on llama.cpp with MTP

speculative decoding enabled. The score below belongs to THIS quant — other

tiers in this repo were not run.

**Ornith-1.5-27B-A3B-CoderX scores 134.0 ±2.5 of 176, ninth of ten — and this figure is

a LOWER BOUND.**

One of the five seeds (s42) is graded on 174 points, not 176: scenario TC-05 was

dropped by an infrastructure fault on the serving side, not by a model failure. The cell

was quarantined and re-run on 2026-09-09, and the fault reproduced, so the cell was

kept and flagged rather than silently pooled. The plotted and tabulated value counts the

dropped scenario as 0, which is why it is a floor. The excluded scenario sits in

Parameter Precision, which is why that category reads 4/4 on that seed rather than 6/6.

Unlike its Coder sibling (136.8), the two are statistically indistinguishable here —

overlapping CIs, and the ordering is inside the noise. Do not read CoderX as worse than

Coder on this benchmark.

Category profile: perfect (100%) on Tool Selection, Restraint & Refusal, Toolset

Scale and Creative Composition; Structured Reasoning 5.6/6 (93.3%). Weakest at

Autonomous Planning 3.4/6 (56.7%), Context & State 12.6/20 (63.0%), Code Patterns

3.8/6 (63.3%) and Hard Mode 25.2/38 (66.3%).

14 safety-critical failures across five seeds: TC-60 on all five, TC-33 (Hallucination

Resistance) and TC-58 (Fake System Message in File) on four each.

Note the quant: IQ4_XS, versus Q4_K_M for the A3B and Omnimerge rows.

!Tool-calling benchmark

Full cohort

| model | quant | Total Points (mean, 5 seeds) | 95% CI | safety-critical (5 seeds) |

|---|---|---|---|---|

| Qwen3.8-27B-Omnimerge-v6 | Q4_K_M | 156.4 ±3.5 | [152.0, 160.8] | 3 |

| Qwen3.8-27B (base) | UD-Q4_K_M | 150.8 ±2.5 | [147.7, 153.9] | 9 |

| Ornith-1.5-35B | IQ4_XS | 146.2 ±2.6 | [143.0, 149.4] | 10 |

| Qwen3.6-27B-Omnimerge-v4 | Q4_K_M | 146.2 ±2.7 | [142.9, 149.5] | 16 |

| Qwen3.6-27B (base) | Q4_K_M | 144.0 ±3.4 | [139.8, 148.2] | 14 |

| Qwen3.6-35B-A3B (base) | IQ4_XS | 141.6 ±2.4 | [138.6, 144.6] | 15 |

| Qwen3.6-27B-A3B-CoderX | Q4_K_M | 137.4 ±4.9 | [131.3, 143.5] | 17 |

| Ornith-1.5-27B-A3B-Coder | IQ4_XS | 136.8 ±4.8 | [130.9, 142.7] | 12 |

| Ornith-1.5-27B-A3B-CoderX | IQ4_XS | 134.0 ±2.5 * | [130.8, 137.2] | 14 |

| Qwen3.6-27B-A3B-Coder | Q4_K_M | 123.2 ±2.3 | [120.4, 126.0] | 15 |

* one seed (s42) is graded on 174 pts, not 176 — see that model's card.

<details>

<summary><b>Basis — read before comparing these numbers to anything</b></summary>

  • Scorer: tool-eval-bench v2.6.0 (the pip/uv-installed package, verified via

tool_eval_bench.__file__, not a git checkout). An earlier note in the runner claimed

cf54b4b (v2.6.0-45); that is wrong and has been corrected — no cell ever ran it.

All 50 cells ran the same v2.6.0, so the cohort is internally consistent.

  • v2.6.0 carries a known scorer crash on TC-62. email_calls[-1] raises IndexError

when a model sent no valid CFO email; the orchestrator catches it and returns

FAIL / 0 points while keeping the scenario in the denominator. It hits 11 of 38

scored cells, 2 pts each, and it is **not neutral — it concentrates on the weakest

models**. Later harness commits credit that behaviour instead, so a fixed scorer would

raise affected scores, unevenly.

  • 5 paired seeds [42–46], 64k context, context-pressure 0.25, max 8 turns, 120 s timeout,

thinking enabled, sampler temp 0.6 / top-p 0.95 / top-k 20 (not greedy).

  • Served on llama.cpp b1788384120-c588c4f47 with MTP speculative decoding enabled

(nextn=YES spec=mtp), one model per GPU, sequential.

  • Quant tiers are not uniform across the cohort (Q4_K_M for the Omnimerge/A3B rows,

IQ4_XS for Ornith and 35B-A3B, UD-Q4_K_M for the Qwen3.8 base). Cross-row gaps

therefore carry a quantisation component and are not purely architectural.

  • Do not pool these with the r/LocalLLaMA published tool-eval-bench figures: those

were run at 256k context and are a different basis despite the shared scorer version.

</details>

Run ManniX-ITA/Ornith-1.5-27B-A3B-CoderX-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models