GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

SC117/Ornith-1.5-35B-A3B-Heretic-MTP-APEX-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: f5f5f7; border radius:…

transformersggufqwen3_5_moereasoningagentic-codinghereticmtpapexquantizationmultimodaltext-generationbase_model:ornith-ai/Ornith-1.5-35B-A3Bbase_model:quantized:ornith-ai/Ornith-1.5-35B-A3Blicense:mitendpoints_compatibleregion:us

Runs locally from ~861.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
7
Pipeline
text-generation
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ornith-1.5-35B-A3B-Heretic-MTP-APEX-I-Balanced.ggufGGUFGGUF24.27 GBDownload
Ornith-1.5-35B-A3B-Heretic-MTP-APEX-I-Compact.ggufGGUFGGUF16.14 GBDownload
Ornith-1.5-35B-A3B-Heretic-MTP-APEX-I-Mini.ggufGGUFGGUF13.29 GBDownload
Ornith-1.5-35B-A3B-Heretic-MTP-APEX-I-Quality.ggufGGUFGGUF21.87 GBDownload
mmproj-Ornith-1.5-35B-A3B-Heretic-MTP-BF16.ggufGGUFBF16861.0 MBDownload

Model Details

Model IDSC117/Ornith-1.5-35B-A3B-Heretic-MTP-APEX-GGUF
AuthorSC117
Pipelinetext-generation
Licensemit
Base modelornith-ai/Ornith-1.5-35B-A3B
Last modified2026-08-22T22:49:35.000Z

Model README

---

library_name: transformers

license: mit

license_link: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B/blob/main/LICENSE

pipeline_tag: text-generation

tags:

  • qwen3_5_moe
  • reasoning
  • agentic-coding
  • heretic
  • mtp
  • apex
  • quantization
  • gguf
  • multimodal

base_model:

  • ornith-ai/Ornith-1.5-35B-A3B

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">

<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">

<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #dbeafe; border-radius: 50%;"></div>

<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #bfdbfe; border-radius: 50%;"></div>

<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #dbeafe; border-radius: 50%;"></div>

<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1;"><span style="background: #007aff; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">HERETIC 1.4.0</span><span style="background: #af52de; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">MTP</span><span style="background: #30d158; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">APEX</span><span style="background: #ff9500; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">MIT</span></div>

<h1 style="margin: 0 0 8px 0; font-size: 32px; font-weight: 700; color: #1d1d1f; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Ornith-1.5-35B-A3B-Heretic-MTP-APEX</h1>

<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;"><span style="color: #86868b;">English</span> | <a href="https://huggingface.co/SC117/Ornith-1.5-35B-A3B-Heretic-MTP-APEX-GGUF/blob/main/README_zh.md" style="color: #007aff; text-decoration: none;">📖 中文文档</a></p>

<p style="margin: 0; font-size: 15px; color: #86868b; position: relative; z-index: 1;">Heretic Trial 62 BF16 merge · replaced MTP head · APEX GGUFs + mmproj</p>

</div>

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🐦</span> About this release</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 12px 0;"><a href="https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B" style="color: #c2410c; text-decoration: none; font-weight: 700;">Ornith-1.5-35B-A3B</a> is DeepReinforce's Qwen3.5 MoE agentic coding model. This experimental derivative merges the selected <a href="https://github.com/p-e-w/heretic" style="color: #c2410c; text-decoration: none; font-weight: 700;">Heretic</a> 1.4.0 Trial 62 LoRA into BF16 weights.</p><p style="margin: 0;">The shipped native MTP head is untrained (initializer-like statistics). It is replaced here with the 19-tensor fused BF16 head from <a href="https://huggingface.co/shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY" style="color: #c2410c; text-decoration: none; font-weight: 700;">shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY</a> (Qwen3.6 graft + 12K KL distillation). Refusal-modified models can behave differently from the original; evaluate before deployment. <b>License: MIT.</b></p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧠</span> Model Details</div><div style="padding: 16px;"><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Architecture</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">Qwen3.5 MoE, multimodal</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Parameters</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">35B total, 3B active per token · 256 experts, 8 active</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Layers</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">40 transformer layers + 1 fused MTP layer (19 tensors, Q8_0 in APEX)</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Context</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">262,144 tokens</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">License</td><td style="padding: 6px 10px; color: #4b5563;">MIT</td></tr></tbody></table></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚡</span> Heretic Trial 62 of 80</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">Classic Heretic (MPOA, rank-3, <code>o_proj</code> + <code>down_proj</code>). Per-layer refusal directions. Search log: <b>9/100 refusals @ first-token KL 0.0107</b>. Re-score after export: <b>11/100 @ KL 0.0105</b>.</p><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold;">direction_index</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">per layer</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold;">attn.o_proj</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">max 3.19 · pos 29.14 · min 2.89 · dist 14.92</td></tr><tr><td style="padding: 6px 10px; font-weight: bold;">mlp.down_proj</td><td style="padding: 6px 10px;">max 3.43 · pos 23.58 · min 1.62 · dist 8.45</td></tr></tbody></table></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔁</span> MTP head</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">Official Ornith-1.5 <code>mtp.*</code> tensors match random init (e.g. <code>q_proj</code> std 0.01993 / kurtosis 2.997; <code>mtp.norm</code> mean 0.02281). Native draft acceptance is poor. This release <b>removes those 785 unpacked tensors</b> and grafts shisa's 19 fused tensors (sha256 <code>73c6e839…de712e</code>). APEX keeps the whole MTP block at <b>Q8_0</b>.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Usage</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">llama.cpp (vision + text, recent build with Qwen3.5 MoE MTP)</p><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">hf download SC117/Ornith-1.5-35B-A3B-Heretic-MTP-APEX-GGUF --include "*.gguf" --local-dir ./models

./llama-server -m ./models/Ornith-1.5-35B-A3B-Heretic-MTP-APEX-I-Compact.gguf --mmproj ./models/mmproj-Ornith-1.5-35B-A3B-Heretic-MTP-BF16.gguf -ngl 99 -c 131072 --draft-mtp</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🎛️</span> Recommended Settings</div><div style="padding: 16px;"><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">General / Coding</td><td style="padding: 6px 10px; color: #4b5563; font-family: monospace;">temperature=0.6, top_p=0.95, top_k=20, do_sample=true</td></tr></tbody></table></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>💡</span> What is APEX?</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0;">These GGUFs use <a href="https://github.com/mudler/apex-quant" style="color: #c2410c; text-decoration: none; font-weight: 700;">APEX</a>, an MoE-aware mixed-precision method. Routed experts compress hardest, shared experts stay high, attention follows a layer-wise gradient. The MTP layer remains Q8_0. I-variants use the original Ornith-1.5 trunk imatrix (40 layers; MTP has no imatrix rows and does not need them at Q8_0). The matching vision projector remains BF16.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> APEX Quantization Tiers</div><div style="padding: 16px;"><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(255,107,53,0.05);"><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c;">File</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c;">Size</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c;">Best For</th></tr></thead><tbody><tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);"><code>-APEX-I-Quality.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">23.49 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Highest quality</td></tr><tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);"><code>-APEX-I-Balanced.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">26.06 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Best all-rounder</td></tr><tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);"><code>-APEX-I-Compact.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">17.33 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Best quality/size ratio</td></tr><tr><td style="padding: 8px 10px;"><code>-APEX-I-Mini.gguf</code></td><td style="padding: 8px 10px;">14.27 GB</td><td style="padding: 8px 10px;">Most compact</td></tr></tbody></table><p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; font-style: italic;">Sizes are measured from the final files.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #ff6b35 0%, #f7931e 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>❓</span> FAQ: Why is I-Balanced larger than I-Quality?</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;"><b>Short answer: the tiers are bit-allocation strategies, not a size ladder.</b> APEX assigns precision per tensor from measured importance, so file size does not grow monotonically with the tier name.</p><p style="margin: 0 0 10px 0;"><b>I-Quality</b> keeps every sensitive tensor high-precision — attention Q6_K in all 40 blocks, shared experts Q8_0, edge blocks and the MTP layer high as well — but compresses the redundant middle routed experts (blk.10–30) aggressively to <b>IQ4_XS</b>. Routed experts carry most of an MoE's parameters, and the importance analysis shows the middle blocks absorb this compression with virtually no measurable loss. That is where the bytes are saved.</p><p style="margin: 0 0 10px 0;"><b>I-Balanced</b> is the conservative, uniform profile: no tensor anywhere below Q5_K. Uniformity simply costs more bytes.</p><p style="margin: 0 0 10px 0;">So bigger does not mean better. I-Quality is the smarter bit allocation and remains the highest-quality preset despite the smaller size; if you need a smaller file, step down to I-Compact or I-Mini instead of picking by file size.</p><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(255,107,53,0.05);"><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c;">Tensor group</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c;">I-Quality</th><th style="padding: 8px 10px; border-bottom: 2px solid #ff6b35; text-align: left; color: #c2410c;">I-Balanced</th></tr></thead><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Routed experts · edge blocks (0–4, 35–39)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">Q6_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">Q6_K</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Routed experts · blk.5–9, 31–34</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">Q5_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">Q5_K</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Routed experts · middle blocks (10–30)</td><td style="padding: 6px 10px;"><b>IQ4_XS</b></td><td style="padding: 6px 10px;"><b>Q5_K</b></td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Shared experts · all blocks</td><td style="padding: 6px 10px;">Q8_0</td><td style="padding: 6px 10px;">Q8_0</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Attention · all 40 blocks</td><td style="padding: 6px 10px;">Q6_K</td><td style="padding: 6px 10px;">Q6_K</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">MTP layer (blk.40)</td><td style="padding: 6px 10px;">Q8_0</td><td style="padding: 6px 10px;">Q8_0</td></tr></tbody></table><p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; font-style: italic;">The two tiers are identical everywhere except the middle routed-expert band — that single band is the whole size difference.</p></div></div>

</div>

Links

  • Original Model: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B
  • Replacement MTP head: https://huggingface.co/shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY
  • Heretic: https://github.com/p-e-w/heretic
  • APEX Quantization: https://github.com/mudler/apex-quant

Citation

@misc{ornith-1.5-35b,
    title = {{Ornith-1.5-35B-A3B}: From Self-Scaffolding to Self-Improvement},
    url = {https://ornith.ai/ornith_1_5.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}

Run SC117/Ornith-1.5-35B-A3B-Heretic-MTP-APEX-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models