zerodigest/Ornith-1.5-35B-Uncensored-YMQ-MTP-GGUF overview
< zerodigest lab banner v2 mid tone <div style="background: 1f1f23; color: f4f4f5; padding: 20px; border: 1px solid 38bdf8; border radius: 8px; margin: 20px 0;…
Runs locally from ~472.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Ornith-1.5-35B-Uncensored-YMQ-L.gguf | GGUF | GGUF | 17.36 GB | Download |
| Ornith-1.5-35B-Uncensored-YMQ-M.gguf | GGUF | GGUF | 15.07 GB | Download |
| Ornith-1.5-35B-Uncensored-YMQ-S.gguf | GGUF | GGUF | 12.49 GB | Download |
| Ornith-1.5-35B-Uncensored-YMQ-XS.gguf | GGUF | GGUF | 11.15 GB | Download |
| mmproj-Ornith-1.5-35B-Uncensored-Q4_K_S.gguf | GGUF | Q4_K_S | 472.9 MB | Download |
| mmproj-Ornith-1.5-35B-Uncensored-Q6_K.gguf | GGUF | Q6_K | 578.7 MB | Download |
Model Details
| Model ID | zerodigest/Ornith-1.5-35B-Uncensored-YMQ-MTP-GGUF |
|---|---|
| Author | zerodigest |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | 0xKitkat/Ornith-1.5-35B-A3B-Uncensored |
| Last modified | 2026-09-04T16:37:38.000Z |
Model README
---
license: apache-2.0
base_model: 0xKitkat/Ornith-1.5-35B-A3B-Uncensored
library_name: gguf
tags:
- text-generation
- gguf
- quantizer
- autoround
- architecture-aware
- moe
- mixture-of-experts
- uncensored
- 35b
---
<!-- zerodigest-lab-banner-v2-mid-tone -->
<div style="background: #1f1f23; color: #f4f4f5; padding: 20px; border: 1px solid #38bdf8; border-radius: 8px; margin: 20px 0; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;">
<h2 style="color: #38bdf8; margin: 0 0 10px 0; font-size: 18px; font-weight: 700; letter-spacing: -0.5px;">⚡ Fuel the Lab: Keep the Optimization Loops Running</h2>
<p style="font-size: 14px; line-height: 1.5; margin: 0 0 16px 0; color: #d4d4d8;">Every single <b>ZeroDigest YMQ-MTP</b> release is handcrafted and manually calibrated via intensive importance-matrix sweeps to protect critical logic pathways. This project is entirely independent research—no automation bots, no corporate backers, and no external funding. Running multi-hour compute arrays consumes massive local infrastructure overhead out-of-pocket. Consider checking out our compiler or supporting our compute costs!</p>
<div style="background: #18181b; padding: 12px; border-radius: 6px; border: 1px solid #3f3f46; font-size: 13px;">
<p style="margin: 0 0 8px 0; font-weight: 600; color: #f4f4f5;">👉 Developer Resources & Support Paths:</p>
<div style="margin-bottom: 6px;">🛠️ <a href="https://github.com/minyor/ymq-compiler" target="_blank" style="color: #38bdf8; text-decoration: underline; font-weight: bold;">View Compiler Source on GitHub (YMQ v2.0)</a></div>
<div style="margin-bottom: 6px;">🚀 <a href="https://runpod.io?ref=dbjmkmeh" target="_blank" style="color: #38bdf8; text-decoration: underline; font-weight: bold;">Deploy weights on RunPod Cloud Compute (Affiliate Link)</a></div>
<div style="margin-bottom: 8px;">☕ <a href="https://ko-fi.com/zerodigest" target="_blank" style="color: #38bdf8; text-decoration: underline; font-weight: bold;">Support the Lab on Ko-fi (One-Time / Monthly)</a></div>
<div style="margin: 8px 0; border-top: 1px solid #3f3f46;"></div>
<div style="margin-bottom: 4px; font-family: monospace;">🧡 <span style="color: #f97316; font-weight: bold;">BTC:</span> <code style="background: #27272a; padding: 1px 4px; border-radius: 4px; color: #f4f4f5;">14Fmic9z3VA1ZU11bWoP6JtU7csTAApwZo</code></div>
<div style="font-family: monospace;">🔷 <span style="color: #6366f1; font-weight: bold;">ETH/USDT:</span> <code style="background: #27272a; padding: 1px 4px; border-radius: 4px; color: #f4f4f5;">0xbe4cdc3adc27c21ef1c27c6b403311db07b35ed2</code></div>
</div>
</div>
Ornith-1.5-35B-Uncensored-YMQ-MTP-GGUF
> Source Model: 0xKitkat/Ornith-1.5-35B-A3B-Uncensored
⚖️ An Architecture-Aware, AutoRound-Inspired MoE Mixed Precision Layout
This repository features advanced, custom architecture-aware quantizations of Ornith-1.5-35B (Uncensored) processed directly from official raw BF16 source files using the custom YMQ-Compiler (v2.0) log-space framework.
These builds natively preserve the delicate Mixture-of-Experts routing topology and utilize high-context optimization parameters tailored for demanding local agent execution environments (such as RooCode/Aider).
<p align="center">
<img src="./logo.jpg" width="800" alt="ZeroDigest YMQ Logo">
</p>
---
📊 Quantization Preset Tier Details
| Preset Tier | Total Size | Target Usage / Memory VRAM Profile | Cognitive Real-World Coding Quality |
| :--- | :--- | :--- | :--- |
| XS | ~11.5 GB | 12GB GPU Entry / Balanced Budget Setup | The 12GB VRAM Champion. Fully re-engineered layout to protect structural tracking in compact footprints. |
| S | ~13.4 GB | Light workspace / Low-VRAM cache headroom | Balanced economy. Linear compression baseline for standard workflows. |
| M | ~16.2 GB | The High-Context Sweet Spot (Recommended) | 🎯 Elite logical stability. Optimizes VRAM to leave a massive headroom buffer for deep agent context loops. |
| L | ~18.6 GB | Premium Single-GPU Processing / Heavy workloads | Near lossless. Solid performance scaling across heavier consumer setups. |
📉 Perplexity Evaluation Metrics (WikiText-2)
The following metrics show the mathematical quality preservation of the YMQ-Compiler log-space cluster analysis compared to standard linear quantization layouts. Tested natively via llama-perplexity over a 4096 context window using the official WikiText-2 test corpus.
| Model Preset Variant | File Size | Perplexity Score (Lower is Better) | Internal Bit Gradient (High ➔ Mid ➔ Low ➔ Default ➔ Floor) |
| :--- | :--- | :--- | :--- |
| XS | ~11.5 GB | 13.5942 | IQ3_S ➔ IQ2_S ➔ IQ2_XS ➔ IQ2_XS (No Floor) |
| S | ~13.4 GB | 13.5454 | IQ4_NL ➔ IQ3_S ➔ IQ3_XXS ➔ IQ2_S (No Floor) |
| M (Recommended)| ~16.2 GB | 12.4310 | Q5_K ➔ IQ4_XS ➔ IQ3_S ➔ IQ3_XXS (No Floor) |
| L | ~18.6 GB | 12.5107 | Q6_K ➔ Q5_K ➔ IQ4_NL ➔ IQ3_S (No Floor) |
💡 The MoE Compression Breakthrough: YMQ vs. Standard Quants
Standard quantization pipelines act like a blunt hammer. They apply a uniform, flat bit-depth across every layer of the model, which completely breaks the delicate routing paths of Mixture-of-Experts (MoE) architectures.
Independent local testing using the official llama-perplexity harness exposes the massive optimization gap between standard, flat layouts and the architecture-aware YMQ-Compiler:
- Standard Flat Q4_K_M (~21.0 GB):
- Hits a high 13.4757 perplexity score on the test corpus.
- Starves the core attention entry channels and native MTP speculator tracking paths of bit-depth resolution.
- Model experiences severe tracking fatigue under stress.
- YMQ-Compiler M Preset (~16.0 GB):
- Scores a spectacular 12.4310 perplexity score on the exact same corpus.
- Surgically isolates and insulates critical logic entry gates behind high-fidelity shields (Q5_K / Q6_K).
- Heavily compresses background expert layers into non-linear matrices.
- Saves a massive 5 Gabytes of VRAM while delivering a drastic leap in logical reasoning clarity.
The Practical Result
The practical result is an elite, filter-free developer build that completely eliminates late-stage context amnesia bugs.
- Flat, uniform community quants suffer from endless reading loops and failed edits past 30k tokens.
- YMQ-M and S presets stream massive 200k+ context windows completely in VRAM.
- Blistering multi-token speculation speeds (80–120 t/s).
- Zero loop failures!
⚖️ YMQ vs. Uniform Quantization (The AutoRound Philosophy)
Standard quantization pipelines apply a blunt, uniform bit-depth across every single layer in a model. This blunt approach completely collapses the delicate routing structures of large Mixture-of-Experts (MoE) architectures, starving critical logic anchors of necessary precision while bloating file sizes with idle parameters.
The YMQ-Compiler implements a post-training optimization philosophy similar to advanced weight-tuning frameworks like Intel's AutoRound:
- Targeted Bit Isolation: Instead of applying a flat matrix mask, YMQ operates strictly in Log-Space. It surgically identifies the core attention entry lanes and high-leverage routing networks, locking them behind heavy
Q5_KandQ6_Khigh-fidelity safety shields. - Expert Layer Flattening: It takes the massive pool of background expert weights (
ffn_*_exps) and compresses them aggressively using dense, non-linear grids (IQ2_XSandIQ4_NL). Because these background parameters make up the majority of the file footprint but are rarely active at the same time, the compiler shaves off gigabytes of background noise without breaking the model's main train of thought.
The result is a highly stable, custom mixed-precision portfolio that matches the low perplexity and high instruction clarity of premium optimized configurations, while allowing 16GB and 24GB single-GPU setups to stream massive context windows with absolute structural peace of mind!
---
🛠️ The YMQ Compilation Architecture
Standard quantization pipelines treat network tensors like a flat dataset, applying destructive blanket low-bit compression to delicate tracking networks. The YMQ-Compiler solves high-context logic decay by parsing model files dynamically via an automated, multi-tiered protection matrix:
- Log-Space Gap Detection Clustering: Instead of flat percentage thresholds, the engine computes statistical cluster variances in log-space, successfully isolating intermediate logical reasoning spikes and elevating them to stable non-linear 4-bit (
IQ4_XS) formats, while compressing idle fact-storage layers to aggressive 2-bit baselines. - Fading Boundary Tapering: Recognizes the extreme fragility of initial token entry data vectors, forcing an input wave cushion (
L00=IQ4_NL→L01=IQ4_XS→L02=IQ3_XXS) that gradually stabilizes parameters before hitting the fallback pools. - Dedicated Gate Insulation: Hard-shields volatile parallel Transformer Multi-Head Attention and MoE expert routing paths, keeping context tracking perfectly noise-free.
- Asymmetric Vocabulary Shielding: Fixes tied-weight boundary errors by mapping the final logit classification exit heads to robust configurations to completely eliminate formatting loops and API tag leakage under deep contexts.
---
🚀 Recommended Runtime Parameters (llama.cpp / llama-server)
Need to scale up? Deploy this exact script on on-demand cloud GPUs via RunPod Cloud Compute.
$./llama-server -m Ornith-1.5-35B-Uncensored-YMQ-M.gguf -ctk q8_0 -ctv q4_0 --ctx-size 245760 --mmproj mmproj-Ornith-1.5-35B-Uncensored-Q6_K.gguf \
--n-cpu-moe 5 --timeout 36000 --checkpoint-min-step 2048 --ctx-checkpoints 4 --spec-type draft-mtp --spec-draft-n-max 2 \
--n-predict -1 --temp 0.6 --top-p 0.95 --top-k 20 --repeat-penalty 1.05 --jinja -fa
🖼️ Vision Projection (--mmproj)
For multimodal vision support, pair these builds with one of the following projection files:
| Variant | File | Size | Notes |
| :--- | :--- | :--- | :--- |
| Full Precision (BF16) | mmproj-Ornith-1.5-35B-Uncensored-BF16.gguf | ~903 MB | Full-precision vision tower, native to the Uncensored abliteration weights. Maximum fidelity for image reasoning tasks. |
| Q6_K | mmproj-Ornith-1.5-35B-Uncensored-Q6_K.gguf (this repo) | ~579 MB | High-fidelity quantized vision tower. Excellent quality-to-size balance with minimal perceptible degradation. |
| Q4_K_S | mmproj-Ornith-1.5-35B-Uncensored-Q4_K_S.gguf (this repo) | ~473 MB | Compact vision projection for VRAM-constrained setups. Retains strong image understanding at reduced footprint. |
Pass via --mmproj <path-to-file> in your llama-server invocation (see example above).
---
☕ Support & Future R&D
If the YMQ-Compiler builds saved your context window from collapsing or optimized your active development cycle speeds, consider buying a coffee to fund further low-level optimization research. Your support keeps the server nodes baking future model scales!
👉 Support ZeroDigest Research on ko-fi
---
📦 Source Framework & Automation Code
The compiler pipeline automation engine, setup thresholds, and structural mapping rules are open-source. To view the implementation details or compile your own custom models natively using this profile layout, visit the official development hub:
Run zerodigest/Ornith-1.5-35B-Uncensored-YMQ-MTP-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models