AtomicChat/ornith-35b-GGUF overview
<center <div style="display:flex; justify content:center; align items:center; gap:2%; max width:560px; margin:0 auto;" <a href="https://atomic.chat" style="fle…
Runs locally from ~19.71 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | AtomicChat/ornith-35b-GGUF |
|---|---|
| Author | AtomicChat |
| Pipeline | text-generation |
| License | mit |
| Base model | deepreinforce-ai/Ornith-1.0-35B |
| Last modified | 2026-07-22T20:06:52.000Z |
Model README
---
license: mit
license_link: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B/blob/main/LICENSE
thumbnail: https://huggingface.co/AtomicChat/ornith-35b-GGUF/resolve/main/hero.png
base_model:
- deepreinforce-ai/Ornith-1.0-35B
base_model_relation: quantized
quantized_by: AtomicChat
pipeline_tag: text-generation
library_name: gguf
tags:
- atomic-chat
- ornith
- deepreinforce-ai
- gguf
- llama.cpp
- quantized
---
<center>
<div style="display:flex; justify-content:center; align-items:center; gap:2%; max-width:560px; margin:0 auto;">
<a href="https://atomic.chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/ornith-35b-GGUF/resolve/main/pill_atomic_v3.png" alt="Atomic Chat" style="width:100%; height:auto; max-width:186px;"></a>
<a href="https://discord.gg/8wGSsvmg4V" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/ornith-35b-GGUF/resolve/main/pill_discord_v3.png" alt="Join Discord" style="width:100%; height:auto; max-width:184px;"></a>
<a href="https://github.com/AtomicBot-ai/Atomic-Chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/ornith-35b-GGUF/resolve/main/pill_github_v3.png" alt="GitHub" style="width:100%; height:auto; max-width:141px;"></a>
</div>
<br/>
<img src="https://huggingface.co/AtomicChat/ornith-35b-GGUF/resolve/main/hero.png" alt="Ornith 1.0 35B" style="width:100%; max-width:100%; height:auto; margin-bottom:0.6em;"/>
<div style="display:flex; justify-content:center; gap:0.5em;">
<a href="https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B"><strong>Base model: deepreinforce-ai/Ornith-1.0-35B</strong></a>
</div>
</center>
Ornith 1.0 35B, self-quantized to GGUF by Atomic Chat. Built straight from DeepReinforce's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
- 0.0B parameters: the weights this repo quantizes.
- Context length: 262,144 tokens (256K), as published by DeepReinforce.
- 40 layers: Mixture-of-Experts.
- Modalities: the base model handles Text, Image; this repo ships text-only quants, it carries no vision projector.
- Full imatrix ladder: every quant is calibrated with an importance matrix.
- State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
- Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
> [!NOTE]
> These GGUFs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
> [!IMPORTANT]
> Always pass --jinja so the Ornith 1.0 35B chat template is applied. Without it the model can emit malformed turns.
Model Overview
| Property | Value |
|---|---|
| Base model | deepreinforce-ai/Ornith-1.0-35B |
| Parameters | 0.0B |
| Layers | 40 |
| Experts | 256 routed (top-8) |
| Context length | 262,144 tokens (256K) |
| Vocabulary | 248,320 |
| Modalities | Text, Image in the base model; text only in this repo, it ships no vision projector |
| Architecture | Mixture-of-Experts, 256 experts (top-8), 16 attention heads over 2 KV heads, Qwen3_5MoeForConditionalGeneration |
| This repo | GGUF quants (imatrix). Quants: Q4_K_M, UD-Q4_K_XL, Q5_K_M, Q6_K, Q8_0 |
<img src="https://huggingface.co/AtomicChat/ornith-35b-GGUF/resolve/main/benchmark.png" alt="Ornith 1.0 35B benchmark scores" style="width:100%; max-width:900px;"/>
Scores are DeepReinforce's published results for the base deepreinforce-ai/Ornith-1.0-35B, not our own measurements. Quantization preserves the large majority of this; Q4_K_M and up stay close to full precision.
Choosing a quant
| Quant | Size | Notes |
|---|---|---|
| Q4_K_M | 21.2 GB | Recommended default. Best balance of size, speed and quality. |
| UD-Q4_K_XL | 21.5 GB | Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint. |
| Q5_K_M | 24.7 GB | Higher quality, low loss. |
| Q6_K | 28.5 GB | Near lossless, noticeably lighter than Q8_0. |
| Q8_0 | 36.9 GB | Effectively lossless, reference quality. |
> [!TIP]
> Pick the largest file that fits your (V)RAM with room for context. Q4_K_M or UD-Q4_K_XL is the sweet spot for most setups; Q6_K or Q8_0 for maximum fidelity.
Get started
Run Ornith 1.0 35B locally with:
- Atomic Chat: the easiest path. Open the app, search
AtomicChat/ornith-35b-GGUF, pick a quant, hit Use this model. - llama.cpp:
llama-server -hf AtomicChat/ornith-35b-GGUF:Q4_K_M --jinja -c 8192 - Ollama:
ollama run hf.co/AtomicChat/ornith-35b-GGUF:Q4_K_M - LM Studio / Jan: search the repo id, download any quant.
Best practices
| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 1.0 |
| top_k | 20 |
DeepReinforce's recommended sampling configuration for deepreinforce-ai/Ornith-1.0-35B.
Run in llama.cpp
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
./llama.cpp/build/bin/llama-server \
-hf AtomicChat/ornith-35b-GGUF:Q4_K_M \
--jinja -ngl 99 -c 8192 -fa on
How these were made
- Download
deepreinforce-ai/Ornith-1.0-35B(original weights). - Convert to f16 GGUF with llama.cpp.
- Build an importance matrix over our calibration corpus.
- Quantize the ladder with
--imatrix. UD-Q4_K_XLadditionally pins the token-embedding and output tensors toQ8_0.
License
Original model by DeepReinforce, released under the MIT license. Full terms: MIT. Quantized by Atomic Chat.
Run AtomicChat/ornith-35b-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models