GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AtomicChat/ornith-9b-GGUF overview

<center <div style="display:flex; justify content:center; align items:center; gap:2%; max width:560px; margin:0 auto;" <a href="https://atomic.chat" style="fle…

ggufatomic-chatornithdeepreinforce-aillama.cppquantizedtext-generationbase_model:deepreinforce-ai/Ornith-1.0-9Bbase_model:quantized:deepreinforce-ai/Ornith-1.0-9Blicense:mitendpoints_compatibleregion:usimatrixconversational

Runs locally from ~4.84 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
10,443
Likes
13
Pipeline
text-generation

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
ornith-9b-IQ4_XS.ggufGGUFIQ4_XS4.84 GBDownload
ornith-9b-Q4_K_M.ggufGGUFQ4_K_M5.24 GBDownload
ornith-9b-Q5_K_M.ggufGGUFQ5_K_M6.02 GBDownload
ornith-9b-Q6_K.ggufGGUFQ6_K6.85 GBDownload
ornith-9b-Q8_0.ggufGGUFQ8_08.87 GBDownload
ornith-9b-UD-Q4_K_XL.ggufGGUFQ4_K_XL5.95 GBDownload

Model Details

Model IDAtomicChat/ornith-9b-GGUF
AuthorAtomicChat
Pipelinetext-generation
Licensemit
Base modeldeepreinforce-ai/Ornith-1.0-9B
Last modified2026-07-22T20:07:32.000Z

Model README

---

license: mit

license_link: https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B/blob/main/LICENSE

thumbnail: https://huggingface.co/AtomicChat/ornith-9b-GGUF/resolve/main/hero.png

base_model:

  • deepreinforce-ai/Ornith-1.0-9B

base_model_relation: quantized

quantized_by: AtomicChat

pipeline_tag: text-generation

library_name: gguf

tags:

  • atomic-chat
  • ornith
  • deepreinforce-ai
  • gguf
  • llama.cpp
  • quantized

---

<center>

<div style="display:flex; justify-content:center; align-items:center; gap:2%; max-width:560px; margin:0 auto;">

<a href="https://atomic.chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/ornith-9b-GGUF/resolve/main/pill_atomic_v3.png" alt="Atomic Chat" style="width:100%; height:auto; max-width:186px;"></a>

<a href="https://discord.gg/8wGSsvmg4V" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/ornith-9b-GGUF/resolve/main/pill_discord_v3.png" alt="Join Discord" style="width:100%; height:auto; max-width:184px;"></a>

<a href="https://github.com/AtomicBot-ai/Atomic-Chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/ornith-9b-GGUF/resolve/main/pill_github_v3.png" alt="GitHub" style="width:100%; height:auto; max-width:141px;"></a>

</div>

<br/>

<img src="https://huggingface.co/AtomicChat/ornith-9b-GGUF/resolve/main/hero.png" alt="Ornith 1.0 9B" style="width:100%; max-width:100%; height:auto; margin-bottom:0.6em;"/>

<div style="display:flex; justify-content:center; gap:0.5em;">

<a href="https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B"><strong>Base model: deepreinforce-ai/Ornith-1.0-9B</strong></a>

</div>

</center>

Ornith 1.0 9B, self-quantized to GGUF by Atomic Chat. Built straight from DeepReinforce's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.

Highlights

  • 0.0B parameters: the weights this repo quantizes.
  • Context length: 262,144 tokens (256K), as published by DeepReinforce.
  • 32 layers: Dense decoder.
  • Modalities: the base model handles Text, Image; this repo ships text-only quants, it carries no vision projector.
  • Full imatrix ladder: every quant is calibrated with an importance matrix.
  • State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.

> [!NOTE]

> These GGUFs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.

> [!IMPORTANT]

> Always pass --jinja so the Ornith 1.0 9B chat template is applied. Without it the model can emit malformed turns.

Model Overview

| Property | Value |

|---|---|

| Base model | deepreinforce-ai/Ornith-1.0-9B |

| Parameters | 0.0B |

| Layers | 32 |

| Context length | 262,144 tokens (256K) |

| Vocabulary | 248,320 |

| Modalities | Text, Image in the base model; text only in this repo, it ships no vision projector |

| Architecture | Dense decoder, 16 attention heads over 4 KV heads, Qwen3_5ForConditionalGeneration |

| This repo | GGUF quants (imatrix). Quants: IQ4_XS, Q4_K_M, UD-Q4_K_XL, Q5_K_M, Q6_K, Q8_0 |

<img src="https://huggingface.co/AtomicChat/ornith-9b-GGUF/resolve/main/benchmark.png" alt="Ornith 1.0 9B benchmark scores" style="width:100%; max-width:900px;"/>

Scores are DeepReinforce's published results for the base deepreinforce-ai/Ornith-1.0-9B, not our own measurements. Quantization preserves the large majority of this; Q4_K_M and up stay close to full precision.

Choosing a quant

| Quant | Size | Notes |

|---|---|---|

| IQ4_XS | 5.2 GB | Excellent quality for size. Recommended low-bit. |

| Q4_K_M | 5.6 GB | Recommended default. Best balance of size, speed and quality. |

| UD-Q4_K_XL | 6.4 GB | Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint. |

| Q5_K_M | 6.5 GB | Higher quality, low loss. |

| Q6_K | 7.4 GB | Near lossless, noticeably lighter than Q8_0. |

| Q8_0 | 9.5 GB | Effectively lossless, reference quality. |

> [!TIP]

> Pick the largest file that fits your (V)RAM with room for context. Q4_K_M or UD-Q4_K_XL is the sweet spot for most setups; Q6_K or Q8_0 for maximum fidelity.

Get started

Run Ornith 1.0 9B locally with:

  • Atomic Chat: the easiest path. Open the app, search AtomicChat/ornith-9b-GGUF, pick a quant, hit Use this model.
  • llama.cpp: llama-server -hf AtomicChat/ornith-9b-GGUF:Q4_K_M --jinja -c 8192
  • Ollama: ollama run hf.co/AtomicChat/ornith-9b-GGUF:Q4_K_M
  • LM Studio / Jan: search the repo id, download any quant.

Best practices

| Parameter | Value |

|---|---|

| temperature | 1.0 |

| top_p | 1.0 |

| top_k | 20 |

DeepReinforce's recommended sampling configuration for deepreinforce-ai/Ornith-1.0-9B.

Run in llama.cpp

git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
./llama.cpp/build/bin/llama-server \
    -hf AtomicChat/ornith-9b-GGUF:Q4_K_M \
    --jinja -ngl 99 -c 8192 -fa on

How these were made

  1. Download deepreinforce-ai/Ornith-1.0-9B (original weights).
  2. Convert to f16 GGUF with llama.cpp.
  3. Build an importance matrix over our calibration corpus.
  4. Quantize the ladder with --imatrix.
  5. UD-Q4_K_XL additionally pins the token-embedding and output tensors to Q8_0.

License

Original model by DeepReinforce, released under the MIT license. Full terms: MIT. Quantized by Atomic Chat.

Run AtomicChat/ornith-9b-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models