GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/NeoHorse-1-9B-GGUF overview

NeoHorse 1 9B GGUF NeoHorse 1 9B https://huggingface.co/TokenRhythm/NeoHorse 1 9B is a 9 billion parameter causal language model from TokenRhythm, post trained…

transformersgguftext-generation-inferencellama-cppagentictool-usecodingreasoning - instruction-followingtext-generationenbase_model:TokenRhythm/NeoHorse-1-9Bbase_model:quantized:TokenRhythm/NeoHorse-1-9Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~3.97 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

10 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
NeoHorse-1-9B.BF16.ggufGGUFGGUF16.69 GBDownload
NeoHorse-1-9B.Q3_K_L.ggufGGUFGGUF4.59 GBDownload
NeoHorse-1-9B.Q3_K_M.ggufGGUFGGUF4.31 GBDownload
NeoHorse-1-9B.Q3_K_S.ggufGGUFGGUF3.97 GBDownload
NeoHorse-1-9B.Q4_0.ggufGGUFGGUF4.95 GBDownload
NeoHorse-1-9B.Q4_K_M.ggufGGUFGGUF5.24 GBDownload
NeoHorse-1-9B.Q4_K_S.ggufGGUFGGUF4.98 GBDownload
NeoHorse-1-9B.Q5_0.ggufGGUFGGUF5.87 GBDownload
NeoHorse-1-9B.Q5_K_M.ggufGGUFGGUF6.02 GBDownload
NeoHorse-1-9B.Q5_K_S.ggufGGUFGGUF5.87 GBDownload

Model Details

Model IDprithivMLmods/NeoHorse-1-9B-GGUF
AuthorprithivMLmods
Pipelinetext-generation
Licenseapache-2.0
Base modelTokenRhythm/NeoHorse-1-9B
Last modified2026-09-09T04:16:54.000Z

Model README

---

license: apache-2.0

base_model:

  • TokenRhythm/NeoHorse-1-9B

pipeline_tag: text-generation

library_name: transformers

language:

  • en

tags:

  • text-generation-inference
  • llama-cpp
  • agentic
  • tool-use
  • coding
  • reasoning

- instruction-following

---

NeoHorse-1-9B-GGUF

> NeoHorse-1-9B is a 9-billion-parameter causal language model from TokenRhythm, post-trained from Qwen3.5-9B as an initial prototype on the path toward recursive self-improvement (RSI), targeting text-based agent harnesses, tool use, coding, and instruction following. Its core innovation is a routing harness that assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and feeds that signal back into shaping the next training mixture — using routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training data, supported by rigorous deduplication, decontamination, and six-dimensional semantic evaluation of training data. This release contains language-model weights only (vision weights excluded, with configuration and tensor keys repackaged for text-only inference) and retains a 262,144-token native context window extensible to 1,010,000. Across a ten-benchmark evaluation against Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and the much larger Muse-Glimmer-30B, NeoHorse-1-9B posts the best overall macro average (69.04 vs. 65.60 for its Qwen3.5-9B base, a +3.44 gain), with standout improvements on agentic benchmarks like VitaBench (+11.00), PinchBench (+7.70), and QwenClawBench (+4.69), alongside strong tau2-Bench (90.82) and BFCL v4 (67.43) scores, though instruction-following gains were mixed. It's servable via SGLang or vLLM with Qwen3-style reasoning and tool-call parsers, and is released under the Apache License 2.0.

Model Files

File Name | Quant Type | File Size | File Link |

|-----------|------------|-----------|-----------|

| NeoHorse-1-9B.BF16.gguf | BF16 | 17.9 GB | Download |

| NeoHorse-1-9B.Q3_K_L.gguf | Q3_K_L | 4.93 GB | Download |

| NeoHorse-1-9B.Q3_K_M.gguf | Q3_K_M | 4.62 GB | Download |

| NeoHorse-1-9B.Q3_K_S.gguf | Q3_K_S | 4.26 GB | Download |

| NeoHorse-1-9B.Q4_0.gguf | Q4_0 | 5.31 GB | Download |

| NeoHorse-1-9B.Q4_K_M.gguf | Q4_K_M | 5.63 GB | Download |

| NeoHorse-1-9B.Q4_K_S.gguf | Q4_K_S | 5.35 GB | Download |

| NeoHorse-1-9B.Q5_0.gguf | Q5_0 | 6.31 GB | Download |

| NeoHorse-1-9B.Q5_K_M.gguf | Q5_K_M | 6.47 GB | Download |

| NeoHorse-1-9B.Q5_K_S.gguf | Q5_K_S | 6.31 GB | Download |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/NeoHorse-1-9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models