prithivMLmods/NeoHorse-1-9B-GGUF overview
NeoHorse 1 9B GGUF NeoHorse 1 9B https://huggingface.co/TokenRhythm/NeoHorse 1 9B is a 9 billion parameter causal language model from TokenRhythm, post trained…
Runs locally from ~3.97 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| NeoHorse-1-9B.BF16.gguf | GGUF | GGUF | 16.69 GB | Download |
| NeoHorse-1-9B.Q3_K_L.gguf | GGUF | GGUF | 4.59 GB | Download |
| NeoHorse-1-9B.Q3_K_M.gguf | GGUF | GGUF | 4.31 GB | Download |
| NeoHorse-1-9B.Q3_K_S.gguf | GGUF | GGUF | 3.97 GB | Download |
| NeoHorse-1-9B.Q4_0.gguf | GGUF | GGUF | 4.95 GB | Download |
| NeoHorse-1-9B.Q4_K_M.gguf | GGUF | GGUF | 5.24 GB | Download |
| NeoHorse-1-9B.Q4_K_S.gguf | GGUF | GGUF | 4.98 GB | Download |
| NeoHorse-1-9B.Q5_0.gguf | GGUF | GGUF | 5.87 GB | Download |
| NeoHorse-1-9B.Q5_K_M.gguf | GGUF | GGUF | 6.02 GB | Download |
| NeoHorse-1-9B.Q5_K_S.gguf | GGUF | GGUF | 5.87 GB | Download |
Model Details
| Model ID | prithivMLmods/NeoHorse-1-9B-GGUF |
|---|---|
| Author | prithivMLmods |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | TokenRhythm/NeoHorse-1-9B |
| Last modified | 2026-09-09T04:16:54.000Z |
Model README
---
license: apache-2.0
base_model:
- TokenRhythm/NeoHorse-1-9B
pipeline_tag: text-generation
library_name: transformers
language:
- en
tags:
- text-generation-inference
- llama-cpp
- agentic
- tool-use
- coding
- reasoning
- instruction-following
---
NeoHorse-1-9B-GGUF
> NeoHorse-1-9B is a 9-billion-parameter causal language model from TokenRhythm, post-trained from Qwen3.5-9B as an initial prototype on the path toward recursive self-improvement (RSI), targeting text-based agent harnesses, tool use, coding, and instruction following. Its core innovation is a routing harness that assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and feeds that signal back into shaping the next training mixture — using routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training data, supported by rigorous deduplication, decontamination, and six-dimensional semantic evaluation of training data. This release contains language-model weights only (vision weights excluded, with configuration and tensor keys repackaged for text-only inference) and retains a 262,144-token native context window extensible to 1,010,000. Across a ten-benchmark evaluation against Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and the much larger Muse-Glimmer-30B, NeoHorse-1-9B posts the best overall macro average (69.04 vs. 65.60 for its Qwen3.5-9B base, a +3.44 gain), with standout improvements on agentic benchmarks like VitaBench (+11.00), PinchBench (+7.70), and QwenClawBench (+4.69), alongside strong tau2-Bench (90.82) and BFCL v4 (67.43) scores, though instruction-following gains were mixed. It's servable via SGLang or vLLM with Qwen3-style reasoning and tool-call parsers, and is released under the Apache License 2.0.
Model Files
File Name | Quant Type | File Size | File Link |
|-----------|------------|-----------|-----------|
| NeoHorse-1-9B.BF16.gguf | BF16 | 17.9 GB | Download |
| NeoHorse-1-9B.Q3_K_L.gguf | Q3_K_L | 4.93 GB | Download |
| NeoHorse-1-9B.Q3_K_M.gguf | Q3_K_M | 4.62 GB | Download |
| NeoHorse-1-9B.Q3_K_S.gguf | Q3_K_S | 4.26 GB | Download |
| NeoHorse-1-9B.Q4_0.gguf | Q4_0 | 5.31 GB | Download |
| NeoHorse-1-9B.Q4_K_M.gguf | Q4_K_M | 5.63 GB | Download |
| NeoHorse-1-9B.Q4_K_S.gguf | Q4_K_S | 5.35 GB | Download |
| NeoHorse-1-9B.Q5_0.gguf | Q5_0 | 6.31 GB | Download |
| NeoHorse-1-9B.Q5_K_M.gguf | Q5_K_M | 6.47 GB | Download |
| NeoHorse-1-9B.Q5_K_S.gguf | Q5_K_S | 6.31 GB | Download |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
Run prithivMLmods/NeoHorse-1-9B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models