GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/NeoHorse-1-4B-GGUF overview

NeoHorse 1 4B GGUF NeoHorse 1 4B https://huggingface.co/TokenRhythm/NeoHorse 1 4B is a 4 billion parameter causal language model from TokenRhythm, post trained…

transformersgguftext-generation-inferencellama-cppagentictool-usecodingreasoning - instruction-followingtext-generationenbase_model:TokenRhythm/NeoHorse-1-4Bbase_model:quantized:TokenRhythm/NeoHorse-1-4Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.93 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

10 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
NeoHorse-1-4B.BF16.ggufGGUFGGUF7.85 GBDownload
NeoHorse-1-4B.Q3_K_L.ggufGGUFGGUF2.26 GBDownload
NeoHorse-1-4B.Q3_K_M.ggufGGUFGGUF2.11 GBDownload
NeoHorse-1-4B.Q3_K_S.ggufGGUFGGUF1.93 GBDownload
NeoHorse-1-4B.Q4_0.ggufGGUFGGUF2.37 GBDownload
NeoHorse-1-4B.Q4_K_M.ggufGGUFGGUF2.52 GBDownload
NeoHorse-1-4B.Q4_K_S.ggufGGUFGGUF2.39 GBDownload
NeoHorse-1-4B.Q5_0.ggufGGUFGGUF2.78 GBDownload
NeoHorse-1-4B.Q5_K_M.ggufGGUFGGUF2.86 GBDownload
NeoHorse-1-4B.Q5_K_S.ggufGGUFGGUF2.78 GBDownload

Model Details

Model IDprithivMLmods/NeoHorse-1-4B-GGUF
AuthorprithivMLmods
Pipelinetext-generation
Licenseapache-2.0
Base modelTokenRhythm/NeoHorse-1-4B
Last modified2026-09-09T04:17:07.000Z

Model README

---

license: apache-2.0

base_model:

  • TokenRhythm/NeoHorse-1-4B

pipeline_tag: text-generation

language:

  • en

library_name: transformers

tags:

  • text-generation-inference
  • llama-cpp
  • agentic
  • tool-use
  • coding
  • reasoning

- instruction-following

---

NeoHorse-1-4B-GGUF

> NeoHorse-1-4B is a 4-billion-parameter causal language model from TokenRhythm, post-trained from Qwen3.5-4B as an initial prototype on the path toward recursive self-improvement (RSI), targeting text-based agent harnesses, tool use, coding, and instruction following. Like its 9B sibling, its core innovation is a routing harness that assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and feeds that signal back into shaping the next training mixture via routing-guided curriculum SFT and on-policy distillation, backed by rigorous deduplication, decontamination, and six-dimensional semantic data evaluation. This release contains language-model weights only (vision weights excluded, repackaged for text-only inference) and retains a 262,144-token native context window extensible to 1,010,000. Across a ten-benchmark evaluation against Qwen3.5-4B, Gemma-4-E4B-it, Nanbeige-4.2-3B, Agents-A1-4B, and Spark-X2.5-4B, NeoHorse-1-4B posts the best overall macro average (64.87 vs. 58.94 for its Qwen3.5-4B base, a +5.93 gain), with the largest improvements on agentic benchmarks like VitaBench (+10.50), WorkBuddy Bench (+9.79), HumanEval (+9.75), and QwenClawBench (+6.21), alongside a strong tau2-Bench score of 88.46 (best in the comparison set). It's servable via SGLang or vLLM with Qwen3-style reasoning and tool-call parsers, and is released under the Apache License 2.0.

Model Files

File Name | Quant Type | File Size | File Link |

|-----------|------------|-----------|-----------|

| NeoHorse-1-4B.BF16.gguf | BF16 | 8.42 GB | Download |

| NeoHorse-1-4B.Q3_K_L.gguf | Q3_K_L | 2.42 GB | Download |

| NeoHorse-1-4B.Q3_K_M.gguf | Q3_K_M | 2.26 GB | Download |

| NeoHorse-1-4B.Q3_K_S.gguf | Q3_K_S | 2.07 GB | Download |

| NeoHorse-1-4B.Q4_0.gguf | Q4_0 | 2.54 GB | Download |

| NeoHorse-1-4B.Q4_K_M.gguf | Q4_K_M | 2.71 GB | Download |

| NeoHorse-1-4B.Q4_K_S.gguf | Q4_K_S | 2.56 GB | Download |

| NeoHorse-1-4B.Q5_0.gguf | Q5_0 | 2.99 GB | Download |

| NeoHorse-1-4B.Q5_K_M.gguf | Q5_K_M | 3.07 GB | Download |

| NeoHorse-1-4B.Q5_K_S.gguf | Q5_K_S | 2.99 GB | Download |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/NeoHorse-1-4B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models