ngquocvinh/NeoHorse-1-9B-GGUF overview
NeoHorse 1 9B GGUF Community GGUF quantizations of TokenRhythm/NeoHorse 1 9B https://huggingface.co/TokenRhythm/NeoHorse 1 9B . <div align="center" style="back…
Runs locally from ~4.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| NeoHorse-1-9B-IQ1_M.gguf | GGUF | IQ1_M | 2.68 GB | Download |
| NeoHorse-1-9B-IQ2_XS.gguf | GGUF | IQ2_XS | 3.06 GB | Download |
| NeoHorse-1-9B-IQ3_M.gguf | GGUF | IQ3_M | 4.11 GB | Download |
| NeoHorse-1-9B-IQ3_S.gguf | GGUF | IQ3_S | 4.07 GB | Download |
| NeoHorse-1-9B-IQ4_NL.gguf | GGUF | IQ4_NL | 5.05 GB | Download |
| NeoHorse-1-9B-IQ4_XS.gguf | GGUF | IQ4_XS | 4.84 GB | Download |
| NeoHorse-1-9B-Q1_0.gguf | GGUF | Q1_0 | 1.83 GB | Download |
| NeoHorse-1-9B-Q2_K.gguf | GGUF | Q2_K | 3.56 GB | Download |
| NeoHorse-1-9B-Q2_K_S.gguf | GGUF | Q2_K_S | 3.44 GB | Download |
| NeoHorse-1-9B-Q3_K_L.gguf | GGUF | Q3_K_L | 4.59 GB | Download |
| NeoHorse-1-9B-Q3_K_M.gguf | GGUF | Q3_K_M | 4.31 GB | Download |
| NeoHorse-1-9B-Q3_K_S.gguf | GGUF | Q3_K_S | 3.97 GB | Download |
| NeoHorse-1-9B-Q4_K_M.gguf | GGUF | Q4_K_M | 5.24 GB | Download |
| NeoHorse-1-9B-Q4_K_S.gguf | GGUF | Q4_K_S | 4.98 GB | Download |
| NeoHorse-1-9B-Q5_K_M.gguf | GGUF | Q5_K_M | 6.02 GB | Download |
| NeoHorse-1-9B-Q5_K_S.gguf | GGUF | Q5_K_S | 5.87 GB | Download |
| NeoHorse-1-9B-Q6_K.gguf | GGUF | Q6_K | 6.85 GB | Download |
| NeoHorse-1-9B-Q8_0.gguf | GGUF | Q8_0 | 8.87 GB | Download |
| calibration/NeoHorse-1-9B-combined.imatrix.gguf | GGUF | GGUF | 4.9 MB | Download |
Model Details
| Model ID | ngquocvinh/NeoHorse-1-9B-GGUF |
|---|---|
| Author | ngquocvinh |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | TokenRhythm/NeoHorse-1-9B |
| Last modified | 2026-09-14T02:52:26.000Z |
Model README
---
license: apache-2.0
base_model: TokenRhythm/NeoHorse-1-9B
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- qwen3.5
- quantized
- text-generation
- agentic
- tool-use
- coding
- reasoning
- instruction-following
---
NeoHorse-1-9B GGUF
Community GGUF quantizations of TokenRhythm/NeoHorse-1-9B.
<div align="center" style="background-color:#f59e0b;color:#ffffff;padding:16px 20px;border-radius:10px;line-height:1.7;">
☕ If this GGUF made your day easier, a coffee would make mine.<br>
<a href="https://ko-fi.com/ngquocvinh" style="color:#ffffff;"><strong style="color:#ffffff;">Send a coffee ☕</strong></a><br>
I build and test these releases myself. Your coffee helps keep me going.<br>
Thank you for supporting this work.
</div>
About NeoHorse-1-9B
NeoHorse-1-9B is an approximately 9B causal language model post-trained from
Qwen3.5-9B for text-based agent harnesses, tool use, coding, reasoning, and
instruction following. This package contains text-only GGUF files; vision
weights are not included. See the official upstream model card
for the model's original details, usage, and limitations.
The upstream card describes a native 262,144-token context and extensibility up
to 1,010,000 tokens. The upstream card also reports a 69.04 ten-benchmark
average versus 65.60 for Qwen3.5-9B under its SGLang evaluation protocol. Those
figures are upstream measurements, not measurements of these GGUF files.
<p align="center">
<a href="https://huggingface.co/TokenRhythm/NeoHorse-1-9B">
<img src="./9B_head_fig.jpg" alt="NeoHorse-1-9B evaluation results" width="100%">
</a>
</p>
<p align="center"><em><a href="https://huggingface.co/TokenRhythm/NeoHorse-1-9B">NeoHorse-1-9B evaluation results</a> (upstream)</em></p>
This is a quantization-only release. No training or fine-tuning was performed
as part of this conversion.
Fidelity measurements
The table compares every published quantization with the converted BF16
reference on the same hold-out split. Evaluation used the first two chunks of
wiki.valid.raw, 2,048 tokens per chunk, and the CUDA build of llama.cpp on an
NVIDIA A10M. The BF16 baseline mean PPL was 5.641047. Values are averaged over
the two evaluated chunks. File sizes use decimal GB (bytes / 1,000,000,000).
Bold rows are practical sweet spots: Q4_K_M is the balanced default, Q5_K_M
and Q5_K_S provide more fidelity headroom, and IQ4_XS is a compact alternative.
| File | Size (GB) | Mean KLD ↓ | Top-1 vs BF16 ↑ | ΔPPL | RMS Δp |
|---|---:|---:|---:|---:|---:|
| NeoHorse-1-9B-Q8_0.gguf | 9.528 | 0.006655 | 98.436% | +0.732% | 0.811% |
| NeoHorse-1-9B-Q6_K.gguf | 7.359 | 0.009702 | 97.312% | +1.120% | 1.408% |
| NeoHorse-1-9B-Q5_K_M.gguf | 6.468 | 0.030926 | 94.673% | +2.158% | 5.841% |
| NeoHorse-1-9B-Q4_K_M.gguf | 5.629 | 0.029467 | 94.428% | +0.951% | 4.429% |
| NeoHorse-1-9B-Q3_K_M.gguf | 4.624 | 0.069233 | 88.368% | +4.573% | 7.326% |
| NeoHorse-1-9B-Q2_K.gguf | 3.827 | 0.199997 | 82.063% | +14.167% | 12.510% |
| NeoHorse-1-9B-IQ2_XS.gguf | 3.285 | 0.342380 | 77.761% | +30.698% | 18.365% |
| NeoHorse-1-9B-IQ1_M.gguf | 2.877 | 0.807561 | 64.565% | +96.977% | 28.701% |
| NeoHorse-1-9B-Q1_0.gguf | 1.966 | 15.570850 | 2.933% | +596289635.160% | 63.365% |
| NeoHorse-1-9B-Q5_K_S.gguf | 6.305 | 0.030149 | 95.112% | +1.381% | 5.817% |
| NeoHorse-1-9B-Q4_K_S.gguf | 5.352 | 0.030445 | 93.842% | +1.040% | 4.639% |
| NeoHorse-1-9B-Q3_K_L.gguf | 4.926 | 0.062186 | 89.541% | +3.793% | 6.790% |
| NeoHorse-1-9B-Q3_K_S.gguf | 4.259 | 0.131594 | 85.386% | +6.965% | 9.989% |
| NeoHorse-1-9B-Q2_K_S.gguf | 3.697 | 0.240966 | 81.183% | +21.317% | 14.223% |
| NeoHorse-1-9B-IQ4_XS.gguf | 5.196 | 0.026534 | 93.891% | +2.125% | 3.807% |
| NeoHorse-1-9B-IQ4_NL.gguf | 5.418 | 0.029899 | 93.695% | +3.189% | 4.788% |
| NeoHorse-1-9B-IQ3_M.gguf | 4.415 | 0.064802 | 88.905% | +2.843% | 7.357% |
| NeoHorse-1-9B-IQ3_S.gguf | 4.371 | 0.081996 | 89.541% | +2.072% | 7.758% |
ΔPPL is the percentage change in perplexity relative to BF16. Lower KLD,
ΔPPL, and RMS Δp, together with higher Top-1 agreement, indicate closer
next-token behavior to BF16. These are next-token fidelity measurements, not a
task benchmark; reasoning, coding, and tool-use behavior can vary by workload.
The machine-readable results are in reproducibility/quality-summary.tsv,
with the evaluation details in reproducibility/manifest.md.
Quick start
Q4_K_M is a practical starting point for local use. Q5_K_M, Q6_K, or Q8_0
provide more headroom for fidelity when memory allows. Q3_K_M and lower are
low-memory experimental options; the measurements above show increasing
distance from BF16.
./llama-cli \
-m NeoHorse-1-9B-Q4_K_M.gguf \
--chat-template-file chat_template.jinja \
--jinja \
--reasoning off \
-p 'Answer briefly in English: What is GGUF and why is it useful for running language models locally?' \
-n 128 -c 4096 -ngl 99
The included chat_template.jinja is retained from the upstream text-only
checkpoint. Adjust context size, GPU layers, and batch settings for the
available hardware.
Reproducibility and validation
- Source: upstream revision
ba5b6e4. - The BF16 GGUF was converted directly from the upstream BF16 safetensors with
--no-nextn, because this text-only checkpoint has no MTP tensors even though
its config retains an MTP layer-count field.
- Every published file was quantized directly from that BF16 GGUF with the
model-specific imatrix included under calibration/; no re-quantization from
another quantized file was used.
- All published files passed load and short English generation smoke tests; the BF16
source additionally passed --check-tensors. Runtime readings are summarized in
reproducibility/runtime-summary.tsv.
- Verify artifacts with
SHA256SUMS.txtand reproduce the
provenance from reproducibility/manifest.md.
Raw conversion, calibration, quantization, smoke-test, fidelity, and benchmark
logs are kept local and are not part of this public model package.
License and attribution
The upstream model and these community quantizations are released under the
Apache License 2.0, with the
upstream LICENSE retained in this repository.
These are community GGUF quantizations, not an official TokenRhythm release or
endorsement.
Run ngquocvinh/NeoHorse-1-9B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models