GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

darioooooo0o/K2-Horizon-0.9B-GGUF overview

K2 Horizon 0.9B GGUF quants X https://img.shields.io/badge/X Follow 000000?logo=x&logoColor=white https://x.com/imdariotoo Requests, questions or suggestions? …

ggufk2-horizonllama.cppk-quantstext-generationbase_model:IFM/K2-Horizon-0.9Bbase_model:quantized:IFM/K2-Horizon-0.9Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~524.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
k2horizon-q3_k_m.ggufGGUFQ3_K_M524.4 MBDownload
k2horizon-q4_k_m.ggufGGUFQ4_K_M635.3 MBDownload
k2horizon-q4_k_s.ggufGGUFQ4_K_S608.7 MBDownload
k2horizon-q5_k_m.ggufGGUFQ5_K_M737.7 MBDownload
k2horizon-q8_0.ggufGGUFQ8_01.07 GBDownload

Model Details

Model IDdarioooooo0o/K2-Horizon-0.9B-GGUF
Authordarioooooo0o
Pipelinetext-generation
Licenseapache-2.0
Base modelIFM/K2-Horizon-0.9B
Last modified2026-09-03T18:56:39.000Z

Model README

---

license: apache-2.0

base_model: IFM/K2-Horizon-0.9B

pipeline_tag: text-generation

library_name: gguf

tags:

  • k2-horizon
  • llama.cpp
  • gguf
  • k-quants

---

K2-Horizon-0.9B GGUF quants

![X](https://x.com/imdariotoo)

Requests, questions or suggestions? Message me on X: https://x.com/imdariotoo

GGUF quantizations of IFM/K2-Horizon-0.9B.

Converted with the official k2-official llama.cpp branch (MBZUAI-IFM port, commit 35999d101).

Files

| File | Quant | Size |

|---|---|---|

| k2horizon-q3_k_m.gguf | Q3_K_M | ~0.5 GB |

| k2horizon-q4_k_s.gguf | Q4_K_S | ~0.6 GB |

| k2horizon-q4_k_m.gguf | Q4_K_M | ~0.6 GB |

| k2horizon-q5_k_m.gguf | Q5_K_M | ~0.7 GB |

| k2horizon-q8_0.gguf | Q8_0 | ~1.1 GB |

Requirements

Use a llama.cpp build from the k2-official branch of MBZUAI-IFM/llama.cpp (or anything that merges that port). Mainline llama.cpp does NOT support the k2-horizon architecture.

Usage

llama-cli -m k2horizon-q4_k_m.gguf -ngl 99 -c 8192

Fits entirely on any modern GPU (even iGPU); no CPU offload needed.

Notes

  • Plain K-quants from BF16, no imatrix.
  • All quants verified loading and generating on RTX 3060 12GB.

Run darioooooo0o/K2-Horizon-0.9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models