ngquocvinh/K2-Horizon-7B-GGUF overview
K2 Horizon 7B GGUF Community GGUF quantizations of IFM/K2 Horizon 7B https://huggingface.co/IFM/K2 Horizon 7B . <div align="center" style="background color: f5…
Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| K2-Horizon-7B-IQ1_M.gguf | GGUF | IQ1_M | 2.49 GB | Download |
| K2-Horizon-7B-IQ2_XS.gguf | GGUF | IQ2_XS | 2.90 GB | Download |
| K2-Horizon-7B-IQ3_M.gguf | GGUF | IQ3_M | 4.10 GB | Download |
| K2-Horizon-7B-IQ3_S.gguf | GGUF | IQ3_S | 4.01 GB | Download |
| K2-Horizon-7B-IQ4_NL.gguf | GGUF | IQ4_NL | 4.99 GB | Download |
| K2-Horizon-7B-IQ4_XS.gguf | GGUF | IQ4_XS | 4.76 GB | Download |
| K2-Horizon-7B-Q1_0.gguf | GGUF | Q1_0 | 1.84 GB | Download |
| K2-Horizon-7B-Q2_K.gguf | GGUF | Q2_K | 3.49 GB | Download |
| K2-Horizon-7B-Q2_K_S.gguf | GGUF | Q2_K_S | 3.31 GB | Download |
| K2-Horizon-7B-Q3_K_L.gguf | GGUF | Q3_K_L | 4.60 GB | Download |
| K2-Horizon-7B-Q3_K_M.gguf | GGUF | Q3_K_M | 4.32 GB | Download |
| K2-Horizon-7B-Q4_K_M.gguf | GGUF | Q4_K_M | 5.21 GB | Download |
| K2-Horizon-7B-Q5_K_M.gguf | GGUF | Q5_K_M | 6.02 GB | Download |
| K2-Horizon-7B-Q6_K.gguf | GGUF | Q6_K | 6.89 GB | Download |
| K2-Horizon-7B-Q8_0.gguf | GGUF | Q8_0 | 8.92 GB | Download |
| reproducibility/k2_horizon_7b_combined.imatrix.gguf | GGUF | GGUF | 5.1 MB | Download |
Model Details
| Model ID | ngquocvinh/K2-Horizon-7B-GGUF |
|---|---|
| Author | ngquocvinh |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | IFM/K2-Horizon-7B |
| Last modified | 2026-09-13T13:46:42.000Z |
Model README
---
license: apache-2.0
base_model: IFM/K2-Horizon-7B
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- k2-horizon
- quantized
- text-generation
---
K2-Horizon-7B GGUF
Community GGUF quantizations of IFM/K2-Horizon-7B.
<div align="center" style="background-color:#f59e0b;color:#ffffff;padding:16px 20px;border-radius:10px;line-height:1.7;">
☕ If this GGUF made your day easier, a coffee would make mine.<br>
<a href="https://ko-fi.com/ngquocvinh" style="color:#ffffff;"><strong style="color:#ffffff;">Send a coffee ☕</strong></a><br>
I build and test these releases myself. Your coffee helps keep me going.<br>
Thank you for supporting this work.
</div>
About K2-Horizon-7B
K2-Horizon-7B is IFM's 7B-core
dense K2-Horizon reasoning model. The upstream model advertises a native
524,288-token context window and supports English, Chinese, coding, reasoning,
long-context, and agentic workloads. See the [official model
card](https://huggingface.co/IFM/K2-Horizon-7B) for the upstream serving stack,
prompt conventions, reasoning settings, and evaluation protocol.

Upstream K2-Horizon-7B benchmark results; image and scores are from the official model card.
This is a quantization-only release. No training, fine-tuning, merging, or
weight modification other than GGUF conversion and quantization was performed.
Fidelity measurements
The table below compares each published GGUF with the BF16 reference on a
held-out WikiText pilot: eight chunks from wiki.test.raw and eight chunks
from wiki.valid.raw, with a 4,096-token context and the same K2 llama.cpp
runtime. Values are averaged across the two splits. Lower Mean KLD, ΔPPL, and
RMS Δp, and higher Top-1 agreement, indicate closer next-token behavior to
BF16. The BF16 reference mean PPL was 8.118333 in this pilot.
| File | Mean KLD ↓ | Top-1 vs BF16 ↑ | ΔPPL | RMS Δp |
|---|---:|---:|---:|---:|
| K2-Horizon-7B-Q8_0.gguf | 0.002534 | 97.673% | +0.069% | 1.366% |
| K2-Horizon-7B-Q6_K.gguf | 0.010005 | 95.362% | +0.813% | 2.579% |
| K2-Horizon-7B-Q5_K_M.gguf | 0.014020 | 94.144% | +0.915% | 3.163% |
| K2-Horizon-7B-Q4_K_M.gguf | 0.027010 | 92.120% | +2.042% | 4.267% |
| K2-Horizon-7B-IQ4_NL.gguf | 0.033619 | 90.987% | +2.631% | 4.769% |
| K2-Horizon-7B-IQ4_XS.gguf | 0.034682 | 90.917% | +2.661% | 4.927% |
| K2-Horizon-7B-Q3_K_L.gguf | 0.106811 | 83.986% | +8.714% | 8.458% |
| K2-Horizon-7B-Q3_K_M.gguf | 0.110649 | 83.741% | +9.319% | 8.619% |
| K2-Horizon-7B-Q2_K.gguf | 0.198297 | 79.156% | +18.618% | 12.053% |
| K2-Horizon-7B-Q2_K_S.gguf | 0.276119 | 76.106% | +26.527% | 14.289% |
| K2-Horizon-7B-IQ2_XS.gguf | 0.422125 | 71.077% | +46.934% | 18.678% |
| K2-Horizon-7B-IQ3_M.gguf | 2.473100 | 37.864% | +1,024.404% | 41.990% |
| K2-Horizon-7B-IQ3_S.gguf | 2.618279 | 37.292% | +1,198.271% | 42.188% |
| K2-Horizon-7B-IQ1_M.gguf | 1.376577 | 49.191% | +266.636% | 33.903% |
| K2-Horizon-7B-Q1_0.gguf | 12.476683 | 0.000% | +24,085,527.083% | 56.752% |
For a general local profile, Q4_K_M is the practical starting point in this
pilot; Q5_K_M and Q6_K retain more BF16-like next-token behavior. IQ4_NL and
IQ4_XS are compact alternatives, while Q3_K_L is the stronger Q3 option here.
The Q2, IQ2, IQ3, IQ1, and Q1 results should be treated as memory-constrained
experimental profiles and checked against the intended workload.
The compact machine-readable results are available in
reproducibility/quality-summary.tsv,
with corpus hashes, evaluation settings, and runtime provenance recorded in
reproducibility/manifest.md. This pilot
measures next-token fidelity; it is not a direct percentage of capabilities
retained and does not replace task-specific evaluation.
Quick start
./llama-cli \
-m K2-Horizon-7B-Q4_K_M.gguf \
--chat-template-file reproducibility/chat_template_smoke_user.jinja \
-p 'Answer briefly in English: What is GGUF, and why is it useful for running language models locally?' \
-n 128 -c 4096 -ngl 99
The included template is the compatible single-turn template used for the
load/generate smoke test. The upstream full tool-aware Jinja template is not
certified by this package.
Reproducibility and validation
The GGUF files were converted directly from the upstream BF16 source and each
ladder member was quantized independently with the model-specific combined
importance matrix. All 15 published files passed the load/generate smoke test.
The public package contains compact reproduction inputs and scripts; raw
conversion, quantization, smoke-test, benchmark, and perplexity logs remain
local and are intentionally not uploaded. Tool calling and the upstream
benchmark suite were not re-evaluated here.
License and attribution
The upstream model is licensed under Apache License 2.0. Preserve upstream
attribution and the license when redistributing these derivative artifacts.
These are community GGUF quantizations, not an official IFM release or
endorsement.
Checksums for published artifacts and reproducibility inputs are in
SHA256SUMS.txt. The source revision, converter/runtime
commit, calibration inputs, and validation settings are in
The compact fidelity results are in
Run ngquocvinh/K2-Horizon-7B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models