GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ngquocvinh/K2-Horizon-7B-GGUF overview

K2 Horizon 7B GGUF Community GGUF quantizations of IFM/K2 Horizon 7B https://huggingface.co/IFM/K2 Horizon 7B . <div align="center" style="background color: f5…

llama.cppggufk2-horizonquantizedtext-generationbase_model:IFM/K2-Horizon-7Bbase_model:quantized:IFM/K2-Horizon-7Blicense:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
5,127
Likes
1
Pipeline
text-generation

Repository Files & Downloads

16 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
K2-Horizon-7B-IQ1_M.ggufGGUFIQ1_M2.49 GBDownload
K2-Horizon-7B-IQ2_XS.ggufGGUFIQ2_XS2.90 GBDownload
K2-Horizon-7B-IQ3_M.ggufGGUFIQ3_M4.10 GBDownload
K2-Horizon-7B-IQ3_S.ggufGGUFIQ3_S4.01 GBDownload
K2-Horizon-7B-IQ4_NL.ggufGGUFIQ4_NL4.99 GBDownload
K2-Horizon-7B-IQ4_XS.ggufGGUFIQ4_XS4.76 GBDownload
K2-Horizon-7B-Q1_0.ggufGGUFQ1_01.84 GBDownload
K2-Horizon-7B-Q2_K.ggufGGUFQ2_K3.49 GBDownload
K2-Horizon-7B-Q2_K_S.ggufGGUFQ2_K_S3.31 GBDownload
K2-Horizon-7B-Q3_K_L.ggufGGUFQ3_K_L4.60 GBDownload
K2-Horizon-7B-Q3_K_M.ggufGGUFQ3_K_M4.32 GBDownload
K2-Horizon-7B-Q4_K_M.ggufGGUFQ4_K_M5.21 GBDownload
K2-Horizon-7B-Q5_K_M.ggufGGUFQ5_K_M6.02 GBDownload
K2-Horizon-7B-Q6_K.ggufGGUFQ6_K6.89 GBDownload
K2-Horizon-7B-Q8_0.ggufGGUFQ8_08.92 GBDownload
reproducibility/k2_horizon_7b_combined.imatrix.ggufGGUFGGUF5.1 MBDownload

Model Details

Model IDngquocvinh/K2-Horizon-7B-GGUF
Authorngquocvinh
Pipelinetext-generation
Licenseapache-2.0
Base modelIFM/K2-Horizon-7B
Last modified2026-09-13T13:46:42.000Z

Model README

---

license: apache-2.0

base_model: IFM/K2-Horizon-7B

base_model_relation: quantized

library_name: llama.cpp

pipeline_tag: text-generation

tags:

  • gguf
  • llama.cpp
  • k2-horizon
  • quantized
  • text-generation

---

K2-Horizon-7B GGUF

Community GGUF quantizations of IFM/K2-Horizon-7B.

<div align="center" style="background-color:#f59e0b;color:#ffffff;padding:16px 20px;border-radius:10px;line-height:1.7;">

☕ If this GGUF made your day easier, a coffee would make mine.<br>

<a href="https://ko-fi.com/ngquocvinh" style="color:#ffffff;"><strong style="color:#ffffff;">Send a coffee ☕</strong></a><br>

I build and test these releases myself. Your coffee helps keep me going.<br>

Thank you for supporting this work.

</div>

About K2-Horizon-7B

K2-Horizon-7B is IFM's 7B-core

dense K2-Horizon reasoning model. The upstream model advertises a native

524,288-token context window and supports English, Chinese, coding, reasoning,

long-context, and agentic workloads. See the [official model

card](https://huggingface.co/IFM/K2-Horizon-7B) for the upstream serving stack,

prompt conventions, reasoning settings, and evaluation protocol.

![K2-Horizon-7B benchmark results](https://huggingface.co/IFM/K2-Horizon-7B)

Upstream K2-Horizon-7B benchmark results; image and scores are from the official model card.

This is a quantization-only release. No training, fine-tuning, merging, or

weight modification other than GGUF conversion and quantization was performed.

Fidelity measurements

The table below compares each published GGUF with the BF16 reference on a

held-out WikiText pilot: eight chunks from wiki.test.raw and eight chunks

from wiki.valid.raw, with a 4,096-token context and the same K2 llama.cpp

runtime. Values are averaged across the two splits. Lower Mean KLD, ΔPPL, and

RMS Δp, and higher Top-1 agreement, indicate closer next-token behavior to

BF16. The BF16 reference mean PPL was 8.118333 in this pilot.

| File | Mean KLD ↓ | Top-1 vs BF16 ↑ | ΔPPL | RMS Δp |

|---|---:|---:|---:|---:|

| K2-Horizon-7B-Q8_0.gguf | 0.002534 | 97.673% | +0.069% | 1.366% |

| K2-Horizon-7B-Q6_K.gguf | 0.010005 | 95.362% | +0.813% | 2.579% |

| K2-Horizon-7B-Q5_K_M.gguf | 0.014020 | 94.144% | +0.915% | 3.163% |

| K2-Horizon-7B-Q4_K_M.gguf | 0.027010 | 92.120% | +2.042% | 4.267% |

| K2-Horizon-7B-IQ4_NL.gguf | 0.033619 | 90.987% | +2.631% | 4.769% |

| K2-Horizon-7B-IQ4_XS.gguf | 0.034682 | 90.917% | +2.661% | 4.927% |

| K2-Horizon-7B-Q3_K_L.gguf | 0.106811 | 83.986% | +8.714% | 8.458% |

| K2-Horizon-7B-Q3_K_M.gguf | 0.110649 | 83.741% | +9.319% | 8.619% |

| K2-Horizon-7B-Q2_K.gguf | 0.198297 | 79.156% | +18.618% | 12.053% |

| K2-Horizon-7B-Q2_K_S.gguf | 0.276119 | 76.106% | +26.527% | 14.289% |

| K2-Horizon-7B-IQ2_XS.gguf | 0.422125 | 71.077% | +46.934% | 18.678% |

| K2-Horizon-7B-IQ3_M.gguf | 2.473100 | 37.864% | +1,024.404% | 41.990% |

| K2-Horizon-7B-IQ3_S.gguf | 2.618279 | 37.292% | +1,198.271% | 42.188% |

| K2-Horizon-7B-IQ1_M.gguf | 1.376577 | 49.191% | +266.636% | 33.903% |

| K2-Horizon-7B-Q1_0.gguf | 12.476683 | 0.000% | +24,085,527.083% | 56.752% |

For a general local profile, Q4_K_M is the practical starting point in this

pilot; Q5_K_M and Q6_K retain more BF16-like next-token behavior. IQ4_NL and

IQ4_XS are compact alternatives, while Q3_K_L is the stronger Q3 option here.

The Q2, IQ2, IQ3, IQ1, and Q1 results should be treated as memory-constrained

experimental profiles and checked against the intended workload.

The compact machine-readable results are available in

reproducibility/quality-summary.tsv,

with corpus hashes, evaluation settings, and runtime provenance recorded in

reproducibility/manifest.md. This pilot

measures next-token fidelity; it is not a direct percentage of capabilities

retained and does not replace task-specific evaluation.

Quick start

./llama-cli \
  -m K2-Horizon-7B-Q4_K_M.gguf \
  --chat-template-file reproducibility/chat_template_smoke_user.jinja \
  -p 'Answer briefly in English: What is GGUF, and why is it useful for running language models locally?' \
  -n 128 -c 4096 -ngl 99

The included template is the compatible single-turn template used for the

load/generate smoke test. The upstream full tool-aware Jinja template is not

certified by this package.

Reproducibility and validation

The GGUF files were converted directly from the upstream BF16 source and each

ladder member was quantized independently with the model-specific combined

importance matrix. All 15 published files passed the load/generate smoke test.

The public package contains compact reproduction inputs and scripts; raw

conversion, quantization, smoke-test, benchmark, and perplexity logs remain

local and are intentionally not uploaded. Tool calling and the upstream

benchmark suite were not re-evaluated here.

License and attribution

The upstream model is licensed under Apache License 2.0. Preserve upstream

attribution and the license when redistributing these derivative artifacts.

These are community GGUF quantizations, not an official IFM release or

endorsement.

Checksums for published artifacts and reproducibility inputs are in

SHA256SUMS.txt. The source revision, converter/runtime

commit, calibration inputs, and validation settings are in

reproducibility/manifest.md.

The compact fidelity results are in

reproducibility/quality-summary.tsv.

Run ngquocvinh/K2-Horizon-7B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models