GraySoft
Projects Models About FAQ Contact Download guIDE →

aessedai/kimi-k2.5-gguf IQ3_S GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aessedai/kimi-k2.5-gguf overview

Updates ### 03/25/2026 I've re-quanted and uploaded new versions for the IQ2XXS, IQ22, and IQ3S quantizations. Those three are using a mixture of @eaddario's target-bpw PR along with some small changes I added to support --tensor-type overrides. The result is that these quants perform better than my previous quants, and when I measured the old quants a couple of days ago it turns out there was some pretty catastrophic issues with the IQ3S and the IQ2S specifically. These new quants measure much better and should serve as better quality replacements. I don't have specific FFN Up / Gate / Down mixtures for the IQ2XXS, IQ2S, and IQ3S quants due to how the bpw budget selection works, but I've kept most of the model in high quality like the rest of my MoE-optimized quants. ### 02/11/2026 Vision support for K2.5 has been merged into llama.cpp's master branch and no longer needs to use the PR branch. ### 02/08/2026 I've updated the PR code to address feedback and updated the mmproj files here to be compatible with the new PR code. ### 02/01/2026 moonshotai has published an updated chat_template.jinja, I have updated the GGUFs in this repository so please re-download the first shard (00001) for your desired quant.

ggufbase_model:moonshotai/Kimi-K2.5base_model:quantized:moonshotai/Kimi-K2.5endpoints_compatibleregion:usimatrixconversational
aessedai/kimi-k2.5-gguf visual
Downloads
1,929
Likes
29
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

43 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Kimi-K2.5-IQ2_S-00001-of-00008.gguf GGUF IQ2_S 6.59 MB Download
Kimi-K2.5-IQ2_S-00002-of-00008.gguf GGUF IQ2_S 44.75 GB Download
Kimi-K2.5-IQ2_S-00003-of-00008.gguf GGUF IQ2_S 46.05 GB Download
Kimi-K2.5-IQ2_S-00004-of-00008.gguf GGUF IQ2_S 46.05 GB Download
Kimi-K2.5-IQ2_S-00005-of-00008.gguf GGUF IQ2_S 46.05 GB Download
Kimi-K2.5-IQ2_S-00006-of-00008.gguf GGUF IQ2_S 45.57 GB Download
Kimi-K2.5-IQ2_S-00007-of-00008.gguf GGUF IQ2_S 46.05 GB Download
Kimi-K2.5-IQ2_S-00008-of-00008.gguf GGUF IQ2_S 37.19 GB Download
Kimi-K2.5-IQ2_XXS-00001-of-00007.gguf GGUF IQ2_XXS 6.59 MB Download
Kimi-K2.5-IQ2_XXS-00002-of-00007.gguf GGUF IQ2_XXS 46.22 GB Download
Kimi-K2.5-IQ2_XXS-00003-of-00007.gguf GGUF IQ2_XXS 46.36 GB Download
Kimi-K2.5-IQ2_XXS-00004-of-00007.gguf GGUF IQ2_XXS 46.36 GB Download
Kimi-K2.5-IQ2_XXS-00005-of-00007.gguf GGUF IQ2_XXS 44.93 GB Download
Kimi-K2.5-IQ2_XXS-00006-of-00007.gguf GGUF IQ2_XXS 46.14 GB Download
Kimi-K2.5-IQ2_XXS-00007-of-00007.gguf GGUF IQ2_XXS 32.74 GB Download
Kimi-K2.5-IQ3_S-00001-of-00010.gguf GGUF IQ3_S 6.59 MB Download
Kimi-K2.5-IQ3_S-00002-of-00010.gguf GGUF IQ3_S 44.90 GB Download
Kimi-K2.5-IQ3_S-00003-of-00010.gguf GGUF IQ3_S 45.94 GB Download
Kimi-K2.5-IQ3_S-00004-of-00010.gguf GGUF IQ3_S 44.14 GB Download
Kimi-K2.5-IQ3_S-00005-of-00010.gguf GGUF IQ3_S 44.14 GB Download
Kimi-K2.5-IQ3_S-00006-of-00010.gguf GGUF IQ3_S 44.14 GB Download
Kimi-K2.5-IQ3_S-00007-of-00010.gguf GGUF IQ3_S 44.14 GB Download
Kimi-K2.5-IQ3_S-00008-of-00010.gguf GGUF IQ3_S 44.14 GB Download
Kimi-K2.5-IQ3_S-00009-of-00010.gguf GGUF IQ3_S 45.18 GB Download
Kimi-K2.5-IQ3_S-00010-of-00010.gguf GGUF IQ3_S 20.76 GB Download
Kimi-K2.5-Q4_X-00001-of-00014.gguf GGUF Q4_X 6.59 MB Download
Kimi-K2.5-Q4_X-00002-of-00014.gguf GGUF Q4_X 44.92 GB Download
Kimi-K2.5-Q4_X-00003-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00004-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00005-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00006-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00007-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00008-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00009-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00010-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00011-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00012-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00013-of-00014.gguf GGUF Q4_X 45.07 GB Download
Kimi-K2.5-Q4_X-00014-of-00014.gguf GGUF Q4_X 2.97 GB Download
mmproj-Kimi-K2.5-BF16.gguf GGUF BF16 Unknown Download
mmproj-Kimi-K2.5-F16.gguf GGUF F16 Unknown Download
mmproj-Kimi-K2.5-F32.gguf GGUF F32 Unknown Download
mmproj-Kimi-K2.5-Q8_0.gguf GGUF Unknown Download

Model Details Live

Model Slug
aessedai/kimi-k2.5-gguf
Author
AesSedai
Pipeline Task
Library
Created
2026-01-28
Last Modified
2026-03-26
Gated
No
Private
No
HF SHA
256f0de191a646e83ee491bf1206a19e003971a5
License
Unknown
Language
Unknown
Base Model
moonshotai/Kimi-K2.5

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "moonshotai/Kimi-K2.5"
    ],
    "frontmatter": {
      "base_model": [
        "moonshotai/Kimi-K2.5"
      ]
    },
    "hero_image_url": "kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\"",
    "summary": "## Updates ### 03/25/2026 I've re-quanted and uploaded new versions for the IQ2_XXS, IQ2_2, and IQ3_S quantizations. Those three are using a mixture of @eaddario's target-bpw PR along with some small changes I added to support --tensor-type overrides. The result is that these quants perform better than my previous quants, and when I measured the old quants a couple of days ago it turns out there was some pretty catastrophic issues with the IQ3_S and the IQ2_S specifically. These new quants measure much better and should serve as better quality replacements. I don't have specific FFN Up / Gate / Down mixtures for the IQ2_XXS, IQ2_S, and IQ3_S quants due to how the bpw budget selection works, but I've kept most of the model in high quality like the rest of my MoE-optimized quants. ### 02/11/2026 Vision support for K2.5 has been merged into llama.cpp's master branch and no longer needs to use the PR branch. ### 02/08/2026 I've updated the PR code to address feedback and updated the mmproj files here to be compatible with the new PR code. ### 02/01/2026 moonshotai has published an updated chat_template.jinja, I have updated the GGUFs in this repository so please re-download the first shard (00001) for your desired quant.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- moonshotai/Kimi-K2.5\n---\n## Updates\n### 03/25/2026\nI've re-quanted and uploaded new versions for the IQ2_XXS, IQ2_2, and IQ3_S quantizations. Those three are using a mixture of @eaddario's [target-bpw PR](https://github.com/ggml-org/llama.cpp/pull/12511) along with some small changes I added to support `--tensor-type` overrides.\n\nThe result is that these quants perform better than my previous quants, and when I measured the old quants a couple of days ago it turns out there was some pretty catastrophic issues with the IQ3_S and the IQ2_S specifically. These new quants measure much better and should serve as better quality replacements.\n\nI don't have specific FFN Up / Gate / Down mixtures for the IQ2_XXS, IQ2_S, and IQ3_S quants due to how the bpw budget selection works, but I've kept most of the model in high quality like the rest of my MoE-optimized quants.\n\n### 02/11/2026\nVision support for K2.5 has been merged into llama.cpp's master branch and no longer needs to use the PR branch.\n\n### 02/08/2026\nI've updated the PR code to address feedback and updated the mmproj files here to be compatible with the new PR code.\n\n### 02/01/2026\n moonshotai has published an updated chat_template.jinja, I have updated the GGUFs in this repository so please re-download the first shard (00001) for your desired quant.\n  - The default system prompt might cause confusion to users and unexpected behaviours, so we remove it.\n  - The token <|media_start|> is incorrect; it has been replaced with <|media_begin|> in the chat template.\n\n## Model\nThis is a text-and-image-only GGUF quantization of moonshotai/Kimi-K2.5. This means that video input is not present in this GGUF, and will not be available until support is added upstream in llama.cpp.\n\nMMPROJ files for image vision input have been provided, and support has been merged into the llama.cpp master branch recently.\n\nThis Q4_X quant is the \"full quality\" equivalent since the conditional experts are natively INT4 quantized directly from the original model, and the rest of the model is Q8_0. I also produced and tested a Q8_0 / Q4_K quant, the model size was identical and the PPL was barely higher. Their performance was about the same so I've only uploaded the Q4_X variant.\n\n| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |\n| :--------- | :--------- | :------- | :------- | :------- | :------- |\n| Q4_X | 543.62 GiB (4.55 BPW) | Q8_0 / Q4_0 | 1.8248 +/- 0.00699 | 0 | 0 |\n| IQ3_S | 377.50 GiB (3.16 BPW) | Q8_0 / varies | 2.116713 ± 0.008620 | +16.0796% | 0.158551 ± 0.001084 |\n| IQ2_S | 311.71 GiB (2.61 BPW) | Q8_0 / varies | 2.433594 ± 0.010455 | +33.4572% | 0.294937 ± 0.001721 |\n| IQ2_XXS | 262.74 GiB (2.20 BPW) | Q8_0 / varies | 3.119876 ± 0.014508 | +71.0926% | 0.540149 ± 0.002570 |\n\n![kld_graph](kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\")\n![ppl_graph](kld_data/02_ppl_vs_filesize.png \"Chart showing Pareto PPL analysis of quants\")",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:moonshotai/Kimi-K2.5",
    "base_model:quantized:moonshotai/Kimi-K2.5",
    "endpoints_compatible",
    "region:us",
    "imatrix",
    "conversational"
  ],
  "likes": 29,
  "downloads": 1929,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-26T06:41:52.000Z",
  "created_at": "2026-01-28T04:36:10.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "6979923add1af3c6590bea4f",
  "id": "AesSedai/Kimi-K2.5-GGUF",
  "modelId": "AesSedai/Kimi-K2.5-GGUF",
  "sha": "256f0de191a646e83ee491bf1206a19e003971a5",
  "createdAt": "2026-01-28T04:36:10.000Z",
  "lastModified": "2026-03-26T06:41:52.000Z",
  "author": "AesSedai",
  "downloads": 1929,
  "likes": 29,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 51
}