GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

batiai/Kimi-K2.6-GGUF overview

Kimi K2.6 GGUF — Quantized by BatiAI <p align="center" <a href="https://flow.bati.ai" <img src="https://img.shields.io/badge/BatiFlow on device%20AI blue?style…

llama.cppggufkimimoonshotquantizedbatiaimixture-of-expertsfrontier1tagenttext-generationenkojazhbase_model:moonshotai/Kimi-K2.6base_model:quantized:moonshotai/Kimi-K2.6license:otherendpoints_compatibleregion:usimatrixconversational

Runs locally from ~12.19 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
124
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

21 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
moonshotai-Kimi-K2.6-IQ3_XXS-00001-of-00009.ggufGGUFIQ3_XXS44.11 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00002-of-00009.ggufGGUFIQ3_XXS44.68 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00003-of-00009.ggufGGUFIQ3_XXS44.69 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00004-of-00009.ggufGGUFIQ3_XXS44.68 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00005-of-00009.ggufGGUFIQ3_XXS42.70 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00006-of-00009.ggufGGUFIQ3_XXS44.68 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00007-of-00009.ggufGGUFIQ3_XXS44.69 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00008-of-00009.ggufGGUFIQ3_XXS44.68 GBDownload
moonshotai-Kimi-K2.6-IQ3_XXS-00009-of-00009.ggufGGUFIQ3_XXS12.19 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00001-of-00012.ggufGGUFIQ4_XS44.03 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00002-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00003-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00004-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00005-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00006-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00007-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00008-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00009-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00010-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00011-of-00012.ggufGGUFIQ4_XS42.25 GBDownload
moonshotai-Kimi-K2.6-IQ4_XS-00012-of-00012.ggufGGUFIQ4_XS42.20 GBDownload

Model Details

Model IDbatiai/Kimi-K2.6-GGUF
Authorbatiai
Pipelinetext-generation
Licenseother
Base modelmoonshotai/Kimi-K2.6
Last modified2026-08-15T02:24:34.000Z

Model README

---

language:

- en

- ko

- ja

- zh

license: other

license_name: modified-mit

license_link: https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE

tags:

- gguf

- kimi

- moonshot

- quantized

- batiai

- mixture-of-experts

- frontier

- 1t

- agent

base_model: moonshotai/Kimi-K2.6

pipeline_tag: text-generation

library_name: llama.cpp

---

Kimi K2.6 GGUF — Quantized by BatiAI

<p align="center">

<a href="https://flow.bati.ai"><img src="https://img.shields.io/badge/BatiFlow-on--device%20AI-blue?style=for-the-badge&logo=apple" alt="BatiFlow"></a>

<a href="https://ollama.com/batiai/kimi-k2.6"><img src="https://img.shields.io/badge/Ollama-batiai%2Fkimi--k2.6-green?style=for-the-badge" alt="Ollama"></a>

</p>

> IQ3_XXS / IQ4_XS quantization of moonshotai/Kimi-K2.6 (1T total / 32B active MoE).

> Quantized directly from official Moonshot FP8 weights by BatiAI.

Why Kimi K2.6?

  • 1T parameters (32B active) — frontier-class open weight model
  • SWE-Bench Pro 58.6 — beats GPT-5.4 xhigh (57.7), Claude Opus 4.6 max (53.4), Gemini 3.1 Pro (54.2)
  • HLE 36.4% (no tools) / 55.5% (w/ tools) — Humanity's Last Exam frontier tier
  • Agent swarm architecture — 300 sub-agents, 4,000 coordinated steps
  • 256K native context (262,144 tokens) via YARN scaling
  • Native tool calling — search, code-interpreter, web-browsing
  • Modified-MIT license — redistribution + fine-tuning allowed
  • Released 2026-04-20 by Moonshot AI

Quick Start

# IQ4_XS (recommended balance, 546GB, M3 Ultra 512GB+)
ollama pull batiai/kimi-k2.6:iq4

# IQ3_XXS (smaller, 394GB, 384GB+ RAM)
ollama pull batiai/kimi-k2.6:iq3

# Q5_K_M (highest quality, 728GB, needs 768GB+ RAM)
ollama pull batiai/kimi-k2.6:q5

Available Quantizations

| Quant | Size | Min RAM | Target Hardware | Notes |

|-------|------|---------|-----------------|-------|

| IQ3_XXS | 394GB | 384GB | M3 Ultra 512GB / H100 node | aggressive compression, imatrix-calibrated |

| IQ4_XS | 546GB | 512GB | M3 Ultra 512GB / 8×A100 80GB | recommended balance |

| Q5_K_M | 728GB | 768GB | 2× M3 Ultra / 8×A100 80GB / H100 node | highest quality, near-original |

> ⚠️ Not for consumer Mac — this is a workstation / server / frontier research model. 16-128GB Macs should use batiai/qwen3.6-35b or batiai/minimax-m2.7 instead (see comparison table below).

Hardware Reality Check

| Your System | IQ3 (394GB) | IQ4 (546GB) | Q5 (728GB) |

|-------------|-------------|-------------|-------------|

| Mac 128GB | ❌ Won't fit | ❌ | ❌ |

| Mac 192GB | ❌ Won't fit | ❌ | ❌ |

| Mac 256GB | ⚠️ Heavy swap (unusable) | ❌ | ❌ |

| Mac 384GB | ⚠️ Tight | ❌ | ❌ |

| Mac M3 Ultra 512GB | ✅ Comfortable | ✅ Usable (tight) | ❌ |

| 2× M3 Ultra (cluster) | ✅ | ✅ | ✅ |

| 8× A100 80GB (640GB total) | ✅ | ✅ Fast | ✅ |

| H100 node (640GB+) | ✅ Fast | ✅ Fast | ✅ Fast |

Numbers based on MoE activation patterns — 32B active params × 4 bytes buffer ≈ 130GB runtime even after quantization, plus shard headers + KV cache (at 256K context, cache alone is 30-80GB).

What BatiAI's Quantization Delivers

| | BatiAI | unsloth / ubergarm |

|---|---|---|

| Source | Direct from official Moonshot FP8 weights | Same (major providers) |

| Quantization flow | FP8 → Q8_0 → IQ3_XXS/IQ4_XS with imatrix (wikitext-2 calibration, 200 chunks) | Similar |

| imatrix | ✅ 200 chunks (quality saturation point) | Varies |

| Tool-calling preservation | ✅ Native template preserved | ✅ |

| Korean validation | ✅ (pending benchmark on target hardware) | ✗ |

| BatiAI signature | ✅ general.author=BatiAI, general.url=https://flow.bati.ai | ✗ |

| Pipeline | Open source — docs/202604-large-moe-quantization.md | Internal |

Model Comparison — BatiAI Model Lineup

Kimi K2.6 is for frontier workstation users. For everyone else:

| Your Hardware | Best BatiAI Model | Size |

|---------------|-------------------|------|

| 16GB Mac | batiai/gemma4-e4b:q4 | 4.9GB |

| 24GB Mac | batiai/gemma4-26b:iq4 | 15GB |

| 48GB Mac | batiai/qwen3.5-35b:iq4 | 22GB |

| 96GB Mac | batiai/qwen3.6-35b:iq4 | 22GB |

| 128GB Mac | batiai/minimax-m2.7:iq3 | 82GB |

| M3 Ultra 512GB / H100 | batiai/kimi-k2.6:iq4 | 509GB |

Benchmarks (source model)

Benchmark numbers from Moonshot AI's official report — validating that aggressive quantization preserves these capabilities is pending on our end (bench.sh on M3 Ultra / H100 target).

| Benchmark | Kimi K2.6 | Comparison |

|-----------|-----------|------------|

| SWE-Bench Pro | 58.6 | GPT-5.4 xhigh 57.7, Opus 4.6 max 53.4 |

| HLE (no tools) | 36.4% | frontier tier |

| HLE (w/ tools) | 55.5% | frontier tier |

| Context | 256K | YARN scaling |

| Native tool use | ✅ | search, code, web |

Technical Details

  • Original Model: moonshotai/Kimi-K2.6
  • Architecture: Mixture of Experts — 1T total / 32B active, 61 layers, 384 experts (8 selected + 1 shared), MLA attention
  • Original storage: FP8 / INT4 hybrid QAT (555GB)
  • License: Modified-MIT
  • Quantized with: llama.cpp
  • Calibration: wikitext-2-raw, 200 chunks (quality saturation)
  • Quantized by: BatiAI

Usage

llama.cpp

./llama-cli -m Kimi-K2.6-IQ4_XS.gguf \
  -p "Your prompt" \
  --ctx-size 65536 \
  --n-gpu-layers 99

Ollama

ollama run batiai/kimi-k2.6:iq4

vLLM / TGI

Not directly compatible — these serve FP8/BF16 safetensors. Use original moonshotai/Kimi-K2.6 for vLLM.

About BatiAI

BatiAI quantizes frontier open weight models with validated quality and transparent provenance. We built BatiFlow — free, on-device AI automation for Mac — and open-source our full quantization pipeline.

The Kimi K2.6 release demonstrates our pipeline handles 1T+ MoE models (most quantization providers stop at 70B). See our Kimi K2.6 quantization notes for the engineering trade-offs.

License

Quantized from moonshotai/Kimi-K2.6. License: Modified-MIT — commercial use + redistribution allowed.

Run batiai/Kimi-K2.6-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models