batiai/Kimi-K2.6-GGUF overview
Kimi K2.6 GGUF — Quantized by BatiAI <p align="center" <a href="https://flow.bati.ai" <img src="https://img.shields.io/badge/BatiFlow on device%20AI blue?style…
Runs locally from ~12.19 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| moonshotai-Kimi-K2.6-IQ3_XXS-00001-of-00009.gguf | GGUF | IQ3_XXS | 44.11 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00002-of-00009.gguf | GGUF | IQ3_XXS | 44.68 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00003-of-00009.gguf | GGUF | IQ3_XXS | 44.69 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00004-of-00009.gguf | GGUF | IQ3_XXS | 44.68 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00005-of-00009.gguf | GGUF | IQ3_XXS | 42.70 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00006-of-00009.gguf | GGUF | IQ3_XXS | 44.68 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00007-of-00009.gguf | GGUF | IQ3_XXS | 44.69 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00008-of-00009.gguf | GGUF | IQ3_XXS | 44.68 GB | Download |
| moonshotai-Kimi-K2.6-IQ3_XXS-00009-of-00009.gguf | GGUF | IQ3_XXS | 12.19 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00001-of-00012.gguf | GGUF | IQ4_XS | 44.03 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00002-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00003-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00004-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00005-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00006-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00007-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00008-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00009-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00010-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00011-of-00012.gguf | GGUF | IQ4_XS | 42.25 GB | Download |
| moonshotai-Kimi-K2.6-IQ4_XS-00012-of-00012.gguf | GGUF | IQ4_XS | 42.20 GB | Download |
Model Details
| Model ID | batiai/Kimi-K2.6-GGUF |
|---|---|
| Author | batiai |
| Pipeline | text-generation |
| License | other |
| Base model | moonshotai/Kimi-K2.6 |
| Last modified | 2026-08-15T02:24:34.000Z |
Model README
---
language:
- en
- ko
- ja
- zh
license: other
license_name: modified-mit
license_link: https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE
tags:
- gguf
- kimi
- moonshot
- quantized
- batiai
- mixture-of-experts
- frontier
- 1t
- agent
base_model: moonshotai/Kimi-K2.6
pipeline_tag: text-generation
library_name: llama.cpp
---
Kimi K2.6 GGUF — Quantized by BatiAI
<p align="center">
<a href="https://flow.bati.ai"><img src="https://img.shields.io/badge/BatiFlow-on--device%20AI-blue?style=for-the-badge&logo=apple" alt="BatiFlow"></a>
<a href="https://ollama.com/batiai/kimi-k2.6"><img src="https://img.shields.io/badge/Ollama-batiai%2Fkimi--k2.6-green?style=for-the-badge" alt="Ollama"></a>
</p>
> IQ3_XXS / IQ4_XS quantization of moonshotai/Kimi-K2.6 (1T total / 32B active MoE).
> Quantized directly from official Moonshot FP8 weights by BatiAI.
Why Kimi K2.6?
- 1T parameters (32B active) — frontier-class open weight model
- SWE-Bench Pro 58.6 — beats GPT-5.4 xhigh (57.7), Claude Opus 4.6 max (53.4), Gemini 3.1 Pro (54.2)
- HLE 36.4% (no tools) / 55.5% (w/ tools) — Humanity's Last Exam frontier tier
- Agent swarm architecture — 300 sub-agents, 4,000 coordinated steps
- 256K native context (262,144 tokens) via YARN scaling
- Native tool calling — search, code-interpreter, web-browsing
- Modified-MIT license — redistribution + fine-tuning allowed
- Released 2026-04-20 by Moonshot AI
Quick Start
# IQ4_XS (recommended balance, 546GB, M3 Ultra 512GB+)
ollama pull batiai/kimi-k2.6:iq4
# IQ3_XXS (smaller, 394GB, 384GB+ RAM)
ollama pull batiai/kimi-k2.6:iq3
# Q5_K_M (highest quality, 728GB, needs 768GB+ RAM)
ollama pull batiai/kimi-k2.6:q5
Available Quantizations
| Quant | Size | Min RAM | Target Hardware | Notes |
|-------|------|---------|-----------------|-------|
| IQ3_XXS | 394GB | 384GB | M3 Ultra 512GB / H100 node | aggressive compression, imatrix-calibrated |
| IQ4_XS | 546GB | 512GB | M3 Ultra 512GB / 8×A100 80GB | recommended balance |
| Q5_K_M | 728GB | 768GB | 2× M3 Ultra / 8×A100 80GB / H100 node | highest quality, near-original |
> ⚠️ Not for consumer Mac — this is a workstation / server / frontier research model. 16-128GB Macs should use batiai/qwen3.6-35b or batiai/minimax-m2.7 instead (see comparison table below).
Hardware Reality Check
| Your System | IQ3 (394GB) | IQ4 (546GB) | Q5 (728GB) |
|-------------|-------------|-------------|-------------|
| Mac 128GB | ❌ Won't fit | ❌ | ❌ |
| Mac 192GB | ❌ Won't fit | ❌ | ❌ |
| Mac 256GB | ⚠️ Heavy swap (unusable) | ❌ | ❌ |
| Mac 384GB | ⚠️ Tight | ❌ | ❌ |
| Mac M3 Ultra 512GB | ✅ Comfortable | ✅ Usable (tight) | ❌ |
| 2× M3 Ultra (cluster) | ✅ | ✅ | ✅ |
| 8× A100 80GB (640GB total) | ✅ | ✅ Fast | ✅ |
| H100 node (640GB+) | ✅ Fast | ✅ Fast | ✅ Fast |
Numbers based on MoE activation patterns — 32B active params × 4 bytes buffer ≈ 130GB runtime even after quantization, plus shard headers + KV cache (at 256K context, cache alone is 30-80GB).
What BatiAI's Quantization Delivers
| | BatiAI | unsloth / ubergarm |
|---|---|---|
| Source | Direct from official Moonshot FP8 weights | Same (major providers) |
| Quantization flow | FP8 → Q8_0 → IQ3_XXS/IQ4_XS with imatrix (wikitext-2 calibration, 200 chunks) | Similar |
| imatrix | ✅ 200 chunks (quality saturation point) | Varies |
| Tool-calling preservation | ✅ Native template preserved | ✅ |
| Korean validation | ✅ (pending benchmark on target hardware) | ✗ |
| BatiAI signature | ✅ general.author=BatiAI, general.url=https://flow.bati.ai | ✗ |
| Pipeline | Open source — docs/202604-large-moe-quantization.md | Internal |
Model Comparison — BatiAI Model Lineup
Kimi K2.6 is for frontier workstation users. For everyone else:
| Your Hardware | Best BatiAI Model | Size |
|---------------|-------------------|------|
| 16GB Mac | batiai/gemma4-e4b:q4 | 4.9GB |
| 24GB Mac | batiai/gemma4-26b:iq4 | 15GB |
| 48GB Mac | batiai/qwen3.5-35b:iq4 | 22GB |
| 96GB Mac | batiai/qwen3.6-35b:iq4 | 22GB |
| 128GB Mac | batiai/minimax-m2.7:iq3 | 82GB |
| M3 Ultra 512GB / H100 | batiai/kimi-k2.6:iq4 | 509GB |
Benchmarks (source model)
Benchmark numbers from Moonshot AI's official report — validating that aggressive quantization preserves these capabilities is pending on our end (bench.sh on M3 Ultra / H100 target).
| Benchmark | Kimi K2.6 | Comparison |
|-----------|-----------|------------|
| SWE-Bench Pro | 58.6 | GPT-5.4 xhigh 57.7, Opus 4.6 max 53.4 |
| HLE (no tools) | 36.4% | frontier tier |
| HLE (w/ tools) | 55.5% | frontier tier |
| Context | 256K | YARN scaling |
| Native tool use | ✅ | search, code, web |
Technical Details
- Original Model: moonshotai/Kimi-K2.6
- Architecture: Mixture of Experts — 1T total / 32B active, 61 layers, 384 experts (8 selected + 1 shared), MLA attention
- Original storage: FP8 / INT4 hybrid QAT (555GB)
- License: Modified-MIT
- Quantized with: llama.cpp
- Calibration: wikitext-2-raw, 200 chunks (quality saturation)
- Quantized by: BatiAI
Usage
llama.cpp
./llama-cli -m Kimi-K2.6-IQ4_XS.gguf \
-p "Your prompt" \
--ctx-size 65536 \
--n-gpu-layers 99
Ollama
ollama run batiai/kimi-k2.6:iq4
vLLM / TGI
Not directly compatible — these serve FP8/BF16 safetensors. Use original moonshotai/Kimi-K2.6 for vLLM.
About BatiAI
BatiAI quantizes frontier open weight models with validated quality and transparent provenance. We built BatiFlow — free, on-device AI automation for Mac — and open-source our full quantization pipeline.
The Kimi K2.6 release demonstrates our pipeline handles 1T+ MoE models (most quantization providers stop at 70B). See our Kimi K2.6 quantization notes for the engineering trade-offs.
License
Quantized from moonshotai/Kimi-K2.6. License: Modified-MIT — commercial use + redistribution allowed.
Run batiai/Kimi-K2.6-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models