mmnga-o/Kimi-K3-REAP50-Width50-UD-gguf overview
Kimi K3 REAP50 Width50 UD GGUF 日本語 ./README jp.md Kimi K3 REAP50 and Width50 overview ./reduction overview.png This is an experimental, reduced GGUF build of K…
Runs locally from ~6.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00001-of-00015.gguf | GGUF | IQ1_M | 6.6 MB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00002-of-00015.gguf | GGUF | IQ1_M | 16.74 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00003-of-00015.gguf | GGUF | IQ1_M | 14.57 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00004-of-00015.gguf | GGUF | IQ1_M | 14.93 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00005-of-00015.gguf | GGUF | IQ1_M | 13.82 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00006-of-00015.gguf | GGUF | IQ1_M | 14.93 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00007-of-00015.gguf | GGUF | IQ1_M | 14.39 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00008-of-00015.gguf | GGUF | IQ1_M | 14.64 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00009-of-00015.gguf | GGUF | IQ1_M | 14.18 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00010-of-00015.gguf | GGUF | IQ1_M | 14.54 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00011-of-00015.gguf | GGUF | IQ1_M | 14.93 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00012-of-00015.gguf | GGUF | IQ1_M | 14.21 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00013-of-00015.gguf | GGUF | IQ1_M | 14.30 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00014-of-00015.gguf | GGUF | IQ1_M | 14.13 GB | Download |
| UD-IQ1_M/Kimi-K3-UD-IQ1_M-00015-of-00015.gguf | GGUF | IQ1_M | 3.10 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00001-of-00014.gguf | GGUF | IQ1_S | 6.6 MB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00002-of-00014.gguf | GGUF | IQ1_S | 17.28 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00003-of-00014.gguf | GGUF | IQ1_S | 14.58 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00004-of-00014.gguf | GGUF | IQ1_S | 14.87 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00005-of-00014.gguf | GGUF | IQ1_S | 14.93 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00006-of-00014.gguf | GGUF | IQ1_S | 14.61 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00007-of-00014.gguf | GGUF | IQ1_S | 14.87 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00008-of-00014.gguf | GGUF | IQ1_S | 14.93 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00009-of-00014.gguf | GGUF | IQ1_S | 14.49 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00010-of-00014.gguf | GGUF | IQ1_S | 14.87 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00011-of-00014.gguf | GGUF | IQ1_S | 14.93 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00012-of-00014.gguf | GGUF | IQ1_S | 14.49 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00013-of-00014.gguf | GGUF | IQ1_S | 14.75 GB | Download |
| UD-IQ1_S/Kimi-K3-UD-IQ1_S-00014-of-00014.gguf | GGUF | IQ1_S | 1.05 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00001-of-00019.gguf | GGUF | Q2_K_XL | 6.6 MB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00002-of-00019.gguf | GGUF | Q2_K_XL | 16.58 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00003-of-00019.gguf | GGUF | Q2_K_XL | 13.69 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00004-of-00019.gguf | GGUF | Q2_K_XL | 13.82 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00005-of-00019.gguf | GGUF | Q2_K_XL | 14.07 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00006-of-00019.gguf | GGUF | Q2_K_XL | 13.23 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00007-of-00019.gguf | GGUF | Q2_K_XL | 13.75 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00008-of-00019.gguf | GGUF | Q2_K_XL | 13.35 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00009-of-00019.gguf | GGUF | Q2_K_XL | 13.69 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00010-of-00019.gguf | GGUF | Q2_K_XL | 13.86 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00011-of-00019.gguf | GGUF | Q2_K_XL | 13.85 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00012-of-00019.gguf | GGUF | Q2_K_XL | 13.69 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00013-of-00019.gguf | GGUF | Q2_K_XL | 13.99 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00014-of-00019.gguf | GGUF | Q2_K_XL | 13.72 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00015-of-00019.gguf | GGUF | Q2_K_XL | 13.91 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00016-of-00019.gguf | GGUF | Q2_K_XL | 13.86 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00017-of-00019.gguf | GGUF | Q2_K_XL | 13.85 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00018-of-00019.gguf | GGUF | Q2_K_XL | 13.69 GB | Download |
| UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00019-of-00019.gguf | GGUF | Q2_K_XL | 6.28 GB | Download |
Model Details
Model README
---
base_model: moonshotai/Kimi-K3
license: other
license_name: kimi-k3
license_link: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
language:
- ja
tags:
- gguf
- kimi-k3
- moe
- reap
---
Kimi-K3 REAP50 Width50 UD GGUF
!Kimi-K3 REAP50 and Width50 overview
This is an experimental, reduced GGUF build of Kimi-K3.
Calibration and selection
Prepared Japanese responses for chat, code generation, reasoning, and tool use were tokenized and evaluated with a forward pass of the original Kimi-K3. Only semantic assistant tokens were scored; prompts and control or structure tokens were excluded.
REAP50: expert axis
For expert e, the score is:
REAP(e) = mean[t routed to e](top-16-renormalized router weight × ||unweighted expert output||₂)
Experts are ranked separately in each of the 92 MoE layers. The top 448 of 896 are retained. No experts are reserved by hand.
Width50: intermediate axis
For every calibration token t routed to expert e, Width50 follows the actual expert MLP:
a(t,e) = SiTU(Wgate,e × h(t), Wup,e × h(t))
y(t,e,b) = router_weight(t,e) × Wdown,e[:,b] × a(t,e)[b]
score(e,b) = Σ[t routed to e] ||y(t,e,b)||₂²
Here, b is one 32-channel slice of the 3072-channel intermediate activation. Each slice is passed through only the matching columns of Wdown, producing a hidden-size output vector. Its router-weighted squared L2 norm is accumulated over semantic assistant tokens.
Eight adjacent 32-channel scores are summed into one physical QK256 score. The highest 6 of 12 QK256 blocks are retained independently for each expert. Low-coverage experts blend this activation score with a weight-based prior. The width map is measured on the original expert IDs and remapped through the REAP50 keep list before GGUF slicing.
Blocks are scored independently; cross-block cancellation and the post-mixture RMSNorm are not part of this ranking.
GGUF build flow
Kimi-K3 weights
→ Quantization
→ Q1 / Q2 GGUF
→ REAP50 expert-axis slice
→ Width50 QK256-block slice
→ final split GGUF
The REAP50 step slices the expert axis of ffn_gate_inp.weight, exp_probs_b.bias, and ffn_{gate,up,down}_exps.weight from 896 to 448 for every MoE layer.
The Width50 step slices the intermediate axis of ffn_gate_exps.weight, ffn_up_exps.weight, and ffn_down_exps.weight from 3072 to 1536. Complete quantization blocks are copied directly, so retained data is not dequantized or requantized.
| | Original | This build |
|---|---:|---:|
| Routed experts per MoE layer | 896 | 448 |
| Routed expert FFN width | 3072 | 1536 |
| Experts used per token | 16 | 16 |
All other tensor data is unchanged. The shared expert metadata is represented as 4 × 1536 instead of 2 × 3072 so that its physical width remains 6144 in llama.cpp.
Files
| Folder | Shards | Size |
|---|---:|---:|
| UD-IQ1_S | 14 | about 181 GiB |
| UD-IQ1_M | 15 | about 194 GiB |
| UD-Q2_K_XL | 19 | about 243 GiB |
Load the first shard in the selected folder.
Usage
Use the Kimi-K3 Width support branch of llama.cpp.
git clone --branch kimi-k3-width-support https://github.com/mmnga/llama.cpp
cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build llama.cpp/build -j --target llama-server
Example using UD-IQ1_S:
./llama.cpp/build/bin/llama-server \
-m ./UD-IQ1_S/Kimi-K3-UD-IQ1_S-00001-of-00014.gguf \
-ot ".*.ffn_.*_exps.*=CPU" \
-ngl 45 \
--ctx-size 8192 \
--flash-attn on \
--jinja \
--override-kv kimi-k3.expert_shared_count=int:2 \
--override-kv kimi-k3.expert_shared_feed_forward_length=int:6144 \
--override-kv kimi-k3.expert_used_count=int:16
The example was tested with a 32 GB RTX 5090 while keeping routed experts on the CPU. Adjust -ngl and context size for your available VRAM and RAM. Load the first shard when using another quantization folder.
This is a heavily reduced experimental model. Quality and stability may differ from the original Kimi-K3.
Run mmnga-o/Kimi-K3-REAP50-Width50-UD-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models