GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

GrEarl/Kimi-K3-GGUF overview

Author Note The author does not own enough hardware to run this 864.81 GiB model. Full model runtime results below were contributed by independent users and ha…

ggufkimi-k3q2_kmoetext-generationconversationalbase_model:moonshotai/Kimi-K3base_model:quantized:moonshotai/Kimi-K3license:otherendpoints_compatibleregion:us

Runs locally from ~637.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
3,813
Likes
29
Pipeline
text-generation
Author

Repository Files & Downloads

94 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Kimi-K3-Q2_K-00001-of-00094.ggufGGUFQ2_K637.2 MBDownload
Kimi-K3-Q2_K-00002-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00003-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00004-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00005-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00006-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00007-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00008-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00009-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00010-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00011-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00012-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00013-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00014-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00015-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00016-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00017-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00018-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00019-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00020-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00021-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00022-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00023-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00024-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00025-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00026-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00027-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00028-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00029-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00030-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00031-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00032-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00033-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00034-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00035-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00036-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00037-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00038-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00039-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00040-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00041-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00042-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00043-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00044-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00045-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00046-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00047-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00048-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00049-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00050-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00051-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00052-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00053-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00054-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00055-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00056-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00057-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00058-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00059-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00060-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00061-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00062-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00063-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00064-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00065-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00066-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00067-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00068-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00069-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00070-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00071-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00072-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00073-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00074-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00075-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00076-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00077-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00078-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00079-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00080-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00081-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00082-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00083-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00084-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00085-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00086-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00087-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00088-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00089-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00090-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00091-of-00094.ggufGGUFQ2_K9.40 GBDownload
Kimi-K3-Q2_K-00092-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00093-of-00094.ggufGGUFQ2_K9.30 GBDownload
Kimi-K3-Q2_K-00094-of-00094.ggufGGUFQ2_K1.78 GBDownload

Model Details

Model IDGrEarl/Kimi-K3-GGUF
AuthorGrEarl
Pipelinetext-generation
Licenseother
Base modelmoonshotai/Kimi-K3
Last modified2026-08-02T21:06:28.000Z

Model README

---

license: other

license_name: kimi-k3

base_model: moonshotai/Kimi-K3

base_model_relation: quantized

pipeline_tag: text-generation

tags:

- gguf

- kimi-k3

- q2_k

- moe

---

Author Note

**The author does not own enough hardware to run this 864.81 GiB model.

Full-model runtime results below were contributed by independent users and have

not been reproduced by the author. They are reported with their environment and

limitations rather than presented as author-run benchmarks.**

Kimi-K3 GGUF — Q2_K experts / Q4_K dense (2.673 bpw)

864.81 GiB in 94 parts. Text model only. The current v3 payload was

promoted on 2026-08-02 UTC. It uses the same

llama.cpp PR #26185

(pwilkin/kimi-k3-text) tensor contract as the earlier builds.

Source: moonshotai/Kimi-K3 at revision

9f62e4e9fffbd0a83ddd60e1c209d828994b3569 (pinned for every download).

| | |

|---|---|

| Parts | 94 (Kimi-K3-Q2_K-000NN-of-00094.gguf) |

| Size | 864.81 GiB (source 1,453.8 GiB → 0.595x) |

| Tensors | 2,573 (matches the contract exactly) |

| Effective rate | 2.673 bpw |

| Architecture | kimi-k3 |

| general.file_type | 10 (MOSTLY_Q2_K) |

| Chat template | embedded, validated byte-exact against K3's own builder |

| Quantization release | v3 |

| Sub-block scale search | 15-step ladder + 18 polished multi-start restarts, unweighted squared error |

| Previous release | v2; archived under v2/ at revision a62667d01793 |

| imatrix | no |

Quantization layout

| ggml type | tensors | size | bpw | what |

|---|---:|---:|---:|---|

| Q2_K | 276 | 832.04 GiB | 2.6250 | routed expert stacks (ffn_{gate,up,down}_exps) |

| Q4_K | 1,067 | 28.58 GiB | 4.5000 | dense 2-D weights, token_embd |

| F32 | 1,112 | 2.25 GiB | 32 | norms, router, ssm_a, ssm_conv1d, fused AttnRes scores |

| Q8_0 | 1 | 1.16 GiB | 8.5 | output.weight |

| F16 | 117 | 0.76 GiB | 16 | attn_k_b / attn_v_b (row length not a multiple of 256) |

The experts hold 2.72 T of the 2.78 T parameters, so they set the overall rate.

Dense weights are kept at Q4_K rather than Q2_K: they are 3.3% of the bytes, and

spending 4.5 bpw there is cheap insurance for the layers that every token passes

through. output.weight is promoted further, to Q8_0, because its error lands

directly on the logits with nothing downstream to average it out. The Q8_0 tensor

is 1.16 GiB and adds 0.54 GiB versus keeping it at Q4_K, or 0.06% of the

artifact; upstream llama.cpp promotes the output head for low-bit file types as

well.

Chat template

K3's tokenizer_config.json ships no chat template — conversation assembly

lives in encoding_k3.py as code, so nothing could simply be copied across. The

template embedded here is [Xenova's Jinja

port](https://huggingface.co/Xenova/Kimi-K3-tokenizer), patched for llama.cpp's

Minja engine and validated against K3's own build_chat_segments():

28 of 28 cases byte-exact. The cases cover normal and multi-turn messages,

thinking_effort, tool declarations, assistant tool calls, tool results, and a

nested response_schema.

The exact template extracted back from this GGUF was also rendered by llama.cpp's

Minja implementation (commit 91f8c9c5): a tools/tool-result case was 1,218 bytes

and a nested-schema case was 621 bytes, both byte-identical to K3's renderer.

Images and batched conversations remain untested.

Provenance and independent conformance. The file in this artifact began with

Xenova's port and was expanded

and patched locally; it is not claimed to be byte-identical to the separate

upstream implementation in Moonshot PR #66.

PR #66 and ChatLint's K3 findings

provide independent, oracle-backed conformance work. The Minja

namespace(items=[]) collision reproduced here was accepted and fixed upstream

in ChatLint commit ecb1a727

and on the PR branch. Against this artifact's 15,053-byte file, ChatLint pinned to

that commit reports 294/294 checks, 0 errors, 0 warnings. ChatLint renders with

transformers-style Jinja2 rather than Minja and checks structural properties, so

it supplements rather than replaces the 28/28 oracle comparison and direct Minja

runs above.

One limitation is shared with PR #66: sandboxed Jinja cannot parse a JSON string

inside tool_calls[].function.arguments. This template emits a valid

<|open|>json type="object" fallback for a non-empty string; callers that require

the reference renderer's per-argument XTML must parse the string into an object

before applying the template.

Measured quantization error

MXFP4 source dequantized to F32, requantized to Q2_K, then dequantized again and

compared against the source. 92 expert tensors sampled, one per routed layer:

| metric | value |

|---|---|

| relative RMSE, mean | 0.263402 |

| cosine similarity, mean | 0.964694 |

The superseded v2 build measured 0.276874 relative RMSE / 0.961095 cosine,

and the original v1 min/max build measured 0.3309 / 0.951450, on the same

92-tensor comparison. Those are historical baselines, not measurements of the

current files.

For reference, the same measurement on the

IQ1_S sibling current

top32/refit2 payload is 0.455936 / 0.890050 at 528.0293 GiB. This repo is

1.638x larger by file size and noticeably closer to the MXFP4 source in this

weight-space measurement. Historical IQ1_S figures of 0.4626 / 0.886692 for the

top8/refit1 payload, and 0.5375 / 0.848230 for the earlier ternary-space snap,

are superseded and do not describe its current files.

The v3 scale search runs in a fused CuPy kernel and keeps each 16-value sub-block

in registers while evaluating the 15-step candidate ladder, polishing its winner,

and testing 18 independent restarts. Assignment and least-squares re-solving are

alternated three times per candidate; the lowest unweighted squared error wins.

Everything downstream remains standard block_q2_K: 84 bytes per 256 elements,

or 2.625 bpw.

The Q2_K bytes are fed to gguf-py's own dequantizer and compared against the

local dequantizer: max absolute difference 0.0, 0 mismatching elements, on

synthetic data and a real K3 expert tensor. Before the full run, the fused search

was also checked against a scalar implementation on 16 real K3 expert tensors

spanning layers 1–91 and experts 0–895; both produced the same measured relative

RMSE, with no sampled tensor regressing against v2.

v1 → v2 → v3

One expert tensor was measured in every one of the 92 rebuilt expert-bearing

parts on identical source inputs:

| release | sub-block search | relative RMSE, mean | cosine, mean |

|---|---|---:|---:|

| v1 | direct min/max | 0.3309 | 0.951450 |

| v2 | 6-candidate unweighted SSE (nstep=5) | 0.276874 | 0.961095 |

| v3 (current) | 15-step ladder + polished 18-restart multi-start | 0.263402 | 0.964694 |

Directly comparing v2 and v3:

| metric | v2 | v3 |

|---|---:|---:|

| relative RMSE, range over 92 parts | 0.275965–0.277595 | 0.263099–0.263567 |

| cosine, range over 92 parts | 0.960898–0.961344 | 0.964649–0.964776 |

  • Relative RMSE falls by 4.87%, equivalent to removing **9.50% of squared

reconstruction error**.

  • Cosine similarity to the source weights rises by 0.360 percentage points.
  • All 92/92 measured parts improved; 0 regressed. Per-part RMSE reduction is

4.66%–5.06%.

  • From v1 to v3, relative RMSE is down 20.4% at the same 2.625 bpw expert

payload size.

  • **File size, tensor count, tensor names, ggml types, and runtime-relevant

metadata are unchanged.** Only the routed-expert Q2_K payload differs; parts 1

and 94 contain no routed experts and are byte-identical to v2.

The expert-free part 1 was copied server-side during promotion, so its custom

provenance KV kimi-k3.conversion.q2k_subblock_search still says

unweighted-sse nstep=5. That value accurately identifies v2 but is stale for

the current expert payload; the authoritative v3 setting is the 15-step ladder

plus 18 polished restarts documented here. This does not affect loading or

inference.

These are weight-space reconstruction measurements. They show that v3 sits

closer to the MXFP4 source, but they do not quantify generation quality.

**No perplexity, standardized benchmark, or source-model output-equivalence

comparison exists for v1, v2, or v3.**

The v3 run used eight NVIDIA L4 workers and rebuilt the 92 expert-bearing parts

in 36.5 minutes for a recorded Modal cost of $2.635. The fused search itself was

verified before publication; every emitted part also passed the no-regression

gate and structural checks for its split number, tensor count, tensor types, and

maximum tensor-name length.

Previous builds

The v2 files were archived under v2/ during promotion. They remain reachable

in the Hugging Face commit history by pinning revision a62667d01793 and

using the v2/ path, even after that folder is removed from the current file

listing. The original v1 min/max build remains reachable at **revision

d76240360965**, where it was stored at the repository root.

Previously downloaded v1 or v2 files remain loadable. The measured difference is

weight-space error only, with no downstream benchmark comparison, so

re-downloading 864 GiB is a quality-versus-bandwidth decision rather than a

compatibility fix.

Runtime status: v2 full load reported; v3 structure verified

kimi-k3 is not merged into released llama.cpp; support exists on

PR #26185

(pwilkin/llama.cpp, branch kimi-k3-text).

A detailed independent report in

Discussion #4

(2026-07-28 UTC) exercised the 94-part Q2_K root available at that time: v2, not

the current v3 payload. The repository root was at

d8aebda6c891e203284fc4ddef5483f2177ee841 when the report was posted; the

reporter did not provide per-file hashes, so this attribution is based on the

repository timeline rather than a cryptographic pin. v3 preserves every tensor

name, type, shape, split field and runtime-relevant metadata, but it has not been

separately loaded as a full 94-part model.

Reported environment and results:

| Item | Third-party v2 report |

|---|---|

| Hardware | AWS p6-b200.48xlarge: 8× NVIDIA B200 / 1.43 TiB aggregate HBM, sm_100 |

| Runtime | pwilkin/llama.cpp kimi-k3-text at 06eec9f5 |

| Load | All 94 parts loaded cleanly, fully VRAM-resident; 869 GiB resident, 97–116 GB per GPU |

| Startup | About 110 seconds from launch to serving from local NVMe |

| Decode | 17.5–17.8 tokens/s, reported stable across tasks |

| Prompt processing | Up to about 330 tokens/s with -ub 4096 |

| Generation checks | At temperature 0, the reporter observed coherent factual answers, instruction following, and correct handling of a simple trap question |

| Tools | The model emitted well-formed tool calls with correctly typed arguments |

| Context | A 262K context allocation ran with little VRAM change and unchanged decode speed; needle retrieval was tested at 26K tokens, not 262K |

| Input rendering | The reporter confirmed that the 28/28-tested prompt template rendered correctly in practice |

These are third-party observations of v2, not author-run measurements and not a

v3 benchmark. They establish full load, serving, generation, raw tool-call

emission, and a 26K retrieval check for v2 on the named setup. They do not

establish v3 generation quality, perplexity, benchmark scores, source-model

equivalence, or 262K retrieval quality. The statement that no quantization damage

was visible is a qualitative user observation, not a formal quality result.

Known runtime limitations from the same report

  • CPU-only currently crashes while loading K3 at

GGML_ASSERT(*cur_backend_id != -1) in ggml_backend_sched_split_graph.

At present use a GPU or GPU+CPU hybrid configuration; the traditional

all-CPU large-RAM/mmap route is not validated.

  • The tested branch lacked a K3 output parser. Without the reporter's

parser PR,

/v1/chat/completions can return empty content, place output in

reasoning_content, and leak raw XTML markers. This is an output-parsing issue,

not evidence that the embedded input template is wrong.

  • special_eos_id is not in special_eog_ids was reported as a benign warning;

raw /completion stopped with stop_type: eos. The metadata has not been

changed solely to silence that warning.

  • Released/stock llama.cpp, CPU-only execution, Metal, Vulkan, and other hardware

configurations remain unverified for this artifact.

What has been verified, and by whom:

| Level | Status |

|---|---|

| GGUF structure, split metadata, tensor names/types, hparams, vocab | Verified by the author on the published files via HTTP Range reads |

| Numeric agreement of custom quant bytes with gguf-py's dequantizer | Verified by the author, exact |

| Chat input template vs K3 build_chat_segments() | Verified by the author, 28/28 byte-exact; tools/schema rendered directly with Minja |

| Full 94-part load, serving, generation, tool emission on 8×B200 | Reported for v2 by an independent user in Discussion #4; v3 not separately tested |

| 262K context allocation / 26K needle retrieval | Reported for v2 by the same user; retrieval was not tested at 262K |

| Perplexity, standardized benchmark, source-model output equivalence | Not measured |

Template compatibility fix

The first template published here used Jinja's namespace({...}) dict-literal

form. Minja accepts only namespace(key=value, ...). Direct Minja execution then

found two more tool/schema-path issues: tojson(sort_keys=true) is not implemented,

and a namespace field named items collides with the object method. The current

15,053-byte template uses recursive dictsort JSON rendering and a non-conflicting

field name. It is embedded in part 1 and also published as chat_template.jinja.

Thanks to the user who caught the original incompatibility.

What was fixed relative to the previous 96-part upload

The earlier Q2_K release in this repo (96 parts, 938.59 GiB) could not have been

loaded even with K3 support present. All four defects required regenerating every

part, so this upload replaces it entirely.

  1. Tensor names were raw Hugging Face names. 668 names across 92 of the 96

parts exceeded GGML_MAX_NAME (64). Now every name follows the PR contract;

longest is 29 characters.

  1. Part 1 had no model metadata. No hparams, no vocabulary. Now part 1 carries

66 KV entries: 25 required hparams plus the full vocabulary (163,840 tokens /

163,328 merges, gpt2 / pre kimi-k2, chkhsh verified).

  1. split.tensors.count was UINT16 and per-part. llama-model-loader.cpp

reads it as a required key and compares it against weights_map.size(),

throwing corrupted model on mismatch. Now INT32 / 2573 in all 94 parts,

equal to the actual sum of per-part tensor counts.

  1. Vision tensors were mixed into the text model. HF shards 95 and 96 hold 168

vision_tower / mm_projector tensors. done_getting_tensors() is called with

partial=false, so a single extra tensor triggers wrong number of tensors.

Those two shards are excluded, which is why there are 94 parts and not 96.

Conversion details

Recorded in the file itself under kimi-k3.conversion.*:

  • a_log_transform = -exp(A_log[:n_head]) — K3 stores A_log as 128 elements

but only the first num_attention_heads = 96 are used

  • kv_b_split = k_b(transposed)+v_bkv_b_proj split into attn_k_b / attn_v_b
  • attn_res_fused = res_norm*res_proj[0] in float32 — the AttnRes norm/proj pair

is fused into one score vector, matching the reference _apply_attn_res()

  • expert_stack_dim = 0 — 896 experts stacked on dim 0
  • source_quant = compressed-tensors/mxfp4-pack-quantized
  • defaults_used = rope_theta — the only value not present in config.json,

taken from the configuration_kimi_k3.py class default (10000.0)

Usage on the PR branch

Pass the first part; llama.cpp follows the split metadata to the rest.

llama-cli -m Kimi-K3-Q2_K-00001-of-00094.gguf -p "..." 

All 94 parts must be present in the same directory.

Run GrEarl/Kimi-K3-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models