GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF overview

Kimi Linear 48B A3B Instruct — APEX GGUF MoE aware, mixed precision APEX quantizations of moonshotai/Kimi Linear 48B A3B Instruct https://huggingface.co/moonsh…

ggufmoeapexquantizedkimi-linearlinear-attentionllama.cpptext-generationbase_model:moonshotai/Kimi-Linear-48B-A3B-Instructbase_model:quantized:moonshotai/Kimi-Linear-48B-A3B-Instructlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~28.23 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
690
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Kimi-Linear-48B-A3B-Instruct-APEX-balanced.ggufGGUFGGUF32.72 GBDownload
Kimi-Linear-48B-A3B-Instruct-APEX-handroll.ggufGGUFGGUF32.72 GBDownload
Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.ggufGGUFGGUF28.23 GBDownload

Model Details

Model IDMyric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF
AuthorMyric
Pipelinetext-generation
Licensemit
Base modelmoonshotai/Kimi-Linear-48B-A3B-Instruct
Last modified2026-07-25T16:58:14.000Z

Model README

---

license: mit

base_model: moonshotai/Kimi-Linear-48B-A3B-Instruct

base_model_relation: quantized

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- moe

- apex

- quantized

- kimi-linear

- linear-attention

- llama.cpp

---

Kimi-Linear-48B-A3B-Instruct — APEX GGUF

MoE-aware, mixed-precision APEX quantizations of

moonshotai/Kimi-Linear-48B-A3B-Instruct

— 48B total / ~3B active, a hybrid linear-attention MoE: most layers use KDA

(Kimi Delta Attention, gated-delta linear attention), a few use full MLA

attention, over a 256-routed + 1-shared expert FFN.

To my knowledge this is the first APEX quant of a linear-attention hybrid MoE.

APEX assigns precision per tensor role and per layer instead of uniformly; here

that meant teaching the recipe about tensor families the stock generator doesn't

know (see Method).

Results

Perplexity on wikitext-2-raw (test, 200×512-token windows), llama-perplexity.

| File | Size | BPW | PPL | Δ vs bf16 |

|------|------|-----|-----|-----------|

| bf16 (reference) | 92 GB | 16.0 | 7.374 | — |

| APEX-balanced (no-imatrix) | 33 GB | 5.72 | 7.377 | +0.04% |

| APEX-handroll (ssm@Q8_0, no-imatrix) | 33 GB | 5.72 | 7.382 | +0.11% |

| APEX-i-quality (imatrix, IQ4_XS mid experts) | 29 GB | 4.94 | 7.399 | +0.34% |

Both no-imatrix tiers land essentially on the bf16 reference (within ~0.1%). The

imatrix-guided i-quality tier (IQ4_XS mid experts, 29 GB) is the smallest here

and is available as Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.gguf: at 4.94 BPW it

trades ~0.34% perplexity for another ~4 GB off balanced.

From a 92 GB bf16 baseline → 33 GB (~2.8× smaller), and it **runs on a 128 GB

unified-memory box** (fits with full GPU offload). Coherent on general and factual

prompts.

The balanced and handroll tiers were built without an imatrix (Q6_K/Q5_K

experts, Q8_0 shared, Q6_K attention — none of which require importance data). The

newer i-quality tier is imatrix-guided (IQ4_XS mid experts). The imatrix itself

(Kimi-Linear-48B-A3B-Instruct.imatrix) is included in this repo, so you can roll

your own deeper lower-bit "I-tier" variants (IQ3/IQ2) — those GGUFs aren't pre-built here.

Note on the two tiers (a null result)

The hand-roll tier pins the KDA recurrence tensors (ssm_conv1d_*,

ssm_f/g_*, ssm_beta) to Q8_0 instead of Q6_K, testing whether protecting

the linear-attention state preserves quality. It doesn't — PPL is identical

within noise (7.382 vs 7.377), at the same size (the ssm tensors are tiny next to

the experts). Use balanced. The hand-roll is kept only to document the

experiment.

Which file

  • APEX-balanced — recommended. Q6_K/Q5_K experts on a layer-depth gradient,

Q8_0 shared experts, Q6_K attention + KDA tensors.

  • APEX-i-quality — imatrix-guided IQ4_XS mid experts (29 GB); the smallest tier

here (PPL 7.399, +0.34% vs bf16). Try it when you want a few GB over balanced.

  • APEX-handroll — experimental (see null-result note below); not recommended.

Usage (llama.cpp)

llama-cli   -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 -p "Hello"
llama-server -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --host 0.0.0.0 --port 8080

Requires a llama.cpp build supporting the kimi_linear architecture and the

kimi-k2 pre-tokenizer.

Method

APEX is a bit-allocation recipe over stock llama-quantize --tensor-type-file.

Kimi-Linear needed two tensor families the stock APEX generator doesn't emit:

  • MLA (full-attention layers): attn_kv_a_mqa, attn_k_b, attn_v_b
  • KDA (linear-attention layers): ssm_conv1d_{k,q,v}, ssm_f_a/f_b,

ssm_g_a/g_b, ssm_beta (norms/1-D state kept F32)

plus a dense layer 0 (--dense-layers 1). The expert intermediate dim is 2048

(256-divisible), so no IQ4_NL workaround was needed. Config generation +

patching: see REPRODUCE.md, patch_kimi_config.py, and

configs/.

Baseline: quantized from

bartowski's bf16 GGUF.

Attribution & licenses

All MIT; see LICENSE and NOTICE.

Unofficial community quantization; not affiliated with or endorsed by Moonshot AI.

Run Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models