Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF overview
Kimi Linear 48B A3B Instruct — APEX GGUF MoE aware, mixed precision APEX quantizations of moonshotai/Kimi Linear 48B A3B Instruct https://huggingface.co/moonsh…
Runs locally from ~28.23 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF |
|---|---|
| Author | Myric |
| Pipeline | text-generation |
| License | mit |
| Base model | moonshotai/Kimi-Linear-48B-A3B-Instruct |
| Last modified | 2026-07-25T16:58:14.000Z |
Model README
---
license: mit
base_model: moonshotai/Kimi-Linear-48B-A3B-Instruct
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- moe
- apex
- quantized
- kimi-linear
- linear-attention
- llama.cpp
---
Kimi-Linear-48B-A3B-Instruct — APEX GGUF
MoE-aware, mixed-precision APEX quantizations of
moonshotai/Kimi-Linear-48B-A3B-Instruct
— 48B total / ~3B active, a hybrid linear-attention MoE: most layers use KDA
(Kimi Delta Attention, gated-delta linear attention), a few use full MLA
attention, over a 256-routed + 1-shared expert FFN.
To my knowledge this is the first APEX quant of a linear-attention hybrid MoE.
APEX assigns precision per tensor role and per layer instead of uniformly; here
that meant teaching the recipe about tensor families the stock generator doesn't
know (see Method).
Results
Perplexity on wikitext-2-raw (test, 200×512-token windows), llama-perplexity.
| File | Size | BPW | PPL | Δ vs bf16 |
|------|------|-----|-----|-----------|
| bf16 (reference) | 92 GB | 16.0 | 7.374 | — |
| APEX-balanced (no-imatrix) | 33 GB | 5.72 | 7.377 | +0.04% |
| APEX-handroll (ssm@Q8_0, no-imatrix) | 33 GB | 5.72 | 7.382 | +0.11% |
| APEX-i-quality (imatrix, IQ4_XS mid experts) | 29 GB | 4.94 | 7.399 | +0.34% |
Both no-imatrix tiers land essentially on the bf16 reference (within ~0.1%). The
imatrix-guided i-quality tier (IQ4_XS mid experts, 29 GB) is the smallest here
and is available as Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.gguf: at 4.94 BPW it
trades ~0.34% perplexity for another ~4 GB off balanced.
From a 92 GB bf16 baseline → 33 GB (~2.8× smaller), and it **runs on a 128 GB
unified-memory box** (fits with full GPU offload). Coherent on general and factual
prompts.
The balanced and handroll tiers were built without an imatrix (Q6_K/Q5_K
experts, Q8_0 shared, Q6_K attention — none of which require importance data). The
newer i-quality tier is imatrix-guided (IQ4_XS mid experts). The imatrix itself
(Kimi-Linear-48B-A3B-Instruct.imatrix) is included in this repo, so you can roll
your own deeper lower-bit "I-tier" variants (IQ3/IQ2) — those GGUFs aren't pre-built here.
Note on the two tiers (a null result)
The hand-roll tier pins the KDA recurrence tensors (ssm_conv1d_*,
ssm_f/g_*, ssm_beta) to Q8_0 instead of Q6_K, testing whether protecting
the linear-attention state preserves quality. It doesn't — PPL is identical
within noise (7.382 vs 7.377), at the same size (the ssm tensors are tiny next to
the experts). Use balanced. The hand-roll is kept only to document the
experiment.
Which file
- APEX-balanced — recommended. Q6_K/Q5_K experts on a layer-depth gradient,
Q8_0 shared experts, Q6_K attention + KDA tensors.
- APEX-i-quality — imatrix-guided IQ4_XS mid experts (29 GB); the smallest tier
here (PPL 7.399, +0.34% vs bf16). Try it when you want a few GB over balanced.
- APEX-handroll — experimental (see null-result note below); not recommended.
Usage (llama.cpp)
llama-cli -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 -p "Hello"
llama-server -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --host 0.0.0.0 --port 8080
Requires a llama.cpp build supporting the kimi_linear architecture and the
kimi-k2 pre-tokenizer.
Method
APEX is a bit-allocation recipe over stock llama-quantize --tensor-type-file.
Kimi-Linear needed two tensor families the stock APEX generator doesn't emit:
- MLA (full-attention layers):
attn_kv_a_mqa,attn_k_b,attn_v_b - KDA (linear-attention layers):
ssm_conv1d_{k,q,v},ssm_f_a/f_b,
ssm_g_a/g_b, ssm_beta (norms/1-D state kept F32)
plus a dense layer 0 (--dense-layers 1). The expert intermediate dim is 2048
(256-divisible), so no IQ4_NL workaround was needed. Config generation +
patching: see REPRODUCE.md, patch_kimi_config.py, and
configs/.
Baseline: quantized from
Attribution & licenses
All MIT; see LICENSE and NOTICE.
- Base: Moonshot AI (@moonshotai) — Kimi-Linear-48B-A3B-Instruct (MIT)
- bf16 GGUF: bartowski (@bartowski) — source
- Engine: llama.cpp (@ggml-org · github) (MIT)
- APEX: Ettore Di Giacinto / LocalAI (@mudler) — localai-org/apex-quant (MIT)
Unofficial community quantization; not affiliated with or endorsed by Moonshot AI.
Run Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models