GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Baekpica/K-EXAONE-236B-A23B-Mixed-Quant-GGUF overview

K EXAONE 236B A23B — Mixed Quant GGUF A 237 billion parameter model, structurally intact, resident on one 128 GB DGX Spark — with its full 262 144 token contex…

ggufk-exaoneexaonemoemixed-quantizationmtpdgx-sparktext-generationenkoesdejavibase_model:LGAI-EXAONE/K-EXAONE-236B-A23Bbase_model:quantized:LGAI-EXAONE/K-EXAONE-236B-A23Blicense:otherendpoints_compatibleregion:usimatrixconversational

Runs locally from ~11.66 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
1,174
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-00001-of-00003.ggufGGUFIQ2XXS36.72 GBDownload
K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-00002-of-00003.ggufGGUFIQ2XXS37.18 GBDownload
K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-00003-of-00003.ggufGGUFIQ2XXS11.66 GBDownload
K-EXAONE-236B-A23B-MXQ-Q2K-Q4Edge-Q8Dense-MTPQ8-pilot-v1-00001-of-00003.ggufGGUFQ2K36.99 GBDownload
K-EXAONE-236B-A23B-MXQ-Q2K-Q4Edge-Q8Dense-MTPQ8-pilot-v1-00002-of-00003.ggufGGUFQ2K36.96 GBDownload
K-EXAONE-236B-A23B-MXQ-Q2K-Q4Edge-Q8Dense-MTPQ8-pilot-v1-00003-of-00003.ggufGGUFQ2K13.90 GBDownload

Model Details

Model IDBaekpica/K-EXAONE-236B-A23B-Mixed-Quant-GGUF
AuthorBaekpica
Pipelinetext-generation
Licenseother
Base modelLGAI-EXAONE/K-EXAONE-236B-A23B
Last modified2026-08-14T08:00:42.000Z

Model README

---

license: other

license_name: k-exaone

license_link: LICENSE

base_model: LGAI-EXAONE/K-EXAONE-236B-A23B

base_model_relation: quantized

language: [en, ko, es, de, ja, vi]

pipeline_tag: text-generation

library_name: gguf

tags: [gguf, k-exaone, exaone, moe, mixed-quantization, mtp, dgx-spark]

---

K-EXAONE-236B-A23B — Mixed-Quant GGUF

**A 237-billion-parameter model, structurally intact, resident on one 128 GB

DGX Spark — with its full 262 144-token context.**

Not a distillation. Not a pruned or expert-dropped variant. Not a

layer-truncated one. Every one of the 128 routed experts is present in every

one of the 47 MoE layers, alongside the shared expert, the dense layer 0, and

the original 1-layer MTP block. The tensor count matches the BF16 source:

781 tensors, 237.10 B parameters. The only thing that changed is the number

of bits each tensor is stored in — assigned by what the tensor does, not by a

global bit budget.

The 250 B-class weight class normally implies a multi-GPU host. This artifact

fits 85.56 GiB of weights and 12.30 GiB of 256K KV cache into a single

GB10's unified memory, measured at 103.95 GiB of 121.6 GiB resident and

serving over an OpenAI-compatible API. That is the result this repository

exists to demonstrate.

| | |

|---|---:|

| Parameters | 237.10 B (A23B active) |

| Routed experts kept | 128 / 128, all 47 MoE layers |

| Tensors | 781 — identical to the BF16 source |

| BF16 size | 441.63 GiB |

| This artifact (v1) | 85.56 GiB — 5.16× smaller |

| Context served on one GB10 | 262 144 tokens |

| Resident at 256K, measured | 103.95 GiB / 121.6 GiB |

Mixed-precision GGUF builds of LGAI-EXAONE/K-EXAONE-236B-A23B, quantized per

module role rather than uniformly, keeping the parts that matter most at 8 bit.

ds4-dfm

This model is served by ds4-dfm, a common ds4 release line for

DFM (독자 파운데이션 모델, 독파모) model families. It was developed to

serve full-scale models on a single NVIDIA DGX Spark with 128 GB of unified

memory through explicit model-family C/CUDA paths and device kernels.

Use the pinned

Baekpica/ds4 v0.5.6.3-dfm

release. One ds4-server command and one HTTP contract cover the integrated

DeepSeek, Solar Open2, K-EXAONE, and Motif-3 families; only the GGUF path and

its matching weight-owner manifest change. The server provides OpenAI Chat

Completions, OpenAI Completions, OpenAI Responses, and Anthropic Messages.

The production GGUF in this repository was loaded through the common release

binary on DGX Spark. That is an integration and API-surface result; context

length, throughput, and quality claims remain limited to the model-specific

evidence below.

Variants

| Variant | Size | Routed gate/up | Routed down | Built with imatrix |

|---|---:|---|---|---|

| v1 …-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1 | 85.56 GiB | IQ2_XXS | Q3_K | yes |

| pilot …-MXQ-Q2K-Q4Edge-Q8Dense-MTPQ8-pilot-v1 | 87.84 GiB | Q2_K | Q2_K | no |

Each is published as three shards (-00001-of-00003 …) because the Hub caps

individual files at 50 GB. Point llama.cpp at the first shard; it loads the

rest automatically. No merge step is needed:

llama-server -m K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-00001-of-00003.gguf \
  -ngl 99 -c 8192

The pilot exists because llama.cpp treats IQ2_XXS without an importance

matrix as a hard error, so it substitutes Q2_K and needs no calibration data.

v1 is the better artifact on both axes — 2.3 GiB smaller and closer to the

Q8_0 reference (see below) — so prefer it unless you specifically want an

artifact built without calibration data.

| sha256 | |

|---|---|

| v1 (unsplit) | 0e93f4bc41db6eb53c3520352ff7ec0be40749948a6608deb4cc2ad0818c94a1 |

| pilot (unsplit) | 2d840ee44b0e10cb2e14ec7cf58d2e7849615de1a92f58b1220790f42310ce39 |

Recipe

Quantization is assigned by what each tensor does, not by a global bit budget.

| Tensor group | Type | Why |

|---|---|---|

| Token embedding, LM head | Q8_0 | multilingual token fidelity; logit sensitivity |

| All norms (incl. QK-norm) | F32 | tiny, and error accumulates through them |

| Router (ffn_gate_inp, exp_probs_b) | F32 | a wrong expert choice costs more than any bit saved |

| Attention Q/K/V/O | Q8_0 | long-context stability |

| Dense layer 0 MLP | Q8_0 | every token passes through it |

| Shared expert | Q8_0 | every token passes through it |

| Routed expert gate/up | IQ2_XXS (pilot: Q2_K) | ~64 % of all parameters; where the compression has to come from |

| Routed expert down | Q3_K (pilot: Q2_K) | weighted accumulation, kept more conservative than gate/up |

| Edge MoE layers 1–4, 44–47 | Q4_K | first and last sparse blocks protected |

| MTP block (blk.48) | Q8_0 | draft quality drives speculative acceptance |

Full recipe: quant-recipe-v1.yaml. Exact per-tensor assignments as fed to

llama-quantize: *.tensor-types.txt. Per-tensor verification against the

recipe: verify-*.json.

Calibration (v1 only)

The importance matrix was built from a corpus covering all six languages the

model serves — Korean weighted heaviest, since routed gate/up goes to the most

aggressive quant in the recipe and Korean capacity is what this artifact exists

to protect. Sources: nvidia/Nemotron-SFT-Multilingual-v2 (ko, ja),

-v1 (es, de), Nemotron-Cascade-SFT-Stage-1,

Nemotron-SFT-Instruction-Following-Chat-v2,

Nemotron-SFT-Competitive-Programming-v2, and Wikipedia for Vietnamese —

neither Nemotron release covers it. Records are rendered with K-EXAONE's own

chat template so the activations the matrix sees match serving time.

Composition: calibration.composition.json.

Provenance

| | |

|---|---|

| Source model | LGAI-EXAONE/K-EXAONE-236B-A23B @ 61e6d578eb102b578e5704e2916ac841df9eca0a |

| Source GGUF | LGAI-EXAONE/K-EXAONE-236B-A23B-GGUF @ 5bd0394e4f42c00df63e207b9c434387523a6b77 |

| BF16 GGUF sha256 | 73be2da8653976df036bf9b6466b011f86cb10f78bab30a47025638ec999d3f8 |

| llama.cpp (quantizer) | ggml-org/llama.cpp @ 6a32c29a746a2e44de463de647f9f6661eb5086b (build b10295) |

| Converter | Baekpica/k-exaone-mixed-ds4 |

| Serving engine (measured below) | Baekpica/ds4 @ b2faf06f1ce2702efa53ac17145b1f56d3fb23b8 |

| — upstream engine | antirez/ds4 |

| — DGX Spark port | Entrpi/ds4-on-spark |

Artifact sha256 and build parameters: *.manifest.json. Tensor-level

verification against the recipe (verify-v1.json): **781 tensors, 85.558 GiB,

0 errors, 0 warnings**, matching the BF16 source's own tensor count.

Model structure

48 transformer layers plus one MTP block stored as blk.48, so

exaone-moe.block_count is 49. Hidden 6144, vocab 153 600, context 262 144.

Attention is GQA — 64 query heads over 8 KV heads at head_dim 128 — with

per-head RMSNorm on Q and K, on an LLLG sliding-window schedule (window 128;

every fourth layer is full attention). Layer 0 is dense (18 432); layers 1–47

are MoE with 128 routed experts, top-8, sigmoid gating with normalized top-k

probabilities, routed_scaling_factor 2.5, plus one 2048-wide shared expert.

The MTP block is a dense layer with its own attention and eh_proj; it shares

the base model's embedding and LM head.

How to run it

Two runtimes serve these files, and they are not interchangeable. **ds4 is the

one this artifact was sized for** and the only one measured here at full

context; llama.cpp runs the same file unmodified, but leaves the MTP block on

the floor.

| | ds4 (Baekpica/ds4, feature/exaone-model-loader) | llama.cpp |

|---|---|---|

| Runs the artifact | yes, unmodified | yes, unmodified |

| MTP block blk.48 | executed, target-verified speculative decoding | ignored — stored, never executed |

| 262 144-token context on one 128 GB device | measured — 103.95 GiB resident | not measured here |

| Multi-turn prefix reuse | yes — a continuation resumes at the divergence point | not measured here |

| Server API | OpenAI / Responses / Anthropic-compatible | llama.cpp HTTP API |

| Validated on GB10 / sm_121 | yes — see below | no |

ds4 — the engine this artifact was sized for

> This section describes software outside this repository. Everything below

> requires

> Baekpica/ds4,

> branch feature/exaone-model-loader. Pin the commit below; the branch moves.

The lineage matters, because almost none of the engine is ours:

| Layer | Repository | What it provides |

|---|---|---|

| Engine | antirez/ds4 | the whole runtime — GGUF loader, sessions, KV, CUDA backend, MoE routing, the OpenAI/Responses/Anthropic server, and the NextN/MTP scheduling contract |

| GB10 port | Entrpi/ds4-on-spark | the sm_121 build target and the aligned-artifact tier that makes mixed-quant MoE weights fast on unified memory |

| This work | Baekpica/ds4 | the exaone-moe model family: GQA + QK-norm attention, the LLLG sliding-window schedule, sigmoid/top-8 routing, and the blk.48 MTP graph |

ds4 was an MLA-only engine; K-EXAONE is plain GQA, so that attention path had to

be written. Neither antirez/ds4 nor Entrpi/ds4-on-spark serves this model as

shipped — use the branch above.

The full serving walkthrough, with the measured numbers, is the next section.

llama.cpp

The artifact is a plain GGUF, so it also runs unmodified on stock llama.cpp —

useful for a quick check, or on hardware where ds4 has no backend.

llama-server -m K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-00001-of-00003.gguf \
  -ngl 99 -c 8192

Point it at the first shard; it loads the other two automatically. A

mixed-quant GGUF needs no special runtime: GGUF stores a type per tensor and

ggml dispatches per tensor, which is how Q4_K_M — itself a mixture of Q4_K,

Q6_K and Q8_0 — already works. This recipe just assigns that mixture more

aggressively, and llama-quantize is what produced the file.

Measured, not assumed: the pilot artifact loaded in llama-server on 4 × RTX

PRO 6000 in 10.2 s and generated 384 tokens of Korean at 78.1 tok/s with

a broken-jamo ratio of 0.000.

Two caveats. llama.cpp ignores the MTP block — those tensors are preserved

in the artifact, not executed. And nothing on this page about 256K context,

resident memory or prefix reuse was measured on it; those are ds4 numbers.

Serving on DGX Spark (GB10 / sm_121) with ds4

Measured on a DGX Spark: NVIDIA GB10, sm_121, 121.6 GiB unified memory,

driver 595.71.05, CUDA 13.3, Linux 6.17.

| | |

|---|---|

| Engine | Baekpica/ds4 |

| Branch | feature/exaone-model-loader |

| Commit | b2faf06f1ce2702efa53ac17145b1f56d3fb23b8 |

| Weights | Baekpica/K-EXAONE-236B-A23B-Mixed-Quant-GGUF, variant v1 |

| Converter / reports | Baekpica/k-exaone-mixed-ds4 |

> **Pin b2faf06 for the measurements on this page. Do not use anything

> earlier than 920427a for long prompts.** Before 920427a, the

> exaone-moe sliding layers allocated a KV ring

> exactly the width of the attention window while prefill ran 2 048-token

> chunks. A chunk writes every row's KV before any row attends, so the ring was

> left holding only the chunk's last 128 positions and all but the final row of

> each chunk attended over slots a later position had overwritten. 36 of the 48

> layers are sliding, so long-prompt comprehension was badly degraded — asked to

> summarise 7 000 tokens of Manzoni's Italian prose, the earlier build answered

> about "Logos" and, on a second passage, about pasta sauce. Short prompts

> (under ~128 tokens) were never affected, which is why the API validation

> suite passed throughout. The same defect gave the two-row MTP verify one stale

> key past depth 128, so the "committed only on an exact match against the

> target's own argmax" guarantee did not hold there either.

>

> This is a serving-engine defect, not an artifact defect: the GGUF files are

> unchanged and the quality numbers below were measured on llama.cpp, not ds4.

1 — get the weights (85.56 GiB across three shards):

hf download Baekpica/K-EXAONE-236B-A23B-Mixed-Quant-GGUF \
  --include 'K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-*.gguf' \
  --local-dir ./K-EXAONE-mixed

2 — build the engine. make cuda-spark is the GB10 target; it forces

CUDA_ARCH=sm_121 across every binary. Building for the wrong architecture is

the single most common way to get wrong kernel results here.

git clone https://github.com/Baekpica/ds4
cd ds4
git checkout b2faf06f1ce2702efa53ac17145b1f56d3fb23b8
make cuda-spark

3 — serve. This is the exact command validated below:

./ds4-server \
  -m ../K-EXAONE-mixed/K-EXAONE-236B-A23B-MXQ-IQ2XXS-Q3K-Q4Edge-Q8Dense-MTPQ8-v1-00001-of-00003.gguf \
  --cuda \
  -c 262144 \
  --host 0.0.0.0 \
  --port 8001

The full 262 144-token context fits on one GB10 with the model resident.

Cold start to listening is about 3 min 45 s, dominated by the one-time

alignment repack. Give the machine ~119 GiB free before starting: the loader

peaks higher than its steady state.

4 — call it. Any OpenAI client works; point base_url at

http://<host>:8001/v1. /v1/models advertises the loaded model as

k-exaone-236b-a23b (thinking on by default) and

k-exaone-236b-a23b-chat (direct answer).

curl -sS http://127.0.0.1:8001/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "k-exaone-236b-a23b-chat",
    "messages": [{"role": "user", "content": "대한민국의 수도는 어디인가요?"}],
    "temperature": 0,
    "max_tokens": 64
  }'

Useful flags: --batched-session N keeps N resident sessions and batches

decode-ready requests (concurrency); --exaone-mtp / --exaone-mtp-timing

enable the MTP path; --kv-disk-dir enables disk KV checkpoints — not

recommended for this model, keep K-EXAONE on in-memory KV.

Budget --batched-session carefully. Each resident session owns its KV and

small session state, but the **1.60 GiB prefill graph workspace is shared once

per server**, not multiplied per slot. The previous per-slot-workspace formula

is obsolete. --batched-session 8 -c 40960 has been boot-validated with about

8.4 GiB free on this host; still leave an operational margin because CUDA's

retained unified-memory pool makes simple host-RAM accounting misleading.

Resident memory at -c 262144

The 85.56 GiB GGUF is mapped once and left unpinned; ds4 then materialises

the weights the CUDA backend actually reads. Those two are alternatives, not

additions — the mapping's pages are handed over, not duplicated:

| Component | Size |

|---|---:|

| Aligned CUDA artifacts (repacked at load) | 39.09 GiB — 78 IQ2 tensors 30.16 GiB + 345 Q8 tensors 8.93 GiB |

| Raw expert cache payload | 45.39 GiB |

| Weights resident on device | 84.48 GiB |

| KV cache, 262 144 tokens (12 full + 36 sliding layers) | 12.30 GiB |

| Graph workspace (shared across sessions) | 1.60 GiB |

| Context buffers (prefill_chunk 2048) | 104.22 MiB |

| nvtop GPU Mem, idle and ready at 256K | 103.95 GiB / 121.6 GiB |

That leaves roughly 18 GiB of headroom on a 121.6 GiB machine with the largest

context the model supports already allocated.

The LLLG schedule is what makes 256K affordable: only 12 of 48 layers keep a

full-context KV, the other 36 keep a 128-position sliding window, so KV costs

48 KiB/token instead of the ~192 KiB/token a fully global GQA stack would

need.

Reading the memory numbers

GB10 is a coherent unified-memory device, so "GPU memory" and "host memory" are

the same physical pool and the usual tools disagree about who owns it:

  • nvidia-smi --query-gpu=memory.used reports [N/A] on GB10. Use

nvtop -s and read the process's gpu_mem_bytes_alloc.

  • Process VmRSS understates residency by design: ds4 leaves the 85.56 GiB

model mmap unpinned and hands the pages to the CUDA cache, so most of the

footprint is CUDA-owned rather than process-anonymous.

  • A small-context run is not comparable to a 256K run. The same build with two

127-token sessions peaks near 90.17 GiB; the 256K server sits at

103.95 GiB. The difference is almost entirely the 12.30 GiB 256K KV.

  • After a clean exit the driver keeps the memory, and that is fine. With

595.71.05, free reports roughly 14 GiB available after ds4-server exits,

and it stays there: the whole of /proc/meminfo accounts for only ~17.7 GiB

of the 127.5 GiB total, so the ~110 GiB is held by the NVIDIA kernel module,

not by page cache. Dropping caches cannot reclaim it — there is nothing in

the page cache to drop.

It also does not need reclaiming. The next CUDA process reuses the driver's

pool directly: a second 256K server booted normally in 230 s with

MemAvailable still showing 14 GiB. **The precondition for booting is that

no other ds4-server is running — not a MemAvailable threshold.** A

readiness check that waits for free memory will wait forever.

Measured throughput

Greedy (temperature: 0), thinking disabled, 128 generated tokens per request,

one cold prompt per measurement over /v1/chat/completions with streaming.

Every frontier uses a disjoint corpus slice. Prefilled excludes the tiny

template prefix already resident in the two short cells; prefill t/s is

prefilled / TTFT. decode t/s is measured between the first and last content

chunk.

| Frontier | Prompt | Prefilled | Prefill t/s | Decode t/s | TTFT |

|---:|---:|---:|---:|---:|---:|

| 2K | 1 451 | 1 387 | 269.6 | 10.75 | 5.14 s |

| 8K | 7 925 | 7 923 | 276.5 | 10.46 | 28.66 s |

| 32K | 31 300 | 31 300 | 245.0 | 9.00 | 127.78 s |

| 64K | 64 663 | 64 663 | 207.3 | 7.27 | 311.90 s |

The previous published engine was roughly 54–56 t/s through 34K. The measured

31.3K cell is now 4.5x faster, and 64K is measured rather than projected. The

2K/8K rows were remeasured on the pinned commit after server warm-up. The

32K/64K cells were taken on the final optimization candidate immediately

before a safety-only invalid-token guard was added to batch embedding; valid

token IDs use the same numerical path. Raw records and exact methodology are

in reports/DGX-SPARK-PREFILL-OPT-2026-08-09.md.

Decode still declines with context depth because 12 full-attention layers read

the growing KV history. Cold 256K prefill has not been measured end to end;

this card intentionally no longer publishes a 256K time extrapolated from the

old 55 t/s kernel stack.

What that means in practice

The 262 144-token context fits, is allocated, and is resident. A cold 31.3K

prompt now reaches first token in 127.8 s and a cold 64.7K prompt in 311.9 s;

decode is 9.0 and 7.3 t/s at those depths. The request-only profile now puts

prefill attention at 2.6% and QK norm/RoPE at 0.6%. The dominant work is the

quantized MoE and dense stack — IQ2 gate/up, Q3 down, dense/shared Q8, then Q4.

Warm continuations still skip almost all of the prefill entirely (below).

Multi-turn chat reuses the prefix; a cold prompt does not. A continuation

resumes at the point where it diverges from what the session already holds, so

only the tail is prefilled:

| Turn | Prompt tokens | Time to first token | Reused |

|---|---:|---:|---:|

| 1 — cold, ~7K document + question | 6 978 | 165.7 s | 0 |

| 2 — same history + the assistant's own reply + a follow-up | 7 083 | 5.9 s | 6 992 |

| 3 — a different document, cold | 6 725 | 137.0 s | 0 |

Turn 2 is 24× faster than the same request without reuse, and turns 1 and 3

are unchanged — an unrelated prompt is not falsely matched onto a live session.

This is worth spelling out because it is the case an all-or-nothing prefix test

gets wrong, and ds4 used to have one. A chat client replays the assistant's

previous reply as text, and re-tokenising it does not reproduce the token IDs

the model sampled. The old test required the new prompt to contain the entire

checkpoint, so a continuation sharing 6 984 of 7 086 tokens — 98.6 % — failed it

and re-prefilled everything, at 143.9 s per turn. It now resumes at the

divergence point instead.

How far back that can reach is a property of the sliding-window KV ring rather

than a tunable: the ring is window + prefill chunk wide, so a divergence

further back than about 2 000 tokens falls back to a cold prefill. Typical chat

divergence is one assistant turn, well inside it. Requires ds4 at the commit

pinned above.

Concurrency now helps. ds4's cross-session row batching (2026-08-08) runs

concurrent decode steps through one pass — the weight-bound stages are read

once for all streams — so aggregate throughput rises with load instead of

staying flat. Steady-state aggregate decode, --batched-session 8, short

prompts, all streams decoding:

| Concurrent streams | Summed decode t/s | Per stream | before row batching |

|---:|---:|---:|---:|

| 1 | 11.5 | 11.5 | 11.1 |

| 2 | 14.8 | ~7.4 | 9.8 |

| 4 | 16.3 | ~4.1 | 10.0 |

| 8 | 18.5 | ~2.3 | 10.8 |

An operator serving several users sees ~15–18 tok/s of total output; a single

user still sees the single-stream rate above. The remaining per-row floor is

mostly the routed experts — concurrent tokens route to largely disjoint

top-8-of-128 sets, so that read genuinely cannot amortise — plus the per-row

attention, which is per-session by construction.

A prefill no longer blocks the batch either: a pending prefill quantum rides

the decode batch's weight sweep (+prefill in the batch log), so admitting a

new long prompt costs the running streams far less than alternating whole

passes did.

One contract changed with row batching: **greedy output across batch widths is

not bit-stable at near-ties**. A request decoded alongside seven others can

pick a different token than the same request alone where the top-2 margin is

tiny, deterministically per batch composition. Sequential (width-1) decode is

unchanged, and the same batch always reproduces the same output.

Where the time goes

The 12 full-attention layers hold 49 152 bytes of KV per context position

(GQA, 8 KV heads × 128 dims, K and V, f16). Decode adds **0.61 µs per context

position** — an effective ~81 GB/s of KV read against roughly 273 GB/s of

device bandwidth. An earlier engine paid 5.97 µs here (~3 % of bandwidth,

one attention block per head); the flash-decode split closed most of that,

and what remains splits between the depth-independent floor (~94 ms/token,

streaming the active weights, near the roofline) and the last ~3× of the

attention read.

The request-only Nsight slice, immediately before the final 64-column-tail

increment, is now: IQ2 aligned gate/up D2R 30.66%, Q3 routed down

23.65%, dense/shared Q8 14.97%, Q4 routed 10.77%, batch embedding

4.90%, prefill attention 2.62%, and QK norm/RoPE 0.60%. That is the next-work

order; attention is no longer the first-order bottleneck.

For scale, the cited Spark Arena vLLM run

reports about 1,233 t/s at pp2048, but it is a different 180B

DeepSeek-V4-Flash checkpoint using MXFP4 experts, FP8 MLA KV, an 8,192-token

batch budget, async scheduling, prefix cache, graphs/compilation, a custom

model modification, and MTP. It is a useful ceiling, not an apples-to-apples

runtime comparison with this 237B top-8 GQA IQ2/Q3/Q4 GGUF.

OpenAI-compatible API

/v1/chat/completions, /v1/completions, /v1/responses and /v1/messages

are served; /v1/models advertises k-exaone-236b-a23b and

k-exaone-236b-a23b-chat. The bare ID keeps thinking on by default; the chat

ID answers directly. Both serve the same loaded GGUF. There is no

llama.cpp-style /health or /props — probe /v1/models plus a real

completion. Validated on this host, greedy (temperature: 0):

| Check | Result |

|---|---|

| GET /v1/models | serves the bare and -chat K-EXAONE IDs |

| non-streaming chat completion | finish_reason=stop, correct Korean answer, usage populated |

| streaming chat completion | SSE chunks, finish_reason, and — with stream_options: {"include_usage": true} — a final usage chunk |

| streamed text == non-streamed text | identical under temperature: 0 |

| thinking mode | reasoning_content arrives in its own delta field, never inlined into content |

| four sequential requests | no state carried between them; repeating the first request reproduces it byte-for-byte |

Two behaviours worth knowing before you benchmark:

  • Thinking is on by default for chat requests. With max_tokens: 64 the

budget is spent inside reasoning_content and content comes back empty —

that is correct, not a hang. Select k-exaone-236b-a23b-chat, or send

"thinking": {"type": "disabled"} / "think": false, for short factual

answers.

  • Streaming usage follows the OpenAI rule: no stream_options.include_usage,

no usage chunk.

Multi-token prediction (blk.48)

ds4 executes the trained MTP block from this same GGUF — no separate draft

model, no second weight copy. It is opt-in and off by default:

--exaone-mtp enables it, --exaone-mtp-timing adds per-cycle counters.

  • Input ordering is the trained one,

enorm(embed(x[p+1])) || hnorm(target_hidden[p]), with the decoder position

explicitly shifted to p + 1.

  • Every draft is verified against the target model's own argmax and

committed only on an exact token-ID match, so speculation cannot change

greedy output. A 64-token identity test passes with plain == MTP,

mismatch = -1.

  • Extra runtime state is 0.50 MiB (a 128-row private f16 KV ring).
  • Speculation runs for greedy requests only (temperature: 0).
  • An automatic loss quench watches the first 12 verifier cycles and disables

speculation for the rest of the session when measured MTP work runs more than

3 % slower. DS4_EXAONE_MTP_NO_QUENCH=1 defeats it, for measurement only.

It is close to a wash now, and still not a win. Measured with the quench

defeated so the whole generation is speculative, on the engine's current

kernels (the two-row verify rides the same aligned-dispatch, small-batch and

flash-decode tiers as everything else):

| Context | Draft acceptance | MTP ms/token | Plain ms/token | vs plain |

|---:|---:|---:|---:|---:|

| 1 387 | 36.6 % | 108.8 | 98.3 | +11 % |

| 7 752 | 60.8 % | 102.5 | 99.7 | +2.8 % |

| 33 914 | 37.0 % | 142.8 | 114.5 | +25 % |

The mechanism is a single ratio: a cycle runs one draft pass plus one

two-row target verify pass; with k the verify's cost relative to a

one-row decode and a the acceptance, a cycle commits 1 + a tokens and

wins exactly when k < 1 + a. An earlier engine paid k ≈ 2 at shallow depth

because its two-row pass re-read the weights per row; that k is now near its

floor, which moved MTP from a 26–50 % loss to the table above. What remains

is acceptance: at the 3 % quench threshold the 8K row is already a wash, and

five to ten more points of acceptance — a warmed MTP ring instead of a cold

128-row one, or corpus luck — is the difference between off and on.

Acceptance numbers are a property of the text (they moved 20 points between

corpus slices in these very measurements); compare k across runs, not

acceptance.

MTP stays off by default. The auto-quench makes --exaone-mtp safe to

try on workloads where drafts land often; nothing here changes greedy output

either way.

Measured quality

32 fixtures, greedy (temperature=0, top_k=1), reasoning off,

max_tokens 768, compared against the same fixtures run on the official

Q8_0 build (234.7 GiB) as reference.

Both sides were run on llama.cpp, on 4 × RTX PRO 6000 (sm_120) — this

table measures the artifact, not the ds4 serving path.

| | pilot Q2_K · 87.84 GiB | v1 IQ2_XXS+Q3_K · 85.56 GiB |

|---|--:|--:|

| word-agreement vs Q8_0, mean | 0.139 | 0.183 |

| — json / tool-call | 0.250 | 0.681 |

| — long-context retrieval | 0.364 | 0.450 |

| identical outputs | 2/32 | 3/32 |

| JSON parses | 4/4 | 4/4 |

| needle retrieved | 3/4 | 3/4 |

| broken-jamo ratio | 0.0001 | 0.0003 |

| repetition (3-gram) | 0.022 | 0.022 |

| decode, 4 × RTX PRO 6000 | 66.5 tok/s | 77.6 tok/s |

v1 tracks the Q8_0 reference more closely than the pilot **while being

smaller** — the importance matrix and Q3_K down are doing real work, most

visibly on structured output. Absolute agreement is low for both because greedy

long-form generation diverges after a single differing token; the pair track for

about 11 words on average before separating. The task-level outcomes (JSON

validity, retrieval, no jamo collapse, no repetition loops) match the Q8_0

reference's own scores.

Limitations

  • IQ2_XXS on routed gate/up is aggressive. The recipe protects embeddings,

attention, router, shared expert, dense layer 0, and the edge MoE layers

specifically to offset it, but expect degradation relative to Q4_K_M on

tasks that lean on rarely-activated experts.

  • Evaluation is a 32-prompt fixture set plus the token-fidelity comparison

above, not a full benchmark suite. Raw results, including the failures, ship

in the converter repository.

  • The full 256K context fits, but cold 256K prefill is not measured. The

deepest cold cell here is 64,663 tokens at 207.3 prefill t/s and 7.27 decode

t/s. Prefix reuse remains the practical way to reach deep contexts without

paying the whole cold admission cost.

  • The remaining prefill gap is mainly quantized linear algebra. In the

request-only profile, IQ2 gate/up + Q3 down + dense/shared Q8 + Q4 account

for about 80% of GPU kernel time; prefill attention is 2.6% and QK

norm/RoPE is 0.6%.

  • The MTP block only runs under ds4, on the pinned branch and commit above.

Under llama.cpp it is inert. There is no third runtime that executes it.

  • MTP is a mild loss (3–25 % by depth and corpus), so it ships off by

default and auto-quenches when enabled. Its verify cost is now near its

floor; the remaining limit is draft acceptance.

  • MTP does not run under --batched-session. ds4 disables speculative

decoding whenever native session batching is active, so concurrency > 1 is

plain decode regardless of the MTP flags.

  • Greedy output across batch widths is not bit-stable at near-ties. A

request decoded alongside others can pick a different token than the same

request alone where the top-2 margin is tiny; the same batch composition

always reproduces the same output, and width-1 decode is unchanged. This is

the standard batched-inference contract.

  • Multi-turn reuse reaches back about 2 000 tokens. The sliding-window KV

ring is window + prefill chunk wide, and a resumed prefill needs the window

that preceded its restart point. A conversation that diverges further back

than that — an edited early message, a re-ordered history — falls back to a

cold prefill. One assistant turn of divergence, the normal case, is well

inside it.

  • Host memory accounting is not usable as a readiness signal. Driver

595.71.05 retains the unified allocation after a clean exit and does not

return it to the kernel. Gate a restart on "no ds4-server process", not on

free.

Acknowledgements

This artifact is only interesting because there is an engine that serves it, and

that engine is almost entirely other people's work.

DwarfStar/ds4, the original engine. Everything here is downstream of it: the

GGUF loader, the session and KV machinery, the OpenAI/Responses/Anthropic

server, the CUDA backend, the MoE routing path, and the NextN/MTP scheduling

contract that the blk.48 work slots into. The exaone-moe family is a new

model family added to his architecture, not a new engine.

DGX Spark fork. The GB10 story rests on it: the sm_121 build target, and

above all the aligned-artifact tier — repacking IQ2/Q8 tensors into

alignment-correct CUDA artifacts — which is what makes a mixed-quant MoE

actually run at speed on GB10's unified memory. That tier was adopted from

this fork, along with its warning that single-stream MTP can lose on Spark,

which our own measurements then confirmed.

Georgi Gerganov and contributors. GGUF, llama-quantize, the importance-matrix

tooling, and the IQ2_XXS/Q3_K/Q4_K/Q8_0 formats are all theirs; this

artifact is a llama-quantize output, and llama.cpp is still the reference

runtime for it.

itself, including the trained MTP block that blk.48 executes.

Mistakes in the recipe, the exaone-moe port, and the measurements here are

ours, not theirs.

License and attribution

Base model © LG Management Development Institute, under the

K-EXAONE AI Model License Agreement (LICENSE, shipped alongside).

Per §2.1, this Agreement is distributed with the artifact and the derivative

name begins with "K-EXAONE". Note §2.2: distributing or sublicensing the model

or derivative works **to third parties for commercial purposes requires a

separate agreement with the Licensor**. This repository is published as a

research artifact; commercial redistribution is not granted by it.

Run Baekpica/K-EXAONE-236B-A23B-Mixed-Quant-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models