Baekpica/K2-Horizon-375B-A23B-Mixed-Quant-GGUF overview
K2 Horizon 375B A23B Mixed Quant MQ87 GGUF Update 2026 09 05 round 6, Prefill 11 + Prefill 12 : on one NVIDIA DGX Spark GB10, sm 121a the round 6 binary measur…
Runs locally from ~3.55 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| K2-Horizon-375B-A23B-MQ87-00001-of-00004.gguf | GGUF | GGUF | 27.85 GB | Download |
| K2-Horizon-375B-A23B-MQ87-00002-of-00004.gguf | GGUF | GGUF | 27.50 GB | Download |
| K2-Horizon-375B-A23B-MQ87-00003-of-00004.gguf | GGUF | GGUF | 27.80 GB | Download |
| K2-Horizon-375B-A23B-MQ87-00004-of-00004.gguf | GGUF | GGUF | 3.55 GB | Download |
Model Details
| Model ID | Baekpica/K2-Horizon-375B-A23B-Mixed-Quant-GGUF |
|---|---|
| Author | Baekpica |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | IFM/K2-Horizon-375B-A23B |
| Last modified | 2026-09-05T12:24:31.000Z |
Model README
---
license: apache-2.0
base_model: IFM/K2-Horizon-375B-A23B
pipeline_tag: text-generation
library_name: gguf
tags:
- k2-horizon
- gguf
- mixed-quant
- iq1_s
- iq1_m
- iq2_xxs
- iq2_xs
- imatrix
- moe
---
K2-Horizon-375B-A23B Mixed Quant MQ87 GGUF
> Update 2026-09-05 (round 6, Prefill 11 + Prefill 12): on one NVIDIA
> DGX Spark (GB10, sm_121a) the round-6 binary measures
> 644.78 / 641.94 tok/s cold prefill and **13.34 / 13.07 tok/s
> decode** on the campaign fixture (first 8,192 tokens of
> promessi_sposi.txt plus 64 greedy tokens, context 8,257; campaign
> baseline 286.55 / 5.52). Both rounds are scheduling only and
> byte-identical to their kill switches and to the round-5 binary.
> Prefill 11
> (8e886f1)
> pipelines the K loop of the compact worklist MMQ kernel that runs
> every routed expert projection (next block's weights prefetched into
> registers behind the MMA, activation halves staged with cp.async,
> two barriers per iteration instead of four): +5.3% vs its kill
> switch on the same binary. Prefill 12
> (12a2e14)
> reads the HMMA prefill-attention fragments with ldmatrix and
> prefetches the next K/V tile: +2.4%. Within one session the two
> add up to 595.79 / 597.52 → 644.78 / 641.94 (+8.0%); the round-5
> numbers below were taken in an earlier session (prefill drifts a few
> percent between sessions on one binary, so only same-binary pairs
> are claimed).
>
> Round 5 (Prefill 9 + Prefill 10): the round-5 binary measures
> 622.48 / 622.48 tok/s cold prefill and **13.42 / 13.35 tok/s
> decode** on the same fixture. Prefill 10
> (fea285f)
> prefills K2 in 1024-token chunks instead of 512: +23.0% vs its
> kill switch on the same binary, an fp-level change (router GEMM
> tiling, attention accumulation order) whose frontier swaps the two
> near-tied top tokens; see the note under the campaign table. Prefill 9
> (6ff7561)
> drops the standalone sanitize pass on the routed-down output (+2.7%,
> bit-identical). Round 4 (b5ca173 IQ1 gate/up pair, dfaa5a8 cached
> RoPE angles, 494.85 / 13.34) and round 3 (3afea1b, 125528a,
> 466.32 / 13.19) were bit-identical scheduling rounds; Prefill 6,
> Decode 2, Decode 3 and Prefill 10 change arithmetic: see the
> numeric-contract notes. Decode sits at the GB10 memory wall (72 ms of
> bandwidth-bound kernels per 76 ms step).
MQ87 is a memory-targeted mixed quant of
designed around an approximately 87 GiB tensor payload for a 128 GB unified-memory
system such as DGX Spark.
The model is converted from the original BF16 checkpoint at revision
d33e3ae45281865ebf9f044b12d3635b1d1e17fe. It is never requantized from FP8.
Support my work
I work on making large language models practical on hardware they were never really designed to fit on — through mixed quantization, inference optimization, custom kernels, and serving experiments.
While much of the development happens on local hardware, calibration, profiling, and large-scale validation often require expensive on-demand GPUs.
Contributions help pay for that compute, storage, and testing infrastructure, so I can keep experimenting and publishing the results openly.
<a href="https://www.buymeacoffee.com/baekpica" target="_blank"><img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" alt="Buy Me a Coffee" style="height: 60px !important;width: 217px !important;"></a> <a href="https://github.com/sponsors/Baekpica" target="_blank"><img src="https://img.shields.io/badge/Sponsor-EA4AAA?style=for-the-badge&logo=githubsponsors&logoColor=white" alt="Sponsor Baekpica on GitHub" style="height: 60px !important;width: 217px !important;"></a>
MQ87 policy
| Tensor group | Type |
|---|---:|
| token embedding and untied output head | Q8_0 |
| all Q/K/V/O attention projections | Q8_0 |
| dense MLP layers 0-2 | Q8_0 |
| all always-active shared experts | Q8_0 |
| router weights/biases, RMSNorm and 1-D control | F32 |
| routed gate/up, layers 7-56 | IQ1_S + imatrix |
| routed down, layers 7-56 | IQ2_XXS + imatrix |
| routed gate/up, layers 3-6 and 57-60 | IQ1_M + imatrix |
| routed down, layers 3-6 and 57-60 | IQ2_XS + imatrix |
The static tensor calculation is 93,080,440,320 bytes, or
86.687915325 GiB, before GGUF metadata/alignment. A fail-closed quantizer
dry-run reproduced exactly 88,768.43 MiB (rounded by llama.cpp to 1.96 BPW)
with all 842 planned tensor types accounted for. Finished-file measurements are
93,091,935,552 bytes across four split files (86.698621094 GiB). The tensor
payload is exactly 93,080,440,320 bytes; the remaining 11,495,232 bytes are
GGUF metadata, tensor descriptors, and alignment.
Expected output tensor-type counts are:
Q8_0: 429F32: 239IQ1_S: 100IQ2_XXS: 50IQ1_M: 16IQ2_XS: 8- Total: 842
Calibration
The importance matrix was measured with the original BF16 model forward pass,
streamed one layer at a time across five H200 GPUs. The forward implementation was
checked tensor-for-tensor against the official Transformers model for a dense layer
and a complete sparse-MoE layer; both checks had zero max/mean/p99 error in BF16.
Calibration size and composition:
| Category | Share | Tokens |
|---|---:|---:|
| coding / tool use | 30% | 1,572,864 |
| math / science reasoning | 25% | 1,310,720 |
| general reasoning | 20% | 1,048,576 |
| long-form / retrieval | 15% | 786,432 |
| general text | 10% | 524,288 |
| Total | 100% | 5,242,880 |
The primary source is the private
Baekpica/Solar-Open2-120B-A15B-REAM-148E-Healing-Mix, supplemented with an
independent multilingual/general corpus. Raw text and the token matrix are not
redistributed. Their pinned hashes and aggregate measurement results are retained
in the calibration report.
Healing-Mix token rows were decoded with its pinned Solar tokenizer and then the
resulting text was re-tokenized with K2-Horizon's own tokenizer
(tokenizer.json SHA-256 2fa69519ff1ee3ee55ac50927b00058e88b9eb90802fabaf3c6b05804601c53b).
The K2 forward pass therefore never interprets Solar token IDs as K2 IDs.
Coverage gate result:
- Exact routes observed:
2,432,696,320(5,242,880 x 58 x top-8). - All empirically reachable layer/expert cells were measured.
- Minimum reachable hit count:
1,726after excluding bias-dominated cells. - Global reachable p05:
61,068; worst per-layer reachable p05:10,310. - Unexplained zero-hit cells:
0. - Three zero-hit cells are mathematically blocked by router bias. They receive a
neutral importance value and are explicitly listed in the report rather than
silently filled. A fourth bias-dominated, endpoint-tie-only cell was measured
770 times and therefore has measured importance.
- Each of the five category partitions independently covered all 11,132
non-bias-dominated layer/expert cells. Their lowest positive counts were 187
(coding/tool), 228 (math/science), 449 (general reasoning), 186
(long/retrieval), and 20 (general text).
Files under calibration/ contain the final imatrix, per-layer/per-expert route
counts, the aggregate report, and BF16-forward validation reports. Private corpus
content is intentionally absent.
Reproduction
The exact first-match tensor policy is in reproduction/mq87.tensor-types.txt.
The final quantization command is equivalent to:
llama-quantize \
--pure \
--imatrix calibration/K2-Horizon-375B-A23B-MQ87.imatrix \
--tensor-type-file reproduction/mq87.tensor-types.txt \
K2-Horizon-375B-A23B-BF16-00001-of-00030.gguf \
K2-Horizon-375B-A23B-MQ87.gguf \
Q8_0 192
The original BF16 GGUF and Q8_0 intermediate are hosted in
Baekpica/K2-Horizon-375B-A23B-GGUF.
Artifact inventory
| File | Bytes |
|---|---:|
| K2-Horizon-375B-A23B-MQ87-00001-of-00004.gguf | 29,905,225,504 |
| K2-Horizon-375B-A23B-MQ87-00002-of-00004.gguf | 29,527,310,912 |
| K2-Horizon-375B-A23B-MQ87-00003-of-00004.gguf | 29,846,576,608 |
| K2-Horizon-375B-A23B-MQ87-00004-of-00004.gguf | 3,812,822,528 |
MQ87-SHA256SUMS contains content hashes generated from the closed local
files. validation/MQ87.audit.json records the split metadata, byte counts,
and exact 842-tensor type inventory.
Runtime status
GGUF creation uses the IFM K2-Horizon llama.cpp branch at commit
35999d101cf2233fc54f09c3c8d599da7303ce02. A single-H200 smoke test loaded all
four shards with full GPU offload and completed a one-token forward pass at
103.0 prompt tok/s with exit code 0. Compatibility with
Baekpica/ds4-dfm-rs support is
available on branch feature/k2-horizon-375b-mq87 at commit d52d964.
On H200, 32 host tests, all synthetic/actual mixed-IQ CUDA checks, and a full
61-layer raw-token forward passed. The forward currently needs diagnostic
DS4_MEMGOV=observe because default enforcement produces a documented
cold-materialization admission disagreement. DGX Spark GB10/sm_121a physical
validation and that memgov follow-up remain required before production serving;
see validation/ds4-k2-h200.audit.json.
Production status
Production and upload are complete. All four public LFS objects were checked
against MQ87-SHA256SUMS: 4/4 files passed, totaling 93,091,935,552 bytes.
validation/MQ87.remote.audit.json records the independent Hub-side size and
SHA-256 verification.
DGX Spark ds4-dfm-rs optimization (2026-09-05 continuation)
New rounds on ds4-dfm-rs branch
feature/k2-horizon-375b-serving, counted from HEAD 2c56fc9. Earlier
IQ2_XXS / IQ1_S worklist work is historical and is not counted here.
Host: one NVIDIA DGX Spark (GB10, sm_121a), CUDA 13.3, Rust ds4-bench,
MQ87 four-shard GGUF. Cold process, default memgov, no MTP, no prefix reuse.
./ds4-bench --cuda \
-m K2-Horizon-375B-A23B-MQ87-00001-of-00004.gguf \
--prompt-file speed-bench/promessi_sposi.txt \
--ctx-start 8192 --ctx-max 8192 --gen-tokens 64
Corpus: raw promessi_sposi.txt (SHA-256
f53e0d80cb2d4492d24ebd63c7000c397b16ae70f9bf09b3763e5d8323ec209f),
first 8,192 prompt tokens, 64 greedy decode tokens, allocated context 8,257.
Load/repack time is excluded from tok/s. Prefill chunk is the engine default
(512 through round 4, 1024 for K2 from Prefill 10 on).
| Round | Change | Commit | Prefill tok/s | vs baseline | Decode tok/s | vs baseline |
|---|---|---|---:|---:|---:|---:|
| baseline | HEAD nsys cell | 2c56fc9 | 286.55 | — | 5.52 | — |
| Prefill 1 | IQ1_M assign-major MMVQ (3-D grid, ncols=1 walk). DS4_MMQ_IQ1M_PREFILL=0 restores the per-token loop. | 096fc9c | 295.04 | +3.0% | 5.47 | -0.9% |
| Prefill 2 | Compact worklist for raw IQ2_XS down | reverted | 272.03 / 272.14 | -5.1% | 5.51 | -0.2% |
| Prefill 3 | IQ1_M slot-loop: one (M, tokens) block reuses the Q8_1 row across top-8. DS4_MMQ_IQ1M_SLOT_LOOP=0 restores the 3-D grid. | 569d13c | 302.94 | +5.7% | 5.50 | -0.4% |
| Prefill 4 | IQ1_M row-tile (4 rows/block) | reverted | 290.18 / 295.28 | -4.2% | 5.50 | -0.4% |
| Prefill 5 | 1024-token prefill chunks | reverted | 307.65 | +7.4% | 5.51 | -0.2% |
| Decode 1 | Share KV across 2 GQA query heads. DS4_EXAONE_ATTN_GQA=0 is one block per head. Full-group tile-6 was 4.13 tok/s and is not used. | c9ebd10 | 270.88 | — | 6.91 | +25.2% |
| Prefill 6 kill switch | DS4_MMQ_IQ1M_WORKLIST=0 on the Prefill 6 binary (assign-major MMVQ, same path as Decode 1) | same binary | 305.11 / 305.04 | +6.5% | 7.14 | +29.3% |
| Prefill 6 | IQ1_M MMQ tile: load_tiles_iq1_m on the per-16 Q3_K/IQ2_XS layout, compact worklist from 256 routed rows. Layer-3 gate kernel 52.98 → 5.68 ms at 512 tokens. | 7d756a9 | 444.50 / 444.65 | +55.1% | 7.14 | +29.3% |
| Decode 2 kill switch | DS4_EXAONE_ATTN_SPLIT=0 on the Decode 2 binary (pair kernel, same path as Prefill 6) | same binary | 445.20 / 444.34 | +55.4% | 7.11 / 7.11 | +28.8% |
| Decode 2 | Split-K decode attention: 256-key chunks, one block per (chunk, KV head), combine kernel; from 2048 keys. Attention kernel 1.214 → 0.584 ms at 8,191 keys. | e055726 | 446.40 / 444.19 | +55.8% | 9.80 / 9.80 | +77.5% |
| Round 3, all defaults | Prefill 7 + Decode 3 | round-3 binary | 466.32 / 466.06 | +62.7% | 13.19 / 13.27 | +139% |
| Prefill 7 kill switch | DS4_MMQ_IQ2XS_WORKLIST=0 (rectangular IQ2_XS schedule) | same binary | 446.93 | +56.0% | 13.26 | +140% |
| Prefill 7 | IQ2_XS down tensors join the compact MMQ worklist (scheduling only, bit-identical). Layer-3 down kernel 15.42 → 7.85 ms at 4,096 rows. | 3afea1b | 466.32 / 466.06 | +62.7% | 13.19 / 13.27 | — |
| Decode 3 kill switch | DS4_EXAONE_ATTN_SPLIT_NATIVE=0 (Solar grouped split kernel, the Decode 2 path) | same binary | 465.77 | +62.5% | 9.81 | +77.7% |
| Decode 3 | The f16 pair kernel runs per 256-key chunk and writes partials for the combine kernel instead of the Solar grouped kernel. Attention kernel 0.587 → 0.159 ms at 8,191 keys (pair kernel 1.327). | 125528a | 466.32 / 466.06 | — | 13.19 / 13.27 | +139% |
| Round 4, all defaults | Prefill 8 + Decode 4 | round-4 binary | 494.85 / 497.37 | +72.7% | 13.34 / 13.26 | +141.7% |
| Prefill 8 kill switch | DS4_MMQ_IQ1_PAIR=0 (two single IQ1 routed calls) | same binary | 484.67 | +69.1% | 13.32 | — |
| Prefill 8 | IQ1_S / IQ1_M gate/up take the K-quant pair path: one expert map and Q8_1 activation for both weights, both compact worklists, no standalone sanitize (consumers zero non-finite values at read). Same kernels, bit-identical. | b5ca173 | 494.85 / 497.37 | +72.7% | 13.34 / 13.26 | — |
| Decode 4 kill switch | DS4_EXAONE_ROPE_TABLE=0 (double-precision trig per head and pair) | same binary | 476.94 | +66.4% | 13.27 | — |
| Decode 4 | One (cos, sin) table per prefill chunk or decode token, shared by every layer's q and k QK-norm/RoPE call; same doubles, same rounding, bit-identical. Removes 40 ms of FP64 trig per 512-token chunk. | dfaa5a8 | 494.85 / 497.37 | +72.7% | 13.34 / 13.26 | — |
| Round 5, all defaults | Prefill 9 + Prefill 10 | round-5 binary | 622.48 / 622.48 | +117.2% | 13.42 / 13.35 | +143.1% |
| Prefill 9 kill switch | DS4_EXAONE_DOWN_SANITIZE=1 (standalone sanitize on the routed-down output) | same binary | 606.33 | +111.6% | 13.40 | — |
| Prefill 9 | The routed-down output only feeds moe_sum, which skips non-finite values at read; the standalone sanitize pass (58 launches per chunk) is dropped. Bit-identical. | 6ff7561 | 622.48 / 622.48 | +117.2% | 13.42 / 13.35 | — |
| Prefill 10 kill switch | DS4_EXAONE_PREFILL_CHUNK=512 (previous chunk) | same binary | 506.08 | +76.6% | 13.28 | — |
| Prefill 10 | K2 prefills in 1024-token chunks: ~42 routed rows per expert instead of ~21, half the per-chunk launches; workspace 0.37 → 0.73 GiB. fp-level change (router GEMM tiling, attention accumulation order). | fea285f | 622.48 / 622.48 | +117.2% | 13.42 / 13.35 | — |
| Prefill 11 kill switch | DS4_MMQ_PIPE=0 (upstream worklist K loop) | same binary | 595.79 / 597.52 | +107.9% | 13.32 / 13.32 | — |
| Prefill 11 | Software-pipelined K loop for the compact worklist MMQ kernel (IQ1_S / IQ1_M / IQ2_XXS / IQ2_XS tiles up to 64 wide): the next block's raw bytes are prefetched into registers behind the MMA phase and both activation halves are staged with cp.async one iteration ahead, two barriers per K iteration instead of four. Same dequantization and dots, bit-identical. Kernel time at 8,192 routed rows: IQ2_XXS down 9.58 → 7.08 ms, IQ2_XS down 11.43 → 8.23 ms, IQ1_S gate 7.29 → 6.74 ms. | 8e886f1 | 627.73 / 629.17 | +119.1% | 13.34 / 13.32 | — |
| Round 6, all defaults | Prefill 11 + Prefill 12 | round-6 binary | 644.78 / 641.94 | +125.0% | 13.34 / 13.07 | +141.7% |
| Prefill 12 kill switch | DS4_FATTN_HMMA_LDSM=0 (scalar shared-memory fragment loads, direct tile fill) | same binary | 629.08 / 628.29 | +119.5% | 13.35 / 13.33 | — |
| Prefill 12 | The GQA-pair HMMA prefill attention kernel reads its K and V mma fragments with ldmatrix / ldmatrix.trans (8 + 8 shared loads per 16-key step and lane instead of 96) and prefetches the next 64-key K/V tile into registers while the current one is consumed. Same tile bytes, same mma order, bit-identical. Synthetic 2048-token full attention 1.594 → 1.259 ms. | 12a2e14 | 644.78 / 641.94 | +125.0% | 13.34 / 13.07 | — |
Prefill 1 vs last accepted (the locked baseline): +8.49 tok/s.
Prefill 3 vs last accepted (096fc9c 295.04): +7.90 tok/s.
Same-binary Prefill 3 kill switch (DS4_MMQ_IQ1M_SLOT_LOOP=0): 294.38 / 5.51.
All 250,624 frontier logits and the 64 greedy IDs are bit-identical across
the kill switch, Prefill 1, and Prefill 3. Decode tok/s is not a claimed gain.
Prefill 2 was bit-identical to its kill switch and faster than that
rectangular path (262.86 → 272.03 tok/s), but both cells sit below the
locked 286.55 baseline, so the patch is not in the release binary.
Prefill 4 was bit-identical but slower than Prefill 3 (same-TU kernel
body change). Prefill 5 reached 307.65 tok/s but moved 56 of 64 greedy
IDs (frontier relative RMS 0.0636), so it is not in the release binary.
Decode 1 same-binary kill switch is 5.51 tok/s. The 270.88 prefill cell
is that compile, not a claimed prefill change. Logits and IDs match
Prefill 3.
Prefill 6 numeric contract. IQ1_M had no MMQ tile upstream, so the eight
edge-layer gate/up tensors ran the decode MMVQ per assignment (28.4% of
the post-Decode 1 8K+64 GPU trace). The new tile keeps the integer dot
exact (int8 = 8·(grid + delta), scale d·(2s+1)/8 per 16 values) but uses
the MMQ Q8_1 activation producer and MMA accumulation, the same tier the
IQ1_S / IQ2_XXS / IQ2_XS layers already use in prefill. Against the
assign-major MMVQ the kernel agrees at relative RMS 1.9e-4 on random
activations (the MMVQ tier stores the Q8_1 scale as fp16) and at
9.5e-8 / 2.6e-7 (synthetic / real layer-3 tensor) when both tiers see
identical Q8_1 activations, i.e. fp32 accumulation order only. On the
8K fixture the frontier
logits differ from Decode 1 by relative RMS 0.0624 (KL 0.022 nats, same
argmax, top-10 overlap 9/10) and the greedy continuation diverges at
token 6; the rejected Prefill 5 (1024-token chunks, no kernel change)
shows the same signature (0.0636, divergence at token 6), so this is the
model's near-tie routing sensitivity to any fp perturbation, not a
kernel fault. Two Prefill 6 processes are byte-identical to each other,
and the kill switch is byte-identical to Decode 1. Decode tok/s is not
a Prefill 6 claim (7.14 on both sides of the switch).
Decode 2 numeric contract. The whole-context pair kernel used 24 blocks
on 48 SMs and read the 8K KV at about 27 GB/s. The split kernel keeps the
per-key math and merges 256-key online-softmax partials in a different
fp32 order: rel RMS 7.1e-7 / 9.6e-7 / 1.8e-6 against the pair kernel at
4,095 / 8,191 / 32,767 keys, within 2e-5 of the CPU reference, and
bit-identical fallbacks below 2,048 keys, for sliding windows and under
the kill switch. Prefill and the frontier logits are unchanged
(byte-identical to Prefill 6); the greedy continuation shares its first 8
tokens with the kill switch and then takes a near-tie branch, both
continuations coherent. Two split runs are byte-identical.
Round 3. Prefill 7 is scheduling only: logits and the 64 greedy IDs are
byte-identical to its kill switch. Decode 3 keeps the pair kernel's
per-key math and the Decode 2 chunking; against the Solar grouped kernel
it agrees at relative RMS ≤ 4.4e-7 (4,095 / 8,191 / 32,767 keys) and
within 2e-5 of the CPU reference, the kill switch reproduces the Decode 2
greedy continuation exactly, the frontier logits are byte-identical, the
greedy continuation takes a near-tie branch at token 7, and the two
all-default runs are byte-identical.
Round 4. Both rounds are scheduling only; every round-4 cell (all
defaults, each kill switch, the repeat) is byte-identical in logits and
greedy IDs to the round-3 binary. The first round-4 build was not: nvcc
had re-associated the RoPE rotation once the angles came from a table,
one ulp per rotated value, enough to move this model's frontier
(relative RMS 0.063); the rotation is pinned to the original
contraction and verified against the original kernel bit for bit.
Round 5. Prefill 9 is scheduling only: byte-identical to its kill
switch, and the 512-chunk cell is byte-identical to round 4. Prefill 10
changes arithmetic at the fp level: against the 512-chunk cell on the
same binary the frontier has relative RMS 0.054, top-10 overlap 9/10, a
byte-identical repeat and coherent text on both sides, but the two
near-tied top tokens swap (" autori" 13.18 / " ese" 12.75 → " ese" 13.70
/ " autori" 13.39; the perturbation's max |Δlogit| is 1.03, the same as
the earlier rounds') and KL is 0.063 against the 0.05 band used so far.
Kernel rows at the 1024-token shapes (prefill attention vs CPU, gate/up
pairs at 8,192 routed rows, IQ worklists at 1,024 tokens / 8,192 rows)
are bit-identical or within tolerance. It is retained on that
assessment; DS4_EXAONE_PREFILL_CHUNK=512 restores the previous
default without a rebuild.
Round 6. Both rounds are scheduling only. Prefill 11 keeps the
upstream dequantization and MMA dots of the worklist kernel and only
changes when the bytes arrive (register prefetch of the next block,
cp.async activation stages); its 16 real-weight kernel cells (IQ1_S,
IQ1_M, IQ2_XXS, IQ2_XS, including the ragged 129-row tile) are
byte-identical to the upstream loop. Prefill 12 keeps the tile bytes,
the mma order and the masked online softmax of the attention kernel
and only changes how the fragments reach the tensor cores; six
synthetic cells (200 / 1,024 / 2,048 tokens, full and window 128) are
byte-identical to the scalar kernel. On the 8K fixture every round-6
cell (all defaults, each kill switch, each repeat) is byte-identical
in logits and greedy IDs to every other and to the round-5 binary.
ncu numbers behind the two rounds: the worklist kernels ran one block
per SM with 61-74% of the issue slots idle and the long scoreboard as
the top stall (memory 31-37%, tensor pipe ~35%); the attention kernel
had the LSU pipe at 54% of peak from scalar fragment loads.
Run Baekpica/K2-Horizon-375B-A23B-Mixed-Quant-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models