AaryanK/Muse-Glimmer-30B-GGUF overview
Muse Glimmer 30B GGUF AK line π I built this line solo the calibration, the per tensor allocations, and the eval harness behind every number below. I'm lookinβ¦
Runs locally from ~2.54 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Muse-Glimmer-30B-AK-Q2_K_XL.gguf | GGUF | Q2_K_XL | 11.60 GB | Download |
| Muse-Glimmer-30B-AK-Q3_K_XL.gguf | GGUF | Q3_K_XL | 12.58 GB | Download |
| Muse-Glimmer-30B-AK-Q4_K_M.gguf | GGUF | Q4_K_M | 14.78 GB | Download |
| Muse-Glimmer-30B-AK-Q4_K_XL.gguf | GGUF | Q4_K_XL | 15.14 GB | Download |
| Muse-Glimmer-30B-AK-Q5_K_M.gguf | GGUF | Q5_K_M | 17.87 GB | Download |
| Muse-Glimmer-30B-AK-Q6_K_XL.gguf | GGUF | Q6_K_XL | 24.44 GB | Download |
| Muse-Glimmer-30B-AK-Q8_K_L.gguf | GGUF | Q8_K_L | 30.07 GB | Download |
| Muse-Glimmer-30B-AK-Q8_K_XL.gguf | GGUF | Q8_K_XL | 32.56 GB | Download |
| dflash-Muse-Glimmer-30B-assistant-Q8_0.gguf | GGUF | Q8_0 | 2.54 GB | Download |
| mmproj-Muse-Glimmer-30B-BF16.gguf | GGUF | BF16 | 3.58 GB | Download |
Model Details
| Model ID | AaryanK/Muse-Glimmer-30B-GGUF |
|---|---|
| Author | AaryanK |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | meta-models/Muse-Glimmer-30B |
| Last modified | 2026-08-13T03:38:17.000Z |
Model README
---
base_model: meta-models/Muse-Glimmer-30B
base_model_relation: quantized
license: apache-2.0
library_name: gguf
pipeline_tag: image-text-to-text
tags:
- gguf
- llama.cpp
- quantized
- imatrix
- muse_glimmer
- conversational
---
Muse-Glimmer-30B - GGUF (AK line)
> π I built this line solo - the calibration, the per-tensor allocations, and the eval harness behind
> every number below. I'm looking for internships in AI agent orchestration and model inference.
> If this work looks relevant to your team: linkedin.com/in/theaaryankapoor
**State-of-the-art GGUF quantizations for
meta-models/Muse-Glimmer-30B.** Eight builds
(27.86 B params, 52 dense layers, GQA 32:2), each with a custom per-tensor bit allocation derived for its
size point - plus the stock BF16 vision encoder.
Benchmarked head-to-head against the Unsloth, Meta and bartowski lines, every file scored on the same
rig against the same BF16 reference: **26 wins, 6 statistical ties, 0 losses across 32 paired comparisons
on two evaluation sets.**
One line per publisher. Log y, bits-per-weight on the secondary axis, and the crowded 16 GB class magnified.
Every comparison with its 95 % interval - blue clears zero, grey is a statistical tie, and the right
column carries the held-out C4 verdict. The full numbers are in the table below.
Comparison set: the three widest-distribution GGUF lines for this model, as published 2026-08-12;
the Method section has everything needed to reproduce any number here.
> File naming. Every quant in this line carries the AK- prefix: these are custom per-tensor
> allocations, not llama.cpp's stock recipes, so AK-Q4_K_M and a stock Q4_K_M are different files.
> mmproj keeps its upstream name.
Which file do I want?
| file | size | bpw | mean KLD β | top-1 β | vs closest rival |
|---|---|---|---|---|---|
| AK-Q2_K_XL | 12.45 GB | 3.576 | 0.056036 | 90.88 % | β27 % KLD |
| AK-Q3_K_XL | 13.51 GB | 3.880 | 0.039079 | 92.30 % | β27 % KLD |
| AK-Q4_K_M | 15.86 GB | 4.556 | 0.013897 | 95.38 % | β6 % KLD |
| AK-Q4_K_XL | 16.26 GB | 4.669 | 0.012286 | 95.65 % | β14 % KLD |
| AK-Q5_K_M | 19.19 GB | 5.512 | 0.004974 | 97.26 % | best measured; β13 % vs Meta dynamic |
| AK-Q6_K_XL | 26.24 GB | 7.536 | 0.000876 | 98.82 % | β4 % KLD |
| AK-Q8_K_L | 32.28 GB | 9.272 | 0.000356 | 99.25 % | β21 % KLD, smaller file |
| AK-Q8_K_XL | 34.96 GB | 10.040 | 0.000316 | 99.30 % | most faithful build |
| mmproj BF16 | 3.85 GB | - | - | vision encoder | stock, unquantized |
AK-Q4_K_XL is the strongest file in the crowded 16 GB class - no published quant of this model at any
comparable size comes within 13 % of it, **including Meta's own official quant, which it beats by 13 %
while being half a GB smaller**. The same story repeats at Q5: AK-Q5_K_M beats Meta's official
19.65 GB kquant-dynamic by 9-13 % on both evaluation sets while being 460 MB smaller. At Q8,
AK-Q8_K_L beats Unsloth's build while being smaller.
llama-server -m Muse-Glimmer-30B-AK-Q4_K_XL.gguf \
--mmproj mmproj-Muse-Glimmer-30B-BF16.gguf -c 8192 -ngl 99
Full measurement table
| publisher | file | bytes | bpw | PPL ratio | mean KLD | p99.9 KLD | top-1 | Ξ vs closest rival |
|---|---|---|---|---|---|---|---|---|
| bartowski | Q2_K_L | 12,348,891,936 | 3.547 | 1.113088 | 0.132355 | 4.5183 | 86.289 % | |
| Unsloth | UD-Q2_K_XL | 12,444,212,256 | 3.574 | 1.065168 | 0.077057 | 2.9476 | 89.240 % | |
| AaryanK | AK-Q2_K_XL | 12,451,267,776 | 3.576 | 1.048590 | 0.056036 | 2.4523 | 90.878 % | β27.3 % [β28.3, β26.3] |
| Unsloth | UD-Q3_K_XL | 13,360,983,072 | 3.837 | 1.047022 | 0.053173 | 2.2479 | 91.175 % | |
| AaryanK | AK-Q3_K_XL | 13,509,095,872 | 3.880 | 1.033090 | 0.039079 | 1.5369 | 92.303 % | β26.5 % [β27.8, β25.2] |
| bartowski | Q3_K_M | 13,962,519,328 | 4.010 | 1.032023 | 0.039487 | 1.6390 | 92.303 % | |
| bartowski | IQ4_XS | 15,435,096,096 | 4.433 | 1.010121 | 0.015440 | 0.6246 | 95.128 % | |
| AaryanK | AK-Q4_K_M | 15,864,857,280 | 4.556 | 1.010896 | 0.013897 | 0.5592 | 95.378 % | β5.6 % [β7.6, β3.3] |
| Unsloth | UD-Q4_K_XL | 15,878,222,368 | 4.560 | 1.010630 | 0.014714 | 0.5871 | 95.249 % | |
| AaryanK | AK-Q4_K_XL | 16,255,873,984 | 4.669 | 1.008957 | 0.012286 | 0.5601 | 95.647 % | β14.0 % [β15.4, β12.7] |
| bartowski | Q4_K_S | 16,320,943,136 | 4.687 | 1.010071 | 0.014293 | 0.5860 | 95.319 % | |
| Meta | kquant-17gb | 16,756,681,056 | 4.812 | 1.009871 | 0.014146 | 0.5918 | 95.297 % | |
| AaryanK | AK-Q5_K_M | 19,191,472,832 | 5.512 | 1.003958 | 0.004974 | 0.1922 | 97.256 % | β2.3 % [β4.8, +0.6] tie |
| Unsloth | UD-Q5_K_M | 19,194,274,848 | 5.513 | 1.004517 | 0.005092 | 0.2027 | 97.157 % | |
| Meta | kquant-dynamic | 19,653,957,984 | 5.645 | 1.004101 | 0.005687 | 0.2206 | 96.965 % | |
| AaryanK | AK-Q6_K_XL | 26,238,366,400 | 7.536 | 1.000831 | 0.000876 | 0.0384 | 98.819 % | β3.8 % [β5.8, β1.9] |
| Unsloth | UD-Q6_K_XL | 26,265,362,976 | 7.543 | 1.000885 | 0.000911 | 0.0386 | 98.867 % | |
| AaryanK | AK-Q8_K_L | 32,283,878,048 | 9.272 | 1.000606 | 0.000356 | 0.0148 | 99.248 % | β20.8 % [β23.8, β17.5] |
| Unsloth | UD-Q8_K_XL | 32,300,651,040 | 9.277 | 1.000728 | 0.000450 | 0.0197 | 99.126 % | |
| AaryanK | AK-Q8_K_XL | 34,958,791,360 | 10.040 | 1.000680 | 0.000316 | 0.0140 | 99.301 % | β29.7 % [β32.4, β26.6] |
Intervals are a paired per-token cluster bootstrap over the 60 evaluation chunks. Ξ is against the
closest-sized non-AaryanK file.
Reading the numbers
Two Q8 builds, two jobs. AK-Q8_K_L is the size-class winner - smaller than Unsloth's Q8 and β20.8 %
KLD, confirmed on every slice tested. AK-Q8_K_XL is the maximum-fidelity build: β29.7 % at +8.2 % bytes
(10.04 bpw vs 9.28), for when the last 2.7 GB of VRAM is cheaper than the last drop of divergence.
PPL ratio and KLD can rank differently at Q3-Q4. PPL scores only the probability of the true next
token; KLD scores the whole distribution - so AK-Q3_K_XL and AK-Q4_K_M lead their size peers on KLD
while trailing by ~0.001 on PPL ratio. Both metrics are in the table; the interval column carries the
verdicts.
Does the margin generalise?
The same files re-measured on six evaluation sets, four held out and audited at zero fragment overlap with
any calibration corpus. 11 of 16 held-out margins exceed the same file's wikitext margin, and every
interval in the chart excludes zero - including all four held-out domains for AK-Q8_K_L, where the
size-matched Q8 lead spans β17 % to β24 %.
Tail behaviour
Method
- Reference: our own BF16 GGUF, converted with llama.cpp pinned at
62bf73d2. The conversion was
checked against every competitor's file across 21 load-bearing KVs, so the comparison measures
quantization rather than a conversion delta.
- Eval:
llama-perplexity --kl-divergence, ctx 4096 Γ 60 chunks β 122,820 scored tokens.
ctx 4096 matters for this architecture - it alternates 3Γ sliding-window (2048) with 1Γ
full-attention NoPE layers, and only at ctx β₯ 4096 does every scored token sit beyond the window.
- Statistics: paired per-token cluster bootstrap at the 2047-token chunk width for every interval.
- Confirmation: the Q4 result was re-run under 9 independent calibration draws across three corpus
families on an untouched slice - beneficial in 9/9, no reversals, every interval excluding zero,
against an MDE fixed before any data was collected.
- Long context: re-measured at ctx 8192; the lead holds and slightly grows.
- Capability: a 130-case tool-calling suite scored as paired agreement with BF16 -
AK-Q4_K_M
matches BF16 on 128 of 130 cases with one flip in each direction: statistically
indistinguishable (exact McNemar p = 1.000).
- Ten pre-registered apparatus gates, all passing, including full-vocabulary agreement with HF
transformers and exact greedy generation agreement (235/235 tokens).
- Scope: the text tower is what is measured and quantized;
mmprojships as the stock BF16 encoder.
KLD values are model-local (this head applies logit soft-capping) - compare within this table only.
Held-out sets: GitHub source (numpy/redis/django/sqlite/nlohmann), OASST dialogue, GSM8K + arXiv
abstracts, and Wikipedia in sixteen non-English languages.
---
*Per-tensor bit allocation derived separately at each target bit-width using importance data from a
diverse in-house calibration set. Base model licence and usage policy unchanged from
meta-models/Muse-Glimmer-30B (Apache-2.0).*
Run AaryanK/Muse-Glimmer-30B-GGUF with guIDE
Download guIDE β the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face Β· Compare models