GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

AaryanK/Muse-Glimmer-30B-GGUF overview

Muse Glimmer 30B GGUF AK line πŸ‘‹ I built this line solo the calibration, the per tensor allocations, and the eval harness behind every number below. I'm lookin…

ggufllama.cppquantizedimatrixmuse_glimmerconversationalimage-text-to-textbase_model:meta-models/Muse-Glimmer-30Bbase_model:quantized:meta-models/Muse-Glimmer-30Blicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~2.54 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
15
Pipeline
image-text-to-text
Author

Repository Files & Downloads

10 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Muse-Glimmer-30B-AK-Q2_K_XL.ggufGGUFQ2_K_XL11.60 GBDownload
Muse-Glimmer-30B-AK-Q3_K_XL.ggufGGUFQ3_K_XL12.58 GBDownload
Muse-Glimmer-30B-AK-Q4_K_M.ggufGGUFQ4_K_M14.78 GBDownload
Muse-Glimmer-30B-AK-Q4_K_XL.ggufGGUFQ4_K_XL15.14 GBDownload
Muse-Glimmer-30B-AK-Q5_K_M.ggufGGUFQ5_K_M17.87 GBDownload
Muse-Glimmer-30B-AK-Q6_K_XL.ggufGGUFQ6_K_XL24.44 GBDownload
Muse-Glimmer-30B-AK-Q8_K_L.ggufGGUFQ8_K_L30.07 GBDownload
Muse-Glimmer-30B-AK-Q8_K_XL.ggufGGUFQ8_K_XL32.56 GBDownload
dflash-Muse-Glimmer-30B-assistant-Q8_0.ggufGGUFQ8_02.54 GBDownload
mmproj-Muse-Glimmer-30B-BF16.ggufGGUFBF163.58 GBDownload

Model Details

Model IDAaryanK/Muse-Glimmer-30B-GGUF
AuthorAaryanK
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelmeta-models/Muse-Glimmer-30B
Last modified2026-08-13T03:38:17.000Z

Model README

---

base_model: meta-models/Muse-Glimmer-30B

base_model_relation: quantized

license: apache-2.0

library_name: gguf

pipeline_tag: image-text-to-text

tags:

- gguf

- llama.cpp

- quantized

- imatrix

- muse_glimmer

- conversational

---

Muse-Glimmer-30B - GGUF (AK line)

> πŸ‘‹ I built this line solo - the calibration, the per-tensor allocations, and the eval harness behind

> every number below. I'm looking for internships in AI agent orchestration and model inference.

> If this work looks relevant to your team: linkedin.com/in/theaaryankapoor

**State-of-the-art GGUF quantizations for

meta-models/Muse-Glimmer-30B.** Eight builds

(27.86 B params, 52 dense layers, GQA 32:2), each with a custom per-tensor bit allocation derived for its

size point - plus the stock BF16 vision encoder.

Benchmarked head-to-head against the Unsloth, Meta and bartowski lines, every file scored on the same

rig against the same BF16 reference: **26 wins, 6 statistical ties, 0 losses across 32 paired comparisons

on two evaluation sets.**

!KL divergence vs file size

One line per publisher. Log y, bits-per-weight on the secondary axis, and the crowded 16 GB class magnified.

!Matched pairs

Every comparison with its 95 % interval - blue clears zero, grey is a statistical tie, and the right

column carries the held-out C4 verdict. The full numbers are in the table below.

Comparison set: the three widest-distribution GGUF lines for this model, as published 2026-08-12;

the Method section has everything needed to reproduce any number here.

> File naming. Every quant in this line carries the AK- prefix: these are custom per-tensor

> allocations, not llama.cpp's stock recipes, so AK-Q4_K_M and a stock Q4_K_M are different files.

> mmproj keeps its upstream name.

Which file do I want?

| file | size | bpw | mean KLD ↓ | top-1 ↑ | vs closest rival |

|---|---|---|---|---|---|

| AK-Q2_K_XL | 12.45 GB | 3.576 | 0.056036 | 90.88 % | βˆ’27 % KLD |

| AK-Q3_K_XL | 13.51 GB | 3.880 | 0.039079 | 92.30 % | βˆ’27 % KLD |

| AK-Q4_K_M | 15.86 GB | 4.556 | 0.013897 | 95.38 % | βˆ’6 % KLD |

| AK-Q4_K_XL | 16.26 GB | 4.669 | 0.012286 | 95.65 % | βˆ’14 % KLD |

| AK-Q5_K_M | 19.19 GB | 5.512 | 0.004974 | 97.26 % | best measured; βˆ’13 % vs Meta dynamic |

| AK-Q6_K_XL | 26.24 GB | 7.536 | 0.000876 | 98.82 % | βˆ’4 % KLD |

| AK-Q8_K_L | 32.28 GB | 9.272 | 0.000356 | 99.25 % | βˆ’21 % KLD, smaller file |

| AK-Q8_K_XL | 34.96 GB | 10.040 | 0.000316 | 99.30 % | most faithful build |

| mmproj BF16 | 3.85 GB | - | - | vision encoder | stock, unquantized |

AK-Q4_K_XL is the strongest file in the crowded 16 GB class - no published quant of this model at any

comparable size comes within 13 % of it, **including Meta's own official quant, which it beats by 13 %

while being half a GB smaller**. The same story repeats at Q5: AK-Q5_K_M beats Meta's official

19.65 GB kquant-dynamic by 9-13 % on both evaluation sets while being 460 MB smaller. At Q8,

AK-Q8_K_L beats Unsloth's build while being smaller.

llama-server -m Muse-Glimmer-30B-AK-Q4_K_XL.gguf \
             --mmproj mmproj-Muse-Glimmer-30B-BF16.gguf -c 8192 -ngl 99

Full measurement table

| publisher | file | bytes | bpw | PPL ratio | mean KLD | p99.9 KLD | top-1 | Ξ” vs closest rival |

|---|---|---|---|---|---|---|---|---|

| bartowski | Q2_K_L | 12,348,891,936 | 3.547 | 1.113088 | 0.132355 | 4.5183 | 86.289 % | |

| Unsloth | UD-Q2_K_XL | 12,444,212,256 | 3.574 | 1.065168 | 0.077057 | 2.9476 | 89.240 % | |

| AaryanK | AK-Q2_K_XL | 12,451,267,776 | 3.576 | 1.048590 | 0.056036 | 2.4523 | 90.878 % | βˆ’27.3 % [βˆ’28.3, βˆ’26.3] |

| Unsloth | UD-Q3_K_XL | 13,360,983,072 | 3.837 | 1.047022 | 0.053173 | 2.2479 | 91.175 % | |

| AaryanK | AK-Q3_K_XL | 13,509,095,872 | 3.880 | 1.033090 | 0.039079 | 1.5369 | 92.303 % | βˆ’26.5 % [βˆ’27.8, βˆ’25.2] |

| bartowski | Q3_K_M | 13,962,519,328 | 4.010 | 1.032023 | 0.039487 | 1.6390 | 92.303 % | |

| bartowski | IQ4_XS | 15,435,096,096 | 4.433 | 1.010121 | 0.015440 | 0.6246 | 95.128 % | |

| AaryanK | AK-Q4_K_M | 15,864,857,280 | 4.556 | 1.010896 | 0.013897 | 0.5592 | 95.378 % | βˆ’5.6 % [βˆ’7.6, βˆ’3.3] |

| Unsloth | UD-Q4_K_XL | 15,878,222,368 | 4.560 | 1.010630 | 0.014714 | 0.5871 | 95.249 % | |

| AaryanK | AK-Q4_K_XL | 16,255,873,984 | 4.669 | 1.008957 | 0.012286 | 0.5601 | 95.647 % | βˆ’14.0 % [βˆ’15.4, βˆ’12.7] |

| bartowski | Q4_K_S | 16,320,943,136 | 4.687 | 1.010071 | 0.014293 | 0.5860 | 95.319 % | |

| Meta | kquant-17gb | 16,756,681,056 | 4.812 | 1.009871 | 0.014146 | 0.5918 | 95.297 % | |

| AaryanK | AK-Q5_K_M | 19,191,472,832 | 5.512 | 1.003958 | 0.004974 | 0.1922 | 97.256 % | βˆ’2.3 % [βˆ’4.8, +0.6] tie |

| Unsloth | UD-Q5_K_M | 19,194,274,848 | 5.513 | 1.004517 | 0.005092 | 0.2027 | 97.157 % | |

| Meta | kquant-dynamic | 19,653,957,984 | 5.645 | 1.004101 | 0.005687 | 0.2206 | 96.965 % | |

| AaryanK | AK-Q6_K_XL | 26,238,366,400 | 7.536 | 1.000831 | 0.000876 | 0.0384 | 98.819 % | βˆ’3.8 % [βˆ’5.8, βˆ’1.9] |

| Unsloth | UD-Q6_K_XL | 26,265,362,976 | 7.543 | 1.000885 | 0.000911 | 0.0386 | 98.867 % | |

| AaryanK | AK-Q8_K_L | 32,283,878,048 | 9.272 | 1.000606 | 0.000356 | 0.0148 | 99.248 % | βˆ’20.8 % [βˆ’23.8, βˆ’17.5] |

| Unsloth | UD-Q8_K_XL | 32,300,651,040 | 9.277 | 1.000728 | 0.000450 | 0.0197 | 99.126 % | |

| AaryanK | AK-Q8_K_XL | 34,958,791,360 | 10.040 | 1.000680 | 0.000316 | 0.0140 | 99.301 % | βˆ’29.7 % [βˆ’32.4, βˆ’26.6] |

Intervals are a paired per-token cluster bootstrap over the 60 evaluation chunks. Ξ” is against the

closest-sized non-AaryanK file.

Reading the numbers

Two Q8 builds, two jobs. AK-Q8_K_L is the size-class winner - smaller than Unsloth's Q8 and βˆ’20.8 %

KLD, confirmed on every slice tested. AK-Q8_K_XL is the maximum-fidelity build: βˆ’29.7 % at +8.2 % bytes

(10.04 bpw vs 9.28), for when the last 2.7 GB of VRAM is cheaper than the last drop of divergence.

PPL ratio and KLD can rank differently at Q3-Q4. PPL scores only the probability of the true next

token; KLD scores the whole distribution - so AK-Q3_K_XL and AK-Q4_K_M lead their size peers on KLD

while trailing by ~0.001 on PPL ratio. Both metrics are in the table; the interval column carries the

verdicts.

Does the margin generalise?

!Domain slices

The same files re-measured on six evaluation sets, four held out and audited at zero fragment overlap with

any calibration corpus. 11 of 16 held-out margins exceed the same file's wikitext margin, and every

interval in the chart excludes zero - including all four held-out domains for AK-Q8_K_L, where the

size-matched Q8 lead spans βˆ’17 % to βˆ’24 %.

Tail behaviour

!Tail percentile

Method

  • Reference: our own BF16 GGUF, converted with llama.cpp pinned at 62bf73d2. The conversion was

checked against every competitor's file across 21 load-bearing KVs, so the comparison measures

quantization rather than a conversion delta.

  • Eval: llama-perplexity --kl-divergence, ctx 4096 Γ— 60 chunks β†’ 122,820 scored tokens.

ctx 4096 matters for this architecture - it alternates 3Γ— sliding-window (2048) with 1Γ—

full-attention NoPE layers, and only at ctx β‰₯ 4096 does every scored token sit beyond the window.

  • Statistics: paired per-token cluster bootstrap at the 2047-token chunk width for every interval.
  • Confirmation: the Q4 result was re-run under 9 independent calibration draws across three corpus

families on an untouched slice - beneficial in 9/9, no reversals, every interval excluding zero,

against an MDE fixed before any data was collected.

  • Long context: re-measured at ctx 8192; the lead holds and slightly grows.
  • Capability: a 130-case tool-calling suite scored as paired agreement with BF16 - AK-Q4_K_M

matches BF16 on 128 of 130 cases with one flip in each direction: statistically

indistinguishable (exact McNemar p = 1.000).

  • Ten pre-registered apparatus gates, all passing, including full-vocabulary agreement with HF

transformers and exact greedy generation agreement (235/235 tokens).

  • Scope: the text tower is what is measured and quantized; mmproj ships as the stock BF16 encoder.

KLD values are model-local (this head applies logit soft-capping) - compare within this table only.

Held-out sets: GitHub source (numpy/redis/django/sqlite/nlohmann), OASST dialogue, GSM8K + arXiv

abstracts, and Wikipedia in sixteen non-English languages.

---

*Per-tensor bit allocation derived separately at each target bit-width using importance data from a

diverse in-house calibration set. Base model licence and usage policy unchanged from

meta-models/Muse-Glimmer-30B (Apache-2.0).*

Run AaryanK/Muse-Glimmer-30B-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models