GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF overview

NVIDIA Nemotron 3.5 Lightning 30B A3B — APEX GGUF Imatrix guided, measured allocation APEX quantizations of nvidia/NVIDIA Nemotron 3.5 Lightning 30B A3B https:…

ggufmoeapexquantizedimatrixnemotronmambahybridmtpspeculative-decodingllama.cpptext-generationbase_model:nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16base_model:quantized:nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16license:otherendpoints_compatibleregion:usconversational

Runs locally from ~13.66 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
602
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-compact.ggufGGUFGGUF16.92 GBDownload
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-i-quality.ggufGGUFGGUF21.53 GBDownload
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-mini.ggufGGUFGGUF13.66 GBDownload

Model Details

Model IDMyric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
AuthorMyric
Pipelinetext-generation
Licenseother
Base modelnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Last modified2026-08-19T13:05:57.000Z

Model README

---

license: other

license_name: openmdw-1.1

license_link: https://openmdw.ai/license/1-1/

base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

base_model_relation: quantized

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- moe

- apex

- quantized

- imatrix

- nemotron

- mamba

- hybrid

- mtp

- speculative-decoding

- llama.cpp

---

NVIDIA Nemotron 3.5 Lightning 30B-A3B — APEX GGUF

Imatrix-guided, measured-allocation APEX quantizations of

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B

— a Nemotron-H hybrid: 52 layers of which only 6 are full attention, the rest split between

Mamba2 SSM and MoE (128 routed experts + shared expert, ~3B active of 31.6B).

Three things distinguish these from the other GGUFs of this model:

  • The MTP head is included. The checkpoint ships a multi-token-prediction head; most quants

drop it, which makes speculative decoding impossible. These keep it (53 blocks, not 52).

  • Built with an importance matrix, generated on this architecture specifically. The others

are stock conversions.

  • Per-tensor bit allocation is measured, not assumed — each tensor was probed for how much

output KL it actually costs at each candidate width, and the budget spent accordingly.

Files

| tier | size | bpw | wikitext PPL | vs bf16 | for | status |

|---|---:|---:|---:|---:|---|---|

| mini | 13.66 GiB | 3.56 | 7.843 | +11.7% | 16 GB card with the full 128k context | available |

| compact | 16.91 GiB | 4.42 | 7.252 | +3.3% | the value pick — 5 GB smaller for 2.8% | uploading |

| i-quality | 21.53 GiB | 5.62 | 7.053 | +0.46% | best quality; 24 GB card or unified memory | uploading |

| bf16 reference | 61.32 GiB | 16.0 | 7.021 | — | (not hosted — measured as the baseline) | — |

mini is up now; compact and i-quality are still being uploaded. All three are built,

measured and gated — the numbers above are from the finished files — but only what the file

listing shows is downloadable yet.

nemotron-lightning.imatrix (56 MB) is included so you can build your own tiers.

Reproducing the perplexity numbers

llama-perplexity -m <tier>.gguf -f wiki.test.raw

wiki.test.raw is the unmodified WikiText-2 raw test split — 1,292,013 bytes, 241,211

words, the file llama.cpp's own perplexity documentation uses, so these numbers are directly

comparable to anyone else's. From

wikitext-2-raw-v1, which shares its

test split with WikiText-103. Identical settings across all four rows above; default context.

The calibration corpus is not WikiText. The imatrix was built on general prose and

scientific text, deliberately, so the reported perplexity is measured on data the quantization

never saw. Calibrating on WikiText train and then scoring on WikiText test flatters the

result — the splits come from the same distribution — and it is an easy mistake to make, since

the obvious calibration file to reach for is often exactly that.

Why mini fits a 16 GB card when a 30B usually doesn't

This model spends 6 KiB per token of KV cache — only 6 of 52 layers are attention, and the

Mamba state is constant-size regardless of context length. So the full 128k context costs

0.75 GiB:

13.66 GiB weights + 0.75 GiB KV @ 128k = 14.41 GiB

For comparison, a conventional 35B-class MoE at ~82 KiB/token would need **10.5 GiB for that

same context** — the cache alone would not fit the card, let alone the weights. That pairing is

the reason this model is interesting at this size point.

Quality gate

Every tier published here passed all three checks. Nothing is uploaded that did not.

| tier | PPL ratio (bar: ≤1.50) | coherent generation | chained tool-calling |

|---|---:|---|---|

| i-quality | 1.00 | pass | 5/5 |

| compact | 1.03 | pass | 5/5 |

| mini | 1.12 | pass | 5/5 |

Tool-calling is a two-turn dependent chain with a distractor tool that must not be called —

a single trivial call is too easy to discriminate between tiers.

Usage

llama-server -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-mini.gguf \
  --ctx-size 131072 -fa on --jinja

The chat template is embedded, so --jinja is enough. Always pass --ctx-size — it

otherwise defaults to the model's trained context.

To use the MTP head for speculative decoding:

llama-server -m ...-APEX-i-quality.gguf --ctx-size 131072 -fa on --jinja \
  --spec-type draft-mtp --spec-draft-n-max 2

Draft depth is model-specific and does not transfer between models — measure it at your own

sampling settings and context length rather than copying a number from elsewhere.

This model emits explicit reasoning traces before its answer. Budget output tokens

accordingly; a short --n-predict will truncate mid-thought.

Needs a llama.cpp new enough to load the MTP head — if you see a tensor-count mismatch on load,

update.

Notes on the architecture

The row dimensions are 2688 (hidden) and 1856 (expert-down), neither divisible by 256. Every

k-quant and IQ type requires a 256-wide superblock, so on this model they are all illegal and

llama-quantize silently substitutes other types. A stock Q4_K_M of this model measures

6.21 bpw against a nominal 4.85, and contains ~1% actual Q4_K. That is why these tiers use

block-32 and block-64 types throughout, chosen deliberately rather than arrived at by fallback.

Attribution

Calibration: a general prose/scientific corpus (no code), matching the other APEX quants in this

collection. Unofficial community quantization; not affiliated with or endorsed by NVIDIA.

Run Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models