Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF overview
NVIDIA Nemotron 3.5 Lightning 30B A3B — APEX GGUF Imatrix guided, measured allocation APEX quantizations of nvidia/NVIDIA Nemotron 3.5 Lightning 30B A3B https:…
Runs locally from ~13.66 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF |
|---|---|
| Author | Myric |
| Pipeline | text-generation |
| License | other |
| Base model | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 |
| Last modified | 2026-08-19T13:05:57.000Z |
Model README
---
license: other
license_name: openmdw-1.1
license_link: https://openmdw.ai/license/1-1/
base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- moe
- apex
- quantized
- imatrix
- nemotron
- mamba
- hybrid
- mtp
- speculative-decoding
- llama.cpp
---
NVIDIA Nemotron 3.5 Lightning 30B-A3B — APEX GGUF
Imatrix-guided, measured-allocation APEX quantizations of
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B
— a Nemotron-H hybrid: 52 layers of which only 6 are full attention, the rest split between
Mamba2 SSM and MoE (128 routed experts + shared expert, ~3B active of 31.6B).
Three things distinguish these from the other GGUFs of this model:
- The MTP head is included. The checkpoint ships a multi-token-prediction head; most quants
drop it, which makes speculative decoding impossible. These keep it (53 blocks, not 52).
- Built with an importance matrix, generated on this architecture specifically. The others
are stock conversions.
- Per-tensor bit allocation is measured, not assumed — each tensor was probed for how much
output KL it actually costs at each candidate width, and the budget spent accordingly.
Files
| tier | size | bpw | wikitext PPL | vs bf16 | for | status |
|---|---:|---:|---:|---:|---|---|
| mini | 13.66 GiB | 3.56 | 7.843 | +11.7% | 16 GB card with the full 128k context | available |
| compact | 16.91 GiB | 4.42 | 7.252 | +3.3% | the value pick — 5 GB smaller for 2.8% | uploading |
| i-quality | 21.53 GiB | 5.62 | 7.053 | +0.46% | best quality; 24 GB card or unified memory | uploading |
| bf16 reference | 61.32 GiB | 16.0 | 7.021 | — | (not hosted — measured as the baseline) | — |
mini is up now; compact and i-quality are still being uploaded. All three are built,
measured and gated — the numbers above are from the finished files — but only what the file
listing shows is downloadable yet.
nemotron-lightning.imatrix (56 MB) is included so you can build your own tiers.
Reproducing the perplexity numbers
llama-perplexity -m <tier>.gguf -f wiki.test.raw
wiki.test.raw is the unmodified WikiText-2 raw test split — 1,292,013 bytes, 241,211
words, the file llama.cpp's own perplexity documentation uses, so these numbers are directly
comparable to anyone else's. From
wikitext-2-raw-v1, which shares its
test split with WikiText-103. Identical settings across all four rows above; default context.
The calibration corpus is not WikiText. The imatrix was built on general prose and
scientific text, deliberately, so the reported perplexity is measured on data the quantization
never saw. Calibrating on WikiText train and then scoring on WikiText test flatters the
result — the splits come from the same distribution — and it is an easy mistake to make, since
the obvious calibration file to reach for is often exactly that.
Why mini fits a 16 GB card when a 30B usually doesn't
This model spends 6 KiB per token of KV cache — only 6 of 52 layers are attention, and the
Mamba state is constant-size regardless of context length. So the full 128k context costs
0.75 GiB:
13.66 GiB weights + 0.75 GiB KV @ 128k = 14.41 GiB
For comparison, a conventional 35B-class MoE at ~82 KiB/token would need **10.5 GiB for that
same context** — the cache alone would not fit the card, let alone the weights. That pairing is
the reason this model is interesting at this size point.
Quality gate
Every tier published here passed all three checks. Nothing is uploaded that did not.
| tier | PPL ratio (bar: ≤1.50) | coherent generation | chained tool-calling |
|---|---:|---|---|
| i-quality | 1.00 | pass | 5/5 |
| compact | 1.03 | pass | 5/5 |
| mini | 1.12 | pass | 5/5 |
Tool-calling is a two-turn dependent chain with a distractor tool that must not be called —
a single trivial call is too easy to discriminate between tiers.
Usage
llama-server -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-mini.gguf \
--ctx-size 131072 -fa on --jinja
The chat template is embedded, so --jinja is enough. Always pass --ctx-size — it
otherwise defaults to the model's trained context.
To use the MTP head for speculative decoding:
llama-server -m ...-APEX-i-quality.gguf --ctx-size 131072 -fa on --jinja \
--spec-type draft-mtp --spec-draft-n-max 2
Draft depth is model-specific and does not transfer between models — measure it at your own
sampling settings and context length rather than copying a number from elsewhere.
This model emits explicit reasoning traces before its answer. Budget output tokens
accordingly; a short --n-predict will truncate mid-thought.
Needs a llama.cpp new enough to load the MTP head — if you see a tensor-count mismatch on load,
update.
Notes on the architecture
The row dimensions are 2688 (hidden) and 1856 (expert-down), neither divisible by 256. Every
k-quant and IQ type requires a 256-wide superblock, so on this model they are all illegal and
llama-quantize silently substitutes other types. A stock Q4_K_M of this model measures
6.21 bpw against a nominal 4.85, and contains ~1% actual Q4_K. That is why these tiers use
block-32 and block-64 types throughout, chosen deliberately rather than arrived at by fallback.
Attribution
- Base model: NVIDIA — NVIDIA-Nemotron-3.5-Lightning-30B-A3B
- APEX recipe & toolkit: LocalAI — localai-org/apex-quant
- Quantization engine: llama.cpp (ggml-org)
Calibration: a general prose/scientific corpus (no code), matching the other APEX quants in this
collection. Unofficial community quantization; not affiliated with or endorsed by NVIDIA.
Run Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models