GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF overview

NVIDIA Nemotron 3 Nano 30B A3B Run with https://llama.app bash llama serve hf ggml org/NVIDIA Nemotron 3 Nano 30B A3B GGUF Source models https://huggingface.co…

ggufquantizedtext-generationbase_model:nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16base_model:quantized:nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16license:otherendpoints_compatibleregion:usconversational

Runs locally from ~20.88 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.ggufGGUFBF1658.84 GBDownload
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4_K_M.ggufGGUFQ4_K_M20.88 GBDownload
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8_0.ggufGGUFQ8_031.28 GBDownload

Model Details

Model IDggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF
Authorggml-org
Pipelinetext-generation
Licenseother
Base modelnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Last modified2026-07-30T13:25:38.000Z

Model README

---

license: other

pipeline_tag: text-generation

tags:

  • gguf
  • quantized

base_model:

  • nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

---

NVIDIA-Nemotron-3-Nano-30B-A3B

Run with https://llama.app

llama serve -hf ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF

Source models

  • https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

TODOs

  • add info
  • add MTP

> [!IMPORTANT]

> This model is automatically converted using https://github.com/ggml-org/convert

Run ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models