GraySoft
Projects Models About FAQ Contact Download guIDE →

ggml-org/nemotron-3-nano-4b-gguf Q8_0 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

ggml-org/nemotron-3-nano-4b-gguf overview

Recommended way to run this model: Links:

ggufbase_model:nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16base_model:quantized:nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16endpoints_compatibleregion:usconversational
ggml-org/nemotron-3-nano-4b-gguf visual
Downloads
792
Likes
2
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

2 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Nemotron-3-Nano-4B-BF16.gguf GGUF BF16 7.41 GB Download
Nemotron-3-Nano-4B-Q8_0.gguf GGUF 3.94 GB Download

Model Details Live

Model Slug
ggml-org/nemotron-3-nano-4b-gguf
Author
ggml-org
Pipeline Task
Library
Created
2026-03-16
Last Modified
2026-03-16
Gated
No
Private
No
HF SHA
12f6af4d6fcafbfe54f29c4c7e26ccb6f4bea0c2
License
Unknown
Language
Unknown
Base Model
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16"
    ],
    "frontmatter": {
      "base_model": [
        "nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16"
      ]
    },
    "hero_image_url": "",
    "summary": "Recommended way to run this model: ``sh llama-server -hf ggml-org/Nemotron-3-Nano-4B-GGUF `` Links:",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16\n---\n\n# Nemotron-3-Nano-4B GGUF\n\nRecommended way to run this model:\n\n```sh\nllama-server -hf ggml-org/Nemotron-3-Nano-4B-GGUF\n```\n\nLinks:\n\n- [Performance of NVIDIA Nemotron 3 models with llama.cpp](https://github.com/ggml-org/llama.cpp/discussions/20421)",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16",
    "base_model:quantized:nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 2,
  "downloads": 792,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-16T19:46:15.000Z",
  "created_at": "2026-03-16T19:40:25.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69b85ca95bff04a265c3b47a",
  "id": "ggml-org/Nemotron-3-Nano-4B-GGUF",
  "modelId": "ggml-org/Nemotron-3-Nano-4B-GGUF",
  "sha": "12f6af4d6fcafbfe54f29c4c7e26ccb6f4bea0c2",
  "createdAt": "2026-03-16T19:40:25.000Z",
  "lastModified": "2026-03-16T19:46:15.000Z",
  "author": "ggml-org",
  "downloads": 792,
  "likes": 2,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 4
}