Abiray/Nemotron-3-Embed-8B-GGUF overview
๐ Nemotron 3 Embed 8B GGUF Quantizations Model Architecture https://img.shields.io/badge/Architecture Transformer blue?style=flat square Context Length https:โฆ
Runs locally from ~3.74 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Nemotron-3-Embed-8B-Q3_K_M.gguf | GGUF | Q3_K_M | 3.74 GB | Download |
| Nemotron-3-Embed-8B-Q4_K_M.gguf | GGUF | Q4_K_M | 4.56 GB | Download |
| Nemotron-3-Embed-8B-Q4_K_S.gguf | GGUF | Q4_K_S | 4.33 GB | Download |
| Nemotron-3-Embed-8B-Q5_K_M.gguf | GGUF | Q5_K_M | 5.30 GB | Download |
| Nemotron-3-Embed-8B-Q5_K_S.gguf | GGUF | Q5_K_S | 5.17 GB | Download |
| Nemotron-3-Embed-8B-Q6_K.gguf | GGUF | Q6_K | 6.08 GB | Download |
| Nemotron-3-Embed-8B-Q8_0.gguf | GGUF | Q8_0 | 7.88 GB | Download |
Model Details
Model README
---
language:
- en
- fr
- de
- es
- it
- pt
- ru
- zh
- ja
license: openmdw-1.1
tags:
- gguf
- embeddings
- retrieval
- semantic-search
- nvidia
- nemotron
- rag
base_model: nvidia/Nemotron-3-Embed-8B-BF16
---
๐ Nemotron-3-Embed-8B (GGUF Quantizations)



Welcome to the GGUF repository for NVIDIA's Nemotron-3-Embed-8B-BF16.
This model is a state-of-the-art, 8-billion parameter multilingual text embedding model optimized for Retrieval-Augmented Generation (RAG), semantic search, and cross-lingual retrieval workflows. It achieved #1 on the multilingual RTEB leaderboard (as of July 2026).
By converting the model to GGUF (GPT-Generated Unified Format), you can run enterprise-grade retrieval locally on consumer hardware (CPUs and GPUs) using tools like llama.cpp, Ollama, or LM Studio.
---
๐ Repository Files & Quantization Options
Below is the directory structure of the available .gguf files in this repository. Choose the quantization level that best fits your VRAM/RAM constraints!
๐ nemotron-3-embed-8b-gguf/
โ
โโโ ๐ README.md
โโโ ๐ config.json
โ
โโโ ๐ฆ Nemotron-3-Embed-8B-Q8_0.gguf (8.46 GB) ๐ข Near zero quality loss
โโโ ๐ฆ Nemotron-3-Embed-8B-Q6_K.gguf (6.53 GB) ๐ข Extremely low quality loss
โโโ ๐ฆ Nemotron-3-Embed-8B-Q5_K_M.gguf (5.69 GB) ๐ก Very low quality loss
โโโ ๐ฆ Nemotron-3-Embed-8B-Q5_K_S.gguf (5.55 GB) ๐ก Very low quality loss
โโโ ๐ฆ Nemotron-3-Embed-8B-Q4_K_M.gguf (4.90 GB) โญ RECOMMENDED - Great balance
โโโ ๐ฆ Nemotron-3-Embed-8B-Q4_K_S.gguf (4.65 GB) ๐ Moderate quality loss
โโโ ๐ฆ Nemotron-3-Embed-8B-Q3_K_M.gguf (4.01 GB) ๐ด High quality loss (Memory constrained only)Run Abiray/Nemotron-3-Embed-8B-GGUF with guIDE
Download guIDE โ the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face ยท Compare models