GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

Abiray/Nemotron-3-Embed-8B-GGUF overview

๐Ÿš€ Nemotron 3 Embed 8B GGUF Quantizations Model Architecture https://img.shields.io/badge/Architecture Transformer blue?style=flat square Context Length https:โ€ฆ

ggufembeddingsretrievalsemantic-searchnvidianemotronragenfrdeesitptruzhjabase_model:nvidia/Nemotron-3-Embed-8B-BF16base_model:quantized:nvidia/Nemotron-3-Embed-8B-BF16license:openmdw-1.1endpoints_compatibleregion:usimatrixfeature-extraction

Runs locally from ~3.74 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
2
Pipeline
โ€”
Author

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Nemotron-3-Embed-8B-Q3_K_M.ggufGGUFQ3_K_M3.74 GBDownload
Nemotron-3-Embed-8B-Q4_K_M.ggufGGUFQ4_K_M4.56 GBDownload
Nemotron-3-Embed-8B-Q4_K_S.ggufGGUFQ4_K_S4.33 GBDownload
Nemotron-3-Embed-8B-Q5_K_M.ggufGGUFQ5_K_M5.30 GBDownload
Nemotron-3-Embed-8B-Q5_K_S.ggufGGUFQ5_K_S5.17 GBDownload
Nemotron-3-Embed-8B-Q6_K.ggufGGUFQ6_K6.08 GBDownload
Nemotron-3-Embed-8B-Q8_0.ggufGGUFQ8_07.88 GBDownload

Model Details

Model IDAbiray/Nemotron-3-Embed-8B-GGUF
AuthorAbiray
Pipelineโ€”
Licenseopenmdw-1.1
Base modelnvidia/Nemotron-3-Embed-8B-BF16
Last modified2026-07-21T07:30:06.000Z

Model README

---

language:

  • en
  • fr
  • de
  • es
  • it
  • pt
  • ru
  • zh
  • ja

license: openmdw-1.1

tags:

  • gguf
  • embeddings
  • retrieval
  • semantic-search
  • nvidia
  • nemotron
  • rag

base_model: nvidia/Nemotron-3-Embed-8B-BF16

---

๐Ÿš€ Nemotron-3-Embed-8B (GGUF Quantizations)

![Model Architecture](#)

![Context Length](#)

![RTEB Rank](#)

![License](#)

Welcome to the GGUF repository for NVIDIA's Nemotron-3-Embed-8B-BF16.

This model is a state-of-the-art, 8-billion parameter multilingual text embedding model optimized for Retrieval-Augmented Generation (RAG), semantic search, and cross-lingual retrieval workflows. It achieved #1 on the multilingual RTEB leaderboard (as of July 2026).

By converting the model to GGUF (GPT-Generated Unified Format), you can run enterprise-grade retrieval locally on consumer hardware (CPUs and GPUs) using tools like llama.cpp, Ollama, or LM Studio.

---

๐Ÿ“‚ Repository Files & Quantization Options

Below is the directory structure of the available .gguf files in this repository. Choose the quantization level that best fits your VRAM/RAM constraints!

๐Ÿ“ nemotron-3-embed-8b-gguf/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“„ README.md
โ”œโ”€โ”€ ๐Ÿ“„ config.json
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q8_0.gguf    (8.46 GB) ๐ŸŸข Near zero quality loss
โ”œโ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q6_K.gguf    (6.53 GB) ๐ŸŸข Extremely low quality loss
โ”œโ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q5_K_M.gguf  (5.69 GB) ๐ŸŸก Very low quality loss
โ”œโ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q5_K_S.gguf  (5.55 GB) ๐ŸŸก Very low quality loss
โ”œโ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q4_K_M.gguf  (4.90 GB) โญ RECOMMENDED - Great balance
โ”œโ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q4_K_S.gguf  (4.65 GB) ๐ŸŸ  Moderate quality loss
โ””โ”€โ”€ ๐Ÿ“ฆ Nemotron-3-Embed-8B-Q3_K_M.gguf  (4.01 GB) ๐Ÿ”ด High quality loss (Memory constrained only)

Run Abiray/Nemotron-3-Embed-8B-GGUF with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models