devquasar/nvidia.llama-3_1-nemotron-ultra-253b-v1-gguf Q4_K_M GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
devquasar/nvidia.llama-3_1-nemotron-ultra-253b-v1-gguf overview
Big thanks to ymcki for updating the llama.cpp code to support the 'dummy' layers. Use the llama.cpp branch from this PR: https://github.com/ggml-org/llama.cpp/pull/12843 if it hasn't been merged yet. Note the imatrix data used for the IQ quants has been produced from the Q4 quant! !image/png 'Make knowledge free for everyone' Quantized version of: nvidia/Llama-3_1-Nemotron-Ultra-253B-v1
Downloads
1,006
Likes
7
Pipeline
text-generation
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
60 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_M-00001-of-00005.gguf | GGUF | IQ1_M | 13.02 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_M-00002-of-00005.gguf | GGUF | IQ1_M | 12.88 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_M-00003-of-00005.gguf | GGUF | IQ1_M | 12.85 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_M-00004-of-00005.gguf | GGUF | IQ1_M | 12.98 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_M-00005-of-00005.gguf | GGUF | IQ1_M | 3.05 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_S-00001-of-00004.gguf | GGUF | IQ1_S | 12.99 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_S-00002-of-00004.gguf | GGUF | IQ1_S | 13.03 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_S-00003-of-00004.gguf | GGUF | IQ1_S | 12.09 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ1_S-00004-of-00004.gguf | GGUF | IQ1_S | 11.86 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ2_XXS-00001-of-00005.gguf | GGUF | IQ2_XXS | 13.02 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ2_XXS-00002-of-00005.gguf | GGUF | IQ2_XXS | 12.96 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ2_XXS-00003-of-00005.gguf | GGUF | IQ2_XXS | 12.94 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ2_XXS-00004-of-00005.gguf | GGUF | IQ2_XXS | 12.86 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ2_XXS-00005-of-00005.gguf | GGUF | IQ2_XXS | 11.04 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00001-of-00008.gguf | GGUF | IQ3_XXS | 12.84 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00002-of-00008.gguf | GGUF | IQ3_XXS | 12.79 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00003-of-00008.gguf | GGUF | IQ3_XXS | 12.79 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00004-of-00008.gguf | GGUF | IQ3_XXS | 12.94 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00005-of-00008.gguf | GGUF | IQ3_XXS | 11.79 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00006-of-00008.gguf | GGUF | IQ3_XXS | 11.47 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00007-of-00008.gguf | GGUF | IQ3_XXS | 12.74 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.IQ3_XXS-00008-of-00008.gguf | GGUF | IQ3_XXS | 3.56 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00001-of-00007.gguf | GGUF | Q2_K | 12.89 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00002-of-00007.gguf | GGUF | Q2_K | 12.99 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00003-of-00007.gguf | GGUF | Q2_K | 13.03 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00004-of-00007.gguf | GGUF | Q2_K | 12.83 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00005-of-00007.gguf | GGUF | Q2_K | 11.65 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00006-of-00007.gguf | GGUF | Q2_K | 11.92 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q2_K-00007-of-00007.gguf | GGUF | Q2_K | 11.68 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00001-of-00010.gguf | GGUF | Q3_K_M | 12.72 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00002-of-00010.gguf | GGUF | Q3_K_M | 12.63 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00003-of-00010.gguf | GGUF | Q3_K_M | 12.96 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00004-of-00010.gguf | GGUF | Q3_K_M | 12.77 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00005-of-00010.gguf | GGUF | Q3_K_M | 12.87 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00006-of-00010.gguf | GGUF | Q3_K_M | 12.42 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00007-of-00010.gguf | GGUF | Q3_K_M | 11.86 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00008-of-00010.gguf | GGUF | Q3_K_M | 12.00 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00009-of-00010.gguf | GGUF | Q3_K_M | 12.94 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_M-00010-of-00010.gguf | GGUF | Q3_K_M | 357.56 MB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00001-of-00009.gguf | GGUF | Q3_K_S | 12.94 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00002-of-00009.gguf | GGUF | Q3_K_S | 12.75 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00003-of-00009.gguf | GGUF | Q3_K_S | 12.76 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00004-of-00009.gguf | GGUF | Q3_K_S | 12.91 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00005-of-00009.gguf | GGUF | Q3_K_S | 11.78 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00006-of-00009.gguf | GGUF | Q3_K_S | 10.65 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00007-of-00009.gguf | GGUF | Q3_K_S | 12.33 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00008-of-00009.gguf | GGUF | Q3_K_S | 13.00 GB | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q3_K_S-00009-of-00009.gguf | GGUF | Q3_K_S | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00001-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00002-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00003-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00004-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00005-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00006-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00007-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00008-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00009-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00010-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00011-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
| nvidia.Llama-3_1-Nemotron-Ultra-253B-v1.Q4_K_M-00012-of-00012.gguf | GGUF | Q4_K_M | Unknown | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": [
"nvidia/Llama-3_1-Nemotron-Ultra-253B-v1"
],
"pipeline_tag": "text-generation",
"frontmatter": {
"base_model": [
"nvidia/Llama-3_1-Nemotron-Ultra-253B-v1"
],
"pipeline_tag": "text-generation"
},
"hero_image_url": "https://raw.githubusercontent.com/csabakecskemeti/devquasar/main/dq_logo_black-transparent.png",
"summary": "Big thanks to ymcki for updating the llama.cpp code to support the 'dummy' layers. Use the llama.cpp branch from this PR: https://github.com/ggml-org/llama.cpp/pull/12843 if it hasn't been merged yet. Note the imatrix data used for the IQ quants has been produced from the Q4 quant! !image/png 'Make knowledge free for everyone' Quantized version of: nvidia/Llama-3_1-Nemotron-Ultra-253B-v1",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model:\n- nvidia/Llama-3_1-Nemotron-Ultra-253B-v1\npipeline_tag: text-generation\n---\n\nBig thanks to ymcki for updating the llama.cpp code to support the 'dummy' layers.\nUse the llama.cpp branch from this PR: https://github.com/ggml-org/llama.cpp/pull/12843 if it hasn't been merged yet.\n\nNote the imatrix data used for the IQ quants has been produced from the Q4 quant!\n\n\n\n[<img src=\"https://raw.githubusercontent.com/csabakecskemeti/devquasar/main/dq_logo_black-transparent.png\" width=\"200\"/>](https://devquasar.com)\n\n'Make knowledge free for everyone'\n\nQuantized version of: [nvidia/Llama-3_1-Nemotron-Ultra-253B-v1](https://huggingface.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1)\n<a href='https://ko-fi.com/L4L416YX7C' target='_blank'><img height='36' style='border:0px;height:36px;' src='https://storage.ko-fi.com/cdn/kofi6.png?v=6' border='0' alt='Buy Me a Coffee at ko-fi.com' /></a>\n",
"related_quantizations": []
},
"tags": [
"gguf",
"text-generation",
"base_model:nvidia/Llama-3_1-Nemotron-Ultra-253B-v1",
"base_model:quantized:nvidia/Llama-3_1-Nemotron-Ultra-253B-v1",
"endpoints_compatible",
"region:us",
"imatrix",
"conversational"
],
"likes": 7,
"downloads": 1006,
"gated": false,
"private": false,
"last_modified": "2025-04-16T07:05:08.000Z",
"created_at": "2025-04-08T03:39:20.000Z",
"pipeline_tag": "text-generation",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "67f49a682389cdb04cb4dc1d",
"id": "DevQuasar/nvidia.Llama-3_1-Nemotron-Ultra-253B-v1-GGUF",
"modelId": "DevQuasar/nvidia.Llama-3_1-Nemotron-Ultra-253B-v1-GGUF",
"sha": "fa034db97ae9c144e5d5a8fd8fc4fefdd363b8db",
"createdAt": "2025-04-08T03:39:20.000Z",
"lastModified": "2025-04-16T07:05:08.000Z",
"author": "DevQuasar",
"downloads": 1006,
"likes": 7,
"gated": false,
"private": false,
"pipeline_tag": "text-generation",
"library_name": "",
"siblings_count": 63
}