wacomctl672/Nemotron-3-Nano-Omni-30B-Abliterated-MM-GGUF overview
GGUF Quantizations for divinetribe/Nemotron 3 Nano Omni 30B Abliterated MM bf16 https://huggingface.co/divinetribe/Nemotron 3 Nano Omni 30B Abliterated MM bf16…
Runs locally from ~22.83 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | wacomctl672/Nemotron-3-Nano-Omni-30B-Abliterated-MM-GGUF |
|---|---|
| Author | wacomctl672 |
| Pipeline | — |
| License | other |
| Base model | divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 |
| Last modified | 2026-08-31T00:09:21.000Z |
Model README
---
license: other
license_name: nvidia-open-model-agreement
license_link: >-
https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
base_model:
- divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16
---
GGUF Quantizations for divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16
tested multimodality with q8_0 mmproj from ggml-org
I will make higher/lower quants if requested, just open an discussion. Or do it yourself with the patch mentioned on the original model card.
should run on anything that can run llama.cpp. tested under windows with cuda
Q4KM is 6.21 BPW because of fallback quantization.
Q8_0 is the usual 8.51 BPW
Disclaimer & Notice
- Uncensored / Abliterated: This model has had safety mitigations and refusal vectors removed. It may generate offensive, harmful, or unfiltered content.
- Use at Your Own Risk: This model is provided "as is" without warranties of any kind. You are solely responsible for how you use the model and any outputs generated.
- No Liability: The uploader is solely providing a converted/quantized GGUF format for research and developer testing and assumes no liability for any damages, legal issues, or misuse.
Run wacomctl672/Nemotron-3-Nano-Omni-30B-Abliterated-MM-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models