GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Brunobkr/2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.gguf overview

<div align="center" <img src="./capa.jpeg" alt="2BKR IQ4 NL hybrid NVIDIA Nemotron 3.5 Lightning 30B A3B" width="100%" style="border radius: 12px; box shadow: …

ggufquantizedtext-generationbase_model:nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16base_model:quantized:nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16license:otherendpoints_compatibleregion:usconversational

Runs locally from ~16.97 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.ggufGGUFIQ4_NL_HYBRID_NVIDIA16.97 GBDownload

Model Details

Model IDBrunobkr/2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.gguf
AuthorBrunobkr
Pipelinetext-generation
Licenseother
Base modelnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Last modified2026-09-03T22:08:42.000Z

Model README

---

license: other

pipeline_tag: text-generation

tags:

  • gguf
  • quantized

base_model:

  • nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

---

<div align="center">

<img src="./capa.jpeg" alt="2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B" width="100%" style="border-radius: 12px; box-shadow: 0 10px 40px rgba(0,0,0,0.6);" />

⚡ 2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B

</div>

Modelo Quantizado

  • Arquivo: 2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.gguf

Comando sugerido (llama-server)

"/home/userk21/llama_server_VULLKAN_MIXED/build/bin/llama-server" \
  -m "/home/userk21/Área de trabalho/userk21/GGUFS/2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.gguf" \
  -ngl 16 \
  -c 50000 \
  -ctk q8_0 \
  -ctv q8_0 \
  -t 4 \
  -tb 4 \
  -b 2048 \
  -ub 1024 \
  -fa on \
  --cpu-strict 1 \
  --parallel 1 \
  --agent \
  --tools all \
  --reasoning auto \
  --kv-unified \
  --load-mode mmap \
  --cors-origins "*" \
  --webui-mcp-proxy \
  --threads-http -1 \
  --port 5173 \
  --host 127.0.0.1

Source models

  • https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

<!--

  • https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
  • https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash
  • https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark

-->

TODOs

  • add info

> [!IMPORTANT]

> This model is automatically converted using https://github.com/ggml-org/convert

Run Brunobkr/2BKR_IQ4_NL_hybrid_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models