GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →

NVIDIA cloud models

18 models tracked via Artificial Analysis. Compare cloud performance, then find local GGUF versions in the GraySoft model catalog.

ModelIntelligenceSpeed (tok/s)
Nemotron 3 Ultra 550B A55B (Reasoning)37.8219.223
NVIDIA Nemotron 3 Super 120B A12B (Reasoning)25.4157.539
Nemotron Cascade 2 30B A3B17.60
Nemotron 3 Nano Omni 30B A3B Reasoning14.9320.368
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)14.2133.94
Llama Nemotron Super 49B v1.5 (Reasoning)12.478.657
Llama 3.3 Nemotron Super 49B v1 (Reasoning)12.20
Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)9.153.866
NVIDIA Nemotron Nano 12B v2 VL (Reasoning)977.003
NVIDIA Nemotron Nano 9B V2 (Reasoning)8.890.739
Llama Nemotron Super 49B v1.5 (Non-reasoning)8.769.926
NVIDIA Nemotron 3 Nano 4B8.70
Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)8.50
Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)8.50
Llama 3.1 Nemotron Instruct 70B7.678.587
NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)7.496.98
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)7.4138.775
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)4.6164.49

Run models locally with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models