GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →

NVIDIA cloud models

19 models tracked via Artificial Analysis. Compare cloud performance, then find local GGUF versions in the GraySoft model catalog.

ModelIntelligenceSpeed (tok/s)
Nemotron 3 Ultra 550B A55B (Reasoning)29.6156.367
Nemotron 3 Super 120B A12B (Reasoning)18.6141.927
Nemotron 3.5 Lightning16.4295.608
Nemotron Cascade 2 30B A3B11.50
Nemotron 3 Nano Omni 30B A3B Reasoning9321.434
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)8.6142.767
Llama Nemotron Super 49B v1.5 (Reasoning)6.618.646
Llama 3.3 Nemotron Super 49B v1 (Reasoning)6.40
Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)3.452.821
NVIDIA Nemotron Nano 12B v2 VL (Reasoning)3.325.47
NVIDIA Nemotron Nano 9B V2 (Reasoning)3.294.506
NVIDIA Nemotron 3 Nano 4B3.10
Llama Nemotron Super 49B v1.5 (Non-reasoning)323.094
Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)2.90
Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)2.80
Llama 3.1 Nemotron Instruct 70B2.119.298
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)1.8136.939
NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)1.8165.785
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)172.097

Run models locally with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models