← All comparisons

DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) vs Claude 4.1 Opus (Reasoning)

Nous Research vs Anthropic — side-by-side benchmark comparison

DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)Claude 4.1 Opus (Reasoning)
Intelligence Index7.642.0
Coding Index36.5
Math Index80.3
Output speed (tok/s)0.044.5
Blended price ($/1M)$0.00$32.81
Time to first token (s)0.00s8.55s
aime0.0%
aime 2580.3%
artificial analysis coding index36.50
artificial analysis intelligence index7.6042.00
artificial analysis math index80.30
gpqa27.0%80.9%
hle4.3%11.9%
ifbench55.4%
lcr66.3%
livecodebench8.5%65.4%
math 50021.8%
mmlu pro36.5%88.0%
scicode9.1%40.9%
tau271.4%
terminalbench hard34.3%

Benchmark data from Artificial Analysis.