← All comparisons

Claude 4.1 Opus (Reasoning) vs DeepSeek R1 Distill Llama 8B

Anthropic vs DeepSeek — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)DeepSeek R1 Distill Llama 8B
Intelligence Index42.012.1
Coding Index36.5
Math Index80.341.3
Output speed (tok/s)44.50.0
Blended price ($/1M)$32.81$0.00
Time to first token (s)8.55s0.00s
aime33.3%
aime 2580.3%41.3%
artificial analysis coding index36.50
artificial analysis intelligence index42.0012.10
artificial analysis math index80.3041.30
gpqa80.9%30.2%
hle11.9%4.2%
ifbench55.4%17.6%
lcr66.3%0.0%
livecodebench65.4%23.3%
math 50085.3%
mmlu pro88.0%54.3%
scicode40.9%11.9%
tau271.4%
terminalbench hard34.3%

Benchmark data from Artificial Analysis.