← All comparisons

Claude 4.1 Opus (Reasoning) vs Devstral Small (May '25)

Anthropic vs Mistral — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)Devstral Small (May '25)
Intelligence Index42.018.0
Coding Index36.512.2
Math Index80.3
Output speed (tok/s)44.50.0
Blended price ($/1M)$32.81$0.00
Time to first token (s)8.55s0.00s
aime6.7%
aime 2580.3%
artificial analysis coding index36.5012.20
artificial analysis intelligence index42.0018.00
artificial analysis math index80.30
gpqa80.9%43.4%
hle11.9%4.0%
ifbench55.4%31.6%
lcr66.3%26.7%
livecodebench65.4%25.8%
math 50068.4%
mmlu pro88.0%63.2%
scicode40.9%24.5%
tau271.4%38.0%
terminalbench hard34.3%6.1%

Benchmark data from Artificial Analysis.