← All comparisons

Claude 4.1 Opus (Non-reasoning) vs Devstral Small (May '25)

Anthropic vs Mistral — side-by-side benchmark comparison

Claude 4.1 Opus (Non-reasoning)Devstral Small (May '25)
Intelligence Index36.018.0
Coding Index12.2
Math Index
Output speed (tok/s)44.70.0
Blended price ($/1M)$32.81$0.00
Time to first token (s)1.63s0.00s
aime6.7%
aime 25
artificial analysis coding index12.20
artificial analysis intelligence index36.0018.00
artificial analysis math index
gpqa43.4%
hle4.0%
ifbench31.6%
lcr26.7%
livecodebench25.8%
math 50068.4%
mmlu pro63.2%
scicode24.5%
tau238.0%
terminalbench hard6.1%

Benchmark data from Artificial Analysis.