← All comparisons

Claude 4.1 Opus (Non-reasoning) vs Devstral Small (Jul '25)

Anthropic vs Mistral — side-by-side benchmark comparison

Claude 4.1 Opus (Non-reasoning)Devstral Small (Jul '25)
Intelligence Index36.015.2
Coding Index12.1
Math Index29.3
Output speed (tok/s)44.7183.4
Blended price ($/1M)$32.81$0.15
Time to first token (s)1.63s0.40s
aime0.3%
aime 2529.3%
artificial analysis coding index12.10
artificial analysis intelligence index36.0015.20
artificial analysis math index29.30
gpqa41.4%
hle3.7%
ifbench34.6%
lcr17.0%
livecodebench25.4%
math 50063.5%
mmlu pro62.2%
scicode24.3%
tau228.4%
terminalbench hard6.1%

Benchmark data from Artificial Analysis.