← All comparisons

Claude 4.1 Opus (Reasoning) vs Devstral Small (Jul '25)

Anthropic vs Mistral — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)Devstral Small (Jul '25)
Intelligence Index42.015.2
Coding Index36.512.1
Math Index80.329.3
Output speed (tok/s)44.5183.4
Blended price ($/1M)$32.81$0.15
Time to first token (s)8.55s0.40s
aime0.3%
aime 2580.3%29.3%
artificial analysis coding index36.5012.10
artificial analysis intelligence index42.0015.20
artificial analysis math index80.3029.30
gpqa80.9%41.4%
hle11.9%3.7%
ifbench55.4%34.6%
lcr66.3%17.0%
livecodebench65.4%25.4%
math 50063.5%
mmlu pro88.0%62.2%
scicode40.9%24.3%
tau271.4%28.4%
terminalbench hard34.3%6.1%

Benchmark data from Artificial Analysis.