← All comparisons

Claude 3.5 Sonnet (Oct '24) vs Devstral Small (Jul '25)

Anthropic vs Mistral — side-by-side benchmark comparison

Claude 3.5 Sonnet (Oct '24)Devstral Small (Jul '25)
Intelligence Index15.915.2
Coding Index30.212.1
Math Index29.3
Output speed (tok/s)0.0183.4
Blended price ($/1M)$6.56$0.15
Time to first token (s)0.00s0.40s
aime15.7%0.3%
aime 2529.3%
artificial analysis coding index30.2012.10
artificial analysis intelligence index15.9015.20
artificial analysis math index29.30
gpqa59.9%41.4%
hle3.9%3.7%
ifbench34.6%
lcr17.0%
livecodebench38.1%25.4%
math 50077.1%63.5%
mmlu pro77.2%62.2%
scicode36.6%24.3%
tau228.4%
terminalbench hard6.1%

Benchmark data from Artificial Analysis.