← All comparisons

Claude 3.5 Sonnet (June '24) vs Devstral Small (Jul '25)

Anthropic vs Mistral — side-by-side benchmark comparison

Claude 3.5 Sonnet (June '24)Devstral Small (Jul '25)
Intelligence Index14.215.2
Coding Index26.012.1
Math Index29.3
Output speed (tok/s)0.0183.4
Blended price ($/1M)$6.56$0.15
Time to first token (s)0.00s0.40s
aime9.7%0.3%
aime 2529.3%
artificial analysis coding index26.0012.10
artificial analysis intelligence index14.2015.20
artificial analysis math index29.30
gpqa56.0%41.4%
hle3.7%3.7%
ifbench34.6%
lcr17.0%
livecodebench25.4%
math 50069.5%63.5%
mmlu pro75.1%62.2%
scicode31.6%24.3%
tau228.4%
terminalbench hard6.1%

Benchmark data from Artificial Analysis.