← All comparisons

Claude 3.7 Sonnet (Reasoning) vs OLMo 2 32B

Anthropic vs Allen Institute for AI — side-by-side benchmark comparison

Claude 3.7 Sonnet (Reasoning)OLMo 2 32B
Intelligence Index34.710.6
Coding Index27.62.7
Math Index56.33.3
Output speed (tok/s)0.00.0
Blended price ($/1M)$0.00$0.00
Time to first token (s)0.00s0.00s
aime48.7%
aime 2556.3%3.3%
artificial analysis coding index27.602.70
artificial analysis intelligence index34.7010.60
artificial analysis math index56.303.30
gpqa77.2%32.8%
hle10.3%3.7%
ifbench48.3%38.1%
lcr60.7%0.0%
livecodebench47.3%6.8%
math 50094.7%
mmlu pro83.7%51.1%
scicode40.3%8.0%
tau254.7%0.0%
terminalbench hard21.2%0.0%

Benchmark data from Artificial Analysis.