← All comparisons

gpt-oss-120b (low) vs Claude 3.5 Sonnet (June '24)

OpenAI vs Anthropic — side-by-side benchmark comparison

gpt-oss-120b (low)Claude 3.5 Sonnet (June '24)
Intelligence Index24.514.2
Coding Index15.526.0
Math Index66.7
Output speed (tok/s)370.00.0
Blended price ($/1M)$0.26$6.56
Time to first token (s)0.49s0.00s
aime9.7%
aime 2566.7%
artificial analysis coding index15.5026.00
artificial analysis intelligence index24.5014.20
artificial analysis math index66.70
gpqa67.2%56.0%
hle5.2%3.7%
ifbench58.3%
lcr43.7%
livecodebench70.7%
math 50069.5%
mmlu pro77.5%75.1%
scicode36.0%31.6%
tau245.0%
terminalbench hard5.3%

Benchmark data from Artificial Analysis.