← All comparisons

gpt-oss-120b (high) vs Claude 3.5 Sonnet (Oct '24)

OpenAI vs Anthropic — side-by-side benchmark comparison

gpt-oss-120b (high)Claude 3.5 Sonnet (Oct '24)
Intelligence Index33.315.9
Coding Index28.630.2
Math Index93.4
Output speed (tok/s)356.80.0
Blended price ($/1M)$0.26$6.56
Time to first token (s)0.51s0.00s
aime15.7%
aime 2593.4%
artificial analysis coding index28.6030.20
artificial analysis intelligence index33.3015.90
artificial analysis math index93.40
gpqa78.2%59.9%
hle18.5%3.9%
ifbench69.0%
lcr50.7%
livecodebench87.8%38.1%
math 50077.1%
mmlu pro80.8%77.2%
scicode38.9%36.6%
tau265.8%
terminalbench hard23.5%

Benchmark data from Artificial Analysis.