← All comparisons

gpt-oss-120b (high) vs Claude 3.5 Sonnet (June '24)

OpenAI vs Anthropic — side-by-side benchmark comparison

gpt-oss-120b (high)Claude 3.5 Sonnet (June '24)
Intelligence Index33.314.2
Coding Index28.626.0
Math Index93.4
Output speed (tok/s)356.80.0
Blended price ($/1M)$0.26$6.56
Time to first token (s)0.51s0.00s
aime9.7%
aime 2593.4%
artificial analysis coding index28.6026.00
artificial analysis intelligence index33.3014.20
artificial analysis math index93.40
gpqa78.2%56.0%
hle18.5%3.7%
ifbench69.0%
lcr50.7%
livecodebench87.8%
math 50069.5%
mmlu pro80.8%75.1%
scicode38.9%31.6%
tau265.8%
terminalbench hard23.5%

Benchmark data from Artificial Analysis.