← All comparisons

gpt-oss-120b (high) vs Claude 4.1 Opus (Reasoning)

OpenAI vs Anthropic — side-by-side benchmark comparison

gpt-oss-120b (high)Claude 4.1 Opus (Reasoning)
Intelligence Index33.342.0
Coding Index28.636.5
Math Index93.480.3
Output speed (tok/s)356.844.5
Blended price ($/1M)$0.26$32.81
Time to first token (s)0.51s8.55s
aime
aime 2593.4%80.3%
artificial analysis coding index28.6036.50
artificial analysis intelligence index33.3042.00
artificial analysis math index93.4080.30
gpqa78.2%80.9%
hle18.5%11.9%
ifbench69.0%55.4%
lcr50.7%66.3%
livecodebench87.8%65.4%
math 500
mmlu pro80.8%88.0%
scicode38.9%40.9%
tau265.8%71.4%
terminalbench hard23.5%34.3%

Benchmark data from Artificial Analysis.