← All comparisons

gpt-oss-120b (high) vs Claude 4.1 Opus (Non-reasoning)

OpenAI vs Anthropic — side-by-side benchmark comparison

gpt-oss-120b (high)Claude 4.1 Opus (Non-reasoning)
Intelligence Index33.336.0
Coding Index28.6
Math Index93.4
Output speed (tok/s)356.844.7
Blended price ($/1M)$0.26$32.81
Time to first token (s)0.51s1.63s
aime
aime 2593.4%
artificial analysis coding index28.60
artificial analysis intelligence index33.3036.00
artificial analysis math index93.40
gpqa78.2%
hle18.5%
ifbench69.0%
lcr50.7%
livecodebench87.8%
math 500
mmlu pro80.8%
scicode38.9%
tau265.8%
terminalbench hard23.5%

Benchmark data from Artificial Analysis.