← All comparisons

gpt-oss-120b (low) vs Claude 4.1 Opus (Reasoning)

OpenAI vs Anthropic — side-by-side benchmark comparison

gpt-oss-120b (low)Claude 4.1 Opus (Reasoning)
Intelligence Index24.542.0
Coding Index15.536.5
Math Index66.780.3
Output speed (tok/s)370.044.5
Blended price ($/1M)$0.26$32.81
Time to first token (s)0.49s8.55s
aime
aime 2566.7%80.3%
artificial analysis coding index15.5036.50
artificial analysis intelligence index24.5042.00
artificial analysis math index66.7080.30
gpqa67.2%80.9%
hle5.2%11.9%
ifbench58.3%55.4%
lcr43.7%66.3%
livecodebench70.7%65.4%
math 500
mmlu pro77.5%88.0%
scicode36.0%40.9%
tau245.0%71.4%
terminalbench hard5.3%34.3%

Benchmark data from Artificial Analysis.