← All comparisons

Claude 4.1 Opus (Reasoning) vs Grok 2 (Dec '24)

Anthropic vs xAI — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)Grok 2 (Dec '24)
Intelligence Index42.013.9
Coding Index36.5
Math Index80.3
Output speed (tok/s)44.50.0
Blended price ($/1M)$32.81$0.00
Time to first token (s)8.55s0.00s
aime13.3%
aime 2580.3%
artificial analysis coding index36.50
artificial analysis intelligence index42.0013.90
artificial analysis math index80.30
gpqa80.9%51.0%
hle11.9%3.8%
ifbench55.4%
lcr66.3%
livecodebench65.4%26.7%
math 50077.8%
mmlu pro88.0%70.9%
scicode40.9%28.5%
tau271.4%
terminalbench hard34.3%

Benchmark data from Artificial Analysis.