← All comparisons

Claude 4.1 Opus (Reasoning) vs Grok 4.20 0309 v2 (Non-reasoning)

Anthropic vs xAI — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)Grok 4.20 0309 v2 (Non-reasoning)
Intelligence Index42.029.0
Coding Index36.522.0
Math Index80.3
Output speed (tok/s)44.5175.2
Blended price ($/1M)$32.81$3.00
Time to first token (s)8.55s0.47s
aime
aime 2580.3%
artificial analysis coding index36.5022.00
artificial analysis intelligence index42.0029.00
artificial analysis math index80.30
gpqa80.9%77.6%
hle11.9%24.2%
ifbench55.4%49.3%
lcr66.3%17.3%
livecodebench65.4%
math 500
mmlu pro88.0%
scicode40.9%32.8%
tau271.4%59.9%
terminalbench hard34.3%16.7%

Benchmark data from Artificial Analysis.