← All comparisons

Claude 4.1 Opus (Reasoning) vs Grok 3 mini Reasoning (high)

Anthropic vs xAI — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)Grok 3 mini Reasoning (high)
Intelligence Index42.032.1
Coding Index36.525.2
Math Index80.384.7
Output speed (tok/s)44.556.8
Blended price ($/1M)$32.81$0.35
Time to first token (s)8.55s0.42s
aime93.3%
aime 2580.3%84.7%
artificial analysis coding index36.5025.20
artificial analysis intelligence index42.0032.10
artificial analysis math index80.3084.70
gpqa80.9%79.1%
hle11.9%11.1%
ifbench55.4%45.9%
lcr66.3%50.3%
livecodebench65.4%69.6%
math 50099.2%
mmlu pro88.0%82.8%
scicode40.9%40.6%
tau271.4%90.4%
terminalbench hard34.3%17.4%

Benchmark data from Artificial Analysis.