← All comparisons

Claude 4.1 Opus (Reasoning) vs Phi-3 Mini Instruct 3.8B

Anthropic vs Microsoft — side-by-side benchmark comparison

Claude 4.1 Opus (Reasoning)Phi-3 Mini Instruct 3.8B
Intelligence Index42.010.1
Coding Index36.53.0
Math Index80.30.3
Output speed (tok/s)44.50.0
Blended price ($/1M)$32.81$0.00
Time to first token (s)8.55s0.00s
aime4.0%
aime 2580.3%0.3%
artificial analysis coding index36.503.00
artificial analysis intelligence index42.0010.10
artificial analysis math index80.3030.0%
gpqa80.9%31.9%
hle11.9%4.4%
ifbench55.4%23.9%
lcr66.3%2.0%
livecodebench65.4%11.6%
math 50045.7%
mmlu pro88.0%43.5%
scicode40.9%9.0%
tau271.4%0.0%
terminalbench hard34.3%0.0%

Benchmark data from Artificial Analysis.