← All comparisons

Claude 4.5 Sonnet (Reasoning) vs Phi-3 Mini Instruct 3.8B

Anthropic vs Microsoft — side-by-side benchmark comparison

Claude 4.5 Sonnet (Reasoning)Phi-3 Mini Instruct 3.8B
Intelligence Index43.010.1
Coding Index38.63.0
Math Index88.00.3
Output speed (tok/s)55.00.0
Blended price ($/1M)$6.56$0.00
Time to first token (s)7.02s0.00s
aime4.0%
aime 2588.0%0.3%
artificial analysis coding index38.603.00
artificial analysis intelligence index43.0010.10
artificial analysis math index88.0030.0%
gpqa83.4%31.9%
hle17.3%4.4%
ifbench57.3%23.9%
lcr65.7%2.0%
livecodebench71.4%11.6%
math 50045.7%
mmlu pro87.5%43.5%
scicode44.7%9.0%
tau278.1%0.0%
terminalbench hard35.6%0.0%

Benchmark data from Artificial Analysis.