← All comparisons

Phi-4 Multimodal Instruct vs Claude 4.5 Sonnet (Reasoning)

Microsoft vs Anthropic — side-by-side benchmark comparison

Phi-4 Multimodal InstructClaude 4.5 Sonnet (Reasoning)
Intelligence Index10.043.0
Coding Index38.6
Math Index88.0
Output speed (tok/s)16.655.0
Blended price ($/1M)$0.00$6.56
Time to first token (s)1.33s7.02s
aime9.3%
aime 2588.0%
artificial analysis coding index38.60
artificial analysis intelligence index10.0043.00
artificial analysis math index88.00
gpqa31.5%83.4%
hle4.4%17.3%
ifbench57.3%
lcr65.7%
livecodebench13.1%71.4%
math 50069.3%
mmlu pro48.5%87.5%
scicode11.0%44.7%
tau278.1%
terminalbench hard35.6%

Benchmark data from Artificial Analysis.