← All comparisons

Phi-4 vs Claude 4.5 Sonnet (Reasoning)

Microsoft vs Anthropic — side-by-side benchmark comparison

Phi-4Claude 4.5 Sonnet (Reasoning)
Intelligence Index10.443.0
Coding Index11.238.6
Math Index18.088.0
Output speed (tok/s)41.155.0
Blended price ($/1M)$0.22$6.56
Time to first token (s)0.50s7.02s
aime14.3%
aime 2518.0%88.0%
artificial analysis coding index11.2038.60
artificial analysis intelligence index10.4043.00
artificial analysis math index18.0088.00
gpqa57.5%83.4%
hle4.1%17.3%
ifbench23.5%57.3%
lcr0.0%65.7%
livecodebench23.1%71.4%
math 50081.0%
mmlu pro71.4%87.5%
scicode26.0%44.7%
tau20.0%78.1%
terminalbench hard3.8%35.6%

Benchmark data from Artificial Analysis.