Microsoft launched several new โopenโ AI models on Wednesday, the most capable of which is competitive with OpenAIโs o3-mini on at least one benchmark.
All of the new pemissively licensed models โ Phi 4 mini reasoning, Phi 4 reasoning, and Phi 4 reasoning plus โ are โreasoningโ models, meaning theyโre able to spend more time fact-checking solutions to complex problems. They expand Microsoftโs Phi โsmall modelโ family, which the company launched a year ago to offer a foundation for AI developers building apps at the edge.
Phi 4 mini reasoningย was trained on roughly 1 million synthetic math problems generated by Chinese AI startup DeepSeekโs R1 reasoning model. Around 3.8 billion parameters in size, Phi 4 mini reasoning is designed for educational applications, Microsoft says, like โembedded tutoringโ on lightweight devices.
Parameters roughly correspond to a modelโs problem-solving skills, and models with more parameters generally perform better than those with fewer parameters.
Phi 4 reasoning, a 14-billion-parameter model, was trained using โhigh-qualityโ web data as well as โcurated demonstrationsโ from OpenAIโs aforementioned o3-mini. Itโs best for math, science, and coding applications, according to Microsoft.
As for Phi 4 reasoning plus, itโs Microsoftโs previously-released Phi-4 model adapted into a reasoning model to achieve better accuracy on particular tasks. Microsoft claims that Phi 4 reasoning plus approaches the performance levels of R1, a model with significantly more parameters (671 billion). The companyโs internal benchmarking also has Phi 4 reasoning plus matching o3-mini on OmniMath, a math skills test.
Phi 4 mini reasoning, Phi 4 reasoning, and Phi 4 reasoning plus are available on the AI dev platform Hugging Face accompanied by detailed technical reports.
Techcrunch event
Berkeley, CA
|
June 5
BOOK NOW
โUsing distillation, reinforcement learning, and high-quality data, these [new] models balance size and performance,โ wrote Microsoft in a blog post. โThey are small enough for low-latency environments yet maintain strong reasoning capabilities that rival much bigger models. This blend allows even resource-limited devices to perform complex reasoning tasks efficiently.โ


