A discrepancy between first- and third-party benchmark results for OpenAIโs o3 AI model is raising questions about the companyโs transparency and model testing practices.
When OpenAI unveiled o3 in December, the company claimed the model could answer just over a fourth of questions on FrontierMath, a challenging set of math problems. That score blew the competition away โ the next-best model managed to answer only around 2% of FrontierMath problems correctly.
โToday, all offerings out there have less than 2% [on FrontierMath],โ Mark Chen, chief research officer at OpenAI, said during a livestream. โWeโre seeing [internally], with o3 in aggressive test-time compute settings, weโre able to get over 25%.โ
As it turns out, that figure was likely an upper bound, achieved by a version of o3 with more computing behind it than the model OpenAI publicly launched last week.
Epoch AI, the research institute behind FrontierMath, released results of its independent benchmark tests of o3 on Friday. Epoch found that o3 scored around 10%, well below OpenAIโs highest claimed score.
OpenAI has released o3, their highly anticipated reasoning model, along with o4-mini, a smaller and cheaper model that succeeds o3-mini.
We evaluated the new models on our suite of math and science benchmarks. Results in thread! pic.twitter.com/5gbtzkEy1B
โ Epoch AI (@EpochAIResearch) April 18, 2025
That doesnโt mean OpenAI lied, per se. The benchmark results the company published in December show a lower-bound score that matches the score Epoch observed. Epoch also noted its testing setup likely differs from OpenAIโs, and that it used an updated release of FrontierMath for its evaluations.
โThe difference between our results and OpenAIโs might be due to OpenAI evaluating with a more powerful internal scaffold, using more test-time [computing], or because those results were run on a different subset of FrontierMath (the 180 problems in frontiermath-2024-11-26 vs the 290 problems in frontiermath-2025-02-28-private),โ wrote Epoch.
According to a post on X from the ARC Prize Foundation, an organization that tested a pre-release version of o3, the public o3 model โis a different model [โฆ] tuned for chat/product use,โ corroborating Epochโs report.
โAll released o3 compute tiers are smaller than the version we [benchmarked],โ wrote ARC Prize. Generally speaking, bigger compute tiers can be expected to achieve better benchmark scores.
Re-testing released o3 on ARC-AGI-1 will take a day or two. Because todayโs release is a materially different system, we are re-labeling our past reported results as โpreviewโ:
o3-preview (low): 75.7%, $200/task
o3-preview (high): 87.5%, $34.4k/taskAbove uses o1 pro pricingโฆ
โ Mike Knoop (@mikeknoop) April 16, 2025
OpenAIโs own Wenda Zhou, a member of the technical staff, said during a livestream last week that the o3 in production is โmore optimized for real-world use casesโ and speed versus the version of o3 demoed in December. As a result, it may exhibit benchmark โdisparities,โ he added.
โ[W]eโve done [optimizations] to make the [model] more cost efficient [and] more useful in general,โ Zhou said. โWe still hope that โ we still think that โ this is a much better model [โฆ] You wonโt have to wait as long when youโre asking for an answer, which is a real thing with these [types of] models.โ
Granted, the fact that the public release of o3 falls short of OpenAIโs testing promises is a bit of a moot point, since the companyโs o3-mini-high and o4-mini models outperform o3 on FrontierMath, and OpenAI plans to debut a more powerful o3 variant, o3-pro, in the coming weeks.
It is, however, another reminder that AI benchmarks are best not taken at face value โ particularly when the source is a company with services to sell.
Benchmarking โcontroversiesโ are becoming a common occurrence in the AI industry as vendors race to capture headlines and mindshare with new models.
In January, Epoch was criticized for waiting to disclose funding from OpenAI until after the company announced o3. Many academics who contributed to FrontierMath werenโt informed of OpenAIโs involvement until it was made public.
More recently, Elon Muskโs xAI was accused of publishing misleading benchmark charts for its latest AI model, Grok 3. Just this month, Meta admitted to touting benchmark scores for a version of a model that differed from the one the company made available to developers.
Updated 4:21 p.m. Pacific: Added comments from Wenda Zhou, a member of the OpenAI technical staff, from a livestream last week.


