Close Menu
Luminari | Learn Docker, Kubernetes, AI, Tech & Interview PrepLuminari | Learn Docker, Kubernetes, AI, Tech & Interview Prep
  • Home
  • Technology
    • Docker
    • Kubernetes
    • AI
    • Cybersecurity
    • Blockchain
    • Linux
    • Python
    • Tech Update
    • Interview Preparation
    • Internet
  • Entertainment
    • Movies
    • TV Shows
    • Anime
    • Cricket
What's Hot

Dubai Real Estate Hits $18.2B in Sales Amid Tokenization Push

June 8, 2025

Malicious Browser Extensions Infect 722 Users Across Latin America Since Early 2025

June 8, 2025

American Psycho Director Mary Harron Surprised Movie Still Relevant

June 8, 2025
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Luminari | Learn Docker, Kubernetes, AI, Tech & Interview Prep
  • Home
  • Technology
    • Docker
    • Kubernetes
    • AI
    • Cybersecurity
    • Blockchain
    • Linux
    • Python
    • Tech Update
    • Interview Preparation
    • Internet
  • Entertainment
    • Movies
    • TV Shows
    • Anime
    • Cricket
Luminari | Learn Docker, Kubernetes, AI, Tech & Interview PrepLuminari | Learn Docker, Kubernetes, AI, Tech & Interview Prep
Home » OpenAI’s o3 AI model scores lower on a benchmark than the company initially implied
AI

OpenAI’s o3 AI model scores lower on a benchmark than the company initially implied

HarishBy HarishApril 20, 2025No Comments4 Mins Read
Facebook Twitter Pinterest LinkedIn Reddit WhatsApp Email
Share
Facebook Twitter Pinterest Reddit WhatsApp Email


A discrepancy between first- and third-party benchmark results for OpenAI’s o3 AI model is raising questions about the company’s transparency and model testing practices.

When OpenAI unveiled o3 in December, the company claimed the model could answer just over a fourth of questions on FrontierMath, a challenging set of math problems. That score blew the competition away — the next-best model managed to answer only around 2% of FrontierMath problems correctly.

“Today, all offerings out there have less than 2% [on FrontierMath],” Mark Chen, chief research officer at OpenAI, said during a livestream. “We’re seeing [internally], with o3 in aggressive test-time compute settings, we’re able to get over 25%.”

As it turns out, that figure was likely an upper bound, achieved by a version of o3 with more computing behind it than the model OpenAI publicly launched last week.

Epoch AI, the research institute behind FrontierMath, released results of its independent benchmark tests of o3 on Friday. Epoch found that o3 scored around 10%, well below OpenAI’s highest claimed score.

OpenAI has released o3, their highly anticipated reasoning model, along with o4-mini, a smaller and cheaper model that succeeds o3-mini.

We evaluated the new models on our suite of math and science benchmarks. Results in thread! pic.twitter.com/5gbtzkEy1B

— Epoch AI (@EpochAIResearch) April 18, 2025

That doesn’t mean OpenAI lied, per se. The benchmark results the company published in December show a lower-bound score that matches the score Epoch observed. Epoch also noted its testing setup likely differs from OpenAI’s, and that it used an updated release of FrontierMath for its evaluations.

“The difference between our results and OpenAI’s might be due to OpenAI evaluating with a more powerful internal scaffold, using more test-time [computing], or because those results were run on a different subset of FrontierMath (the 180 problems in frontiermath-2024-11-26 vs the 290 problems in frontiermath-2025-02-28-private),” wrote Epoch.

According to a post on X from the ARC Prize Foundation, an organization that tested a pre-release version of o3, the public o3 model “is a different model […] tuned for chat/product use,” corroborating Epoch’s report.

“All released o3 compute tiers are smaller than the version we [benchmarked],” wrote ARC Prize. Generally speaking, bigger compute tiers can be expected to achieve better benchmark scores.

OpenAI’s own Wenda Zhou, a member of the technical staff, said during a livestream last week that the o3 in production is “more optimized for real-world use cases” and speed versus the version of o3 demoed in December. As a result, it may exhibit benchmark “disparities,” he added.

“[W]e’ve done [optimizations] to make the [model] more cost efficient [and] more useful,” Zhou said. “We still hope that — we still think that — this is a much better model.”

Granted, the fact that the public release of o3 falls short of OpenAI’s testing promises is a bit of a moot point, since the company’s o3-mini-high and o4-mini models outperform o3 on FrontierMath, and OpenAI plans to debut a more powerful o3 variant, o3-pro, in the coming weeks.

It is, however, another reminder that AI benchmarks are best not taken at face value — particularly when the source is a company with services to sell.

Benchmarking “controversies” are becoming a common occurrence in the AI industry as vendors race to capture headlines and mindshare with new models.

In January, Epoch was criticized for waiting to disclose funding from OpenAI until after the company announced o3. Many academics who contributed to FrontierMath weren’t informed of OpenAI’s involvement until it was made public.

More recently, Elon Musk’s xAI was accused of publishing misleading benchmark charts for its latest AI model, Grok 3. Just this month, Meta admitted to touting benchmark scores for a version of a model that differed from the one the company made available to developers.

Updated 4:21 p.m. Pacific: Added comments from Wenda Zhou, a member of the OpenAI technical staff, from a livestream.





Source link

Share. Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Email
Previous ArticleBitget detects irregularity in VOXEL-USDT futures, rolls back accounts
Next Article What to Stream: ‘Andor,’ ‘Babygirl’ and new Wu Tang Clan music
Harish
  • Website
  • X (Twitter)

Related Posts

Lawyers could face ‘severe’ penalties for fake AI-generated citations, UK court warns

June 7, 2025

Trump administration takes aim at Biden and Obama cybersecurity rules

June 7, 2025

Week in Review: Why Anthropic cut access to Windsurf

June 7, 2025

Will Musk vs. Trump affect xAI’s $5 billion debt deal?

June 7, 2025

Building More Scalable GenAI Applications for Startups and Developers

June 7, 2025

2025 will be a ‘pivotal year’ for Meta’s augmented and virtual reality, says CTO

June 6, 2025
Add A Comment
Leave A Reply Cancel Reply

Our Picks

Dubai Real Estate Hits $18.2B in Sales Amid Tokenization Push

June 8, 2025

Malicious Browser Extensions Infect 722 Users Across Latin America Since Early 2025

June 8, 2025

American Psycho Director Mary Harron Surprised Movie Still Relevant

June 8, 2025

Why Gerard Butler Returned for Live-Action ‘How to Train Your Dragon’

June 8, 2025
Don't Miss
Blockchain

Dubai Real Estate Hits $18.2B in Sales Amid Tokenization Push

June 8, 20253 Mins Read

Dubai’s real estate market surged in May, posting record sales volumes and transaction values that…

Bitcoin market of 2025 driven by stablecoin regulation: Finance Redefined

June 6, 2025

How to Earn Passive Income with Peer-to-Peer Lending

June 6, 2025

Mass data deletion by governments is accelerating.

June 6, 2025

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

About Us
About Us

Welcome to Luminari, your go-to hub for mastering modern tech and staying ahead in the digital world.

At Luminari, we’re passionate about breaking down complex technologies and delivering insights that matter. Whether you’re a developer, tech enthusiast, job seeker, or lifelong learner, our mission is to equip you with the tools and knowledge you need to thrive in today’s fast-moving tech landscape.

Our Picks

Lawyers could face ‘severe’ penalties for fake AI-generated citations, UK court warns

June 7, 2025

Trump administration takes aim at Biden and Obama cybersecurity rules

June 7, 2025

Week in Review: Why Anthropic cut access to Windsurf

June 7, 2025

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

Facebook X (Twitter) Instagram Pinterest
  • Home
  • About Us
  • Advertise With Us
  • Contact Us
  • DMCA Policy
  • Privacy Policy
  • Terms & Conditions
© 2025 luminari. Designed by luminari.

Type above and press Enter to search. Press Esc to cancel.