People are using Super Mario to benchmark AI now

March 3, 2025

53

Thought Pokémon was a tough benchmark for AI? One group of researchers argues that Super Mario Bros. is even tougher.

Hao AI Lab, a research org at the University of California San Diego, on Friday threw AI into live Super Mario Bros. games. Anthropic’s Claude 3.7 performed the best, followed by Claude 3.5. Google’s Gemini 1.5 Pro and OpenAI’s GPT-4o struggled.

It wasn’t quite the same version of Super Mario Bros. as the original 1985 release, to be clear. The game ran in an emulator and integrated with a framework, GamingAgent, to give the AIs control over Mario.

Image Credits:Hao Lab

GamingAgent, which Hao developed in-house, fed the AI basic instructions, like, “If an obstacle or enemy is near, move/jump left to dodge” and in-game screenshots. The AI then generated inputs in the form of Python code to control Mario.

Still, Hao says that the game forced each model to “learn” to plan complex maneuvers and develop gameplay strategies. Interestingly, the lab found that reasoning models like OpenAI’s o1, which “think” through problems step by step to arrive at solutions, performed worse than “non-reasoning” models, despite being generally stronger on most benchmarks.

One of the main reasons reasoning models have trouble playing real-time games like this is that they take a while — seconds, usually — to decide on actions, according to the researchers. In Super Mario Bros., timing is everything. A second can mean the difference between a jump safely cleared and a plummet to your death.

Games have been used to benchmark AI for decades. But some experts have questioned the wisdom of drawing connections between AI’s gaming skills and technological advancement. Unlike the real world, games tend to be abstract and relatively simple, and they provide a theoretically infinite amount of data to train AI.

The recent flashy gaming benchmarks point to what Andrej Karpathy, a research scientist and founding member at OpenAI, called an “evaluation crisis.”

“I don’t really know what [AI] metrics to look at right now,” he wrote in a post on X. “TLDR my reaction is I don’t really know how good these models are right now.”

At least we can watch AI play Mario.

Source link

A post-Trump restoration is still possible

Amazon shares fall as it prepares $200bn AI spending blitz

What happens next in Venezuela?

Trump’s Venezuela punt could turn into an oil-drilling own goal

Asda and Morrisons’ private equity owners raise £6.5bn in property deals

France vs Norway FIFA world cup 2026 live match

Top 10 highest goal scorer in the world cup history

Top 10 Youngest Players Ever to Play in the FIFA World Cup

Canadian Ice Hockey Legend Claude Lemieux Dies at 60

Too soon? Bears, 49ers among early bets to win next year’s Super Bowl

People are using Super Mario to benchmark AI now

Former Tesla product manager wants to make luxury goods impossible to fake, starting with a chip

What to know about Netflix’s landmark acquisition of Warner Bros.

YouTube rolls out an AI playlist generator for Premium users

LEAVE A REPLY Cancel reply

Most Popular

France vs Norway FIFA world cup 2026 live match

Top 10 highest goal scorer in the world cup history

Top 10 Youngest Players Ever to Play in the FIFA World Cup

Canadian Ice Hockey Legend Claude Lemieux Dies at 60

Recent Comments

EDITOR PICKS

France vs Norway FIFA world cup 2026 live match

Top 10 highest goal scorer in the world cup history

Top 10 Youngest Players Ever to Play in the FIFA World Cup

POPULAR POSTS

France vs Norway FIFA world cup 2026 live match

Top 10 highest goal scorer in the world cup history

Top 10 Youngest Players Ever to Play in the FIFA World Cup

POPULAR CATEGORY

ABOUT US

FOLLOW US