Technology
Humans Outsmart AI at 2025 International Math Olympiad
Despite impressive performances from leading artificial intelligence models developed by Google and OpenAI, human contestants have once again outperformed machines at the 2025 International Mathematical Olympiad (IMO), held this month in Queensland, Australia.
In a breakthrough moment for AI, Google’s advanced Gemini chatbot and OpenAI’s experimental reasoning model each achieved gold-level scores, marking the first time generative AI systems have reached such heights at the prestigious global competition. However, neither model managed a perfect score, unlike five of the event’s young human participants.
Google confirmed on Monday that its Gemini model solved five out of the six challenging IMO problems. “We can confirm that Google DeepMind has reached the much-desired milestone, earning 35 out of a possible 42 points, a gold medal score,” said IMO president Gregor Dolinar. “Their solutions were astonishing in many respects. IMO graders found them to be clear, precise and most of them easy to follow.”
Out of the 641 contestants from 112 countries, only about 10 percent earned gold medals. Notably, five students achieved flawless performances with full 42-point scores.
OpenAI also reported that its own generative AI system reached the same gold-level benchmark. “The result achieved a longstanding grand challenge in AI,” said OpenAI researcher Alexander Wei. “We evaluated our models on the 2025 IMO problems under the same rules as human contestants,” he explained. “For each problem, three former IMO medalists independently graded the model’s submitted proof.”
Last year, Google reached a silver-medal score at the IMO in Bath, United Kingdom, solving four out of six problems. That attempt took two to three days of computation. In contrast, this year’s Gemini model completed its solutions within the strict 4.5-hour contest window, the same time limit given to human competitors.
According to the IMO, the AI models were privately tested using this year’s official problems, which were the same ones tackled by the student contestants. “It is very exciting to see progress in the mathematical capabilities of AI models,” noted Dolinar.
However, he also offered a word of caution: contest organizers could not independently verify the amount of computing power used in the AI trials or confirm the complete absence of human assistance in generating the results.
Despite AI’s growing capabilities, this year’s IMO showed that in the realm of abstract reasoning and mathematical problem-solving, human brilliance continues to lead.
PUNCH






















