AI news, models and products, with sources中文
NewsLandmark Research21 Jul 2025

AI reaches gold-medal level at the Maths Olympiad

Experimental models from Google DeepMind and OpenAI each solve five of six problems at the International Mathematical Olympiad within the time limits, writing proofs in natural language. DeepMind's result is officially graded by the IMO.

Google DeepMind chief executive Demis Hassabis speaking at a lectern
Google DeepMind chief executive Demis Hassabis at the Nobel Lectures in December 2024. File photo. Photo: Jay Dixit / Wikimedia Commons, CC BY-SA 4.0

What happened

On Saturday 19 July 2025, OpenAI researcher Alexander Wei announced on X that an experimental OpenAI reasoning model had scored 35 points on the 2025 IMO problems, a gold-medal score.

Two days later, on 21 July, Google DeepMind said that an advanced version of Gemini Deep Think had also solved five problems and scored 35, and that this result had been certified by IMO coordinators. The IMO president, Gregor Dolinar, was quoted in DeepMind’s announcement: “We can confirm that Google DeepMind has reached the much-desired milestone, earning 35 out of a possible 42 points — a gold medal score.”

According to Reuters, it was the first time AI systems had crossed the gold-medal threshold at the IMO.

First, what is the IMO?

The 2025 IMO took place from 10 to 20 July on the Sunshine Coast in Queensland, Australia. According to the IMO’s website:

Contestants
630
Countries and territories
110
Gold medals (35+ points)
72
Perfect scores of 42
5

Source: IMO website, 2025 edition page. Silver cut-off 28 points (104 medals); bronze 19 points (145 medals).

So the AI score of 35 was exactly that year’s gold line. Seventy-two students reached or beat it, and five got every problem right.

IMO 2025: AI scores and the medal cut-offsOut of 42 points
  • Top student score (5 students)42
  • Gemini Deep Think (official marking)35
  • OpenAI experimental model (unofficial marking)35
  • Gold cut-off35
  • Silver cut-off28
  • Bronze cut-off19

OpenAI's score was assessed by three former IMO medallists it engaged, not by the IMO. Sources: IMO website; Google DeepMind blog, 21 Jul 2025; AFP and Reuters reports

What each company achieved

Google DeepMind. A year earlier, in 2024, DeepMind reached silver-medal standard with two specialist maths systems, AlphaProof and AlphaGeometry 2: four problems solved, 28 points. But experts first had to translate the problems from ordinary language into a “formal language” such as Lean, and translate the proofs back, and the computation took two to three days.

In 2025 DeepMind switched to Gemini’s general-purpose “Deep Think” mode. It says the model worked from the official problem statements and wrote rigorous proofs in natural language, all within the 4.5-hour time limit. It used “parallel thinking”: exploring several possible solutions at once and combining them before answering, instead of following a single line of reasoning. DeepMind also says this version received extra reinforcement-learning training, was given a curated set of high-quality solutions to maths problems, and had general tips on approaching IMO problems added to its instructions.

IMO graders described its solutions as “clear, precise and most of them easy to follow”.

OpenAI. According to Reuters, OpenAI used an experimental model built around massively scaling up “test-time compute”: letting the model think for longer and running many lines of reasoning in parallel. OpenAI researcher Noam Brown declined to say how much computing power it used, but called it “very expensive”.

Wei wrote that OpenAI “evaluated our models on the 2025 IMO problems under the same rules as human contestants”, and that “for each problem, three former IMO medalists independently graded the model’s submitted proof.” He added that OpenAI did not plan to release anything with this level of maths capability for several months.

Reactions

What it means. Junehyuk Jung, a Brown University maths professor and visiting researcher at DeepMind, told Reuters the result suggested AI was less than a year away from helping mathematicians crack unsolved research problems. “I think the moment we can solve hard reasoning problems in natural language will enable the potential for collaboration between AI and mathematicians,” he said. Jung won an IMO gold medal himself as a student in 2003.

A row over timing. Reuters reported that this was the first year the IMO coordinated officially with some AI developers. It certified the results of those companies, including Google, and asked them to publish on 28 July. DeepMind chief executive Demis Hassabis wrote on X that DeepMind had “respected the IMO Board’s original request” to wait until the official results were verified and the students had “rightly received the acclamation they deserved”. OpenAI announced first, on Saturday 19 July, and told Reuters it had permission from an IMO board member to publish after the closing ceremony. IMO president Dolinar told Reuters the competition allowed cooperating companies to publish on the Monday.

The students. AFP’s headline was “Humans beat AI gold-level score at top maths contest”: neither model scored full marks, while five young contestants did.

Things to keep in mind

  • OpenAI’s score is not official. OpenAI was not part of the IMO’s official evaluation; its score was assessed by former medallists it engaged. DeepMind’s was marked by IMO coordinators against the student criteria.
  • Certification covered the answers only. DeepMind notes that the IMO confirmed its submitted answers were complete and correct, but that this review “does not extend to validating our system, processes, or underlying model”.
  • Computing power and human involvement were not checked. AFP reported that Dolinar cautioned the organisers could not verify how much computing power the models used or whether there had been human involvement.
  • Conditions were not identical to the students’. DeepMind gave its model curated solutions and problem-solving tips, which contestants in the exam hall do not have.
  • Both were unreleased experimental models. OpenAI said nothing at this level would be released for several months; DeepMind said it would first give a version to trusted testers, including mathematicians, and then to Google AI Ultra subscribers.
  • 35 was exactly the cut-off, and neither system solved all six problems.

What it means for ordinary users

IMO problems cannot be solved by plugging numbers into a formula; they need step-by-step proofs that convince expert markers. An AI doing this in plain language within the time limit shows how quickly reasoning models were improving at rigorous logic. DeepMind said at the time that such systems would become tools for mathematicians, scientists and engineers.

Things moved fast from there. In September 2025, AI also featured at the world finals of the International Collegiate Programming Contest (see AI solves all 12 problems at the world programming finals). In September 2026 OpenAI announced that an internal model had found a proof for a version of a Millennium Prize problem, which set off a bitter dispute over credit (see OpenAI publishes a proof of a Millennium Prize problem).

For everyday users, one point is worth remembering: competition results come from specially tuned experimental versions using a lot of computing power. The models in everyday chat apps may not perform the same way, so check every step when you use AI for maths or logic.

Sources