Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad
Google DeepMind announced that an advanced version of Gemini with Deep Think achieved gold-medal standard at the International Mathematical Olympiad 2025, solving five of six problems for 35 points. Unlike the prior year's silver-medal result, this model operated end-to-end in natural language within the 4.5-hour competition time limit. The results were officially graded and certified by IMO coordinators using the same criteria as for student solutions.
Key Takeaways
- An advanced version of Gemini Deep Think solved five of six IMO 2025 problems perfectly, earning 35 points and gold-medal level performance.
- The results were officially graded and certified by IMO coordinators using the same criteria as student solutions.
- Last year, the combined AlphaProof and AlphaGeometry 2 systems achieved silver, solving four of six problems for 28 points.
- This year's model operated end-to-end in natural language, producing proofs directly from official problem descriptions within the 4.5-hour limit.
- Deep Think uses parallel thinking to explore and combine multiple solutions rather than a single linear chain of thought.
- The model was trained with novel reinforcement learning techniques and given a curated corpus of high-quality math solutions.
Stats & Key Facts
- #Five of six IMO 2025 problems solved perfectly
- #35 total points this year
- #28 points and four of six problems last year
- #4.5-hour competition time limit
- #Approximately 8% of contestants receive a gold medal
- #Two to three days of computation required last year
What the IMO is
- ›The International Mathematical Olympiad is the world's most prestigious competition for young mathematicians, held annually since 1959.
- ›Each participating country is represented by six elite, pre-university mathematicians.
- ›They solve six exceptionally difficult problems in algebra, combinatorics, geometry, and number theory.
Medals go to the top half of contestants, with about 8% receiving a gold medal. The IMO has also become an aspirational test for AI systems' mathematical reasoning.
This year's result
- ›An advanced version of Gemini Deep Think solved five of six problems perfectly.
- ›It earned 35 total points, achieving gold-medal level performance.
- ›The solutions are available online.
DeepMind says it was among an inaugural cohort to have model results officially graded and certified by IMO coordinators using the same criteria as student solutions.
A step beyond last year
The 2025 result advances on the prior year.
- ›At IMO 2024, AlphaProof and AlphaGeometry 2 achieved silver, solving four of six problems for 28 points.
- ›Last year's approach required experts to translate problems into domain-specific languages such as Lean and took two to three days of computation.
- ›This year, the model operated end-to-end in natural language within the 4.5-hour competition time limit.
How Deep Think works
- ›Deep Think is an enhanced reasoning mode for complex problems.
- ›It uses parallel thinking to simultaneously explore and combine multiple possible solutions before a final answer.
- ›This replaces a single, linear chain of thought.
DeepMind also trained this version with novel reinforcement learning techniques that use more multi-step reasoning, problem-solving, and theorem-proving data, and provided a curated corpus of high-quality solutions plus general hints on approaching IMO problems.
What comes next
- ›A version of the Deep Think model will go to a set of trusted testers, including mathematicians, before rolling out to Google AI Ultra subscribers.
- ›DeepMind continues work on formal systems AlphaGeometry and AlphaProof.
- ›It expects agents combining natural language fluency with rigorous, verified reasoning to become valuable tools for mathematicians and scientists.
Frequently Asked Questions
How did Gemini perform at IMO 2025?
An advanced version of Gemini Deep Think solved five of six problems perfectly for 35 points, achieving gold-medal level performance.
How does this compare to last year?
Last year, AlphaProof and AlphaGeometry 2 achieved silver, solving four of six problems for 28 points and requiring two to three days of computation.
Did the model work in natural language?
Yes. This year's model operated end-to-end in natural language, producing proofs directly from official problem descriptions within the 4.5-hour limit.
Were the results independently graded?
Yes. IMO coordinators officially graded and certified the results using the same criteria as for student solutions.
What is Deep Think mode?
It is an enhanced reasoning mode that uses parallel thinking to explore and combine multiple solutions before answering, rather than a single linear chain of thought.
Gemini Deep Think's officially certified gold-medal result, achieved end-to-end in natural language within the time limit, marks a clear advance over the prior year's silver.
Continue Learning
Comments
Sign in to join the conversation