Xiaohongshu's dots-note-3.0 Wins IMO Gold with a Perfect Score, Marking China's First AI Triumph at the Mathematical Olympiad

Technology22.Jul.2026 01:096 min read

The results of the 67th International Mathematical Olympiad (IMO 2026) have been announced, with Xiaohongshu's dots-note-3.0 earning a gold medal after solving all six problems for a perfect score of 42 points. It is the first Chinese large language model to receive official IMO gold-level recognition and, following Google's Gemini, the second AI model worldwide to reach this benchmark—while also becoming the first to achieve a perfect score.

Xiaohongshu's dots-note-3.0 Wins IMO Gold with a Perfect Score, Marking China's First AI Triumph at the Mathematical Olympiad

The results of the 67th International Mathematical Olympiad (IMO 2026) have now been released, and Xiaohongshu’s large model, dots-note-3.0, delivered a standout performance: it solved all six problems correctly and achieved the maximum possible score of 42 points, earning a gold-medal result.

This marks a first for China’s large-model ecosystem. According to the reported outcome, dots-note-3.0 is the first Chinese model to reach an IMO gold-medal level under official evaluation, and only the second model globally to do so after Google’s Gemini. More notably, while Gemini had previously reached a gold-medal standard by solving 5 of the 6 problems, dots-note-3.0 went a step further and finished with a perfect score.

Why the IMO matters so much for AI reasoning

The International Mathematical Olympiad is widely regarded as one of the toughest mathematical competitions in the world for elite high school students. It has long been a proving ground for exceptional mathematical talent, and many of today’s most prominent mathematicians—including a significant share of Fields Medal winners such as Terence Tao—once competed on this stage.

The format itself explains why the IMO is such a demanding benchmark. The contest runs across two days, with contestants facing three problems per day and receiving 4.5 hours each day to complete them. The questions span algebra, combinatorics, geometry, and number theory. Each problem is worth 7 points, for a total of 42.

What makes the IMO especially valuable for evaluating AI systems is that it is not a test of final answers alone. Participants must present complete, rigorous proofs that can withstand detailed scrutiny. That requirement raises the bar significantly: a model must do far more than produce an answer that looks plausible. It has to understand the problem, develop a valid line of argument, and express that reasoning in a form that can be checked step by step.

How dots-note-3.0 compares with Gemini

Google’s Gemini provides a useful point of comparison for measuring progress in this area. In its first IMO attempt in 2024, Gemini scored 28 points, enough for a silver-medal result. At that stage, it reportedly relied on converting problems into formal languages such as Lean, and some solutions took days to complete.

By 2025, Gemini Deep Think had improved substantially. It was able to produce end-to-end proofs in natural language and solve 5 problems within the official 4.5-hour time limit, reaching gold-medal level.

dots-note-3.0’s IMO 2026 result pushes that ceiling even higher. Solving all six problems correctly under these conditions is a stronger outcome than a 5-of-6 gold-level finish, and it suggests another meaningful leap in the mathematical reasoning ability of large models.

That achievement carries extra weight because IMO problems are created by experts and kept under strict confidentiality before the contest. In other words, models cannot be specifically trained on the exact questions in advance. Performances in this setting therefore say more about real-time comprehension, exploration, and disciplined reasoning than many other benchmark results do.

Why the perfect score stands out

The strength of the result becomes even clearer when viewed against this year’s scoring threshold. For IMO 2026, the gold-medal cutoff was 29 points, which was 6 points lower than last year’s 35. That drop indicates that this year’s paper was significantly harder overall.

Under the scoring rules, reaching 29 points could mean fully solving four problems and then collecting just 1 additional point from the remaining two. dots-note-3.0 did far more than that: it completed the contest with the full 42 points, finishing 13 points above the gold line.

In a competition where even top human contestants often leave points on the table, that margin is difficult to ignore.

What this result suggests about the model

First, the perfect score points to a high level of stability in agentic reasoning. Success at the IMO requires a long chain of capabilities: understanding natural-language statements, identifying the key constraints, generating intermediate claims, testing possible approaches, and then organizing everything into a proof that remains logically sound from start to finish. A single misread condition or unjustified leap can undermine an entire solution.

Because dots-note-3.0 achieved full marks across six different problems, the performance appears less like a one-off flash of insight and more like evidence of consistent reasoning quality across multiple mathematical domains.

Second, the model reportedly handled the problems end to end in natural language, rather than depending on a separate formalization pipeline. That matters because it suggests a more complete loop from problem understanding to final explanation, and potentially stronger transferability beyond narrowly structured theorem-proving setups.

Why its third solution drew particular attention

During the competition, dots-note-3.0 was said to operate in an agent-style framework, combining natural-language reasoning with Python code execution and using recursive self-critique to examine its own arguments.

Among its six solutions, the answer to Problem 3 has attracted the most discussion. The problem belonged to combinatorial game theory. A more common route, according to descriptions of the solution, would be to transform the original question into one about graph connectivity and then build the proof from there.

Instead, dots-note-3.0 chose an inductive approach. It focused on the deeper structural core of the problem and selected an induction object and proposition that fit the task particularly well. That decision appears to have impressed experienced human competitors.

Liu Hanzuo, a two-time CMO gold medalist, described the solution as “correct and beautiful,” praising its compact structure and natural logic, and noting that it would count as a relatively concise and elegant IMO-style proof.

Another CMO gold medalist, Wang Qiantong, said the method was “clear, concise, and directly aimed at the essence of the problem,” adding that while each step felt natural in hindsight, it would be difficult for human contestants to think of that angle during the contest.

What it means for Xiaohongshu

For Xiaohongshu, the result is more than just a competition headline. The company has traditionally been associated in public perception with its content community, recommendation systems, and content understanding capabilities. In the foundation-model space, however, it has lacked a similarly visible and authoritative demonstration of technical strength.

An IMO perfect score changes that conversation quickly. A result like 42 out of 42 is easy for the industry to understand and difficult to dismiss. It offers a highly direct form of technical validation and positions Xiaohongshu as a foundation-model player that deserves much closer attention.

According to the report, dots-note-3.0 is expected to be open-sourced soon.