REASONING

Two AI Systems Posted a Perfect Score at the 2026 Math Olympiad

Huawei's Celia and RedNote's dots-note-3.0 each scored a perfect 42 out of 42 at the 2026 International Mathematical Olympiad under the competition's own grading — a first. Four more frontier models reportedly matched that score days later, but only in an unofficial test graded by another AI, not by IMO judges.

An empty examination hall lit in amber, desks covered with proof-diagram papers under a shadowed judges' podium.
An empty examination hall lit in amber, desks covered with proof-diagram papers under a shadowed judges' podium.

Two AI systems — Huawei's Celia and RedNote's dots-note-3.0 — solved all six problems at the 2026 International Mathematical Olympiad and received a perfect 42 out of 42 from the competition's own graders, the first time a language model has cleared the IMO's official judging process at full marks.

The exam was held in Shanghai on July 15-16, 2026, alongside 666 human contestants; the closing ceremony followed on July 22. Only seven of those students matched the AI systems' score, a reminder of how rare a clean sweep is even among the world's strongest teenage mathematicians.

The IMO's AI track runs under strict conditions: companies receive the six problems only after the human competition has finished, submit written solutions inside a fixed window, and are barred from any human intervention during that time. Submissions are graded by IMO's own coordinators against the same rubric applied to student papers — one that rewards a complete, gap-free proof, not just a correct final answer.

That distinction matters because a year earlier, in 2025, Google DeepMind and OpenAI each reported 35 out of 42 at the same competition — gold-medal level, but short of perfect — using verification methods outside the IMO's own channel. The jump from 35 to 42 in a single year marks the point where the exam's hardest six-problem, two-day format produced its first flawless AI papers under the organizers' own scoring.

A parallel, unofficial track shows why the word "official" carries weight here. Days after the results, investor Deedy Das ran the same problem set through four more frontier systems — Claude Fable 5, GPT-5.6 Sol, Kimi K3, and the Lean-based prover AxiomProver — and reported that all four also reached 42 out of 42. But those scores were graded by Claude-based agents in Das's own harness, not by IMO judges; Das's own notes describe the results as "strong but not authoritative."

The gap between the two tracks is the actual story. An official IMO grade means a human coordinator read a natural-language proof against the same bar used for a sixteen-year-old's paper and found no gap in logic. A self-administered grade means one chatbot decided whether another chatbot's proof holds up — a lower and far less tested bar, whatever the raw number reads.

RedNote said it plans to open-source dots-note-3.0 "in the near future," without giving a date. The model is still described as the lightest, beta-stage entry in the company's dots3 family, alongside two larger variants aimed at heavier workloads.