Newsroom

News from the edge of what is proved

Mathematics changed shape in the last eighteen months: proofs now arrive from machines, and the argument about what counts is being had in public. We collect what actually happened — every entry with its primary source attached — and the papers worth reading behind it.

  • As of August 20, 2026
  • 15 stories
  • 12 publications

Dispatches

1 entry

Newest first. Every entry links the primary source first — the wiki, the blueprint, the abstract — and the reporting second.

  1. 42 out of 42 — the Olympiad stops being a measurement

    At the 2026 IMO in Shanghai several frontier systems were officially graded a perfect 42/42, among them RedNote's dots-note 3.0 and Huawei's Celia, with models from OpenAI, Anthropic, Axiom Math and Moonshot AI reported at the same score. Seven of 666 human contestants managed full marks. Two years after a silver medal and one after gold, the contest has saturated.

    SourcesFrance 24South China Morning Post

Publications

3 entries

The papers behind the headlines. Titles and author lists are reproduced as published; the line underneath says why it is on this list.

  1. First Proof Second BatchMohammed Abouzaid, Nikhil Srivastava, Rachel Ward, Lauren WilliamsTen research-level problems, methods, model attempts, human solutions, referee reports and logs in one public record. It is a high-quality capability evaluation, but only for its selected tasks and systems.arXiv:2606.18119June 16, 2026
  2. LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean AutoformalizationYuanhe Zhang, Yuekai Sun, Taiji Suzuki, Jason D. Lee, Fanghui LiuShort proofs are solved. The open question is whether a system can hold a formalization together over days. This benchmark measures that.arXiv:2606.05400June 4, 2026
  3. FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?Nikil Ravi, Kexing Ying, Vasilii Nesterov, Rayan Krishnan, Elif Uskuplu, Bingyu Xia, Janitha Aswedige, Langston NasholdGraduate-level statements, machine-checked answers: a benchmark that a convincing-sounding proof cannot pass.arXiv:2603.26996March 31, 2026