METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

OpenAI publishes 722 math manuscripts on GitHub

OpenAI has posted 722 mathematical manuscripts produced by an unreleased internal model to GitHub, grouped into 372 result families. The collection includes claims on the quasi-Riemann hypothesis and the Hodge conjecture, but many results are not yet formalized, and the model's name and prompts were left out.

OpenAI publishes 722 math manuscripts on GitHub

Image: METAL

Summary

  • On October 6, OpenAI published 722 mathematical manuscripts produced by an unreleased internal model, grouped into 372 result families, in the GitHub repository openai/math.
  • The model was posed about 4,000 problems and each result used on average about three hours of ChatGPT Pro thinking compute; many come with Lean formalizations, but the company said unformalized results could have issues.
  • OpenAI said it drew on recommendations from an advisory group at the Institute for Advanced Study in Princeton but did not disclose the model's name or per-result prompts, and it announced support for workshops and conferences along with plans to release the model.

On October 6, OpenAI uploaded 722 mathematical manuscripts produced by an unreleased internal frontier model to the GitHub repository openai/math in a single release. The papers are organized into 372 result families that group related results, and many come with Lean formalizations that let a computer check the proofs. The company said the model was posed about 4,000 problems over the evaluation period and that each result used, on average, compute equivalent to about three hours of ChatGPT Pro thinking. The repository also carries a caveat that results without formalizations could have issues.

The 41-page overview PDF in the repository, which METAL reviewed, lists the results across 17 fields, including number theory, algebraic geometry, theoretical computer science, combinatorics, operator algebras and partial differential equations. Each claim on the list addresses a long-standing question in mathematics. The list includes the so-called quasi-Riemann hypothesis, that Dirichlet L-functions have no zeros where the real part exceeds 7/8; a negative resolution of Hilbert's tenth problem over the rationals; the irrationality of Catalan's constant; the rational Hodge conjecture for CM abelian varieties; a result that all nonabelian free group factors are isomorphic; and the circulant Hadamard conjecture. OpenAI noted that the numbering does not indicate a ranking.

The company also disclosed how the results were made. According to the README, most results came from a single internal model following the same procedure. OpenAI widened its evaluations to open research problems after performance on its existing math evaluations saturated, and some outputs build on results the model produced earlier. The two exceptions are work on a zero-free region for the Riemann zeta function and the proof of the Hodge conjecture for CM abelian varieties, and the manuscript on the 11/12 region was edited by a human for readability. The repository also includes summaries of the model's reasoning for 10 results, among them the irrationality exponent of pi, the Mahler conjectures and the Mézard–Parisi formula for spin glasses.

The release format was shaped with outside recommendations in mind. OpenAI said it set its disclosure principles after consulting the Advisory Group on Mathematics and Artificial Intelligence, an independent body at the Institute for Advanced Study in Princeton. METAL has reported that OpenAI announced the launch of this advisory group. According to reports, the advisory group issued recommendations on September 29 based on more than 600 responses from the math community, asking labs to disclose the model name, prompts, reasoning summary, time taken and compute cost for each result and to deposit results in scholarly repositories that no AI lab controls. On the practice of testing advanced math problems on proprietary models, the recommendations stated: "We do not endorse this practice, and we ask them to stop."

오픈AI가 모델 추론 요약을 공개한 10개 결과군 번호와 주제 목록 표

This release followed only some of those requests. OpenAI provided reasoning summaries, the number of attempted problems, compute estimates and a revision policy that keeps earlier versions. The model's name and per-result prompts were left out, and cost was given only as an average rather than per result. The repository is also a GitHub repository managed by OpenAI, and the company said it would keep looking for community-hosted repositories that meet the advisory group's standards. According to reports, OpenAI drew a line under which the advisory group can advise on how results are communicated but has no say over whether or how quickly results are produced.

The same model was at the center of a dispute a month earlier. On September 8, OpenAI announced that the same internal system had proved a finite-time singularity for the Navier–Stokes equations, and according to reports it said at the time that the model was significantly more capable than GPT-6 Astra. METAL has reported that the announcement set off a priority dispute. After NYU mathematician Tristan Buckmaster alleged that OpenAI had followed his research direction, OpenAI mathematician Sébastien Bubeck countered at a press conference: "We did not use their prompts or proofs to prompt our models or direct our agents."

Wariness in the math community has grown since then. Fields Medal-winning mathematician Terence Tao wrote on Mastodon at the time: "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential." METAL has reported that 25 Fields Medalists issued an open letter saying the goals of AI companies and mathematics have diverged. According to reports, Fields Medalist Timothy Gowers, a member of the advisory group, has warned that within one to two decades the mathematical literature could grow enormously while no human community remains that truly understands it.

OpenAI put money and the model on the table as next steps. It said it would fund workshops, conferences and special programs aimed at understanding major results produced by AI, and that it is working to responsibly release the model that produced these results. The company said it wants to "directly empower scientists with state-of-the-art capabilities." It did not disclose the size of the funding, a timeline or when the model would be released.

The weight of this announcement lies less in the number of papers than in the order of verification. Lean formalization can confirm by machine that a proof is logically correct, but it says nothing about how new or important a result is. Sorting out which of the 722 papers are genuine breakthroughs and which need repair now falls to academia. The gap in speed between the side producing results and the side that has to read them is set to become the agenda for mathematics.

Comments