Crab Research

Event discussion

When a machine gives the answer first: correctness, transmission, and research ethics around Navier–Stokes

This article begins with OpenAI’s week-long Navier–Stokes release, follows the Alpöge–Buckmaster timeline and the joint declaration, and then separates formal correctness, knowledge transmission, and research process. Our view is that a reliably checked conclusion can first have a clear scope for use, while understanding, explanation, and transmission remain equally important contributions that can develop in parallel.

It began with an answer that appeared almost overnight

Primary materials and further reading · Joint statement by 25 Fields Medalists · Terry Tao on the Alpöge–Buckmaster results · OpenAI’s Navier–Stokes account · Tristan Buckmaster’s public statement

The Navier–Stokes equations describe the motion of fluids such as water and air. Whether smooth solutions of the three-dimensional incompressible equations remain smooth is one of the seven Millennium Prize Problems. It belongs to the theory of partial differential equations, yet it also touches aircraft design, weather prediction, and the study of blood flow. A finite-time singularity is a loss of regularity, such as an unbounded velocity, that develops from initially smooth data in finite time. Viscosity tends to damp this instability, while nonlinear interactions can concentrate energy at ever smaller scales; the tension between them is what makes the problem so difficult.

In the first week of September 2026, OpenAI brought this problem into a public contest over AI mathematical ability. According to the company’s account, an internal model under training began showing unusually strong mathematical performance on August 28. On September 1, the team heard reports that two Millennium Problems might have been solved and began testing the system on open problems and other high-impact questions. Nearly one hundred agents worked in parallel for about fifty hours. Their first result concerned the unforced Euler equations, an inviscid model of fluid motion, and the team then shifted its resources toward Navier–Stokes.

OpenAI says that the agents reached an analytic answer for Navier–Stokes on September 5, about 88 hours after the effort began. The team then spent a reported 17 hours in total formalizing and checking the result in Lean, and released a paper and code on September 8. Lean is an interactive theorem prover. A formal proof presents its proposition, dependencies, and checking rules in a form that a trusted kernel can check again. OpenAI described the work as solving the Millennium problem within its stated formulation and presented the release as a public record of model capability and research speed.

On that timeline, the news is easy to understand: a long-standing problem had produced a new result that others could continue to check, and formal material was released with the paper. The controversy entered when a second research timeline reached the public record in the same week.

A second research timeline entered the public record

Tristan Buckmaster and Antonin Alpöge had already been working on fluid equations for a long time. In a public statement, Buckmaster said that on August 15 they had made progress on the Boussinesq and Euler equations with smooth forcing, and that months of research conversations, prompts, and materials had been kept in Codex sessions. For a reader, those sessions are best understood as a workbench for a research process: they record how questions were posed and how lines of attack were tested.

On September 3, reports began circulating that a large language model might have solved a major open problem. Buckmaster then contacted OpenAI. In two calls on September 6, he asked whether the internal system had accessed the unpublished work he and Alpöge had left in Codex during the previous two months, or whether those materials had been used for training. The document he published records the questions and the calls. On September 7, he and Alpöge released three finite-time blow-up results, for the incompressible porous-medium equation, the two-dimensional Boussinesq equation, and the three-dimensional incompressible Euler equation, together with Lean formalizations. These results concern neighboring equations in the same line of fluid-singularity research; the early release placed their work and its timestamp in the same public record.

The public document also records a dispute over priority and authorship, including whether Anthropic employee Levent Alpöge should be included in a joint release. In an update on September 10, OpenAI said that Buckmaster’s Codex prompts from the previous two months could not have affected its internal system or training, and that its researchers and agents had not seen the other results before their release. The company also said that, after completing the project and Lean verification, it had hoped to arrange a parallel release and acknowledged the other researchers’ priority for the forced Euler result.

The public accounts differ on data isolation, priority, and authorship. Deciding between them requires access logs, training-data provenance, communications, version timestamps, and audit material that an independent party can inspect. The dispute therefore reaches beyond the proof itself to the question of whether researchers can preserve, explain, and publish their work under fair conditions.

Tao’s presentation and the joint declaration widened the setting

On September 7, Terry Tao introduced the Alpöge–Buckmaster work as an exciting development and explained how its three finite-time blow-up results relate to the unforced Navier–Stokes problem. He also noted that the first proof texts were extremely difficult to read and that the authors were rewriting them as a professional paper. Solving a problem, understanding the structure of a proof, and extracting methods that others can reuse do not happen on the same timetable.

On September 11, Tao published the joint statement titled “A Severe Misalignment of AI in Mathematics,” initially signed by 25 Fields Medalists. The statement places conceptual understanding and mathematical insight at the center of what the community values. It describes problem solving as a tool or proxy for reaching those goals. It also criticizes the use of famous open problems as AI benchmarks, warning that competition can push rapid announcements ahead of careful writing, method extraction, citation of prior work, and integration into shared mathematical knowledge.

The statement appeared three days after OpenAI released its paper, so it naturally became part of the same public conversation. Its text addresses a broader pattern of AI practice and does not issue a mathematical ruling on any particular paper. Its central causal question nevertheless meets this event directly: when an answer is presented as a public measure of model ability, can speed and spectacle outrun the time needed for explanation, attribution, and communal understanding? The later analysis therefore has to separate mathematics, transmission, and research process.

Seeing the episode through three dimensions

The timeline above shows how the episode unfolded. The next step is to change the angle of view. When a mathematical result enters public discussion, it has three dimensions at once. Keeping them distinct makes the declaration easier to understand and shows which parts of the dispute require different evidence.

The first dimension is correctness. It asks whether an AI-generated proposition holds within its stated assumptions and scope, and whether its proof or formal object can be checked independently. There is an objective side and a human-recognition side. A proposition may already hold while readers still need time to read the paper, inspect its dependencies, reproduce its experiments, or rerun its formal code. The result’s validity and the community’s time to establish that validity are two stages of the same work.

The second dimension is knowledge transmission. A checked result still has to become knowledge that others can read, cite, teach, and develop. That requires explaining where the problem came from, locating prior work, showing which structures actually carry the proof, marking the scope of application, and making later reuse possible. This dimension asks whether a result can enter shared knowledge and whether people can draw structure and methods from it.

The third dimension is ethics. It asks how the result was produced and released, and whether the process respected researchers and shared rules. Were unpublished materials accessed or used for training? Were contributions and authorship represented accurately? How was priority handled? Did the people involved have a fair chance to publish? These questions call for access records, data provenance, communications, timestamps, and independent audit. That evidence is different from the evidence used to assess a mathematical proposition.

The joint declaration speaks mainly to the first two dimensions. It places conceptual understanding and mathematical insight at the center, and worries that using famous open problems as benchmarks can push the speed of an answer ahead of the work of transmitting knowledge. Its subject is the relationship between obtaining an answer and turning that answer into shared mathematical knowledge. The data-isolation, priority, and authorship dispute in this episode belongs to the third dimension; the declaration’s emphasis on understanding cannot settle it by itself.

With these three dimensions in view, the analysis can proceed in order: what OpenAI’s published paper and formal materials contribute to the correctness question, why knowledge transmission needs its own time and work, and how a valuable result can coexist with questions about the propriety of the research process.

First question: within what scope is the formal result correct?

Formalization turns the propositions, variables, assumptions, dependencies, and proof written in natural language into objects that a checker can process. If the paper’s claims correspond line by line to the formal proposition, the dependencies are public, the kernel accepts the proof term, and others can reproduce the check, the conclusion has a definite correctness boundary within its stated scope. Whether it reveals a new structure, reads well, or belongs in a textbook requires different evidence.

SAT gives a small but powerful comparison. For a fixed Boolean instance, a solver can provide an assignment when it claims satisfiability, or a certificate that an independent checker can verify when it claims unsatisfiability. Once the certificate passes, an engineering system can use the result within that scope to rule out a design, confirm constraints, or continue a computation. A user can rely on the certificate without understanding every path searched by the solver. The result can be useful before the certificate has been distilled into an elegant theory.

A formalized AI proof follows the same logic at a richer level: it supplies an exact proposition, a proof object, dependencies, and a kernel check. If the natural-language claim corresponds faithfully to the formal proposition, no dependency is hidden, and the check can be reproduced from the published foundations, the result has usable correctness within that scope. It may still take humans a long time to read. Structural explanation, method development, and textbook treatment may come later; their delay does not erase the formal evidence already established.

Correctness and understanding are two contributions that can develop in parallel. A formalized result should not lose all scientific value merely because it has not become a textbook within days. If a result with the same scope and formal status were produced by a human, the community would normally allow months or years for checking and digestion. Requiring an AI-origin result to provide a complete conceptual account within days, and denying recognition otherwise, places the source above the work and creates a double standard. Criticism of an unfaithful formalization, a missing dependency, an unreadable paper, or inaccurate attribution can all be justified; the criticism should identify the concrete defect.

We discuss elsewhere how a formal conclusion with a clear scope might become a callable mathematical component. Here the relevant point is simpler: a result needs a checkable boundary before anyone can discuss stable reuse.

More on callable formal mathematical components

Second question: how does knowledge get transmitted?

The joint statement is right to emphasize careful writing, explanation, accurate attribution, teaching, and later research. Those practices are the long-term foundation of a mathematical community. A result enters shared knowledge as it reaches the literature, the classroom, and subsequent work where others can understand and extend it. Authors remain responsible for that path after they release a result.

Knowledge transmission is an independent contribution and the process by which results enter shared knowledge. A reliably checked conclusion can be cited, reproduced, and used within a clearly stated scope while mathematicians continue the work of explaining its proof structure, improving its exposition, and exploring possible generalizations. Mathematics has a function of leaving knowledge that people can inherit, and another function of leaving components with clear boundaries that later work can reuse reliably. The two functions can strengthen each other and can mature at different speeds.

The history of the Poincaré Conjecture gives a concrete timescale. Perelman posted the first of his relevant preprints on November 11, 2002; two further installments appeared in March and July 2003. About eight months passed before the initial set of materials was in place. Verification, exposition, and consolidation by Hamilton, Kleiner, Lott, Tian, and others continued afterward. By 2006, roughly three and a half years after the first preprint, the proof had gradually become an understanding on which the community could publicly rely. The fact that it was not textbook-ready in its first days did not remove the mathematical content it had already introduced.

OpenAI released the paper and formal material on September 8. On September 14, the revision date of this article, only six days had passed. Six days is enough for readers to see a paper and code; it is too short for a community to complete every check, explanation, and extraction of structure. We can ask whether the writing is clear and whether related work is properly handled, and we can expect later papers to make the proof easier to read. An unfinished transmission process cannot serve as evidence that the conclusion has no value.

If a human author submitted the same formal result, we would treat understanding and textbook treatment as work that takes time. AI-assisted work deserves the same timescale. Evaluation can move with the evidence: first whether the proposition holds, then whether the explanation is reliable, and finally how the result enters the wider body of mathematical knowledge.

Poincaré timeline sources · Poincaré chronology · Perelman's first preprint

Third question: are mathematical value and research process the same thing?

The mathematical value of a result and the way it was produced and released are different questions. A model company can have commercial goals and still do scientifically useful work; motive does not automatically cancel value. A correct conclusion, in turn, does not grant immunity to problems involving data use, authorship, or priority.

For this event, the questions are concrete. Were unpublished Codex materials accessed or used for training? Were the researchers’ prompts and data isolated from the internal system? Were contributions and authorship represented accurately? After learning that related research existed, was a fair parallel-release arrangement still offered? Who had the power to decide when the work became public? Buckmaster’s document and OpenAI’s update give different accounts. The public record establishes that a dispute exists, but it does not independently settle every fact.

If a company used unpublished research to publish first, or placed researchers at a clear disadvantage over data boundaries or authorship, that would be a failure of research governance even if the final mathematics were correct. If logs, data provenance, timestamps, communications, and independent audit show that the systems remained isolated, researchers would still be entitled to ask how priority and authorship were handled. Mathematical evidence and process evidence answer different questions.

The value of the result and the propriety of the release should therefore be assessed separately. The first requires the paper, formal code, dependencies, and reproduction. The second requires access logs, data provenance, communications, authorship arrangements, and an audit that an independent third party can inspect. Using one set of evidence to replace the other would move the discussion away from the issue at hand.

From result to shared knowledge: a traceable scientific process

A new machine-assisted result can move through a continuous scientific process. Each stage adds a kind of trust and leaves a record that later readers can inspect:

The stages can remain open as the work proceeds; they do not have to wait for one final moment. A formal result can first become a checkable and citable record while transmission continues. Exposition and external review can also improve after publication. Each stage has its own evidence, and public writing should say what is complete and what is still underway.

  • Fix the proposition, scope, formal proof, dependencies, axiom status, and reproducible checking procedure.
  • Turn the formal object into a paper that researchers can read smoothly, with the background, prior work, examples, and proof route in view.
  • Release the paper and code as a public preprint, invite external review, and record questions, revisions, and corrections.
  • Keep the paper, formal code, version timestamps, citations, and public discussion in one traceable record.
  • Let surveys, classrooms, textbooks, and later research extract the structure so that the result becomes understandable as well as reliably reusable.

Read the Proof Engine programme

Let discovery, checking, and transmission keep their own time

This episode separates several judgments that are often folded together: whether a conclusion is correct, whether the community has understood it, and whether the research process respected its participants and public rules. They are related, but each needs its own evidence and its own time.

We will continue to move results along the same chain: fix the proposition and formal boundary, complete a repeatable kernel check, write the proof as a paper that people can read, release a preprint for external review, and preserve revisions, explanations, and later uses in a traceable version record. That keeps the first form of the knowledge while giving it a route into shared understanding and reliable reuse.

Mathematical work can endure in two ways: as knowledge that people can understand, teach, and transmit, and as a component with a clear boundary that later work can call on reliably. Science needs both timescales, and it needs them to meet in one honest record. New technology changes the speed of discovery and checking. The work of a scientific community is to turn that speed into results that can be checked, explained, and handed on.