OpenAI's Navier-Stokes proof has a mismatch between human and machine versions

Mathematicians say OpenAI's Lean formalization of its Navier-Stokes proof does not match the natural-language version it published. The mismatch does not necessarily mean the proof is wrong, but it raises doubts about relying on AI to auto-formalize mathematical results. The researchers argue that human review remains necessary for such proofs.
OpenAI released its Navier-Stokes work in two forms: a natural-language proof and a Lean 4 formalization intended for machine checking. Researchers at Cambridge, including Anders Hansen and Fabian Circelli, examined the pair and found they are not equivalent. The discrepancy centers on Lemma 8.6: the human-readable version bounds a quantity by m+4, while the Lean code uses m+5, a weaker condition.
The team says auto-formalization can drift when software seeks compilable code, potentially altering the original argument. OpenAI reportedly acknowledged the mismatch but maintained that neither version is necessarily invalid. The mathematicians stress they are not declaring the natural-language proof right or wrong; their concern is that machine verification alone cannot replace human reading.
The mismatch may affect mathematicians, AI developers, and journals evaluating machine-generated proofs. It could reinforce calls for human review before accepting AI-assisted results, potentially slowing claims of automated discovery. It may also encourage clearer standards for linking natural-language arguments to formal code. For the public, the episode could temper confidence in AI mathematical breakthroughs while highlighting the continued role of expert scrutiny.