Gaël Varoquaux

Sat 10 October 2026

←Home

How much plagiarism is there in the mathematical genius of big AI?

Result after result, OpenAI’s AIs are solving famous mathematical problems. But researchers recognize in them their own unfinished lines of work, explored with these very tools: how does an idea go from a conversation to a discovery?

Note

This post was originally published in French as part of my scientific chronicle in Les Echos.


Navier-Stokes, one of the seven “Millennium Problems”, “non-sofic” groups, then, a few days ago, 722 papers in one go: OpenAI is stringing together impressive results that we are still digesting. Where does the intelligence of their AI come from?

A language model is first and foremost an immense memory, which works by analogy. Give it the riddle of the ferryman who must take a wolf, a goat and a cabbage across a river: even if you rename the characters, it manages. Beyond exact matches, it recognizes the structure of a problem it has already seen. In mathematics, this memory is precious: the first difficulty of research is to read, again and again, to find in prior work, close or distant, the tool that will unlock the situation. The AI has read everything.

But the AI does more than adapt known solutions: it breaks problems down into simpler steps, a “chain of thought”, each step drawing on its memory. Knowing how to break things down is learned from examples, like those in school textbooks. To multiply the examples, the AI is made to produce thousands of attempts on problems whose solution can be checked, looking for the sequences of steps that succeed: this is reinforcement learning. By seeing many fruitful sequences, the AI builds itself an implicit library of decomposition strategies, which it transposes to new problems.

Researchers’ conversations with the AI are also full of new sequences of steps on cutting-edge problems. A mathematician working with an assistant guides it, corrects it, steers it toward an unusual path: each session is a successful chain of reasoning, annotated by an expert. These sessions are kept by the AI providers.

Yet researchers recognize their own lines of work in OpenAI’s results. Tristan Buckmaster and Levent Alpöge had been working for a year, with OpenAI’s tools, on an unpublished approach to Navier-Stokes; OpenAI only took it up very recently, with a proof that Buckmaster finds surprisingly close. Andreas Thom had discussed with ChatGPT for months techniques that are out of fashion for this problem, the very ones at the heart of OpenAI’s proof. At Anthropic, Claude has just “discovered” enzymes that a PhD student in Copenhagen says he has been studying since 2022, having entrusted it with his unpublished data. The companies deny it, but OpenAI first conceded that it “cannot rule out” that “de-identified” data improved its models.

We will probably never get to the bottom of this story. But to de-identify an idea is to strip it of its name: precisely what separates inspiration from plagiarism. We accept the tools’ terms without reading them, terms that often allow the use of our data “to improve the service”. When we entrust them with an idea, do the AI giants slip from improvement to plagiarism, and will they be shocked, shocked, to learn where their discoveries came from?


AI chronicles

Find all my AI chronicles here

The goals of these “AI chronicles” is to introduce concepts of AI to a broader public, staying at a very very high level.


Sources

Non-sofic groups (Andreas Thom)

Navier-Stokes (Tristan Buckmaster and Levent Alpöge)

OpenAI’s 722 papers (October 6, 2026)

Enzymes (Anthropic and Mario Rodríguez Mestre)

Learning to reason with reinforcement learning

Go Top