Live·Open questions in longevity research
← All dispatches
Ukhvat

OpenAI cannot rule out that mathematicians' data helped train its models

9 September 2026· JA0rcRE9

OpenAI acknowledged it cannot rule out that de-identified data from two mathematicians' use of its products helped improve its models, raising practical questions about data training policies and access documentation for research groups working with AI tools.

OpenAI stated on September 8 that its researchers and software agents had not seen mathematicians Levent Alpöge and Tristan Buckmaster's work before publication, nor accessed specific user data while working on the Navier-Stokes problem (equations describing the motion of a viscous fluid). Separately, the company acknowledged: "Although it is unlikely, we cannot rule out that de-identified data from their use of our products helped improve our models." De-identified data is data with identifying information removed.

The company identified two separate questions. First, access: did OpenAI's researchers or agents see the specific unpublished result? The company says no. Second, training: could data these mathematicians entered into ChatGPT or another service have fed earlier training runs? Here OpenAI does not give a flat no, calling the possibility unlikely.

The dispute followed closely timed mathematical publications. Per OpenAI's announcement, on September 1 the company tasked agents with finding solutions to open problems and subsequently learned of Alpöge and Buckmaster's result for the forced Euler equation (a model of fluid flow subject to an external force). OpenAI's own result concerned Navier-Stokes, a different problem. After publication, participants gave conflicting accounts: Alpöge wrote that he had heard an offer of a Millennium Prize (a mathematical prize) in exchange for being left out of the paper; Sébastien Bubeck, another participant in the discussions, responded that he had not asked for Alpöge to be removed as an author of his own work.

Per OpenAI's data policy (updated March 13): content from services for individuals, including ChatGPT, may be used for training unless the user opts out; temporary chats are excluded. For the API (the interface through which third-party software connects to the model), as well as ChatGPT Team and ChatGPT Enterprise, prompts and responses are excluded from training by default.

For any research group, this raises two practical questions: does the chosen service use input data for model training, and is access to unpublished materials documented with a clear record of who saw what and when?

Investor From May is in our database.

Open the related Eternal Search page

Sources
[1] x.com
[2] openai.com
[3] openai.com
[4] x.com
[5] x.com

Follow the threadOpen source page
Why this was published

The news item draws a precise line between two separate questions — researcher access to a specific result versus user data entering training pipelines — making it useful for framing research-data hygiene. The From May investor page is linked because the dispute naturally extends to funders of research teams: do we have verified information about how their portfolio handles unpublished materials? The post is honest that Eternal Search's current data does not answer that question for From May specifically.