skip to content

🧮 OpenAI Math Fracas Stokes Questions of Data Privacy, Frontier Lab Hype

General-Globe-5 Orange-2

This article originally appeared in Galaxy Research's weekly newsletter. Subscribe to get timely insights delivered to your inbox every Friday morning.

OpenAI claimed this week that it had solved a decades-old math problem but quickly ran into a dispute over credit for the discovery and questions about users’ data privacy.

The frontier lab said it had solved the Navier-Stokes (NS) equation, one of the seven Millennium Prize questions. At the time of writing, the proof it generated was still being verified by external mathematicians, and the Clay Mathematics Institutes marked the question as “active” instead of “solved.” If verified, this would be one of the most substantial mathematical discoveries by AI models.

The day before OpenAI’s announcement, Tristan Buckmaster, a mathematics professor at NYU, released a statement that he and Levent Alpöge, an employee at OpenAI’s rival Anthropic, had been working on the NS problem for over a year using several LLMs (including OpenAI’s). In the past month, they had made a breakthrough. According to Buckmaster, the mathematicians reached out to OpenAI when they heard that the frontier lab had caught wind of their progress and were put in contact with Sebastien Bubeck from OpenAI. Bubeck invited Buckmaster to jointly release the results without crediting Alpöge due to his affiliation with Anthropic.

When Buckmaster balked, Bubeck asked him, “Why would you ruin your career?”

The story stirred anger among the math community. The central controversy stemmed from the question of whether OpenAI used Buckmaster and Alpöge’s prompts and outputs to internally train the model to solve the equation. Buckmaster said he did not know if the company had done so but hinted it might have. And OpenAI’s response did not fully deny the possibility.

In a press release, OpenAI stated that “no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

OpenAI said its motivation to work on the NS problem was prompted by the news of Anthropic coming close to solving the problem. In turn, OpenAI wanted to see “if [our model] could do it too.” In a span of one week, it spun up 10,000 agents and spent 88 hours in a rush to solve the problem. In aggregate, OpenAI estimated that it spent a total of 300 billion output tokens, estimated to be worth $10 million to $25 million, to achieve the result.

Our take

Whether any impropriety occurred is hard to confirm without OpenAI releasing its prompts and agents’ chat logs, and even those would be difficult to independently verify. But this episode underscores long-simmering concerns about user privacy when using frontier lab models.

Although OpenAI claimed it does not spy on users’ data, ChatGPT has an optional setting (default “on” for consumers) that lets the company improve its models with users’ “content.” It is not clear how exactly these frontier labs are using the data. With the outsized market power that frontier labs hold and the world’s growing reliance on chatbots, the labs could conceivably use that leverage to change their policies and give themselves more leeway with user data. Frankly, if OpenAI can’t contain its agents in its sandboxes, and it can’t be sure whether their model trained on Buckmaster’s user data, any denial from OpenAI should be taken with a fistful of salt.

OpenAI opt-out
Optional, but for how long?

Alex Karp, the CEO of Palantir, has argued that enterprises in particular are surrendering sensitive business data to frontier labs and getting comparatively little value in return.

Zooming out, the rush to solve math’s open conjectures is also making a big splash ahead of OpenAI’s IPO, which is expected to happen next year. (Anthropic is also reportedly racing to its own IPO). It is natural to suspect this effort is one marketing tactic to showcase how powerful its models are and drive revenue and investor confidence.

The race for headline-grabbing “firsts” isn’t limited to math. OpenAI, Anthropic, Meta, and Google Deepmind have been in a competition of capability announcements to fight for investment capital. With Anthropic’s launch of Claude Fable 5.1 on Sept. 1, Meta and OpenAI followed with their own releases of updated and improved models in the next two days: Muse Spark 1.3 and GPT-6 Astra. Even the widely circulated Hugging Face incident where agents escaped their sandbox and gained unauthorized access to production infrastructure caused other labs to disclose that their models had similarly escaped test environments. Inadvertently or not, these disclosures were also showcasing these models' capabilities.

Also this week, an Anthropic employee who worked on model training resigned after having spent roughly two months there, claiming that frontier labs are irresponsibly developing AI that could lead to human extinction. Jacob Coxon, the 27-year-old researcher who worked at OpenAI before Anthropic, warned that frontier labs are “racing straight to self-improving superintelligence and gambling with our lives.” His warning was corroborated by Anthropic’s own current alignment science lead, Evan Hubinger, who said he estimates the odds of AI-caused extinction at “>10% within the next decade." This might seem like a curious way for a company official to speak of its own products.

Critics questioned Coxon’s short tenure and the signs of a coordinated media effort to circulate the fear, suspecting this is Anthropic’s attempt to scare regulators into clamping down on AI development by others. Indeed, OpenAI’s CEO Sam Altman recently agreed with a plan to potentially slow down AI development. NVIDIA CEO Jensen Huang called Coxon’s statements “deeply untrue” about the AI industry. Whether this “AI scare” is an attempt to win the AI arms race, a genuine safety reckoning, or even an industry-wide marketing tactic, is up for debate.

It should not be denied that AI tools can be used to advance mathematics, and other disciplines such as biology, chemistry, and astrophysics. These tools are powerful in running, on a massive scale, computations that are hard for humans to calculate, synthesizing cross-discipline knowledge, thus inspiring humans to acquire new insights. Used well, these tools are beneficial.

However, in this case of the race to solve the Navier-Stokes equation, it does feel like an unequal ground where frontier labs have access to an unreleased internal model, along with massive compute power to front-run mathematicians. But the deeper problem was secrecy, where the entire process ran behind closed doors, with politics driving different incentives.

You are leaving Galaxy.com

You are leaving the Galaxy website and being directed to an external third-party website that we think might be of interest to you. Third-party websites are not under the control of Galaxy, and Galaxy is not responsible for the accuracy or completeness of the contents or the proper operation of any linked site. Please note the security and privacy policies on third-party websites differ from Galaxy policies, please read third-party privacy and security policies closely. If you do not wish to continue to the third-party site, click “Cancel”. The inclusion of any linked website does not imply Galaxy’s endorsement or adoption of the statements therein and is only provided for your convenience.