OpenAI quit the Caltech Mathathon in a day. The letter that asked for a suspension is still standing
Caltech mathematicians asked organizers to cancel the Mathathon. OpenAI quit its sponsorship within a day. The organizers stayed, and added new rules.

On September 10, a coalition of current and former Caltech mathematicians published an open letter asking the organizers of the Caltech Mathathon to cancel it. The Mathathon is a two-round event starting October 30 in which 100 teams spend 40 hours prompting large language models at open research problems, funded by roughly $2 million in AI credits from OpenAI and Anthropic. That same evening OpenAI withdrew its sponsorship. The organizers replied the same day, tightened their rules, and kept the event running.
A student hackathon, a letter, an exit
The Mathathon started as a student project. By September 10 the organizers had received over a thousand applications from undergraduates, PhD students, postdocs and faculty in the US and abroad. The first round runs 40 hours, October 30 to November 1, and per Business Insider each team receives $20,000 in tokens; teams whose answers judges find "promising and well explained" advance to a six-month verification round with more credits. The organizers list support from OpenAI, Anthropic, DARPA expMath and Cognition, among others.
The open letter treated that design as the problem rather than the answer. Signatories must be Caltech affiliates, including JPL, or researchers at PhD level and above; the letter listed 771 names at publication. Its request is not hedged. After noting that some signatories had already proposed restructuring the event "to unsatisfactory effect," it closes by asking organizers to "give up this harmful enterprise entirely." The letter's softest claim and its sharpest one sit close together: the event "is likely to have destructive impacts for the mathematical community," and AI companies are "engaged in research misconduct," a charge it supports by linking a Scientific American report on OpenAI's August batch of ten AI-generated math results, where experts alleged uncited use of prior work.
What the letter actually claims
The letter's five numbered concerns are not a checklist; they build one argument about where verification labor lands. Undergraduates may lack the expertise to check what their models emit, no competition rule compels arXiv publication, and the checking therefore falls on the community in the weeks after. Forty hours, the letter says, "is not enough time to deeply understand, solve, and communicate a problem." It also names the incentive: AI companies "will take credit for the effort of talented undergrads," while human mathematicians proving the same result may be forced to abort research their approaches illuminate. On careers, it argues that association with Anthropic and OpenAI "could plausibly negatively impact the future reputations of participants."
Behind the five concerns sits a claim about sequence. The major labs, the letter writes, are "engaged in an arms race to declare themselves the first to prove prominent open conjectures," and in announcing results they have skipped the practices that keep a field healthy: sharing methods, mentoring students and postdocs, and taking care not to scoop colleagues "at the last mile." The letter is explicit that the labor of cleaning up afterward is unpaid: after each drop, mathematicians "are compelled to step in to properly verify, disseminate, and sometimes discredit entirely the claimed results."
The $20,000 figure becomes an equity argument rather than a spec. "There is no reasonable future model of mathematical research in which every research mathematician receives 20000 USD in AI credits to prove a result," the letter says. It also objects to a line then on the event site, asking what the role of a mathematician is when AI can solve conjectures faster, calling the claim both motivated and empirically dubious. The organizers have since rewritten the event's public description.
The organizers answer, and keep the event
The response letter concedes ground and holds its line. "We think several of these concerns are valid," the organizers write, and after discussions with faculty advisors they made arXiv publication an explicit requirement of the verification round. That round now runs six months and asks for a preprint, a video presentation and an auto-formalized solution. A separate educational track produces course notes, explainers and more elegant proofs of existing theorems, with prizes the organizers say are comparable to the main track.
They also contest the framing. The primary organizers "are undergraduates, the majority of whom plan to pursue theoretical mathematics," they note, and most applicants are PhD students with publication records. Their own description of the event is deliberately small: "If successful, Mathathon will not be a list of 100 problems with convoluted solutions and occasional Lean verifications. Instead, it will be an experiment in what responsible AI usage for mathematics could look like." Their bluntest concession is the one the letter will test: "We cannot guarantee that all work produced at this event will meet the mathematical community's standards." On corporate influence, they argue the event is itself a venue for critical engagement, with debates and panels between AI critics and supporters planned. They also account for the three proposals they say they received from signatories and how each was answered: the educational track, which they are hosting; the arXiv requirement, which they upgraded to a condition of the verification round; and an opening-ceremony speaking slot for a named signatory, which they welcomed but say they have not assigned and cannot promise as a headline role.
The backdrop that made this land
None of this arrived in a vacuum. On September 8, OpenAI said its agents had solved the Navier-Stokes existence and smoothness problem, a Clay Millennium Prize question (our analysis of the claim). NYU mathematician Tristan Buckmaster then said OpenAI may have drawn on his Codex conversations, and OpenAI said it could not rule out that "de-identified data derived from their usage of our products helped improve our models." On September 10, The Verge reported that TU Dresden group theorist Andreas Thom accused the company of "dishonesty" over his own private ChatGPT exchanges and demanded disclosure of training data. The Clay Mathematics Institute has said it will review the Navier-Stokes situation in detail.
Read against that week, a hackathon offering $2 million in industry credits and a 40-hour clock looks different than it would have a month earlier. The dispute is the Buckmaster question projected onto a student event: who verifies AI mathematics, who gets credited, and who absorbs the cost of the checking.
What would settle it
Little here is adjudicated, and the two documents disagree about what the event even is, a training ground or an advertisement. Three things would sharpen the picture. First, whether Anthropic follows OpenAI out the door. It is the other named sponsor and did not respond to Business Insider's request for comment. OpenAI's statement came from research lead Dan Roberts on X rather than to press: "We recognize that the rapid progress of AI in mathematics is disruptive." Second, whether the organizers restructure again or suspend. They have already moved once under pressure, which suggests the letter's practical demands are not immovable; a second concession would test whether the two sides are negotiating or merely trading statements.
Third, and slowest, whether the event's outputs survive review. The verification round's preprint, video and formalization are checkable artifacts. If the six-month results hold up to referees, the organizers will have built a case the letter cannot answer. If they do not, the letter's central prediction, that verification labor gets displaced onto a community that did not ask for it, will have been demonstrated in public. Slop that nobody checks is easy to ignore. Slop carrying an arXiv identifier is not.


