A mathematician sits down to work on a problem. She knows the rough outline of the proof. First, she gets her LLM to check the literature to make sure that she hasn’t been scooped recently, then feeds the intermediate lemmas into the LLM to get its opinion on whether the general shape of the proof is correct. For the third lemma, it looks pretty general, so she asks the LLM if it’s been proved before, and it returns a paper from the 1990s with something very similar, but not exactly, the same; the proof is an easy adaptation of the cited literature, and it saves her days of checking. The LLM writes a quick program to check a few examples and finds a counterexample to one of the steps, and she saves weeks of trying to prove that false result. Her program to close the final result is taking a long time, so she puts it into her LLM, which uses standard algorithms from computer science to speed it up and it finishes. The theorem is proved.
This is a benign version of how LLMs could help speed up mathematics. Every line of the proof is human, the strategy is human, and the paper is human. But with the LLM as a helper, it was done in half the time. Even if we ignore the more problematic generation of AI papers, in whole or in part, there is a potential productivity boost here. Previous rounds of improvements in the way mathematicians work have been absorbed. Older mathematicians working in the pre-digitization age would try to find a reference by picking up the 1954 copy of Mathematical Reviews, looking through it to find the papers written by the author they know wrote a paper, then going back and picking up the 1955 copy instead. The arXiv and search engines suddenly made it a lot easier to find pre-existing literature, and to read it outside of the university library, and indeed outside of the university system.
This production leap is much larger than those because pre-LLMs, we were still constrained by the speed of thought and typing. Another bottleneck in mathematics creation has been removed. What does that mean for the community?
Après Mai, le déluge
Around June 2026, commercial-tier LLMs started to become useful in a mathematical setting. Their utility has only become more apparent over the months that followed, and we are now at a time when I can press a button and receive a respectable 20-page paper on a minor problem in my area. I would have to read it to check if it is true, but with improvements to the models and an understanding of how to prompt it correctly, my trust in AI-generated and human-generated mathematics are converging. In a given day, I could probably create two papers on minor, but useful, areas of my field, and read them carefully enough that I would believe any errors in them are quite subtle.
I am obviously not the only person that can do this. In fact, everyone can do this. So the number, and indeed quality, of AI-assisted, or fully AI-generated papers will increase day by day, flooding the arXiv, and then a few weeks later journals, with submissions. And in a few months with further improvements, the submissions will be on reasonable problems with plausible mathematics.
We can already see the impact on the arXiv, with papers in the mathematics subfield going from 4040 in January 2026 to 6307 in August 2026. September is on track to be another high-output month.

(Naturally I improved my productivity by having AI produce that image.)
The current journal system cannot absorb a 50% increase in the number of submissions. And this is just the start of the revolution, because adoption is not instant. It’s not unreasonable to suggest that January 2027 will have more than double the submissions of January 2026.
What is the point of journals?
Why do I submit my articles to journals? Primarily because I need to for my job. Were I retired, I would still want to publish papers, but I would have a lot less incentive to do so in peer-reviewed journals. Historically, journals were the only way to easily get mathematical ideas out into the wider world, but with the advent of the Internet, the arXiv, and institutional repositories that is no longer the case. Indeed, so divorced is mainstream mathematics from journals now that the arXiv has become the primary means of distribution of results.
The journal serves two purposes: first, a human has (hopefully!) read your paper and made suggestions for improvement; second, your paper has been deemed of sufficient correctness and importance to warrant a particular stamp. The best stamps, festooned with gold leaf, say “Journal of the American Mathematical Society” or “Annals of Mathematics” on them. The cachet of the stamp changes depending on the journal, but a stampless paper is, in eyes of the people who matter for a working mathematician’s career funders, promotion committees, REF panel members in the UK a worthless paper.
The imprimatur of journal publication is a benefit that accrues to, and is necessary for, human mathematicians. A graduating PhD student or postdoc needs publications to get a job. Publish or perish is not a motto but a reality under the current system.
OpenAI is not a human and does not need to get a job. It needs funding, and got $122bn of it in March 2026. Its motivations are different; a press release is far better for it than a journal publication. Since LLMs don’t need the journal credit, and humans do, why should the output of LLMs be submitted to a journal at all?
You referee my paper, I’ll referee yours
For referees the situation is bleaker still. Every paper needs a referee, but the explosion in papers has not resulted in a secondary explosion of qualified people who can be referees. Perversely, the explosion in AI use will make things worse for the most qualified people. How does an editor choose a referee? He picks either from people he knows personally or does a quick Google search to try to find people who have published in the area, particularly if the paper is in an area he is not familiar with. The pool in the first case will specifically exclude people considered to be LLM factories, as although they are producing mathematics, and indeed worthwhile mathematics, they will not be considered capable of refereeing the work in the eyes of the editor. In the second, if the editor looks at the papers and sees they are AI-generated he will hopefully exclude that person as a potential referee, or worse, will not know, and someone who is not an expert will be asked to referee a paper. The better of these two scenarios results in human experts being completely overwhelmed with referee requests. They will either accept too many or start rejecting at an even higher rate than before the AI explosion.
It also exposes a flaw in the refereeing process. At the moment referees perform a generally thankless task because it is necessary. I want my papers to be refereed, so I will referee other people’s, and indeed slightly more because early-career mathematicians will generally not be asked to referee as frequently. LLMs are not considered to be referees, so the reciprocity that comes from refereeing is broken. Why should I referee your AI-generated paper?
It gets worse for anyone who agrees to referee an AI-generated paper. Suppose I read one, and find a fatal issue on page 10. I set out my objections and respond to the editor, who passes on my comments. The next day the paper comes back, completely different, with an extra ten pages of argument and a set of computer programs in three different languages, some of which appear just to check the first 1000 cases of a result proved in full generality in the text. Humans would take a few months to make a local correction, keeping the global structure the same. AI might do that, or decide to completely alter the structure of the article. Either way, it’s back on my desk the very next day, longer and with more parenthetical remarks about how certain approaches don’t work, or how a lemma was proved using exact arithmetic rather than floating-point calculations or estimates.
Humans only beyond this point
It therefore seems to make sense that journals adopt the policy: “No AI-generated papers”. I would go further, requiring papers to have a substantial input from a human. I certainly want to allow proof strategies like the Four Colour Theorem, for example, where a human reduction and LLM case checking proves the result. There is a human guiding mind, the human reduction, and a machine is deployed to clean up the resulting mess. In the Four Colour Theorem it was a computer program that did that. If an LLM had written the actual code instead of Appel and Haken, that would not have changed the result. They are mathematicians, not coders: the mathematics is what is being considered, not the programs.
Of course there would need to be judgement calls, but that is literally an editor’s job. Completely AI-generated papers should be instantly rejected. Papers where the authors (plaintively?) certify that they have checked the mathematics should be rejected, because they are authorless. Refereeing a paper does not confer ownership. Mathematicians often have work published posthumously, but the person who submits the paper on their behalf does not become a coauthor, even if they have checked it. Papers where AI performed much of the mathematics but under the guiding instruction of a human are trickier, but for those the ‘author’ would have to make a strong case for why the paper should be considered the product of a human mind.
For those journals that do not follow this prescription, the prognosis is poor. Imagine a situation where the number of papers has doubled, but half of the journals are refusing AI-generated papers. What happens to those that remain?

There are other coping strategies that could be deployed, such as rate-limiting authors. An author submits a paper to a journal and is on ‘cooldown’, unable to submit again for a set period of time. Such a strategy is used by the European Research Council to prevent being flooded with submissions. The cooldown can be altered: the paper being accepted turns off the cooldown, whereas having significant flaws (rather than being rejected for significance reasons) could increase the cooldown as the author clearly is producing substandard work. Submitting AI-generated slop should turn the cooldown dial to the maximum setting, and in the most egregious cases could result in a ban.
People don’t have a human right to be published in a journal. And submitting AI-generated mathematics with your name on the top, pretending you had any contribution, is tantamount to academic fraud.
Sorry I can’t referee that article, I’m washing my hair for the next three months
Referees will need to be blunt as well. Even the benign scenario I outlined at the start, where the human still does the mathematics but is just faster now, the bar for significance is now raised, and the requests will increase substantially. As a referee, my personal rules are to outright refuse to look at mathematics that looks AI-generated. This applies to the paper, because I don’t particularly want to read AI-generated text. I will then run the paper through an AI to see what it says about it. Some journals expressly forbid me from doing this, and to those I will simply refuse to referee. Refereeing is not a labour of love, but an unpaid chore. Sometimes I learn something interesting from a paper, but I am unlikely to learn anything from a paper that can be torn to shreds by the latest AI model. If the paper passes muster from an AI inspection, then I will decide if it’s significant, and then I will read it to decide if I believe it.
My time is finite, and LLMs offer a way to shorten the refereeing process. If the LLM says that Lemma 3.4 has a counterexample, and gives me the counterexample, then I don’t need to read any of the paper. Refereeing is not a creative endeavour that we need to maintain as the preserve of humanity. Anything to make it more efficient should be done. The most obvious example is to have an LLM check every citation to make sure that it actually says what the author claims it says.
But this should be done before it lands on my desk. Before submitting any paper, authors should subject it to an LLM refereeing. It’s cheap, quick, and will catch many typos and problems that a referee picks up. You can ask it just to flag issues rather than try to correct them if you wish to preserve the sanctity of your work (and if your work up until then was LLM-free then you probably should). This level of LLM usage should be encouraged, not disavowed.
Things are going to have to change around here
Finally, the incentive structure for mathematicians is going to have to change. It’s trivially easy to lie. I can get AI to generate a paper, read it and understand it, and then create it myself. As long as it’s plausible that I wrote it, and the writing itself is not AI-generated, it will be incredibly hard to detect. If my career depended on producing papers, that sort of temptation would be incredibly hard to resist, especially once it becomes commonplace among a researcher’s competitors. Worse for the brightest, any exceptional PhD student is going to instantly face allegations of LLM use when she submits a great paper. In the near future, if LLMs surpass humans at production of mathematical structure, rather than calculational mathematics that it now specializes in, how are young mathematicians to be judged?
In the UK the first and most obvious target is REF 2029. This octennial (roughly) exercise asks universities to submit the best papers of their researchers, which are then graded and the institutions rewarded points for the best research. And by points I mean money. Can it survive the AI explosion unscathed?
Received 1 September 2026, revised 15 September 2026.
Add to the discussion