Where will all the papers go?

David Craven, professor at the University of Birmingham

A mathematician sits down to work on a problem. She knows the rough outline of the proof. First, she gets her LLM to check the literature to make sure that she hasn’t been scooped recently, then feeds the intermediate lemmas into the LLM to get its opinion on whether the general shape of the proof is correct. For the third lemma, it looks pretty general, so she asks the LLM if it’s been proved before, and it returns a paper from the 1990s with something very similar, but not exactly, the same; the proof is an easy adaptation of the cited literature, and it saves her days of checking. The LLM writes a quick program to check a few examples and finds a counterexample to one of the steps, and she saves weeks of trying to prove that false result. Her program to close the final result is taking a long time, so she puts it into her LLM, which uses standard algorithms from computer science to speed it up and it finishes. The theorem is proved.

This is a benign version of how LLMs could help speed up mathematics. Every line of the proof is human, the strategy is human, and the paper is human. But with the LLM as a helper, it was done in half the time. Even if we ignore the more problematic generation of AI papers, in whole or in part, there is a potential productivity boost here. Previous rounds of improvements in the way mathematicians work have been absorbed. Older mathematicians working in the pre-digitization age would try to find a reference by picking up the 1954 copy of Mathematical Reviews, looking through it to find the papers written by the author they know wrote a paper, then going back and picking up the 1955 copy instead. The arXiv and search engines suddenly made it a lot easier to find pre-existing literature, and to read it outside of the university library, and indeed outside of the university system.

This production leap is much larger than those because pre-LLMs, we were still constrained by the speed of thought and typing. Another bottleneck in mathematics creation has been removed. What does that mean for the community?

Après Mai, le déluge

Around June 2026, commercial-tier LLMs started to become useful in a mathematical setting. Their utility has only become more apparent over the months that followed, and we are now at a time when I can press a button and receive a respectable 20-page paper on a minor problem in my area. I would have to read it to check if it is true, but with improvements to the models and an understanding of how to prompt it correctly, my trust in AI-generated and human-generated mathematics are converging. In a given day, I could probably create two papers on minor, but useful, areas of my field, and read them carefully enough that I would believe any errors in them are quite subtle.

I am obviously not the only person that can do this. In fact, everyone can do this. So the number, and indeed quality, of AI-assisted, or fully AI-generated papers will increase day by day, flooding the arXiv, and then a few weeks later journals, with submissions. And in a few months with further improvements, the submissions will be on reasonable problems with plausible mathematics.

We can already see the impact on the arXiv, with papers in the mathematics subfield going from 4040 in January 2026 to 6307 in August 2026. September is on track to be another high-output month.

(Naturally I improved my productivity by having AI produce that image.)

The current journal system cannot absorb a 50% increase in the number of submissions. And this is just the start of the revolution, because adoption is not instant. It’s not unreasonable to suggest that January 2027 will have more than double the submissions of January 2026.

What is the point of journals?

Why do I submit my articles to journals? Primarily because I need to for my job. Were I retired, I would still want to publish papers, but I would have a lot less incentive to do so in peer-reviewed journals. Historically, journals were the only way to easily get mathematical ideas out into the wider world, but with the advent of the Internet, the arXiv, and institutional repositories that is no longer the case. Indeed, so divorced is mainstream mathematics from journals now that the arXiv has become the primary means of distribution of results.

The journal serves two purposes: first, a human has (hopefully!) read your paper and made suggestions for improvement; second, your paper has been deemed of sufficient correctness and importance to warrant a particular stamp. The best stamps, festooned with gold leaf, say “Journal of the American Mathematical Society” or “Annals of Mathematics” on them. The cachet of the stamp changes depending on the journal, but a stampless paper is, in eyes of the people who matter for a working mathematician’s career \text{\textendash} funders, promotion committees, REF panel members in the UK \text{\textendash} a worthless paper.

The imprimatur of journal publication is a benefit that accrues to, and is necessary for, human mathematicians. A graduating PhD student or postdoc needs publications to get a job. Publish or perish is not a motto but a reality under the current system.

OpenAI is not a human and does not need to get a job. It needs funding, and got $122bn of it in March 2026. Its motivations are different; a press release is far better for it than a journal publication. Since LLMs don’t need the journal credit, and humans do, why should the output of LLMs be submitted to a journal at all?

You referee my paper, I’ll referee yours

For referees the situation is bleaker still. Every paper needs a referee, but the explosion in papers has not resulted in a secondary explosion of qualified people who can be referees. Perversely, the explosion in AI use will make things worse for the most qualified people. How does an editor choose a referee? He picks either from people he knows personally or does a quick Google search to try to find people who have published in the area, particularly if the paper is in an area he is not familiar with. The pool in the first case will specifically exclude people considered to be LLM factories, as although they are producing mathematics, and indeed worthwhile mathematics, they will not be considered capable of refereeing the work in the eyes of the editor. In the second, if the editor looks at the papers and sees they are AI-generated he will hopefully exclude that person as a potential referee, or worse, will not know, and someone who is not an expert will be asked to referee a paper. The better of these two scenarios results in human experts being completely overwhelmed with referee requests. They will either accept too many or start rejecting at an even higher rate than before the AI explosion.

It also exposes a flaw in the refereeing process. At the moment referees perform a generally thankless task because it is necessary. I want my papers to be refereed, so I will referee other people’s, and indeed slightly more because early-career mathematicians will generally not be asked to referee as frequently. LLMs are not considered to be referees, so the reciprocity that comes from refereeing is broken. Why should I referee your AI-generated paper?

It gets worse for anyone who agrees to referee an AI-generated paper. Suppose I read one, and find a fatal issue on page 10. I set out my objections and respond to the editor, who passes on my comments. The next day the paper comes back, completely different, with an extra ten pages of argument and a set of computer programs in three different languages, some of which appear just to check the first 1000 cases of a result proved in full generality in the text. Humans would take a few months to make a local correction, keeping the global structure the same. AI might do that, or decide to completely alter the structure of the article. Either way, it’s back on my desk the very next day, longer and with more parenthetical remarks about how certain approaches don’t work, or how a lemma was proved using exact arithmetic rather than floating-point calculations or estimates.

Humans only beyond this point

It therefore seems to make sense that journals adopt the policy: “No AI-generated papers”. I would go further, requiring papers to have a substantial input from a human. I certainly want to allow proof strategies like the Four Colour Theorem, for example, where a human reduction and LLM case checking proves the result. There is a human guiding mind, the human reduction, and a machine is deployed to clean up the resulting mess. In the Four Colour Theorem it was a computer program that did that. If an LLM had written the actual code instead of Appel and Haken, that would not have changed the result. They are mathematicians, not coders: the mathematics is what is being considered, not the programs.

Of course there would need to be judgement calls, but that is literally an editor’s job. Completely AI-generated papers should be instantly rejected. Papers where the authors (plaintively?) certify that they have checked the mathematics should be rejected, because they are authorless. Refereeing a paper does not confer ownership. Mathematicians often have work published posthumously, but the person who submits the paper on their behalf does not become a coauthor, even if they have checked it. Papers where AI performed much of the mathematics but under the guiding instruction of a human are trickier, but for those the ‘author’ would have to make a strong case for why the paper should be considered the product of a human mind.

For those journals that do not follow this prescription, the prognosis is poor. Imagine a situation where the number of papers has doubled, but half of the journals are refusing AI-generated papers. What happens to those that remain?

Source: https://xkcd.com/1170/. Licence CC-BY-NC 2.5.

There are other coping strategies that could be deployed, such as rate-limiting authors. An author submits a paper to a journal and is on ‘cooldown’, unable to submit again for a set period of time. Such a strategy is used by the European Research Council to prevent being flooded with submissions. The cooldown can be altered: the paper being accepted turns off the cooldown, whereas having significant flaws (rather than being rejected for significance reasons) could increase the cooldown as the author clearly is producing substandard work. Submitting AI-generated slop should turn the cooldown dial to the maximum setting, and in the most egregious cases could result in a ban.

People don’t have a human right to be published in a journal. And submitting AI-generated mathematics with your name on the top, pretending you had any contribution, is tantamount to academic fraud.

Sorry I can’t referee that article, I’m washing my hair for the next three months

Referees will need to be blunt as well. Even the benign scenario I outlined at the start, where the human still does the mathematics but is just faster now, the bar for significance is now raised, and the requests will increase substantially. As a referee, my personal rules are to outright refuse to look at mathematics that looks AI-generated. This applies to the paper, because I don’t particularly want to read AI-generated text. I will then run the paper through an AI to see what it says about it. Some journals expressly forbid me from doing this, and to those I will simply refuse to referee. Refereeing is not a labour of love, but an unpaid chore. Sometimes I learn something interesting from a paper, but I am unlikely to learn anything from a paper that can be torn to shreds by the latest AI model. If the paper passes muster from an AI inspection, then I will decide if it’s significant, and then I will read it to decide if I believe it.

My time is finite, and LLMs offer a way to shorten the refereeing process. If the LLM says that Lemma 3.4 has a counterexample, and gives me the counterexample, then I don’t need to read any of the paper. Refereeing is not a creative endeavour that we need to maintain as the preserve of humanity. Anything to make it more efficient should be done. The most obvious example is to have an LLM check every citation to make sure that it actually says what the author claims it says.

But this should be done before it lands on my desk. Before submitting any paper, authors should subject it to an LLM refereeing. It’s cheap, quick, and will catch many typos and problems that a referee picks up. You can ask it just to flag issues rather than try to correct them if you wish to preserve the sanctity of your work (and if your work up until then was LLM-free then you probably should). This level of LLM usage should be encouraged, not disavowed.

Things are going to have to change around here

Finally, the incentive structure for mathematicians is going to have to change. It’s trivially easy to lie. I can get AI to generate a paper, read it and understand it, and then create it myself. As long as it’s plausible that I wrote it, and the writing itself is not AI-generated, it will be incredibly hard to detect. If my career depended on producing papers, that sort of temptation would be incredibly hard to resist, especially once it becomes commonplace among a researcher’s competitors. Worse for the brightest, any exceptional PhD student is going to instantly face allegations of LLM use when she submits a great paper. In the near future, if LLMs surpass humans at production of mathematical structure, rather than calculational mathematics that it now specializes in, how are young mathematicians to be judged?

In the UK the first and most obvious target is REF 2029. This octennial (roughly) exercise asks universities to submit the best papers of their researchers, which are then graded and the institutions rewarded points for the best research. And by points I mean money. Can it survive the AI explosion unscathed?


Received 1 September 2026, revised 15 September 2026.

5 responses to “Where will all the papers go?”

  1. bengreen Avatar

    Your last point: I very much hope not.

  2. Anonymous Avatar
    Anonymous

    I think your last paragraph about incentive structures mean that any desire to remove AI usage is unrealistic. Many academics are in extremely precarious life situations, and the incentives to tell a lie for the reward of security are too high. On the other hand, if the result is good enough for a top journal, doesn’t it still contribute to mathematics?

    On the other hand, there can be a large amount of nepotism in the journal structures. There have been many instances I’ve seen of a not so interesting result published in a top journal because an author is a big name.

    I personally think it’s okay to use AI as long as the mathematics is good. If papers can be written more quickly with AI, they can also be reviewed more quickly with AI. Perhaps, as a result of the influx, the standard of journals will increase, but this can only be a good thing.

  3. hopeful_mathematician Avatar
    hopeful_mathematician

    I think we will have to redefine what a research mathematics contribution is, and many people already are. My thoughts thus far have led me to think a few things. My experience is only in a couple areas of mathematics, so obviously there is a risk in generalizing.

    1. I don’t see how a human-track separate from an AI-track is plausible. No one will read a new paper that develops some mathematics they could have obtained by simply prompting an existing LLM, nor would such a paper be a contribution to mathematics (except perhaps in terms of alternative explanations, exposition, etc.). Hence, any attempt to restrict journals to non-AI papers seems pretty futile.

    2. For mathematical and computational papers without external interaction with the world (i.e. no wider science aspect), peer review seems totally pointless. No one will want to review all the papers, but also one no longer needs humans to review the correctness of a paper (it can be reviewed several times approximately correctly by readers pointing an AI at it, trying to use it, or by providing formal proofs). The only role of peer review would be a subjective stamp of approval, but that will be increasingly unreliable and unnecessary as we can just see what was useful post hoc.

    3. I do think some kind of unit of contribution or “nice” mechanism for dissemination is still probably useful, like a paper, but I don’t see why it would be “published” after any kind of editorial process, rather than just being available in a repository.

    I suspect that journals in mathematics and some computational sciences will disappear. We should really spend energy coming up with some good ideas about how we should store, search and communicate mathematical knowledge, accepting that diversity and resilience are a good thing so we may want several ways to store it, view it, etc..

    I have no idea how we will attempt to properly assign credit for new discoveries, or if we should even try. The distinction between dissemination and credit assignment was neatly wrapped together until now but I don’t see how it can be salvaged. This is really distressing in terms of how we have structured our profession, and how we are now unable to give much useful advice to the next generation. Even though this is scary, I don’t see how any attempt to retain anything resembling the status quo will be successful.

    If we want to continue pushing the frontier of mathematics, and I don’t see why we cannot do so with the assistance of AI and the development of structures (including human ones) to assist us, we will simply have to adapt. As another one of the posts on this website said, there is a difference between mathematics and the profession of being a mathematician. Mathematics is not under threat, and on the contrary is about to make explosive progress. The specific role mathematicians and universities will play is, however, a very open question.

    1. Hume Chang Avatar
      Hume Chang

      You are a bit dismissive to the button pusher. Nowadays everything can be prompted, but knowing what to prompt and the act of actually prompting it is not easy. Your contempt reminds me of a history anecdote.

      Columbus was dining with Spanish noblemen who downplayed his discovery of the New World, claiming anyone could have done it. Columbus then challenged them to make an egg stand on its tip. When everyone tried and failed, he tapped the egg on the surface so it stayed upright, proving that hard tasks are only easy after someone shows you how.

      Moreover, out of the myriad of prompts they can do, like generating a modern America themed Elder Scroll 5 short video:
      https://youtu.be/o1_RxGSBBAY?is=5ExAO48N5ZDIOVDR
      they choose to envision and make a reality the beautiful theorem now presented before you. I won’t dismiss it as “no big deal, I can do it too”
      In fact, I don’t think you are able to generate a video like above that really made my day.

  4. Hume C. Avatar
    Hume C.

    It is meaningless to publish. We already know AI scraps chat logs. Given now everyone uses it, AI will know every math results ever existed. Once AI solves a problem, AI internalizes the result. The next time someone asks a related/further question, AI will use the previously obtained result (that it’s without proper citations or credit may be by design, else AI would say “chats with user_id xxxx inspired me to derive this result”.). That AI doesn’t discriminate results by journal stamps on it may actually be an improvement of the current system.

    For a researcher to keep up-to-date, just ask AI agent to curate weekly new results it has learned relevant to the user, based on its understanding of the users research interest.

Add to the discussion

New posts by email.

Prefer a feed reader? Subscribe by RSS.

Also on Mathstodon.

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading