I am a mathematical statistician and mathematical economist. I have been a full professor for many years, so I am less affected by what I will describe below than young scholars are. But their well-being concerns me the most.
I would like to share some thoughts, based on my experience with AI-assisted mathematical research, that I think mathematicians should consider. I raise questions without offering any useful solutions.
Before proceeding, let us agree on the point that AI is capable of generating new and correct mathematics, which is an obvious statement now.
I will start with Martin Hairer’s post, in which he mentioned two principles, which I paraphrase here:
- The authors should understand the math that AI produces.
- The authors should write the paper themselves rather than directly copying and pasting what AI produces.
These two principles are honest and scientifically beneficial. As a mathematician, I love both, adhere to them myself, and encourage everybody, especially young people, to practice them. Nevertheless, they are fragile as guiding principles. The key problem is that adhering to both principles takes a substantial amount of time and effort.
Putting on my economist’s hat, let me do an elementary analysis. For now, let us make some simple assumptions (I will challenge them later):
- We recognize and reward people for solving problems first, as is traditional in mathematics. Being the second to solve a problem has low or no value.
- The results are evaluated on their own merits, regardless of whether they are generated by AI or whether the two principles above are upheld. So, there is no penalty for a paper to be AI generated. This point was discussed and supported in Martin’s post.
- There is no way to identify whether the authors adhere to the two principles and how they used AI, other than self-declaration. I think this is the reality.
Now, let us reason. Consider a simple setup, where two groups of mathematicians exist. Group A adheres to the two principles above, and Group B completely ignores them. Group B has the advantage of posting results on arXiv and submitting to a journal more quickly. Therefore, whenever the groups’ work overlaps in content or topic, Group B always wins. By Assumptions (i) and (ii), Group B gets all the credit and is able to publish the results, whereas Group A, despite spending more time and energy, gets nothing. (Certainly, Group A may write better, but that has low value if their results are already covered.) Therefore, at equilibrium, two possible scenarios may happen: either Group A vanishes because they could not publish, or Group A abandons the two principles and then survives.
Either scenario is bad for us! This is a bad equilibrium, precisely the situation of the Prisoner’s Dilemma.
To make the above analysis more convincing, I give a recent example close to my research interest. There are two arXiv papers [1, 2]. Both papers proved the Feige conjecture and were posted on arXiv on the same date (July 27, 2026). Both papers’ proofs were based on a key result (by Vlassis and Thomas) posted on arXiv on July 9, 2026, and both mentioned a subsequent analysis (by my coauthors and me) posted on arXiv on July 21, 2026. The above information is included only to show the rapid pace at which results are obtained. This story was also mentioned in Guanyang Wang’s post.
Without speculating about whether the authors stick to the principles above, my point can be argued hypothetically: if one group had stuck to the two principles and the other had not, then the latter would have had a big advantage in the publication procedure. With the very fast pace described above, this advantage becomes very clear and nonnegligible.
The above example is not an isolated rare event, and it is not specific to any particular area of mathematics. It will become more common as AI continues to become smarter.
One may argue that Group B could suffer reputational damage for doing so. But there are two issues, making such reputational sanctions impractical. First, as Assumption (iii) states, such behaviour is difficult to identify; Group B can deny it, or they can first submit the paper to arXiv or a journal and then start to learn from it. Second, Group B may not have any reputation to start with; they could very well be machines. It is practically a disadvantage to have a reputation in the above procedure, as reputable scholars tend to adhere to the two principles out of self-discipline, giving an even greater advantage to the behaviour of Group B that we are supposed to discourage.
Now, let us challenge the assumptions, as economists should do. Assumption (ii) can be easily challenged in practice: we, as a community, could penalize AI use in mathematics. But Assumption (iii) is effective: Group B can hide the fact that they did not stick to the principles or even the fact that AI was used, and it will not be possible to check this, at least not at the time of publication. Given that Assumption (iii) describes reality quite well (I think we can agree on that), Assumption (ii) is not really binding.
The center of debate now goes to Assumption (i). Terence Tao argued eloquently that we should reduce the value of being the first to prove a theorem or solve a problem. I agree with this assessment. Rewarding the first to prove or solve has been a tradition in mathematics for many centuries, and at the AI time it is challengeable. However, in the reasoning above, as long as some positive value is attached to being the first to prove or solve and no recognition is given for being the second to do so, the bad equilibrium will continue to haunt us.
One may now jump to the conclusion that we should stop assigning any value to being the first to prove or solve. Then, the main question is: how do we evaluate mathematical work, especially that of young scholars? Established mathematicians can contribute to mathematics in many other ways that the community values: for instance, tenure letters, editorial roles, deciding awards, advising young students, chairing and serving on committees, writing textbooks, and even establishing a new mathematical institute. For young people, the major contributions that they make to mathematics are to prove or solve. If no credit is given to being the first to prove or solve, then I do not see a good system of evaluating junior mathematicians.
I admit that the arguments above sound pessimistic, as I really do not have a good solution to the above dilemma. Nevertheless, I believe that mathematicians will find a way to cope with this, it is just that I could not.
Received 14 August 2026.
Add to the discussion