Writing mathematics in the age of AI

Martin Hairer Professor of Pure Mathematics at EPFL and Imperial College London

With this post, I would like to start a discussion about best practices for writing a mathematics article in which some (or possibly all) of the results were obtained with the help of AI. (I will use ‘AI’ or ‘AI system’ as a generic term for LLM harnesses like Codex, Fable or similar tools.)

The Leiden declaration provides a good baseline ethical framework for this. However, it addresses many different issues and tends to stay at a rather general abstract level. Furthermore, it is not always clear how to implement its recommendations in practice. Here, I want to focus specifically on the aspect of writing an article and to provide concrete recommendations that working mathematicians can use straight away. I am intentionally not addressing the question of how to use AI in order to produce mathematical results, first because I do not feel qualified for this and second because this is something that is still evolving from month to month.

This post is by no means intended to be definitive or exhaustive. This new era is still in its infancy and we, as a community, are still in the process of figuring out what should and shouldn’t be considered best practice. Instead, this is intended as a personal take and as the start of a conversation. I hope that the comment sections for this post will be useful for us to try to find some consensus as a community.

Mathematical content

In my view, the most dangerous trap that a mathematician risks falling into is to use an argument provided by an AI system without fully understanding it and therefore running the risk of it being wrong. We have all been guilty of providing proofs “by intimidation” where a step is “easy to see” or “immediate”, when it really isn’t to all but a small number of subject experts. Unfortunately, AI turns out to be very good at taking advantage of these rhetorical devices to try to fool us into believing that an argument is sounder than it really is in order to collect the “reward” of having solved the problem it was given.

While we all make honest mistakes, the mathematical community is sufficiently small so that genuine dishonesty (trying to fool the community into believing that a result was proven when it wasn’t) exacts a very steep career-destroying price and is therefore vanishingly rare. As a result, we are not used to be on the lookout for deceptive arguments. Because of this, it is fair that an argument produced by AI should be subject to more scrutiny (and more adversarial proof-reading) than one produced by a trusted colleague.

One good thing is that an AI will not get annoyed at its interlocutor or mock their lack of knowledge. If there is any doubt whatsoever about the validity of a step in its reasoning, one should therefore have no qualms about asking repeatedly for clarification until a satisfactory argument is produced.

Disclosure

In line with the Leiden declaration, I believe that if the use of AI contributed meaningfully to the mathematical content of the article, then it should be disclosed. It is very important to note that this is not so that the author can be ‘punished’ by having their result less valued. In fact, it is important precisely in order to create a culture where articles are judged on their merits, independently of the tools being used to produce them. It is also important to flag the use of AI to the reader so they are aware of the potential pitfalls discussed in the previous section.

The form of the disclosure should be as informative as possible, so “The proof of Lemma 3.5 was obtained by ChatGPT” or “The idea of considering the widget of Proposition 4 as a multi-faceted smurf was suggested by Fable” is much more helpful than a boilerplate “AI was used in the preparation of this article”. I slightly disagree with the Leiden declaration regarding the disclosure of prompts, thinking traces, and energy and token usage for the simple pragmatic reason that these are not always available and the output typically isn’t reproducible anyway. If they are available though, providing a link to the thinking trace can be very valuable, especially if a sophisticated harness was used.

On the other hand, I do not believe that it is particularly useful to disclose the use of AI if it was only used for menial tasks like searching for references, proofreading a draft of the paper for mathematical or linguistic typos, searching for grammatical mistakes, etc. After all, there is no expectation to disclose the use of search engines, spellcheckers, etc that have long been used for such tasks, it is simply expected that the authors would have made use of such tools.

Attribution

A major problem with AI usage is that it appears to be extremely good at putting together several mathematical ideas in order to solve a problem, but quite bad at providing references to the articles where these ideas originated in the first place. It is somewhat ironic though that, when asked specifically about it, AI system do actually tend to be very good at finding such references.

If AI output contains what appears to be a new idea, it is therefore very important to try to get to the bottom of where it actually comes from. In many cases, this can be done by simply copying the output into a new context, together with a prompt along the lines of: “Someone submitted the following mathematical argument, but references seem to be missing (or incomplete). Could you find out where these ideas first appeared in the mathematical literature?”

If you are developing a harness geared towards mathematical research, please do seriously consider including such a step, even if it increases token usage.

Writing style

Here is a personal plea: please do not copy / paste the output of an AI system into your article. Instead, for the reasons discussed i the first section, digest the argument and then explain it in your own words. Some reasons for this are:

  • Often, AI systems have a somewhat meandering style that focuses on following an argument step by step and loses track of the big picture. The reader however is typically interested primarily in the big picture and the way in which the main ideas interlock.
  • In the process of digesting the argument, you may well realise that it can be presented in a clearer and simpler way.
  • The reader is interested in your thoughts. Do not insult them by spending less time to write a paragraph than it takes them to read it.
  • It is jarring for the reader to have a significant break in style between the human-written parts and the AI written parts of the paper. Even for technical writing, I personally find the AI writing style to be grating and soulless.
  • In my experience, AI systems tend to be more verbose than necessary and spend a lot of time explaining in detail relatively basic material while glossing over more subtle points. Good technical writing should to the precise opposite.
  • It is easy to overlook parts where its notations differs from those of the rest of the paper.
  • It makes it easy to overlook logical inconsistencies (see also the first section).

A corollary of this is that I would strongly discourage the use of AI to provide “filling material”. I would rather read a dense but well-written paragraph than a full page containing the same information, interspersed with buzzwords and empty phrases.

Conclusion and disclaimer

This post should not be construed as any opinion regarding the ethics of AI usage, whether in mathematical research or otherwise. Some mathematicians have taken public positions against the use of AI while some have gone all in. I believe that positions are defensible in good faith (albeit starting with quite different assumptions and priorities), but this is a debate that should be reserved for a different post, possibly also on this site.

4 responses to “Writing mathematics in the age of AI”

  1. […] Analysing the CoT can also give more insight into the lineage of new ideas, and to avoid plagiarism and missing citations (even if accidental). Sometimes, models aren’t that effective at doing this by themselves, as touched upon in Martin Hairer’s blog post. […]

  2. lovingboldly29683bb2c8 Avatar
    lovingboldly29683bb2c8

    “It is very important to note that this is not so that the author can be ‘punished’ by having their result less valued.”

    This claim on “evaluation being neutral to AI use” is an important statement in your post. It raises two delicate questions: how that can be achieved, and why it is a good idea. Spoiler: I believe that the answers are: it can’t, and it’s not.

    How? Evaluation of mathematical research is not a centralized process where universal criteria can be easily established. It is a diffuse responsibility shared by journals, departments, funding agencies, hiring committees… Many institutions will issue their own reccomendations on evaluation of AI generated mathematics, and it is unlikely that they will all coincide. Even assuming this will be the case, evaluation is ultimately performed by people with personal convinctions and unconscious biases. I do not believe that it is realistic to expect that AI use will not be taken into account in eavluations in the foreseeable future.

    Why? There are plenty of good reason why AI generated mathematics might be evaluated differently in some situations. No matter the direction we take, there will be a long transition phase in which access to advanced models is highly impaired. Many papers from pre AI age will be under review, “competing” with a flood of AI generated ones. For many purposes (grants, hiring..), evaluation of a mathematicians record track is used as a projection on their future achievements. Even after the transition is over, there are good reason to evaluate differently a result obtained with or without the help of AI. The friction and timeframe associated to the traditional proving process have side effects on the community that are difficult to measure. A departement where most faculties make heavy use of AI might feel a legitimate need to hire a profile more inclined in the opposite direction.

    The only solid reason that I see to ignore AI use for evaluation purposes, is to avoid creating an incentive not to disclose it. It is a good pragmatic reason, but a very bad reason in principle. Ethical behavior should not depend on how easy it is to enforce it. AI use should be disclosed for ethical reasons, because it is an important piece of information, and that should be enough.

    I see one (very partial) solution to this caveat: journals might ask to sistematically include an AI use statement, even to merely state that no AI was used at all. Of course people might still lie. But lying in print is very different from omitting information. An explicit lie would be unanimosly considered an ethical and sciencific fraud. That can be a deterrent for many people.

    A corollary of this is that I disagree with your point that cosmetic AI uses (reference search, typo and math checking..) need not be disclosed. Disclosing them is useful because it carries information on which AI uses where *not* made, avoiding grey areas.

    For now, I have started to include systematically an AI use statement even in the case of no use or minor use. I encourage everyone to do the same.

    1. AS Avatar

      The solid reason to ignore AI use for evaluation purposes is to prevent reviewers’ prejudices from influencing the process. If a reviewer feels “insulted” by the mere fact that a result was obtained with partial use of AI, valuable and impactful work will not get published.

      The reality is that the publication process is already subjective, with plenty of conscious and unconscious biases. We don’t need to add more.

      Please leave considerations of grants and hiring for grants and hiring committees. Economic considerations are better addressed elsewhere. The scientific record belongs to the public as a whole, and the public is best served when every impactful result gets published.

  3. Alexis Marchand Avatar

    I would like to add one point to the paragraph on disclosure: I think that any paper relying on large language models should also include a conflict of interest statement. For example, if the authors were given free access to a certain LLM model, or if they received funding from an LLM company, or even more, if they are employees or shareholders of an LLM company, then I think that this should be disclosed. While papers including a statement on LLM usage are getting more and more common, I have not seen any such declaration. This is not part of our culture (yet), but I think that this needs to change given the economic stakes.

Leave a Reply

New posts by email.

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading