With this post, I would like to start a discussion about best practices for writing a mathematics article in which some (or possibly all) of the results were obtained with the help of AI. (I will use ‘AI’ or ‘AI system’ as a generic term for LLM harnesses like Codex, Fable or similar tools.)
The Leiden declaration provides a good baseline ethical framework for this. However, it addresses many different issues and tends to stay at a rather general abstract level. Furthermore, it is not always clear how to implement its recommendations in practice. Here, I want to focus specifically on the aspect of writing an article and to provide concrete recommendations that working mathematicians can use straight away. I am intentionally not addressing the question of how to use AI in order to produce mathematical results, first because I do not feel qualified for this and second because this is something that is still evolving from month to month.
This post is by no means intended to be definitive or exhaustive. This new era is still in its infancy and we, as a community, are still in the process of figuring out what should and shouldn’t be considered best practice. Instead, this is intended as a personal take and as the start of a conversation. I hope that the comment sections for this post will be useful for us to try to find some consensus as a community.
Mathematical content
In my view, the most dangerous trap that a mathematician risks falling into is to use an argument provided by an AI system without fully understanding it and therefore running the risk of it being wrong. We have all been guilty of providing proofs “by intimidation” where a step is “easy to see” or “immediate”, when it really isn’t to all but a small number of subject experts. Unfortunately, AI turns out to be very good at taking advantage of these rhetorical devices to try to fool us into believing that an argument is sounder than it really is in order to collect the “reward” of having solved the problem it was given.
While we all make honest mistakes, the mathematical community is sufficiently small so that genuine dishonesty (trying to fool the community into believing that a result was proven when it wasn’t) exacts a very steep career-destroying price and is therefore vanishingly rare. As a result, we are not used to be on the lookout for deceptive arguments. Because of this, it is fair that an argument produced by AI should be subject to more scrutiny (and more adversarial proof-reading) than one produced by a trusted colleague.
One good thing is that an AI will not get annoyed at its interlocutor or mock their lack of knowledge. If there is any doubt whatsoever about the validity of a step in its reasoning, one should therefore have no qualms about asking repeatedly for clarification until a satisfactory argument is produced.
Disclosure
In line with the Leiden declaration, I believe that if the use of AI contributed meaningfully to the mathematical content of the article, then it should be disclosed. It is very important to note that this is not so that the author can be ‘punished’ by having their result less valued. In fact, it is important precisely in order to create a culture where articles are judged on their merits, independently of the tools being used to produce them. It is also important to flag the use of AI to the reader so they are aware of the potential pitfalls discussed in the previous section.
The form of the disclosure should be as informative as possible, so “The proof of Lemma 3.5 was obtained by ChatGPT” or “The idea of considering the widget of Proposition 4 as a multi-faceted smurf was suggested by Fable” is much more helpful than a boilerplate “AI was used in the preparation of this article”. I slightly disagree with the Leiden declaration regarding the disclosure of prompts, thinking traces, and energy and token usage for the simple pragmatic reason that these are not always available and the output typically isn’t reproducible anyway. If they are available though, providing a link to the thinking trace can be very valuable, especially if a sophisticated harness was used.
On the other hand, I do not believe that it is particularly useful to disclose the use of AI if it was only used for menial tasks like searching for references, proofreading a draft of the paper for mathematical or linguistic typos, searching for grammatical mistakes, etc. After all, there is no expectation to disclose the use of search engines, spellcheckers, etc that have long been used for such tasks, it is simply expected that the authors would have made use of such tools.
Attribution
A major problem with AI usage is that it appears to be extremely good at putting together several mathematical ideas in order to solve a problem, but quite bad at providing references to the articles where these ideas originated in the first place. It is somewhat ironic though that, when asked specifically about it, AI system do actually tend to be very good at finding such references.
If AI output contains what appears to be a new idea, it is therefore very important to try to get to the bottom of where it actually comes from. In many cases, this can be done by simply copying the output into a new context, together with a prompt along the lines of: “Someone submitted the following mathematical argument, but references seem to be missing (or incomplete). Could you find out where these ideas first appeared in the mathematical literature?”
If you are developing a harness geared towards mathematical research, please do seriously consider including such a step, even if it increases token usage.
Writing style
Here is a personal plea: please do not copy / paste the output of an AI system into your article. Instead, for the reasons discussed i the first section, digest the argument and then explain it in your own words. Some reasons for this are:
- Often, AI systems have a somewhat meandering style that focuses on following an argument step by step and loses track of the big picture. The reader however is typically interested primarily in the big picture and the way in which the main ideas interlock.
- In the process of digesting the argument, you may well realise that it can be presented in a clearer and simpler way.
- The reader is interested in your thoughts. Do not insult them by spending less time to write a paragraph than it takes them to read it.
- It is jarring for the reader to have a significant break in style between the human-written parts and the AI written parts of the paper. Even for technical writing, I personally find the AI writing style to be grating and soulless.
- In my experience, AI systems tend to be more verbose than necessary and spend a lot of time explaining in detail relatively basic material while glossing over more subtle points. Good technical writing should to the precise opposite.
- It is easy to overlook parts where its notations differs from those of the rest of the paper.
- It makes it easy to overlook logical inconsistencies (see also the first section).
A corollary of this is that I would strongly discourage the use of AI to provide “filling material”. I would rather read a dense but well-written paragraph than a full page containing the same information, interspersed with buzzwords and empty phrases.
Conclusion and disclaimer
This post should not be construed as any opinion regarding the ethics of AI usage, whether in mathematical research or otherwise. Some mathematicians have taken public positions against the use of AI while some have gone all in. I believe that positions are defensible in good faith (albeit starting with quite different assumptions and priorities), but this is a debate that should be reserved for a different post, possibly also on this site.
Leave a Reply