With the onset of “LLM driven math turbulence”, various guiding principles that seemed for granted are being shaken up a little. I can imagine that today someone could upload a new preprint introducing a rich theory that prepares all kinds of exciting unknown terrain for future exploration, and on the next day, before a word was talked about that theory with anyone, LLM driven agents may already have analyzed and used material from that preprint in the process of generating proofs to a list of conjectures which someone else (or some company) requested from a chatbot. Whether this would be perceived as good or bad likely depends on who one asks under which circumstances, but it seems natural to wonder: would the only guaranteed way for ruling out such a scenario (if one wanted to) consist of not uploading new research as preprints? And,
would a majority of mathematicians like their preprints (including the LaTeX files) being crawled by arbitrary automated robots?
There is hardly one answer fitting all or a majority. For instance, researchers may enjoy seeing their not yet peer-reviewed proofs appear in LLM generated answers. Some may want to actively support LLM driven research or at least have it stabilized at an equilibrium as quickly as possible. Others may not want to contribute new ideas immediately to automated LLM driven research or to external companies with other interests. Also, there can be nuances, as some may be fine with sharing titles and abstracts of preprints with robots, but would like to reserve the proofs mostly for mathematicians during the peer-review stage.
Nevertheless, there appear not many options on the table, as the preprint server mostly used by the math community operates based on principles such as the open distribution of knowledge, and openness is something that hardly anyone would like to change. But what means open? How to choose the topology that’s best for sustained mathematical science development? Is there a natural mechanism that ensures sufficient openness of preliminary research findings but still provides a new lever for the community to influence slightly the balance between the amount of digested mathematics and the number of math manuscripts?
One hypothetical thought could be that some authors may choose to encrypt their freshly indexed preprints, keeping the title and a possibly long abstract available to the public, with automated decryption of the main body to occur after a well-defined embargo period, e.g. 1-2 years (or any time scale larger than that of LLM development). Such a preprint would be compiled by the server after upload and then encrypted, retaining a readable copy only for private access or moderation. It would be time stamped and versioned, just like done by the arXiv right now, hence the authors could without worries start talking to others about the novelties of their new theory and send readable versions to journals and colleagues. To prevent meaningless encrypted preprints, there could be requirements for using this particular service: e.g., a limit of encrypted uploads per verified researcher per year. Other researchers interested in learning about the details contained in an encrypted preprint could invite the authors to visit them and give a talk, invite them to online seminars, engage in private exchange or to take this thought ad absurdum, someone could prompt an LLM
The drawbacks of voluntary encryption of preprints may outweigh the positives, but it could still be valuable to debate about the role of preprints in a changing research environment. Our way of handling preprints potentially influences peer-review, knowledge digestion, and the ability of external forces to act on our profession.
Received 8 September 2026.
Add to the discussion