What about the preprints?

Manuel Rissel, tenure-track faculty at ShanghaiTech University

With the onset of “LLM driven math turbulence”, various guiding principles that seemed for granted are being shaken up a little. I can imagine that today someone could upload a new preprint introducing a rich theory that prepares all kinds of exciting unknown terrain for future exploration, and on the next day, before a word was talked about that theory with anyone, LLM driven agents may already have analyzed and used material from that preprint in the process of generating proofs to a list of conjectures which someone else (or some company) requested from a chatbot. Whether this would be perceived as good or bad likely depends on who one asks under which circumstances, but it seems natural to wonder: would the only guaranteed way for ruling out such a scenario (if one wanted to) consist of not uploading new research as preprints? And,

would a majority of mathematicians like their preprints (including the LaTeX files) being crawled by arbitrary automated robots?

There is hardly one answer fitting all or a majority. For instance, researchers may enjoy seeing their not yet peer-reviewed proofs appear in LLM generated answers. Some may want to actively support LLM driven research or at least have it stabilized at an equilibrium as quickly as possible. Others may not want to contribute new ideas immediately to automated LLM driven research or to external companies with other interests. Also, there can be nuances, as some may be fine with sharing titles and abstracts of preprints with robots, but would like to reserve the proofs mostly for mathematicians during the peer-review stage.

Nevertheless, there appear not many options on the table, as the preprint server mostly used by the math community operates based on principles such as the open distribution of knowledge, and openness is something that hardly anyone would like to change. But what means open? How to choose the topology that’s best for sustained mathematical science development? Is there a natural mechanism that ensures sufficient openness of preliminary research findings but still provides a new lever for the community to influence slightly the balance between the amount of digested mathematics and the number of math manuscripts?

One hypothetical thought could be that some authors may choose to encrypt their freshly indexed preprints, keeping the title and a possibly long abstract available to the public, with automated decryption of the main body to occur after a well-defined embargo period, e.g. 1-2 years (or any time scale larger than that of LLM development). Such a preprint would be compiled by the server after upload and then encrypted, retaining a readable copy only for private access or moderation. It would be time stamped and versioned, just like done by the arXiv right now, hence the authors could without worries start talking to others about the novelties of their new theory and send readable versions to journals and colleagues. To prevent meaningless encrypted preprints, there could be requirements for using this particular service: e.g., a limit of encrypted uploads per verified researcher per year. Other researchers interested in learning about the details contained in an encrypted preprint could invite the authors to visit them and give a talk, invite them to online seminars, engage in private exchange…\ldots or to take this thought ad absurdum, someone could prompt an LLM…\ldots

The drawbacks of voluntary encryption of preprints may outweigh the positives, but it could still be valuable to debate about the role of preprints in a changing research environment. Our way of handling preprints potentially influences peer-review, knowledge digestion, and the ability of external forces to act on our profession.


Received 8 September 2026.

7 responses to “What about the preprints?”

  1. Nicholas Braun Rodrigues Avatar
    Nicholas Braun Rodrigues

    Just one thought on preprints (papers in general) in the era of AI-generated mathematics: Why publish? If one uses IA to generate a result, probably that AI model now knows your result, or, if someone else want to know it, they can just prompt themselves into AI. So why keep publishing results? To claim authorship? The meaning of authorship will also have to be updated at each new AI model. These are questions that the community of AI-users mathematicians will have to answer eventually. Regarding your point, I’m considering on not posting my preprints to arxiv anymore (at least this is my current position). And I’m not using AI to do research, and I intend on continuing doing so.

    1. Manuel Rissel Avatar
      Manuel Rissel

      Thats very interesting, I did not know about this, thanks very much for sharing! But in any case, I believe that even if arXiv does not allow bots to download papers, they cannot prevent them from doing it either. If a file is encrypted well enough, good luck to the bots. Certainly it can leak through untrusted parties with whom the file was shared privately, but that would not be a systematic flaw and to a larger degree controllable by the author.

  2. nimatt Avatar
    nimatt

    You might want to have a look at this:

    https://math.mit.edu/~elmos/petition.html

  3. Dominique Manchon Avatar

    On November 25th 2025 the new “english mandatory” policy was announced on arXiv’s blog, resulting in numerous reactions and protests under the blog post. The new policy took place on Februray 11th 2026. On March 26th, I posted the the text reproduced below on arXiv’s blog, which got no answer since then…

    “Je me demande si la migration d’arXiv hors de l’université Cornell vers un statut d’organisation à but non lucratif indépendante, mais dans le giron de grandes organisations comme la fondation Simons et la fondation Schmidt Sciences, cette dernière très branchée IA, n’est pas la véritable raison de cette nouvelle politique.

    https://tech.cornell.edu/arxiv/

    Des éclaircissements seraient bienvenus.

    (I wonder whether arXiv’s move from a Cornell University component into an independent nonprofit organization in the orbit of respectable charity organizations, among which the Simons Foundation and the very AI-inclined Schmidt Sciences, could be the very reason why this new policy has been imposed to us. Some explanations would be welcome.)”

  4. Hume Avatar
    Hume

    One thing I don’t understand about your presumption is why would a person release something as a conjecture without having their LLM agents to take a stab on it first? And if their agents fail, they don’t need to worry too much it gets scooped cuz it wasn’t within their reach in the first place anyway.

    1. Manuel Avatar

      I have somehow the opposite question: would really everyone use LLMs to do research or share novel ideas with them?

      Lets make two assumptions on which likely most would agree
      – human consciousness \neq LLM structure,
      – number of results all LLMs together can prove in X years is finite.

      Necessarily, there are thus infinitely many results any LLM interaction will not produce in a fixed time, while humans could discover it on an LLM free route, simply since we are not the same as LLMs. Once you get your feedback from an LLM you also think along different routes, thus LLM free routes may provide progress that would not be obtained otherwise at a given time. This may particularly apply to the creation new theories and concepts. So it seems not entirely unreasonable that some people would not use them (for everything) and would also prefer some options for sharing their preliminary findings.

      But there are also various other reasons for thinking about the questions I asked. For instance, look at peer review:

      – increasing number of papers could mean increasingly hasty or indifferent peer review. If papers get generated or results are largely found by LLMs, do the authors of these works still read other papers and spend time with peer-review or just ask LLMs to do this, as well? If a paper is not publicly on arXiv, someone intending to do so would likely have to violate a clear journal rule like ”do not upload this manuscript to an LLM”. Thus, for some researchers who would like to deliberately choose a slower (or is it a faster one?) path of understanding, it may be interesting to some to share their preliminary progress in a way that would minimize LLMs deciding about its fate (e.g. during peer-review).
      – If someone generates results en mass using LLM agents, there seem less incentives to share it first in an encrypted way and wait for formal publication, relying on seminar talks and mini-courses to transmit the ideas to a wider audience during peer-review stage. Thus, preprints carefully shared in an encrypted way may gain some trust from serious reviewers and may actually be read more carefully or with more passion.

Comments are moderated. Read our comment policy.

Add to the discussion

New posts by email.

Prefer a feed reader? Subscribe by RSS.

Also on Mathstodon.

Latest comments across the site.

  1. Ethan has made many excellent points. In addition, the term "understanding" for humans is not at all well defined. What...

  2. This is a wonderful post, reflecting a much more clear-eyed view than many others here, in my opinion. The unfortunate...

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading