Only Anatevka

Ethan Sussman, postdoc at Northwestern University

At the moment, the mathematics community is facing an unprecedented misalignment brought on by the advent of LLMs. According to the standard telling, the basic tension is between our traditional values (whatever those may be) and those of the corporations that make and market LLMs. The purpose of this brief essay is to point out that some (arguably many) examples of that claimed misalignment boil down to a preexisting mismatch between the values we profess and the incentive structure that we have set up. We are therefore blaming AI companies for something which is our responsibility to fix.

A good illustration comes from a recent petition. The exact content of that petition is unimportant for the present discussion. I want to focus on a single line:

Human mathematicians who are simultaneously proving the same results [as an LLM] may be forced to abruptly abort these research projects, regardless of the additional insights their unique approaches may bring.

Q. Forced by whom?

Of our ostensible traditional values, the most important is supposed to be understanding. Numerous pronouncements have been made that this thing is the true goal of mathematical labor. Moreover, problem solving and the production of novel theorems are necessary but not sufficient steps in this direction. This becomes a criticism of LLMs and the conduct of their makers when it is implied that, by prioritizing answer-generation, they ultimately hurt mathematical understanding. To me, this feels like a case of the pot calling the kettle black. To what extent is our current system really set up to maximize the production of understanding?

Partially, of course. But we cannot reward understanding directly, so instead we reward measurable proxies thereof and incentivize their production. For various reasons, the main proxy has become the research paper, published with the imprimatur of some prestigious journal. Some reasons are good: e.g. the need for an objective metric on which hiring committees can base their decisions.1 Some are not: e.g. the prejudicial view (famously expressed by Hardy) that the expository arts are for “second-rate minds.” The question becomes the extent to which our preferred proxies measure the thing they are supposed to, and that is open for debate.

I believe that most of the misalignment mentioned above has to do with LLMs exposing and widening that gap. Consider the example above, and the question raised. Suppose a human –\text{\textendash} say Buckmaster, Córdoba, or Martínez-Zoroa –\text{\textendash} is carrying out a unique research program along a unique line, via an approach that is expected to bring additional insights. Why should the fact that an LLM had previously proven a big theorem in the domain force them to “abruptly abort” their work? AI companies possess no formal authority to decide what mathematicians spend their time on, nor do they have that inclination. The subtext in the open letter seems to be that, without the carrot-on-a-stick of priority to claim, there will be little incentive for mathematicians to carry on. But if their work really stands to deliver understanding, then our incentives have failed to incentivize the right thing.

Ultimately, it is mathematicians who decide what mathematical activity is worthwhile. Buckmaster and Córdoba both have tenure. Martínez-Zoroa is on tenure track. They are free to work on what they please —\text{\textemdash} and we, as mathematicians, are free to reward them for doing so. Mathematicians decide what sort of papers appear in top journals, what sort of work garners prizes, and ultimately who is hired. If OpenAI solving the Navier–\text{\textendash}Stokes millennium problem (or prematurely finishing someone else’s research program) decreases our understanding, that reflects a failing on our part, irrespective of any wrongdoing on theirs.

Our emphasis on one particular form of mathematical activity –\text{\textendash} proving novel theorems and the ancillary activity of paper-writing –\text{\textendash} is to the detriment of all others. These second-class mathematical activities include refereeing, code generation and formalization, distillation and digestion, simplification, codification, exposition, teaching, and learning. All are in service of understanding. Perhaps Hardy was right that these activities are for second-rate minds, but the unfortunate reality is, in the age of the LLM and massive agentic swarms (the “Large Agent Collider”), we are all second-rate minds. If we wish for mathematical practice to flourish in the upcoming era, then the responsibility falls on us to redesign our incentive structure to better reward the whole spectrum of mathematical activity.  We can do little about the Cossacks at the gates.

We can renegotiate our tradition to better align with our needs. Perhaps the new world won’t be so bad.

  1. For mathematics, this is the LpL^p norm of paper strength, for some p≫1p\gg 1. For theoretical physics, with a smaller selection of journals to differentiate the strength of individual papers, it is the L1L^1 norm. ↩︎

Received 12 September 2026.

8 responses to “Only Anatevka”

  1. Jen Avatar
    Jen

    This is a wonderful post, reflecting a much more clear-eyed view than many others here, in my opinion. The unfortunate thing is that senior mathematicians are the ones with the power to lead this much-needed change in our incentive structure, and they have the least amount of stakes in the discussion. Hopefully tenured professors with a sense of social responsibility can step up to lead this change.

    1. just different Avatar

      Not just social responsibility, but *professional* responsibility. The decisions made in the coming months by senior mathematicians and policymakers will determine the mathematical capacity for the next several decades. If the decisions are bad or dilatory, the field will experience a massive COVID-style “learning loss” that might take generations to remedy.

  2. technicallybeard69aeef4272 Avatar
    technicallybeard69aeef4272

    Ethan has made many excellent points. In addition, the term “understanding” for humans is not at all well defined. What is the quantitative measure of human understanding? One could define it roughly as the ability to actually do something with the result that you supposedly understand. But do what? Prove more theorems? Cook up new questions, conjectures, tools, concepts, primitives for math? Pass a test about this theorem? Based on this definition of “understanding,” LLMs have deeper and broader and better understanding of their proofs than humans do; think about it. Moreover, a “digestion” of an LLM proof by a human mathematician amounts to boiling down the proof into a shorter form in terms of standard human math concepts and the subset of the math literature that is known to humans, but this does not imply that the reader actually “understands” the proof or the digestion of the proof. How many math papers and math talks do you actually “understand”? I suggest that human digestions of proofs and human math talks and human math papers attempt to create the illusion of “understanding,” but typically they do not result in actual understanding (whatever that means). Furthermore, we can easily train LLMs to produce shorter more elegant proofs using concepts and results that are more familiar to humans. Gauss was not satisfied with only one proof of a theorem, but rather he strived for many different proofs of a given theorem, each giving different insights.

    1. Michael Rozynski Avatar
      Michael Rozynski

      “In addition, the term “understanding” for humans is not at all well defined.”

      I don’t understand.

      1. AI-ON Avatar
        AI-ON

        “In mathematics you don’t understand things. You just get used to them.”
        (John von Neumann)

  3. AI-ON Avatar
    AI-ON

    The slogan so far has been ‘publish or perish’, which could soon become ‘prompt or perish’. But why does the alternative ‘perish’ have to be so drastic?

    “Our emphasis on one particular form of mathematical activity..is to the detriment of all others.”

    Apparently, ‘perish’ subsumes all those second-class activities listed above and is seen as an equivalent of being somehow ‘non-existent’ in academia, even if it does not mean leaving academia altogether. How about replacing ‘perish’ with ‘prosper’? I think it’s really worth considering.

    I’ve read Hardy’s apology completely a couple of weeks ago, although I’d dipped into this essay earlier. In fact, I bought both Hardy’s Mathematician’s Apology and Harris’ Mathematics without Apologies simultaneously, in order to compare these two works. I am reading the book by Michael Harris now.

    After reading Hardy I came to the conclusion that one should be thankful for his views and opinions but should not take them all too seriously, since those come out of the murky depths of a deeply disturbed personality; this becomes manifestly evident towards the end of his essay.

    “Judged by all practical standards, the value of my mathematical life is nil; and outside mathematics it is trivial anyhow.” (G.H. Hardy, last page)

    In the end, despite his marvellous achievements, as a person Hardy was ‘perished’. Should he really serve as a shining example to others now or in future, or should one better leave him in the past as an equally shining ‘counterexample’?

  4. Manuel Rissel Avatar
    Manuel Rissel

    One main obstacle is the following:

    how to better measure work of mathematicians in a realistic way, that is feasible and natural enough to be adopted globally?

    It’s clearly important to realize the flaws of the current system, but it’s another to find a good solution. That could be more difficult than solving some millenium problem (and maybe it is the real millenium problem) so it is hard to expect senior mathematicians to just easily find good answers and coordinate them globally in a way people agree on and adopt.

    There are certainly many more or less idealistic views of how it should be, but I haven’t seen a concrete one that actually seems to have the chance being adopted by a majority.

    So far, my current idea, which is not perfect at all (I am not fully satisfied by the points I am listing below) is as follows:

    – Credit could go to major lines of study, 1 good holistic contribution per two years could be an optimum. A working paper with properly proven results could be put to arXiv, maybe in a way not accessible to robots (a thought I recently played with in another post here) and then be further refined and extended until a final paper emerges after 1-X years. If one has a few collaborations and a student, maybe the number of contributions increases a bit. But if one is responsible about ones work, too many collaborations are likewise not possible, as after all being author means one contributed a notable portion.

    – Brute force harvesting of low hanging fruit should not be incentivized. As a researcher, one has a specialization and a program and within this program one tries to advance understanding. Just to try cashing out on an imperfect system that is currently under transition is hardly responsible and should strongly impact reputation (and I believe it already does). One can do it, its not illegal, but it should not lead to a research position or promotion.

    – Selection of results to be published could be based largely on presentations and lectures gives during peer-review stage, and/or on concrete applicability and other values brought by the study. There could be some service where one can document in an encrypted way the progress of ongoing working projects, keeping project scope and title public, so that if other similar results appear it can be made public with time stamps at at which time who had already which progress to demonstrate a coherent timeline of findings, allowing a judgement whether results obtained by different groups were obtained in a sense independently. It would also make public the order of milestones achieved and make the project more credible. Such a service could be used like arXiv by most researchers to establish a public program on which they work, so that others know about this. Then, multiple results of similar nature could still be accepted by different journals, and those who provide better exposition, better lectures and talks, and so on may end up as the preferred one, or several different studies on the same issue end up as equally preferred by readers/journals.

    One aspects of the points above is transparency. If everyone who wishes to compete in math research environment, e.g for jobs, has to transparently state own working projects starting from early stage and update them (possibly in an encrypted way that can be encrypted later), probably several current issues could be tackled. Results that suddenly appear outside of well documented working projects could count negative. In the best case this also reflects the scope of grants one has and so on.

    The idea of working papers is something I see sometimes with economists, who often appear to publish quite less than mathematicians. That looks somewhat more healthy to me.

  5. a Avatar
    a

    I also think that we should be looking at other disciplines carefully to see how they evaluate academic work. (This seems a naural thing to do but I have not seen such a study in this recent literature).

Comments are moderated. Read our comment policy.

Add to the discussion

New posts by email.

Prefer a feed reader? Subscribe by RSS.

Also on Mathstodon.

Latest comments across the site.

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading