Only Anatevka

Ethan Sussman, postdoc at Northwestern University

At the moment, the mathematics community is facing an unprecedented misalignment brought on by the advent of LLMs. According to the standard telling, the basic tension is between our traditional values (whatever those may be) and those of the corporations that make and market LLMs. The purpose of this brief essay is to point out that some (arguably many) examples of that claimed misalignment boil down to a preexisting mismatch between the values we profess and the incentive structure that we have set up. We are therefore blaming AI companies for something which is our responsibility to fix.

A good illustration comes from a recent petition. The exact content of that petition is unimportant for the present discussion. I want to focus on a single line:

Human mathematicians who are simultaneously proving the same results [as an LLM] may be forced to abruptly abort these research projects, regardless of the additional insights their unique approaches may bring.

Q. Forced by whom?

Of our ostensible traditional values, the most important is supposed to be understanding. Numerous pronouncements have been made that this thing is the true goal of mathematical labor. Moreover, problem solving and the production of novel theorems are necessary but not sufficient steps in this direction. This becomes a criticism of LLMs and the conduct of their makers when it is implied that, by prioritizing answer-generation, they ultimately hurt mathematical understanding. To me, this feels like a case of the pot calling the kettle black. To what extent is our current system really set up to maximize the production of understanding?

Partially, of course. But we cannot reward understanding directly, so instead we reward measurable proxies thereof and incentivize their production. For various reasons, the main proxy has become the research paper, published with the imprimatur of some prestigious journal. Some reasons are good: e.g. the need for an objective metric on which hiring committees can base their decisions.1 Some are not: e.g. the prejudicial view (famously expressed by Hardy) that the expository arts are for “second-rate minds.” The question becomes the extent to which our preferred proxies measure the thing they are supposed to, and that is open for debate.

I believe that most of the misalignment mentioned above has to do with LLMs exposing and widening that gap. Consider the example above, and the question raised. Suppose a human –\text{\textendash} say Buckmaster, Córdoba, or Martínez-Zoroa –\text{\textendash} is carrying out a unique research program along a unique line, via an approach that is expected to bring additional insights. Why should the fact that an LLM had previously proven a big theorem in the domain force them to “abruptly abort” their work? AI companies possess no formal authority to decide what mathematicians spend their time on, nor do they have that inclination. The subtext in the open letter seems to be that, without the carrot-on-a-stick of priority to claim, there will be little incentive for mathematicians to carry on. But if their work really stands to deliver understanding, then our incentives have failed to incentivize the right thing.

Ultimately, it is mathematicians who decide what mathematical activity is worthwhile. Buckmaster and Córdoba both have tenure. Martínez-Zoroa is on tenure track. They are free to work on what they please —\text{\textemdash} and we, as mathematicians, are free to reward them for doing so. Mathematicians decide what sort of papers appear in top journals, what sort of work garners prizes, and ultimately who is hired. If OpenAI solving the Navier–\text{\textendash}Stokes millennium problem (or prematurely finishing someone else’s research program) decreases our understanding, that reflects a failing on our part, irrespective of any wrongdoing on theirs.

Our emphasis on one particular form of mathematical activity –\text{\textendash} proving novel theorems and the ancillary activity of paper-writing –\text{\textendash} is to the detriment of all others. These second-class mathematical activities include refereeing, code generation and formalization, distillation and digestion, simplification, codification, exposition, teaching, and learning. All are in service of understanding. Perhaps Hardy was right that these activities are for second-rate minds, but the unfortunate reality is, in the age of the LLM and massive agentic swarms (the “Large Agent Collider”), we are all second-rate minds. If we wish for mathematical practice to flourish in the upcoming era, then the responsibility falls on us to redesign our incentive structure to better reward the whole spectrum of mathematical activity.  We can do little about the Cossacks at the gates.

We can renegotiate our tradition to better align with our needs. Perhaps the new world won’t be so bad.

  1. For mathematics, this is the LpL^p norm of paper strength, for some p≫1p\gg 1. For theoretical physics, with a smaller selection of journals to differentiate the strength of individual papers, it is the L1L^1 norm. ↩︎

Received 12 September 2026.

2 responses to “Only Anatevka”

  1. Jen Avatar
    Jen

    This is a wonderful post, reflecting a much more clear-eyed view than many others here, in my opinion. The unfortunate thing is that senior mathematicians are the ones with the power to lead this much-needed change in our incentive structure, and they have the least amount of stakes in the discussion. Hopefully tenured professors with a sense of social responsibility can step up to lead this change.

  2. technicallybeard69aeef4272 Avatar
    technicallybeard69aeef4272

    Ethan has made many excellent points. In addition, the term “understanding” for humans is not at all well defined. What is the quantitative measure of human understanding? One could define it roughly as the ability to actually do something with the result that you supposedly understand. But do what? Prove more theorems? Cook up new questions, conjectures, tools, concepts, primitives for math? Pass a test about this theorem? Based on this definition of “understanding,” LLMs have deeper and broader and better understanding of their proofs than humans do; think about it. Moreover, a “digestion” of an LLM proof by a human mathematician amounts to boiling down the proof into a shorter form in terms of standard human math concepts and the subset of the math literature that is known to humans, but this does not imply that the reader actually “understands” the proof or the digestion of the proof. How many math papers and math talks do you actually “understand”? I suggest that human digestions of proofs and human math talks and human math papers attempt to create the illusion of “understanding,” but typically they do not result in actual understanding (whatever that means). Furthermore, we can easily train LLMs to produce shorter more elegant proofs using concepts and results that are more familiar to humans. Gauss was not satisfied with only one proof of a theorem, but rather he strived for many different proofs of a given theorem, each giving different insights.

Comments are moderated. Read our comment policy.

Add to the discussion

New posts by email.

Prefer a feed reader? Subscribe by RSS.

Also on Mathstodon.

Latest comments across the site.

  1. Ethan has made many excellent points. In addition, the term "understanding" for humans is not at all well defined. What...

  2. This is a wonderful post, reflecting a much more clear-eyed view than many others here, in my opinion. The unfortunate...

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading