AI and the mismeasure of science

Mario Pasquato, astrophysicist at IASF-Milano, and author of Picking Flowers in the Garden of Forking Paths

Most of us are here because of the events of this summer, namely, large language models (LLMs) taking on mathematics. Since I am not a mathematician (I am a physicist and work at the intersection of astrophysics and machine learning), my contribution will address a broader issue, largely building on essays that I published on my blog a few months before the proverbial excrement hit the fan.

Astrophysics has been impacted by AI in a different fashion from mathematics: less publicly, less spectacularly, but earlier and, if anything, even more strongly. The writing has been on the wall for at least a year, but even back when GPT-3 landed, it was pretty obvious that the face of our discipline, for better or worse, would have to change permanently, at least to anyone able and willing to connect the dots.

One person who is particularly good at connecting such dots is Prof. David Hogg. He wrote a position paper back in February on the future of astrophysics in a world where LLMs become ever better at doing data analysis and writing papers. In May, I wrote a piece that shares his diagnosis but not his prognosis. This is essentially an updated version of that piece, taking into account what happened to mathematics this summer.

Hogg argues that, since astrophysics essentially coincides with the astrophysics literature (more on this later), and since a major portion of our research is essentially data science and writing, tasks at which LLMs are becoming increasingly competent and, at any rate, much faster than human researchers, we should either restrict LLM use or accept that LLMs will come to dominate astrophysics.

To borrow the names introduced by political scientist Jim Dator and his group at the University of Hawaii in their framework for envisioning future scenarios, the first scenario, where the astrophysics community manages to restrict LLM use, would be called Discipline, and the second, where people keep using LLMs and nothing else changes, Continued Growth. Hogg calls them more colorfully “ban-and-punish” and “let-them-cook”.

Under “ban-and-punish,” we would basically be doing research as in, say, 2019, using hand-written code that at least one human researcher has spent some time proofreading and understanding, while also hand-typing all of our scientific papers. In this scenario, we would use LLMs only for menial tasks, such as reformatting a table, or as a glorified spell checker. Hogg points out two obvious downsides to this. First, it would have to be enforced, because at least some researchers would want to defect and use LLMs in some covert fashion. This would violate academic freedom, even if it were possible. At any rate, the academic, publishing, and grant-making environment is too fragmented for meaningful enforcement. The second downside is that we would be leaving utility on the table, so to speak: if LLMs, or human-LLM centaurs, are more efficient at doing research than humans alone, this would severely slow the pace of scientific discovery relative to what it could be in an unrestricted setting.

In the “let-them-cook” scenario we would collectively let our LLMs loose on at least the ever-expanding corpus of publicly available data, automatically producing analyses that would be turned, perhaps with some human supervision, into publishable papers. The details of the refereeing pipeline under this scenario are left as an exercise for the reader in Hogg’s white paper, but it is pretty clear that any arrangement that allows even a relatively tiny fraction of these potential papers to be published will flood the literature, largely crowding out whatever human production is left. A researcher who wants to keep up with the literature would then be forced to play a losing game of endless catch-up. Since reading about astrophysics is no substitute for actually practicing it, this would lead to a world where astrophysics still exists, but is no longer practiced by humans.

In his paper, Hogg does not make any specific policy recommendation, but it seems clear that he is looking for a middle ground between these two extremes. What I would like to propose is a new scenario, one that Dator would call Transformation. Where Continued Growth is basically an extrapolation of current trends and Discipline is an artificial slowdown, Transformation is a qualitatively different outcome altogether. Not only is this not a middle ground between the two extremes outlined by Hogg, it is also, in my opinion, more likely to play out than either of them. The key move to access this scenario is to reject the main assumption Hogg makes in his piece, i.e. that astrophysics roughly coincides with the published astrophysical literature:

Astrophysics is (roughly) the astrophysics literature. […\ldots] I am occasionally asked to speak on the idea that software is just as important as other kinds of scientific outputs. People are surprised when I don’t completely agree. I believe that (in astrophysics) software is written to support the astrophysics literature, and that every important piece of software should have an associated paper in the astrophysics literature. That paper describes the intellectual contributions encoded in the software, and provides a standardized mechanism for tracing intellectual provenance, giving credit, and coordinating criticism. […\ldots] Nothing we have has the track record that the traditional journals do for long term dissemination and preservation of knowledge; we don’t know what computing hardware platforms we will have in 15 years, let alone 200, but I can pretty-much guarantee that—\text{\textemdash}if human civilization is still in existence—\text{\textemdash}we will still be able to read Vera Rubin’s papers about the dark matter […\ldots]. My point is that astrophysics is the astrophysics literature.

Note that Hogg is a deservedly successful scientist, and his opinions on these matters are likely shared by many in the community. I would hate to misrepresent his position here, but it really seems to me that he is saying that “science = papers”.

There may be some cultural bias at work here: Americans and other anglophones may be more prone to identifying peer-reviewed literature with science because the dominance of their language on the international scientific scene is relatively recent and roughly coeval with the ascent of peer review. When new results were instead disseminated through learned societies and private correspondence, the languages used by scientists were more diverse.

At any rate, it is the notion that “science = papers” that makes LLM usage both inevitable and problematic. Researchers engage with a handful of papers a day, at the very most. Today’s astro-ph had 93 submissions [at the time of writing the original post in May]. Even if we do away with the accelerating factor represented by LLMs via a “ban-and-punish” approach, and the literature keeps growing in size at 2019 rates, I would expect many researchers to eventually use LLMs to find and summarize papers rather than read them directly. While you would not be writing with an LLM, you would be writing for one!

Of course, LLMs also make data analysis and the production of scientific prose much cheaper, but this merely exacerbates a problem that was largely already there. The paper used to be an artisanally crafted good, and LLMs are about to turn it into an industrially produced one. Now, if we, as astrophysicists, were exclusively in the business of producing papers, we would indeed face one of the two scenarios outlined by Hogg: either our guild somehow restricts the application of this new technology to preserve the status quo, or we succumb and become alienated, as Marxists would have it.

The key to a way out is to admit that the main job of scientists is not, after all, producing papers, because the equation “science = papers” does not hold. This should have been obvious in retrospect: the NASA abstract system lists 236,208 refereed papers for the 2020-present period, while the entire corpus of pre-1950 consists of 132,600 papers, yet it would be hard to convince me that the amount of new knowledge we actually generated in the last six and a half years is greater, let alone by almost a factor of two.

Far from tracking the generation of new knowledge, the current paper explosion is the unintended result of unrealistic legibility requirements imposed on the scientific community by policy makers. In addition to being a way to disseminate new findings, papers have become a signaling mechanism: researchers who are good at writing papers are also good at doing science, or so the key decision makers hope.

LLMs are breaking this signaling mechanism because you can’t do costly signaling with cheap stuff, but the runaway explosion in the number of papers was already problematic, and LLMs are merely showing that the king is naked. Writing many papers is easier than achieving real scientific results, so the incentive structure scientists respond to promotes a decoupling between one activity and the other.

Because science is a strong-link problem, that is, it lives or dies with the best output of any given scientist, it is fine if most papers a scientist writes are essentially ignored in the long term. The key, lasting contributions, those that end up in textbooks, are few and far between and may occur only once in a scientist’s lifetime. Or never. This makes scientific activity hard to evaluate, and the shortcut of using a proxy measurement such as papers is very tempting.

Still, keeping scientists around is worthwhile both for the incremental results they produce and, crucially, because some of them sometimes have a great idea that changes the face of their discipline. Using the proxy of paper production tracks, at best, incremental progress of the kind that is now likely to become increasingly more automated by AI.

It is helpful here to remember the distinction introduced by Kuhn between normal science, where incremental puzzle solving is king, and science in times of crisis, which eventually usher in scientific revolutions. Under normal science, it is tempting to treat scientists as a standardized commodity whose worth can be measured in terms of a standardized product: papers. In the not-so-unlikely scenario in which AI automates most of normal science, or at least its proxy of paper production, that commodity can be easily replaced. But if AI, for whatever reason, cannot compete with scientists at the different game of coming up with crazy ideas and seeing what sticks that characterizes paradigm changes, then following this temptation may lead us to a period of prolonged stagnation.

The bottom line is that there are times when it’s clear that scientific output is not measured in papers: in a crisis, it’s fairly easy to appreciate that papers are not the fire that powers the engine of science. At best, they are the smoke coming out of the exhaust pipe, a byproduct that merely reveals that the engine is on. We just happen not to be in a crisis right now.

With this understanding in mind, there is a simple recipe that would trigger a Transformation future: change the way we evaluate science to make it clear that scientists are not a commodity. My proposal is simple: we need to cap the number of papers a given author can publish per year. This obviously removes the incentive to use LLMs to publish an inordinate number of papers, but lets us use LLMs to do science.

Papers will still be authored by humans, in the sense that a human ultimately signs off on and takes responsibility for a paper’s contents, but besides that, anything goes. You can spin up as many agents as you like and have them write all the code you need. You can use an LLM to write your introduction, or the whole paper, for that matter. But you get to do this only twice a year1.

Unlike Hogg’s “ban-and-punish” scenario, a yearly cap on publications is enforceable: just get the main journals to accept no more than two papers per year per researcher. If this is, for any reason, not feasible, have granting agencies and tenure committees refuse to consider more than two papers per year per researcher when evaluating an application. Some details may need ironing out, such as whether to cap all papers or only first-author papers, and what to do with unused publishing credits, but the main idea is quite simple. These details need to be adapted to the scientific community of each discipline: for instance, mathematicians already publish at a lower rate than astrophysicists and usually sign their papers in alphabetical order.

No matter the subject, though, absent the incentive to publish more papers, researchers will concentrate on quality. Quantity would have lost any signaling value anyway, so it is reasonable to restrict it in a controlled fashion before it ends up flooding the literature with noise. Limiting the number of papers makes it easy for humans to evaluate them in depth for quality, possibly with the help of AI, as long as we agree on what quality is. Two dimensions of quality that come to mind are impact, not in terms of citations, but as judged subjectively by other humans, and, for experimental disciplines, reproducibility. The latter is incidentally now much easier to check automatically thanks to AI, as long as the datasets are available.

I expect this scenario to yield much longer papers, full of consolidated results. Think of it as the reverse of the infamous salami-slicing technique, in which researchers split a result into multiple papers to boost their stats. Researchers will write only when they have something meaningful to say, and will want to say it in a form that maximizes its intelligibility, so they will also have an incentive to be pedagogical and perhaps to introduce elements of intellectual branding on the model of, say, continental philosophers or literary scholars, which would further decommoditize the role of the scientist. You may say that I reinvented books here, but I prefer to think of distill.pub , an online journal whose distills were interactive adventures into a topic that combined elements of quality outreach with the presentation of state-of-the-art results, minus the part in which the journal shut down because distills ended up being too ambitious to curate by hand.

  1. arXiv has just announced a new rate limit policy of two submissions per calendar month: it’s a start. ↩︎

Received 3 September 2026, revised 1 October 2026.

One response to “AI and the mismeasure of science”

  1. Anonymous Avatar
    Anonymous

    It is certainly an interesting idea, could you develop more on the following objections?
    1) wouldn’t this inevitably increase the number of predatory journals?
    2) would young researchers stand a chance in publishing in respectable journals? I feel that as long as such young person is not “part of the club” it could be even worse.
    3) longer and deeper papers sounds excellent, but peer review would have to be boosted dramatically. I am aware of papers already taking 6 years in a peer-review process.

    Thank you

Comments are moderated. Read our comment policy.

Add to the discussion

New posts by email.

Prefer a feed reader? Subscribe by RSS.

Also on Mathstodon.

Latest comments across the site.

  1. It is certainly an interesting idea, could you develop more on the following objections? 1) wouldn’t this inevitably increase the...

  2. The ancients knew that it is fruitless to argue taste: "de gustibus non disputandum est," or "Chacon a son gout."...

  3. One computer science conference proceedings editor recently interviewed ten authors of suspicious papers, to see how well they understood their...

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading