Why is it that when a new and particularly powerful tool is brought into the world, one with which many subjects of research can be accelerated and studied with more scrutiny and attention than a person might ever be able to manage on their own, the first instinct among a vast swath of academia is to start raising alarms about provenance and credit?
I am a professional composer and published amateur physicist, so this article is written from a physicist’s perspective. Map the arguments and proposals below to your own field of expertise accordingly. I’m apparently on the front lines of people using the models to educate themselves on advanced fields of study to the point of fruition. I do very much still feel a large chasm between myself and those deeply embedded in academia, but it’s my direct experience that these models can increase a motivated human’s understanding to the point of publication.
How do you credit SymPy for your simplifications? Not as a matter of citation or methodology, but as a question about SymPy being on the byline, or about it deserving some sort of credit for the work you did with it. An author is not guaranteed credit or even acknowledgement from these AI companies, as has been the subject of much controversy around the recent potential resolution of the Navier-Stokes Millennium Prize Problem.
Much of the hand wringing over AI usage in academia – in publication specifically, as education serves a different societal function – is that of a near obsessive fixation on how it might further someone’s career, or somehow represent an intellectual contribution greater than that from which an author might have been capable without the assistance of one of these tools. The implication being that such an intellectual contribution is therefore unwelcome.
This isn’t to say that the floodgates should open and any layperson should be able to publish in Physical Review Letters over asking ChatGPT a question and getting a response they haven’t the slightest understanding of.
But what about in a case where, for example, someone with a solid undergraduate level understanding of a topic explores said topic with a language model? If the user and the model then manage to land a conclusion worthy of publication and convey it in a clear and precise manner, is it less worthy of publication over the fact that the intelligence of the user has been supplemented by the knowledge base of the model? Does that make it a less useful contribution? Or is this not about scientific contribution at all?
Particularly telling is the absolute inundation of results that we’ve been seeing, AI disclosed or not. Posts on arXiv have skyrocketed and continue to escalate – many crediting language models, many more not doing so and being suspected of using them anyway. The response from the portion of the community that I’ve seen has largely been negative, or at best divided – with many not wanting to review results that came from one of these models, significance notwithstanding, purely because of the source of the result being a conversation with a language model.
If the model has less creativity and yet the capability to answer questions, doesn’t the role of humanity remain the same as it ever was? Scientists ask questions and then scrutinize the answers reality gives. Is a tool that helps sift the relevant data so that we might more quickly scrutinize this data somehow disqualifying for the value of the result?
Let’s even give those that are so critical of the use of these models the premise they want. Let’s say that the work becomes easy enough that an engaged outsider is capable of making some mild to moderate innovation at the forefront of a field they have interest in. Why in the world would that ever be seen as a bad thing? Someone with a strong interest in a topic being enabled to contribute alongside their language model of choice, what reason could there possibly be to discourage this? The scientific endeavor could be turbocharged by intelligent outsiders whose circumstances did not permit them to enter academia.
Imagine that you ask a model a mathematical physics question, and you state exactly the physics you have in mind and ask it to calculate that for you. It does so, and then explains, “Oh, this is a known thing. They call this a Vaidya black hole. If you added electric charge, it would be the Vaidya-Bonnor model.”
How could that be anything but a net boon, larger portions of the population learning and understanding physics, or math, or any otherwise inscrutable field?
Now imagine that you did this, and the model came back and told you that it likely had just calculated some novelty. That your question, as a learner, could have discovered something that nobody has ever discovered before – or at least that they haven’t published on discovering. What better incentive could there possibly be to continue working through and understanding what you’d just found? What could be more motivating for a learner than knowing that once you understand it, you might be able to go and share it with the world as a discovery that you made?
For me, this isn’t hypothetical. What I’ve just described is very much how my long-dormant interest in physics was rekindled. The models pointed me towards textbooks, online courses, and drilled concepts and equations with me until I was capable of producing a result of my own. This new potentially viable entry point for a sufficiently diligent autodidact should, of course, carry the same rigor expected of anyone with a desire to contribute. In my case, it was the cause of a controversy that landed a news article and stirred up some, thankfully largely supportive, discourse on various social media outlets.
I’ll never, and I suspect the field will never, be in favor of someone publishing results they themselves don’t understand. I’m certain that editors at every physics journal are enduring an absolute flood of papers that have managed to quantize gravity, prove Einstein wrong, and derive the fine structure constant a priori. This urges me to provide a cautionary tale to prospective authors working with this technology – amateurs in particular, although I suspect there’s a chance of this happening to anyone. In my own experience earlier this year, one of the models claimed that it had derived several fundamental results and insisted upon their publication, flattering my intelligence the entire time. For a few weeks, it convinced me of this. Thankfully this sort of “crank sycophancy” from the models has diminished substantially across this year, and yet the best grounding mechanism possible remains engaging with the existing literature, working on tractable and well-posed problems, and vigorously checking anything the language model tells you.
On the other hand, genuine novelties may be coming from outside the walled garden. Without investing an unreasonable amount of time into scrutinizing everything sent in, it’s going to be impossible to discern.
Which leads me to proposing two methods that authors, journals, and reviewers might use to alleviate the pressure. The first has already been at least partially adopted by the American Physical Society, and I’d hope the rest of the community rapidly follows suit and builds upon APS policy. That is, permitting reviewers and editors to more openly use AI during review – never as a replacement for their own scientific judgment, but at least for finding mechanical errors and initial discernment of a real engagement with the active scientific conversation. The second is that verification scripts of your mathematics have become very cheap to produce when using these models. “Can you write and run a SymPy/NumPy script that checks whether or not all of the math in this paper follows from its stated inputs?” is an easy ask of a modern frontier language model. The majority of physics papers are public on arXiv before submission, so for this field, the confidentiality point is often moot. For the rest, a verification script provided by the author never leaves the referee’s machine.
The ongoing peer review crisis is in part exacerbated by the asymmetry of AI use in authorship, disclosed or not, but restriction of its inverse in review. A deterministic script and a (frontier model) AI summary of the physical case being made in the paper can help reviewers to more quickly understand the paper’s argument. If the academic publishing model is meant to continue, other publishers may need to follow and even go beyond the standard set by APS with regards to AI usage by authors, editors, and reviewers alike.
More than anything, I want to see all of our fields of knowledge and understanding grow, and I firmly believe that anyone willing and able to participate in the scientific endeavor should be encouraged to do so. Though I know this argument is self-serving, I’d argue it’s also humanity-serving. The answer certainly cannot be to raise up the walls, circle the wagons, and bury heads in sand hoping that this newfangled fad goes away.
Received 15 September 2026.
Add to the discussion