As Yogi Berra said, it’s tough to make predictions, especially about the future.
So I decided to write about my present experience, which I think I have a better chance of being right about. At the least, it’ll provide a snapshot of a tiny piece of a large moving picture.
For most of the past year, I have been using various AIs for mathematical discussions — sometimes to test their abilities, sometimes because I want their help. (For the purposes of this post, I will use AI and LLM as synonyms.) There is no doubt that things are changing in their abilities to be valuable interlocutors – alas, I am not confident that the process is because they are generally getting better at being right. I wonder what it is that they are getting better at.
More precisely, they are clearly getting better at something, and in math, it is correlated with being right on pretty-well-understood phenomena, and therefore our experience using it is better. But, for example, it could just be optimized for making mistakes that are much harder to find, and it’s getting to the point that we don’t find them.
Currently, it seems like they tend to distil “conventional wisdom” and bring to bear broad acquaintance with many things I am only vaguely aware of which is certainly valuable to me, but they often do not execute it correctly or “really understand it”. A few months ago, I asked many LLMs for which primes does 2 have a cube root? They all told me it was a problem in class field theory, which is right; they all gave me criteria, to which they all found counterexamples when prompted. (I thought this was a softball problem: the answer to this problem being a conjecture of Euler proved by Gauss.) They all offered to and then did a reasonable job explaining elementary class field theory to me. One of them apologized and told me that they only intended that their criterion be applied to sufficiently large primes, and then told me it was ineffective about how high it would have to go for its criterion to hold! (Then it helpfully offered to explain to me why number theoretic theorems are sometimes ineffective.) Here the LLM was confabulating and very obviously doing so. But that was a few months ago, and it would do better now. (They do better – I just checked!)
Most of the time, the process is extremely frustrating – and it produces arguments that, depending on my expertise in the other area, are either easy or difficult to vet. Switching from generating ideas to vetting mainly bad arguments is not a lot of fun, but it can be educational. It’s not all bad, but it hasn’t made me more productive. It has made me much less polite in interacting with the LLM — and I hope that doesn’t spill over to my interactions with humans…. With the frustration level so high, I rarely keep at it for more than an hour or so.
But, I do expect this to change with time. Even over the short time I’ve been using them, I have seen progress in what they can do, and I have learnt things about problems I really cared about, through interactions that do not look any different on a screen from one I could have had with a collaborator (Mr. Turing). Whatever it is that LLMs can do, they surely can do better than we do, since AI’s don’t have human limits like small working memory, and they have the capacity to work for unbounded lengths of time without clear positive feedback. (If one again insists on more precision: they have larger bounds on the amount of resources that they can spend exploring a direction without positive feedback). They can also be prompted and taught to subject their responses to tough internal interrogation before answering. Whether there will be more qualitative changes than these is something that I view as wide open – but even quantitative changes have enormous impact and they are deservedly absorbing a lot of attention and causing much concern.
It was outside of theorem proving or problem solving that I had my most positive experience with an LLM. I had written a memo in support of something I was somewhat ambivalent about. The LLM rewrote it, and I got angry. I said something like “I will not write such BS – when I don’t believe it.” and went on about the need for integrity and explained what the goals of the memo were in more detail – namely supporting a plan with some reservations. We went through about eight hours or so of revisions and in the end wrote a memo that was much stronger than what I had originally intended, but which, at this point I was comfortable with. To be clear, at the end of the process, I more strongly supported the plan suggested in the memo than I had when I started.
Which got me thinking. Why did I spend eight hours on a memo that I thought a small fraction of the time was what I wanted to spend? Obviously, I found the process engaging.
What was so engaging about this process? I was able to see progress in every draft. The memo became more nuanced with every iteration. (This is something which still isn’t happening a lot with my mathematical discussions with LLMs – which is why I tend to use it more limitedly in my research.)
And finally, why did my support for this initiative increase through the process? Here there are two possible explanations, and I think they both have some truth. I spent a lot of time refining the argument, and it got to the point where it was better. And secondly, LLMs are really good at being convincing. Having critiqued various arguments as being too strong to be correct etc., the final version ended up a persuasive and nuanced argument or, perhaps an argument adapted to convincing me, as I was now much more psychologically invested in it.
So, ironically, it is out of my most positive experience that one of my most serious worries grows, and it is a problem that I fear will not resolve over time unless we think hard about what to do about it. LLMs are extremely convincing. We developed a human technology for this task for dealing with the fact that humans can be convincing yet incorrect: That technology is the social construct we call academe. We talk to each other, criticize each other, and write careful papers (ideally) which are carefully refereed (when the referees do so, and this is a harder and harder task). All of which, when it works, prevents “model collapse” which happens when humans (or AI’s) are able to promulgate ideas without robust criticism.
In math, we might be able to solve the problem (that of the convincing incorrect argument) because we have the idea of a proof, an idea that has taken millennia to refine, and we can insist on formalization (or, short of that, really insisting on understanding every point in excruciating detail). Of course, asking LLMs to only answer us in formalizable ways might make them too taciturn to engage with in a valuable give and take….I think different people will want different thresholds of this, but it is a technical problem.
In any case, this will not solve the problem outside of mathematics. The extent to which reason is based on the rational as opposed to the logical is the same extent to which an articulate, knowledgeable, experienced interlocutor, which can well be an LLM, can convince one of everything.
When I do talk to LLMs about things outside of mathematics, it is similar to my experience with journalists — about the the things of which I know, I notice misstatements, some serious, and about the others, I feel grateful to the authors from whom I’ve learnt. It can be very hard to be critical about things of which one is ignorant — and it does not come easily to many.
The implications of this will be different for theoretical sciences, for experimental sciences, technology and politics — and, needless to say, for therapy and the arts. For all of these, I am way more concerned than I am for mathematics.
By the way, the proposal in the memo was rejected, and that’s for the best. My deep ambivalence returned a few days after I submitted it.
Received 11 August 2026.
Add to the discussion