The question of reasoning traces

Segev Gonen Cohen, PhD student at ETH Zurich

I want to raise a specific question that I think has not received enough attention, and that might become more important over time. As will be clear, it is intimately related to many other pertinent questions, and as such can lead to interesting conversations. 

This post should be read as one long chain of thought (as such, it is written in the first person) and is not meant to be definitive or complete. It doesn’t contain many policy proposals or recommendations, mostly as there are others far more qualified for that, but also because I don’t feel that I have many suitable answers.

In this post, ‘model’ is a catch-all term that includes LLMs, general generative AIs, harnesses (e.g. Codex, Claude Code, …..), agents, and any future potential breakthroughs (e.g. world models, …..). I mention some questions about data originality later, a more intricate analysis should clearly distinguish between data sourced during training-time and inference-time – I don’t do this as I do not currently think it fundamentally changes the questions asked. 

I thank Simon Machado for asking me to write this following a related conversation, and Johannes Schmitt for helpful conversations around it. I wish to congratulate Raphael, Francesco, Simon, and Anand on the launch of an interesting, and hopefully successful, initiative. 

What is Chain of Thought reasoning?

Very roughly, Chain of Thought (CoT) reasoning is a technique that allows models to split up tasks, rather than giving an immediate answer. It was shown to be an effective method by a team in Google Brain in 2022, was utilised in OpenAI’s o1 model in September 2024, and first accessible to many non-paying costumers in DeepSeek’s R1 model release in January 2025, causing the (in)famous ‘DeepSeek moment’. 

Most model calls now (whether on chat-based interfaces or via API calls) automatically apply a chain-of thought process if deemed necessary, and in many cases the user is given access to part of it (loosely phrased as ‘thinking’). More sophisticated harnesses clearly break down tasks into steps and inform the user, often allowing for additional guidance throughout the process. To my understanding, for proprietary models there is no way to access the full CoT (it seems that sometimes a model accidentally leaks it, and it can be bizarre), more on this shortly. 

If we view models as doing ‘thinking’ in a way that is familiar to us, then CoT can be viewed as ‘deep thought’ and ‘reasoning’ (in the sequel I loosely refer to this as Human Chain of Thought (HCoT)). This is of course a weak analogy, and one could also consider tool calls, context, harnesses, specific agent workflows, ‘skills’ developed over time, and countless other factors as relevant to the discussion. I choose to bundle these together to not get bogged down in the details.  

The question I want to ask, in the specific context of AI for mathematics, is this: should this be freely released and publicly available? Should it be provided to interested third parties upon request? Or is there an inherent and unalienable right to keep this hidden? 

I will try to touch upon multiple viewpoints (without endorsing one or the other), I hope this will lead to interesting and useful conversations. 

The potential benefits to accessing the CoT

As I see it, there are some immediate benefits to accessing the CoT. In the case of an AI provided proof or (counter)example, it may provide more insight into the ideas, the methods, and the innovations involved. This can be immensely profitable, often more than the result itself. Having access to the CoT might let us concretely discern the sources of inspiration. This is often very difficult (or impossible) with HCoT. In theory, tracking down the origins of genuinely novel and creative strategies could lead to multiple breakthroughs and leaps in our knowledge. However, this is part of a much broader conversation about what true inspiration is, and if models are genuinely capable of whatever it might mean.  

Analysing the CoT can also give more insight into the lineage of new ideas, and to avoid plagiarism and missing citations (even if accidental). Sometimes, models aren’t that effective at doing this by themselves, as touched upon in Martin Hairer’s blog post.

If, in some instances, the ‘inspiration’ really is merely not citing previously done work, the CoT could reveal it. If it is more significant, amounting to more blatant theft (even if accidental), the CoT might reveal this too. There is clearly a wide range of how serious this is – for example, if one overhears a conversation, doesn’t register it, but this leads to a spark, it would be difficult to claim that they are acting maliciously. If however an AI company trains models on users’ conversation data, and this leads to a breakthrough, many would consider this a far greater infringement. But again, this leads to the question of where inspiration comes from more generally, even with humans. 

One could compare the use of the ‘Acknowledgements’ section in a paper, where authors attempt to track down and thank colleagues that helped their work, also when the specific input isn’t clearly discernible or traceable. 

Another huge benefit to CoT access is that it can be helpful in finding errors in claimed proofs, especially in instances where the model is reward-hacking in some way (for example, knowingly lying to satisfy the user, or hallucinating). Having access to the models’ thoughts can certainly make finding mistakes easier. We are already familiar with this idea, in the same way that having students provide ‘working out steps’ is often beneficial to a grader (also in awarding partial credit, as failed proof attempts are often interesting in their own right). 

How does this compare to current mathematical practices? 

Of course, when reading papers, we don’t normally have access to a human authors’ thinking process. In talks, or more expository texts (including surveys, books, and lecture notes), often more context and framing is given. Whether or not this is comparable to the full HCoT may depend on the person involved, the setting, and the particular listener (an expert is likely to infer far more context into the presenters’ internal thought process than a relative beginner). Oftentimes, the presenter even has an incentive to keep parts of their own thinking hidden, to avoid being usurped to a publication. 

In many instances individual communication with suitable authors can provide far more insight into their own HCoT. Many mathematicians are willing and keen to share specific and often very personal insights, failed attempts, and breakthroughs (and where they came from). This is however immensely constrained by geography, seniority, familiarity,  and other countless social and communal structures. Supervisors might give preference to their students, and collaborators might (and most likely will) share far more with each other than with an interested third party.  

Given the huge disparity in approaches to this question amongst humans, it is unlikely  that a broad consensus will be formed anytime soon. However, as stressed above, there  isn’t a great mechanism (yet, as far as I am aware) to really explain a HCoT. A models’ CoT might be different, already being explicit ‘in silica’, as it were.

Is there a reasonable case to keep CoT hidden? 

Let me start with the answer I expect some in the industry will give. The CoT of proprietary models is part of their intellectual property, and is a product of intensive research, training, and expensive compute allocation. It is natural for companies to restrict the access to this, to avoid aiding their competitors. One could go as far as framing this question as an integral part of the Open vs Closed weights debate, which is an incredibly active one already in the industry (see for example the recently formed ‘alliance for open and secure AI’, and the response by Anthropic). Should the CoT, as a part of how the model functions, be open source or trade secrets that the company can choose to restrict? 

This is even more potent given the (claimed) prevalence of distillation attacks, the process by which an (either adversarial or friendly) organisation uses the responses of a given model to train another at a reduced cost, essentially piggybacking off the initial one. Access to the CoT could reasonably make this much more effective.

One could try to make the claim that having access to the CoT as part of mathematics is not adversarial use and will not be used to extract gain; although it can potentially be used for marketing, or for applications of mathematics that do come with direct financial (or other) benefits. 

Many in the industry are wary of access to the CoT, specifically training on it, due to the potential loss-of-control scenarios associated. The classic analogy is this; suppose that a parent punishes a child (in their formative years) for every minor infraction. Over time, the child may learn to lie and deceive the parent, without necessarily improving their behaviour. This has been observed in models, and will immediately make many queasy about giving away access to the CoT. Of course, loss-of-control scenarios are necessary to discuss, albeit far beyond the scope of this post.  

In what format should the CoT be supplied? 

In cases where CoT is available, it is often a garbled confusing mess, written in ‘Neuralese’. Understanding the CoT is a difficult challenge in-and-of itself, barring some dramatic breakthrough in the field of Mechanistic Interpretability. Some research [Anthropic, 2025] shows that in fact even the CoT doesn’t always faithfully reflect the factors that lead to the models’ answer. 

A potential solution is to have another model read the CoT and write an exposition on it, akin to the approach OpenAI took alongside their recent announcement of ‘ten advances in mathematics and theoretical computer science’. This of course assumes that the author has access to the CoT, and that we take it in good faith that both the author and the model strive to represent this accurately and honestly; I doubt this will always be the case. 

Perhaps a solid compromise could be to release the full prompting history, in the hope that the results would be somewhat reproducible. For example, it was claimed by an Anthropic employee that their own model reproduced solutions to 5 of the 10 aforementioned problems, within 24 hours. While to me this is far less impressive than solving them in the first case, these glimmers of reproducibility (and the ability to glean information from the partial reasoning your own model provides) could be vital in the future. 

Tool calls, search history, and sources accessed could be also included in such a package. However, if we view prompting and harnessing as a skill comparable to a talent, or the product of experience (especially when it involves problem-specific insights from the user or from another model), this might also be controversial.

What next?

Part of the reason that it might be difficult to suggest any concrete policies regarding access to the CoT, aside from technical and feasibility constraints, is that it cuts into the deeper questions of AI personhood. Where one sits on the questions of assigning credit to the models, of holding them liable for errors, or the role of the prompter in the ‘discovery’, will undoubtedly affect their opinions about the CoT. I hope that as conversations into all of these continue, discussions of the role of the CoT are included. 

It seems that currently the most prominent response of the mathematical community as a whole to AI is the Leiden declaration, signed by 3496 signatories (at the time of writing), and endorsed by the IMU and a plethora of prominent members of the community. Unfortunately, they make no mention of the CoT and of potential approaches to the challenges and opportunities it brings; although this can reasonably be interpreted in the framework they propose.

It seems to me that the question of access to the CoT might depend on many factors, and that the community might accept that in many, if not most, instances, providing the CoT is unnecessary. The necessity might be tested on the significance of the proved claim, the novelty of the method, the identity of the associated human authors, the proportion of the work directly attributable to the model, the reproducibility of the prompting process, the general access to the model, and the general environment surrounding the claims made. If the author is a PhD student and the paper is a solid but expected result, the CoT is probably unnecessary. In bigger breakthroughs however (for example on the level of the construction of a non-sofic group), it might be expected that the CoT is supplied. This is especially the case when proofs are provided by prominent AI companies, using models not widely available to the public; and when it is likely that the results are used (at least partially) for marketing purposes, and include contentious  phrasing (that might have been written by the models). 

A separate but closely related issue is that of CoT for unsuccessful runs. It was stated by Noam Brown, a prominent research scientist at OpenAI, that the runs leading to the aforementioned ten problems cost around $2000. Conspicuously missing are the other runs – one might guess that multiple (perhaps even orders of magnitude more) similar attempts at other problems were unsuccessful (see for example Tobias Osborne’s post). Is this something that should be disclosed to the community as part of a result? It seems clear that a human shouldn’t have to disclose multiple failed attempts at a proof; should this be viewed differently? At what scale does this make a difference, or does it only matter when there is something to be gained by a nefarious prompter (although one could argue that there is always something personal to be gained from successful results anyway)? 

There is also a conversation to be had regarding who should have access to the CoT, even hypothetically. Perhaps it should only be accessible to editors and referees, to trusted third-party organisations, or to anyone that specifically requests it (especially if the individual has concerns that their own work was unfairly used)? Would it be fair in some instances for at least the prompters themselves to have access to the CoT? Maybe there is space for a CoT verification body that (probably using its own models) ensures that credit is correctly attributed, as part of a broader overhaul of how AI is used in Mathematics? 

I look forwards to continuing this discussion, and many others, on this virtual common room. 

Disclaimer: In the article I reference the Leiden declaration, of which I am not currently  a signatory. This should not be read as taking any specific position, rather I have not yet  given (to my own fault) sufficient thought to the declaration to decide if I should sign. I  anticipate that this will change in the near future.


Received 11 August 2026.

Comments are moderated. Read our comment policy.

Add to the discussion

Discover more from Proofs and Prompts

Subscribe now to keep reading and get access to the full archive.

Continue reading