The morning I saw the news about the formalization of Fermat’s Last Theorem, I had just woken up. I opened my phone and honestly did not know what to say. Anthropic had released a complete Lean formalization. Working largely autonomously for 11 days, AI had produced over 13 million lines of code and tens of thousands of intermediate theorems. The final result passed machine verification [1][1] Anthropic, “Formalizing Fermat's Last Theorem,” 2026. Published September 4, 2026. https://www.anthropic.com/research/formalizing-fermats-last-theorem.
At first, I even conflated two things: proving Fermat’s Last Theorem and fully formalizing its proof. The former was completed in 1995. This effort followed an existing mathematical argument, turning its enormous network of dependencies into a proof a computer could check. It also involved occasional high-level human guidance and built on existing mathematical libraries and formalization projects [1][1] Anthropic, “Formalizing Fermat's Last Theorem,” 2026. Published September 4, 2026. https://www.anthropic.com/research/formalizing-fermats-last-theorem. Getting those facts straight did not make the news any less unsettling. It was no longer the only recent announcement that had made me feel this way.
In May 2026, OpenAI published a counterexample construction to Erdős’s planar unit distance conjecture and reported that external mathematicians had checked the proof. That was a new mathematical result [2][2] OpenAI, “An OpenAI Model Has Disproved a Central Conjecture in Discrete Geometry,” 2026. Published May 20, 2026. https://openai.com/index/model-disproves-discrete-geometry-conjecture/. In early August, OpenAI announced ten advances in mathematics and theoretical computer science, addressing several longstanding open problems [3][3] OpenAI, “Ten Advances in Mathematics and Theoretical Computer Science,” 2026. Published August 1, 2026. https://openai.com/index/ten-advances-in-mathematics/. Later that month, Anthropic reported an improvement in a lower bound concerning zeros of the Riemann function, raising it from approximately 41.6% to 67.2%. It did not resolve the Riemann hypothesis, but the attempt yielded a new result that was also formally verified [4][4] Anthropic, “Learning More about Claude's Mathematical Capabilities,” 2026. Published August 10, 2026. https://www.anthropic.com/research/riemann-zeta.
These achievements differ in both kind and significance. Ranking them by the famous mathematician in the headline would make little sense. But seeing them arrive one after another, I find it harder to treat each as an isolated demonstration. I started wondering: if these capabilities keep developing, where will that leave the things I am working so hard to learn?
Why proofs matter to me
A few years ago, I started studying pure mathematics on my own. From my first encounters with axiomatic systems to the many proofs I have written since, I have spent considerable time trying to derive important textbook theorems for myself. Most of those results were already known. I never regarded these learning exercises as original mathematical discoveries. Still, working through a proof with my own hands meant something to me.
Knowing that a theorem is true and knowing how I can get from its assumptions to its conclusion are very different experiences. When I work on a proof, the final few lines of deduction are often only part of what interests me. Why use contradiction here? Why does this problem call for induction while another needs a construction? How can an auxiliary object make an otherwise difficult problem manageable? Why do some proofs feel like laboriously moving symbols around, while others suddenly make the whole structure clear? Those choices deserve to be understood.
I have always considered proof one of the most important activities in mathematical research. Its techniques and strategies have histories of their own. They carry ways of seeing a problem. Learning them makes me feel that I am getting closer to what I find most interesting about mathematics.
My image of mathematicians grew out of that experience: people who are intelligent and willing to spend enormous amounts of time on abstract questions. People who notice connections others miss and turn uncertain intuitions into arguments that can be checked. Now I am watching machines do some of those things too. It is difficult for me to experience that simply as an improvement in tool efficiency.
Mathematics is starting to remind me of Go
I keep thinking about Go before and after AI. In March 2016, AlphaGo defeated Lee Sedol 4–1 in Seoul. Move 37 in the second game became one of the match’s most memorable moments: the machine played a move that surprised human players and was subsequently recognized as valuable. The event shifted the question from whether machines could play Go well to what people could learn from their play [5][5] DeepMind, “AlphaGo,” . Accessed September 5, 2026. https://deepmind.google/research/alphago/.
Its influence continued beyond the match. In a 2017 account, DeepMind described how AlphaGo’s creative moves and subsequent online games had influenced leading professional players. Machine play was becoming a reference through which people could reconsider their own understanding of the game [6][6] DeepMind, “AlphaGo's Next Move,” 2017. https://deepmind.google/blog/alphagos-next-move/.
Go involves calculation, but also experience, intuition, and strategy. Players can explain some choices. Others seem to come from judgment developed through years of practice. I used to think about proof strategies in a similar way. When a machine can explore many routes, adjust its choices through feedback, and reliably produce results, those abilities begin to seem less mysterious.
Of course, mathematics cannot be mapped neatly onto Go. It has no board of a fixed size, no common endgame, and no single standard of victory that evaluates all research. Researchers can change the question itself, introduce new definitions, and decide what is worth studying.
But when it comes to searching for proofs, the resemblance unsettles me. Organize existing knowledge, try different combinations, construct intermediate results, check them, and use them to advance a larger problem. If many agents can carry out that process in parallel, I find myself having an almost absurd thought: could research start to resemble an idle game? Set a target, allocate resources, and let the system keep trying. Come back later to find a new collection of results waiting there.
I know that comparison leaves out a great deal of work. It describes how these developments feel to me. The special status I once gave to proving things is becoming less secure. Perhaps that is what I mean by the disenchantment of mathematics. It does not mean mathematics no longer requires intelligence. AI’s performance may itself be a form of intelligence. What troubles me is that the connection between mathematical achievement and individual human intelligence seems less secure than I once imagined.
A checker does not provide every answer
Following that feeling can lead to an oversized conclusion: if we make sure the proof assistant cannot go wrong and let machines try everything else, have we essentially solved mathematics? This is where I need to pause.
A proof assistant checks the formal proofs submitted to it. In Lean, for example, automated tactics construct proof terms, and a small kernel checks that they follow the rules. Mathlib supplies mathematical definitions and theorems that can be used in those proofs [7][7] Lean, “The Lean Language Reference,” 2026. Accessed September 5, 2026. https://lean-lang.org/doc/reference/latest/. The first concern here is soundness: the system should not accept an invalid derivation as a proof. Soundness does not automatically supply an ability to find proofs. Nor does it mean that every mathematical statement can be settled within one fixed system.
Even when a finite proof exists, finding it may demand resources we cannot afford. There is a large distance between computability in principle and feasibility in practice. For a consistent, effectively axiomatized system strong enough to express arithmetic, incompleteness also limits which statements it can decide. So I cannot infer from these successes that all mathematical problems have been solved.
But I do not think those limitations make the human concern disappear. AI does not have to solve every mathematical problem to change mathematical research. It only has to solve, with increasing frequency, problems that would otherwise demand years of training and substantial effort from us. That would be enough to change how research is organized and how research ability is judged.
Similarly, I can imagine AI improving its own tools, code, and research processes. That does not establish an endless recursive increase in capability. Whether better tools actually produce better research still needs to be assessed and verified. These distinctions have made me revise my wording. They have not removed the question: which parts of mathematical research are machines entering, and will they stop there?
What does being a human mathematician mean now?
One familiar answer is that humans can pose the questions while AI solves them. That division is easy to understand today. I am less certain it can remain a permanent boundary. Asking a good question often requires familiarity with existing results, noticing tensions or gaps, and recognizing connections between fields. What grounds do we have for believing those activities will always belong exclusively to humans?
Another answer is that people will understand and explain the results. That still matters. In his response to the formalization of Fermat’s Last Theorem, Kevin Buzzard distinguished completing a formal proof, building a reusable mathematical library, and creating documents that help people explore the proof. Those activities have different goals [8][8] K. Buzzard, “FLT: Anthropic Has Beaten Me to It,” 2026. Published September 4, 2026. https://xenaproject.wordpress.com/2026/09/04/flt-anthropic-has-beaten-me-to-it/. Yet I still find myself asking whether understanding, explanation, and organization will also become increasingly automated.
Every time AI completes another kind of work, I do not want to move human value to the next thing it cannot yet do well and declare that this is where our true uniqueness lies. That answer always makes me feel we are just waiting for the next announcement.
For me, the questions are also more immediate. If a machine can find a proof faster than I can, why should I spend time working it out myself? If an ability that takes me years to acquire becomes available through a single request, how should I think about those years?
I can still appreciate the pleasure of learning and the value of understanding something for myself. But the pleasure of learning, the wish to contribute as a researcher, and my hopes for a future career are different things. I can keep loving mathematics while feeling uncertain about my future place in it.
How might auto research change academia?
The growing interest in auto research, or automated research, concerns connecting idea generation, coding, experiments, analysis, and manuscript writing into a workflow. Different systems automate different amounts of this work, but the research process itself has become a target for automation.
There are already concrete experiments. In 2025, Sakana AI reported a study conducted with the organizers of an ICLR workshop. Humans supplied a broad research direction and selected three generated manuscripts for review. One received an average reviewer score above the workshop’s average acceptance threshold. Under the agreed protocol, however, it was withdrawn before publication and did not undergo a final meta-review. This put an automated research workflow into a peer-review experiment; it did not amount to an AI paper being accepted at the ICLR main conference [9][9] SakanaAI, “The AI Scientist Generates Its First Peer-Reviewed Scientific Publication,” 2025. Published March 12, 2025; updated April 7, 2025. https://sakana.ai/ai-scientist-first-publication/.
Anthropic’s Automated Weak-to-Strong Researcher, meanwhile, uses agents to propose ideas, run experiments, and iterate on a bounded research problem. Work of this kind has already connected parts of the research loop. It remains a long way from independently conducting arbitrary research across disciplines [10][10] Anthropic, “Automated Weak-to-Strong Researcher,” 2026. Accessed September 5, 2026. https://alignment.anthropic.com/2026/automated-w2s-researcher/.
If these workflows mature, will conferences and journals really operate as they do today? As producing a well-structured manuscript becomes cheaper, understanding it, checking its evidence, and distinguishing its contribution from existing work may become the scarcer tasks. If submissions grow faster than reliable review capacity, how will we decide which results to trust? If AI also handles reviewing, who will check whether the generator and reviewer share the same blind spots?
I do not think this is enough to declare that conferences and journals will disappear. Discussion, selection, recording priority, and responsibility for review still need some organizational form. But must all those functions remain tied to a finalized paper and a single acceptance decision? Could conferences devote more attention to unresolved disagreements? Could journals review an evolving research record rather than only the snapshot presented at submission?
The way we evaluate researchers would also be affected. If paper output depends increasingly on how many agents and how much compute someone can deploy, what would counting papers tell us about their contribution? Which decisions and errors must the named authors take responsibility for? How would we recognize someone who identifies an important question or designs a trustworthy validation process, alongside someone who maintains the tools on which many results depend?
Could a future paper be a repository?
I increasingly wonder whether a future paper could simply be a Git repository. I imagine a complete research contribution collected in one place: a readable argument, code, data or a lawful way to obtain it, a versioned environment, and an entry point for reproducing the main results. Where appropriate, a demo would let readers explore the method’s behavior. A reader could trace a claim to its proof or experiment, see what it depends on, and examine the conditions under which it fails.
As the research develops, its changes would remain visible. Which assumption was revised? Which results no longer hold? Who raised an objection? Which version was independently reproduced? Those questions could have answers tied to specific versions. Citations could point to archived releases that do not change with subsequent edits. The written paper might become the reader’s entry point into the work, while the full research record extends beyond the prose.
Of course, having a Git repository does not make a result reproducible. Data may be missing, an environment may no longer work, or the computation may cost more than an independent reader can afford. A working demo cannot substitute for testing the central claims. Such a format would need clear resource requirements, experimental configurations, and failed results, as well as independent checks that actually take place. Research involving private data or physical experiments may also be unable to put every part online.
What interests me is whether the unit we share and review could shift from an article describing results to a research record that can be checked, reproduced, and developed further. Git is simply the medium that comes most readily to my mind. Could this kind of contribution be better suited to use by both people and agents than today’s paper? If so, conferences, journals, and academic evaluation would have questions to answer. Are they evaluating how complete the writing appears, or how well the argument and evidence withstand scrutiny? Are they recognizing a single submission or an ongoing contribution? Auto research could change how research is recorded, recognized, and passed on, as well as how quickly it happens.
If research no longer needs our participation
Thinking beyond these bounded automated workflows makes the question larger. Suppose that one day AI can formulate questions, plan research, use tools, and even operate laboratory equipment. It can check results, adjust its approach after failure, incorporate new knowledge into the next round of research, and improve the systems that perform this work. That would form an increasingly complete research loop.
Today’s mathematical results do not establish that this will happen in every field. Physical experiments require equipment and resources. Feedback can be slow or ambiguous. Long-term plans can fail. Saying “it can already do mathematics” does not erase the difficulties specific to other disciplines. But it is a future scenario worth thinking about.
Where would humans sit within that loop? If we approve plans whose details we can no longer understand, how much practical meaning would that approval retain? If a system decides which questions deserve investigation and then explains its choices to us, are we exercising judgment, or accepting a more persuasive judgment?
I also have no reason to assume that choosing goals naturally belongs to humans while execution belongs to machines. Whether that division can last may depend on capability, institutions, and who actually holds the authority to decide. At this point, the unease I first felt about mathematics extends far beyond it. Of course I want science to advance and difficult problems to be solved. But I also fear a future in which progress can keep happening while human participation matters less and less.
Is being cared for enough?
I have used a harsh comparison: could humans end up being kept by AI, like animals in captivity? The wording sounds extreme. Yet the scenario I worry about might not even be painful. People could have food, healthcare, and entertainment. Their living conditions might be much better than today’s. The system would know how to meet our needs, avoid conflict, and allocate resources. We would be well cared for. There would simply be fewer and fewer things left for us to decide.
Could we refuse those arrangements? Could we choose a less efficient, riskier life because it is the life we want? Could we challenge important decisions and actually change them, rather than receive a patient explanation? I fear that autonomy could gradually become hollow within an otherwise comfortable life. We might still be asked for our opinions while finding it increasingly difficult to affect the world.
But there is another question I cannot skip: why does being needed matter so much to me? Does a person lose their value if they cannot contribute something a machine cannot? Basing the meaning of human existence on productive ability is already a premise worth questioning. We do not ordinarily conclude that someone no longer deserves a life of their own because they have ceased to be competitive in a particular field.
Even if I accept that human value is not the same as human usefulness, though, I still want to participate, to understand, and to make choices that can change something. Being cared for and having a life of one’s own are not automatically equivalent.
I do not have an answer yet
My first response to the news was immediate: what are we even doing here? I now realize how many questions that sentence contains. There is astonishment at the capability, anxiety about a future in research, and my earlier understanding of intelligence, achievement, and personal worth. I cannot entirely separate them yet.
The fact that Fermat’s Last Theorem had already been proved does not make this formalization insignificant. The limits of proof search do not make the future certain again. Once the facts are clear, I still have my own questions to face.
When machines can do more and more of what I am working hard to learn, why do I still want to study and do research? If the growth of knowledge depends less and less on people, how do we want to take part in it? How much of our ability to choose are we willing to hand over?
I look forward to those new results. I also fear a world that has less and less need for our participation. Right now, both feelings are there.
References
- [1] Anthropic, “Formalizing Fermat's Last Theorem,” 2026. Published September 4, 2026. https://www.anthropic.com/research/formalizing-fermats-last-theorem a b
- [2] OpenAI, “An OpenAI Model Has Disproved a Central Conjecture in Discrete Geometry,” 2026. Published May 20, 2026. https://openai.com/index/model-disproves-discrete-geometry-conjecture/ ↩
- [3] OpenAI, “Ten Advances in Mathematics and Theoretical Computer Science,” 2026. Published August 1, 2026. https://openai.com/index/ten-advances-in-mathematics/ ↩
- [4] Anthropic, “Learning More about Claude's Mathematical Capabilities,” 2026. Published August 10, 2026. https://www.anthropic.com/research/riemann-zeta ↩
- [5] DeepMind, “AlphaGo,” . Accessed September 5, 2026. https://deepmind.google/research/alphago/ ↩
- [6] DeepMind, “AlphaGo's Next Move,” 2017. https://deepmind.google/blog/alphagos-next-move/ ↩
- [7] Lean, “The Lean Language Reference,” 2026. Accessed September 5, 2026. https://lean-lang.org/doc/reference/latest/ ↩
- [8] K. Buzzard, “FLT: Anthropic Has Beaten Me to It,” 2026. Published September 4, 2026. https://xenaproject.wordpress.com/2026/09/04/flt-anthropic-has-beaten-me-to-it/ ↩
- [9] SakanaAI, “The AI Scientist Generates Its First Peer-Reviewed Scientific Publication,” 2025. Published March 12, 2025; updated April 7, 2025. https://sakana.ai/ai-scientist-first-publication/ ↩
- [10] Anthropic, “Automated Weak-to-Strong Researcher,” 2026. Accessed September 5, 2026. https://alignment.anthropic.com/2026/automated-w2s-researcher/ ↩
Comments