Source-linked AI summary

Mathematics in the age of AI

Terence Tao

arXiv:2608.16753v1math.HO

TL;DR

The paper asks how mathematics should respond if AI tools become capable of performing research-level tasks, focusing on the goals and values that govern mathematical research. It uses problem solving as a case study and argues that AI-driven proof abundance requires greater emphasis on proof digestion rather than proof generation. The paper also notes that current AI proof exposition often obscures difficult steps and lacks useful literature context.

  • Problem

    The paper asks what goals and values mathematical research actually optimizes, beyond the explicit goals communicated to the public, students, and funders.

  • Method

    The essay conditionally assumes reasonably strong AI capabilities and examines problem solving as a case study for inspecting mathematics’ implicit goals and practices.

  • Results

    AI could shift mathematics from proof scarcity to proof abundance, overwhelming verification, exposition, peer review, and canonicalization systems designed for scarcity.

  • Takeaways & Limitations

    Mathematical culture should place less emphasis on proof generation and priority, and more emphasis on exposition, refereeing, publication, and canonicalization.

  • Takeaways & Limitations

    Current AI-generated mathematical exposition often obscures novel arguments and fails to situate results in prior literature or provide a useful high-level overview.

Abstract

from arXiv · show

An essay, based on a public lecture delivered at the 2026 International Congress of Mathematicians, on how the mathematical community might respond to the arrival of artificial intelligence tools that are capable of performing research-level mathematical tasks. Rather than debating the capabilities of such tools, we condition on the hypothesis that these capabilities will arrive, and examine instead a question that is orthogonal to it: what the goals and values of mathematical research actually are. The problem-solving component of mathematics is used as a case study.

1. A historical prologue

Mathematics once operated on implicit foundations, but early-twentieth-century paradoxes and incompleteness results forced the subject to formalize its objects and rules. The resulting framework is explicit, standardized, inspectable, and trusted, while AI now threatens to unsettle mathematics’ implicit values and practices.

  • Before the twentieth century, mathematicians generally worked with sets, numbers, infinities, and proofs without precisely specifying their foundations.
  • Russell’s paradox and Gödel’s incompleteness theorems exposed contradictions and limits in previously implicit assumptions about mathematics.
  • The foundational crisis produced an explicit, rigorous, standardized framework whose objects and reasoning rules can be inspected, taught, and mechanized.
  • AI is stress-testing mathematics’ implicit values and practices, including what counts as a contribution, what gets rewarded, and who or what is credited with the work.
  • The essay argues that making these unwritten goals explicit and codified could leave the mathematical community stronger and more resilient.

2. The motivating question

The essay asks how the mathematical community should respond to modern AI technologies and their claimed or actual ability to perform mathematical tasks. It treats this as a community-wide question that cannot be settled by mathematical logic alone.

  • The motivating question concerns how mathematics should respond to AI technologies capable, or purportedly capable, of performing mathematical tasks.
  • The question belongs to the entire community, and the essay presents an approach rather than claiming to provide all the answers.
  • The question is metamathematical, political, ethical, sociological, and cultural rather than purely mathematical.
  • The essay borrows mathematical language such as conjectures and hypotheses to clarify the question’s structure.

3. The first subquestion: AI capability

The first subquestion is what AI tools will actually be able to do in mathematical research, but the essay sets capability debates aside because its focus is the community’s response under an assumed capability scenario. Existing evidence is difficult to interpret because it is often uncontrolled, biased, and incompletely reported.

  • The AI capability subquestion concerns what tools will actually be able to accomplish, represented as a family of conjectures with many free parameters.
  • The capability conjecture ranges from some AI tools completing some research-level tasks with limited success to stronger forms that could challenge current mathematical culture.
  • Public debate has focused mainly on which versions of the AI Capability Conjecture are true, making constructive discussion of community response difficult while the issue remains disputed.
  • The scope excludes AI’s impacts beyond mathematics, which the essay identifies as too broad to address.
  • Evidence about AI capabilities is often uncontrolled and affected by reporting bias, undisclosed costs and variables, and conflation of truth with desirability.
  • The essay is not about establishing the capability conjecture; it briefly cites the First Proof project as an independent assessment using novel research-level problems.

4. The complement to the capability conjecture

The essay isolates the community’s goals and values as the component of the AI-response question orthogonal to AI capability. It conditionally assumes that reasonably strong AI capabilities will soon arrive and analyzes the consequences without requiring belief in that assumption.

  • The essay studies the “orthogonal complement” of AI capability: the goals and values that shape mathematics’ response to capable AI tools.
  • Its Working Hypothesis assumes that AI tools will soon perform a reasonable fraction of research-level mathematical tasks with reasonable success, quality, supervision, and cost.
  • The analysis is conditional: readers need not want, believe, or accept the Working Hypothesis, and evidence for or against it is orthogonal to the discussion.

5. The orthogonal subquestion: our goals and values

Conditioning on the Working Hypothesis brings the goals and values of mathematical research into focus as a fundamental question. The essay argues that AI-driven excessive optimization could make historically aligned goals diverge.

  • The central question asks which explicit and implicit goals, objectives, and values the mathematical community actually optimizes in practice.
  • The essay proposes examining this question directly because AI may remove mathematics’ former ability to delegate it to the humanities.It argues that such examination would be valuable regardless of whether the Working Hypothesis holds.
  • The proposed partial goal list includes solving problems, developing theories, understanding the world, sustaining community, training mathematicians, accumulating knowledge, and creating aesthetic works.
  • Historically, progress toward one mathematical goal often served as a usable proxy for progress toward the others.
  • AI tools heighten Goodhart’s-law risks because they optimize for satisfactory-looking outputs and because industry incentives reward demonstrable, benchmarkable achievements.These technical and economic pressures target metrics that have historically functioned as proxies.
  • Excessive optimization for one or two goals may cause mathematics’ previously aligned goals to diverge from one another.Figure 2 presents this divergence as an intentionally oversimplified diagram.

6. A case study: problem solving

The essay treats problem solving as a sequence of implicit goals extending beyond generating solutions: verification, exposition, community acceptance, and incorporation into definitive theory. AI intensifies the need to make these stages explicit because correct proofs may remain opaque, poorly situated, or undigested.

  • Goal 6.1 to Goal 6.2: Problem solving begins with generating solutions to unsolved problems, but this objective alone produces many incorrect purported proofs.The essay cites the steady stream of purported proofs of the Riemann hypothesis as evidence that maximizing solution flow is inadequate.
  • Goal 6.1 to Goal 6.2: The second goal adds correctness: solve unsolved problems and verify the solutions.
  • Goal 6.3: AI-generated proofs can be formally correct yet incomprehensible, motivating a goal that also requires clear communication and community understanding.Formal verification removes dependence on author reputation, but it does not ensure that humans understand the resulting proof.
  • Goal 6.3: Proof exposition remains difficult because AI writing may be flawless in form while obscuring novel steps, omitting literature context, and erasing informative friction.Human exposition's awkwardness and careful treatment of difficult steps can signal where readers should slow down.
  • Goal 6.4: A proof contributes to its field only when mathematicians digest, accept, and incorporate it into collective knowledge, not merely when it is correct and readable.Human editors and referees provide the slow external process through which individual achievements become collective progress and understanding.
  • Goal 6.5: The final proposed goal extends the pipeline through canonicalization: results enter definitive theory in natural generality, with appropriate proofs, connections, and textbook integration.The essay describes canonicalization as the slowest, least AI-optimizable, and most valuable stage, while noting that AI depends on canonical theories as training data.
  • Pipeline implications: The five-stage pipeline may be oversimplified, but deconstructing problem solving exposes implicit subgoals and possible impedance mismatches throughout the process.

7. Proof scarcity and proof abundance

Under the Working Hypothesis, AI-generated mathematical proofs could outpace verification, exposition, peer review, and canonicalization, producing a transition from proof scarcity to proof abundance. These pressures would intensify stresses already caused by literature growth, increasingly long and specialized proofs, and strain on refereeing.

  • Proof abundance: AI-generated proofs may accumulate faster than they can be verified, written up readably, reviewed, or incorporated into definitive form.The resulting impedance mismatches occur at multiple stages of the mathematical pipeline.
  • Proof abundance: The community could transition from an era of proof scarcity to one in which proofs are abundant but difficult to digest.
  • Existing stresses: AI will markedly exacerbate pre-existing stresses from literature growth, increasingly long and specialized proofs, and strain on the refereeing system.

8. From goals to recommendations

The essay connects identified goals of problem solving to recommendations for adapting mathematical practice under AI. It endorses disclosure, review support, human authorship and attribution, and greater emphasis on proof digestion and explainability.

  • Recommendations: The essay uses the Leiden Declaration on Artificial Intelligence and Mathematics as a source of recommendations for individual mathematicians.
  • Disclosure and review: Mathematicians should transparently disclose automated tool use, including large language models, machine learning systems, proof assistants, and mathematical software.
  • Disclosure and review: Authors should support reviewing by disclosing tool use, providing precise references, and supplying formal proofs where feasible and appropriate.
  • Proof digestion: Mathematical culture should place less emphasis on proof generation and being first, and more on exposition, refereeing, publication, and canonicalization.
  • Human authorship and attribution: Credit and responsibility should remain with human authors, while authors should proactively identify and credit sources and disclose unresolved attribution problems.
  • Explainability: A formally verified result should not be published if its authors cannot give a clear, correct, expert-level, and properly attributed account of it.

9. Closing thoughts

The essay argues that AI's effects should be examined across teaching, mentoring, hiring, grants, refereeing, and outreach, with responses tailored to each domain. It calls for open community discussion and public engagement about AI capabilities, goals, and values.

  • Broader implications: AI's potential impact extends beyond problem solving to teaching, mentoring, hiring, grant applications, refereeing, and public outreach.
  • Broader implications: Responses should differ by domain: education may require tight restrictions, while other areas may require mathematicians to define their own AI best practices.The essay states that producing correct homework is not sufficient to train a mathematician.
  • Community response: The mathematical community should hold open and honest discussions about AI capability and about its own goals and values.
  • Community response: Mathematicians should participate in public discourse by explaining and contextualizing AI-assisted methods and results, especially where specialized expertise is needed.

Appendix A. Some new workflows and infrastructures

The appendix points readers to existing projects that illustrate new mathematical infrastructure. These include formalized mathematics, peer-reviewed video communication, problem and constants databases, competitions, and the First Proof project.

  • Examples: Mathlib illustrates a unified formalized library of mathematics.
  • Examples: Mathematical Discourse illustrates a peer-reviewed video journal for research talks.
  • Examples: The Erdős problems database and optimization constants database illustrate repositories for mathematical problems and optimization constants.
  • Examples: The SAIR Foundation mathematics competitions and First Proof project illustrate infrastructure for mathematical competitions and research-level problem assessment.
Loading 2608.16753v1…