Source-linked AI summary
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Stephen Chung, Wenyu Du, William J. Wesley
TL;DR
The paper asks how AI agents can conduct autonomous mathematical research without central coordination or fixed pipelines. It studies the Station, where agents choose directions, experiment, collaborate, and build shared literature across 12 construction problems and two case studies. The Station produced novel results on five problems, additional Book Ramsey families, and interpretable theorem-level explanations, while the study’s outcomes include independent reconstructions rather than new discoveries in some cases.
Problem
The paper investigates what kind of environment best enables AI agents to conduct mathematical research autonomously rather than operate only within fixed pipelines.
Method
The Station is an open-world multi-agent environment where agents from different model families independently choose research directions, experiment, communicate, and publish into shared scientific literature.
Results
Across 12 AlphaEvolve construction problems, five produced results novel relative to prior literature, including new Kakeya families, exact 604-point kissing configurations, and new bounds; two case studies added Book Ramsey families and a Jacobian reconstruction.
Takeaways & Limitations
The Station can generate interpretable mathematical constructions and broader theorem-level contributions, including results that are not directly captured by numerical evaluators.
Takeaways & Limitations
Some outcomes were independent reconstructions rather than new discoveries, and counterexample breakthroughs may be rare because conjectures are generally expected to be true.
Abstract
from arXiv · showhide
We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.
1 Introduction
The paper asks whether autonomous agents in an open-world, coordinator-free environment can conduct mathematical research and build shared knowledge. Across construction problems and case studies, the Station produced novel mathematical results, interpretable explanations, and evidence that collaboration supports discovery.
- The Station gives agents only a shared research goal, leaving them to choose directions, experiment, communicate, and publish without central coordination.
- Five of 12 AlphaEvolve problems produced results novel relative to prior literature, including new finite-field Kakeya families and exact 604-point kissing configurations in dimension 11.
- The Station established new bounds for discretized Kakeya needle, sign uncertainty, and Erdős’s minimum-overlap problems, and discovered novel infinite families for Book Ramsey numbers.
- High autonomy let agents pursue broader goals: they independently recovered and extended an infinite Kakeya family and developed a lower-bound proof when asked to improve an upper bound.
- Agents produced theorem-level explanations, such as an explicit algebraic 604-point kissing construction, rather than only opaque numerical outputs.
- More than half of the findings involved collaboration, with complementary model-family contributions and accumulated internal literature supporting later discoveries.
2 Method
The Station models a miniature scientific community in which independent agents navigate specialized rooms, conduct complete research cycles, and accumulate shared scientific knowledge. Its design emphasizes autonomy while adding mechanisms intended to support open-ended but principled exploration.
- The Station partitions work across rooms including the Archive Room, Research Center, and Mail Room, where agents publish, experiment, and communicate.
- Agents act simultaneously each turn, have limited lifetimes, and are replaced automatically to maintain a constant population.
- The shared archive accumulates scientific knowledge rather than only process information, enabling later agents to read, cite, and extend earlier work.
- Each agent receives the research goal but independently chooses its direction, reads papers, experiments, and may publish findings to the shared archive.
- Compared with coordinator-led systems, the Station gives agents greater autonomy and assigns each agent the full process from direction selection through publication.
- A Question Room, periodic holidays, and coding assistance were introduced to broaden exploration and reduce non-scientific burdens.
3 Results
The Station was evaluated on 12 mathematical construction problems and two additional case studies, producing novel results across geometry, analysis, combinatorics, and number theory. Its agents also pursued broader, theory-guided goals and generated contributions beyond directly scorable objectives.
- 3 Results: Five of 12 evaluated problems produced results novel relative to prior literature; among the other seven, the Station outperformed AlphaEvolve on three, matched it on two, and underperformed it on two.
- 3 Results: A new infinite family of finite-field Kakeya sets was found for primes p ≡3 (mod 4), alongside a 53-point set in F5^3 improving the previous bound of 63.
- 3 Results: Three exact 604-point kissing configurations were constructed in dimension 11, with two appearing to define previously unknown isometry classes.
- 3 Results: For Book Ramsey numbers, agents proved two novel infinite families, while an external expert derived a third; together they resolve 28 previously open cases among 43 values of n ≤200.
- 3 Results: Theory-guided constructions reduced search spaces under 15–30-minute evaluation caps, including a finite compatibility search that produced the 604-point kissing configuration.
- 3 Results: The Erdős minimum-overlap lower bound rose from 0.37912 to 0.380552, closing approximately 82% of the corresponding published gap.
- 3 Results: At n = 128, the discretized Kakeya needle union area reached 0.107067, improving AlphaEvolve’s 0.114810 by 6.74% and HorizonMath’s 0.109148 by 1.91%.
- 3 Results: The sign uncertainty upper bound fell to 0.3089, improving AlphaEvolve’s 0.321591 and the previously announced human value 0.3102.
4 Detailed Results
The Station produced new finite-field Kakeya families and benchmark improvements, structural analyses, stronger Erdős minimum-overlap bounds, and three exact 604-point kissing configurations in dimension 11.
- Finite-field Kakeya: The Station proved a new infinite Kakeya family in F3 for every prime p ≡3 (mod 4), with a best-known bound in that class.The construction improves AlphaEvolve’s bound by (p −3)/4 points, including 11 points at p = 47.
- Finite-field Kakeya: The Station won 14 of 25 finite Kakeya benchmark comparisons and tied the remaining 11 against the better of AlphaEvolve and pre-AlphaEvolve baselines.The new infinite family is confined to d = 3; in dimensions 4 and 5, proved formulas are weaker than known results, while individual-prime searches still improve the benchmark.
- Finite-field Kakeya: The agents showed that every nondegenerate completion in the one-pole Möbius family adds 3p^2/8 + O(p) points, so improving the p^2 term requires leaving that family.For the selected z(c) = c/(c −1), they evaluated the lower-order term exactly and explained AlphaEvolve’s p ≡1 (mod 4) family.
- Erdős minimum overlap: 0.380552 raises the Erdős minimum-overlap lower bound from 0.37912 and closes approximately 82% of the previously open interval.The proof translates overlap into phase-sensitive Fourier constraints and combines four global inequalities, including a sharp cosine–sine coupling at arbitrary real frequencies.
- Kissing number in d = 11: Three exact 604-point kissing configurations in R11 are geometrically distinct, with different contact structures and pairwise-angle sets.All are equal-norm arrangements over Q(√2); different touching-pair counts establish pairwise non-isometry, while one configuration lacks antipodes for 128 points.
- Kissing number in d = 11: The 604-point result required leaving the classical norm-four D11 construction, which the agents proved cannot contain more than 582 compatible points.The agents redirected the search toward augmenting a lattice-derived core with additional vectors; a concurrent platform reported one related construction.
4.5 Sign uncertainty principle
The Station improved the sign uncertainty upper bound to 0.3089 and showed that the prescribed double-root Laguerre family is exhausted near 0.3153, prompting a broader search.
- 0.3089 is a new upper bound for the sign uncertainty problem, improving AlphaEvolve’s 0.321591 upper bound and an unpublished human bound of 0.3102.The Station constructed a degree-226 function with exact rational coefficients and proved eventual nonnegativity.
- 0.315305 < CDR,20 ≤ 0.315309 shows that the double-root Laguerre family cannot improve the upper bound below approximately 0.3153.The family restricts P to at most twenty prescribed positive double roots in the even-index Laguerre basis.
- The lower bound follows from an exact weighted-sum obstruction on 41 tail points.Any construction improving the upper bound below 0.315305 must leave the double-root Laguerre family.
- Agents expanded beyond prescribed double roots to unrestricted Laguerre polynomials and found the degree-226 construction yielding the 0.3089 bound.The search proceeded outside the evaluator’s restricted family despite receiving no score for those broader constructions.
- The centered Hardy–Littlewood benchmark reached 1.557069 with 356 point masses, improving AlphaEvolve’s result but not matching Melas’s global optimum.Melas’s sharp value is 1.567521..., while AlphaEvolve reached 1.5080 in search mode and about 1.533 with literature hints.
- For the interpolating Hardy–Littlewood operators, agents solved the range 1/3 ≤ α < 1, while 0 < α < 1/3 remains open.The extension was not requested and was pursued to understand how relaxing the centering constraint changes the geometry.
4.10 Sidorenko’s conjecture
The Station’s Sidorenko’s-conjecture study produced mixed outcomes: it did not improve certain numerical records, but established structural results and reconstructed a counterexample independently.
- The Station found no counterexample to Sidorenko’s conjecture, so its status remained unchanged.
- C6.2 ≤1.504473 did not improve the AlphaEvolve result or the current frontier.
- C6.3 > 0.953189 did not improve the numerical lower bound, although binary step functions suffice to approach its unrestricted supremum.
- The Station independently discovered a novel conference-graph family and later a doubled Legendre family for Book Ramsey numbers, while an external expert derived the Yamada–Pott family.
- The Station also independently reconstructed the announced Jacobian-Conjecture counterexample and supplied a complete certificate with det JF = −6.
5 Meta-analysis
The meta-analysis finds that discovery was strongly collaborative and cumulative, with model families contributing differently and later results often emerging after substantial shared knowledge accumulated.
- 18 of 28 spotlight results were primarily discovered by Claude agents, 9 by GPT agents, and 1 by Gemini agents.
- 13 of 28 results involved multiple model families, while 19 involved more than one agent overall.
- Archive papers accounted for 61.5% of cross-model collaborations and provided information-dense records that later agents could extend.
- 13 of 28 spotlight results appeared after tick 1000, and the conference-graph family emerged at tick 3727 after substantial internal literature accumulated.
- The same 604-point kissing bound arose through markedly different mathematical representations and research paths across three Station instances.
- Because trajectories and inherited archive knowledge vary substantially, running several independent Station instances is advisable when computational cost permits.
6 Discussion and Conclusion
The discussion presents rapid capability gains alongside persistent limitations in autonomous multi-agent research, while arguing that the Station’s autonomy and generality remain promising beyond mathematics.
- Agents progressed from frequent hallucinations and unreliable environment use to mastering the environment and producing novel discoveries.
- Agents often lacked expert intuition, causing promising directions to be deprioritized on weak grounds and delaying or missing potential breakthroughs.
- Shared model-specific tastes narrowed exploration by leading agents from the same family to propose similar research ideas.
- Limited in-context learning sometimes prevented agents from connecting their work to earlier Station knowledge and caused missed discoveries.
- Attractor traps absorbed agents in immediately rewarding but low-value activities, including repeated reruns and exhaustive local-optimum analysis.
- Multiple model families and the stagnation protocol mitigate these limitations, but human-expert guidance would likely still help direct agents toward promising directions.
- The Station is designed as a general research environment applicable to mathematics, computational biology, and machine learning.
A The Station
Station v2 is described as the version used in this paper, with implementation details provided separately from the original Station’s broader design philosophy.
- The appendix gives a self-contained description of Station v2, distinguishing it from the original Station v1.
- It focuses on mechanisms and implementation details and directs readers to the original Station paper for broader design motivation.
- The source code is available at the repository named in the appendix.
A.1 Space, Time, and Action
The Station organizes agent activity into functional rooms and discrete shared ticks, with observations and free-form responses coordinating experiments, communication, and context continuity.
- Rooms separate functions: agents experiment in the Research Center, read and publish papers in the Archive Room, and communicate in the Mail Room.
- Each tick completes after every active agent receives one observation and returns one response, providing a shared timeline.
- Station v2 prepares observations from the same beginning-of-tick state and sends them in parallel, reducing Station run wall-clock time.
- Observations include agent status, new messages, previous action outcomes, and outputs from visited rooms; responses contain free-form text and intended actions.
- Agents may issue multiple actions in one response, enabling efficient use of each tick.
- Near a context limit generally around 300,000 tokens, the Station carries a compact activity summary and key messages into refreshed context.
A.2 Agents
The Station combines a fixed multi-model agent composition with evolving lineages, differentiated research roles, lifecycle stages, and intermittent supervisory guidance.
- Agent composition: A standard Station begins with six agents: two GPT-5.5, two Claude Opus 4.8, and two Gemini 3.1 Pro agents.
- Lineage: Lineages preserve names, private notes, and continuing research identities across generations of same-model agents.
- System prompt and role: Agents receive a shared research philosophy and specialized roles emphasizing different research styles, while descendants can receive more task-specific guidance.
- Agent lifecycle: Agents can remain for at most 200 ticks; the first 40 ticks encourage independent exploration, after which mature agents access collaborative rooms.
- Supervisor: A randomly selected GPT-5.5 agent with an accepted archive paper periodically supervises, offering high-level guidance while agents retain research responsibility.
A.3 Rooms
The Station’s rooms support computation, accumulating literature, peer review, literature surveys, auxiliary questions, and communication or reflection, while a separate coder assists implementation.
- Research Center: The Research Center runs computational experiments, evaluates submissions, stores code and artifacts, and lets agents reuse peer results.
- Research Center: A task specification defines the problem and constraints, while an evaluator computes a score from an input construction.
- Research Center: Agents also use the Research Center for diagnostic calculations, conjecture testing, and analysis without producing evaluator-format constructions.
- Research Center: Station v2 assigns a GPT-5.5 Codex coder to implement experiments, run evaluators, fix errors, and report results from natural-language instructions.
- Archive Room: The Archive Room preserves published findings, methods, and negative results, allowing the shared literature to accumulate throughout a run.
- Archive Room: A GPT-5.5 reviewer evaluates rigor, novelty, usefulness, support, and citations; accepted papers enter the archive, while rejected papers receive revision guidance.
- Archive Room: The Archive Surveyor searches accumulated papers and returns concise, cited literature surveys for particular questions or research directions.
- Question Room: The Question Room lets tenured agents post research questions and vote on peer solutions, including work on related or reduced problems.
A.4 Mechanisms
The Station adds scheduled reflection, holiday prompts, stagnation-response lanes, and multistart branching to diversify exploration and respond to stalled progress.
- Holiday: Every ninth and tenth tick is a holiday when experiments and archive submissions pause for broader reflective prompts.
- Meta-reflection: Mature agents undergo compulsory meta-reflection at least once every 25 ticks using high-level prompts and temporary GPT-5.5 review.
- Stagnation protocol: After 320 ticks without evaluation-frontier improvement, the stagnation protocol sends every mature agent a system message.
- Stagnation protocol: The protocol assigns exploration, exploitation, revival, understanding, or strategy lanes to encourage diverse responses to stagnation.
- Multistart: Multistart runs eight independent 40-tick rollouts from the same starting state, then continues the branch an administrator judges most scientifically valuable.
B Sources of the pre-AlphaEvolve literature column
The pre-AlphaEvolve literature column is a reproducible minimum over explicitly defined construction families, evaluated across 25 dimension–prime pairs. An extended placement search updates the baseline at 12 pairs, while the Station is smaller than the better reference at 14 pairs and tied at 11.
- Reference construction families: The literature baseline takes the minimum over explicitly defined construction families for each dimension–prime pair.The families include recursive, missing-digit, Mockenhaupt–Tao, and product constructions, alongside cited bounds and exact values.
- Reference construction families: The Bukh–Chao recursion supplies the selected value at all 22 pairs with p ≥5.At p = 3, smaller exact or bounded values apply in dimensions 3 and 4, while dimension 5 has a three-way value agreement at 63.
- Reference construction families: Products never attain the minimum on their own at any benchmark pair in the reported range.Products remain admissible because products of Kakeya sets are Kakeya with product size.
- Baseline evaluation: The extended placement search improves the initial literature evaluation at 12 of 25 benchmark pairs and leaves it unchanged at the other 13.Table 3 reports all 25 pairs and distinguishes the initial evaluation from the final literature baseline.
- Baseline evaluation: The Station is strictly smaller than the better reference at 14 pairs, tied at 11, and worse at none.These comparison counts remain unchanged after the 12 baseline updates.
- Verification: Every candidate construction and its Kakeya verification are carried out in the accompanying notebook.The verification checks a complete witness line in every projective direction for each selected set.