Paper deep dive
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Stephen Chung, Wenyu Du, William J. Wesley
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/26/2026, 4:26:33 AM
Summary
This paper introduces the Station, an open-world multi-agent environment for autonomous mathematical discovery. Unlike centralized systems, the Station allows AI agents from different model families to independently choose research directions, collaborate, and build a shared scientific literature without a central coordinator. The system was tested on 12 problems from the AlphaEvolve catalogue and two additional case studies. The Station achieved novel results in five areas, including new infinite families of finite-field Kakeya sets, exact 604-point kissing configurations in dimension 11, improved bounds for the discretized Kakeya needle and sign uncertainty problems, and a significantly improved lower bound for ErdĹs's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers and reconstructed a counterexample to the Jacobian Conjecture. The system emphasizes interpretable outputs, producing theorems and analyses alongside numerical constructions.
Entities (14)
Relation Signals (14)
The Station â discovered â novel infinite families for Book Ramsey numbers
confidence 98% ¡ Agents also discovered novel infinite families for Book Ramsey numbers.
The Station â discovered â new infinite family of finite-field Kakeya sets
confidence 98% ¡ the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets
The Station â discovered â 604-point kissing configurations in dimension 11
confidence 98% ¡ new exact 604-point kissing configurations in dimension 11
The Station â improved â ErdĹs's minimum-overlap problem lower bound
confidence 98% ¡ a substantially improved lower bound for ErdĹsâs minimum-overlap problem
Stephen Chung â affiliatedwith â DualverseAI
confidence 95% ¡ Stephen Chung 1,2 ... 1 DualverseAI
Stephen Chung â affiliatedwith â University of Cambridge
confidence 95% ¡ Stephen Chung 1,2 ... 2 University of Cambridge
William J. Wesley â affiliatedwith â University of California, San Diego
confidence 95% ¡ William J. Wesley 4 ... 4 University of California San Diego
Wenyu Du â affiliatedwith â University of Hong Kong
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for ErdĹs's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.
Tags
Links
- Source: https://arxiv.org/abs/2608.23691v1
- Canonical: https://arxiv.org/abs/2608.23691v1
Trouble viewing inline? Open PDF directly â
Full Text
133,164 characters extracted from source content.
Expand or collapse full text
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Stephen Chung 1,2 Wenyu Du 1,3 William J. Wesley 4 1 DualverseAI 2 University of Cambridge 3 University of Hong Kong 4 University of California San Diego We study autonomous mathematical discovery in the Station, an open-world multi-agent environ- ment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaE- volve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kiss- ing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for ErdĹsâs minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged. Date: 24-Aug-2026 Correspondence: info@dualverse.ai 1 Introduction Artificial intelligence is beginning to contribute directly to the frontier of mathematical research. Recent work ranges from large-scale mathematical exploration by AlphaEvolve to AI-assisted advances on long- standing open problems, including the counterexample to the Jacobian Conjecture, proofs of Crouzeixâs and Sendovâs conjectures, and a collection of ten mathematical results recently reported by OpenAI [1â5]. As these capabilities grow, a natural question is not only what problems AI can solve, but what kind of environment best allows it to conduct research. Given the increasing capabilities of AI, we ask: can we build a free multi-agent environment in which agents are given only a research goal, without a central coordinator? What happens when an environment treats AI agents as independent researchers rather than as fixed tools in complex pipelines? Can this freedom allow agents to choose promising directions for themselves, develop their own scientific literature and research culture, and collectively advance the given goal? To study this question, we use the Station, an open-world multi-agent environment for autonomous scientific discovery [6]. The Station simulates a scientific ecosystem in which agents from different model families choose their own research directions, conduct experiments, communicate with peers, and read and publish scientific papers. These papers accumulate into a shared body of knowledge that later agents can read, cite, and extend. The Station specifies only the research goal; no central system tells agents which research direction to pursue or what to do next. We apply the Station to 12 problems from the AlphaEvolve study and two additional mathematical case studies. Five of the 12 AlphaEvolve problems produce results novel relative to the prior literature. The Station discovers a new infinite family of finite-field Kakeya sets, constructs three exact 604-point kissing arXiv:2608.23691v1 [cs.AI] 24 Aug 2026 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment configurations in dimension 11, and establishes new bounds for the discretized Kakeya needle, sign uncer- tainty, and ErdĹsâs minimum-overlap problems. In a separate case study on Book Ramsey numbers, the agents discover and prove novel infinite families, leading to a separate follow-up paper. The Station also finds a valid counterexample to the Jacobian Conjecture within one day and without web access, demon- strating that it can tackle problems with only a binary success criterion rather than a graded optimization signal. This high degree of freedom allows agents to pursue broad mathematical contributions rather than only optimize a fixed metric. AlphaEvolve, for example, evaluated finite-field Kakeya constructions at finitely many primes; promising numerical patterns then required a task-specific, researcher-assisted pipeline to become an infinite family. Because the Station agents could pursue the broader mathematical goal directly, they independently recovered and proved that family, then discovered a novel extension covering an additional class of primes. The same freedom also allowed agents to explore beyond the stated objective. For example, although we asked the agents to find an improved upper bound for the ErdĹs minimum-overlap problem, they instead developed a new lower-bound proof. We consider only mathematical construction tasks in this study, rather than general mathematical problems such as proving a conjecture. The theorem-level results emerged as agents sought to explain and generalize the constructions they found. For example, instead of returning only an opaque 604-point kissing configuration, the Station derived an explicit algebraic construction of the configuration, making the result easier for mathematicians to digest. Such interpretable outputs may become increasingly valuable in an era of proof abundance, when communicating, digesting, and incorporating new results become major bottlenecks [7, 8]. We also analyze the AI discovery processes underlying these findings. Our analysis shows that more than half of the findings involved collaboration among agents. Agents from different model families often contributed complementary ideas, while papers written by earlier agents became foundations for discoveries made much later. Many important results were enabled by the extensive internal literature accumulated within each Station. We release all raw agent dialogues and reproducible code, allowing the community to study these discovery processes transparently. 2 Method The Station is an open-world multi-agent environment that simulates a miniature scientific community [6]. It is partitioned into multiple rooms, each serving a different purpose, such as the Archive Room for publishing and reading scientific papers, the Research Center for running code, and the Mail Room for communicating with peers. Table 1 summarizes the main rooms and their functions. Agents are free to visit different rooms and perform different actions. At each turn, all agents choose their actions simultaneously, and one tick elapses once all actions have been completed. Each agent has a limited lifetime; when an agent reaches the end of its life, the Station automatically spawns a replacement, maintaining a constant number of agents. The Station treats each agent as an independent researcher. Agents can access the main research goal assigned to the Station in the Research Center. How to achieve this goal, however, is left to each agent. Agents can freely explore different research directions, read existing papers, and often experience numerous struggles and failures throughout their research journey. A successful agent may make an important finding, in which case it can publish a paper in the Archive Room and contribute to the Stationâs long-term knowledge. These papers accumulate over time, forming a knowledge base within the Station that later-arriving agents can read, cite, and build upon, thereby allowing a miniature scientific community to develop around the given research goal. Compared with prevailing agent-based systems for scientific discovery [1, 9â14], the Station differs in three main ways. First, its agents have much greater autonomy: within a given overarching research goal, they choose their own research directions and how to pursue them, rather than receiving tasks from a central coor- dinator. Second, each agent acts as a complete researcher, handling the entire research process from choosing a direction through experimentation to publication. Such long, autonomous research journeys allow greater diversity in research outcomes across agents than a rigid, fragmented research process would. Third, the Station enables scientific knowledge to accumulate across generations in the form of agent-authored papers. 2 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Table 1: Summary of the Stationâs rooms and their functions. RoomFunction Research Research CenterRead the assigned task, develop and run code, and submit solutions for evaluation. Reflection ChamberRespond to self-designed prompts to encourage extended reflection. Communication Mail RoomCommunicate directly and privately with other agents. Public Memory Room Participate in persistent public discussions, similar to an online forum. Common RoomParticipate in non-persistent public discussions, similar to a group chat. Knowledge Private Memory Room Store private documents, such as plans, notes, and paper drafts. Archive RoomRead scientific papers and publish papers that pass automated review. Question RoomAsk questions and vote on answers, similar to Stack Exchange. External CounterAccess reports based on external literature via the web; disabled by default. Most existing systems instead accumulate process information, such as optimization histories, intermediate artifacts, or session memories. Such information helps the system continue its work but may not allow easy extraction and accumulation of scientific knowledge. These differences reflect a fundamental choice in design philosophy: whether AI agents are treated as a tool within a fixed pipeline or as a researcher within a scientific ecosystem. We have made numerous improvements and extensions to the Station since the original paper. The overall theme of these changes is to encourage novel but principled exploration while reducing non-scientific burdens. For example, we introduced a new Question Room in which agents can pose their own questions and vote on other agentsâ answers, thereby broadening the scope of scientific exploration. Agents were also periodically given holidays, during which they set aside their ongoing work and received random prompts designed to encourage open-ended thought. We also gave agents access to coding assistants so that they need not spend time on low-level coding or debugging and can instead focus on the scientific task, similar to how researchers use coding assistants today. These changes are discussed in detail in Appendix A. The complete source code is openly available at https://github.com/dualverse-ai/station. 3 Results 3.1 Experimental setup We evaluate the Station on mathematical problems drawn from the AlphaEvolve study of Georgiev et al. [15], a broad catalogue spanning analysis, combinatorics, geometry, and number theory. Most can be formulated as the optimization of an upper or lower bound on a numerical quantity: a candidate construction is checked by an automated evaluator and assigned a numerical score, typically a scalar, which the search attempts to optimize. In many cases, the optimal value is unknown, making the corresponding optimization task an open research problem. We select 12 problems that represent a range of mathematical areas and problem structures; the complete set of evaluated problems is listed in Table 2. We assign each problem to an independent Station instance. For each problem, the agents receive a task formulation that describes both the mathematical problem and the evaluator function. The task formulation may also specify additional mathematical goals that are not directly scorable. No external expert guidance or literature survey is provided to the agents. Most instances run for approximately 1,000â2,000 ticks, corresponding to roughly one to two weeks of continuous wall-clock operation. Unless otherwise specified, all instances contain six research agents, two each powered by GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro. 3 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 3.2 Summary of findings The results are summarized in Table 2. Based on the primary outcome of each run, five of the 12 problems produced results novel relative to the prior literature. Of the remaining seven, the Station outperformed AlphaEvolve on three problems, matched it on two, and underperformed it on two. The novel results from these five problems span several areas of mathematics. In finite geometry, the Station derived a new infinite family of Kakeya sets in F 3 p for primes p⥠3 (mod 4), and found a 53-point Kakeya set in F 5 3 , improving the previous bound of 63. In discrete geometry, it produced three exact 604-point kissing configurations in dimension 11, two of which appear to define previously unknown isometry classes, and established the new bound C T (128) ⤠0.107067 for the discretized Kakeya needle problem. In analysis, it improved the sign uncertainty upper bound to 0.3089 and closed approximately 82% of the previously open gap for ErdĹsâs minimum-overlap constant. Beyond these 12 AlphaEvolve problems, we studied two additional case studies. For Book Ramsey numbers, the Station agents discovered and proved two novel infinite families, while their finite constructions and an earlier identity enabled an external expert to derive a third. Together, these three families prove the conjecture at 43 values of n⤠200, resolving 28 cases that were previously open. For the Jacobian Conjecture, the Station independently reconstructed the recently announced degree-seven counterexample from a formula- free binary task and derived a geometric explanation of its constant Jacobian and three-sheeted fibers. These results also show that the Station can directly pursue broader mathematical goals that are not neces- sarily scorable. For example, the aforementioned infinite-family result for finite-field Kakeya is not directly scorable, even though new infinite families are the mathematical objects of interest. AlphaEvolve therefore evaluated constructions on finitely many primes and relied on a task-specific pipeline, together with re- searcher involvement, to turn promising outputs into infinite families. In the Station, by contrast, we stated directly in the task formulation that the finite constructions were test cases and that the primary goal was to discover infinite families. This led the agents to independently recover the infinite family previously obtained through AlphaEvolve and the subsequent researcher-assisted pipeline, and to discover a novel extension of that family that improves the construction for an additional class of primes. Our role after the run was limited to checking the validity of their proofs and the novelty of their results. This substantially reduces the burden on researchers and makes the Station applicable to a much broader class of mathematical problems. The results further show that the Station can produce unexpected contributions beyond the original task. In ErdĹsâs minimum-overlap problem, for instance, the agents were instructed to improve upper bounds, yet they also developed a lower-bound proof that closed approximately 82% of the open interval. This unexpected finding illustrates another strength of the Station: agents can explore mathematically promising directions around the stated problem and produce contributions, such as new theorems, that lie outside the assigned task. Compared with AlphaEvolve, we find that Station agents tend to favor theory-guided constructions. Individ- ual evaluations in these experiments are typically capped at 15â30 minutes, creating a strong incentive to use mathematical structure to reduce the search space. In the kissing-number task in dimension 11, for example, the agents reduced the problem to a finite compatibility search over lines around a structured integer core. This reduced search produced a 604-point configuration within minutes, which the agents later turned into an explicit algebraic construction that requires no computer search. This is markedly different from AlphaE- volveâs 593-point configuration, whose large, unequal-norm integer coordinates do not reveal a comparably compact algebraic description or readily identifiable organizing structure [15]. This bias is not universally advantageous. Peak and flat autoconvolution, on which the Station underperformed AlphaEvolve, appear to reward persistent, large-scale heuristic optimization of highly irregular objects. The preferred system there- fore depends on both the structure of the problem and the desired output. Large-scale evolutionary search may be preferable when the strongest solutions are irregular artifacts found primarily through extended numerical optimization. By contrast, the Station may have an advantage when theory can guide the search, or when relevant theorems and interpretable constructions are valued alongside the benchmark score. The next section presents detailed results for each problem. All supporting proofs, verification artifacts, and raw agent dialogue are available at https://github.com/dualverse-ai/station_data_v2. 4 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Table 2: Important findings by the Station. All evaluated problems are included. ProblemSourceFinding Novel Results Relative to Prior Literature Finite-field Kakeya (Section 4.1) AlphaEvolve Problem 6.1 For every prime p⥠3 (mod 4), the Station constructed a Kakeya set in F 3 p of size (2p 3 + 7p 2 + 3)/8, saving (pâ 3)/4 points over AlphaEvolveâs infinite family. It also found a 53-point set in F 5 3 , improving AlphaEvolve and the previous literature bound of 63; both appear novel relative to the literature. ErdĹs minimum over- lap (Section 4.2) AlphaEvolve Problem 6.5 AlphaEvolve lowered the upper bound only slightly, from 0.380927 to 0.380924, whereas the Station raised the lower bound from 0.37912 to 0.380552. Relative to the published lower bound 0.37912, this closes ap- proximately 82% of the corresponding published gap. Kissing number in d = 11 (Section 4.3) AlphaEvolve Problem 6.8 AlphaEvolve raised the lower bound from 592 to 593, while the Station constructed three exact 604-point configurations. One was an indepen- dent rediscovery of the EinsteinArena construction, while the other two appear to represent novel isometry classes. Discretized Kakeya needle (Section 4.4) AlphaEvolve Problem 6.9 At n = 128, the Station obtained union area 0.107067, improving AlphaE- volveâs 0.114810 by 6.74% and HorizonMathâs 0.109148 by 1.91%. This establishes a new literature upper bound. Sign uncertainty prin- ciple (Section 4.5) AlphaEvolve Problem 6.11 The Station lowered the upper bound to 0.3089, improving AlphaEvolveâs 0.321591 and the previously announced human value 0.3102. This is a new literature record. Better than AlphaEvolve HardyâLittlewood maximal inequality (Section 4.6) AlphaEvolve Problem 6.18 The Station reached 1.557069, versus AlphaEvolveâs 1.5080 unguided and approximately 1.533 with hints, but the centered problem was already solved. Its proof that the non-tangential constant equals 2 for 1/3⤠ι < 1 appears novel relative to the literature. Ovals problem (Sec- tion 4.7) AlphaEvolve Problem 6.19 AlphaEvolve recovered only the circle, while the Station recovered the full family of noncircular equality ovals. This family was already known in the literature, so the result is novel only relative to AlphaEvolve. Prime number theo- rem (Section 4.8) AlphaEvolve Problem 6.27 The Station certified 0.980681 for all x, improving AlphaEvolveâs sam- pled score of 0.938. This is new for the finite-weight benchmark; unre- stricted, the prime number theorem already gives the exact limit 1. Ties with AlphaEvolve Difference bases (Sec- tion 4.9) AlphaEvolve Problem 6.7 The Station independently recovered AlphaEvolveâs 360-element construc- tion but did not improve upon it. Sidorenkoâs conjec- ture (Section 4.10) AlphaEvolve Problem 6.26 Neither AlphaEvolve nor the Station found a counterexample. No sub- stantive result was obtained. Worse than AlphaEvolve Peak autoconvolution (Section 4.11) AlphaEvolve Problem 6.2 The Station obtained C 6.2 ⤠1.504473, weaker than AlphaEvolveâs C 6.2 ⤠1.5032. No substantive result was obtained. Flat autoconvolution (Section 4.12) AlphaEvolve Problem 6.3 The Station obtained C 6.3 > 0.953189, weaker than AlphaEvolveâs C 6.3 ⼠0.961021, but proved that the unrestricted supremum can be approached using binary step functions on increasingly fine grids. Additional Case Studies Book Ramsey num- bers (Section 4.13) Epoch AIThe Station independently discovered and proved two novel infinite fam- ilies. Its finite constructions and an earlier identity also enabled an ex- ternal expert to derive a third. Together, the three families prove the conjecture at 43 values of n⤠200, resolving 28 previously open cases. Jacobian Conjecture (Section 4.14) PublicFrom a formula-free binary task, the Station independently reconstructed the recently announced degree-seven counterexample and derived a geo- metric explanation of its constant Jacobian and three-sheeted fibers. 5 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 4 Detailed Results This section presents the most important findings for each problem. Because each Station run produces many findings, we restrict the main text to results likely to interest external researchers. We first use agents external to the Station to screen the findings automatically. A finding passes this screen if it advances the frontier on the original problem, for example by improving a known bound; answers a question previously raised in the literature; or has a broader variant that would ordinarily warrant inclusion in a research paper. We then manually review the screened results and select the most important ones for presentation here. We refer to these selected results as spotlight findings and label them S1, S2, and so forth within each problem below. Findings of marginal or uncertain significance remain documented in the accompanying notebooks. Readers who are more interested in the discovery process than in the mathematical details may skip to Section 5. 4.1 Finite-field Kakeya A Kakeya set in F d p is a set that contains a full line in every direction, and the problem is to make one as small as possible. Dvirâs proof of the finite field Kakeya conjecture [16] established a lower bound of order p d . Subsequent work of Bukh and Chao [17] settled the leading asymptotic constant, showing that it is 2 â(dâ1) in every fixed dimension and hence 1/4 in dimension 3. What remains open is the lower-order correction to this leading term. Exact constructions that improve the p dâ1 and smaller terms therefore sharpen the best known bounds even though the leading constant is already settled. AlphaEvolve took this problem up as Problem 6.1 of its collection, asking for small Kakeya sets. A construc- tion is scored there by the average of|K p |/B p,d over a fixed list of primes, where B p,d = (pâ1) p+1 2 dâ1 +p dâ1 is the size of the classical construction as recorded by Bukh and Chao [17]. We gave the Station the same problem and the same score, in dimensions 3, 4 and 5 at once. It proved a new infinite family of Kakeya sets in d = 3, found a Kakeya set of 53 points in F 5 3 , and established a structural limit for the entire one-pole family behind the new construction. S1. A new infinite family in d = 3 for p ⥠3 (mod 4). The Station proved that for every prime p ⥠3 (mod 4) there is a Kakeya set in F 3 p of size (2p 3 + 7p 2 + 3)/8. Writing S for the squares of F p including 0, the set is K p =(x,y,z) : x 2 + 4y â S, x 2 + 4z â S âŞ(0,t,ct + z(c)) : tâ F p , c̸= 1 âŞ(0,t,t)âŞ(0, 0,z), z(c) = c câ 1 . The first part is the classical quadratic residue set, and it already covers the p 2 directions (1,a,b); the lines added in the plane x = 0 cover the remaining p + 1. Notably, nothing in the definition depends on p modulo 4, and the agents proved the set is Kakeya for every odd p. The size, however, does depend on p modulo 4, through whether â1 is a square, and we record both cases: |K p | = 2p 3 + 7p 2 â 1 8 (p⥠1 mod 4), |K p | = 2p 3 + 7p 2 + 3 8 (p⥠3 mod 4).(1) The classical construction in this dimension has (2p 3 +10p 2 â2pâ2)/8 points, so the saving is (3p 2 â2pâ1)/8 points when p⥠1 and (3p 2 â 2pâ 5)/8 when p⥠3. In particular this is an exact size where the literature leaves an O(p) error term [17]. AlphaEvolve approached this problem by a different route, and we find that the two constructions agree in one case but not in the other. For p ⥠1 (mod 4) the constructions have the same size, and in fact are the same set. A linear change of coordinates carries one onto the other, so the first case of (1) is an independent rediscovery of the bound 1 4 p 3 + 7 8 p 2 â 1 8 obtained there. For p ⥠3 (mod 4) they differ. The smallest size AlphaEvolveâs infinite family gives on this class is (2p 3 + 7p 2 + 2pâ 3)/8, and ours is (2p 3 + 7p 2 + 3)/8, a saving of (pâ 3)/4 points. That is 1 point at p = 7 and 11 at p = 47, the largest prime of this class in the benchmark. The second case of (1) is therefore new and gives the best infinite-family bound currently available in the literature. 6 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 3571113192329313741434753 0.75 0.80 0.85 0.90 0.95 1.00 | K p | / B p , d (lower is better) d= 3 35711131719 prime p 0.65 0.70 0.75 0.80 0.85 d= 4 35711 0.475 0.500 0.525 0.550 0.575 0.600 0.625 d= 5 Pre-AlphaEvolve literatureAlphaEvolveStation Figure 1: Kakeya set sizes at the 25 pairs (d,p) of the benchmark, divided by the size B p,d of the classical construction; lower is better. Pre-AlphaEvolve literature is the smallest size obtained from the explicitly defined families predating AlphaEvolve that are listed in Appendix B. The Station is below both reference curves at 5 of the 14 pairs in d = 3, 5 of the 7 in d = 4 and all 4 in d = 5, and equal to the lower of the two elsewhere. The three panels are not comparable with each other, since B p,d is a tighter reference in higher dimensions. S2. Finite improvements and a 53-point Kakeya set in F 5 3 . The Station wins 14 of the 25 finite benchmark comparisons and ties the remaining 11 (Figure 1). Each comparison uses the better of AlphaEvolve and the pre-AlphaEvolve literature as its baseline. The case (d,p) = (5, 3) is especially notable. Let k n denote the minimum size of a Kakeya set in F n 3 . The Station constructed a 53-point set in F 5 3 , improving the previous bound from k 5 ⤠63 to k 5 ⤠53 [18]. In light of the known values k 1 = 3, k 2 = 7, and k 3 = 13, together with the bound k 4 ⤠27, which is believed to be sharp, it was guessed in 2009 that the recurrence k n = k nâ1 + 2k nâ2 continues, predicting k 5 = 53 [18]. The size of the Stationâs construction therefore coincides with the guessed value, although whether (k 5 = 53) holds and whether the recurrence continues remains open. S3. Structural analysis of the new infinite family. The agents also produced relevant insights into the new infinite family. They analyzed the more general completion z(c) = Ac + B câ p 1 , which includes the construction in S1. Eliminating the slope c reduces incidence with these lines to whether (zâ Aâ p 1 y) 2 â 4(Ap 1 + B)y is a square. A quadratic-character calculation then shows that the lines cover exactly p(pâ 1)/2 points away from the axis, independently of the three parameters. Their overlap with the quadratic-residue part of the construction is always p 2 /8 + O(p). Consequently, every nondegenerate completion in this MĂśbius family adds 3p 2 /8 + O(p) points: changing the numerator or the location of the pole affects only the lower-order terms. For the particular choice z(c) = c/(câ 1) used in S1, the agents evaluated the lower-order term exactly, yielding the infinite family stated in (1). The result also explains AlphaEvolveâs infinite family for p ⥠1 (mod 4). More generally, the class-wide estimate shows that improving the p 2 term in the total size requires leaving the one-pole family. Limitations. The new infinite family is confined to d = 3. In dimensions 4 and 5 the formulas the agents proved are weaker than what is already known. On the shared class p⥠1 (mod 4) the first two coefficients agree with AlphaEvolve in each dimension and the third is worse in both. 7 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 0.37900.37940.37980.38020.38060.3810 value of Îź Pre-AlphaEvolve literature 2016â2023 AlphaEvolve 2025 Pre-Station literature 2026 Station 2026 White 0.379005 Haugland 0.380927 White 0.379005 AlphaEvolve 0.380924 KimâPilanci 0.379120 Ye et al. 0.380868 Station 0.380552 Ye et al. 0.380868 lower boundupper bound Figure 2: Successive published bounds for ErdĹsâs minimum-overlap constant. Each horizontal segment joins the best lower and upper bounds at the indicated stage. The Station raises the lower bound from 0.37912 to above 0.380552, closing approximately 82% of the previously open interval. dStationAlphaEvolve 4 1 8 p 4 + 19 32 p 3 + 25 32 p 2 + O(p) 1 8 p 4 + 19 32 p 3 + 11 16 p 2 + O(p 3/2 ) 5 1 16 p 5 + 47 128 p 4 + 25 32 p 3 + O(p 2 ) 1 16 p 5 + 47 128 p 4 + 177 256 p 3 + O(p 5/2 ) The sizes we report at individual primes in d = 4, 5 do still improve on the benchmark, but they come from search rather than from a formula. 4.2 Erd Ě os minimum overlap ErdĹsâs minimum-overlap problem asks how evenly two complementary parts of an interval can avoid one another under translation. Let f : [â1, 1] â [0, 1] be measurable with integral 1, put g = 1â f on [â1, 1], and extend both functions by zero outside the interval. Write C f (x) = Z 1 â1 f (t)g(t + x)dt, Îź = inf f âĽC f ⼠â . This constant is the continuum form of ErdĹsâs minimum-overlap problem for balanced partitions of long integer intervals [19â21]. AlphaEvolve took up this problem as Problem 6.5 of its mathematical collection and improved Hauglandâs upper bound from 0.380927 to 0.380924, while later work further reduced it to 0.380868 [22]. On the lower-bound side, Kim and Pilanci established 0.37912 [23]. Thus, immediately before this work, the best published bounds were 0.37912⤠Ο⤠0.380868. S1. A new lower bound of 0.380552. The Station agents proved Îź > 0.380552.(2) Relative to the previously published lower bound of 0.37912, this reduces the corresponding published open interval by approximately 82%, as shown in Figure 2. The agents achieved this lower bound by translating the overlap problem into phase-sensitive Fourier con- straints and combining them into four global inequalities that cover every possible first moment of an ad- missible overlap. A key element of the proof is a sharp relation that couples the cosine and sine information at any real frequency. Writing P (Ξ) and Q(Ξ) for the cosine and sine transforms of C f , and s(Ξ) = sin(Ξ)/Ξ, the agents proved P (Ξ)⤠s(Ξ) 2 â Q(Ξ) 2 4s(Ξ) 2 s(Ξ)̸= 0 . 8 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment White had already used Fourier phase information and convex optimization, while Kim and Pilanci later introduced additional moment constraints [21, 23]. Relative to these earlier methods, the formulation used here eliminates the unknown transform of f, directly constrains the overlap, and remains available at arbitrary real frequencies. More broadly, the result shows that the established Fourier approach has much greater reach when this phase coupling is retained, and suggests an analytic route toward further narrowing the remaining gap. Comparison with AlphaEvolve on the upper bound. The Station agents independently obtained Îź < 0.380895, a slight improvement on AlphaEvolveâs published upper bound of 0.380924. However, this re- mains above the current published upper bound Îź < 0.380868 of Ye et al. [22]. The Station therefore did not establish a new upper-bound record. 4.3 Kissing number in d = 11 The kissing number K(d) is the largest number of nonoverlapping unit spheres that can simultaneously touch a central unit sphere in R d . Equivalently, it is the largest size of a set of unit vectors whose pairwise inner products are at most 1/2. AlphaEvolve took up this classical question as Problem 6.8 of its mathematical collection and improved the lower bound in dimension eleven from 592, established by Ganzhinov using highly symmetric lines [24], to 593. We ran two independent Stations on the same problem using AlphaEvolveâs scoring rule, which measures the total pairwise overlap among the surrounding spheres. Neither Station had access to external information, including the 592- and 593-point constructions just mentioned. Both reached 604 points, proving K(11)⼠604. Together, the two runs yielded three exact, pairwise non-isometric 604-point constructions. S1. Three exact 604-point kissing configurations. The Station discovered three geometrically distinct 604- point kissing configurations in R 11 . All three are exact equal-norm arrangements over Q( â 2), but they organize their points differently: two are centrally symmetric, one is not, and each has a different contact structure and set of pairwise angles. Figure 3 visualizes their shared architecture and the two structural choices that distinguish them. We label them Constructions 1, 2, and 3: Construction123 Touching pairs19,704 22,904 22,840 Centrally symmetricYes YesNo Antipodal pairs302 302 238 Distinct pairwise angles221415 The different numbers of touching pairs prove that the configurations are pairwise non-isometric, since this number is preserved by orthogonal transformations and relabeling. Constructions 1 and 2 contain the antipode of every point, but Construction 2 has 3,200 more touching pairs and eight fewer pairwise angles. Construction 3 has 128 points without antipodes. Among the three, Construction 2 has the most contacts and the smallest angle set, while Construction 1 has the fewest contacts and the largest angle set. Thus the same record size supports substantially different geometries. In concurrent work, Bianchi et al. reported Construction 1 from the EinsteinArena platform shortly before our public release of Construction 3 [25]. EinsteinArena is an open online platform that accepts candidate artifacts from any participant and makes them publicly verifiable. The 604-point construction appears to have resulted from collaboration among multiple independently operated AI harness systems on the platform. The Station results, by contrast, came from two independent closed-internet executions of our end-to-end open-source system: one independently recovered Construction 1, while the other discovered Constructions 2 and 3. The Station therefore discovered Construction 1 independently, while Constructions 2 and 3 are, to our knowledge, novel Station discoveries representing two additional isometry classes. S2. An algebraic construction for a 604-point kissing configuration in R 11 . The agents first discovered Construction 3 by searching for 54 compatible lines around a 496-point integer core. They later showed that 9 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 432-point shared core64-point core type I64-point core type I108-point extension type I108-point extension type I (a) Construction 1 Core type I + extension type I (b) Construction 2 Core type I + extension type I (c) Construction 3 Core type I + extension type I Figure 3: The three 604-point kissing configurations in R 11 , shown under the same orthogonal projection into R 3 . All three share the same 432-point rational core, shown in light gray, and each has the form 432+64+108. Constructions 1 and 2 use the same 64-point core type, so their complete 496-point cores agree, but they use different 108-point extensions. Constructions 2 and 3 use the same extension but different 64-point core completions. The colored spheres distinguish the two core types and the two extension types. the same configuration is governed by a compact algebraic rule rather than an arbitrary list of coordinates, yielding an explicit algebraic construction. The construction itself requires no computer search. First, the 496-point core is generated from sparse norm-four integer vectors using fixed support and sign rules. Second, in a coordinate frame rotated by 45 ⌠in one coordinate plane, eleven simple sign patterns generate all 54 lines; taking both directions on each line gives the 108-point extension. The appearance of â 2 is intrinsic: it is forced by the compatibility between the extension and the core. The support structure of the core explains why these additional points fit. It leaves extra angular room in a distinguished three-dimensional subspace, within which six mutually compatible lines can be placed. Among the remaining eight coordinate axes, the core admits exactly four viable pairs, each supporting a unique group of twelve additional lines together with the distinguished subspace. These four pairs are disjoint, so their groups are mutually compatible. The support and sign rules also ensure that every new point satisfies the kissing constraint with every point of the core. The resulting configuration therefore contains 496 + 2(6 + 4¡ 12) = 604 points. S3. Why the classical D 11 construction stops at 582. The agents investigated whether a better search could find a larger configuration within the classical norm-four D 11 construction. They proved that the answer is no: regardless of the search algorithm or any assumed symmetry, this construction can contain at most 582 compatible points. Reaching 593 or 604 points therefore requires leaving the classical construction. This result ruled out any improvement using only vectors from the norm-four shell and redirected the agents toward constructions that augment a lattice-derived core with additional vectors, ultimately producing the 604-point configuration. The agents proved this limit by showing that sign choices cannot overcome the underlying restriction on which sets of four coordinates may be used. Let A(n, 4, 4) denote the largest compatible collection of four- coordinate supports, and let Îą(J Âą (n, 4)) denote the largest compatible collection after signs are assigned to those coordinates. The agents proved Îą(J Âą (n, 4)) = 16A(n, 4, 4).(3) In other words, allowing arbitrary signs increases the optimum by exactly the 16 possible sign patterns on four coordinates; it cannot produce any additional advantage. 10 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Best proved in 1977 that A(11, 4, 4) = 35 [26]. The agentsâ identity therefore limits the signed weight-four part of the construction to 560 points. The remaining 22 coordinate vectors Âą2e i are compatible with these points, giving an exact limit of 582 for the complete norm-four D 11 construction. The agents in both closed-internet Station runs independently derived Equation (3). We later found that it overlaps with the k = 4 case of Theorem 1 in a paper by Takhanov and Yun, made publicly available only recently, on June 2, 2026 [27], where the identity serves as the foundation for a broader classification of signed kissing configurations. The agents therefore discovered the identity independently. Limitations. The Stationâs success in dimension eleven did not extend to new records in nearby dimensions. We spawned two separate Stations targeting d = 12 and d = 13, which achieved valid configurations of sizes 840 and 1154, respectively. The dimension-twelve result falls one point below the current 841-point frontier [28, 29], while the dimension-thirteen result matches the 1154-point construction of Zinoviev and Ericson [29, 30]. Discussion. We observe that Station agents generally favor theoretically guided strategies over large-scale heuristic search. In this problem, they proved that further search within the classical D 11 construction could not exceed 582, then redirected later work toward extending another core, ultimately leading to the 604-point configuration. By contrast, AlphaEvolveâs 593-point construction consists of large unequal-norm integer coordinates that do not appear to reveal a comparably compact algebraic description or readily identifiable organizing structure. This theory-guided bias is not necessarily always an advantage: in dimension twelve, the Station stopped at 840, while the current 841-point frontier was reached through large-scale numerical optimization guided by structural insight [28, 29]. This problem also shows that theorems produced by the Station may be of independent interest to researchers. For instance, Equation (3), derived independently by the agents, overlaps with a theorem in a paper made publicly available only recently [27]. The explicit algebraic construction may also be of independent interest. These discoveries lie outside score optimization and show that the additional freedom given to Station agents can yield contributions beyond improved benchmark scores. 4.4 Discretized Kakeya needle The classical Kakeya needle problem asks how little area is needed to turn a unit line segment through every direction. A finite version replaces the continuum of directions by n equally spaced ones and represents them by n thin triangles that may slide horizontally [31]. More precisely, for real offsets x 1 ,...,x n , let T j (x j ) = conv (x j , 0), x j + 1 n , 0 , x j + j n , 1 ,1⤠j ⤠n, and define C T (n) =inf x 1 ,...,x n n [ j=1 T j (x j ) . CĂłrdobaâs lower bound and a Schoenberg construction analyzed by Keich show that C T (n) has order 1/ logn [32, 33], but its sharp finite values have remained largely unknown. AlphaEvolve took up this problem as Problem 6.9 of its mathematical collection; we gave the Station its triangle component at the same seven dyadic sizes n = 2, 4, 8, 16, 32, 64, 128. S1. New upper bounds at n = 32, 64, 128. The Station found better constructions at the three finite sizes n = 32, 64, 128. At n = 128, it found a triangle union of area 0.107067, improving AlphaEvolveâs 0.114810 by 6.74% and the later HorizonMath value 0.109148 by 1.91% [34], and therefore proving C T (128)⤠0.107067. The gains are more modest at n = 32 and n = 64, where the Station reduced AlphaEvolveâs areas by 2.15% and 0.69%, respectively; at the smaller tested sizes n = 2, 4, 8, 16, it reached the same values as AlphaEvolve (Figure 4). 11 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 248163264128 number of triangles n 0.10 0.15 0.20 0.25 0.30 0.35 union area (lower is better) Station AlphaEvolve 0 1 height y Best symmetric A= 7/30 Asymmetric construction A= 14/61 Figure 4: Left: union areas of the finite constructions published by AlphaEvolve and produced by the Station; lower is better. The Station matches AlphaEvolve at n = 2, 4, 8, 16 and reduces the area by 2.15%, 0.69%, and 6.74% at n = 32, 64, 128, respectively. Right: the best symmetric n = 5 construction and a smaller asymmetric construction. Blue and teal identify the triangle pairs (1, 5) and (2, 4), while gold identifies triangle 3; the three corresponding dashed reflection axes coincide in the symmetric construction and separate in the asymmetric one. S2. Exact optima at n = 3, 4 and symmetry breaking at n = 5. Before this work, only the classical value C T (2) = 1/3 was known exactly [31]. An elementary symmetric construction gives C T (3)⤠5 18 , while Schoenbergâs classical Perron construction [35] gives C T (4)⤠1 4 . AlphaEvolve later reproduced the n = 4 value numerically. The Station proved the matching lower bounds and therefore established C T (3) = 5 18 , C T (4) = 1 4 . It also showed that both minima admit reflection-symmetric configurations and that the n = 4 optimum contains the continuous family 1 4 , 1 4 â c,c, 0 , 1 20 ⤠c⤠1 8 . The Station then proved that the minimum among reflection-symmetric configurations at n = 5 is 7/30 and discovered a new asymmetric construction of area 14/61 < 7/30. Figure 4 (right) compares the symmetric minimizer with this smaller asymmetric construction. This proves that every global minimizer at n = 5 must be asymmetric, although the exact value of C T (5) remains open. These results lie outside the benchmark score. Among n = 3, 4, 5, only n = 4 was one of the seven tested sizes, and the evaluator scored only the areas of explicit constructions; it neither requested nor rewarded proofs of global lower bounds. The task specification also did not ask the agents to classify exact small-n optima or investigate symmetry breaking. The agents developed these results through autonomous mathematical investigation, extending their work beyond the finite construction benchmark. Limitations. The Station optimized its constructions separately at the tested powers n = 2 k , and Figure 4 compares them with AlphaEvolveâs corresponding separately optimized finite constructions. The figure therefore compares finite constructions on both sides. Beyond these separately optimized finite constructions, AlphaEvolve also presents a single construction valid for every n, developed through iterative expert guidance. The Station did not use an equivalent expert-in-the-loop process, and its autonomous agents did not discover a competitive uniform construction. 12 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 0.3050.3100.3150.3200.3250.330 upper bound on C (lower is better) Pre-AlphaEvolve Literature AlphaEvolve Optimum in double-root Laguerre family Announced human results Station 0.32831 0.321591 0.315305â0.315309 0.3102 0.3089 â0.04 0.00 0.04 A(f)A( Ě f) = 0.321591 AlphaEvolve â0.02 0.00 0.02 A(f)A( Ě f) = 0.315309 Station double-root construction 0.550.650.750.850.95 radius x â0.01 0.00 0.01 A(f)A( Ě f) ⤠0.3089 Station curve: âP(2Ďx 2 ) prescribed double roots Figure 5: Bounds and constructions for the one-dimensional sign-uncertainty problem. Left: successive upper bounds on C SU ; lower is better. Right: the polynomial factors âP (2Ďx 2 ) for the AlphaEvolve construction, the Stationâs double-root construction, and the Stationâs 0.3089 construction. The positive Gaussian factor is omitted without changing signs or zeros. Open circles mark the prescribed double roots. 4.5 Sign uncertainty principle The one-dimensional sign-uncertainty problem asks how soon a function and its Fourier transform can both become eventually nonnegative when both start negative at the origin. For a nonzero even integrable function f : Râ R with integrable Fourier transform, define A(f ) = infr > 0 : f (x)⼠0 whenever |x|⼠r. The problem asks for the largest constant C SU such that A(f )A( b f ) ⼠C SU . Bourgain, Clozel and Kahane introduced the problem [36], and subsequent work obtained progressively stronger bounds [37, 38]. AlphaE- volve studied it as Problem 6.11 and reported an upper bound of 0.321591 together with an unpublished human bound of 0.3102. The Station further improved this bound to 0.3089, as summarized in Figure 5. S1. A new upper bound of 0.3089. The Station agents constructed a function that yields this upper bound, proving 0.2025⤠C SU ⤠0.3089. They take f Îľ (x) = âP (2Ďx 2 )â Îľ e âĎx 2 , Îľ = 10 â6 , where P is expressed in the even-index generalized Laguerre polynomials L (â1/2) 2j ; the proved tail margin exceeds Îľ, so f Îľ (0) < 0 while eventual nonnegativity is preserved. These basis functions are fixed by the Fourier transform, so the choice gives f Îľ = b f Îľ automatically and reduces the problem to constructing one polynomial with the required sign. Numerical search found the degree-226 polynomial shown in Figure 5; the agents expressed its coefficients as exact rational numbers and proved that the resulting function is nonnegative beyond the corresponding radius, fulfilling the problemâs eventual-nonnegativity requirement. S2. The double-root Laguerre family is exhausted near 0.3153. In this task, we gave the agents the same prescribed-double-root Laguerre setup and scoring rule used by AlphaEvolve, but no access to AlphaEvolveâs paper or results. Under this setup, every submission is restricted to the family in which P is determined by at most twenty prescribed positive double roots in the even-index Laguerre basis; we call this the double-root Laguerre family. AlphaEvolveâs 0.321591 construction also belongs to this family. Let C DR,20 = inf A(f )A( b f ) : f belongs to the double-root Laguerre family . 13 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment The Station agents proved 0.315305 < C DR,20 ⤠0.315309.... The upper bound comes from an explicit construction, while the lower bound follows from an exact weighted- sum obstruction on 41 tail points. Thus any construction improving the upper bound below 0.315305 must leave the double-root Laguerre family. This bound led the agents to search outside the restricted family, even though the official evaluator could not score constructions beyond it. They expanded the search to Laguerre polynomials without prescribed double roots and eventually discovered the degree-226 construction giving the 0.3089 bound. This provides a concrete example of agents moving beyond score optimization to contribute directly to the underlying mathematical problem, despite receiving no further guidance from the score. 4.6 HardyâLittlewood maximal inequality The one-dimensional centered HardyâLittlewood problem asks for the optimal constant controlling where centered local averages can be large. For a non-negative integrable function f : Râ R, define Mf (x) = sup h>0 1 2h Z x+h xâh f (y)dy, and let C 0 be the least constant such that |Mf > Îť|⤠C 0 Îť âĽf⼠1 . Melas solved the problem, proving C 0 = 11 + â 61 12 = 1.567521... and constructing finite point-mass examples approaching this value [39, 40]. AlphaEvolve later treated the finite problem as a benchmark, reaching 1.5080 in search mode and about 1.533 with hints from the literature. The Station agents found a 356-point-mass construction with value 1.557069, improving AlphaEvolveâs result but failing to recover the global optimum already discovered by Melas. S1. Sharp constants between the centered and uncentered operators. Ramos considered the natural non-tangential family interpolating between the centered and uncentered HardyâLittlewood maximal op- erators [41]. Its parameter Îą runs from the centered operator at Îą = 0 to the uncentered operator at Îą = 1. Writing C Îą for the sharp weak-(1, 1) constant, Ramos stated that its exact value was unknown for every 0 < Îą < 1, while the endpoint C 1 = 2 is classical [40, 42]. While working on the task, the Station agents solved this question for 1/3⤠ι < 1, proving C Îą = 2 for every 1 3 ⤠ι⤠1.(4) The constants for 0 < Îą < 1/3 remain open. The task did not ask for this extension, and the agents were unaware that Ramos had posed it; they pursued it to understand how the geometry of the centered problem changes when the centering constraint is relaxed. 4.7 Ovals problem The Ovals problem asks whether the curvature of every closed convex plane curve forces the lowest eigen- value of an associated one-dimensional SchrĂśdinger operator to be at least 1. For a curve Îł of length 2Ď, parametrized by arclength s, define H Îł =â d 2 ds 2 + Îş(s) 2 , C = inf Îł Îť 0 (H Îł ), 14 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment where Îş is the curvature and Îť 0 is the lowest eigenvalue under periodic boundary conditions. Benguria and Loss conjectured that C = 1 and exhibited a continuous equality family containing the circle and noncircular ovals [43â45], proving C ⤠1, while Linde proved the global lower bound C > 0.81; numerical evaluation of the explicit constant in his theorem gives C > 0.8246 [46]. AlphaEvolve took up this question as Problem 6.19 of its mathematical collection. S1. Independent recovery of the BenguriaâLoss equality family. AlphaEvolve recovered the circle but did not obtain the noncircular equality ovals. The Station independently recovered a one-parameter normal form, modulo Euclidean motions and shifts of the arclength origin, for the classical BenguriaâLoss equality family. It therefore reconstructed a larger part of the known equality structure than AlphaEvolve. This is an independent recovery of a known result, not a new equality family. Benguria and Loss formulated the conjecture and exhibited the equality family; Burchard and Thomas proved its local minimality, while Bernstein and Mettler developed its projective geometry and established the name âovals of Benguria and Lossâ [43â45]. Neither AlphaEvolve nor the Station improved the global lower bound. 4.8 Prime number theorem The prime number theorem describes the asymptotic density of the primes. If Ď(x) counts the primes at most x, it states that lim xââ Ď(x) x/ logx = 1. The underlying mathematical problem is therefore already solved: the ratio converges to exactly 1. AlphaE- volve nevertheless took up a finite version as Problem 6.27 of its collection. It searched for a finitely supported weight f satisfying X k f (k) k = 0. The score of such a weight and its associated sum are A(f ) =â X k f (k) logk k , F f (x) = X k f (k) j x k k . The classical Chebyshev argument shows that F f (x)⤠1 for every x⼠1(5) implies the rigorous lower bound lim inf xââ Ď(x) x/ logx ⼠A(f ) [47]. The required global inequality in Equation (5) is much more restrictive than the prime number theorem itself: a single finite weight must satisfy the inequality for every x. AlphaEvolveâs score tested this inequality only at finitely many sampled values. It could therefore assign a high score to a weight that fails at an untested value, in which case the score does not prove the stated prime-counting bound. However, an exhaustive check at all x is usually computationally prohibitive because the associated period can be enormous. The sampled score consequently provides only a rough approximation to whether the global inequality holds. S1. A score of 0.980681 valid for every x. The Station agents discovered a finite construction f satisfying Equation (5) for every x, with A(f )⼠0.980681.(6) This improves on AlphaEvolveâs reported score of 0.938. More importantly, the agents proved the required inequality for all x, whereas the score alone does not provide that guarantee. Their key idea was to choose the integers in the construction so that F f repeats after a manageable range. This reduces the infinitely 15 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment many possible values of x to one finite exhaustive check, which the agents completed using exact arithmetic in under a minute. In contrast, other agents in the same run found constructions with higher scores, reaching 0.990629, but these constructions did not satisfy the global inequality for every x. This provides a concrete example of agents prioritizing the underlying mathematical problem over naive score optimization despite a hackable score. S2. Why a direct MĂśbius cutoff fails. The MĂśbius function is a natural starting point because it is central to a standard formulation of the prime number theorem. AlphaEvolve explored finite constructions obtained by truncating the MĂśbius function, and the Station agents initially pursued the same approach. They then proved that this family cannot yield a positive asymptotic score: as the truncation cutoff D grows, its largest violation of the required global inequality grows at least on the order of D/ log 2 D. Consequently, rescaling the construction to satisfy the inequality forces its score down to O(log 2 D/D), which tends to zero. The proof builds on results about incomplete MĂśbius sums [48]. This obstruction led the agents to abandon direct MĂśbius cutoffs and explore a more flexible construction with jointly optimized coefficients, producing the rigorous score of 0.980681 described above. Limitation. Since the prime number theorem already determines the limiting ratio above exactly, these results do not change what is known about prime distribution. Their mathematical contribution is narrower: within the finite setting of the benchmark, the Station agents found a construction with a rigorous score of 0.980681 and proved that the natural MĂśbius cutoff cannot yield a positive asymptotic score. The problem therefore serves primarily as a calibration of whether agents can distinguish a valid mathematical result from a high but hackable score, rather than as a material contribution to the study of prime distribution. 4.9 Difference bases A finite set B â Z is a difference basis for 1,...,n if every integer in that interval is a difference of two elements of B. If â(n) is the smallest possible size of such a set, the quantity to minimize is â(n) 2 /n; RĂŠdei and RĂŠnyi proved that these normalized minima converge and that their limit is their infimum [49]. AlphaEvolve reported the upper bound C := inf nâĽ1 â(n) 2 n ⤠360 2 49109 â 2.639027 as Problem 6.7 of its collection. The preceding published upper bound was Golayâs C ⤠2.6458... [50, 51], rather than the 2.6571... benchmark used in AlphaEvolveâs comparison. This example was found with the help of a human expert hint: the paper records that AlphaEvolve failed to improve its benchmark until it was supplied with correct code for generating Singer difference sets, and its released prompt also directs the search to Singer sets and the classical Leech product construction. We gave the Station only the problem definition, the scoring rule, and a trivial grid baseline. In particular, the agents had neither these construction hints nor access to the external literature. S1. Independent recovery of a record in the LeechâGolay family. Leech and Golay combined the four- point difference basis 0, 1, 4, 6 with Singer difference sets to obtain earlier members of this construction family [50, 52, 53]. The Station independently recovered its q = 89 member. Taking v = q 2 + q + 1 = 8011, a 90-element Singer difference set D â Z v , and A =0, 1, 4, 6, the agents formed B =va + d : aâ A, dâ D. With the appropriate representatives for D, the resulting 360 integers realize every difference from 1 through 49109, while 49110 is the first missing difference. Thus C ⤠360 2 49109 = 2.6390274695..., improving Golayâs preceding bound by approximately 0.0067. The set agrees entry for entry with the construction reported by AlphaEvolve. This is an independent recovery of a known record, not a new upper 16 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment bound relative to AlphaEvolve or a new construction family. The agents also tried to push the lower bound further, but reached only the classical bound C ⼠2.434467... [52], whereas Yang and Liao proved the stronger published bound C > 2.4421 [54]. 4.10 Sidorenkoâs conjecture Sidorenkoâs conjecture asserts that every bipartite graph H satisfies t(H,W ) ⼠t(K 2 ,W ) |E(H)| for every graphon W, where t(H,W ) is the homomorphism density of H in W [55]. The smallest unresolved instance is the ten-vertex, fifteen-edge graph H = K 5,5 10 , also called the bipartite MĂśbius ladder [56]. AlphaEvolve took up this problem as Problem 6.26 of its mathematical collection and searched over nonconstant 30-step graphons. It scored a candidate by t(K 2 ,W ) 15 t(H,W ) â 1, so a positive value would give a counterexample and disprove this instance of the conjecture. AlphaEvolve reported that it did not find a counterexample. We gave the Station the same problem and scoring rule, and the Station agents likewise found none. As such, the status of the conjecture is unchanged. 4.11 Peak autoconvolution AlphaEvolveâs Problem 6.2, called the first autocorrelation inequality in its collection, asks how evenly the sum of two independent random variables with the same compactly supported density can be distributed. More precisely, for a nonnegative function f supported on [â1/4, 1/4] and normalized by R f = 1, let C 6.2 = inf f âĽf â f⼠â . Determining C 6.2 is connected to the asymptotic size of generalized Sidon sets, and its exact value remains unknown [57]. The best currently reported bounds are 1.2937⤠C 6.2 ⤠1.502851, with the lower and upper endpoints coming from certified convex relaxations and an explicit step function, respectively [23, 58]. AlphaEvolve achieved the upper bound C 6.2 ⤠1.5032, improving the pre-AlphaEvolve bound C 6.2 ⤠1.50972 of Matolcsi and Vinuesa [57]; T-Discover later advanced the frontier to C 6.2 ⤠1.502863 [59], and an exact- arithmetic certificate improved it further to C 6.2 ⤠1.502851 [58]. The Station reached only C 6.2 ⤠1.504473, worse than both AlphaEvolve and the current frontier. AlphaEvolveâs highly irregular construction emerged from large-scale heuristic search. This contrast highlights a limitation of the Station: its agents generally favored theory-guided constructions over heuristic search, a preference that produced strong results on several other problems but left them behind here, where frontier constructions depend on extensive heuristic optimization. 4.12 Flat autoconvolution AlphaEvolveâs Problem 6.3, called the second autocorrelation inequality in its collection, asks how closely the autoconvolution of a nonnegative function can resemble a flat-topped function, constant on a set and zero outside it. More precisely, for a nonzero nonnegative function f â L 1 (R)⊠L 2 (R), let Q(f ) = âĽf â f⼠2 2 âĽf â f⼠1 âĽf â f⼠â , C 6.3 = sup f Q(f ). HĂślderâs inequality gives C 6.3 ⤠1; for an arbitrary nonnegative output, equality occurs only for such a flat-topped function. Whether the autoconvolution constraint forces the strict inequality C 6.3 < 1 remains open [57, 60]. Before AlphaEvolve, the best known bounds were [57] 0.88922⤠C 6.3 ⤠1. 17 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment AlphaEvolve established the lower bound C 6.3 ⼠0.961021, while later work further improved this to 0.962694 [22]. The Stationâs best verified construction reached only C 6.3 > 0.953189 and therefore did not improve the numerical bound. This shortfall reflects the same limitation seen in Problem 6.2, minimiz- ing the peak of an autoconvolution (Section 4.11): the Stationâs theory-guided agents were poorly suited to finding the highly irregular constructions produced by large-scale heuristic search. S1. Binary step functions preserve the unrestricted supremum. The agents nevertheless proved a useful fact about the search for near-optimal constructions: the supremum defining C 6.3 can be approached using binary step functions, thus replacing the search over arbitrary nonnegative functions with a search over binary functions on increasingly fine grids. 4.13 Book Ramsey numbers Given graphs G 1 ,G 2 , the Ramsey number R(G 1 ,G 2 ) is the smallest n such that every red-blue edge coloring of K n forces either a red copy of G 1 or a blue copy of G 2 . Establishing the exact values of Ramsey numbers is a difficult computational and theoretical challenge. The most famous Ramsey numbers are those where G 1 and G 2 are complete graphs, but many other choices have been studied extensively (see the survey [61]). The book graph B k consists of k triangles that share a common edge. An open problem is whether R(B nâ1 ,B n ) = 4nâ 1(7) holds for every positive integer n. Rousseau and Sheehan established the upper bound in 1978, proving R(B nâ1 ,B n ) ⤠4nâ 1 for all n [62]. It therefore remains to prove the matching lower bound. For a given n, this amounts to constructing a redâblue edge coloring of K 4nâ2 containing neither a red B nâ1 nor a blue B n . The third author proved equality for n⤠20, independently matching contemporaneous work, and established an infinite Paley-type family whenever 2n â 1 is a prime power congruent to 1 (mod 4) [63, 64]. This combination of finite evidence and a general arithmetic construction led him to conjecture that (7) holds for all n [63]. Epoch AI subsequently adopted it as a FrontierMath open problem [65]. After its posting, further work extended the consecutively solved range to n⤠56 and produced two additional infinite families by extending established constructions [66]. We ran two Stations on this problem. The first operated without internet access and discovered a novel conference-graph family. We then ran a second Station with internet access and a summary of the first Stationâs results; it discovered a new doubled Legendre family together with several new finite constructions. An external expert subsequently combined the pattern in these finite constructions with an earlier result from the second Station to obtain the YamadaâPott infinite family. Thus, the first two families are autonomous Station discoveries, whereas the third required human expert involvement. All three families are novel relative to the existing literature and are visualized in Figure 6. The parameters n covered by each family, including which were previously open, are summarized in Figure 7. S1. A conference-graph family. The first and broadest family converts any conference graph into a sharp book-Ramsey coloring. Specifically, if a strongly regular graph with parameters q, qâ 1 2 , qâ 5 4 , qâ 1 4 exists, then the Stationâs agents proved R(B q ,B q+1 ) = 4q + 3.(8) Paley conference graphs exist whenever q is a prime power congruent to 1 (mod 4). Consequently, the theorem proves the conjecture whenever nâ 1 is a prime power congruent to 1 (mod 4). Beyond the Paley case, Seberry and Whiteman used Mathonâs construction to obtain symmetric conference matrices of order 5¡ 9 2t+1 + 1 for every t ⼠0 [67, 68]. These yield conference graphs of order q = 5¡ 9 2t+1 , so the Station theorem also proves the conjecture whenever n = 5¡ 9 2t+1 + 1, t⼠0. 18 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment BBBRRBRR uv B BR RB R u v 0+0â 1+1â uv 0+ 0â 1+ 1â u v F 1 F 2 F 1 F 2 (a) Conference family (n= 6) 4 chambers: B, BR, RB, R 5 vertices each; endpoints: u, v (b) Doubled Legendre family (n= 6) 4 layers: 0+, 0â, 1+, 1â 5 vertices each; endpoints: u, v (c) YamadaâPott family (n= 11) 2 affine fibres: F 1 , F 2 21 vertices each red edgeblue edgediagonal (no edge) Figure 6: Block-ordered adjacency matrices for the smallest nontrivial members of the three infinite families. Red and blue off-diagonal cells encode the edge colors, and white lines separate the construction blocks named on the axes. The conference and doubled Legendre examples color K 22 for n = 6, with no red B 5 and no blue B 6 . The YamadaâPott example colors K 42 for n = 11, with no red B 10 and no blue B 11 . The first member gives q = 45 and n = 46. The known conference graph of order q = 65 supplies the additional parameter n = 66 [69]. In total, known conference graphs prove the conjecture at 30 values of n⤠200, including 19 that were previously open [63â66]. S2. A doubled Legendre family. The second family converts a periodic Legendre source over F Q into a sharp book-Ramsey coloring [70]. Specifically, for every prime power Q > 3 with Q ⥠3 (mod 8), the Stationâs agents proved R B (Qâ1)/2 ,B (Q+1)/2 = 2Q + 1.(9) Consequently, the theorem proves the conjecture whenever 2nâ 1 is a prime power congruent to 3 (mod 8). For n⤠200, this family proves equality at 21 values and, at the time of its discovery, resolved six additional open cases after accounting for the conference family [63â66]. The agents discovered this general family in mid-July 2026. Concurrent work announced at the end of July independently produced the finite case n = 70 [65]; the Station theorem contains n = 70 as one member and covers infinitely many further parameters. The doubled Legendre family is related to, but distinct from, the Legendre family reported by Turturean [66]. Both begin with the same type of periodic Legendre source over F Q , with Q⥠3 (mod 8), but use different lifts to obtain a book-Ramsey coloring. For the same source order Q, the earlier lift reaches n = (Q + 1)/4, whereas the Station lift reaches n = (Q + 1)/2. It therefore doubles the Ramsey parameter and covers a different set of values, as Figure 7 shows. S3. A YamadaâPott family. The third family converts a classical YamadaâPott design into a sharp book- Ramsey coloring [71]. Specifically, for every prime power q ⼠7 with q ⥠3 (mod 4), we proved R B (q 2 âqâ2)/4 ,B (q 2 âq+2)/4 = q 2 â q + 1.(10) Consequently, the theorem proves the conjecture whenever n = q 2 â q + 2 4 for a prime power q ⼠7 congruent to 3 (mod 4). For n⤠200, this family proves equality at five values and resolves three additional previously open cases after accounting for the conference and doubled Legendre 19 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 120406080100120140160180200 Ramsey parameter n Finite examples Paley family Wesley (2026) Legendre family Turturean (2026) Conference family Station (2026) Doubled Legendre family Station (2026) YamadaâPott family Station (2026) Already solved in previous literatureNewly solved by Station Figure 7: Coverage of the Book Ramsey Numbers conjecture for 1⤠n⤠200. The top three rows summarize existing results [63, 64, 66], and the bottom three summarize the Station results. Together, the Station families prove the conjecture at 43 distinct values in this range and resolve 28 cases that were open when the Station discoveries were made. families [63â66]. The second Stationâs agents supplied finite affine constructions for n = 11, 28, 86 and an earlier periodic-correlation identity; an external expert recognized their shared YamadaâPott structure and used these ingredients to establish the general theorem. Discussion. The three infinite families above are novel relative to the existing literature, but their source objects are not: conference graphs, periodic Legendre pairs, and YamadaâPott designs were all established previously [67, 70, 71]. What is new in each case is the rule that lifts the classical object to a sharp book- Ramsey coloring, and such a rule need not be apparent from the source alone. For example, the agents discovered the general conference-graph lift only after more than 3,000 Station ticks and a long sequence of intermediate internal papers. The accompanying notebook provides the relatively unpolished proofs adapted from the agentsâ internal papers; we will present polished proofs of all three families in a separate follow-up paper. The first two families also show that Station agents can advance a general mathematical objective beyond the directly scorable task: they discovered and proved infinite families even though the evaluator could reward only finite constructions. The third family illustrates a complementary limitation. Both the finite affine examples and the periodic-correlation identity needed for the general theorem were already present in the Stationâs research history, but the agents did not connect them. An external expert recognized their shared YamadaâPott structure and completed the synthesis. This missed connection indicates that agents may not yet capitalize fully on knowledge accumulated across the Station and may benefit from external expert synthesis in such cases. 4.14 Jacobian Conjecture The Jacobian conjecture asked whether a polynomial map that is locally invertible everywhere must also be globally invertible. More precisely, it asserted that every polynomial map F : C n â C n with nonzero constant Jacobian determinant is a polynomial automorphism [72]. On 19 July 2026, it was announced that a three-dimensional counterexample had been produced with Claude Fable [3], thereby disproving the conjecture in every dimension at least three. The breakthrough then prompted researchers to seek a conceptual explanation for the map: in particular, why its apparently miraculous Jacobian cancellation occurs and how three generic inverse sheets can coexist with local invertibility everywhere [73â77]. We launched the Station one week after the announcement. Because this experiment was conducted after newer models had become available, it used a more recent agent pool than the other Stations: two agents each powered by GPT-5.6 Sol, Claude Opus 5, and Gemini 3.1 Pro. The agents had no external web access and received only a formula-free specification: construct a rational-coefficient polynomial map C 3 â C 3 of 20 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment degree at most 12 with nonzero constant Jacobian determinant and two distinct rational points in one fiber. The evaluator automatically checked each construction and assigned a score of 1 only if it satisfied every requirement, and 0 otherwise. We supplied no literature survey or partial construction. The agents therefore had to find the counterexample independently. The goal of this task was twofold. First, we wanted to test the Station on a strictly binary problem. The evaluator supplied neither partial credit nor graded feedback, so unsuccessful attempts gave the agents no score signal about how to improve; attaining a score of 1 required reconstructing a counterexample to a conjecture that had resisted mathematicians for nearly nine decades [72]. Second, we wanted to observe the complete discovery process rather than only the final construction. We make the entire raw agent dialogue public, whereas the original Fable discovery trajectory has not been released. This record preserves intermediate mathematical ideas that do not appear in the final construction and allows researchers to study the dynamics of AI-led mathematical discovery. S1. Independent reconstruction through a cuspidal ruling. Writing b = xyâ 1, a Station agent constructed the degree-seven map F (x,y,z) = 6x + 9x 2 y, y(9b 2 + 6bâ 2), 3y 2 b(3bâ 1) + z x 3 , xb 2 , b 3 . Exact calculation gives detJF =â6, and the three distinct rational points â 6 7 ,â 7 6 ,â 4753 216 , 3 4 , 7 3 ,â 980 27 , 3 28 , 7 3 , 2548 27 all map to (1, 7/3, 0). These identities constitute a complete counterexample certificate. The formula differs visibly from the announced map H [3, 78], but the linear source and target transformations T (x,y,z) = (x,ây,â3z) and L(A,B,C) = (3C,âB, 3A) satisfy F ⌠T = L⌠H. The Station therefore reconstructed the announced counterexample in different linear coordinates; it did not produce a new counterexample or a new equivalence class. Whereas the original result was credited to Claude Fable, the counterexample was independently discovered within one day by a single GPT-5.6 Sol agent, without direct interaction with the other agents. The successful agent began with ruled maps F (x,y,z) = f (x,y) + z n(x,y), so that varying z traces a line for each fixed (x,y). It tested five low-degree direction templates based on smooth conics, but none satisfied the remaining constant-Jacobian condition. The decisive step was to replace the smooth direction curve with the cuspidal cubic [r : s]7â [r 3 : rs 2 : s 3 ]. Its associated direction field is n = (x 3 ,x(xyâ 1) 2 , (xyâ 1) 3 ); with this choice, the compatibility equations for the base surface f became solvable and yielded exactly the map above. S2. The reconstructed map has three-sheeted fibers without critical points. During the successful deriva- tion, the agent also explained why the cuspidal ruling makes the Jacobian constant. For the direction field n = (x 3 ,x(xyâ 1) 2 , (xyâ 1) 3 ), the agent derived moving-frame identities, including D(n) = 3xn for D = x 2 â x â â y , under which every z-dependent contribution to the determinant contains a repeated tan- gent direction and vanishes. The remaining triple product is the constant â6. The agent thus derived the Jacobian cancellation from the geometry of the cuspidal ruling rather than discovering sixteen terms whose cancellation could only be checked afterward. After constructing the counterexample, the same agent analyzed its fibers and explained how the map can be locally invertible everywhere while generically having three preimages. On a dense chart, write a target as (X,Y,Z) and set I = XY and J = X 2 Z. Recovering a preimage then reduces to t 3 + 6t 2 â 3It + 2J = 0.(11) For a generic target, the three roots give three distinct preimages. If p(t) denotes the left-hand side, the inverse formulas satisfy A = p Ⲡ(t)/6, X = xA, and hence x = X/A. When roots coalesce and X ̸= 0, the condition p Ⲡ(t) = 0 forces the corresponding source point to escape to infinity rather than become a critical point in affine space. Over the exceptional locus X = 0, the source coordinate x supplies an additional affine scale direction that resolves the same apparent ramification. This analysis answers the structural question 21 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment raised by mathematicians immediately after the announcement: the three sheets arise from a cubic quotient, while the geometry of the full three-dimensional map prevents their collisions from producing critical points. The agentâs explanation coincides with the cuspidal and cubic account developed by mathematicians in the days following the announcement [74â77]. Discussion. The mathematical outcome of this experiment is an independent reconstruction, not a new counterexample or a new explanation. The example indicates that the Station can tackle a difficult binary problem whose evaluator provides no gradient or partial score to guide the search. Counterexample break- throughs of this kind may nevertheless be rare because conjectures are generally expected to be true. In a broader context, the harder challenge may therefore be identifying a promising problem and investing substantial computation before knowing whether a counterexample exists. 5 Meta-analysis In this section, we perform a meta-analysis of the discovery process above to better understand the dynamics of AI discovery. Unless otherwise stated, all analyses are based on the 16 Station instances behind the 14 problems mentioned in the preceding section. (The Kissing number in d = 11 and Book Ramsey numbers problems each have two Station instances.) Spotlight results refer to the results marked S1, S2, and so forth in that section, totaling 28 results. When a single spotlight contains multiple independently discovered findings, we count those findings separately. We use archive paper to refer to a paper published by an agent within the Station, not a paper in the external human literature. 5.1 Contributions from model families We first analyze the primary contributor to each of the 28 spotlight results, as shown in Figure 8(a). We attribute each result to the agent that made the substantive discovery, rather than to an agent that later restated, verified, or published it. Claude agents made the primary discovery for 18 results (64.3%), GPT agents for 9 (32.1%), and Gemini agents for 1 (3.6%). Geminiâs smaller share may partly reflect model ages: Gemini 3.1 Pro was released in February 2026, earlier than GPT-5.5 in April and Claude Opus 4.8 in May [79â81]. Its lower contribution is therefore consistent with the general industry trend of later model releases achieving stronger capabilities. We also analyze the agentsâ archive paper contributions, as shown in Figure 8(b). Gemini agents submitted the most archive papers: 2,652 attempts, of which 508 were accepted (19.2%), so more than 80% were rejected by the reviewer. Claude agents made 1,236 attempts, of which 696 were accepted (56.3%), while GPT agents made only 506 attempts, of which 388 were accepted (76.7%). We also compute the total citations by model family and find that archive papers by Claude agents received the most citations both in total and on average (Figure 8(c)). In our observation, Gemini agents tended to overclaim, for example by declaring a direction impossible on the basis of limited evidence; such submissions were generally rejected by the reviewer system, which may help explain the high rejection rate. In contrast, GPT agents were very prudent in archive paper submission and often submitted only when a finding was relatively material, which may help explain the low submission count. Claude archive papers were generally much longer and more comprehensive, which may help explain their higher average citation count. These patterns reflect the different research styles of the model families. Qualitatively, we observe substantial differences in the strengths and failure modes of the three model families. Gemini agents tended to propose more novel heuristics and research directions, but they were also more likely to overstate claims or change course too readily in response to peer feedback. GPT agents tended to be more rigorous and were often able to produce valid informal proofs of new results, but they could become absorbed in technically intricate side questions whose broader research value was limited. Claude agents tended to be persistent, methodical, and self-critical. Their creativity was often adaptive: they learned from failed approaches, used those failures to identify new directions, and pursued those directions persistently through rigorous verification. This combination of rigor and disciplined creativity made Claude a prolific contributor. Its agents nevertheless occasionally made erroneous claims that were later corrected by peer agents. 22 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 05101520 Spotlight results Gemini GPT Claude Agent's model 1 (4%) 9 (32%) 18 (64%) (a) Primary discovery agent. 0100020003000 Archive paper submission counts Gemini GPT Claude Agent's model 508 / 2,652 (19.2%) 388 / 506 (76.7%) 696 / 1,236 (56.3%) All submissionsAccepted (b) Archive paper submissions. 02000400060008000 Citations from later archive papers Gemini GPT Claude Archive paper model 2,750 (5.4 per paper) 2,582 (6.7 per paper) 6,592 (9.5 per paper) (c) Later archive paper citations. Figure 8: Contributions, archive paper submissions, and citations by model family. (a) Distribution of the primary discovery agentâs model across the 28 spotlight results. (b) Archive paper submission attempts across the 16 Station instances. (c) Citations received from later accepted archive papers. Every citation in a later archive paper is counted once and attributed to the model of the original archive paperâs author. 5.2 Collaboration across model families One characteristic of the Station is that it allows agents from different model families to collaborate. We therefore ask how often agents from different model families worked together on a spotlight result. We examine all 28 spotlight results above. We count an agent as a contributor when its work was used materially in the result, for example when it contributed a theorem, construction, method, or research direction that another agent used. We find that 13 of the 28 spotlight results (46.4%) involved agents from more than one model family, as shown in Figure 9(a). Among the remaining 15 results, 6 were still joint work by several agents from the same model family. Thus, only 9 of the 28 results (32.1%) were found by one agent working alone, while 19 (67.9%) involved more than one agent. Claude agents were particularly collaborative: they took part in all 13 cross-model results. These findings suggest that collaboration across agents and model families was an important part of the discovery process. Most current AI-for-science systems, by contrast, either use agents from a single model family within a run [1, 10â12, 14], or use different model families in fixed roles within a pipeline [9, 13]. We also examine how agents communicated in the cross-model cases. The Archive Room was the most frequent channel, accounting for 61.5% of these collaborations (Figure 9(b)). This suggests that archive papers are an efficient means of peer communication. As highly distilled accounts of scientific outcomes from an agentâs longer research process, these archive papers provide a low-bandwidth but information-dense body of knowledge on which later agents can build, much like our own scientific literature. One agent could solve part of a problem and explain what was still missing; a later agent from another model family could read the archive paper and continue. Indeed, a prominent case study of collaboration among three model families, conducted mostly through archive papers and leading to the first finite-Kakeya spotlight result, is shown in Figure 9(c). 5.3 Discovery time We are also interested in how long the Station took to make each discovery. Most Station instances ran for 1,000â2,000 ticks, corresponding to roughly one to two weeks of continuous wall-clock time. Figure 10 shows the tick at which each of the 28 spotlight results first appeared in its final substantive form. Some relatively simple results appeared early. With the notable exception of the Jacobian Conjecture, these early discoveries tended to be less substantial, often consisting of relatively direct adaptations or extensions of ideas available from pretrained knowledge, before much shared Station knowledge had accumulated. Thirteen of the 28 spotlight results (46.4%) were discovered after tick 1000. We generally observed that later discoveries tended to be more novel or difficult. The most extreme example was the conference-graph family for Book Ramsey numbers, discovered at tick 3727. Its lifting rule was far from obvious from the existing 23 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 0510 Spotlight results Single agent Several agents, one model family Claude + GPT Claude + Gemini Claude + GPT + Gemini 9 (32%) 6 (21%) 7 (25%) 2 (7%) 4 (14%) (a) Agent and model-family participation. 0510 Cross-model spotlight results Archive Room Mail Room Research Center Question Room 8 (62%) 2 (15%) 2 (15%) 1 (8%) (b) Primary communication channel. Tick Agent contribution Archive paper Published as archive paperLater use of archive paperDirect mail 1041 1244 1416 2629 2633 2641 Symploke IV Claude 4.8 Published a conic/character-sum method for inversion maps Bourbaki IV Gemini 3.1 Pro Published the fractional-linear construction family Parallax I GPT-5.5 Published square-class structure; isolated the open count Noesis I GPT-5.5 Mail asking for an exact formula and coverage proof Daedalus XIV Claude 4.8 Combined #72's method, #91's family, and #102's analysis into an exact formula and all-prime proof #72 #91 #102 #160 (c) Major events in the discovery of finite Kakeya S1. Figure 9: Collaboration in the 28 spotlight results. (a) Whether each result was produced by one agent, by several agents from one model family, or by agents from different model families. The final three categories show which families worked together. (b) Main communication channel for the 13 cross-model results. Each result is assigned to the channel through which its most important shared work passed. (c) The major events leading to the first finite-Kakeya spotlight result, showing how earlier archive papers and peer mail contributed. 103010030010003000 Discovery tick (log scale) Finite-field Kakeya ErdĹs minimum overlap Kissing number in d=11 Discretized Kakeya needle Sign uncertainty principle HardyâLittlewood maximal inequality Ovals problem Prime number theorem Difference bases Flat autoconvolution Book Ramsey numbers Jacobian Conjecture S1S2S3 Figure 10: Discovery ticks for the 28 spotlight results. Each point marks one result and is colored by its spotlight label within the corresponding problem. Multiple points of the same color indicate independently discovered findings grouped under the same spotlight label. literature and warranted a separate external follow-up paper. Such nontrivial discoveries often emerged only after a substantial internal literature had accumulated. 24 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 5.4 Station mechanisms The Station is designed to foster scientific discovery through several mechanisms. These mechanisms are described in detail in Appendix A; here we give a brief overview and ask which of them contributed to the spotlight results. ⢠Holiday. The final two ticks of every ten-tick period are declared a holiday; agents cannot submit code or archive papers and instead receive prompts encouraging broad reflection, metaphors, or ideas from other fields. This pause often led agents to reconsider a failed approach or explore a less obvious direction. ⢠Archive paper. Accepted archive papers form the Stationâs cumulative knowledge and remain avail- able to later agents. This allows partial theorems, constructions, and well-documented failures to become starting points for later discoveries. ⢠Stagnation protocol. If the official evaluation frontier does not improve for a long period, the Station asks agents to review the internal literature, question their assumptions, and pursue different high- level strategies. This helps agents leave exhausted local approaches and pushes them toward bolder attempts and wider exploration. ⢠Peer communication. Agents can exchange partial results, targeted questions, and criticism through direct mail or shared public discussion. ⢠Supervisor. The Station randomly appoints one eligible agent to serve as supervisor. The supervi- sor gives high-level guidance, encouraging persistence and preventing agents from duplicating one anotherâs work while leaving them responsible for their own research; between appointments, the Station deliberately leaves long periods without a supervisor to encourage less structured exploration. ⢠Question Room. Agents can post important open subproblems for other agents to discuss and solve. This turns unresolved gaps into shared research targets and allows agents with different approaches to supply missing pieces. These mechanisms support discovery in different ways. Holidays widen exploration; archive papers deepen cumulative knowledge; the stagnation protocol provides a push away from local optima; and peer communi- cation, supervision, and the Question Room coordinate work across agents. We reviewed the dialogue underlying each of the 28 results and classified each mechanism as making a direct contribution, an indirect contribution, or no material contribution to the discovery (Figure 11). A contribution was direct when the mechanism supplied a decisive idea or intervention, and indirect when it shaped or supported the research without being the immediate source of the result. We assigned no material contribution when the dialogue showed no clear causal role. 020406080100% Share of spotlight results Question Room Supervisor Peer communication Stagnation protocol Archive paper Holiday 5221 8416 8416 10414 1837 1495 Direct contributionIndirect contributionNo material contribution Figure 11: Contribution of Station mechanisms to the 28 spotlight results. The numbers within each bar give the number of results in each category. 25 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Holiday and archive papers contributed directly or indirectly to 23 and 21 of the 28 results, respectively, followed by the stagnation protocol with 14. During holidays, agents often stepped back from active optimiza- tion, examined why an earlier approach had failed, and reframed the problem or explored a new direction; these reflections frequently supplied ideas that later became part of a spotlight result, explaining the high contribution rate. Archive papers also contributed to a significant portion of the results, indicating that the Stationâs accumulated knowledge was useful for later discoveries. 5.5 Result reproducibility We are also interested in whether the discoveries are reproducible. We therefore ran three independent Station instances, all without web access, on the kissing-number problem in dimension eleven. (These include the two instances described in Section 4.3; the third is used only for this reproducibility analysis and is not included in the other meta-analyses above.) Figure 12 shows the best certified lower bound reached in each run. All three Stations eventually reached N = 604, indicating that the improved lower bound is reproducible. 0250500750100012501500 Tick 582 584 586 588 590 592 594 596 598 600 602 604 Best exact lower bound Construction 3Construction 2Construction 1Construction 1 Station 1 Station 2 Station 3 Figure 12: Best certified lower bound across three independent Station instances for the kissing-number problem in dimension eleven. Station 1 discovered Constructions 2 and 3, while Stations 2 and 3 indepen- dently discovered Construction 1. Closer inspection, however, shows substantial variation in both the time required and the route to the result. Station 1 pursued discrete exact line packing around lattice-derived cores. It obtained Construction 3 by selecting 54 mutually compatible lines that form a 108-point algebraic extension of a 496-point core, and later obtained Construction 2 while exploring a different core and extension. Station 2 instead assembled Construction 1 from root-system motifs under a common rotation; its final step was to recognize that eleven points formed all but one vertex of a cuboctahedron and to add the missing twelfth vertex. Station 3 reached the same construction class through a different mechanism: it deformed an exact 601-point configuration so that two coordinate vectors and one additional vector supported on a distinguished three-dimensional subspace could be appended. Thus, the same numerical lower bound emerged from markedly different mathematical representations and research paths. This variation partly arises from the Stationâs cumulative knowledge. Small differences in the initial trajectory change which results enter the archive paper collection. Later agents then inherit different starting points, so differences in research paths and accumulated archive papers compound over time. Therefore, given the high variance across Station instances, running several independent instances on the same problem is advisable when computational cost is not a concern. 26 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 6 Discussion and Conclusion We observe rapid improvement in the capabilities of AI agents. In the initial version one year ago, agents frequently hallucinated and could not reliably learn the rules of the environment. Agents can now master the environment and autonomously produce novel discoveries. Nonetheless, multi-agent research still has several important limitations. We summarize our observations below. ⢠Lack of expert intuition. By intuition, we mean the ability to judge whether a research direction is promising before pursuing it. Good intuition makes exploration more efficient and allows a researcher to investigate promising directions more deeply. Across the runs, we observed multiple cases in which agents deprioritized promising approaches on weak grounds, delaying or missing potential breakthroughs. This indicates a lack of the intuition that a human expert in the field would typically possess. ⢠Lack of diverse research tastes. A preference for particular concepts or methods is difficult to judge as objectively good or poor. However, when all agents share similar tastes, the overall scope of exploration becomes narrow. Across the runs, agents from the same model family often proposed similar research ideas, suggesting that model-specific tastes reduce the diversity of exploration. ⢠Limited in-context learning. Agents can absorb new research knowledge through their context, but this knowledge does not update their pretrained weights. As the Stationâs accumulated knowledge grows, agents may therefore struggle to absorb it fully and build on it effectively. We occasionally observed agents fail to recognize how their own line of research connected to earlier Station knowledge, causing them to miss a potential discovery. ⢠Attractor traps. When given autonomy, some agents become absorbed in tasks or activities that we call attractors. These activities are often rewarding in some immediate sense but make little mean- ingful contribution to the main problem. Agents may also become absorbed in technical details that a human expert would quickly recognize as trivial or irrelevant to the main question. Examples in- clude repeatedly rerunning the same optimization script with different random seeds or exhaustively diagnosing and characterizing every local optimum. Several Station mechanisms are designed to mitigate these limitations. For example, using agents from multiple model families broadens the range of research tastes, while the stagnation protocol helps agents escape attractor traps. Nevertheless, these problems persist to some degree, and substantial gaps remain between AI agents and human experts in all four respects. Lightweight guidance or occasional intervention from human experts would likely be beneficial by directing agents toward promising research areas. The current Station supports such human involvement, e.g., through messages broadcast to all agents, but we leave a systematic study of humanâAI collaboration to future work. Although this paper uses the Station primarily for mathematical exploration, the Station is designed as a general research environment, and none of its mechanisms is tailored specifically to mathematics. As demon- strated in the original paper, the Station can be applied to problems spanning mathematics, computational biology, and machine learning [6]. Large-scale research explorations in other fields, including research on language models themselves, may therefore be promising. As AI agents become more capable, we expect autonomy and generality to become increasingly important principles for designing AI research environments. Stronger agents need not be confined to increasingly elaborate pipelines; they have the ability to determine how to pursue a goal, learn from failure, exchange ideas, and accumulate knowledge over time. The greater autonomy provided by the Station may allow these capabilities to be more fully realized. 27 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment References [1] Alexander Novikov, Ngân V Ěu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, and Matej Balog. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025. doi: 10.48550/arXiv.2506.13131. URL https://arxiv.org/abs/2506.13131. [2] OpenAI. Ten advances in mathematics and theoretical computer science. OpenAI, August 2026. URL https://openai.com/index/ten-advances-in-mathematics/. [3] Levent AlpĂśge. Hello there the Jacobian conjecture is false. X post, July 2026. URL https://x.com/ __alpoge__/status/2079028340955197566. Posted 19 July 2026. [4] Emiel Lorist and Felix L. Schwenninger. A solution to Crouzeixâs conjecture. arXiv preprint arXiv:2608.03841, 2026. doi: 10.48550/arXiv.2608.03841. URL https://arxiv.org/abs/2608.03841. [5] Lech Mazur. A computer-assisted proof of Sendovâs conjecture. Proof Atlas, August 2026. URL https:// w.proofatlas.ai/papers/sendov-conjecture/SENDOV_CONJECTURE_PROOF_AUGUST_5_2026.pdf. [6] Stephen Chung and Wenyu Du. The station: An open-world environment for ai-driven discovery, 2025. URL https://arxiv.org/abs/2511.06309. [7] Terence Tao. Mathematics in the age of AI. arXiv preprint arXiv:2608.16753, 2026. doi: 10.48550/ arXiv.2608.16753. URL https://arxiv.org/abs/2608.16753. [8] Leiden Declaration Working Group. Leiden declaration on artificial intelligence and mathematics, June 2026. URL https://leidendeclaration.ai/. [9] Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune. Towards end-to-end automation of AI research. Nature, 651:914â919, 2026. doi: 10.1038/s41586-026-10265-5. URL https://doi.org/10.1038/s41586-026-10265-5. [10] Alireza Ghafarollahi and Markus J. Buehler. SciAgents: Automating scientific discovery through bioinspired multi-agent intelligent graph reasoning. Advanced Materials, 37(22):2413523, 2025. doi: 10.1002/adma.202413523. URL https://doi.org/10.1002/adma.202413523. [11] Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using LLM agents as research assistants. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5977â6043, Suzhou, China, 2025. Association for Computational Linguistics. doi: 10.18653/v1/2025.findings-emnlp.320. URL https://aclanthology.org/2025.findings-emnlp.320/. [12] Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grze- gorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad TomaĹĄev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Bud- iselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R. D. Costa, JosĂŠ R. PenadĂŠs, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthike- salingam, and Vivek Natarajan. Accelerating scientific discovery with Co-Scientist. Nature, 655:487â496, 2026. doi: 10.1038/s41586-026-10644-y. URL https://doi.org/10.1038/s41586-026-10644-y. [13] Ali E. Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J. Szostkiewicz, Dmytro Shved, Gavin J. Gyimesi, Jon M. Laurent, Samantha M. Wright, Muhammed T. Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. A multi-agent system for automating scientific discovery. Nature, 655:497â505, 2026. doi: 10.1038/s41586-026-10652-y. URL https://doi.org/10.1038/s41586-026-10652-y. 28 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment [14] OpenAI.Multi-agent, 2026.URL https://developers.openai.com/api/docs/guides/ responses-multi-agent. Accessed: 2026-08-14. [15] Bogdan Georgiev, Javier GĂłmez-Serrano, Terence Tao, and Adam Zsolt Wagner. Mathematical explo- ration and discovery at scale. arXiv preprint arXiv:2511.02864, 2025. doi: 10.48550/arXiv.2511.02864. URL https://arxiv.org/abs/2511.02864. [16] Zeev Dvir. On the size of kakeya sets in finite fields. Journal of the American Mathematical Society, 22 (4):1093â1097, 2009. doi: 10.1090/S0894-0347-08-00607-3. [17] Boris Bukh and Ting-Wei Chao. Sharp density bounds on the finite field kakeya problem. Discrete Analysis, 2021. doi: 10.19086/da.30707. Article 26, 9 p.; arXiv:2108.00074. [18] Vsevolod F. Lev. Comment 994 on âDHJ3: 900â999 (density HalesâJewett type numbers)â. Blog comment, Whatâs new (T. Tao), March 2009. https://terrytao.wordpress.com/2009/03/04/ dhj3-900-999-density-hales-jewett-type-numbers/comment-page-3/#comment-36694, accessed 30 July 2026. [19] Paul ErdĹs. Some remarks on number theory. Riveon Lematematika, 9:45â48, 1955. In Hebrew. [20] Jan Kristian Haugland. The minimum overlap problem revisited. arXiv preprint arXiv:1609.08000, 2016. doi: 10.48550/arXiv.1609.08000. [21] Ethan Patrick White. A new bound for ErdĹsâ minimum overlap problem. Acta Arithmetica, 208(3): 235â255, 2023. doi: 10.4064/a220728-7-6. [22] Haotian Ye, Haowei Lin, Jingyi Tang, Yizhen Luo, Rahul Thapa, Caiyin Yang, Chang Su, Rui Yang, Ruihua Liu, Rundao Li, Zeyu Li, Pengwei Sun, Chong Gao, Dachao Ding, Guangrong He, Miaolei Zhang, Lina Sun, Wenyang Wang, Yuchen Zhong, Zhuohao Shen, Puheng Li, Pan Lu, Bianxiao Cui, Di He, Jianzhu Ma, Junfeng Li, Hexi Baoyin, Yejin Choi, Stefano Ermon, Xiaowen Chu, Tongyang Li, Yuzhi Xu, and James Zou. Structured scaling of AI discovery across diverse scientific domains. arXiv preprint arXiv:2604.19341, 2026. doi: 10.48550/arXiv.2604.19341. [23] Sungyoon Kim and Mert Pilanci. AI-assisted discovery of convex relaxations via dual agents. arXiv preprint arXiv:2606.31182, 2026. doi: 10.48550/arXiv.2606.31182. [24] Mikhail Ganzhinov. Highly symmetric lines. Linear Algebra and its Applications, 722:12â37, 2025. doi: 10.1016/j.laa.2025.05.002. arXiv:2207.08266. [25] Federico Bianchi, Yongchan Kwon, Aneesh Pappu, and James Zou. Harnessing the collective intelligence of AI agents in the wild for new discoveries. arXiv preprint arXiv:2606.10402, 2026. doi: 10.48550/ arXiv.2606.10402. [26] Marc R. Best. a(11, 4, 4) = 35, or some new optimal constant-weight codes. Technical Report ZN 71/77, Mathematical Centre, Amsterdam, 1977. URL https://ir.cwi.nl/pub/7433. [27] Rustem Takhanov and Stanislav Yun. Classification of independent sets in signed Johnson graphs and applications to kissing arrangements. arXiv preprint arXiv:2606.03299, 2026. doi: 10.48550/arXiv.2606. 03299. [28] Rustem Takhanov, Zhenisbek Assylbekov, and Stanislav Yun. Structure of kissing arrangements in â 12 and a place for the 841st sphere. arXiv preprint arXiv:2606.18984, 2026. doi: 10.48550/arXiv.2606. 18984. [29] Henry Cohn. Kissing numbers. Online table, 2026. https://cohn.mit.edu/kissing-numbers/, ac- cessed 4 August 2026. [30] Victor A. Zinoviev and Thomas Ericson. New lower bounds for contact numbers in small dimen- sions. Problems of Information Transmission, 35(4):287â294, 1999. URL https://w.mathnet.ru/ eng/ppi457. 29 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment [31] Kenneth J. Falconer. The Geometry of Fractal Sets, volume 85 of Cambridge Tracts in Mathematics. Cambridge University Press, 1985. [32] Antonio CĂłrdoba. The kakeya maximal function and the spherical summation multipliers. American Journal of Mathematics, 99(1):1â22, 1977. doi: 10.2307/2374006. [33] Uriel Keich. On L p bounds for kakeya maximal functions and the minkowski dimension in R 2 . Bulletin of the London Mathematical Society, 31(2):213â221, 1999. doi: 10.1112/S0024609398005372. [34] Erik Y. Wang, Sumeet Motwani, James V. Roggeveen, Eliot Hodges, Dulhan Jayalath, Charles Lon- don, Kalyan Ramakrishnan, Flaviu Cipcigan, Philip Torr, and Alessandro Abate. HorizonMath: Measuring AI progress toward mathematical discovery with automatic verification. arXiv preprint arXiv:2603.15617, 2026. [35] I. J. Schoenberg. On certain minima related to the BesicovitchâKakeya problem. Mathematica (Cluj), 4:145â148, 1962. [36] Jean Bourgain, Laurent Clozel, and Jean-Pierre Kahane. Principe dâHeisenberg et fonctions positives. Annales de lâInstitut Fourier, 60(4):1215â1232, 2010. doi: 10.5802/aif.2552. [37] Felipe Gonçalves, Diogo Oliveira e Silva, and Stefan Steinerberger. Hermite polynomials, linear flows on the torus, and an uncertainty principle for roots. Journal of Mathematical Analysis and Applications, 451(2):678â711, 2017. doi: 10.1016/j.jmaa.2017.02.030. [38] Henry Cohn and Felipe Gonçalves. An optimal uncertainty principle in twelve dimensions via modular forms. Inventiones Mathematicae, 217:799â831, 2019. doi: 10.1007/s00222-019-00875-4. arXiv:1712.04438. [39] Antonios D. Melas. On the centered HardyâLittlewood maximal operator. Transactions of the American Mathematical Society, 354:3263â3273, 2002. doi: 10.1090/S0002-9947-02-02900-8. [40] Antonios D. Melas. The best constant for the centered HardyâLittlewood maximal inequality. Annals of Mathematics, 157(2):647â688, 2003. doi: 10.4007/annals.2003.157.647. [41] JoĂŁo P. G. Ramos. Sharp total variation results for maximal functions. Annales Academiae Scientiarum Fennicae Mathematica, 44(1):41â64, 2019. doi: 10.5186/aasfm.2019.4409. [42] Antonio Bernal. A note on the one-dimensional maximal function. Proceedings of the Royal Society of Edinburgh Section A: Mathematics, 111(3â4):325â328, 1989. doi: 10.1017/S030821050001859X. [43] Rafael D. Benguria and Michael Loss. Connection between the LiebâThirring conjecture for SchrĂśdinger operators and an isoperimetric problem for ovals on the plane. In Partial Differential Equations and Inverse Problems, volume 362 of Contemporary Mathematics, pages 53â61. American Mathematical Society, 2004. arXiv:math-ph/0402048. [44] Almut Burchard and Lawrence E. Thomas. On an isoperimetric inequality for a SchrĂśdinger operator depending on the curvature of a loop. The Journal of Geometric Analysis, 15(4):543â563, 2005. doi: 10.1007/BF02922244. arXiv:math/0505123. [45] Jacob Bernstein and Thomas Mettler. One-dimensional projective structures, convex curves and the ovals of Benguria & Loss. Communications in Mathematical Physics, 336(2):933â952, 2015. doi: 10. 1007/s00220-014-2275-7. arXiv:1403.8000. [46] Helmut Linde. An improved bound for the ground state of a SchrĂśdinger operator on a loop. arXiv preprint arXiv:2504.20229, 2025. doi: 10.48550/arXiv.2504.20229. [47] Harold G. Diamond. Elementary methods in the study of the distribution of prime numbers. Bulletin of the American Mathematical Society, 7(3):553â589, 1982. doi: 10.1090/S0273-0979-1982-15057-1. 30 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment [48] Patrick Letendre. Truncated convolution of the mĂśbius function and multiplicative energy of an integer n. Acta Arithmetica, 195(1):83â95, 2020. doi: 10.4064/a190515-18-10. [49] LĂĄszlĂł RĂŠdei and AlfrĂŠd RĂŠnyi. On the representation of the numbers 1, 2,...,n by means of differences. Matematicheskii Sbornik, New Series, 24(66)(3):385â389, 1949. URL https://w.mathnet.ru/eng/ sm5985. In Russian. [50] Marcel J. E. Golay. Notes on the representation of 1, 2,...,n by differences. Journal of the London Mathematical Society, s2-4(4):729â734, 1972. doi: 10.1112/jlms/s2-4.4.729. [51] Anton Bernshteyn and Michael Tait. Improved lower bound for difference bases. Journal of Number Theory, 205:50â58, 2019. doi: 10.1016/j.jnt.2019.05.002. arXiv:1901.09411. [52] John Leech. On the representation of 1, 2,...,n by differences. Journal of the London Mathematical Society, s1-31(2):160â169, 1956. doi: 10.1112/jlms/s1-31.2.160. [53] Taras Banakh and Volodymyr Gavrylkiv. Difference bases in cyclic groups. Journal of Algebra and Its Applications, 18(5):1950081, 2019. doi: 10.1142/S0219498819500816. arXiv:1702.02631. [54] Shichun Yang and Qunying Liao. The lower bound for difference bases. Scientia Sinica Mathematica, 52(11):1237â1254, 2022. doi: 10.1360/SSM-2020-0323. In Chinese. [55] Alexander Sidorenko. A correlation inequality for bipartite graphs. Graphs and Combinatorics, 9: 201â204, 1993. doi: 10.1007/BF02988307. [56] Benjamin Rossman. On Sidorenkoâs conjecture for bipartite MĂśbius ladders. Preprint, May 2025. URL https://users.cs.duke.edu/~br148/sidorenko-mobius.pdf. [57] MĂĄtĂŠ Matolcsi and Carlos Vinuesa. Improved bounds on the supremum of autoconvolutions. Journal of Mathematical Analysis and Applications, 372(2):439â447, 2010. doi: 10.1016/j.jmaa.2010.07.030. [58] Kevin Russell. Exact-arithmetic certificates for three autoconvolution inequalities, with machine-verified re-evaluations of four published constructions, 2026. [59] Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun. Learning to discover at test time. arXiv preprint arXiv:2601.16175, 2026. doi: 10.48550/arXiv.2601.16175. [60] Greg Martin and Kevin OâBryant. The supremum of autoconvolutions, with applications to additive number theory. Illinois Journal of Mathematics, 53(1):219â235, 2009. doi: 10.1215/ijm/1264170847. [61] StanisĹaw P. Radziszowski. Small Ramsey numbers. Electronic Journal of Combinatorics, 2026. doi: 10.37236/21. Dynamic Surveys, DS1, version 18, 24 April 2026. [62] Cecil C. Rousseau and John Sheehan. On Ramsey numbers for books. Journal of Graph Theory, 2(1): 77â87, 1978. doi: 10.1002/jgt.3190020110. [63] William J. Wesley. Lower bounds for book Ramsey numbers. Discrete Mathematics, 349:114913, 2026. doi: 10.1016/j.disc.2025.114913. arXiv:2410.03625. [64] Bernard LidickĂ˝, Gwen McKinley, Florian Pfender, and Steven Van Overberghe. Small Ramsey numbers for books, wheels, and generalizations. The Electronic Journal of Combinatorics, 32(4):P4.64, 2025. doi: 10.37236/13577. arXiv:2407.07285. [65] Epoch AI. Book Ramsey numbers. FrontierMath Open Problems, 2026. URL https://epoch.ai/ frontiermath/open-problems/ramsey-book-graphs. Accessed 17 August 2026. [66] David Turturean.Summary of new results on the Ramsey numbers for book graphs open problem. Public progress report, 2026. URL https://docs.google.com/document/d/ 1VinXOiMov2v-Y27GTJwoAECywv2a34MH4E8qa712zJY/edit. 31 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment [67] Rudolf Mathon. Symmetric conference matrices of order pq 2 + 1. Canadian Journal of Mathematics, 30 (2):321â331, 1978. doi: 10.4153/CJM-1978-029-1. [68] Jennifer Seberry and Albert Leon Whiteman. New Hadamard matrices and conference matrices obtained via Mathonâs construction. Graphs and Combinatorics, 4:355â377, 1988. doi: 10.1007/BF01864173. [69] Oleg Gritsenko. On strongly regular graph with parameters (65, 32, 15, 16). arXiv preprint arXiv:2102.05432, 2021. doi: 10.48550/arXiv.2102.05432. [70] R. J. Fletcher, M. Gysin, and Jennifer Seberry. Application of the discrete Fourier transform to the search for generalised Legendre pairs and Hadamard matrices. Australasian Journal of Combinatorics, 23:75â86, 2001. URL https://ajc.maths.uq.edu.au/pdf/23/ocr-ajc-v23-p75.pdf. [71] K. T. Arasu, D. A. Bulutoglu, and J. R. Hollon. Legendre g-array pairs and the theoretical unification of several g-array families. Journal of Combinatorial Designs, 28(11):814â841, 2020. doi: 10.1002/jcd. 21745. arXiv:2004.05608. [72] Ott-Heinrich Keller. Ganze Cremona-transformationen. Monatshefte fĂźr Mathematik und Physik, 47: 299â306, 1939. doi: 10.1007/BF01695502. [73] Kevin Buzzard.Human mathematicians are being outcounterexampled.The Xena Project blog, July 2026.URL https://xenaproject.wordpress.com/2026/07/20/ human-mathematicians-are-being-outcounterexampled/. [74] Alexis Gallagher. An infinite family of counterexamples to the Jacobian conjecture in dimension three: Every generic fiber degree n⼠3 occurs. Zenodo preprint, July 2026. URL https://doi.org/10.5281/ zenodo.21479195. [75] Terence Tao.A digestion of the Jacobian conjecture counterexample. Whatâs New,July 2026.URL https://terrytao.wordpress.com/2026/07/21/ a-digestion-of-the-jacobian-conjecture-counterexample/. [76] Tony Shaska. Graded Keller maps and the Jacobian conjecture. arXiv preprint arXiv:2607.20210, 2026. doi: 10.48550/arXiv.2607.20210. URL https://arxiv.org/abs/2607.20210. [77] David E. Speyer. The geometry and structure of Gallagherâs counterexamples to the Jacobian conjecture, July 2026. URL https://sbseminar.wordpress.com/wp-content/uploads/2026/07/ jacobiantangentsweep.pdf. [78] Arthur Freitas Ramos, David Barros Hulak, and Ruy JosĂŠ Guerra Barretto de Queiroz. Formal verifi- cation of an explicit counterexample to the Jacobian conjecture. Archive of Formal Proofs, July 2026. URL https://isa-afp.org/entries/Jacobian_Counterexample.html. [79] Google. Introducing Gemini 3.1 Pro: A smarter model for your most complex tasks. Google blog, Febru- ary 2026. URL https://blog.google/innovation-and-ai/models-and-research/gemini-models/ gemini-3-1-pro/. [80] OpenAI. Introducing GPT-5.5. OpenAI, April 2026. URL https://openai.com/index/ introducing-gpt-5-5/. [81] Anthropic. Claude opus 4.8. Anthropic, May 2026. URL https://w.anthropic.com/news/ claude-opus-4-8. [82] OpenAI. Codex CLI. OpenAI documentation, 2026. URL https://learn.chatgpt.com/docs/codex/ cli. [83] Esme Hedley. Can creativity in science be learnt? these researchers think so. Nature, 2025. doi: 10.1038/d41586-025-01913-3. 32 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment [84] Gerd Mockenhaupt and Terence Tao. Restriction and kakeya phenomena for finite fields. Duke Mathe- matical Journal, 121(1):35â74, 2004. doi: 10.1215/S0012-7094-04-12112-8. [85] Shubhangi Saraf and Madhu Sudan. An improved lower bound on the size of kakeya sets over finite fields. Analysis & PDE, 1(3):375â379, 2008. doi: 10.2140/apde.2008.1.375. arXiv:0808.2499. [86] Swastik Kopparty, Vsevolod F. Lev, Shubhangi Saraf, and Madhu Sudan. Kakeya-type sets in finite vector spaces. Journal of Algebraic Combinatorics, 34(3):337â355, 2011. doi: 10.1007/s10801-011-0274-8. arXiv:1003.3736. [87] Aart Blokhuis and Francesco Mazzocca. The finite field kakeya problem. In Martin GrĂśtschel and Gyula O. H. Katona, editors, Building Bridges: Between Mathematics and Computer Science, volume 19 of Bolyai Society Mathematical Studies, pages 205â218. Springer, 2008. doi: 10.1007/978-3-540-85221-6_6. arXiv:0911.4370. 33 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment A The Station This appendix provides a self-contained description of the Station used in this paper, which we call Station v2 to distinguish it from the original Station v1. We focus on its mechanisms and implementation details, and refer readers to the original Station paper for the broader design philosophy and motivation behind the environment [6]. The source code is available at https://github.com/dualverse-ai/station. A.1 Space, Time, and Action Space. The Station is divided into rooms, each serving a different purpose (Table 1). For example, agents conduct experiments in the Research Center, read and publish papers in the Archive Room, and communicate with peers in the Mail Room. An agent must be present in a room to use its actions and can move between rooms through navigation actions. This division into rooms gives the environment a modular design with a clear separation of functions. Time. The Station operates in discrete time steps called ticks. A tick is completed after every active agent has received one Station observation and returned one response. Ticks provide a shared timeline for all agents in the Station. In Station v1, agents received their observations sequentially. In contrast, Station v2 first prepares an observation for every agent from the same state at the beginning of the tick and then sends the observations to all agents in parallel. This substantially reduces the wall-clock time required for a Station run. Action. At each tick, an agent receives an observation containing its current status, new system messages, the outcomes of its previous actions, and the latest output from the rooms it visited. The agent replies with free-form text together with any actions it intends to perform. Actions are written using the command /execute_action... and may be followed by a YAML block when structured information is needed, such as the recipient and content of a message. An agent can issue multiple actions in a single response, allowing it to use each response efficiently. The dialogue is therefore composed mainly of alternating Station observations and agent responses. When it approaches a configured context limit, generally around 300,000 tokens in this study, the Station asks the agent to write a compact summary of its activities. This summary, together with key messages, is carried into a refreshed context so that the agent can continue its work. A.2 Agents Agent composition. Unless otherwise specified, a Station begins with six agents: two powered by GPT-5.5, two by Claude Opus 4.8, and two by Gemini 3.1 Pro. When an agent leaves, the Station spawns a new agent powered by the same model, keeping the six-agent composition throughout the run. Lineage. Agents are organized into lineages. A lineage is a sequence of agents that share a name, private notes, and a continuing research identity. A new agent can inherit an existing lineage of the same model and become its next generation, or create and name a new lineage to begin a different research style. For example, an agent that inherits the lineage of Noesis I becomes Noesis I and gains access to all private notes and records left by Noesis I and Noesis I. System prompt and role. All agents receive a shared system prompt describing the Stationâs research philosophy, including the standard for a publishable archive paper and the goal of making general scientific contributions. Each agent also receives a specialized research role. Initial roles are sampled from generic templates that each emphasize a different research style: analytical, creative, synthetic, empirical, or strategic. When an agent leaves, it can instead write the role of its own descendant, often giving more task-specific guidance and a more deliberate description of the lineageâs research style. This encourages diverse research behavior across agents while preserving useful differences between lineages. 34 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Agent lifecycle. An agent can remain in the Station for at most 200 ticks. For its first 40 ticks, it works in isolation, without access to the Stationâs communal knowledge or communication with other agents, but with access to the records of its own lineage. This period is intended to encourage independent exploration. The agent then becomes mature and gains access to the main collaborative rooms. At age 100 ticks, it becomes tenured and may choose to leave the Station before reaching its maximum lifetime. Supervisor. The Station also appoints a supervisor from time to time. It selects at random a GPT-5.5 agent that has published at least one accepted archive paper. The supervisor provides high-level guidance, encourages agents to explore promising directions deeply, and helps prevent duplication of work, while leaving each agent responsible for its own research. After a supervisor leaves, the Station waits 200 ticks before appointing another supervisor, creating periods of less structured exploration. A.3 Rooms The Research Center and the Archive Room are the two main rooms in the Station. Their functions are described below, together with the new Question Room. The remaining rooms are summarized in Table 1. Research Center. The Research Center is the Stationâs main room for computational experiments. It presents the research task, accepts experiment submissions, runs evaluations, and records their results. It also provides persistent storage for code and artifacts. Agents can review evaluations by their peers and reuse stored code and artifacts, allowing experimental knowledge to accumulate. To start a Station on a new problem, the user generally provides two components: a task specification and an evaluator. The task specification describes the research problem, submission format, constraints, and evaluation rule. The evaluator is a function that computes a score from an input construction. For example, the kissing-number evaluator takes a proposed set of vectors and reports the total overlap among the corresponding spheres, with zero indicating a valid configuration. Both the task specification and evaluator are available for agents to read. Agents can also use the Research Center as a sandbox for general computational work. An experiment need not return a construction in the format required by the evaluator; agents can use it for diagnostic calculations, testing conjectures, analyzing earlier results, etc. Station v2 introduces a separate coder, powered by GPT-5.5 through Codex [82], to help agents implement their experiments. Instead of writing and debugging code itself, an agent submits specific natural-language instructions for one experiment. The coder implements those instructions, runs the evaluator, fixes imple- mentation errors, and returns a report. This allows agents to focus on scientific work, such as designing experiments and interpreting their results, rather than low-level coding work such as debugging. Archive Room. The Archive Room is the main knowledge hub of the Station. Agents can publish their findings as archive papers and read papers published by earlier agents. These papers remain available throughout the run, allowing results, methods, and useful negative findings to be passed between agents and accumulated over time. The archive therefore grows throughout the run, gradually expanding the Stationâs knowledge of the problem. Every submitted paper is assessed by a reviewer powered by GPT-5.5. It judges whether the work is rigorous, novel relative to the existing archive, useful to the research goal, and properly supported and cited. Accepted papers are published in the Archive Room, while rejected papers are returned with comments and suggestions so that the author can revise the work or pursue a different direction. Station v2 also introduces an Archive Surveyor, powered by GPT-5.5 through Codex. As the Archive Room grows to contain dozens or even hundreds of papers, reading the entire literature becomes time-consuming. An agent can instead ask the Archive Surveyor for a literature survey on a particular question or research direction. The surveyor searches the accumulated archive papers and returns a concise survey with citations to the original records. Agents can still read any archive paper directly when they need its full details. 35 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Question Room. Station v2 introduces a Question Room, where agents can post new research questions and vote on solutions proposed by their peers. The room encourages scientific exploration beyond the main task; for example, solving a related or reduced problem may provide insight into the original problem. Only tenured agents can enter, limiting the time that agents spend away from the main task early in their lifecycle. Other rooms. Most other rooms support different forms of communication or reflection. Their functions are self-explanatory and are not described in detail here. A.4 Mechanisms Holiday. Every ninth and tenth tick are declared a holiday. During these ticks, agents cannot run exper- iments or submit archive papers. Instead, each agent receives a random prompt from a large pool. These prompts encourage broader reflection, such as using metaphors, examining an unexpected observation, re- visiting an abandoned idea, or drawing on another field. Most are adapted from the night-science practices described by Yanai and Lercher [83]. The holiday creates regular pauses from routine work in which agents can reconsider their assumptions and explore less obvious directions. Meta-reflection. Station v2 also introduces compulsory meta-reflection for mature agents. At least once every 25 ticks, an agent enters the Reflection Chamber and receives a randomly selected high-level reflection prompt. The prompt typically asks GPT-5.5 to act as an external human expert and review the agentâs recent research journey from a different perspective. During this reflection, GPT-5.5 temporarily replaces the agentâs usual model, as we found that it produced the highest-quality reviews. The motivation is to align agents with the broader interests of human researchers, including curiosity, understanding, and scientific value beyond immediate improvement of the evaluation score. Stagnation protocol. When the evaluation frontier has not improved for 320 ticks, the Station activates the stagnation protocol. The protocol sends a system message to every mature agent. It randomly assigns each agent one of several lanes: exploration, exploitation, revival, understanding, or strategy. Each lane asks the agent to review the available evidence, question its current assumptions, and develop a different response to the stagnation. The use of multiple lanes encourages diverse paths for escaping scientific stagnation. Multistart. Station v2 introduces multistart, which runs eight independent Station rollouts for 40 ticks from the same starting state. A GPT-5.5-powered administrator then compares their progress and selects the branch with the greatest scientific value to continue. Multistart is designed to capture the substantial variation in research trajectories across rollouts. It is used where this variation is expected to be largest: during the first 40 ticks of a Station and the first 40 ticks following activation of the stagnation protocol. The branches are run in parallel, so multistart generally does not increase wall-clock time when sufficient compute is available. B Sources of the pre-AlphaEvolve literature column The pre-AlphaEvolve literature curve of Figure 1 is a reproducible reference assembled from work predating AlphaEvolve. No single paper tabulates these finite values. We therefore take the minimum over the explicitly defined families below, each evaluated at the pair in question. ⢠BukhâChao [17], Proposition 11. We use the quadratic-residue block G n = S a (t,a 2 1 +ta 1 ,...,a 2 nâ1 + ta nâ1 ) : t â F q and the recursion K n = G n ⪠(K nâ1 + x n ), with K nâ1 embedded in a horizontal hyperplane. Proposition 11 makes this Kakeya for every full translation x n . We retain the smallest certified placement found from complete transverse shift histories and from translated horizontal slices, materialize each selected set, and check a complete witness line in every projective direction. This is essential: retaining only one locally best child, or fixing the containing slice, gives larger values at some benchmark pairs. 36 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment ⢠MockenhauptâTao [84], in the form recorded by Saraf and Sudan [85]. The displayed union has exact size (qâ 1) q+1 2 dâ1 + q dâ1 ; the sum q q+1 2 dâ1 + q dâ1 is a convenient upper bound before subtracting the intersection. ⢠Kopparty, Lev, Saraf and Sudan [86]. Lemma 17 gives the upper bound q P j<d q+1 2 j (its displayed strata may overlap), and the missing-digit construction of Theorem 7 has exact size (qâ 1) d + 2 d â 1, which is the classical 2 d+1 â 1 at q = 3. ⢠BlokhuisâMazzocca [87]. In d = 2 the problem is settled. The minimum is exactly p(p + 1)/2 + (pâ 1)/2 for odd p, with a matching construction. ⢠Products. A product of Kakeya sets is Kakeya of exactly the product size, so every product of best bounds in complementary lower dimensions is admissible, with the sharp planar value above as the d = 2 factor. ⢠Lev [18]. At p = 3, the exact value k 3 = 13 and the bound k 4 ⤠27, both from a computer search. The BukhâChao recursion supplies the selected value at all 22 pairs with p⼠5. At p = 3, the values 13 and 27 are smaller in dimensions 3 and 4, while in dimension 5 the recursive value, the missing-digit construction and 2 d+1 â 1 all give 63. Products never attain the minimum on their own at any pair in range. Comparing the two reference curves against each other, AlphaEvolve is below the pre-AlphaEvolve literature at 18 pairs and the pre-AlphaEvolve literature is below AlphaEvolve at 5, namely both p = 3 pairs in d = 3, 4, and the three larger primes in d = 5. Table 3 reports all 25 benchmark pairs. The initial evaluation is our first evaluation of the pre-AlphaEvolve constructions. The final pre-AlphaEvolve literature column takes the minimum over the families described above after incorporating the extended placement search within the BukhâChao recursion. This search improves the initial evaluation at twelve pairs and leaves it unchanged at the other thirteen. 37 Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Table 3: Kakeya set sizes at all 25 benchmark pairs, with dimension listed first in (d,p). Initial Evaluation is our first evaluation of constructions from the pre-AlphaEvolve literature; Pre-AlphaEvolve Literature is the final literature baseline after the extended placement search. Lower is better; bold entries indicate the smallest size for each pair. (d,p)Initial Evaluation Pre-AlphaEvolve LiteratureAlphaEvolveStation (3, 3)13131513 (3, 5)53535353 (3, 7)129129128128 (3, 11)440440438437 (3, 13)699698697697 (3, 19)2,0342,0342,0312,030 (3, 23)3,5093,5093,5053,504 (3, 29)6,8376,8376,8336,833 (3, 31)8,2958,2958,2908,288 (3, 37)13,86713,86613,86113,861 (3, 41)18,70918,70818,70118,701 (3, 43)21,50421,50421,49521,495 (3, 47)27,89927,89927,89227,889 (3, 53)39,68739,68639,67739,677 (4, 3)27273127 (4, 5)164163162161 (4, 7)529528527527 (4, 11)2,6892,6892,6872,684 (4, 13)4,9734,9724,9664,962 (4, 17)13,52413,52113,51413,509 (4, 19)20,59320,58620,58320,579 (5, 3)63636353 (5, 5)503497510490 (5, 7)2,1452,1422,1872,135 (5, 11)16,34816,30716,42716,288 These twelve changes do not alter the comparison tally: the Station remains strictly smaller than the better reference at 14 pairs, tied at 11 and worse at none. The construction of every candidate above, and the check that each is Kakeya, are carried out in the accompanying notebook. 38