Paper deep dive
Can AI Follow In Einstein's Footsteps?
Michael Shalyt, Nathan Regev, Marin SoljaÄiÄ, Ido Kaminer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/1/2026, 2:20:32 AM
Summary
The paper argues that AI's trajectory in physics discovery is reversing the historical human progression from pattern prediction to principle-based theory. While early AI focused on symbolic regression (equation discovery), modern frontier AI (e.g., AlphaFold, GraphCast) prioritizes high-accuracy black-box prediction over theoretical understanding. The authors contend that to achieve paradigm-level discoveries like quantum gravity, AI must develop the ability to pose fundamental questions, invent principles, and construct new mathematical frameworks rather than merely optimizing predictive performance.
Entities (10)
Relation Signals (9)
AI Discovery â mirrorsinreverse â Human Discovery
confidence 95% · most visible AI contributions to physics discovery appear to mirror the historical development of physics, but in reverse
GraphCast â exemplifies â Black-box Prediction
confidence 90% · GraphCast ... treat complex phenomena like fluid simulations as high-dimensional statistical prediction problems
AlphaFold â exemplifies â Black-box Prediction
confidence 90% · AlphaFold ... demonstrates that neural networks can achieve predictive performance without explicit symbolic formulas
Black-box Surrogates â lacks â Theoretical Understanding
confidence 90% · powerful predictors such as AlphaFold and GraphCast, which can be remarkably accurate yet do not provide clear theoretical understanding
Human Discovery â progressedfrom â Pattern Prediction
confidence 90% · Human discovery in physics progressed, in broad strokes, from ancient pattern prediction, through phenomenological laws ... to principle-based universal theories
Human Discovery â progressedto â Principle-Based Theory
confidence 90% · Human discovery in physics progressed ... to principle-based universal theories such as relativity and the Standard Model
Einstein â used â Principle-Based Theory
confidence 90% · Einstein did not build general relativity using data ... He started with abstract principles of symmetry
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI is accelerating physics discovery, but perhaps away from Einstein-level theory building. To understand this gap, we must recognize a striking trend: while being very successful, the most visible AI contributions to physics discovery appear to mirror the historical development of physics, but in reverse. Human discovery in physics progressed, in broad strokes, from ancient pattern prediction, through phenomenological laws such as Kepler's, to principle-based universal theories such as relativity and the Standard Model. On the AI side, prominent contributions to physics discovery point in the opposite direction: early milestones emphasized explicit equation-discovery methods, such as symbolic regression, whereas more recent frontier contributions are powerful predictors such as AlphaFold and GraphCast, which can be remarkably accurate yet do not provide clear theoretical understanding. If this trend continues, AI would become extraordinarily good at prediction but may struggle to ever propose its first serious contender to quantum gravity or other paradigm-level theories. We review the current landscape of AI for physics discovery and highlight a critical missing skill: the ability to pose the right questions or invent the right principles to guide the development of new theories and the tests to falsify them. This mode of discovery has driven many of the deepest advances since the 17th century, where symmetry, simplicity, and new mathematical frameworks guided theory construction before experimental tests. Equipping AI systems with such skills could move them from predicting within known frameworks to proposing the next paradigm-level discovery in physics.
Tags
Links
- Source: https://arxiv.org/abs/2607.27794v1
- Canonical: https://arxiv.org/abs/2607.27794v1
Trouble viewing inline? Open PDF directly â
Full Text
67,933 characters extracted from source content.
Expand or collapse full text
CAN AI FOLLOW IN EINSTEINâS FOOTSTEPS? Michael Shalyt 1 Nathan Regev 1 Marin Solja Ë ci Ì c 2 Ido Kaminer 1â 1 Technion - Israel Institute of Technology, Haifa 3200003, Israel. 2 Massachusetts Institute of Technology, Cambridge MA 02139, USA. ABSTRACT AI is accelerating physics discovery, but perhaps away from Einstein-level theory building. To understand this gap, we must recognize a striking trend: while being very successful, the most visible AI contributions to physics discovery appear to mirror the historical development of physics, but in reverse. Human discovery in physics progressed, in broad strokes, from ancient pattern prediction, through phe- nomenological laws such as Keplerâs, to principle-based universal theories such as relativity and the Standard Model. On the AI side, prominent contributions to physics discovery point in the opposite direction: early milestones empha- sized explicit equation-discovery methods, such as symbolic regression, whereas more recent frontier contributions are powerful predictors such as AlphaFold and GraphCast, which can be remarkably accurate yet do not provide clear theoretical understanding. If this trend continues, AI would become extraordinarily good at prediction but may struggle to ever propose its first serious contender to quantum gravity or other paradigm-level theories. We review the current landscape of AI for physics discovery and highlight a critical missing skill: the ability to pose the right questions or invent the right principles to guide the development of new the- ories and the tests to falsify them. This mode of discovery has driven many of the deepest advances since the 17th century, where symmetry, simplicity, and new mathematical frameworks guided theory construction before experimental tests. Equipping AI systems with such skills could move them from predicting within known frameworks to proposing the next paradigm-level discovery in physics. 1AUTOMATION OF SCIENTIFIC SKILLS In Douglas Adamsâ The Hitchhikerâs Guide to the Galaxy, the supercomputer Deep Thought fa- mously calculates the âAnswer to the Ultimate Question of Life, the Universe, and Everythingâ to be 42. Adams captured the idea that real progress in physics often requires figuring out the right question rather than just finding an accurate answer. Current developments in AI for science, and especially in AI for physics, risk heading toward the Deep Thought conundrum in real life. The Per- spective below reviews the history of advances in AI for physics to highlight a growing emphasis on black-box approaches â often at the expense of simpler solutions and deeper physical understanding â before proposing paths forward. Recent years have shown major strides in the ability of AI systems to accelerate scientific discovery across most fields of science 1â3 . New algorithms increasingly demonstrate skills previously associ- ated with expert scientists 4,5 . Early milestones include automated proof construction, exemplified by the computer-assisted proof of the four-color theorem 6 and interactive theorem provers such as Lean 7 . Subsequent work simulated creative mathematical reasoning through automated conjecture generation in graph theory (Graffiti 8 ), number theory (the Ramanujan Machine 9,10 ), knot theory 11 , and matrix multiplication 12 . In physics, algorithmic tools long ago began automating core components of scientific reasoning. Pattern-recognition skills, such as anomaly detection in cosmology and high-energy physics ex- periments, are now routinely delegated to machine learning systems 13 . AI systems also automate the construction of predictive symbolic formulas from data with symbolic regression systems such as BACON 14,15 , Eureqa 16 , and SINDy 17 . Recent advances in this area include the recovery of physical equations from the Feynman Lectures on Physics by the AI Feynman Project 3 , derivation of complex â Corresponding author: kaminer@technion.ac.il. 1 arXiv:2607.27794v1 [physics.hist-ph] 30 Jul 2026 ecological and climate equations by knowledge-guided deep symbolic regression 18 , and discovery of an analytical formula candidate for predicting the concentration of dark matter from the mass distribution of nearby cosmic structures 19 . Although these and other early AI achievements mim- icked human-like pattern recognition, more recent prominent contributions of AI to physics follow a different trajectory 20 . Neural networks expanded the skill of generalization and no longer search for explicit symbolic formulas: The most prominent success of AI for science to date, protein folding by AlphaFold 21â23 , demonstrates that neural networks can achieve predictive performance without explicit symbolic formulas, yet providing accuracy surpassing decades of human effort. Conceptually similar efforts have increasing impact in additional fields, such as predictive models trained on weather data 24 . In contrast to the approaches above, when the governing formulas are known, they can help the training process by constraining the learning objective or by shaping the inductive bias of the neural network, as in physics-informed neural networks (PINNs) 25,26 . In other settings, established physical formulas can be replaced by learned surrogate models 27â30 that emulate complex dynamics while reducing computational costs. In all of the above, algorithms have successfully simulated specialized skills. More recently, large language models (LLMs) and their reasoning variants gradually automate general research skills 31â34 . Examples include Olympiad-level mathematical reasoning 35â37 , unifying mathematical knowledge representations 38 , autonomous proofs of open questions 39â43 , and analytical representa- tions of scattering amplitudes in high-energy physics 44,45 . Although much of the discussion around AI for science treats disciplines uniformly, physics has a unique relationship with AI. Physical platforms such as integrated photonics inspire new hard- ware architectures for accelerating computation 46 . Concepts from statistical mechanics inform re- search on explainability 47 , while other concepts inspire changing model architecture 48,49 . The fact that many physical systems are governed by compact, universal equations makes them particularly amenable to approaches such as PINNs, which embed these equations into the learning objective 50 . Agentic AI robots are being developed to automate experimental skills from building setups to per- forming experiments 51,52 . These interfaces are reviewed elsewhere 53â60 . Essential questions remain: What skills are required for AI to eventually create its first paradigm- shifting discovery in physics? How far are we from an AI-generated competitor to the Standard Model of particle physics or to ÎCDM of cosmology? The bottom line of the Perspective is that achieving such breakthroughs requires more than better prediction skills: It requires AI systems capable of identifying the right questions, sustaining high- risk hypotheses long enough to test them, and constructing new theoretical frameworks guided by abstract principles, possibly through the development of new kinds of mathematics. 2THE REVERSE EPISTEMIC TRAJECTORY There is no single story describing the discovery process in the history of physics, but until re- cently, it followed a trajectory towards deeper levels of understanding. AI is now following a similar trajectory, but apparently in reverse epistemic order (Fig. 1). The scientific revolution established a new standard for understanding nature: derivation of universal mathematical laws from empirical observation, as Newton famously framed in the Principia 61 . His achievement was to identify mathematical laws that govern phenomena across scales, from falling apples to orbiting planets. Since then, the formulation of universal laws from empirical observation has been regarded as the hallmark of physical âunderstandingâ. From the perspective of AI, symbolic regression aims to automate precisely this process: finding mathematical laws from empirical data. Yet more recent developments in AI for science, such as GNoME 62 , provide accurate predictions without explicit mathematical laws, echoing physics discoveries before the scientific revolution. It is in this sense that AI discovery in physics is following the history of human discovery in reverse order. This reversal is not a universal law of AI for science, nor a chronology of AI architectures. Rather, it reflects human incentives and practical choices that have recently shifted the fieldâs center of gravity: many of the most visible and influential AI-for-science successes are now produced by systems optimized for predictive power rather than for discovery of explicit, human-interpretable theories. 2 These successes coexist with â but increasingly overshadow â efforts to make AI discover organizing principles and mathematical structures, and thereby become a participant in theory building. Figure 1: The trajectories of human vs AI physics discovery â reversal in explainability. The ar- rows summarize example contributions representing an apparent trend. Many milestone discoveries throughout history followed a trend: from pattern prediction to phenomenological laws and finally to principle-based theories, the level of âunderstandingâ is increasing. In contrast, the center-of-gravity of AI contributions to physics has followed this trend in the reverse direction: from discovery of explicit symbolic formulas to neural network predictive systems with decreasing explainability. At a coarse level, the history of physics may be viewed as progressing through three broad stages: Pattern Prediction - Ancient systems, such as the Metonic cycle or Mayan astronomy, predicted celestial events with high accuracy via cycles and correlations, without explicit mathematical laws. Phenomenology - Keplerâs three laws and Planckâs early black-body radiation formula provided mathematical equations that fitted the data but lacked the underlying principles to explain it. Principle-Based Theory - Newton, Maxwell, Heisenberg, Schrödinger, Einstein, and others es- tablished fundamental principles and a new mathematical language, from which specific laws could be deduced. The development of successful AI contributions to physics has mirrored the two earlier stages of human physics, but in reverse order. Early symbolic-regression models such as BACON, Eureqa, and SINDy extract equations from data â akin to discovery of phenomenological laws. More recent AI systems based on neural networks 24 increasingly abandon explicit equations and function as high-accuracy predictive oracles â closer in spirit to pre-scientific pattern prediction. This historical sketch is necessarily approximate, not a strict historical timeline. One caveat is worth highlighting. In the eyes of some ancient scientists, their theory was also principle-based, even if their principles were often metaphysical rather than based on mathematical laws. Ptolemaic astron- omy, for instance, aimed to represent celestial motion guided by abstract principles about cosmic harmony and the perfection of idealized circular motion. Kepler rejected the idea that his laws are mere calculational devices fitting Tycho Braheâs data, instead explicitly framing his seminal book Astronomia nova as âcelestial physicsâ, seeking principles he regarded as genuinely explanatory. They did not see their work as phenomenological. In contrast, certain human discoveries that are seen today as symbols of principle-based advances were discovered without those principles, or even using the wrong ones. Maxwellâs route to field theory initially relied on mechanical scaffold- ings (âmolecular vorticesâ) that were later discarded while the equations survived 63 . This caveat does not overturn the broader trend. What matters in modern view is not whether his- torical actors believed they had principles, but whether the resulting discovery compressed more phenomena into a simpler and more powerful mathematical structure. That is the kind of progress 3 exemplified by relativity, quantum theory, and gauge theory. In contrast, much of current AI-driven discovery improves prediction while retreating from the desire for mathematical simplicity. What recent history reveals is a shift in ideals and âoptimization targetâ regarding AI physics discovery â from prediction with understanding to prediction without it. 3AI PREDICTION WITHOUT UNDERSTANDING The shift of many high-impact AI applications toward black-box surrogates reactivates a pre- scientific-revolution mode of discovery, albeit one operating with superhuman efficacy. A growing number of frontier AI for physics systems prioritize utility over the use of explicit mathematical laws, and indeed provide a lower level of (human) understanding 64,65 . This prioritization provides an immediate benefit in predictive payoff. Certain problems in physics still benefit from a partial reliance on explicit formulas and laws, as common in PINNs 25 . These AI models embed known physical laws, typically partial differential equations (PDEs), directly into the networkâs loss function to act as structured regularizers 66,67 . Similarly, design priors can be used to enforce exact physical constraints 68 , such as embedding Lagrangian mechanics to guarantee energy conservation 69 . While such neural networks are highly effective, their scope is still limited to particular domains, as they require an externally-supplied theoretical foundation â which is not discovered by the networks. In contrast, increasingly popular approaches now use purely black-box surrogates such as Google DeepMindâs GraphCast 24 and NVIDIAâs Project NeRD 70 . Rather than enforcing known physical laws (e.g., Navier-Stokes equations), these surrogate systems treat complex phenomena like fluid simulations as high-dimensional statistical prediction problems 71 . By training on massive datasets, these systems often surpass the accuracy of traditional numerical solvers, yet they do so without an explicit representation of the fluid dynamics or mechanical principles involved. Historically, this predictive surrogate approach resembles ancient astronomical traditions such as the Antikythera mechanism 72 or Mayan calendrical systems. These were masterpieces of forecasting, capable of predicting eclipses and planetary positions with high precision, yet they possessed no understanding of gravity or orbital mechanics. Figure 2: Evolution of scientific theories. A schematic representation of the iterative, upward growth of scientific understanding. Every extension of an established theory provides predictions that are tested against observations, either refuting this theory ex- tension (tree stump) or upholding it (con- tinued growth of the trunk). Every exten- sion of the theory provides further applica- tions (fruits) and increased understanding. Paradigm Shifts, in the sense of Kuhn, could be considered as the sudden revelation of a separate tree providing the previous appli- cations and understanding, as well as new ones that cannot be otherwise accessed. We ask what developments would enhance the ability of AI systems to mimic the entire iterative process, and potentially even assist the discovery of unexpected paradigm shifts. From the broader perspective of the AI community, the fact that a model does not mirror the under- lying physics is no longer viewed as a problem. This shift in priorities can be attributed to the Bitter 4 Lesson 73,74 of AI development: methods that leverage massive computation and data often outper- form those relying on human-designed theoretical structures. In this view, the value of a model lies in whether it predicts accurately and efficiently and not in whether it recovers the true mechanism. In fact, a model can be extremely useful even if its internal âunderstandingâ is physically wrong. An analogous situation appears in the history of optics. The optical theories of Euclid, and later Ptolemy and Ibn al-Haytham, produced quantitatively successful descriptions of refraction and the passage of light through materials, despite operating on fundamentally incorrect premises regarding the na- ture of vision. Neural network black-box models recreate this situation almost inevitably. They are not optimized to recover the true underlying mechanism, but to compute accurate outputs with an efficient universal architecture based on matrix multiplications and nonlinear activation functions. This architecture can often accurately predict the data at lower computational cost or greater speed than explicit mathematical laws, while being oblivious to the true underlying equations. But we should not let this apparent advantage lead us to abandon the search for simpler, explicit theories. The Bitter Lesson explains why black-box predictors can dominate in data-rich regimes. Yet paradigm-level physics often begins in the opposite regime, where data is sparse or conceptually misinterpreted. In such cases, historical progress has often depended not on prediction alone, but on hypothesized mathematical structures that expose new variables, constraints, and principles. This is precisely the limitation identified by Bunge in âA General Black Box Theoryâ 75 : a theory that deals solely with the relationship between inputs (stimuli) and outputs (responses) lacks âcon- ceptual handlesâ that can be interrogated or proven wrong. A paradigm shift, in the Kuhnian sense 76 , requires a crisis within the underlying theory; however, a black-box predictor with no explicit struc- ture or epistemological interpretation cannot undergo such a crisis â it can only be re-optimized. If that AI model then undergoes a re-optimization granting it unprecedented predictive capabilities, on the level of a paradigm shift in their application â would it count as a true paradigm shift? Caveat (âIpcha Mistabraâ): Do we Really Need to Understand AI Discoveries? There are reasons to believe AI would also be better if it was always trying to come up with simplified, explainable representations of the world around it. But this is not necessarily so, and a provocative possibility is that some AI discoveries will rely on new types of mathematics that the AI will not be able to explain to humans but be able to communicate at a higher level (e.g., between AI agents) â so attempts at human understanding might just slow down progress. Related phenomena have already appeared in studies of emergent communication, where coop- erating agents can drift away from natural language toward task-specific communication pro- tocols 77,78 . More recently, LLMs have begun to explore reasoning in continuous latent spaces rather than through explicit natural-language chains of thought 79 . If such systems were to gen- erate powerful but non-human-interpretable representations, a central challenge would become how to recognize when a significant discovery had occurred at all (potentially by quantifying compression or complexity of the predictive structure). Could such an advance be considered a paradigm shift? Whatever status one assigns to non- human-interpretable advances, they do not remove the need for explicit, falsifiable theory in physics. The central agenda of this Perspective is therefore to argue for AI systems that do more than predict: they should help formulate principles, identify the right questions, and direct the experiments that can test them. Explainability in physics for AI alignment. AI models that predict without explaining how they do so are harder to supervise, falsify, verify, map their limitations, and eventually trust. Such trust is becoming vital as practical applications increasingly rely on black-box predictions and results. Even if alignment and accuracy are ensured by other means, there remains a need for human-level explainability to provide feedback and direct the AI research towards types of discoveries that so far remain inaccessible by AI systems. Shifting AI development from prediction without understanding to enhanced explainability is a prerequisite for future human-AI collaboration. 4CATEGORIES OF PHYSICS DISCOVERIES: HUMAN VS AI The deepest scientific breakthroughs often require more than new equations: they introduce new organizing principles, and sometimes entirely new mathematical language. In fundamental physics, 5 especially over the past century, many paradigm shifts have come precisely from this move toward new mathematical frameworks. This thinking motivates the classification of discoveries in physics into three rough categories. These categories allow us to verbalize the epistemological leaps required for discovery, and do not serve as strict mutually exclusive classifications. Sorting Physics Discoveries by Levels of Mathematical Abstraction CategoryCategory ACategory BCategory C Typenew solution or capabilitynew equation or lawnew math framework Description prediction, device, applica- tion theory, ansatz, formulamathematicallanguage, symmetry principle Human Discovery Examples ballistics and orbital pre- diction;photonic devices and antennas; transistors and lasers; collider phenomenol- ogy Keplerâs laws; Lorentz force and the Biot-Savart Law; semiconductor band theory andstimulatedemission; Standard Model and Higgs mechanism Newtonian mechanics and calculus; field theory and vectorcalculus;Hilbert space and operator formal- ism;renormalization and gauge theories in quantum field theory AI Discovery Examples Solid and fluid material simulationsusinggraph networks;protein folding and weather prediction using transformers; data-efficient approximators for problems in quantum mechanics and hydrodynamics using PINN analytic formula for dark matter concentration using symbolic regression; analyt- ical representations of scat- tering amplitudes by LLMs; âinvertingâ the process of ef- fective field theory by com- putational means ? To contribute to paradigm shifts, AI should move beyond making predictions and even beyond discovering new solutions and equations, toward inventing new mathematical frameworks. These categories also apply outside of physics. A classic example is the discovery of formulas for fundamental constants such as Ï and e by mathematicians over the centuries (Category A), many of them found to be special cases of the theory of continued fractions (Category B), recently found to be captured and unified by the mathematical framework of conservative matrix fields (Category C) 10,38 . Nevertheless, we should not expect every discovery to fall neatly within one of the categories A-B-C. General relativity is clearly Category C, whereas the Schwarzschild metric is technically a particular solution and thus closer to Category A. Yet some solutions have consequences so profound â reveal- ing black holes, horizons, singular behavior â that they reshape how we understand the framework from which they arise, and even point the way toward future Category C theory developments, such as modern efforts in quantum gravity. This blurring of categories points to the core issue: the most profound gap in current AI for physics discovery is not the absence of Category C outputs, but its inability to replicate the mode of dis- covery that produced many of the great advances of 20th-century physics. Current AIs (including LLMs and symbolic regression) excel at induction: evaluating how well a hypothesis matches the data. They also possess growing capabilities in deduction: computing the consequences of a given physical theory or framework (via simulation or logic). However, many great leaps of modern physics (e.g., general relativity, the Dirac equation, the Higgs mechanism) were neither inductive nor deductive. Einstein did not build general relativity using data about Mercuryâs orbit. He started with abstract principles of symmetry (equivalence, covariance), abducted a candidate theory satis- fying those principles, and only then sought empirical verification. This is the Popperian ideal: to construct bold theories, deduce empirical predictions, and subject those predictions to experimental tests aimed at falsifying them. This discovery pattern seems not to have been captured yet by current AI systems for physics. 6 Historical Examples of Principle-Guided Human Discoveries Laughlinâs theory for the fractional quantum Hall effect (Category A) found an explicit many- body wavefunction structure fixed by symmetry and simplicity in the lowest Landau level 80 ; de- duced incompressibility and fractionally charged excitations, capturing the essence of a new quan- tum fluid. Bethe ansĂ€tze (Category A) solved the one- dimensional spin chain by guessing a constrained eigenstate form and enforcing consistency condi- tions on the phases 81 ; strong constraints on trans- lation invariance and particle exchange converted the problem into closed-form equations. GinzburgâLandau theory of superconductivity (Category B) introduced an order-parameter field and free-energy functional consistent with sym- metries 82 ; postulating the right variables and in- variances predicted macroscopic superconducting behavior, later shown to work nearT c 83 . Parisiâs theory of spin glasses (Category C) pre- sented replica-symmetry breaking for disordered systems 84 , replacing a single macroscopic order parameter by a hierarchy of overlaps between pure states; resulting representation of equilibrium state space predicted many universal properties of com- plex systems. These physics discoveries all seem to rely on âwell-educated guessesâ, which have a shared structure. Even when based on established frameworks, they all impose an additional structured hypothesis, such as a certain ansĂ€tze, a symmetry, or gauge â enforcing simplicity before deducing falsifiable consequences. 5CAN AI AUTOMATE THE RARE INSIGHTS OF TOP PHYSICISTS? Could it be that AI is fundamentally unable to achieve human-level creativity? While this had been a long-standing question, a growing body of evidence suggests that there is no fundamental limit on AI creativity 85 . Among many examples, especially visible milestones include DeepMindâs highly unconventional âMove 37â by AlphaGo 86 that upended centuries of human expert intuition, and OpenAIâs disproof of Paul Erd Ì osâs famous planar unit-distance conjecture by an internal model 87 that combined techniques from two different fields of mathematics. An even more recent example is Anthropicâs counterexample to the Jacobian conjecture by Fable 88 that provided an initially opaque construction, subsequently reverse-engineered and generalized to human-understandable geometric terms 89 . These examples represent cases in which AI systems went beyond prior knowledge and made unexpected choices, which created new opportunities for subsequent human understanding. If the gap is not creativity, what missing skills, then, prevent AI from making Category C discov- eries? Unlike human physicists that have successfully driven progress across all three categories, AI-powered breakthroughs to date remain concentrated in Categories A and B, with the most visible recent successes increasingly falling on the predictive, low-explainability side of Category A â a trend that shows no signs of slowing. AI for physics of Category A solves physics problems or predicts the outcome of physical processes. Such AI systems include surrogate models trained on simulated data, where an expensive simulator is replaced by an efficient neural network 90,91 . When reliable simulations are unavailable, too expen- sive, or insufficiently accurate, neural networks can instead be trained directly on experimental data. Both simulation-trained and experiment-trained approaches can be enhanced using known physical formulas to improve training-data efficiency, robustness, and generalization. AI for physics of Category B provides an explicit theory or formula. Such AI systems include symbolic regression techniques usually building on experimental data. Other approaches rely on symbolic computation to make analytical predictions in search for more general theories consistent with experiments. These approaches are now enhanced by LLMs. Strikingly, no AI has made a discovery of Category C, involving the successful use of a new principle or the invention of a novel mathematical framework. Why is it so hard? The creation of a Category C discovery presents a challenge that current AI models are ill-equipped to meet. The primary bottleneck is not just producing a novel âcreativeâ idea. Novelty is cheap in a sufficiently large search space. The harder task is estimating the scientific payoff of each novel idea: judging how hard it is to execute, whether it is likely to become useful, and whether it yields falsifiable tests. This is where a payoff-per-effort ratio matters. Many top-tier human researchers possess a rare âgut feelingâ for this selection problem â a talent for identifying 7 research questions and elegant approximations that are both profound and practically executable. Part of this ability comes from experience, and part of it may be a rare form of scientific taste. Because this âgut feeling" is such a unique talent and is only rarely written down as a procedure, current AI systems have little direct training data for it. It is therefore unsurprising that current top AI models struggle to imitate the kind of judgment behind Category C discoveries. As long as we cannot train AI systems to imitate this undocumented feeling directly, our leading approach to AI- assisted progress in physics remains to build AI systems that can generate, test, and refine candidate principles until some acquire the marks of a useful theory. While it remains unclear how Category C breakthroughs could be fully automated, the following section sketches several strategies for steering AI for physics toward this higher-level mode of dis- covery. 6HOW TO TEACH AI BETTER MODES OF DISCOVERY? Several complementary strategies could help AI systems move toward Category C physics discover- ies. This section is not intended as a comprehensive survey or a closed roadmap; the field is develop- ing rapidly, and new approaches will likely emerge that are difficult to anticipate today. Rather, we highlight representative directions that seem especially relevant for principle-guided theory building and for raising the level of understanding and explainability in AI for physics discovery. Propose questions rather than answers. As famously shown in the Deep Thought story, it is (much) more important in physics to figure out the right question rather than merely finding the answer. Indeed, the formulation of the question is rarely a simple preliminary step â it often consti- tutes the discovery itself. Agentic AI systems can be trained to prioritize proposing hypotheses and research questions, then running closed-loop simulations to stress-test and simplify their proposals. In AI models already used for physics discovery, these exploratory preferences can be reinforced via modular skill definitions (such as markdown-based skill specifications [92]) and eventually inte- grated into a unified agent harness applying additional cognitive styles proposed below. Training to the tail and developing persistent âobsessionsâ. Current generative AI is architected to optimize for mainstream human thought (the statistical mode) rather than outliers (the statisti- cal tail). Through next-token prediction and reinforcement learning from human feedback (RLHF), LLMs are trained to smooth away âweirdnessâ in favor of the plausible consensus. Yet, paradigm- shifting breakthroughs are, by definition, radical outliers. Human discovery in physics has histori- cally often relied on a capacity for persistent âobsessionâ â a cognitive style that allows a researcher to hold onto an outlier idea over years. Einstein chewed on relativity for a decade; Bednorz and MĂŒller went against prevailing expectations for years before discovering cuprate superconductors. In comparison, leading AI models are currently trained to produce outputs that look like good think- ing from inside the consensus of a field. Revolutionary work, by definition, looks wrong from inside the consensus until it changes that consensus. LLMs are optimized to do the exact opposite: notice when the field disagrees and update toward the center. To move beyond this, we must build archi- tectures capable of proposing anti-consensus positions and sustaining them over time, while guiding searches to gradually accumulate evidence that can support or refute these positions. Striving for simple interpretable theories. There are strong reasons to believe that AI for physics would be more successful if trained to derive simpler theories and explicit equations from the pro- vided data. The inherent advantage of such explicit structures is evident in recent works: Incorpo- rating explicit laws into the AI model as inductive bias enables training with less data 93â95 and, more importantly, improves generalization capabilities 96â98 . Therefore, alongside prediction-driven AI models that will continue to be developed aggressively, we should invest in AI systems that search for lower-complexity theories. This approach is impor- tant even for the discovery of approximate laws, which remain scientifically powerful even when they are not exact: they can reveal the dominant variables, improve robustness, and support out- of-distribution generalization. More generally, whenever a simpler theory exists beneath a complex dataset, finding it is valuable not only for human understanding but also for automated extrapolation and generalization. The practical target is not only elegance for its own sake, but compression that makes the theory easier to use or test, and its ideas easier to extend or transfer across domains. 8 The Simplicity of Physical Laws: Guiding Star or Trap? âThe miracle of the appropriateness of the language of mathematics for the formulation of the laws of physics is a wonderful gift which we neither understand nor deserve.â Wigner 1960 99 One possible explanation for the âgiftâ Wigner described is the separation of scales. Natural phenomena containing behaviors on different scales of energy, distance, or time, often admit compact effective descriptions in terms of fewer parameters. Historically, this separation en- abled discovery of laws one scale at a time, a process that guided many discoveries of funda- mental laws of physics over the years. But the historical success of physics should not make us assume that an elegant theory lies un- derneath all domains of physics and other fields of science. For one, the electronic description of matter contains contributions that often cannot be separated in scale, causing inherent com- plexity in condensed matter physics. Similar challenges are fundamental to many domains of chemistry and biology. Where the scales cannot be separated, the search for an elegant underlying theory may itself be a hopeless pursuit. In such domains, accurate AI predictions without transparent explanation may in fact be the best one can hope for. The question is therefore not whether explainability is always preferable, but when the lack of explainability signals an unfinished theory and when it reflects genuine irreducible complexity. With the rise of AI-assisted discovery, learning to make this distinction becomes essential: when should we accept predictive accuracy as the best available outcome, and when should we con- tinue searching for a simpler underlying theory? In the case of the success of AlphaFold: is the lack of a closed-form formula for protein folding inherent to this process, or have we just missed the underlying structural laws? Either way, the predictive success of recent AI models should not tempt us to stop questioning whether a simpler structure remains hidden underneath. Enhancing theory building using symbolic computing. To move AI from prediction without understanding to the invention of new principle-based physical frameworks, AI systems require an environment where their abductive âguessesâ can be rigorously tested. Such an environment should be designed to cultivate in AI the aforementioned âgut feelingâ of top-tier human researchers. Symbolic computing tools could be used for this purpose, providing theory-building sandboxes. Computer algebra systems (CAS) could enable leading AI models and agentic LLM systems to write and execute code relevant to theory building in physics. AI agents connected to symbolic engines can manipulate exact mathematical objects, ensuring that a proposed theory is structurally coherent before it is tested against data. Despite the potential of this neurosymbolic approach, efforts in this direction have so far been lim- ited (with most examples being in AI for mathematics 100,101 ). Pioneering efforts have used sym- bolic computing tools to insert symmetries into AI pipelines. For instance, recent frameworks have demonstrated how to enforce or discover continuous symmetries via specialized regulariza- tions based on Lie derivatives 102 , embed fundamental space-time symmetries directly into network structures 103 , or integrate theoretical axioms into hypothesis generation to simultaneously satisfy data and physical laws 104 . Broader efforts in this direction utilize symmetry constraints in symbolic regression algorithms 105,106 . Existing methods primarily treat symmetry as a constraint for data fitting or model training. There remains a crucial step of âsymmetry abductionâ â toward principle-guided theory building, requiring development of closed-loop neurosymbolic architectures. We see early prospects of this concept in high-energy physics, where transformers have been trained on data generated by advanced symbolic tools to complete missing integer structures in scattering amplitudes 107,108 . Looking forward, AI must move from imposing symmetries and predicting amplitudes to propos- ing the Lagrangians themselves. A concrete route toward this goal is to convert theory-building into a constrained algorithmic search problem. Early efforts in this direction have been devel- oped in the Standard Model effective field theory (SMEFT). In SMEFT, candidate theories of ânew physicsâ are encoded through the most general higher-dimensional operator expansion consistent with gauge symmetries, organized by a strict simplicity criterion (power counting in 1/Î) 109,110 . 9 This approach has provided a consistent language for combining empirical constraints with beyond- Standard-Model scenarios 111 . The physics community has already built part of the required CAS infrastructure.For ex- ample, effective field theory (EFT) matching is increasingly automated via tools such as matchmakereft 112 and Matchete 113 . Building on this logic, recent computational work has demonstrated the inversion of the usual EFT workflow: starting from a target low-energy (IR) La- grangian and systematically searching for high-energy (UV) theories whose lower energy limits match it 114 . By automating part of the search over symmetries, gauge groups, and field properties, this approach illustrates how symbolic infrastructure can support the discovery of new candidate theories in domains such as quantum gravity. To achieve this at scale, AI systems should be trained to use established symbolic tools in these areas of physics. Future âAI physicistsâ could be trained to manipulate tensors and symmetries via xAct 115 and Cadabra 116,117 , process large symbolic expressions in FORM 118,119 , automate quan- tum many-body derivations in electronic-structure theory with SeQuant 120 , and handle Feynman diagrams using FeynCalc 121 and FeynRules 122 . By integrating such CAS code into agentic workflows, we can create the desired theory-building sandbox for AI systems. In it, the AI can gen- erate novel mathematical frameworks, filter out redundant or mathematically inconsistent ideas, and optimize for the rare, high âpayoff-per-effortâ insights that characterize the greatest human discov- eries. Artificial intuition via physics-aware world models. Human physicists are constantly driven to compress complex observations into simplified, consistent mental representations of reality. This internal âworld modelâ is a source of physical intuition, allowing scientists to filter out contradic- tory hypotheses and identify elegant approximations. This physical intuition is instrumental for the construction of model-based Gedankenexperiments that famously guided the development of new domains of physics. To develop a similar intuition in AI, there are strong reasons to believe we must architect systems that are explicitly forced to come up with simplified representations of their environment. Developing an artificial world model 123 that learns the underlying dynamics of how the world evolves could be the key to granting AI genuine physical intuition. Currently, efforts toward learning a world model are highly fragmented. Separate systems achieve superhuman accuracy in protein folding, global weather prediction, etc. It may be that each task- specific system already holds an underlying intuition of thermodynamics or fluid dynamics. This intuition could be shared among AI systems by direct communication between AI systems, or by one AI using others as tools. A promising direction for overcoming this fragmentation lies in training scientific foundation models on massive, heterogeneous corpora of direct physical measurements â encompassing microscopic, spectroscopic, astronomical imaging, and sensor data. Early indicators of this broader approach are emerging in AI models designed to unify multiple scientific data types within a shared learning framework 124 . The hypothesis is that scaling up multi-domain physical data will force the AI system to find trans- ferable structures or principles common across seemingly different domains of physics. If an AI is forced to predict both the Newtonian trajectories of planetary motion and the effect of gravitational lensing using the same latent space, it may be pushed to find a unifying framework, e.g., Einsteinâs relativity in this case. The hope is for such unification to lead to a compressed representation that can be translated back into explicit equations and human-understandable theory. 10 Is the Bottleneck AI, or Physics Itself? Some may argue that AI has failed to find the next general relativity because humans may not find one either: perhaps the era of simple, experimentally testable theoretical revolutions in physics is over. The past few decades could be read as evidence for this view. We believe it is too early to announce the end of theoretical physics. This concern should in- stead inform how we test AI systems for physics discovery. We can give AI systems artificial worlds with hidden laws and test whether they can invent simple theories inside those worlds. Alternatively, we can design historical reconstruction tasks such as Demis Hassabisâ â1911 cut- offâ test 125 , recently attempted in âMachina Mirabilisâ 126 : remove general relativity from the training distribution, provide only pre-relativistic observations plus consistency constraints, and evaluate whether the system can invent principles similar to Einsteinâs original insights. If AI systems can discover simple theories in artificial worlds, or rediscover known theories when the historical answer is withheld, but still fail to produce comparable advances in the real world â that failure would itself be important. It would suggest that the bottleneck may lie not only in AI architecture, but also in the intrinsic complexity or experimental untestability of the remaining undiscovered theories. Either outcome would be valuable. Success would show a path toward AI-assisted theory discovery, while failure under controlled conditions would teach us something profound about both AI and the present state of theoretical physics. 7FORMALIZING THE PARADIGM SHIFT If AI is to synthesize new theories, we must ask: what is the native language of physics discovery? Currently, generative AI relies predominantly on natural language, simply because of the sheer abundance of textual data. Natural language in the form of physical data is highly expressive, but it lacks the structure required for mathematical physics. Conversely, the mathematics community is increasingly turning to Interactive Theorem Provers (ITPs) such as Lean 7 to formalize proofs. Recent efforts suggest that Lean could also be adapted to formalize certain domains of physics 127 . However, traditional ITPs may only be sufficient for Category A and B discoveries â deriving predictions and phenomenological laws from established frameworks. i.e., Lean is designed to prove that a conclusion follows from fixed axioms, but it is not designed to realize that the axioms themselves need to be replaced. Therefore, bringing formalization to Category C discovery will require a fundamentally new type of formal system: a âReverse ITPâ. Instead of navigating from fixed axioms to a conclusion, this sys- tem should formalize the abductive process itself as an organized search over provisional principles. A generative AI system would propose âtemporary axiomsâ â bold physical ansĂ€tze, symmetries, or other mathematical structures â and use symbolic tools to deduce their falsifiable consequences. It would then test these consequences against data or consistency constraints, modifying parameters or the provisional axioms themselves when contradictions arise. In this sense, theoretical contra- dictions become analogous to a loss signal for refining the provisional axioms of the candidate framework. The output would not be a proof alone, as in conventional ITPs, but a set of candi- date frameworks, each with provisional axioms and derived consequences that can be automatically tested and falsified. Human scientists communicate ideas through representational languages that trade precision for expressive range. At one end is formal mathematics: equations, definitions, and theorems can be transmitted with very high precision, but only after an idea has been cast into a sufficiently rigid form. Natural language occupies a middle ground: it can express hypotheses, analogies, intuitions, and incomplete arguments, but at the cost of increased ambiguity. At the other end are artistic modes of communication, such as images, music, and metaphor, which can convey broader forms of intuition and experience, but with even less determinate meaning. Physics has historically relied on all parts of this spectrum. Although its final results are often expressed mathematically, the reasoning that leads to them often involves âcontrolled vaguenessâ: diagrams, analogies, ansĂ€tze, scaling arguments, effective descriptions, and provisional concepts whose meaning becomes sharper only through use. 11 This line of thinking suggests that a Lean-like formal system may be too rigid for physics, especially in the early stages of theory formation. Physics may need a representational medium that is perhaps just a bit more flexible than formal mathematics but clearly more structured than ordinary natural language or art. The relevant goal may therefore be a semi-formal language of physics discovery: precise enough to support symbolic manipulation, consistency checks, and derivation of falsifiable consequences, yet flexible enough to represent approximate principles, incomplete analogies, emer- gent variables, and candidate mathematical structures before they have been fully formalized. Such a language may need to preserve something art-like: the ability to communicate rich, partially formed intuitions before they have been compressed into equations. If AI is to contribute to paradigm-shifting physics, we should help teach it the mode of discovery that shaped modern physics from Einstein to Higgs (and since): not just fitting patterns, but constructing principle-constrained frameworks and deducing their falsifiable consequences. Only then will we finally know whether AI can actually discover the next Standard Model, or merely keep predicting around it. 12 REFERENCES 1. Wang, H. et al. Scientific discovery in the age of artificial intelligence. Nature 620, 47â60 (2023). 2. Kramer, S., Cerrato, M., Brugger, J., DĆŸeroski, S. & King, R. D. Automated Scientific Dis- covery: From Equation Discovery to Autonomous Discovery Systems. Machine Learning 115, 109 (2026). 3. Udrescu, S.-M. & Tegmark, M. AI Feynman: A physics-inspired method for symbolic regres- sion. Science Advances 6 (2020). 4. Kitano, H. Artificial intelligence to win the Nobel Prize and beyond: Creating the engine of scientific discovery. AI magazine 37, 39â49 (2016). 5. Carleo, G. et al. Machine learning and the physical sciences. Reviews of Modern Physics 91, 045002 (2019). 6. Appel, K. & Haken, W. Solution of the four color map problem. Scientific American 237, 108â121 (1977). 7. De Moura, L., Kong, S., Avigad, J., Van Doorn, F. & von Raumer, J. The Lean theorem prover (system description) in International Conference on Automated Deduction (2015), 378â388. 8. Fajtlowicz, S. On Conjectures of Graffiti. Annals of Discrete Mathematics 38, 113â118 (1988). 9. Raayoni, G. et al. Generating conjectures on fundamental constants with the Ramanujan Ma- chine. Nature 590, 67â73 (2021). 10. Elimelech, R. et al. Algorithm-assisted discovery of an intrinsic order among mathematical constants. Proceedings of the National Academy of Sciences 121, e2321440121 (2024). 11. Davies, A. et al. Advancing mathematics by guiding human intuition with AI. Nature 600, 70â74 (2021). 12. Fawzi, A. et al. Discovering faster matrix multiplication algorithms with reinforcement learn- ing. Nature 610, 47â53 (2022). 13. Belis, V., Odagiu, P. & Aarrestad, T. K. Machine learning for anomaly detection in particle physics. Reviews in Physics 12, 100091 (2024). 14. Langley, P. BACON: A production system that discovers empirical laws in Proceedings of the 5th International Joint Conference on Artificial Intelligence (1977), 344. 15. Langley, P., Simon, H. A., Bradshaw, G. L. & Zytkow, J. M. Scientific Discovery: Computa- tional Explorations of the Creative Process (MIT Press, 1987). 16. Schmidt, M. & Lipson, H. Distilling Free-Form Natural Laws from Experimental Data. Sci- ence 324, 81â85 (2009). 17. Brunton, S. L., Proctor, J. L. & Kutz, J. N. Discovering governing equations from data by sparse identification of nonlinear dynamical systems. Proceedings of the National Academy of Sciences 113, 3932â3937 (2016). 18. Li, Q. et al. Advancing symbolic regression for earth science with a focus on evapotranspira- tion modeling. npj Climate and Atmospheric Science 7, 275 (2024). 19. Cranmer, M. et al. Discovering Symbolic Models from Deep Learning with Inductive Biases. Advances in Neural Information Processing Systems 33, 17429â17442 (2020). 20. Langley, P. The Computational Gauntlet of Human-Like Learning. Proceedings of the 36th AAAI Conference on Artificial Intelligence (2022). 21. Senior, A. W., Evans, R., Jumper, J., et al. Improved protein structure prediction using poten- tials from deep learning. Nature 577, 706â710 (2020). 22. Jumper, J. et al. Highly Accurate Protein Structure Prediction with AlphaFold. Nature 596, 583â589 (2021). 23. Abramson, J., Adler, J., Dunger, J., et al. Accurate structure prediction of biomolecular inter- actions with AlphaFold 3. Nature 630, 493â500 (2024). 24. Lam, R. et al. Learning skillful medium-range global weather forecasting. Science 382, 1416â 1421 (2023). 25. Raissi, M., Perdikaris, P. & Karniadakis, G. E. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics 378, 686â707 (2019). 13 26. Raissi, M., Yazdani, A. & Karniadakis, G. E. Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations. Science 367, 1026â1030 (2020). 27. Peurifoy, J. et al. Nanophotonic particle simulation and inverse design using artificial neural networks. Science Advances 4, eaar4206 (2018). 28. Pestourie, R., Mroueh, Y., Nguyen, T. V., Das, P. & Johnson, S. G. Active learning of deep surrogates for PDEs: application to metasurface design. npj Computational Materials 6, 1â7 (2020). 29. Jiang, R. & Willett, R. Embed and emulate: learning to estimate parameters of dynamical systems with uncertainty quantification in Proceedings of the 36th International Conference on Neural Information Processing Systems (Curran Associates Inc., 2022). 30. Elorza Casas, C. A., Ricardez-Sandoval, L. A. & Pulsipher, J. L. A comparison of strategies to embed physics-informed neural networks in nonlinear model predictive control formulations solved via direct transcription. Computers and Chemical Engineering 198, 109105 (2025). 31. Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. Ad- vances in Neural Information Processing Systems 35, 24824â24837 (2022). 32. Boiko, D. A., MacKnight, R. & Gomes, G. Emergent autonomous scientific research capa- bilities of large language models arXiv: 2304.05332 [physics.chem-ph] (2023). 33. Bran, A. M. et al. Augmenting large language models with chemistry tools. Nature Machine Intelligence 6, 525â535 (2024). 34. Xu, F. et al. Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models arXiv: 2501.09686 [cs.AI] (2025). 35. Lewkowycz, A. et al. Solving Quantitative Reasoning Problems with Language Models in Advances in Neural Information Processing Systems (2022). 36. Zheng, K., Han, J. M. & Polu, S. miniF2F: a cross-system benchmark for formal Olympiad- level mathematics in International Conference on Learning Representations (2022). 37. Trinh, T. H., Wu, Y., Le, Q. V., He, H. & Luong, T. Solving olympiad geometry without human demonstrations. Nature 625, 476â482 (2024). 38. Raz, T. et al. From Euler to AI: Unifying Formulas for Mathematical Constants in Advances in Neural Information Processing Systems 38 (Curran Associates, Inc., 2025), 123959â124040. 39. Polu, S. & Sutskever, I. Generative Language Modeling for Automated Theorem Proving arXiv: 2009.03393 [cs.LG] (2020). 40. Reddy, C. K. & Shojaee, P. Towards scientific discovery with generative AI: progress, op- portunities, and challenges in Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelli- gence and Fifteenth Symposium on Educational Advances in Artificial Intelligence (AAAI Press, 2025). 41. Sothanaphan, N. Resolution of Erdos Problem 728: a writeup of Aristotleâs Lean proof arXiv: 2601.07421 [math.NT] (2026). 42. Chen, E. et al. Felâs Conjecture on Syzygies of Numerical Semigroups arXiv: 2602.03716 [math.CO] (2026). 43. Ma, J. & Tang, Q. An Erdos problem on random subset sums in finite abelian groups arXiv: 2602.05768 [math.CO] (2026). 44. Guevara, A., Lupsasca, A., Skinner, D., Strominger, A. & Weil, K. Single-minus gluon tree amplitudes are nonzero arXiv: 2602.12176 [hep-th] (2026). 45. Guevara, A., Lupsasca, A., Skinner, D., Strominger, A. & Weil, K. Single-minus graviton tree amplitudes are nonzero arXiv: 2603.04330 [hep-th] (2026). 46. Shen, Y. et al. Deep learning with coherent nanophotonic circuits. Nature Photonics 11, 441â 446 (2017). 47. Bahri, Y., Kadmon, J., Ganguli, S. & Sohl-Dickstein, J. Statistical mechanics of deep learning. Annual Review of Condensed Matter Physics 11, 501â528 (2020). 48. Toscano, J., Oommen, V., Varghese, A., et al. From PINNs to PIKANs: recent advances in physics-informed machine learning. Machine Learning and Computational Science & Engi- neering 1, 15 (2025). 49. Liu, Z. et al. KAN: Kolmogorov-Arnold Networks in International Conference on Learning Representations (ICLR) (2025). 14 50. Raissi, M., Perdikaris, P., Ahmadi, N. & Karniadakis, G. E. Physics-Informed Neural Net- works and Extensions. arXiv:2408.16806 (2024). 51. Burger, B. et al. A mobile robotic chemist. Nature 583, 237â241 (2020). 52. Boiko, D. A., MacKnight, R., Kline, B. & Gomes, G. Autonomous chemical research with large language models. Nature 624, 570â578 (2023). 53. Karagiorgi, G., Kasieczka, G., Kravitz, S., Nachman, B. & Shih, D. Machine learning in the search for new fundamental physics. Nature Reviews Physics 4, 399â412 (2022). 54. Thiyagalingam, J., Shankar, M., Fox, G. & Hey, T. Scientific machine learning benchmarks. Nature Reviews Physics 4, 413â420 (2022). 55. Krenn, M., Pollice, R., HĂ€se, F., Friederich, P. & Aspuru-Guzik, A. On scientific understand- ing with artificial intelligence. Nature Reviews Physics 4, 761â769 (2022). 56. Gentine, P., Eyring, V. & Beucler, T. Machine learning for the physics of climate. Nature Reviews Physics 6, 249â251 (2024). 57. Jiao, L. et al. AI meets physics: A comprehensive survey. Artificial Intelligence Review 57, 1â70 (2024). 58. Makke, N. & Chawla, S. A Perspective on Symbolic Machine Learning in Physical Sciences in NeurIPS 2024 Workshop on Machine Learning and the Physical Sciences (2024). 59. Wetzel, S. J., Ha, S., Iten, R., Klopotek, M. & Liu, Z. Interpretable Machine Learning in Physics: A Review arXiv: 2503.23616 [physics.comp-ph] (2025). 60. Ray, S. Generative Metascience: A Review of AI as the Next Scientific Instrument and the Emerging Paradigm of Algorithmic Discovery. MetaScientia: Journal of the History and Phi- losophy of Science 1, 249â285 (2025). 61. Newton, I. Philosophiae naturalis principia mathematica (Jussu Societatis Regiae ac Typis Josephi Streater, 1687). 62. Merchant, A. et al. Scaling deep learning for materials discovery. Nature 624, 80â85 (2023). 63. Maxwell, J. C. On Physical Lines of Force. Part I.âThe Theory of Molecular Vortices applied to Magnetic Phenomena. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science. 4 21, 161â175 (1861). 64. Messeri, L. & Crockett, M. J. Artificial intelligence and illusions of understanding in scientific research. Nature 627, 49â58 (2024). 65. Langley, P. Integrated Systems for Computational Scientific Discovery in Proceedings of the AAAI Conference on Artificial Intelligence (2024). 66. Cai, S., Mao, Z., Wang, Z., et al. Physics-informed neural networks (PINNs) for fluid me- chanics: a review. Acta Mechanica Sinica 37, 1727â1738 (2021). 67. Karniadakis, G. E. et al. Physics-informed machine learning. Nature Reviews Physics 3, 422â 440 (2021). 68. Loh, C., Christensen, T., Dangovski, R., Kim, S. & Solja Ë ci Ì c, M. Surrogate- and invariance- boosted contrastive learning for data-scarce applications in science. Nature Communications 13, 4223 (2022). 69. Cranmer, M. et al. Lagrangian Neural Networks arXiv: 2003.04630 [cs.LG] (2020). 70. Xu, J. et al. Neural Robot Dynamics in 9th Annual Conference on Robot Learning (2025). 71. Sanchez-Gonzalez, A. et al. Learning to Simulate Complex Physics with Graph Networks in Proceedings of the 37th International Conference on Machine Learning 119 (PMLR, 2020), 8459â8468. 72. Seiradakis, J. & Edmunds, M. Our current knowledge of the Antikythera Mechanism. Nature Astronomy 2, 35â42 (2018). 73. Sutton, R. The Bitter Lesson (Mar. 2019). w.incompleteideas.net/IncIdeas/BitterLesson.html. 74. Yousefi, M. & Collins, J. Learning the Bitter Lesson: Empirical Evidence from 20 Years of CVPR Proceedings in Proceedings of the 1st Workshop on NLP for Science (NLP4Science) (Association for Computational Linguistics, 2024), 175â187. 75. Bunge, M. A General Black Box Theory. Philosophy of Science 30, 346â358 (1963). 76. Kuhn, T. S. The Structure of Scientific Revolutions (University of Chicago Press, 1962). 77. Chaabouni, R. et al. Emergent Communication at Scale in International Conference on Learning Representations (2022). 15 78. Lee, J., Cho, K. & Kiela, D. Countering Language Drift via Visual Grounding in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (Asso- ciation for Computational Linguistics, 2019), 4385â4395. 79. Hao, S. et al. Training Large Language Models to Reason in a Continuous Latent Space arXiv: 2412.06769 [cs.CL] (2025). 80. Laughlin, R. B. Anomalous Quantum Hall Effect: An Incompressible Quantum Fluid with Fractionally Charged Excitations. Physical Review Letters 50, 1395â1398 (May 1983). 81. Bethe, H. Zur Theorie der Metalle. I. Eigenwerte und Eigenfunktionen der linearen Atom- kette. Zeitschrift fĂŒr Physik 71, 205â226 (1931). 82. Ginzburg, V. L. & Landau, L. D. On the theory of superconductivity. Zh. Eksp. Teor. Fiz. 20, 11 (1950). 83. Gorâkov, L. P. Microscopic Derivation of the Ginzburg-Landau Equations in the Theory of Superconductivity. Soviet Physics JETP 9, 1364â1367 (1959). 84. Parisi, G. Infinite Number of Order Parameters for Spin-Glasses. Phys. Rev. Lett. 43, 1754â 1756 (1979). 85. Si, C., Yang, D. & Hashimoto, T. Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers arXiv: 2409.04109 [cs.CL] (2024). 86. Silver, D. et al. Mastering the game of Go with deep neural networks and tree search. Nature 529, 484â489 (2016). 87. Alon, N. et al. Remarks on the disproof of the unit distance conjecture arXiv: 2605.20695 [math.CO] (2026). 88. Alpöge, L. Counterexample to the Jacobian Conjecture (July 19, 2026). https://x.com/ __alpoge__/status/2079028340955197566. 89. Gallagher, A. An infinite family of counterexamples to the Jacobian Conjecture in dimension three: every generic fiber degree nâ„ 3 occurs (July 2026). 90. Ziv, R. et al. Unsupervised Machine Learning for Experimental Detection of Quantum-Many- Body Phase Transitions arXiv: 2512.01091 [quant-ph] (2025). 91. Regev, N., Shultzman, A., Loignon-Houle, F., Roques-Carmes, C. & Kaminer, I. Neural network inverse design of nanophotonic scintillators in Conference on Lasers and Electro- Optics/Europe (CLEO/Europe 2025) and European Quantum Electronics Conference (EQEC 2025) (Optica Publishing Group, 2025). 92. See for example the skills.md repository https://skills.sh. 93. Zhu, Y., Zabaras, N., Koutsourelakis, P.-S. & Perdikaris, P. Physics-constrained deep learning for high-dimensional surrogate modeling and uncertainty quantification without labeled data. Journal of Computational Physics 394, 56â81 (2019). 94. Sun, L., Gao, H., Pan, S. & Wang, J.-X. Surrogate modeling for fluid flows based on physics- constrained deep learning without simulation data. Computer Methods in Applied Mechanics and Engineering 361, 112732 (2020). 95. Ma, A. & Solja Ë ci Ì c, M. Learning simple heuristic rules for classifying materials based on chemical composition arXiv: 2505.02361 [cond-mat.mtrl-sci] (2025). 96. Kim, S. et al. Integration of Neural Network-Based Symbolic Regression in Deep Learning for Scientific Discovery. IEEE Transactions on Neural Networks and Learning Systems 32, 4166â4177 (2021). 97. Ma, A. et al. Topogivity: A Machine-Learned Chemical Rule for Discovering Topological Materials. Nano Letters 23, 772â778 (2023). 98. Ma, A., Dugan, O. & Solja Ë ci Ì c, M. Predicting band gap from chemical composition: A sim- ple learned model for a material property with atypical statistics arXiv: 2501.02932 [cond-mat.mtrl-sci] (2025). 99. Wigner, E. P. The Unreasonable Effectiveness of Mathematics in the Natural Sciences. Com- munications on Pure and Applied Mathematics 13, 1â14 (1960). 100. Romera-Paredes, B. et al. Mathematical discoveries from program search with large language models. Nature 625, 468â475 (2023). 101. Novikov, A. et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery tech. rep. (Google DeepMind, May 2025). 16 102. Otto, S. E., Zolman, N., Kutz, J. N. & Brunton, S. L. A Unified Framework to Enforce, Dis- cover, and Promote Symmetry in Machine Learning. Journal of Machine Learning Research 26, 1â83 (2025). 103. Xingyu, G. et al. General framework for E(3)-equivariant neural network representation of density functional theory Hamiltonian. Nature Communications 14, 2848 (2023). 104. Cory-Wright, R., Cornelio, C., Dash, S., Fardad, M. & Horesh, L. Evolving scientific discov- ery by unifying data and background knowledge with AI Hilbert. Nature Communications 15, 5922 (2024). 105. Chen, A. J., Yang, J. & Yu, R. Governing Equation Discovery with Relaxed Symmetry Con- straints in Advances in Neural Information Processing Systems (NeurIPS 2025) (2025). 106. AbdusSalam, S., Abel, S. & RomĂŁo, M. C. Symbolic regression for beyond the standard model physics. Phys. Rev. D 111, 015022 (2025). 107. Cai, T. et al. Transforming the bootstrap: using transformers to compute scattering amplitudes in planar N = 4 super YangâMills theory. Machine Learning: Science and Technology 5, 035073 (2024). 108. Cai, T. et al. Recurrent Features of Amplitudes in PlanarN = 4 Super Yang-Mills Theory. Journal of High Energy Physics 2025, 143 (2025). 109. Grzadkowski, B., Iskrzy Ì nski, M., Misiak, M. & Rosiek, J. Dimension-Six Terms in the Stan- dard Model Lagrangian. Journal of High Energy Physics 2010, 85 (2010). 110. Brivio, I. & Trott, M. The Standard Model as an Effective Field Theory. 793, 1â98 (2019). 111. Ellis, J., Murphy, S.-F., Sanz, V. & You, T. Interpreting the LHC Run 2 Higgs, Diboson, and Electroweak Data. Journal of High Energy Physics 2018, 146 (2018). 112. Carmona, A., Lazopoulos, A., Olgoso, P. & Santiago, J. Matchmakereft: automated tree-level and one-loop matching. SciPost Physics 12, 198 (2022). 113. Fuentes-Martin, J., Konig, M., Pages, J., Thomsen, A. E. & Wilsch, F. Matchete: An auto- mated tool for matching effective theories. European Physical Journal C 83 (2023). 114. Lifshitz, A. & Kaminer, I. Emergent Curvature from Flat-Space QFTs: A Computational Search for Quantum Gravity (2026). In preparation. 115. MartĂn-GarcĂa, J. M. et al. xAct: Efficient tensor computer algebra for the Wolfram Language w.xact.es (2002â2026). 116. Peeters, K. Cadabra: a field-theory motivated symbolic computer algebra system. Computer Physics Communications 176, 550â558 (2007). 117. Peeters, K. Cadabra2: computer algebra for field theory revisited. Journal of Open Source Software 3, 1118 (2018). 118. Vermaseren, J. A. M. New features of FORM arXiv: math-ph/0010025 [math-ph] (2000). 119. Kuipers, J., Ueda, T., Vermaseren, J. & Vollinga, J. FORM version 4.0. Computer Physics Communications 184, 1453â1467 (2013). 120. Gaudel, B. et al. SeQuant framework for symbolic and numerical tensor algebra. I. Core capabilities. The Journal of Chemical Physics 164, 142502 (Apr. 14, 2026). 121. Shtabovenko, V., Mertig, R. & Orellana, F. FeynCalc 9.3: New features and improvements. Computer Physics Communications 256, 107478 (2020). 122. Alloul, A., Christensen, N. D., Degrande, C., Duhr, C. & Fuks, B. FeynRules 2.0 - A complete toolbox for tree-level phenomenology. Computer Physics Communications 185, 2250â2300 (2014). 123. LeCun, Y. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Review 62 (2022). 124. Bodnar, C., Bruinsma, W., Lucic, A. & et al. A foundation model for the Earth system. Nature 641, 1180â1187 (2025). 125. OfficeChai Team. A Test Of AGI Could Be If A System Trained Till 1911 Data Could Discover General Relativity: Google DeepMind CEO Demis Hassabis https://officechai.com/ai/a-test- of-agi-could-be-if-a-system-trained-till-1911-data-could-discover-general-relativity-google- deepmind-ceo-demis-hassabis (Feb. 2026). 126. Hla, M. Machina Mirabilis https://michaelhla.com/blog/machina-mirabilis.html (Mar. 2026). 17 127. Breen, B. et al. Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics arXiv: 2510.12787 [cs.AI] (2025). 18