Paper deep dive
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee, Stephen DeWitt, Sagar Dolas, Carrie Eckert, David Elbert, Ian Foster, Tirthankar Ghosal, Anna Giannakou, Tom Gibbs, Leslie Hamilton, Glenn Lockwood, Theresa Mayer, Ben Mintz, Raffi Nazikian, Sal Nimer, Amanda Randles, Woong Shin, Sreenivas Rangan Sukumar, FrĂŠdĂŠric Suter, Mitra Taheri, Michela Taufer, Draguna Vrabie
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/15/2026, 4:36:59 AM
Summary
This paper updates a community roadmap for autonomous science, highlighting a transition from isolated demonstrations to validated discoveries driven by multi-agent systems and self-driving laboratories. It identifies a defining capability-reliability gap where verifying discoveries is harder than generating them. The roadmap elevates trust, verification, reproducibility, safety, security, and governance to first-class dimensions, outlining a two-year horizon focused on interface standardization (MCP, A2A), verification scaffolding, and governance to connect national, international, and commercial platforms through the AISLE grassroots network.
Entities (10)
Relation Signals (9)
Trust, Verification, and Reproducibility â elevatedto â first-class status
confidence 93% ¡ elevating two former cross-cutting concerns, trust, verification, and reproducibility, and safety, security, and governance, to first-class status
Safety, Security, Integrity, and Governance â elevatedto â first-class status
confidence 93% ¡ elevating two former cross-cutting concerns, trust, verification, and reproducibility, and safety, security, and governance, to first-class status
Genesis Mission â centers â autonomous experimentation
confidence 92% ¡ the Genesis Mission has placed autonomous experimentation at the center of U.S. federal science strategy
Capability-Reliability Gap â limits â autonomous science
confidence 91% ¡ this asymmetry, more than raw model capability, is what limits autonomous science today
AISLE â organizes â Self-Driving Laboratories
confidence 90% ¡ proposed the Autonomous Interconnected Science Lab Ecosystem (AISLE), a grassroots network organized around five critical dimensions for connecting them
Multi-agent systems â produce â experimentally validated hypotheses
confidence 89% ¡ Multi-agent systems have produced experimentally validated hypotheses
Agent2Agent Protocol â connects â agents across institutions
confidence 88% ¡ the Agent2Agent (A2A) protocol connects agents to one another across vendor and institutional boundaries
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown more interoperable and orchestrated, reasoning-trained and domain foundation models have raised the capability ceiling, and the Genesis Mission has placed autonomous experimentation at the center of U.S. federal science strategy, with industry emerging as a primary actor. Progress has met a sobering counter-current, including a corrected flagship discovery result, benchmarks showing that agents which rival experts on closed-ended questions still complete only a fraction of open-ended research, and fabricated citations surfacing at leading venues. We read this as the defining tension of the field. Producing a candidate discovery is no longer the hard part, but verifying it is, and this asymmetry now limits autonomous science more than raw model capability. We update the roadmap around seven dimensions, revisiting the original five and elevating two former cross-cutting concerns, trust, verification, and reproducibility, and safety, security, and governance, to first-class status. We assess the original milestones (M1 through M14) as achieved, partially achieved, reframed, or open, add four new milestones (M15 through M18), and scope the path forward to a two-year horizon. The first year concentrates on interfaces, protocol adoption, and the scaffolding of verification, and the second targets federation, zero-trust coordination, and governance. Throughout, we position the grassroots network as the interoperability fabric that lets national programs, international initiatives, and commercial platforms connect rather than re-silo.
Tags
Links
- Source: https://arxiv.org/abs/2607.12113v1
- Canonical: https://arxiv.org/abs/2607.12113v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
88,956 characters extracted from source content.
Expand or collapse full text
See pages - of autonomous_science_cover.pdf Disclaimer. This report describes products of research sponsored by the U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research under Contract No. DE-SCL0000175, âA Testbed for Multi-Agent Autonomous Science: From Lab Bench to Supercomputerâ. This report was prepared as an account of work sponsored by agencies of the United States Government. Neither the United States Government nor any agency thereof, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. License. This report is made available under a Creative Commons Attribution 4.0 International Public license (https://creativecommons.org/licenses/by/4.0). Preferred citation R. Ferreira da Silva, M. Abolhasani, P. Beaucage, L. Biven, M. Bussmann, K. Chard, R. Coffee, S. DeWitt, S. Dolas, C. Eckert, D. Elbert, I. T. Foster, T. Ghosal, A. Giannakou, T. Gibbs, L. Hamilton, G. Lockwood, T. Mayer, B. Mintz, R. Nazikian, S. Nimer, A. Randles, W. Shin, S.R. Sukumar, F. Suter, M. Taheri, M. Taufer, D. Vrabie, âToward Trustworthy Autonomous Science: A Two-Year Community Roadmapâ, Technical Report, ORNL/TM-2026/4663, July 2026. ⏠@techreportautonomousscience2026roadmap, author = Ferreira da Silva, Rafael and Abolhasani, Milad and and Beaucage, Peter and Biven, Laura and Bussmann, Michael and Chard, Kyle and Coffee, Ryan and DeWitt, Stephen and Dolas, Sagar and Eckert, C. and Elbert, David and Foster, Ian T. and Ghosal, Tirthankar and Giannakou, A. and Gibbs, Tom and Hamilton, Leslie and Lockwood, Glenn and Mayer, Theresa and Mintz, Benjamin and Nazikian, Raffi and Nimer, Salahudin and Randles, Amanda and Shin, Woong and Sukumar, Sreenivas Rangan and Suter, Fr\âed\âeric and Taheri, Mitra and Taufer, Michela and Vrabie, Draguna, title = Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap, year = 2026, number = ORNL/TM-2026/4663, institution = Oak Ridge National Laboratory Executive Summary A year ago, we argued that autonomous laboratories remained isolated islands, and we proposed a grassroots network, the Autonomous Interconnected Science Lab Ecosystem (AISLE), organized around five critical dimensions for connecting them. The field has since moved faster than that roadmap anticipated. Multi-agent systems have produced experimentally validated hypotheses; self-driving laboratories have grown markedly more interoperable, increasingly coordinated by shared orchestration software; AI agents, built on reasoning-trained models and scientific foundation models, have grown markedly more capable of multi-step reasoning and tool use; and a national mobilization, the Genesis Mission, has placed autonomous experimentation at the center of U.S. federal science strategy. Most strikingly, industry has become a primary actor, as well-capitalized entrants build robotic discovery laboratories at a pace that public programs cannot match. Progress has encountered a sobering counter-current. A published correction to a flagship autonomous-discovery result retracted its novelty claims and removed a training-data leak; a growing body of benchmarks shows that agents that rival experts on closed-ended questions still complete only a small fraction of open-ended research; and fabricated citations have surfaced even in papers accepted at leading venues. Tellingly, no system has yet made and experimentally self-validated a genuinely novel discovery from end to end. We read this not as a transient growing pain but as the defining tension of the field. Producing a candidate discovery is no longer the hard part. Verifying it is. This asymmetry, more than raw model capability, is what limits autonomous science today. Accordingly, this report organizes the roadmap around seven dimensions, the five we revisit and the two we elevate from cross-cutting afterthoughts to first-class dimensions. ⢠Instrument and cyberinfrastructure integration. Vendor-neutral access to increasingly interoperable self-driving laboratories, with AI inference moving onto the instruments themselves. ⢠Agent-driven data management. Provenance and quality enforced at the point of capture, where scarce data are made AI-ready by construction. ⢠Agent-driven orchestration. Reasoning agents that plan and coordinate specialized scientific methods, and that explain why a result holds. ⢠Interoperable agent interfaces. The consolidating MCP and A2A stack, now joined by an emerging skills layer. ⢠Education and workforce development. The judgment to resist the homogenizing pull of automation. â Trust, verification, and reproducibility. (elevated) Verification and validation across the lifecycle as a first-class requirement. â Safety, security, integrity, and governance. (elevated) Screening and accountability where digital design meets physical execution. For each dimension, we assess the original milestones (M1âM14), classifying each one as achieved, partially achieved, reframed, or open, and we add four new milestones (M15âM18) for the two elevated dimensions. We scope the path forward to a two-year horizon, with the first year concentrating on interfaces, protocol adoption, and the scaffolding of verification, and the second targeting federation, zero-trust coordination, and governance. Throughout, we position the grassroots network as the interoperability fabric that allows national programs, international initiatives, and commercial platforms to connect rather than re-silo. Table of Contents 1 Introduction 2 A Brief State of Autonomous Science 2.1 Shift 1: From Isolated Demonstrations to Validated Discoveries 2.2 Shift 2: Self-Driving Laboratories Enter a Second Generation 2.3 Shift 3: The Interface Layer Begins to Consolidate 2.4 Shift 4: Reasoning and Foundation Models as a New Substrate 2.5 Shift 5: Benchmarks Expose the Capability-Reliability Gap 2.6 Shift 6: Trust and Governance Move to the Foreground 2.7 Shift 7: Industry Becomes a Primary Actor 2.8 Relation to Prior Roadmaps and Surveys 3 Critical Dimensions of the Roadmap 3.1 Instrument and Cyberinfrastructure Integration 3.2 Agent-Driven Data Management 3.3 Agent-Driven Autonomous Orchestration 3.4 Interoperable Agent Interfaces 3.5 Education and Workforce Development 3.6 Trust, Verification, and Reproducibility 3.7 Safety, Security, Integrity, and Governance 4 The National and Global Ecosystem 5 Conclusion and Revised Roadmap References Introduction The scientific discovery process is being reshaped by the convergence of automation, robotics, machine learning (ML), and artificial intelligence (AI). Modern instruments generate data at rates that outpace human analysis, and the cadence of human decision-making is increasingly mismatched with the speed at which experiments can be planned, executed, and interpreted [23]. Autonomous science addresses this mismatch by closing the loop between hypothesis, experiment, and analysis, with the goal of compressing discovery cycles that once took years or decades into months or weeks. One year ago, we argued that this promise was being held back by fragmentation, in that autonomous laboratories operated as isolated islands, unable to communicate across institutional or disciplinary boundaries [23]. To address this, we proposed the Autonomous Interconnected Science Lab Ecosystem (AISLE), a grassroots network organized around five critical dimensions: (1) instrument and cyberinfrastructure integration, (2) agent-driven data management with FAIR compliance [88], (3) agent-driven autonomous orchestration, (4) interoperable agent interfaces, and (5) education and workforce development. For each dimension, we surveyed the state of the art, identified open challenges, and proposed a set of milestones (M1âM14) to guide community efforts. The intervening year has been unusually eventful, and on balance, it has moved faster than the original roadmap anticipated. Multi-agent systems built on large language models (LLMs) have generated hypotheses that were subsequently validated in the laboratory, from SARS-CoV-2 nanobody design [73] to drug-repurposing candidates in oncology and ophthalmology [26, 25]. The interface âplumbingâ that the original roadmap called for has begun to consolidate, with the Model Context Protocol (MCP) [5] and the Agent2Agent (A2A) protocol [41] emerging as complementary standards, and with science-specific middleware layering them onto high-performance computing (HPC) and experimental facilities [36, 58]. Most consequentially for a roadmap aimed at national-scale coordination, the launch of the Genesis Mission [77] has placed robotic laboratories and autonomous experimentation at the center of U.S. federal science strategy, alongside parallel efforts in other regions [22, 2]. Progress has been accompanied by a sobering counter-current. A published correction to a widely cited autonomous materials-discovery result walked back its central novelty claims, re-characterizing ânovelâ compounds as novel to a prediction platform rather than to science, and removing one compound that had leaked from the training data [74]. A growing body of benchmarks shows that agents which approach expert performance on closed-ended scientific questions still complete only a small fraction of open-ended, end-to-end research tasks [11, 16, 78], and fabricated citations have begun to appear even in papers accepted at leading venues [4]. We claim that these are not transient growing pains but the defining tension of the field. We can now generate candidate discoveries faster than we can verify them. Consequently, an updated roadmap cannot treat trust, verification, safety, and governance as cross-cutting concerns folded into other dimensions; they must become first-class dimensions of the roadmap itself. In this report, we present an updated community roadmap for interconnected autonomous science, one year after AISLE, scoped deliberately to a two-year horizon. Given how quickly the field is moving, we favor a two-year roadmap over a longer-range vision, so that the milestones we propose remain concrete and accountable rather than speculative. We group them into targets for the first year and targets for the second. This work makes the following contributions: 1. We characterize how the landscape of autonomous science has changed over the past year, organized around seven shifts that bear directly on the original roadmap (Section 2). 2. We assess progress against the original AISLE milestones (M1âM14), classifying each as achieved, partially achieved, reframed, or open, and we refine the five original dimensions accordingly (Section 3). 3. We elevate two concerns to first-class dimensions of the roadmap, namely trust and verification (Section 3.6), and safety, security, and governance (Section 3.7), and we propose milestones for each. 4. We position the grassroots AISLE network within the new federal and international landscape (Section 4), arguing that top-down mobilization and bottom-up coordination are complementary rather than redundant. Figure 1 shows the resulting structure in two complementary views. Figure 1(a) renders the architecture as a closed discovery loop in which federated AI agents, drawing on a shared data fabric, hypothesize and design, execute on self-driving laboratories and instruments, capture and curate data, and analyze and verify results, with MCP and A2A as the layers that connect agents to tools and to one another, all enclosed by the two dimensions we elevate to first-class, trust and verification on the inside, and safety, security, and governance on the outside, above an education and workforce foundation. Figure 1(b) recasts the same roadmap as a two-year trajectory that climbs the autonomy ladder from tool to analyst to scientist, grouping representative milestones into successive phases so that trustworthy autonomy rises to meet, rather than outrun, frontier model capability. Safety, Security & Governance Trust, Verification & Reproducibility Education & Workforce Development Federated AI Agents ++ Data Fabric Hypothesize & Design Execute on SDLs & Instruments Capture & Curate Data Analyze & Verify iterative discovery loopA2AMCP (a) Interconnected, closed-loop architecture Closing the capabilityâreliability gap over two yearstimeautonomy â model capabilityToolAnalystScientisttrustworthy autonomy0â8 mo8â16 mo16â24 mo â Vendor-agnostic interfaces (M1) â MCP/A2A adoption (M10) â AI metadata (M5) â Federated data mesh (M6) â Verification & validation (M15) â Zero-trust comms (M11) â Self-discovering networks (M12) â Agent identity & governance (M18) â National framework (M4) (b) From assistance to trustworthy autonomy Figure 1: The updated AISLE roadmap. (a) A closed discovery loop, in which federated AI agents over a distributed data fabric hypothesize and design, execute on self-driving laboratories and instruments, capture and curate data, then analyze and verify, is coordinated by two protocol layers, MCP for agent-to-tool access and A2A for agent-to-agent collaboration across institutions, and is enclosed by the two dimensions this paper elevates to first-class, namely trust and verification and safety, security, and governance, over an education and workforce foundation. (b) Across a two-year horizon, the roadmap climbs an autonomy ladder from tool to analyst to scientist, with representative milestones per phase, so that trustworthy autonomy rises to close the gap with frontier model capability. The remainder of this report is structured as follows. Section 2 characterizes the current state of autonomous science. Section 3 develops the seven dimensions of the roadmap, the five we revisit, and the two we elevate. Section 4 situates the roadmap within national and global initiatives. Finally, Section 5 concludes the report and outlines a revised milestone roadmap. A Brief State of Autonomous Science Before revisiting the roadmap, we characterize how the landscape has changed since the creation of the AISLE grassroots network. We organize the discussion around seven shifts and adopt two lenses that the community converged on during the past year. The first is an autonomy ladder that distinguishes degrees of agency. Recent surveys have largely abandoned the binary vision of automated versus not automated in favor of graded taxonomies of autonomy. A representative formulation distinguishes three levels, tool, analyst, and scientist, according to whether the system executes well-defined tasks under direct supervision, conducts analysis within human-set boundaries, or formulates hypotheses and proposes new lines of inquiry on its own [92, 87]. The practical value of such a ladder is that it locates a given system, and a given roadmap milestone, at a specific rung rather than asserting wholesale autonomy. The second is a persistent capability-reliability gap that separates what systems can propose from what can be trusted, that is, the gap between performance on closed-ended scientific questions and performance on open-ended research. We claim that this gap, rather than raw model capability, is the binding constraint on autonomous science today, and we return to it throughout the report. Shift 1: From Isolated Demonstrations to Validated Discoveries A year ago, the strongest claims for autonomous discovery rested on self-driving laboratories optimizing within narrow design spaces. Since then, multi-agent LLM systems have produced hypotheses that were subsequently validated experimentally. A virtual laboratory of AI agents designed SARS-CoV-2 nanobodies that were synthesized and shown to bind variant targets [73], an AI co-scientist generated drug-repurposing and target-discovery hypotheses confirmed in vitro by collaborating laboratories [26], and a multi-agent system proposed a therapeutic candidate for an ophthalmic indication that was validated in follow-up assays [25]. These are genuine advances. In every case, however, the physical experiments were executed by humans, the problems were human-selected, and at least one celebrated result amounted to re-deriving a mechanism that a human laboratory had already established but not yet published. At the time of writing, there is no verified instance of an agent autonomously making and experimentally self-validating a genuinely novel discovery end to end. The best-funded recent efforts each miss a different aspect of autonomous discovery, as the systems that reason most autonomously still rely on humans to run the physical experiments [25], the systems that synthesize most autonomously have had their novelty claims contested [74], and the systems that close the full loop without human intervention do so only in computational settings with no wet-lab validation [90]. We revisit both the agents that produce such results and the verification their claims still demand in Section 3. Shift 2: Self-Driving Laboratories Enter a Second Generation The self-driving laboratory (SDL) community has begun to describe its own trajectory as a move from a first generation of narrow, hand-tuned, poorly interoperable platforms toward what it calls a second generation, or SDL 2.0, that is interoperable, orchestrated, safe, and capable of hypothesis generation [38, 80]. The concrete advance so far is interoperability, with the other properties still largely aspirational. By orchestrated, the community means coordinated by a software layer that sequences instruments, robots, and computational steps into managed, restartable campaigns rather than hand-scripted one-off runs. This shift directly concerns the first three dimensions of the original roadmap, which we revisit in Section 3, where we detail the orchestration frameworks and robotic platforms that make it concrete. Shift 3: The Interface Layer Begins to Consolidate When the original roadmap was written, agents reached tools and one another through point-to-point and proprietary interfaces, and we described the requirements in generic terms. Within a year, two complementary standards have emerged. The Model Context Protocol (MCP) connects a single agent to many tools and data sources, a vertical concern [5], while the Agent2Agent (A2A) protocol connects agents to one another across vendor and institutional boundaries, a horizontal concern [41]. For science specifically, federated-agent middleware and protocol adapters now expose HPC, data-movement, and instrument services to LLM agents through these interfaces [36, 58]. We treat this consolidation as the most concrete update to the original roadmap, even as the pattern by which agents consume these protocols keeps changing, and we develop it in Section 3.4. Shift 4: Reasoning and Foundation Models as a New Substrate The underlying models have also changed. Reasoning-oriented LLMs trained with reinforcement learning now approach expert performance on graduate-level scientific question answering [28], and domain foundation models have produced experimentally corroborated designs in materials and structural biology [91, 1]. These models are the substrate on which orchestration (Section 3.3) is increasingly built. They also sharpen the reliability question because their fluency makes unsupported outputs harder, not easier, to detect. Shift 5: Benchmarks Expose the Capability-Reliability Gap A wave of benchmarks has made the second framing lens quantitative. Agents that perform well on isolated, closed-ended tasks complete only a small fraction of open-ended, end-to-end research tasks [11, 16], and research-grade coding benchmarks curated by scientists remain largely unsolved by frontier models [78]. Reproducibility-focused agent benchmarks report accuracies that, in some settings, fall below random guessing [68]. Deployment evidence points the same way, as the first systematic study of agents in production finds that teams keep them deliberately bounded and human-supervised, with most running only a handful of steps before a human intervenes and the majority still gated by human evaluation, and with reliability named as the top obstacle [59]. The lesson for the roadmap is that progress should be measured against open-ended, verifiable tasks rather than against exam-style proxies that are rapidly saturating. Shift 6: Trust and Governance Move to the Foreground The past year also supplied concrete evidence that trust and governance can no longer be treated as secondary. A published correction to a flagship autonomous-discovery result re-characterized its novelty claims and removed a training-data leak [74], fabricated citations appeared in papers at leading venues [4], analyses found that journal AI-disclosure policies have done little to curb undisclosed AI-assisted writing [31], and the convergence of capable models with remotely accessible laboratories raised dual-use concerns that the existing policy apparatus does not address [79]. These developments motivate the two dimensions we add to the roadmap in Sections 3.6 and 3.7. Shift 7: Industry Becomes a Primary Actor Finally, the past year changed who builds autonomous science. A field that had been driven largely by academic and national-laboratory prototypes acquired well-capitalized industrial entrants, and the change is large enough to bear on every dimension of the roadmap. Startups founded by senior industry researchers and dedicated to autonomous discovery raised sums that dwarf typical academic budgets, with Lila Sciences assembling roughly half a billion dollars to build robotic âAI science factoriesâ [40] and Periodic Labs raising a $300 million seed round to pursue closed-loop materials discovery [3]. The major cloud and chip vendors moved in parallel, shipping agentic research platforms [44] and scientific foundation models [91], and positioning large-scale compute and physics-based simulation as the substrate for autonomous experimentation [54]. Commercial cloud laboratories matured into a credible delivery model for remote, programmatic experimentation, a development we take up in Section 3.1. We read this influx as a double-edged development. It supplies capital, engineering, and infrastructure that the academic community cannot match, yet the distance between the capability these entrants claim and the evidence they have published is itself an instance of the capability-reliability gap, since several of the best-funded autonomous-discovery efforts have released little peer-reviewed validation as of this writing [40, 3]. A roadmap for interconnected science must therefore treat industry as a primary actor while asking of it the same open interfaces, provenance, and verification that it asks of public laboratories. Relation to Prior Roadmaps and Surveys The past year also produced a wave of surveys and roadmaps for autonomous science, including taxonomies of autonomy [92], broad framings of agentic discovery [87], and domain-oriented catalogs of agentic systems [27]. These works are valuable as maps of the literature, and we draw on them for the two lenses above. Our contribution is different in kind. Rather than surveying the field anew, we update a specific, milestone-bearing roadmap [23, 24] against one year of evidence, score its milestones, and revise its structure where the evidence demands. We believe that this form of accountable revision, in which a community commits to milestones and then publicly assesses them, is a useful complement to survey-style framings and one that the field currently lacks. We intend it as a recurring practice rather than a one-off, revisiting and rescoring these milestones as the field advances so that the roadmap stays honest about what has and has not been achieved. Critical Dimensions of the Roadmap We organize the roadmap around seven dimensions. Rather than enumerate them in isolation, Figure 1 situates them within a single closed-loop architecture, where the discovery-loop stages carry orchestration, instruments, and data, the coordinating protocol layers carry the agent interfaces, the enclosing bands carry trust and governance, and the foundation carries education and workforce development. The first five revisit and update the dimensions of the original AISLE roadmap [23] in light of the past year, while the last two, trust and verification (Section 3.6) and safety, security, and governance (Section 3.7), are elevated here from cross-cutting concerns to first-class dimensions. For each dimension, we summarize the state of the art, identify the challenges that remain, state research priorities, and record the status of the associated milestones in the scorecard of Table 1. The subsections that follow justify these entries dimension by dimension, and we interpret the overall pattern in the conclusion (Section 5). Table 1: Status of the original AISLE milestones (M1âM14) and proposed new milestones (M15âM18), each with a target year (Y1 or Y2) on the two-year roadmap. Statuses are provisional and subject to confirmation against the cited evidence. # Milestone (abbreviated) Status By M1 Common instrument interfaces, HAL Partial Y1 M2 End-to-end cross-institution workflows Partial Y2 M3 Open compute fabric, fault tolerance, digital twins Open Y2 M4 Scalable national instrument framework Open Y2 M5 AI-driven metadata and annotation Partial Y1 M6 Federated data mesh, FAIR governance Open Y2 M7 Near-real-time processing and AI provenance Partial Y2 M8 Hierarchical LLM orchestration, verification Partial Y1 M9 Cross-facility knowledge integration Open Y2 M10 Standardized cross-vendor agent interfaces Reframed Y1 M11 Zero-trust, sub-second agent coordination Open Y2 M12 Self-discovering agent networks Partial Y2 M13 National education consortium Open Y1 M14 Virtual labs, human-AI assessment Open Y2 M15 Verification and validation across the lifecycle New Y1 M16 Reproducibility and efficiency benchmark New Y1 M17 Screening at digital-physical interface New Y1 M18 Cross-institution agent identity, governance New Y2 Instrument and Cyberinfrastructure Integration The first dimension concerns how autonomous agents orchestrate diverse experimental equipment and computational resources across institutional boundaries. Increasingly, the two are inseparable, as instruments, edge devices, HPC, cloud, and digital twins form a single distributed cyber-physical system in which AI models, simulation, and laboratory automation execute as one coordinated workflow whose reliability determines the trustworthiness of the result. Over the past year, the defining development has been the communityâs move toward the second-generation self-driving laboratories (SDLs) described in Section 2, more interoperable and orchestrated than the narrow, hand-tuned systems that dominated earlier demonstrations [38, 80]. Brief state of the art. Orchestration software has matured from bespoke scripts toward reusable frameworks, such as MADSci for modular discovery campaigns [8] and ChemOS 2.0 for chemical SDLs [70], while mobile and multi-robot platforms now approach human throughput on real synthesis tasks [12]. Beneath the agent layer, device-control standards such as SiLA 2 provide vendor-neutral, typed instrument communication [69], and standards bodies have begun to target the interfaces that a modular autonomous laboratory requires [47]. At the facility scale, programs that connect instruments, robotic laboratories, and HPC into shared workflows have continued to grow [55]. A complementary commercial route to instrument access has matured in parallel, namely cloud laboratories that expose physical instruments to remote programmatic control, so that an experiment specified as code is executed by robots and technicians at a central facility [20]. This model has reached academia through the first university cloud lab [14], and it has already been driven by a language-model agent that learned the facilityâs scripting language from documentation and carried out cross-coupling reactions end to end [10]. National programs have begun to move this model from isolated facilities toward a networked resource, most concretely a federal test-bed for a network of programmable cloud laboratories linked by shared networking and data standards (Section 4) [49]. A further shift concerns where computation happens relative to the instrument. As the data rates of upgraded light sources, particle detectors, and radio arrays began to outpace the store-then-analyze pipeline, reaching exabyte-scale annual volumes at the largest facilities [64, 32], machine-learning inference has moved toward the instrument edge so that detector data are processed as they stream rather than after they are written to storage [18]. Demonstrations include real-time ptychographic reconstruction from a streaming detector on an edge accelerator [9] and machine-learning triggers that decide within microseconds which events to retain [33]. Challenges. The heterogeneity that motivated the original roadmap persists, as most instruments still ship without a standard control interface, retrofitting drivers is labor intensive, and the majority of deployed SDLs operate at the lower rungs of the autonomy ladder of Section 2. The past year added a sharper concern at the instrument level, namely that conclusions drawn from automated characterization can be wrong in ways that human inspection would catch, as the corrected analysis of an autonomous materials laboratory made clear [74]. Organizational barriers compound the technical ones, since intellectual-property and liability questions arise the moment an instrument in one institution is driven by an agent in another. A deeper constraint is that the hardware has advanced more slowly than the software around it. Robotics, sample handling, and reconfigurable experimental setups remain comparatively rigid, and where the apparatus cannot be recomposed under program control, the reach of an agentic workflow is bounded by a largely fixed experiment, which optimization and campaign-driven science can accommodate but open-ended discovery cannot. Research priorities. We reaffirm the need for vendor-agnostic hardware abstraction layers and self-describing instruments that expose their capabilities semantically, and we add the validation of automated characterization as a priority. In effect, this extends the MCP-style capability discovery of the agent interface layer (Section 3.4) down to the instrument and its data-acquisition system, so that an agent can discover and drive a detector as readily as any other software tool. This capability sits above the typed device-control interfaces that standards such as SiLA 2 already provide [69]. We also add modular, agile, and reconfigurable experimental hardware as a priority in its own right, since vendor-neutral software interfaces deliver little if the apparatus beneath them cannot be recomposed. Recent work pursues exactly this for the structural refinement that the corrected materials result called into question, building a rubric-bounded agent for Rietveld analysis of X-ray diffraction whose scoring rewards credible fits and flags out-of-scope cases rather than forcing a pattern fit [67]. Physics-aware digital twins should be used to test autonomous workflows before they touch physical instruments, as in self-driving laboratories that validate a workcell in simulation before any physical run [45], and instrument-level outputs should carry the provenance needed to audit downstream claims (Section 3.6). We further prioritize real-time, edge-side inference and data reduction so that analysis and on-the-fly experiment steering keep pace with instruments whose output exceeds what centralized pipelines can absorb, and so that raw streams are turned into analysis-ready data at the point of acquisition (Section 3.2) [18]. Above the individual instrument, the cyberinfrastructure half of this dimension calls for an open compute fabric that spans edge, HPC, cloud, and heterogeneous accelerators, exposing vendor-neutral interfaces for model execution, accelerator scheduling, workflow portability, and provenance capture, much as MCP and A2A do for tools and agents (Section 3.4), so that workflows move across platforms while hardware and software ecosystems evolve independently beneath them [82]. The status of the original milestones M1 through M4 is summarized in Table 1, and the national framework envisioned by M4 is now partly scaffolded by federal mobilization [77]. Agent-Driven Data Management The second dimension concerns the shift from centralized repositories to distributed systems in which autonomous agents curate, validate, and federate scientific data, enforcing FAIR principles at the point of capture [23, 88]. The past year reframed this dimension around trust. As autonomous campaigns generate data that humans never inspect, provenance and quality assessment become prerequisites for credible discovery rather than after-the-fact bookkeeping. Brief state of the art. Federated data services coordinate movement and processing across many facilities [15], and pass-by-reference systems allow large datasets to be shared without duplication in distributed agent workflows [61]. Provenance models such as PROV-O provide a vocabulary for traceability [37], and standards bodies have begun to target the data and knowledge-management gaps specific to autonomous laboratories [47]. What remains missing is the autonomous enforcement of these capabilities, i.e., agents that curate and annotate data as it is produced rather than leaving it to downstream human stewardship. Challenges. Beyond the format and schema heterogeneity identified previously, four issues have moved to the foreground. First, autonomous agents must negotiate schema evolution when they encounter new experiment types, without manual intervention. Second, data quality is now a verification problem because contaminated or leaked data can propagate silently through AI-driven decision chains, as illustrated by a training-data leak in a flagship autonomous-discovery result [74]. Agents, therefore, require mechanisms to assess reliability from the experimental context, rather than treating all data as equally trustworthy. Third, automation multiplies the volume of raw data, yet the high-quality, well-characterized data needed to train reliable models remains scarce, and how that data is represented can determine whether a model learns the underlying physics at all [89, 29]. Much of what separates reusable data from raw output is well-curated metadata, the theoretical and experimental context, conditions, and provenance that make a measurement findable, interpretable, and trustworthy [88], and capturing it at the moment of acquisition rather than reconstructing it afterward is what keeps data quality tractable as volume grows. Fourth, the move to online, in-transit data reduction (Section 3.1) sharpens the problem because signals that are filtered or summarized at the instrument edge can never be re-examined. The FAIR lifecycle and its provenance must therefore be preserved through reduction, recording not only what was kept but also what was removed and why, lest reproducibility be quietly lost as the data is reduced. Research priorities. We prioritize federated data-mesh architectures in which each laboratory maintains a node with standardized interfaces and global discovery indices, support for both explicit and implicit schemas so that agents can infer structure from heterogeneous sources, and the embedding of provenance frameworks [37] into instrument middleware so that every autonomous decision is traceable across facilities and timescales. We also prioritize making data AI-ready by construction, with readiness levels that track a dataset from raw capture to a form fit for training [13], and community standards for autonomous-laboratory data that play, for machine consumption, the role that FAIR played for sharing [88, 34]. Data curated to this standard are not only an audit trail but also the substrate for surrogate models that compress expensive experiments and carry materials knowledge toward manufacturing timescales [76, 48]. More broadly, as autonomous agents become the dominant consumers of these data, the systems that serve them are better designed to be agent-first from the outset, anticipating the high-volume, exploratory access patterns of agents rather than retrofitting interfaces built for human analysts [42]. The status of milestones M5 through M7 is summarized in Table 1. Agent-Driven Autonomous Orchestration The third dimension is the cognitive core of interconnected autonomous laboratories, namely the agents that navigate scientific decision spaces while remaining aligned with physical and domain knowledge [23]. The substrate for this dimension changed markedly over the past year, as reasoning-oriented LLMs and domain foundation models matured into orchestrators that coordinate specialized methods such as Bayesian optimization, uncertainty quantification, and reinforcement learning. Brief state of the art. Multi-agent systems that generate, debate, and evolve hypotheses have produced experimentally validated results [26], and agentic tree-search systems now carry an idea from conception to a written manuscript [90]. The models beneath them have improved on two fronts. Reasoning-trained LLMs approach expert performance on graduate-level science questions [28], and domain foundation models produce experimentally corroborated designs in materials and structural biology [91, 1]. A complementary line of work casts the LLM as an orchestrator of established tools rather than a replacement for them, for example, by exposing expert-designed chemistry tools to a language model [43]. Challenges. The probabilistic nature of LLM-based agents remains in tension with the determinism that reproducible science assumes, and it is unclear how to guarantee reproducible outcomes or grounding in physical law, when an orchestrator is non-deterministic, higher-latency, and difficult to verify. The capability-reliability gap of Section 2 makes this concrete, since strong performance on closed-ended tasks has not carried over to open-ended research. The orchestrator is therefore best understood as one component of the ecosystem, rather than a replacement for the established methods it coordinates. A recurring response is to keep those methods in the loop, combining data-driven learning with physical and chemical constraints and checking an agentâs proposals against known laws before they reach an instrument [27]. In this view, the probabilistic reasoning of an LLM is an asset for exploration and a liability for commitment, and the design problem is to route each to where it belongs. Research priorities. We reaffirm three thrusts. The first is hierarchical architectures in which LLM-based agents orchestrate established scientific methods abstracted as actuators. The second is a verification and validation infrastructure that uses digital twins and formal or symbolic methods to enforce physics-based constraints as hard boundaries, which we develop in Section 3.6. The third is distributed, real-time knowledge integration that keeps agents grounded across facilities and long campaigns. Grounding of this kind rests on a primitive the original roadmap did not name, namely durable agent memory. We mean policy-bound, provenance-rich, scoped persistence, spanning the working, episodic, semantic, and procedural forms long studied for cognitive agents, that lets an agent carry context, decisions, and lessons across a long campaign instead of rebuilding them inside a prompt, and that is distinct from both the scientific data of Section 3.2 and the model weights beneath it [72]. Early systems realize this idea by managing a small in-context working set against external stores, much as an operating system pages physical memory against a larger virtual address space [57]. Cutting across all three, the construction of an agent is worth treating as a workflow stage in its own right, so that a scientist authors a durable contract, a version-controlled rubric, a graded curriculum, and a curated knowledge base, that bounds an automated builder and survives the churn of the underlying models. This has been demonstrated for the orchestration of a crystallographic refinement tool [67]. We also hold that an orchestrator should yield understanding, not only predictions, since an explanation of why a design or result holds is what makes an autonomous finding actionable, and interpretability of this kind is increasingly expected of scientific machine learning [63, 56]. A further fragility is dependence on the models themselves, since most deployed agents rely on a small set of proprietary frontier models [59], an uneven foundation for reproducible science and one not equally available across regions. Open-weight models accompanied by clear provenance of their training data are therefore of growing interest, because reproducibility requires knowing what a model learned from, and because scientific sovereignty requires not depending on a single vendor. Once agents depend on them throughout a campaign, models are best treated as persistent scientific infrastructure, versioned, evaluated, and retired with the discipline applied to instruments and scientific software rather than swapped silently beneath a running workflow. The status of milestones M8 and M9 is summarized in Table 1. Interoperable Agent Interfaces The fourth dimension concerns the interfaces and standards through which autonomous agents reach the tools, data, and capabilities they depend on, and through which they coordinate with one another across institutional and disciplinary boundaries. This dimension has changed more than any other since the original roadmap, which described its requirements in generic terms because no widely adopted standard yet existed [23]. Within a year, a recognizable two-axis stack has emerged. The Model Context Protocol (MCP) standardizes how a single agent connects to many tools and data sources, while the Agent2Agent (A2A) protocol standardizes how agents discover, authenticate, and delegate to one another across vendor and institutional boundaries [5, 41]. We treat these as complementary, vertical and horizontal respectively, rather than competing [19]. More recently, a third layer has begun to form above these protocols, namely agent skills: composable bundles of instructions, scripts, and resources that an agent discovers and loads on demand. While MCP and A2A govern how agents reach tools and one another, skills capture what an agent needs to know to carry out a task [7]. Brief state of the art. Both protocols moved under neutral governance within months of each other, which lowers the risk of relying on a single vendor for betting infrastructure [41, 19]. For science specifically, federated-agent middleware deploys stateful agents across HPC, data, and experimental resources [36], and thin adapters now expose facility services, such as data movement and remote function execution, to LLM-based agents through MCP [58]. These layers sit above the instrument-control standards of Section 3.1 [69], so that a typed instrument interface can be wrapped for agent access without discarding existing control software. The way agents consume these interfaces is itself in flux. Loading every tool definition into the context window and routing each intermediate result back through the model scales poorly, in both token cost and error surface, once an agent faces dozens of tools. A growing practice instead has the agent write code that calls the tools and discovers their definitions on demand, which repositions the protocol as a substrate beneath code execution rather than a direct call interface [6]. A further step, pursued in industry though likely beyond our two-year horizon for science, is a runtime that surfaces tools and agents dynamically, exposing each step of a workflow only to the capabilities relevant to its goal and folding discovery, scoping, and least-privilege access into the interface layer itself. We read this churn not as the demise of any one protocol but as a reason to standardize the interface while leaving the calling pattern free to evolve, which is the layered stance we take in the priorities below. Challenges. The most-cited blocker is identity and delegated authorization at the facility scale because agents must act on a scientistâs behalf across institutions without holding overly broad, long-lived credentials [58]. A second mismatch is structural, as the request-response tool model fits poorly with the stateful, long-running jobs that characterize scientific campaigns [36, 58]. Semantic interoperability across domains, the provenance of agent-to-agent decisions, and the tension between protocol fragmentation and convergence all remain open. The security posture of these protocols also trails their adoption, as national-security guidance warns that placing tool descriptions, control flow, and data in one shared context blurs trust boundaries, overlapping contexts can leak state across tasks, and unverified dynamic tool discovery widens the reach of a compromised component [52]. Research priorities. We prioritize layered protocol architectures that separate networking, message formatting, semantic interpretation, and coordination, agent identity with attribute-based access control designed for multi-institutional collaboration, provenance-carrying messages that record which agent did what with which tool on which data, and self-describing capability discovery validated on multi-facility testbeds. On identity in particular, the same runtime that scopes which tools a step may see can also scope its authority, issuing signed delegations on demand and limiting their lifetime to a single invocation rather than to the agent as a whole, a pattern commercial systems already implement and one that facility identity-and-access infrastructure will need to grow to support. Recording the prompts, decisions, and tool calls exchanged between agents is a solved problem in agent-observability products, so the open question for science is not whether to log these interactions but what a scientifically adequate record must contain, one that ties each decision to the data, tools, and model versions behind it and serves reproducibility and cross-facility reuse rather than operational monitoring alone. We develop that provenance model, and its use in verification, in Section 3.6 [71]. In the same spirit, and because no widely used skill library yet encodes scientific procedures, we see an opportunity to package validated experimental and analysis workflows as portable, self-describing skills that move across laboratories, which would give capability discovery concrete content to advertise and reuse [7]. The original milestone M10 named specific transport protocols, and we reframe it around the MCP and A2A consolidation. The status of M10 through M12 is summarized in Table 1. Education and Workforce Development The fifth dimension concerns preparing researchers for environments increasingly shaped by autonomous systems, where competencies span AI/ML methods, workflow thinking, human-machine collaboration, and ethical reasoning [23]. The past year added empirical urgency, as a large-scale analysis of AI-engaged researchers found that they publish more and are cited more, even as the range of topics the community collectively studies contracts [30]. Workforce development must therefore cultivate not only fluency with autonomous tools but also the judgment to resist their homogenizing pull. Brief state of the art. Scientific education still largely treats AI/ML as supplementary rather than integral, which leaves competency gaps for autonomous-laboratory settings. National and international programs have begun to fund workforce development as an explicit component of their AI-for-science strategies [77, 22], but systematic curriculum redesign remains limited across institutions. Challenges. The central tension is balancing automation with foundational understanding since over-reliance risks producing scientists who cannot critically evaluate automated results. Compounding this, current assessment methods do not measure the ability to collaborate with AI, many educators lack the relevant expertise, and access to autonomous-laboratory infrastructure for hands-on training is uneven, which raises equity concerns. Research priorities. We prioritize modular curricula that integrate AI/ML competencies with scientific reasoning across disciplines, immersive virtual environments that simulate autonomous laboratories where physical access is limited, and assessment methodologies that evaluate trust calibration and the interpretation of AI decisions, with ethical reasoning integrated throughout. As agentic workflows take over routine tasks and extend what a single researcher can attempt, the design of the human-machine interface becomes a research problem in its own right. Interfaces must let a scientist supervise, interrogate, and override an agent, calibrate trust in its output, and be augmented rather than sidelined by it. The status of milestones M13 and M14 is summarized in Table 1. Trust, Verification, and Reproducibility We elevate trust, verification, and reproducibility from a cross-cutting concern, as it appeared in the original roadmap, to a first-class dimension. The motivation is the capability-reliability gap introduced in Section 2. The ability of agents to propose hypotheses, designs, and analyses has outrun the infrastructure needed to verify them. The past year supplied concrete evidence of the cost of this gap, from the corrected materials-discovery result we examine below [74], to fabricated citations surfacing in papers accepted at leading venues [4], to reproducibility benchmarks on which agents perform near or below chance [68]. We claim that an interconnected ecosystem without commensurate verification infrastructure would amplify, not contain, these failure modes. Brief state of the art. Evaluation has shifted from closed-ended question answering toward reproduction and end-to-end research, with benchmarks that ask agents to reproduce published computational results [68] or to carry a task through the full discovery process [11]. Early verification mechanisms are appearing, including uncertainty quantification, neurosymbolic constraints, and provenance that classifies each claim by its epistemic source, but they are not yet standard components of autonomous workflows. Provenance for agentic workflows, in particular, is beginning to mature, with recent work extending the W3C PROV standard over the Model Context Protocol to capture agent prompts, responses, and decisions as near-real-time, end-to-end workflow provenance, so that an erroneous output can be traced as it propagates from one agent to the next across edge, cloud, and HPC resources [71]. Challenges. Verification in this setting is hard for four reasons. First, the non-determinism of LLM-based orchestration is in tension with run-to-run reproducibility. Second, fluent outputs make hallucinated results and citations harder to detect, not easier. Third, distinguishing genuine novelty from novelty relative to a model or database requires clean train and test separation that current pipelines do not guarantee [74]. The same difficulty recurred for a generative materials model whose experimentally highlighted compound was argued to be isostructural with a phase known for decades and already present in the modelâs training distribution [35], which shows that the problem is not confined to a single laboratory or method. Fourth, verifying long campaigns distributed across facilities and agents is qualitatively harder than checking a single result. A cautionary case. The corrected materials-discovery result illustrates the failure mode in miniature. An autonomous laboratory reported dozens of newly synthesized compounds, and subsequent analysis showed that many were ordered versions of already-known disordered phases, that the automated structural refinement was below the quality a human expert would accept, and that one reported compound had leaked from the training data [74, 75]. None of these errors required a flaw in the robotics. They followed from treating an automated pipelineâs output as a discovery without independent verification. We read this not as an indictment of autonomous laboratories, but as evidence that verification must be designed in from the start, especially once results from one laboratory begin to feed the agents of another. Research priorities. We prioritize verification-and-validation architectures spanning the experiment lifecycle, in which physics-based and logic-based constraints act as hard boundaries rather than soft preferences, end-to-end provenance that links each claim to the data and tool executions that produced it, automated verification of citations and reported results at the point of submission [4], and reproducibility metrics for autonomous workflows that make run-to-run variation measurable. We use the two terms deliberately. Verification asks whether a workflow was built and executed correctly, and validation asks whether its result corresponds to physical reality, the second being the harder problem in the physical sciences, where a simulation is a poor substitute for an experiment and only measurement can settle the question. It is also why the strongest claims of autonomous discovery have so far come from mathematics and computation, where a result can be checked by a machine, rather than from physics, chemistry, or biology, where it cannot. Because an agent can reason soundly from a mistaken picture of the world, a high-consequence action should additionally be checked against independent measurements that the agent does not control. Two further practices sharpen this agenda. First, the natural unit of reproducibility for an agentic workflow is the run itself, an auditable record of the plan, the tool calls, the data and model versions, the human approvals, and the evaluation outcomes that produced a result, rather than the result alone [71]. That record must reach the computational execution as well, what we call AI provenance, capturing model identities and versions, inference parameters, retrieved sources and prompt context, uncertainty estimates, and the accelerators and execution environments behind a result, so that a finding can be reproduced as models, data, and computing environments evolve. Second, evaluation must judge the process and not only the answer, since tool choice, recovery from failure, and respect for budget and policy all bear on whether a result can be trusted [46]. As autonomous laboratories move toward continuous operation, efficiency itself becomes a measure of scientific productivity, so community benchmarks should report not only discovery rate and accuracy but systems-level costs, from energy per validated experiment and compute per optimization cycle to data-movement overhead, end-to-end latency, and the overhead of verification, enabling reproducible comparison across heterogeneous computing environments [65]. We propose two milestones, M15 and M16, summarized in Table 1. Safety, Security, Integrity, and Governance We add safety, security, integrity, and governance as the seventh dimension of the roadmap. The original roadmap noted intellectual-property and liability concerns in passing, and the past year made a broader set of governance questions unavoidable. The convergence of capable models with remotely accessible laboratories makes it easier for non-experts to run sophisticated experiments, which raises dual-use concerns that current biosecurity policy does not address [79, 60]. Human interaction with self-driving labs and their LLM interfaces continues to have unexpected and unintended actions resulting in, for example, disclosure of personal information [62]. In parallel, the integrity of the scholarly record has come under strain, as analyses find that journal disclosure policies have done little to curb undisclosed AI-assisted writing [31], and as community consensus holds that AI systems cannot be accountable authors. Brief state of the art. Technology-and-policy reviews have begun to map the governance landscape for autonomous laboratories [79], work on AI models and agents has started to catalog capabilities of concern [60, 66], and national strategy now frames autonomous experimentation as a security as well as a scientific priority [77]. Publisher norms and disclosure requirements exist, but evidence indicates that they are widely unobserved [31], and screening at the boundary between digital design and physical execution remains largely voluntary. Analyses of commercial cloud laboratories have made the threat concrete, observing that an experiment can be designed on a laptop in one jurisdiction and executed by robots in another, and have proposed know-your-customer screening together with a cloud-lab security consortium modeled on the existing self-regulation of commercial gene synthesis [39]. Challenges. Dual-use risk concentrates where capable agents meet cloud and self-driving laboratories because the same automation that lowers the barrier for legitimate researchers also lowers it for misuse [79, 60]. Security, safety, and ethical concerns can arise at the human interface with agents and self-driving laboratories when there is physical proximity between humans and equipment or dangerous materials; when human-derived data is accessible to agents; or when there is the potential for misalignment of objectives and ethical norms [66, 53]. Accountability is diffuse when agents act across institutions, complicating intellectual property, liability for cross-institutional failures, and the question of who is answerable for an autonomous decision. As campaigns increasingly span public and commercial partners, unresolved questions of data ownership and value capture become a near-term barrier rather than a distant one, and they need to be settled alongside the technical interfaces, not after them [79]. Throughout, governance lags capability, and a networked ecosystem widens the gap by increasing both reach and speed. These are precisely the conditions under which a small number of failures, or a single misuse, can erode public trust in autonomous science as a whole. Research priorities. We prioritize human-in-the-loop guardrails with reliable override, screening, and audit at the digital-to-physical interface for networked SDLs, accountable agent identity and provenance building on Sections 3.4 and 3.6, governance frameworks that preserve institutional autonomy while enabling collaboration, and disclosure-by-construction, in which AI-assisted contributions are recorded automatically rather than self-reported. Trust in a distributed autonomous system also needs roots below the application layer, in trusted execution environments, cryptographically verifiable model and agent identities, signed and immutable workflow and provenance records, and hardware-rooted attestation, which together let one institution rely on anotherâs computation without surrendering control of it. Underlying these safeguards is a simple principle, that the autonomy granted to an agent should match the cost, risk, and reversibility of the action it would take, and should be raised only as auditable evidence of reliability accumulates rather than asserted at the outset [46, 59]. We propose two milestones, M17 and M18, summarized in Table 1. The National and Global Ecosystem The original roadmap argued for a grassroots, bottom-up network on the premise that no coordinated national program existed to connect autonomous laboratories. That premise has partly changed. The launch of the Genesis Mission has mobilized U.S. national laboratories around an integrated platform that couples high-performance computing, scientific foundation models, datasets, and automated laboratory systems, with explicit near-term objectives for robotic laboratories and autonomous experimentation [77]. In the months since launch, that mandate has acquired concrete form, as the program named twenty-six national science and technology challenges, one of which, achieving AI-driven autonomous laboratories, targets the same automated experimentation that the present roadmap addresses [85]. Related efforts, such as the Trillion Parameter Consortium [81], the European strategy for AI in science [22], and the Acceleration Consortium [2], constitute a dense and growing institutional landscape. That landscape now has a substantial private layer as well since the Genesis Mission enlisted two dozen industrial partners, among them chip and cloud vendors and venture-funded autonomous-discovery startups [84], and since those same vendors and startups are building platforms and laboratories of their own (Section 2). Within the United States, the mobilization reaches well beyond the U.S. Department of Energy. The National Science Foundation has established a Directorate for Technology, Innovation, and Partnerships and a milestone-funded X-Labs initiative for breakthrough science, with early topics that include scientific instrumentation [50, 51], the Defense Advanced Research Projects Agency treating AI-driven autonomous experimentation as a national-security capability [17], and the National Institute of Standards and Technology anchoring the measurement science and data standards that autonomous laboratories require, building on the Materials Genome Initiative [34, 48]. A concrete near-term pathway is a DOE call for robotics and automation testbeds for autonomous scientific discovery, framed as reusable community infrastructure for the national laboratories and their industry partners [83]. Closest of all to this roadmapâs own premise, a National Science Foundation test-bed program aims to build a network of programmable cloud laboratories, remotely accessible autonomous facilities to be linked by computational networking and shared data and AI standards [49]. This multi-agency picture strengthens, rather than weakens, the case for a coordinating fabric because each program risks building its own island unless interfaces and standards are held in common. The European setting shows how much this division of labor depends on local conditions. Public compute for large-scale AI is concentrated in a few sites, such as the EuroHPC exascale system JUPITER at JĂźlich and the national centres of the Gauss Centre for Supercomputing, while compute at the laboratories themselves is more modest, and high-speed networking between laboratories and compute sites is not yet in place [21]. Federation across facilities, which much of this roadmap presumes, is therefore gated by infrastructure that is unevenly available. Access to the most widely used proprietary models is likewise not guaranteed in every region, and European alternatives remain sparse, which is one more reason that open-weight models with transparent provenance (Section 3.3) matter beyond reproducibility alone. We include this less as a digression than as a reminder that a roadmap written largely against U.S. mobilization must be read and adapted against the compute, data, and network realities of each region that adopts it. We argue that top-down mobilization and bottom-up coordination are complementary rather than redundant. Large programs supply what a grassroots network cannot, including compute at scale, foundation models, a security framing, and funding at scale [77, 81]. A grassroots network, in turn, supplies what large programs and commercial vendors tend to underweight, namely vendor-neutral interfaces and cross-institutional standards that prevent each new program or proprietary platform from becoming another silo, together with the deliberate inclusion of resource-constrained institutions through portable, low-footprint laboratory modules. In this division of labor, an interconnected network such as AISLE is best understood not as an alternative to federal mobilization but as the interoperability fabric and community-standards layer that allows these initiatives to connect to one another and to the broader research community, rather than to interconnect only internally. We therefore see the roadmap of this paper as a contribution that is independent of but synergistic with the national and global programs now taking shape. A practical consequence is that the milestones we revise below should be read as community commitments that any of these programs can adopt, instrument, and report against, rather than as the agenda of a single institution. As concrete evidence that such adoption is feasible, the Genesis Mission consortium has organized its own work into groups for robotics and automation, data integration and standards, model development and validation, and computing infrastructure, which align respectively with the instrument, data, trust, and orchestration dimensions of Section 3 [86]. Conclusion and Revised Roadmap In this paper, we presented an updated community roadmap for interconnected autonomous science, one year after the original AISLE roadmap [23]. We characterized seven shifts in the landscape, namely validated discoveries, a second generation of self-driving laboratories, a consolidating interface layer, reasoning and foundation models as a new substrate, benchmarks that expose a capability-reliability gap, the move of trust and governance to the foreground, and the entry of industry as a primary actor. We used two lenses, an autonomy ladder and that same gap, to interpret them. We refined the five original dimensions in light of this evidence, and we elevated two concerns, trust and verification (Section 3.6) and safety, security, and governance (Section 3.7), to first-class dimensions of the roadmap. Our milestone-by-milestone assessment is collected in the scorecard of Table 1, presented with the dimensions in Section 3. The pattern is consistent in that the dimensions closest to raw model capability have advanced the fastest, while those that require cross-institutional infrastructure, verification, and governance remain largely open. Orchestration (Section 3.3) and agent interfaces (Section 3.4) moved the furthest; the former carried by reasoning and foundation models and the latter by the consolidation of agent protocols. Data management, education, and the federated aspects of instrument integration moved the least because they depend on coordination that no single model improvement can supply. The two dimensions we add, trust and governance, do not appear on the original scorecard precisely because the past year revealed them to be prerequisites rather than refinements. Read as a two-year plan, the first year concentrates on interfaces, protocol adoption, AI-driven metadata, and the scaffolding of verification, while the second year targets federation, zero-trust coordination, cross-facility knowledge integration, and governance. We deliberately keep the horizon short because, at the current pace of change, a longer-range plan would be obsolete before it could be acted upon. In future work, we plan to convene the community around this two-year roadmap to develop reference implementations that exercise the consolidated interface layer (Section 3.4) across multiple facilities and to prototype the verification, validation, and governance mechanisms that the past year has shown to be prerequisites, rather than refinements, for trustworthy autonomous discovery. References [1] J. Abramson, J. Adler, J. Dunger, R. Evans, T. Green, A. Pritzel, O. Ronneberger, L. Willmore, A. J. Ballard, J. Bambrick, et al. (2024) Accurate structure prediction of biomolecular interactions with alphafold 3. Nature 630 (8016), p. 493â500. Cited by: §2.4, §3.3. [2] Acceleration Consortium (2023) The acceleration consortium. Note: https://acceleration.utoronto.ca/ Cited by: §1, §4. [3] Andreessen Horowitz (2025) Investing in Periodic Labs. Note: https://a16z.com/announcement/investing-in-periodic-labs/ Cited by: §2.7. [4] S. Ansari (2026) Compound deception in elite peer review: a failure mode taxonomy of 100 fabricated citations at neurips 2025. arXiv preprint arXiv:2602.05930. Cited by: §1, §2.6, §3.6, §3.6. [5] Anthropic (2024) Introducing the model context protocol. Note: https://w.anthropic.com/news/model-context-protocol Cited by: §1, §2.3, §3.4. [6] Anthropic (2025) Code execution with MCP: building more efficient agents. Note: https://w.anthropic.com/engineering/code-execution-with-mcp Cited by: §3.4. [7] Anthropic (2025) Equipping agents for the real world with Agent Skills. Note: https://w.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills Cited by: §3.4, §3.4. [8] Argonne AD-SDL (2025) The modular autonomous discovery for science (MADSci) framework. Note: https://github.com/AD-SDL/MADSci Cited by: §3.1. [9] A. V. Babu, T. Zhou, S. Kandel, T. Bicer, Z. Liu, W. Judge, D. J. Ching, Y. Jiang, S. Veseli, S. Henke, et al. (2023) Deep learning at the edge enables real-time streaming ptychographic imaging. Nature Communications 14, p. 7059. External Links: Document Cited by: §3.1. [10] D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes (2023) Autonomous chemical research with large language models. Nature 624 (7992), p. 570â578. External Links: Document Cited by: §3.1. [11] J. Bragg, M. DâArcy, N. Balepur, D. Bareket, B. Dalvi, S. Feldman, D. Haddad, J. D. Hwang, P. Jansen, V. Kishore, et al. (2025) Astabench: rigorous benchmarking of ai agents with a scientific research suite. arXiv preprint arXiv:2510.21652. Cited by: §1, §2.5, §3.6. [12] E. J. Brass, S. Veeramani, Z. Zhou, H. Fakhruldeen, J. S. Manzano, R. Clowes, I. Akpinar, M. R. Ward, J. W. Ward, and A. I. Cooper (2026) A mobile robotic process chemist. Digital Discovery 5 (3), p. 1363â1371. Cited by: §3.1. [13] W. Brewer, P. Widener, V. Anantharaj, F. Wang, T. Beck, A. Shankar, and S. Oral (2025) Data readiness for scientific AI at scale. External Links: 2507.23018 Cited by: §3.2. [14] Carnegie Mellon University (2021) Carnegie mellon to build first-of-its-kind university cloud lab. Note: https://w.cmu.edu/news/stories/archives/2021/august/first-academic-cloud-lab.html Cited by: §3.1. [15] R. Chard, J. Pruyne, K. McKee, J. Bryan, B. Raumann, R. Ananthakrishnan, K. Chard, and I. T. Foster (2023) Globus automation services: research process automation across the spaceâtime continuum. Future Generation Computer Systems 142, p. 393â409. Cited by: §3.2. [16] Z. Chen, S. Chen, Y. Ning, Q. Zhang, B. Wang, B. Yu, Y. Li, Z. Liao, C. Wei, Z. Lu, et al. (2025) Scienceagentbench: toward rigorous assessment of language agents for data-driven scientific discovery. In International Conference on Learning Representations, Vol. 2025, p. 96934â96990. Cited by: §1, §2.5. [17] Defense Advanced Research Projects Agency (2026) Biological technologies office (BTO). Note: https://w.darpa.mil/about/offices/bto Cited by: §4. [18] A. M. Deiana, N. Tran, J. Agar, M. Blott, G. Di Guglielmo, J. Duarte, P. Harris, S. Hauck, M. Liu, M. S. Neubauer, et al. (2022) Applications and techniques for fast machine learning in science. Frontiers in Big Data 5, p. 787421. External Links: Document Cited by: §3.1, §3.1. [19] A. Ehtesham, A. Singh, G. K. Gupta, and S. Kumar (2025) A survey of agent interoperability protocols: MCP, ACP, A2A, and ANP. External Links: 2505.02279 Cited by: §3.4, §3.4. [20] Emerald Cloud Lab (2024) Emerald cloud lab and the Symbolic Lab Language. Note: https://w.emeraldcloudlab.com/ Cited by: §3.1. [21] EuroHPC Joint Undertaking (2025) JUPITER: launching europeâs exascale era. Note: https://w.eurohpc-ju.europa.eu/jupiter-launching-europes-exascale-era-2025-09-05_en Cited by: §4. [22] European Commission (2025) A european strategy for AI in science and the RAISE initiative. Note: https://research-and-innovation.ec.europa.eu/ Cited by: §1, §3.5, §4. [23] R. Ferreira da Silva, M. Abolhasani, D. A. Antonopoulos, L. Biven, R. Coffee, I. T. Foster, L. Hamilton, S. Jha, T. Mayer, B. Mintz, et al. (2025) A grassroots network and community roadmap for interconnected autonomous science laboratories for accelerated discovery. In Workshop Proceedings of the 54th International Conference on Parallel Processing, p. 142â150. Cited by: §1, §1, §2.8, §3.2, §3.3, §3.4, §3.5, §3, §5. [24] R. Ferreira da Silva, R. Moore I, B. Mintz, R. Advincula, A. Alnajjar, L. Baldwin, C. A. Bridges, R. Coffee, E. Deelman, C. Engelmann, et al. (2024) Shaping the future of self-driving autonomous laboratories workshop. Technical report Technical Report ORNL/TM-2024/3714, Oak Ridge National Laboratory. External Links: Document Cited by: §2.8. [25] A. E. Ghareeb, B. Chang, L. Mitchener, A. Yiu, C. J. Szostkiewicz, J. M. Laurent, M. T. Razzak, A. D. White, M. M. Hinks, and S. G. Rodriques (2025) Robin: a multi-agent system for automating scientific discovery. arXiv preprint arXiv:2505.13400. Cited by: §1, §2.1. [26] J. Gottweis, W. Weng, A. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger, K. Rong, R. Tanno, et al. (2025) Towards an AI co-scientist. arXiv preprint arXiv:2502.18864. Cited by: §1, §2.1, §3.3. [27] M. Gridach, J. Nanavati, K. Z. E. Abidine, L. Mendes, and C. Mack (2025) Agentic ai for scientific discovery: a survey of progress, challenges, and future directions. arXiv preprint arXiv:2503.08979. Cited by: §2.8, §3.3. [28] D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, et al. (2025) Deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948. Cited by: §2.4, §3.3. [29] M. Haghighatlari, J. Li, F. Heidar-Zadeh, Y. Liu, X. Guan, and T. Head-Gordon (2020) Learning to make chemical predictions: the interplay of feature representation, data, and machine learning methods. Chem 6 (7), p. 1527â1542. External Links: Document Cited by: §3.2. [30] Q. Hao, F. Xu, Y. Li, and J. Evans (2025) Artificial intelligence tools expand scientistsâ impact but contract scienceâs focus. Nature. External Links: Document Cited by: §3.5. [31] Y. He and Y. Bu (2026) Academic journalsâ ai policies fail to curb the surge in ai-assisted academic writing. Proceedings of the National Academy of Sciences 123 (9), p. e2526734123. Cited by: §2.6, §3.7, §3.7. [32] HEP Software Foundation (2019) A roadmap for HEP software and computing R&D for the 2020s. Computing and Software for Big Science 3, p. 7. External Links: Document Cited by: §3.1. [33] Z. Jiang, B. Carlson, A. Deiana, J. Eastlack, S. Hauck, S. Hsu, R. Narayan, S. Parajuli, D. Yin, and B. Zuo (2024) Machine learning evaluation in the Global Event Processor FPGA for the ATLAS trigger upgrade. Journal of Instrumentation 19, p. P05031. External Links: Document Cited by: §3.1. [34] H. Joress, Z. Trautt, A. McDannald, B. DeCost, A. G. Kusne, and F. Tavazza (2024) Driving U.S. innovation in materials and manufacturing using AI and autonomous labs. Technical report Technical Report NIST Special Publication 1320, National Institute of Standards and Technology. External Links: Document Cited by: §3.2, §4. [35] M. Juelsholt et al. (2026) Continued challenges in high-throughput materials predictions: MatterGen predicts compounds from the training dataset. Materials Horizons. External Links: Document Cited by: §3.6. [36] A. Kamatar, J. G. Pauloski, Y. Babuji, R. Chard, M. Sakarvadia, D. Babnigg, K. Chard, and I. Foster (2025) Empowering scientific workflows with federated agents. arXiv preprint arXiv:2505.05428. Cited by: §1, §2.3, §3.4, §3.4. [37] T. Lebo et al. (2013) PROV-O: the PROV ontology. Note: W3C Recommendation, https://w.w3.org/TR/prov-o Cited by: §3.2, §3.2. [38] H. Lee, H. J. Yoo, H. S. Jang, B. Park, Y. J. Park, and S. S. Han (2026) Toward self-driving laboratory 2.0 for chemistry and materials discovery. Materials Horizons 13 (10), p. 4712â4739. Cited by: §2.2, §3.1. [39] Y. J. Lee and B. Del Castello (2024) Documenting cloud labs and examining how remotely operated automated laboratories could enable bad actors. Technical report Technical Report PE-A3851-1, RAND Corporation. External Links: Document Cited by: §3.7. [40] Lila Sciences (2025) Lila sciences raises $235m series a to advance scientific superintelligence and autonomous laboratories. Note: https://w.lila.ai/ Cited by: §2.7. [41] Linux Foundation (2025) Agent2Agent (A2A) protocol. Note: https://a2a-protocol.org/ Cited by: §1, §2.3, §3.4, §3.4. [42] S. Liu, S. Ponnapalli, S. Shankar, S. Zeighami, A. Zhu, S. Agarwal, R. Chen, S. Suwito, S. Yuan, I. Stoica, M. Zaharia, A. Cheung, N. Crooks, J. E. Gonzalez, and A. G. Parameswaran (2025) Supporting our AI overlords: redesigning data systems to be agent-first. External Links: 2509.00997 Cited by: §3.2. [43] A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller (2024) Augmenting large language models with chemistry tools. Nature machine intelligence 6 (5), p. 525â535. Cited by: §3.3. [44] Microsoft (2025) Transforming R&D with agentic AI: introducing Microsoft Discovery. Note: https://azure.microsoft.com/en-us/blog/transforming-rd-with-agentic-ai-introducing-microsoft-discovery/ Cited by: §2.7. [45] S. K. Moore (2025) This self-driving lab uses a digital twin to speed discovery. Note: IEEE Spectrum, https://spectrum.ieee.org/autonomous-lab-argonne-polybot Cited by: §3.1. [46] National Institute of Standards and Technology (2023) Artificial intelligence risk management framework (AI RMF 1.0). Note: NIST AI 100-1, https://w.nist.gov/itl/ai-risk-management-framework Cited by: §3.6, §3.7. [47] National Institute of Standards and Technology (2025) Development of standards to support a modular and autonomous laboratory ecosystem. Note: https://w.nist.gov/programs-projects/development-standards-support-modular-and-autonomous-laboratory-ecosystem Cited by: §3.1, §3.2. [48] National Science and Technology Council (2021) Materials genome initiative strategic plan. Note: https://w.mgi.gov Cited by: §3.2, §4. [49] National Science Foundation (2025) Test bed: toward a network of programmable cloud laboratories (PCL test bed). Note: NSF 25-541, https://w.nsf.gov/funding/opportunities/pcl-test-bed-test-bed-toward-network-programmable-cloud-laboratories Cited by: §3.1, §4. [50] National Science Foundation (2026) Directorate for technology, innovation and partnerships (TIP). Note: https://w.nsf.gov/tip/about-tip Cited by: §4. [51] National Science Foundation (2026) NSF X-Labs. Note: https://w.nsf.gov/funding/initiatives/nsf-x-labs Cited by: §4. [52] National Security Agency (2026) Model context protocol (MCP). Note: Cybersecurity Information Sheet U/O/6030316-26, https://w.nsa.gov/Portals/75/documents/Cybersecurity/CSI_MCP_SECURITY.pdf Cited by: §3.4. [53] I. C. Ngong, K. Murugesan, S. Kadhe, J. D. Weisz, A. Dhurandhar, and K. N. Ramamurthy (2026) AgentSCOPE: evaluating contextual privacy across agentic workflows. External Links: 2603.04902, Link Cited by: §3.7. [54] NVIDIA (2025) NVIDIA and U.S. government to boost AI infrastructure and r&d investments. Note: https://blogs.nvidia.com/blog/nvidia-us-government-to-boost-ai-infrastructure-and-rd-investments/ Cited by: §2.7. [55] Oak Ridge National Laboratory (2025) INTERSECT: interconnected science ecosystem. Note: https://w.ornl.gov/intersect Cited by: §3.1. [56] F. Oviedo, J. L. Ferres, T. Buonassisi, and K. T. Butler (2022) Interpretable and explainable machine learning for materials science and chemistry. Accounts of Materials Research 3 (6), p. 597â607. External Links: Document Cited by: §3.3. [57] C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez (2023) MemGPT: towards LLMs as operating systems. External Links: 2310.08560 Cited by: §3.3. [58] H. Pan, R. Chard, R. Mello, C. Grams, T. He, A. Brace, O. P. Skelly, W. Engler, H. Holbrook, S. Y. Oh, et al. (2025) Experiences with model context protocol servers for science and high performance computing. arXiv preprint arXiv:2508.18489. Cited by: §1, §2.3, §3.4, §3.4. [59] M. Z. Pan et al. (2025) Measuring agents in production. External Links: 2512.04123 Cited by: §2.5, §3.3, §3.7. [60] J. Pannu, D. Bloomfield, R. MacKnight, M. S. Hanke, A. Zhu, G. Gomes, A. Cicero, and T. V. Inglesby (2025) Dual-use capabilities of concern of biological ai models. PLoS computational biology 21 (5), p. e1012975. Cited by: §3.7, §3.7, §3.7. [61] J. G. Pauloski, K. Rydzy, V. Hayot-Sasson, I. Foster, and K. Chard (2024) Accelerating Python applications with Dask and ProxyStore. arXiv preprint arXiv:2410.12092. Cited by: §3.2. [62] O. G. S. Project (2025) LLM02:2025 â sensitive information disclosure. Note: https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/OWASP Top 10 for LLM Applications Cited by: §3.7. [63] R. Roscher, B. Bohn, M. F. Duarte, and J. Garcke (2020) Explainable machine learning for scientific insights and discoveries. IEEE Access 8, p. 42200â42216. External Links: Document Cited by: §3.3. [64] A. M. M. Scaife (2020) Big telescope, big data: towards exascale with the Square Kilometre Array. Philosophical Transactions of the Royal Society A 378, p. 20190060. External Links: Document Cited by: §3.1. [65] R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni (2020) Green AI. Communications of the ACM 63 (12), p. 54â63. External Links: Document Cited by: §3.6. [66] N. Shapira, C. Wendler, A. Yen, G. Sarti, K. Pal, O. Floody, A. Belfki, A. Loftus, A. R. Jannali, N. Prakash, J. Cui, G. Rogers, J. Brinkmann, C. Rager, A. Zur, M. Ripa, A. Sankaranarayanan, D. Atkinson, R. Gandikota, J. Fiotto-Kaufman, E. Hwang, H. Orgad, P. S. Sahil, N. Taglicht, T. Shabtay, A. Ambus, N. Alon, S. Oron, A. Gordon-Tapiero, Y. Kaplan, V. Shwartz, T. R. Shaham, C. Riedl, R. Mirsky, M. Sap, D. Manheim, T. Ullman, and D. Bau (2026) Agents of chaos. External Links: 2602.20021, Link Cited by: §3.7, §3.7. [67] W. Shin, C. A. Bridges, M. T. McDonnell, and R. Ferreira da Silva (2026) Fantastic scientific agents and how to build them: AgentBuild for Rietveld refinement. arXiv preprint arXiv:2606.12834. Cited by: §3.1, §3.3. [68] Z. S. Siegel et al. (2024) CORE-Bench: fostering the credibility of published research through a computational reproducibility agent benchmark. External Links: 2409.11363 Cited by: §2.5, §3.6, §3.6. [69] SiLA Consortium (2024) SiLA 2: standardization in lab automation. Note: https://sila-standard.com/ Cited by: §3.1, §3.1, §3.4. [70] M. Sim, M. G. Vakili, F. Strieth-Kalthoff, H. Hao, R. J. Hickman, S. Miret, S. Pablo-GarcĂa, and A. Aspuru-Guzik (2024) ChemOS 2.0: an orchestration architecture for chemical self-driving laboratories. Matter 7 (9), p. 2959â2977. Cited by: §3.1. [71] R. Souza, A. Gueroudji, S. DeWitt, D. Rosendo, T. Ghosal, R. Ross, P. Balaprakash, and R. F. Da Silva (2025) PROV-AGENT: unified provenance for tracking ai agent interactions in agentic workflows. In 2025 IEEE International Conference on eScience (eScience), p. 467â473. Cited by: §3.4, §3.6, §3.6. [72] T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths (2024) Cognitive architectures for language agents. Transactions on Machine Learning Research. Note: arXiv:2309.02427 Cited by: §3.3. [73] K. Swanson, W. Wu, N. L. Bulaong, J. E. Pak, and J. Zou (2025) The virtual lab of ai agents designs new sars-cov-2 nanobodies. Nature 646 (8085), p. 716â723. Cited by: §1, §2.1. [74] N. J. Szymanski, B. Rendy, Y. Fei, R. E. Kumar, T. He, D. Milsted, M. J. McDermott, M. Gallant, E. D. Cubuk, A. Merchant, et al. (2026) Author correction: an autonomous laboratory for the accelerated synthesis of inorganic materials. Nature 650 (8100), p. 1â1. Cited by: §1, §2.1, §2.6, §3.1, §3.2, §3.6, §3.6, §3.6. [75] N. J. Szymanski et al. (2023) An autonomous laboratory for the accelerated synthesis of novel materials. Nature 624 (7990), p. 86â91. External Links: Document Cited by: §3.6. [76] D. P. Tabor, L. M. Roch, S. K. Saikin, C. Kreisbeck, D. Sheberla, J. H. Montoya, S. Dwaraknath, M. Aykol, C. Ortiz, H. Tribukait, et al. (2018) Accelerating the discovery of materials for clean energy in the era of smart automation. Nature Reviews Materials 3, p. 5â20. External Links: Document Cited by: §3.2. [77] The White House (2025) Launching the genesis mission. Note: Executive Order 14363, https://w.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/ Cited by: §1, §3.1, §3.5, §3.7, §4, §4. [78] M. Tian, L. Gao, S. D. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y. Li, et al. (2024) Scicode: a research coding benchmark curated by scientists. Advances in Neural Information Processing Systems 37, p. 30624â30650. Cited by: §1, §2.5. [79] A. V. Tobias and A. Wahab (2025) Autonomous âself-drivingâlaboratories: a review of technology and policy implications. Royal Society Open Science 12 (7), p. 250646. Cited by: §2.6, §3.7, §3.7, §3.7. [80] G. Tom, S. P. Schmid, S. G. Baird, Y. Cao, K. Darvish, H. Hao, S. Lo, S. Pablo-GarcĂa, E. M. Rajaonson, M. Skreta, et al. (2024) Self-driving laboratories for chemistry and materials science. Chemical Reviews 124 (16), p. 9633â9732. Cited by: §2.2, §3.1. [81] Trillion Parameter Consortium (2024) Trillion parameter consortium (TPC). Note: https://w.anl.gov/cels/trillion-parameter-consortium Cited by: §4, §4. [82] U.S. Department of Energy, Office of Science (2023) Integrated research infrastructure architecture blueprint activity: final report 2023. Technical report U.S. Department of Energy. External Links: Document Cited by: §3.1. [83] U.S. Department of Energy, Office of Science (2026) Robotics and automation testbeds for autonomous scientific discovery. Note: DOE National Laboratory Announcement LAB 26-3601, https://science.osti.gov/grants/Lab-Announcements/Open Cited by: §4. [84] U.S. Department of Energy (2025) Energy department announces collaboration agreements with 24 organizations to advance the Genesis Mission. Note: https://w.energy.gov/articles/energy-department-announces-collaboration-agreements-24-organizations-advance-genesis Cited by: §4. [85] U.S. Department of Energy (2026) Energy department announces 26 Genesis Mission science and technology challenges. Note: https://w.energy.gov/articles/energy-department-announces-26-genesis-mission-science-and-technology-challenges Cited by: §4. [86] U.S. Department of Energy (2026) Genesis Mission collaboration. Note: https://w.energy.gov/undersecretaryforscience/genesis-mission/genesis-mission-collaboration Cited by: §4. [87] J. Wei, Y. Yang, X. Zhang, Y. Chen, X. Zhuang, Z. Gao, D. Zhou, G. Wang, Z. Gao, J. Cao, et al. (2025) From ai for science to agentic science: a survey on autonomous scientific discovery. arXiv preprint arXiv:2508.14111. Cited by: §2.8, §2. [88] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J. Boiten, L. B. da Silva Santos, P. E. Bourne, et al. (2016) The FAIR guiding principles for scientific data management and stewardship. Scientific data 3 (1), p. 1â9. Cited by: §1, §3.2, §3.2, §3.2. [89] P. Xu, X. Ji, M. Li, and W. Lu (2023) Small data machine learning in materials science. npj Computational Materials 9, p. 42. External Links: Document Cited by: §3.2. [90] Y. Yamada, R. T. Lange, C. Lu, S. Hu, C. Lu, J. Foerster, J. Clune, and D. Ha (2025) The ai scientist-v2: workshop-level automated scientific discovery via agentic tree search. arXiv preprint arXiv:2504.08066. Cited by: §2.1, §3.3. [91] C. Zeni, R. Pinsler, D. ZĂźgner, A. Fowler, M. Horton, X. Fu, Z. Wang, A. Shysheya, J. CrabbĂŠ, S. Ueda, et al. (2025) A generative model for inorganic materials design. Nature 639 (8055), p. 624â632. Cited by: §2.4, §2.7, §3.3. [92] T. Zheng, Z. Deng, H. T. Tsang, W. Wang, J. Bai, Z. Wang, and Y. Song (2025) From automation to autonomy: a survey on large language models in scientific discovery. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 17744â17761. Cited by: §2.8, §2.