Paper deep dive
Beyond the Desk: Barriers and Future Opportunities for AI to Assist Scientists in Embodied Physical Tasks
Irene Hou, Alexander Qin, Lauren Cheng, Philip J. Guo
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/23/2026, 12:05:29 PM
Summary
This study investigates the integration of AI into 'beyond the desk' scientific work, such as lab and field research. Through interviews with 12 scientific practitioners, the authors identify three primary barriers to AI adoption: high-stakes experimental risks, constrained physical environments, and the inability of AI to replicate human tacit knowledge. The paper proposes speculative designs for AI assistants that function as background infrastructure to support, rather than replace, human expertise.
Entities (5)
Relation Signals (3)
Scientific Practitioners â identifiedbarriers â AI Adoption
confidence 95% ¡ found three barriers to AI adoption in these settings
Scientific Practitioners â performed â Embodied Physical Tasks
confidence 95% ¡ scientific practitioners doing hands-on lab and fieldwork
AI â supportedby â Speculative Design
confidence 90% ¡ Participants then developed speculative designs for future AI assistants
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:More scientists are now using AI, but prior studies have examined only how they use it 'at the desk' for computer-based work. However, given that scientific work often happens 'beyond the desk' at lab and field sites, we conducted the first study of how scientific practitioners use AI for embodied physical tasks. We interviewed 12 scientific practitioners doing hands-on lab and fieldwork in domains like nuclear fusion, primate cognition, and biochemistry, and found three barriers to AI adoption in these settings: 1) experimental setups are too high-stakes to risk AI errors, 2) constrained environments make it hard to use AI, and 3) AI cannot match the tacit knowledge of humans. Participants then developed speculative designs for future AI assistants to 1) monitor task status, 2) organize lab-wide knowledge, 3) monitor scientists' health, 4) do field scouting, 5) do hands-on chores. Our findings point toward AI as background infrastructure to support physical work rather than replacing human expertise.
Tags
Links
- Source: https://arxiv.org/abs/2603.19504v1
- Canonical: https://arxiv.org/abs/2603.19504v1
Trouble viewing inline? Open PDF directly â
Full Text
96,534 characters extracted from source content.
Expand or collapse full text
Beyond the Desk: Barriers and Future Opportunities for AI to Assist Scientists in Embodied Physical Tasks Irene Hou UC San DiegoLa Jolla, CAUSA ihou@ucsd.edu 0009-0008-0511-7685 , Alexander Qin UC San DiegoLa Jolla, CAUSA aqin@ucsd.edu 0009-0009-5624-5625 , Lauren Cheng UC San DiegoLa Jolla, CAUSA lacheng@ucsd.edu 0009-0007-6488-7500 and Philip J. Guo UC San DiegoLa Jolla, CAUSA pg@ucsd.edu 0000-0002-4579-5754 (2026) Abstract. More scientists are now using AI, but prior studies have examined only how they use it âat the deskâ for computer-based work. However, given that scientific work often happens âbeyond the deskâ at lab and field sites, we conducted the first study of how scientific practitioners use AI for embodied physical tasks. We interviewed 12 scientific practitioners doing hands-on lab and fieldwork in domains like nuclear fusion, primate cognition, and biochemistry, and found three barriers to AI adoption in these settings: 1) experimental setups are too high-stakes to risk AI errors, 2) constrained environments make it hard to use AI, and 3) AI cannot match the tacit knowledge of humans. Participants then developed speculative designs for future AI assistants to 1) monitor task status, 2) organize lab-wide knowledge, 3) monitor scientistsâ health, 4) do field scouting, 5) do hands-on chores. Our findings point toward AI as background infrastructure to support physical work rather than replacing human expertise. scientific practice, embodied physical work, AI, speculative design Accepted to CHI 2026. This is the authorsâ preprint version. â isbn: 978-1-4503-X-X/26/02â journalyear: 2026â conference: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems; April 13â17, 2026; Barcelona, Spainâ booktitle: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI â26), April 13â17, 2026, Barcelona, Spainâ copyright: noneâ ccs: Human-centered computing Empirical studies in HCI Figure 1. Via interviews with 12 scientific practitioners working in lab and field settings, we found three barriers to adopting AI for embodied physical tasks: (A) Neuroscientists like P1 have experimental setups (e.g., surgeries to implant sensors in rodent brains) that are too high-stakes to risk AI making errors. (B) The challenging environments of lab and field sites, like a biochemistâs clean-room lab bench with delicate equipment (P12), make it hard to access AI tools. (C) Scientists feel AI cannot match the tacit knowledge and contextual judgment of human experts like a field roboticist improvising materials to get their experimental rig working at a remote volcano site (P2). (Image credit: all illustrations were drawn by human artist Lauren Cheng without AI assistance.) Three illustrations labeled (A), (B), and (C) depict scientists working in a lab or in a field setting. Image (A) shows two scientists sitting at a table performing rodent brain surgery, with one appearing exhausted as the clock shows 8:00pm. In (B), a cellular biologist stands at a lab bench preparing reagents with a test tube in her left hand and a pipette in her right hand. She appears confused as she relies on scattered sticky notes and analog protocols. In (C), a scientist working at remote terrain near a volcano is stuck fixing an unresponsive robot with limited tools and no support. He holds the broken robot leg in his right hand as the quadrupedal robot lies next to him with tangled wires. 1. Introduction Figure 2. To find barriers and opportunities for AI adoption, we visited workplaces and performed on-site interviews with 12 scientific practitioners in a range of lab and field settings such as (A) behavioral neuroscience labs with surgery and bench areas (P1, P5), (B) field site for a primate behavioral scientist (P8), (C) cellular biochemistry lab (P12). Four images of various scientific settings are labeled with (A), (B), (C), and (D). Image (A) shows a behavioral neuroscience lab bench with a microscope, tools, and scattered notes for soldering a neurobiology implant device. In (B), multiple scientists in cleanroom suits stand around computers and machinery in a nuclear fusion laboratory. In (C), a macaque in a dry landscape interacts with an apparatus made out of acrylic panels. Image (D) provides a closer look at a lab bench, featuring many reagent bottles and sticky notes on the shelves and a notebook on the table top. As AI continues to advance, global investment toward AI for science and the automation of knowledge discovery have led to widely publicized, high-profile breakthroughs. In 2024, teams were awarded the Nobel Prize in Chemistry for AlphaFold, an AI system that uses deep learning to predict 3D protein structures (Jumper et al., 2021), along with the Nobel Prize in Physics for the theoretical foundations that led to deep neural networks (44). Other fields have seen similar innovations, from drug discovery (Wallach et al., 2015; Medicine, 2024) to climate prediction (Lam et al., 2023). Tech companies have also recently invested in using AI for science, such as Googleâs AI Co-Scientist to aid in hypothesis deliberation, literature review, and data synthesis (Gottweis et al., 2025). However, most of these initiatives have centered on computation, simulation, data analysis, and other forms of knowledge work that take place âat the desk.â While AI systems are promising for the computational and simulation-heavy areas of science, this is only one aspect of scientific practice. Decades of STS (Science and Technology Studies) and HCI research have shown that discovery frequently involves material interaction (Latour and Woolgar, 1986) and embodied physical labor in labs and field sites (Cetina, 1999). A lot of critical scientific labor happens âbeyond the deskâ in high-stakes and physically-demanding environments. For instance, a nuclear physicist must don a full-body cleanroom suit to enter a high-security lab space where they fabricate delicate fuel capsules, a task where fine-grained physical coordination is vital and a momentâs lapse of attention can derail months of hard work and put others in danger. A biochemist similarly labors under intense timing constraints, running between incubators, centrifuges, and a series of sterile workspaces to keep cell cultures alive while following a choreographed experimental protocol. A field scientist hikes through the forests of the Republic of Congo to study chimpanzee behavior, improvising tools on-the-go in the face of unpredictable terrain and limited infrastructure. These are all real stories from the 12 scientific practitioners111In this paper, we use the term scientific practitioners to specifically refer to people engaged in the day-to-day practices that produce scientific knowledge, including experimental setup, data collection, troubleshooting, interpretation, and the embodied labor of laboratory and field environments. Although workers of all ages and experience levels can engage in such practitioner labor, oftentimes much of this labor is taken on by junior lab staff (e.g., graduate students, postdocs, lab technicians) while senior scientific staff (e.g., professors, industry lab PIs) may focus more on âat the deskâ knowledge work such as writing grant proposals, framing research papers, and giving talks to academic and broader public audiences. we interviewed for this paper (see Figures 1 and 2), and they illustrate how scientific practice across a range of domains can require physical presence, quick judgment, and expert hand-eye coordination. This sort of âbeyond the deskâ labor currently exceeds what can be codified or simulated on the computer since it involves intimate contact with the physical world. Despite âbeyond the deskâ science contributing critical findings to research and innovation, little is known about how contemporary AI tools fit into these workflows. For instance, HCI research in AI often focuses on its capabilities for screen-bound knowledge workflows such as programming, data science, creative writing, or UI/UX/visual design (Guo et al., 2025; Laban et al., 2024; Vaithilingam et al., 2022; Zheng et al., 2025; Liu et al., 2023; Shi et al., 2023; Suh et al., 2024). However, scientific work is often constrained by physical realities and unpredictable environments, involves improvised non-repeatable procedures, and is shaped by the multi-sensory tacit knowledge222Tacit knowledge is hands-on, domain-specific knowledge that is hard to precisely articulate in words, and thus rarely written down (Polanyi, 1966). In scientific settings this may involve a nuclear physicist knowing how to finely manipulate a fuel capsule âby feelâ using their sensory intuition. By definition this unwritten knowledge cannot be in the training sets for text-based AI systems like LLMs. One can imagine training AI with video data, but even those cannot capture senses like how some fuel capsules âfeelâ right or wrong when manipulated with precision handheld tools, and how to adjust on-the-fly. of domain experts. Moreover, the stakes of error are high, ranging from wasted months of research effort and materials to putting humans in danger. These conditions may pose unique barriers to AI adoption, which led us to the following question that, to our knowledge, we are the first to raise: How do scientific practitioners who work in embodied, improvisational, and high-stakes lab and field environments perceive the relevance, limitations, and future potential of AI assistance? To address this question, we visited lab and field sites of 12 scientific practitioners and conducted in-situ interviews in their workplaces. Our participants worked in laboratory and field environments across biology, neuroscience, animal cognition, nuclear physics, materials science, and field robotics. They had diverse backgrounds, from small university research labs to large industrial organizations, and a range of work experience levels ranging from junior university researchers to senior staff scientists. Participants reflected on their current use of AI and articulated the stakes and limitations of AI assistance in their domains. At the end of each interview, they each engaged in a speculative design activity where they envisioned an imagined âidealâ future AI assistant. Our study revealed three barriers that constrain AI adoption in scientific practice, shown in Figure 1: (A) experimental setups are too high-stakes to risk AI errors, (B) constrained environments make it hard to access AI, and (C) AI cannot match the tacit knowledge (Polanyi, 1966), contextual judgment, and embodied physical skill of human experts. Thus, instead of seeking AI that âdoes the scienceâ for them, they envisioned future tools that improve human memory, keep track of documentation, and help prevent costly errors due to lapses in human attention. They sketched a series of speculative designs for future AI systems, summarized in Figure 4, that can 1) monitor task status, 2) organize lab-wide knowledge, 3) monitor scientistsâ health, 4) do field scouting, 5) do hands-on chores. Participants consistently favored passive, context-aware systems that could blend into existing workflows and support the conditions of human-scientific reasoning rather than AI that deprives them of the opportunity to âdo the actual work [myself].â (P11) We hope our findings inform ongoing conversations within HCI and the broader Human-AI interaction community about how to design tools that respect the unique constraints of real-world scientific labor. In doing so, we also surface a broader opportunity to expand the design space of AI beyond raw productivity and automation, and toward supporting the intangible human conditions that make knowledge creation possible. Our findings can inform AI design for other physically-based domains outside of science, such as medical practitioners, emergency first responders, artisanal craft workers, or field technicians working in demanding outdoor conditions. We make the following contributions to HCI: ⢠The first study examining how scientific practitioners perceive AI tools in embodied lab/field work âbeyond the deskâ ⢠Three sets of barriers that constrain AI use in materially grounded settings such as lab and field science (See Figure 1) ⢠Five speculative design concepts reflecting scientific practitionersâ future visions for AI support (See Figure 4) 2. Related Work Our paper extends the long lineage of research on studies of scientistsâ workflows in two novel ways: 1) by emphasizing how they envision modern AI helping them, and 2) by focusing on under-studied physical tasks âbeyond the desk.â 2.1. Studies of Scientistsâ Workflows Scientific workflows and collaborative practices have been long-standing topics of interest in HCI, CSCW, and Science and Technology Studies (STS). For instance, ethnographic studies have examined how scientists conduct their work, from the day-to-day of laboratory life (Latour and Woolgar, 1986) to the âshop workâ and âshop talkâ through which technical practices, material artifacts, and social processes coordinate to produce scientific knowledge (Lynch, 1979). Viewed through this lens, scientific practiceâfrom the lab to the fieldâis not only the execution of formal protocols, but also a complex negotiation of material and social conditions. Researchers have documented how tacit knowledge (Collins, 1974), embodied physical skill (Cetina, 1999), infrastructure (Vertesi, 2014), and environmental constraints (Yeh et al., 2006) shape everything from experimental design to collaboration. Starting in the 2000s, more contemporary work analyzed scientistsâ programming and data-centric workflows âat the desk.â Studies have examined how scientists develop and maintain software (Hannay et al., 2009; Carver et al., 2007), rely on computational tools (Huang et al., 2025), manage file dependencies (Gori et al., 2020), and organize high performance computing (HPC) scripts, directories, and datasets (Strong et al., 2011). Although prior work has offered rich accounts of how scientists work to produce scientific knowledge, most of it predates the widespread availability of modern AI, especially its use âbeyond the deskâ out in lab and field settings. Our interview study advances this lineage of research by being the first to examine how scientific practitioners engaged in physical lab and field work imagine modern AI assistance within their day-to-day embodied workflows. 2.2. Human-Centered AI in Science With recent advances in deep learning and generative AI, researchers have explored how these technologies can automate aspects of scientific discovery in tasks like molecular structure prediction, chemical synthesis, and hypothesis generation (Ramos et al., 2025; Jumper et al., 2021; Reddy and Shojaee, 2025; Baek et al., 2024; Rapp et al., 2024). While pushing the frontiers of what AI can do for science, most efforts center around automating processes of modeling, simulation, and algorithmic optimization (Lu et al., 2024; Jablonka et al., 2023) with GPU-enhanced computation. There is an emerging line of work on robotic arms to automate some physical lab processes like pipetting and diluting microfluids; however, these systems are currently expensive and inflexible, so they are limited more to larger-scale industrial production settings and not seen in most academic labs, where these tasks are done by hand (Arnold, 2024). Semi-autonomous UAVs and field robots have considerable potential to support field scientists in tasks such as interacting with remote environments (47), terrain imaging (Seifert et al., 2019), and environmental monitoring (Koukouvelas et al., 2023). Emerging work in robotics for fieldwork has begun to explore opportunities in this direction, highlighting challenges in animal species and individual identification, site access, and data handling (Pringle et al., 2025). However, existing systems are typically deployed for predefined missions, rather than multi-step or improvisational scientific workflows, and research in this area remains focused on technical and algorithmic capabilities. Recognizing the limits of full automation, HCI researchers have focused on Human-AI collaboration, where human scientists drive inquiry and AI assists (Schmidgall et al., 2025; OâDonoghue et al., 2023). These systems support tasks like experimental planning and validation (Schmidgall et al., 2025; OâDonoghue et al., 2023), data processing (Jablonka et al., 2023), and research question ideation (e.g. CoQuest (Liu et al., 2024), PersonaFlow (Liu et al., 2025)). Most work here has been on building new system prototypes; to our knowledge, there have been only three prior studies of how scientists use AI in their workflows. The first in 2023 presented interviews with professors and tech-industry scientists (often at more senior levels) about their desired use cases for AI in science; participants envisioned only AI support for knowledge workflows âat the deskâ, with potential in science education, data wrangling, literature review, coding, and technical writing (Morris, 2023). Two more recent studies at CHI 2025 focused again only on desk-based workflows: one studied how scientists and operations staff use an internal ChatGPT-style chatbot at a U.S. national laboratory (e.g., for writing emails, reports, and manuscripts) (Wagman et al., 2025), and the other on how scientists use AI for programming and data analysis (OâBrien, 2025). However, despite the rise of computational techniques, much of modern science still happens in the lab or out in the field, and thus exercise scientistsâ physical skills and embodied cognitive abilities that go âbeyond the desk.â Our work differs from these prior studies by focusing on the perspectives of scientific practitioners in improvisational, uncertain, and noisy environments such as wet labs or remote outdoor field sites. Our study provides a situated, embodied perspective of practitioners ranging from graduate students to experienced full-time lab staff, which complement the findings of prior studies focused on the desk-based computer-centric workflows of scientists (Morris, 2023; Wagman et al., 2025). We extend their findings by situating scientific practitionersâ perceptions, barriers to adoption, and speculative visions of AI within physical contexts. Taken together, our findings can be combined with those from prior work on scientistsâ use of AI at-the-desk to give the field a more holistic view of how lab and field science might be done end-to-end in the future with human-centered AI support. 2.3. Human-Centered AI in General Knowledge Work Zooming out farther beyond applications to science, consumer-available AI tools capable of programming, writing, and generating imagery have impacted many types of professional expert stakeholders, especially software engineers (Zheng et al., 2025; Vaithilingam et al., 2022; Khojah et al., 2024; Liu et al., 2023), designers (Subramonyam et al., 2025; Zhu et al., 2018; Shi et al., 2023), educators (Prather et al., 2025), writers (Guo et al., 2025; Ippolito et al., 2022), artists (Chang et al., 2023; Tang et al., 2024; Suh et al., 2021), and knowledge workers across areas like law, medicine, and business (Amershi et al., 2019; Fok et al., 2024). Recent studies have been directed towards understanding how AI could change or improve these domain workflows in terms of productivity, creativity, or collaboration (Vanukuru et al., 2025; Pu et al., 2025; Suh et al., 2024). Studies have also investigated how professionals use AI as a co-creative or assistive partner, such as surfacing tensions in how writers choose to integrate AI assistance (Guo et al., 2025). However, the aforementioned domains and studies typically involve âdesk work.â One notable exception (Kernan Freire et al., 2023) examined how LLM-based cognitive assistants can support factory workers in physically-demanding environments, for example, in helping workers resolve mechanical issues with production lines. Similarly, Kernan et al. proposes a set of design guidelines that emphasizes real-time data and domain-specific knowledge integration for AI assistants in manufacturing (Kernan Freire et al., 2023). While this line of research considers factory-style production labor, our study spotlights AI in the context of the behind-the-scenes physical work that powers science. 3. Methods To understand how scientists are engaging with AI tools in laboratory and field environments, including current use patterns, barriers to adoption, and future speculative design ideas, we conducted situated on-site interviews with 12 scientific practitioners who work in lab and field settings. While interviews were planned for 45â60 minutes, many continued longer based on voluntary participant interest. Participants received $30 USD gift cards, and this protocol was approved by our Institutional Review Board. 3.1. Participants We sought out a range of scientific practitioners (5 female, 7 male), in particular those who worked in settings that required physical presence at the lab bench or in the field (Table 1). Participants were recruited via direct email outreach, word-of-mouth, and snowball sampling. This resulted in 12 participants, spanning R1 university (34) labs to industrial research organizations (mostly in the U.S. but a few in the U.K. and Switzerland). Many operated in high-stakes or physically-demanding contexts, such as nuclear science, remote field robotics, applied neuroscience, and ecological fieldwork. Participants represented a broad range of expertise. Their roles ranged from junior university researchers to senior staff scientists (2 to 28 years of experience in their field). In the âRoleâ column of Table 1, the term âScientistâ specifically refers to participants who are formally employed under that job title and often have a doctorate degree. But note in interviews all participants self-identified as âscientistsâ and were paid practitioners engaged in scientific labor. Table 1. We interviewed 12 scientific practitioners at research institutions across multiple fields about the barriers and opportunities for AI tools in their daily working environment. The âYoEâ column indicates how many years of experience they have in their respective field. âR1 Public/Privateâ refers to U.S. high research output (R1) universities (34). âAI Useâ is self-reported frequency of AI usage. ID YoE Gender Field Institute Lab Description Role AI Use P1 3 M Behavioral Neuroscience R1 Public Neural activity of rodent behavior PhD student Sporadic P2 5 M Field Robotics R1 Private Robot locomotion and navigation dynamics PhD student Daily P3 2 F Field Robotics R1 Private Robot locomotion and navigation dynamics Undergrad researcher Daily P4 5 M Nuclear Fusion Industry Inertial fusion materials design and development Scientist Sporadic P5 3 M Behavioral Neuroscience R1 Public Neural activity of rodent behavior PhD student Sporadic P6 10 M Nuclear Fusion Industry Inertial fusion materials design and development Scientist Sporadic P7 3 F Nuclear Fusion Industry Inertial fusion materials design and development Data Scientist Sporadic P8 8 M Cognitive Science R1 Public Social behavior and cognition of primates PhD student Daily P9 28 F Materials Science Industry Medical device bioengineering Staff Scientist Rarely P10 5 M Neuroengineering Industry Biohybrid neural systems, cognitive preservation Scientist/Entrepreneur Daily P11 6 F Experimental Psychology R1 Public Spatial memory and cognition PhD student Daily P12 4 F Biochemistry/Cell Biology R1 Public Cellular physiology in cancer and kidney disease PhD student Sporadic 3.2. Design Rationale for Our Situated Interview Protocol Our interview study drew on the method of naturalistic inquiry (Lincoln, 1985) to guide data collection situated within participantsâ real-world scientific environments. Our protocol consisted of having participants walk through how they typically work on-site (e.g., at their lab bench) and mention where AI plays a role, discuss barriers to adopting AI in their workplace, then come up with speculative designs for how future AI tools could help them. Our protocol includes: 1) Situated on-site interviews: We held interviews directly in participantsâ workspaces when possible, such as at their lab bench. This allowed them to demonstrate workflow components while talking and pointing to surrounding tools, which elicited nuanced qualitative insights. In the few cases when in-person access was not feasible (e.g., visitors not permitted inside a clean-room), participants were asked to bring photos of their workspace to âvirtuallyâ walk through them. This enabled them to reference concrete tasks, tools, and features of their work environment. 2) What does âAIâ mean? An important consideration in this study is defining what the term âAIâ means. In accordance with our naturalistic inquiry (Lincoln, 1985) approach, we started each interview by asking each participant how they use âAIâ without explicitly defining the term for them or showing upfront examples. We adopted an emic approach (as opposed to etic) (14) by having participants self-define this term based on their own perceptions. In practice, unsurprisingly, most discussed ChatGPT or other chat-based, text-generating LLMs from the 2020s; but some nuclear scientists talked about 2010-era ML algorithms that had been developed in-house for classifying their bespoke datasets. Nobody mentioned GOFAI (i.e., symbolic or logic-based AI that predated the rise of machine learning in the 2000s) (16). 3) Speculative design activity: The final part of each interview (15â20 minutes) drew on methods from the field of speculative design (Auger, 2013; Hoffman, 2022) to elicit system ideas unconstrained by current technologies or practices. Participants sketched designs for an âidealâ future AI tool to assist their specific workflow. To help participants think beyond current practices and familiarize them with modern AI tools, the interviewer first asked guiding questions grounded in the workflows each participant described, then demonstrated relevant modern AI capabilities such as multimodal vision and language models they may not have encountered. This enabled them to expand upon their emic self-definition (14) of âAIâ that they initially provided to us. One scenario we presented was how someone building a custom PC can upload photos of a computer motherboard to ChatGPTâs vision model and ask it to identify a LED pin that they needed to use.333This scenario was intentionally chosen to be simple and domain-agnostic so as not to bias participants toward coming up with any specific kinds of speculative designs in their own field. It was also presented only after participants described the challenges inherent to their own physical environment and workflow. While this domain-agnostic scenario helped orient participants unfamiliar with multimodal AI, its primary purpose was to encourage participants to think beyond text-based interfaces. Speculative designs emerged mainly through contextual discussion of the barriers previously surfaced and through physical walkthroughs of their lab spaces, which primed participants to think in terms of their own domain. Participants were prompted to sketch their lab spaces while they came up with ideas, using the physical layout to imagine where and how AI could assist their embodied work. 3.3. Data Analysis Interviews were audio-recorded with consent then transcribed by an online service. Two authors reviewed this data independently to familiarize themselves with transcribed content. Transcripts were manually reviewed for errors and analyzed with an inductive approach using thematic analysis (Braun and Clarke, 2006). Our team iteratively generated and refined codes that reflected emerging patterns across multiple interviews. Codes were grouped by present-day AI use cases, perceptions of current barriers, and speculative design categories. For the speculative design activity, participantsâ sketches were coded alongside transcript data, which allowed the integration of design ideas into the emerging thematic structure. Throughout the four-month interview and analysis period, the research team met regularly to discuss interpretations and resolve disagreements before finalizing the set of themes presented in Sections 4-6. Here is a representative example of our iterative analysis process at work: initially, we mapped the raw data of reported use cases and barriers to tasks located in specific lab or field sites. However, the diversity of labs and field sites made this approach too narrow. Thus, we abstracted out to higher-level themes (e.g. âat the deskâ versus âbeyond the deskâ) that cut across contexts, which let us capture a wider range of experiences while preserving situational relevance. To analyze speculative designs, we applied a similar procedure. We began by categorizing designs that were tied to discrete physical spaces and types of tasks (e.g. pipetting assistance, surgery protocol assistance). Then, we merged them into broader archetypes, which allowed us to focus on participantsâ higher-level task rationale. For instance, pipetting assistance would be considered seeking support in âhands-on physical chores.â See Figure 3 for the thematic mappings that emerged from our analyses. We include two thematic maps to illustrate the structure of analyses across both empirical and speculative components of our study. The top panel summarizes three higher-level barriers that emerged from inductively coding participantsâ descriptions of lab and fieldwork. The bottom panel summarizes the five speculative AI assistant archetypes and their mid-level clusters, showing how participantsâ design ideas were organized into broader functional groupings. Figure 3. Thematic maps summarizing (top) the three categories of barriers that discourage lab and field scientific practitioners from adopting AI, and (bottom) five speculative AI assistant archetypes organized into mid-level conceptual clusters. The figure two thematic maps, one titled âBarriers that Discourage Lab and Field Scientific Practitioners from Adopting AI in their Workflows,â and the other titled âFive Speculative Designs for Imagined Future AI Scientific Assistants.â Each column represents a major theme and includes the relevant codes within. 3.4. Study Scope and Limitations We intentionally scoped our interview protocol and speculative design activity to focus on the day-to-day âon the groundâ work of scientific practice. Thus, our study does not cover higher-level considerations such as how organizations set AI policies and cope with systemic risks; see Wagman et al.âs CHI 2025 paper for thoughtful coverage of organizational and social issues around AI adoption at a large U.S. national laboratory (Argonne) (Wagman et al., 2025). We also did not cover broader issues such as AI ethics or philosophical objections to AI usage. Note that our participants were likely self-selected to be more open to using AI in their work, so we are lacking the perspectives of those who are strongly opposed to AI use. The ethics and norms around AI use in science continues to be debated, especially over challenges related to bias and lack of transparency (Ding and Li, 2025). Morris previously identified barriers to adoption including scientistsâ concerns with inaccuracy, hallucination, and falsified results (Morris, 2023); other fears include privacy, security vulnerabilities, and plagiarism (K. B. Wagman, M. T. Dearing, and M. Chetty (2025); 24). While some participants mentioned these topics in passing, they were not the focus of our study. Although we sampled participants across a range of disciplines (e.g., nuclear science, biochemistry, primate field cognition, neuroscience), we cannot make claims about how universally applicable our findings are across all of science. It is likely that we did not cover fields that may be more pro-AI or anti-AI than our sample. Additionally, while participants represented a range of expertise, our sample skewed towards early-career scientists-in-training and student researchers. Thus, we cannot claim broader generality across all scientists (especially professors and lab PIs who may spend more time on grant-writing, advising, and giving talks to academic and broader public audiences). We scoped the focus of our study on scientific practitionersâthe often-junior staff who do the day-to-day physical labor in lab and field settings. We direct the reader to findings from interviews with professors and industry lab PIs for complementary perspectives from more senior scientists (Wagman et al., 2025; Morris, 2023). Additionally, all of our participants were based in the U.S. or parts of Western Europe (U.K. and Switzerland). Our findings may not be representative of AI perceptions in geographical regions with different scientific infrastructures, funding models, or cultural norms (See Table 1 for participant demographics). Lastly, while we were allowed to visit and conduct interviews in-situ at lab and field sites, we did not directly observe these scientists doing their tasks live due to the time-sensitive and high-stakes nature of their work (e.g., we could not observe a rodent neural implantation surgery). Instead they walked us through a âsimulationâ of what they would ordinarily do while they pointed to their physical tools, which limited the fidelity of our observations. And while we showed participants unfamiliar with AI a domain-agnostic multimodal AI example to demonstrate capabilities beyond text-based chatbots, this may have subconsciously anchored some of their speculative designs. Additional illustrative examples may have primed participants to ideate more broadly. 4. Findings: How Lab and Field Scientific Practitioners Currently Use AI Tools In the following three sections, we first examine how scientific practitioners working in lab and field settings currently incorporate AI tools into their work (Section 4). We then describe the barriers they identified that limit adoption in practice (Section 5). Finally, we present five archetypal speculative designs for future AI assistants that participants came up with, which illustrate how they envision, in the ideal case, AI supporting both their day-to-day and long-term research (Section 6). Below we present two categories of current AI use cases: at the desk, and beyond the desk. 4.1. AI Tools for Knowledge Work âAt the Deskâ Of the 12 participants we interviewed, all had heard of AI tools, particularly recent Generative AI (genAI) tools. However, AI usage patterns varied widely, ranging from little-to-no use (P4, P7, P9) to sporadic reliance (P1, P5, P6, P12) to using on a daily basis (P2, P3, P8, P10, P11). The most commonly-mentioned AI tool was ChatGPT, followed by other text-based LLMs such as Claude, Google Gemini, and Perplexity AI-powered search. For most participants we interviewed, AI use remained peripheralâused intermittently on the computer (i.e., âat their deskâ) at the beginning or end stages of their workflows, such as during experimental planning, data analysis, or manuscript writingâand was contextually disconnected from their day-to-day physical work in the lab and the field. The most prevalent AI use cases were for programming and debugging (P1, P2, P3, P5, P6, P8, P10, P12), followed by brainstorming and ideation around experiment protocols (P1, P2, P3, P8, P10, P11). Several participants used AI for clarifying and organizing their thoughts (P3, P10, P11), while others had integrated or considered integrating it into their scholarly communication practices through paper writing assistance (P2, P4, P7, P12), figure generation (P2, P10, P11), and visualization support for presentations (P10). Information-seeking behaviors included using AI for literature review tasks (P1, P8), personal tutoring for learning new concepts (P3, P5, P11), and general knowledge queries (P1, P2, P3, P6, P7, P8, P10, P11, P12). Other use cases included interpreting technical manuals (e.g. setting up trail cameras) (P8), verifying experimental calculations (P5, P12), writing emails, or assisting with paperwork (P7, P11, P12). When participants walked the interviewer through their work-spaces, many drew clear distinctions between computational and physical spaces. P1, who works in a wet lab studying the neural activity of rodent behavior through in vivo electrophysiology, stated that he âmainly only uses [AI] for data analysis.â P1 only uses AI tools on his computer at home or at the computer station in an office next to the wet lab. P5, who works in the same field, also expressed that he âuses [AI] seldomly. I probably donât utilize it as much as I could for my benefit, but itâs because itâs mainly for coding.â Note that both of these participants spend significant amounts of time directly interfacing with animals: training, feeding, socializing them, surgically implanting custom-built âdrivesâ to measure neural behavior, and running experiments. In these contexts, where research spans years and involves hands-on, dynamic interaction with living systems in-situ, AI tool utility was perceived as limited and more suited to purely computational work back at their desks. In sum, the majority of participantsâ current AI usage was for computer-based knowledge work at their desks, which confirms findings from the three known prior studies of scientists and AI use (Morris, 2023; Wagman et al., 2025; OâBrien, 2025), along with broader studies of other types of knowledge workers (Guo et al., 2025; Khojah et al., 2024; Pu et al., 2025; Liu et al., 2023). Where our study goes beyond prior work is uncovering use cases, barriers, and design opportunities for scientific AI use âbeyond the desk,â which we present below. 4.2. AI Tools for Physical Work âBeyond the Deskâ Several described using AI tools in their lab or field environments, such as voice-activated assistance when their hands were busy (P8), help during DIY (Do-It-Yourself) engineering tasks (P8), and hardware troubleshooting (P2, P3, P11). For instance, P8, a field researcher who studies primates (e.g., macaques, orangutans, chimpanzees) in remote locations around the world such as rainforests, often needs to engineer experimental setups on-site to conduct his research. These apparatuses cannot be pre-constructed due to their size and are sometimes built with salvaged or improvised materials on-site. P8 recounted a logistically-challenging experience in the Republic of Congo, where he had to construct an experimental apparatus from âdriftwood that [he] found and random objects that [he] uses.â P8 once asked ChatGPTâs voice assistant mode to help him construct an apparatus made of acrylic, wood, and epoxy to observe primate inhibitory control: âI can directly ask [ChatGPT] questions, like trivial stuff when I build something, which Iâm not great at. I sort of figure stuff out on the go, but if I ask it, âOh, I need to cut acrylic glass, what tool should I use, whatâs the right blade?â ⌠for even stuff like that, I can use it in the field.â This use case illustrates how AI tools can help participants âfigure out tools and solutionsâ (P8) and serve as low-friction, just-in-time support for improvisational DIY problem-solving. In another case, P11, a doctoral student researcher in the field of experimental psychology works with virtual reality and game-controller equipment in her experiments and uses ChatGPT to troubleshoot hardware problems. She once nearly canceled a day of scheduled study participants due to issues with a computer graphics card âbut then I gave it to ChatGPT, [and asked] what are all the possible issues for this? What are some of the workarounds, can I replace the cable? Like, going through all the different diagnostics I can do on my end before I take that equipment to a specialist or get outsourced help.â She later said that she found this form of AI assistance empowering: âIâm working on a computer vision project right now, which I have zero expertise in, but with ChatGPT, I feel Iâm more empowered to do stuff. I donât have to know every single thing.â These examples show how scientific practitioners need to pick up a wide range of improvised DIY skills on-the-job, often in domains outside their formal academic training. In P8âs case, although his expertise was in primate observation, he pointed out that he needed domain knowledge in mechanical engineering and woodwork. For P11, although she conducts research in the field of cognitive psychology, in order to run her experiments, she needed to have domain knowledge in computer hardware. 5. Findings: Barriers that Discourage Lab and Field Scientific Practitioners from Adopting AI in their Workflows Refer to Figure 1 on Page 1 to see three barriers that hindered participants from using AI for lab and field work. 5.1. Scientific practitioners do not want to face risk of AI errors since physical experimental setups are too high-stakes The most frequently-mentioned barrier to AI usage was the trade-off between potential time saved and the devastating cost of AI-induced errors for physical setups that are hard to setup and maintain. This concern spanned labs both large and small, especially those that worked on projects that required substantial investments of time and money. For example, P1, a doctoral student neuroscience researcher, described the extensive, months-long cycle of his labâs experiments: acquiring and training rodents (spans 6 months), hand-building delicate custom electronic âdrivesâ designed to measure neural activity (spans 3 months), surgically implanting these tiny drive devices into rodents, constructing bespoke experimental setups, and running behavioral trials to collect data. In particular, pre-experiment surgeries were especially high-stakes and time-consuming: âIf we are hitting one [brain] region in a surgery weâve done before, it can usually take eight hours. The first surgery for [labmateâs project] took, I want to say closer to 16 hours, because it was multiple [regions]â (P1). In this context, P1 feared relying on AI for informational support because even a single mistake could cost him âthe 10 hours that I was in surgery at that point, but also the months that I had spent training this ratâif the surgery goes wrong and the rat dies.â P1 summed up the risk in one question: âHow much time does [AI] save versus how much time does it cost me if itâs wrong?â Figure 1A illustrates this example. Participants described similar stakes at larger industrial labs, with an additional element of errors being dangerous and/or amplifying due to AI inaccuracy. P4, P6, and P7âs scientific work revolve around designing fuel capsules for experiments at U.S. national labs aiming to achieve nuclear fusion. These experiments rely on rare and costly âbeamtimeâ (Traweek, 2009) opportunities to access expensive lab-wide shared laser systems. Collecting sufficient data for even a single experiment can span years, and failed trials cannot easily be repeated (P7). P4 further emphasizes that: âItâs imperative that every data set that we acquire is correct [âŚ] thereâs a lot more at stake in that measurement. If youâre limited to 10 opportunities, if you mess up one, then you canât get that back. Any delays that we incur are going to affect everyone downstream from this as well; in severe cases, that can jeopardize the shot.â That is why P4 preferred human judgment and human validation over AI tools: âBecause of the gravity of the data that weâre producing, people are much more comfortable trusting another person to do that.â P6 added that it was hard to trust AI due to its inability to express nuance: âChatGPTâs writing style makes it sound like it is 100% correct. Like itâl confidently tell you things that are actually wrong, and thatâs really dangerous [âŚ] if Iâm proposing a hypothesis or giving my opinion, I want to also have the nuance of saying, âHey, this is what I think, but it might not be entirely true.ââ 5.2. Physical environments of lab and field sites make it hard to use current AI tools Embodied research often takes place in challenging environments that require participants to, for example, collect observational data outdoors (P2, P8) or work in highly-controlled settings to protect sensitive experimental conditions (P1, P4, P6), such as in nuclear fusion materials development. Because a substantial portion of their time is spent in these environments, many participants could not physically access electronic devices that run AI tools like ChatGPT. For instance, during a tour of P7âs nuclear fusion lab, she gestured toward a notebook placed outside of lab doors and explained that all external equipment, including phones, laptops, and paper notebooks, were prohibited inside due to contamination risks, which meant no way to access AI assistants inside. Scientists jotted down notes before entering the lab, where full-body clean-room suits are required: hood, hair bonnet, face veil, gloves, coveralls, and boots. P12, who works with cancer cells, must also work in similar clean-room environments (see Figure 1B) using laminated paper documentation rather than electronic devices: âYou always have your gloves on with some reagents on it. Maybe some cells are even on it [âŚ] you take [the laminated paper] with you wherever you go, so you can still follow the protocol. At the end of the day, you will spray it, clean it.â In another instance, P1 noted that he could not have cell phones or laptops in the experiment room because electromagnetic waves, such as Wi-Fi, may interfere with electrode data collection. In the field AI-enabled devices were avoided not because of contamination risk, but to prevent behavioral interference with the animals being studied such as when brightly-lit devices or voice dictation might distract primates and harm data collection for P8. He remarked that while he did occasionally use his phone to consult ChatGPT, it often was âmessy when you do testing. Iâm wearing gloves. I am in a sandy, grimy environment, where I donât want to take out the phone to type. Itâs definitely a little bit more rugged and reliable to have pen and paper and a clipboard.â 5.3. AI cannot match the tacit knowledge and contextual judgment of human scientists Laboratory and field science involve lots of tacit knowledge consisting of hard-to-capture âunwrittenâ expertise and embodied decision making (Polanyi, 1966), where scientists rely on physical senses such as sight, touch, and even smell to evaluate problems in situ. Many participants were skeptical of AIâs ability to match humans in these skills. For instance, Figure 1C shows how P3 found ChatGPT unable to help troubleshoot mechanical issues with a robot while out at experimental field sites in deserts and volcanos: âBecause itâs hard to explain it to the AI in a way that makes sense. Like, I canât really explain how the robot moves; itâs difficult to use my words.â P10âs workplace once debated whether to use AI to automate the placing of fragile organoids beneath a microscope. However, they decided it was more efficient and reliable for a human to do this work, in part because the lab can be a very ânoisyâ environment that makes current AI tools infeasible: âSay you have your pipette and it falls on the groundâŚitâs dirty, but you have something in your hand, so you put your stuff down, you take the pipette up and replace it. Itâs contaminated. Maybe you should change your gloves. Thatâs so much noise for just one action that happens with something falling on the ground [âŚ] A human brain is actually unmatchable, that biological brain, because weâre made exactly for tasks like that [âŚ] We can very easily adjust based on previous knowledge, gut feeling, intuition.â (P10) Moreover, due to the trial-and-error nature of such work, P10 added that some tasks may only be performed once, which makes training a task-specific AI replacement not worthwhile. P10 described this threshold as a âgolden zoneâ where it is âbetter to let the human do it.â A field researcher also shared how his tacit expertise could not be replicated by AI. Over time, P8 has developed the ability to distinguish between chimpanzee subjects by sight: âFor a group of like, 15 chimpanzees, I can tell them apart. Others cannot [âŚ] at some point, you have an intuition.â Despite attempts to improve computer vision models, AI tools to identify animal subjects âis still way behind to be good enough for the kind of complex data we have.â P12 explained that AI could only assist with piecemeal tasks, such as drafting a protocol for âspecific positions, or like small pieces of experiment planning.â However, planning an entire experiment end-to-end requires sequencing tasks across multiple physical stations, a process that relies on oneâs detailed tacit knowledge of their labâs architectural layout accumulated from firsthand experience. As P12 put it, âFirst I go into the cell culture⌠then I do my RNA extraction⌠then I go to the other room⌠AI cannot do this [because it] cannot connect the different [lab] stations. It doesnât know the timeline, the context. It cannot plan a whole experiment for me because it doesnât know how an experiment is set up.â Lastly, P4 mentioned how evaluating scientific work could be a surprisingly subjective process, which is what separates an expert from a novice and cannot be feasibly replicated by current AI tools. He explained, âThere isnât necessarily a clear right or wrong in terms of data quality [âŚ] is this within the realm of acceptability? Are there negative influences from, letâs say, the environment or instrument that would warrant me to have to retake the data set?â P4 later used the term âartisanâ to describe the nature of scientific procedures that often evolve over time, which highlights the interpretive, experience-driven aspects of experimental work. 6. Findings: Five Speculative Designs for Imagined Future AI Scientific Assistants Figure 4. Five archetypes that are representative of speculative designs that our 12 participants envisioned for how future AI could help lab and field scientists: (A) as the labâs collective knowledge keeper, (B)-(C) distributed lab-wide task status monitor, (D) real-time monitor for scientistsâ cognitive and physical health, (E) mobile scout for fieldwork, (F) collaborator for hands-on physical chores. (Image credit: all illustrations were drawn by human artist Lauren Cheng without AI assistance.) A larger illustration shows an overhead view of a lab space with a distributed AI system across multiple rooms labeled (A), (B), and (C), while three illustrations underneath labeled (D), (E), and (F) demonstrate scientists using AI to assist them in their work. In the chemical reagent room (A), a screen located on the lab bench wall houses relevant protocols, notes from colleagues, reminders, images, videos, observations, and warnings specific to the lab. A camera on the device is directed at the work surface. In room (B), cell incubators are highlighted to signify detection of a drop in temperature. In another room (C), a scientist focused on a separate task is alerted about the cell incubators by the screen at her lab bench and by an audio warning. Illustration (D) depicts a scientist wearing a cleanroom suit facing a wall-mounted screen displaying âscanningâŚ,â health information, and suggested tasks. Next to them is a door with a biohazard symbol. In (E), a field scientist working in a remote area deploys stimuli to far away primates via a drone, monitoring their behavior on his device. Illustration (F) shows a woman directing a humanoid robot to pick up a fifty kilogram glass panel used in a muddy terrain tank in her robotics lab. As each participant was conveying barriers to AI adoption in their workplace, that provided a natural segue into our speculative design activity, where we had them sketch out what an âidealâ future AI assistant would look like for typical tasks. They drew a map of their lab or field space on paper and pointed out how an AI might help. Their design ideas varied depending on the type of task and its location within the lab or out on the field (e.g. lab bench, fume hood, incubation room, surgery room, experiment clean-rooms, outdoors, see Figure 2 for examples). Due to space limitations here, we will not present every single design idea that participants had. Instead we grouped similar-themed ideas together into a set of five archetypal speculative designs and present them here. 6.1. Archetype 1: AI as the labâs collective knowledge keeper Across nearly all participants, human limitations in knowledge transfer and documentation emerged as persistent bottlenecks. Participants spoke about limitations in their ability to recall and pass down expertise (P1, P2, P3, P4, P5, P6, P7, P9, P10, P12). P3 summed it up as: âNo one likes to do documentation when youâre actually working. You just want to work [âŚ] but afterwards you donât remember certain parts. Itâd be super helpful if Generative AI could do it for you.â Experienced scientists struggled to pass down contextual knowledge and expertise to junior colleagues due to fragmented documentation and overwhelming context (P1, P2, P4, P6, P9, P12); and junior scientists struggled to navigate lab-specific, idiosyncratic, tacit practices without clear guidance when a mentor was not nearby (P3, P5, P4). Some labs maintained shelves of note-filled binders that often went unread, while others had abandoned multiple attempts to create usable computer documentation systems or digital archives. Even when documentation did exist, it was frequently described as dense, confusing, and constantly evolving (P1, P3, P4, P7), exacerbated by each lab and organization having their own custom protocols, tools, equipment: âThereâs a lot of tribal knowledgeâ (P4). Thus, in our design sessions participants envisioned context-aware AI systems that could alleviate these gaps by capturing, summarizing, and situating knowledge in real time. For example, P2 imagined a mounted AI camera system that could monitor his robotics experiments through video, digest the streamed footage, and âprovide corrections and live feedback in-contextâ for the experiment. He envisioned being able to query the AI, which would then refer back to specific video frames when generating responses. P1 imagined scenarios where future AI could actively participate in data collection and notetaking: âThereâs a ton of digital information that it could automatically take the notes from, âHey, this thing co-occurred with that thingâ [âŚ] or integrating behavioral tracking in my note taking like, âAt what point does the rat do this?â Or if I just set up a boundary, âHereâs the area I want to keep a track of and hereâs the rat. Make a note when the rat leaves this area.â P12 combined a similar idea with lab-based tacit knowledge-sharing, where context was automatically captured and passed on to colleagues: âBecause we have a couple of people in the lab, if AI could note that I made a mistake there, and for everyone else that has to use this protocol again later, tell them âPlease be aware this happened to me. Thatâs how you solve it.â Then, [AI] could pop up and tell you, âSomeone had a problem here at this step.ââ See Figure 4 (A). 6.2. Archetype 2: AI as a distributed lab-wide task status monitor Whether in the lab or out in the field, participants frequently identified cognitive limitations that frustrated them, such as difficulty recalling specific procedural knowledge (P1, P2, P3, P4, P5, P10, P11, P12) and lapses in memory or attention in the midst of doing hands-on tasks (P1, P4, P5, P8, P11). These moments of cognitive load were seen as opportunities for AI intervention as a persistent memory aid, context-aware feedback provider, or auxiliary observer. For example, P1 faced attentional limitations during live experiments with rodents since he needed to juggle multiple simultaneous processes: evaluating data (audio and visual), attending to electrophysiology data, and managing other active tasks, such as âpresenting a certain stimulus at a specific time, taking the rats out, cleaning the arena, setting it back up, putting them back in, baiting a maze [âŚ]â He also needed to simultaneously attend to equipment failures or apparatus malfunctions, then react quickly. P1 imagined that a camera-based AI assistant would be able to offload some of this burden, acting as an auxiliary monitorâan extra pair of eyesâthat can provide attentional scaffolding when cognitive load exceeds human capacity: âI could totally see it being useful in tracking the things that I might normally be looking at but that I just canât during this time because Iâm too focused on this other part or something going wrong.â P12 designed for similar attentional limitations when multi-tasking in a cell biology wet-lab with many ongoing tasks and sensitive equipment. Her idea in Figure 4 (B-C) illustrates experimental protocols distributed across screens in different rooms, and AI monitoring multiple ongoing tasks for her. P12 explained how this system could prevent error by surfacing contextually important information for her at the right time, âImagine you walk around and you have a screen there like, âAh, Iâm at this step. I need this and this at this station.â Then, at every screen, every station, you follow the next steps.â P12 added that this could help prevent careless mistakes; she shared a story where a labmate accidentally left a cell incubator door open, which ended the experiment and delayed their work for weeks. A system could detect the open incubator door in the equipment room (where people usually are not present), illustrated in Figure 4 (B), and then send an audible alert to a scientist working in a different lab room, as shown in Figure 4 (C). P5âs hands-free voice-interaction design was informed by his workspace, where access to input devices like computer keyboards was not possible: âMy biggest concern in this space is my hands are often full in some way.â During the design session P5 demonstrated how he constructs a rodent brain implant drive under time pressure to check on other tasks: âLike, Iâm holding this [tetrode wire that is thinner than a human hair], and I finally have the position I want, but Iâm thinking, âHey, is this actually gonna fit?â or Iâm not sure how much I have on my alarm left before I need to check on the rats. If I could ask AI on the spot without me having to stop and go to a computer, like a distributed system throughout the lab.â The voice-activated AI monitoring system he envisioned (similar in form to Figure 4 (C)) aligned with his vision of being able to attend to multiple aspects of the lab concurrently while reducing the friction of task-switching. 6.3. Archetype 3: AI as a real-time monitor for scientistsâ cognitive and physical health Besides monitoring lab tasks, participants also envisioned AI systems that act as status monitors for the actual scientists who must do high-stakes and high-intensity work. They brought up concerns that dangerous memory and attention lapses stemmed not only from stressful workloads, but from cognitive and physiological fluctuations in their bodies. In P4âs lab, he and his colleagues were encouraged to self-evaluate oneâs own headspace and physical condition before certain kinds of intensive work with nuclear materials. Drinking too much caffeine the night before could be reason for pause. P4 added, âYou donât want to be in autopilot [âŚ] I make sure not to do this [task] unless I feel like Iâm focused.â P6, who shares this concern, speculatively designed an AI assistant that could screen for these physical and cognitive fluctuations and alert him accordingly before entering the experiment room: âIf it notices over the past two days, your heart rate was elevated more than normal, itâl warn you.â See Figure 4 (D) for an illustration of this idea. More generally, several participants (P1, P4, P5, P10, P12) worried that lapses in memory or careless mistakes could lead to unwanted harm and cascading consequences: âIf your work rhythm gets thrown off, you can forget things [âŚ] if you forget to arm [activate] that thing, you donât collect any data. Thatâs commonâ (P5). P5 envisioned querying AI before running an experiment as a quick cognitive status check, âHey, Iâm about to start, am I forgetting anything?â 6.4. Archetype 4: AI as a mobile scout for fieldwork Participants working in the field envisioned AI assistants that were mobile, lightweight, compatible with remote or rugged environments (P2, P3, P8), and powered by local AI models with offline access (P2, P3). Unlike lab-based researchers, these participants often moved between locations and conducted work that required environmental adaptability. Because voice interactions with AI could disrupt naturalistic conditions or alter subject behavior (e.g., for animals being observed in close proximityâP8 and P11), they explored alternative modalities in speculative designs. P8, who often collects data at primate sanctuaries, imagined using small autonomous flying drones to reduce the amount of time spent trekking across long distances: âOftentimes, there is time wasted to figure out where even the animals are.â P8 speculated on future capabilities of these drones: âYou can probably efficiently collect a lot of information about where animals are and what theyâre doing by using a drone.â Beyond passive observation, P8 envisioned active intervention, such as deploying stimuli or food via drone and recording the resulting behaviors: âWhy couldnât the drone do that for me and just also then film it and see what happens?â These designs, shown in Figure 4 (E), illustrate how highly-mobile and non-verbal AI modalities align more with the constraints of field-based research. 6.5. Archetype 5: AI as a collaborator for hands-on physical chores Many participants struggled with technical chores that were not core to their research questions but still essential to running their experiments: soldering (P1, P5, P3), epoxying (P8), sawing (P8), wiring (P2, P3, P11), cementing (P1, P5), mixing chemical substances (P1, P6, P9, P12), performing surgical procedures (P4, P6, P10). These skills were often self-taught or learned through trial and error, since they had no formal professional training in those fields. In some cases, recently-acquired skills were only useful for a single experiment. As P8 noted, investing significant time in mastering such skills could slow research momentum, moreso when it can be slow to get feedback on them: âMy PI [principal investigator] would get a chance to look at my [hardware devices] a week after I built them and say, âOkay, this one looks really good. This one isnât sturdy enough. This one isnât wound tight enoughâ [âŚ] but Iâd forgotten what I did already.â After surfacing these pain points, participants imagined multimodal AI assistants that could provide scaffolded, live support during such embodied hands-on work. Both P1 and P12 suggested modalities that allowed them to visually learn while they picked up new protocols: âIf I could have some audio thing saying, âWhatâs the next step?â Or have a screen that displays the next step instruction or visual instruction [explaining] what this part is [âŚ]â (P1). They shared how audio interaction with AI could provide the best hands-free assistance, alongside visual instructions or diagrams on the screen to serve as persistent references, offering clarity during unfamiliar or error-prone procedures. Human sensory and physical limitations were also surfaced in the interviews. P12 noted that, âWe have things you cannot see by eye [âŚ] you have to train your muscle memory.â Multiple participants reported challenges related to the steadiness of their hands (P1, P3, P5, P12) or sustained attention and focus (P1, P4, P12), compounding the difficulty of performing hand-eye coordination tasks that required precision. These participants tended to extend their visions into more embodied, robotic forms of AI assistance; their imagined robots were designed to compensate for human sensory or physical limitations like acting as an extra pair of eyes and hands to help them in these chores. For example, P10 and P3 envisioned humanoid robots that could help lift, transport, or stabilize materials, whether delicate tiny components like neural cells or heavy equipment like glass panels during experiments (Figure 4 (F)). For P3, this also made certain parts of experiments more accessible: âAn extra pair of hands would help most when you have to take glass panels out [âŚ] itâs really heavy.â Scientific practitioners preferred passive AI, with proactive intervention for critical errors. One recurring theme we observed across design sessions was that most participants (P1, P3, P4, P6, P7, P8, P9, P12) leaned toward passive rather than proactive assistance. Passive assistance stays in the background, offering support only when directly queried by the user. In contrast, proactive assistance engages with users by issuing real-time alerts or suggestions based on human behavior. Some participants were more strongly averse to proactive assistance than others. P9 felt it would be disruptive to her workflow and wanted AI systems to notify users through subtle cues: âI would prefer the robot to raise its hand or flash until I touch it. I donât want it to interact and directly talk to me like, âHey, you missed that!â That would drive me nuts.â Modalities of AI intervention that came up most frequently in sessions were wall-mounted screens and hands-free, voice-interactive systems. Notably, wall-mounted screens without audio or voice controls were a preferred modality for participants who were concerned that audio could be a potential confound or distraction when running experiments (P1, P11, P12). In these scenarios, participants preferred AI assistants to provide suggestions, instructions, and warnings via text on large mounted screens, or intervene through minimal visual cues, like flashing red lights (P9). More generally, P11 did not want future AI systems to deprive her of the opportunity to âdo the actual work [myself].â That said, some (P1, P2, P5, P10, P11, P12) favored a mix of passive and proactive. They emphasized that although passive systems were preferred, they wanted AI to intervene when a mistake was imminent, especially in tasks where mistakes could be costly or hard to detect (P2, P12). P5 wanted control over this threshold of AI intervention: âIdeally, thereâd be different preferences. For one, because you obviously donât want [AI] to every 30 seconds go, âHey, do this. Hey, do that.â But in emergency situations, like, âHey, you forgot to arm a system,â I want that.â 7. Discussion Our participants judged AI less by raw capability than by how well it fit the embodied, high-stakes conditions of their work. They found the most value in systems that eased cognitive load through memory, attention, and documentation support, but preferred these to remain passive in the background unless errors became critical. More broadly, this suggests a shift in framing; rather than designing AI as a collaborator that directly participates in scientific reasoning, AI may be more effective when treated as infrastructure that supports the conditions under which scientific reasoning happens. 7.1. Limitations of AI for Physical and Embodied Scientific Work Despite rapid advances in recent tools, AI still remains peripheral among our participants in the physically-oriented work of lab and field science. Corroborating the findings of prior studies (Wagman et al., 2025; Morris, 2023), it is used mostly âat the deskâ for planning, literature review, data analysis, or writing. Despite the status quo, participants were able to use our speculative design activity to envision deeper, context-aware AI integration into their future workflows. This gap highlights structural mismatches, not necessarily technical limitations. As an analogy, current AI assistants resemble âgeneralistâ programmers or technicians who can write boilerplate code but lack the bespoke domain expertise to identify which approaches are the most insightful within a scientific field. Current AI tools are also ill-suited for improvisational and materially grounded practices that characterize much of day-to-day scientific work âbeyond the desk.â While extensive prior work has posited that knowledge-based professions are more readily augmented by, or even displaced by, AI (Tomlinson et al., 2025; Brynjolfsson et al., 2025; Maslej et al., 2025; Sherman, 2025; De Cremer et al., 2023), our study findings suggest that scientific domains involving physical manipulation, situated judgment, and tacit expertise currently resist such displacement of human labor. Tasks requiring bodily coordination and tool improvisation (e.g., meticulous pipetting techniques, DIY repairs in the field) exemplify the kinds of uncertain dynamic settings where scientists viewed human adaptability as irreplaceable. Iterative experimental work often requires scientists to learn new skills or obtain new materials and equipment. Automation needs therefore vary widely depending on individual situations and cannot be easily routinized with AI. Our findings complement recent analyses on AIâs impact on various occupations (Tomlinson et al., 2025), which shows that roles involving manual labor, operating machinery, or other physically-oriented tasks currently have the least potential for AI collaboration. While Tomlinson et al. (Tomlinson et al., 2025) show that AI aligns best with knowledge work and communication-heavy tasks, our findings reveal how even within knowledge-intensive scientific roles, AI adoption is uneven, concentrated to desk-based activities, and largely absent from lab and fieldwork practices. Although we focused on scientific practitioners working in physical domains, we view their work as a proxy for other occupations with similar core elements such as, say, hospital surgeons or emergency first responders. 7.2. Future Opportunities for AI in Scientific Practice Speculative design activities at the end of each interview revealed directions for how AI could be integrated in ways that augment scientistsâ capabilities without displacing their judgment (Figure 4). From our findings, we derive four higher-level design opportunities for AI, which address recurrent challenges that our participants faced. These opportunities build directly on the five speculative archetypes described in Section 6. (1) Design to scaffold memory, attention, documentation, and knowledge transfer in physical tasks. Archetype 1 (AI as a collective knowledge keeper) and Archetype 2 (AI as a distributed task monitor) show how participants struggle with fragmented documentation and lapses in attention during complex embodied work, as well as difficulty capturing and understanding tacit knowledge. These archetypes point toward designs of real-time, context-aware recording and summarization tools to reduce cognitive load âin the moment,â preserve expertise, and support knowledge transfer. Unlike desk-based tools that require users to stop and externalize knowledge, support should operate in the background during embodied activity. This distinction matters because physical workflows make it difficult to shift attention or focus away from the task at hand; knowledge must be captured and organized in situ rather than through retrospective, at-desk effort. (2) Tune AI proactivity to relative stakes. Archetype 2 and Archetype 3 (AI as a monitor for scientistsâ cognitive/physical health) illustrate how participants determine interruptions by AI to be acceptable when errors or safety risks carry meaningful consequences, such as open incubator doors or unsafe physical or cognitive conditions. Scientific practitioners otherwise preferred systems that stayed passive, intervening only when errors had high stakes. Thus, enable AI to adapt its intervention level to task context, with user-configurable thresholds (e.g., use only visual cues in settings that are sensitive to sound, such as with animal subjects). (3) Expand interaction modalities beyond the desk. Arche-type 2, Archetype 4 (AI as a mobile scout for fieldwork), and Arche-type 5 (AI as a collaborator for physical chores) illustrate that scientists often work with their hands occupied, across multiple rooms and workspaces, or in remote outdoor settings where desktop/tablet interfaces are impractical. Participants envisioned solutions ranging from voice interfaces, large ambient displays, and physical indicators, to flying drones and modular robotic assistants. AI systems need to be embedded in and responsive to the environmental and spatial realities of lab and field work. (4) Embrace digital naturalism in designs. Archetype 4 highlights the need for AI systems that can adapt to local environments and are able to operate under mobility and limited infrastructure. Many future AI designs could benefit from elements of digital naturalism, a design philosophy that advocates for modular, field-adaptable, and open-ended scientific tools (Makery, 2019). Quitmeyer cautions against stripping away the rich, local contexts of scientific practice and environments (Quitmeyer, 2024). These tools should be materially embedded and co-configured on site, becoming extensions of scientistsâ senses and mental models, and helping them form deeper connections to their work instead of abstracting it away âat the desk.â Taken together, these opportunities suggest designing AI as infrastructure integrated into scientistsâ sites of use. 7.3. AI as Infrastructure instead of Collaborator: Supporting the conditions for human scientific reasoning Current public conversation around AI in science has focused on using it to automate reasoning and discovery, as exemplified by highly-funded corporate projects such as Google DeepMindâs AlphaFold, which won the 2024 Nobel Prize for Chemistry (Jumper et al., 2021), along with follow-up systems like AI Co-scientist, âa multi-agent AI system [âŚ] to accelerate the clock speed of scientific and biomedical discoveriesâ (Gottweis et al., 2025). Despite the attention given to these large-scale automation efforts, our study participants described the day-to-day of lab and field science as surprisingly âartisanâ (in the words of P4), requiring hands-on skill, tacit knowledge, and situated improvisation while doing physical chores in the lab. We saw that many challenges our participants faced were less about enhancing reasoning and more about sustaining the complex, tightly-coupled physical conditions that make deep scientific reasoning possible. Participants derived fulfillment from discovery; it was the mental and procedural burdens of minimizing mistakes within their physical environments that they struggled with. Participants consequently did not envision AI as a collaborator in reasoning or decision-making. They instead imagined intelligent, context-sensitive infrastructureâ tools that could passively monitor workflows, intervene to prevent errors, and reduce time wasted on unproductive troubleshooting attempts. The high stakes of error were surfaced as a major constraint to AI adoption. Mistakes meant weeks of lost labor or, in extreme cases, catastrophic consequences. With these conditions in mind, participants were open to AI systems that prevented lapses in attention or memory, but were skeptical towards AI tools that produced confident outputs without accountability or transparency. While pattern recognition and workflow monitoring were welcome, interpretive AI systems were seen as brittle, lacking the embodied, tacit knowledge that scientific judgment relies on. These findings suggest a reframing of AI integration in scientific contexts. While much of recent work has focused on building AI systems that can reason, hypothesize, or automate decision-making, our study participants had different unmet needs. Instead, they imagined intelligent tools that strengthen the underlying infrastructures that make scientific reasoning possible: monitoring experiments, stepping in rapidly to prevent human lapses, which includes bodily limitations in memory, attention, and focus, and preserving context across their local communities, spaces, and time. This reframing aligns with prior critiques in human-centered computing, such as concerns that current computational systems risk imposing solutions that prioritize abstraction and detachment from physical contextââreinventing virtually every other site of practice in its own imageâ (Agre, 1997). Suchman and Agre argue that all research practices, âeven the most analytic, [are] fundamentally concrete and embodiedâ (Suchman, 1987), and our findings reflect this. The scientific practitioners we interviewed did not see reasoning as something that could be separated from the physical and practical conditions of their work, and were skeptical of AI tools that treated it that way. They advocated for more âdefensiveâ speculative AI designsâsystems that mitigate risk and preserve stability instead of maximizing speed and efficiency. Despite advances in technology, many scientific practitioners still rely on paper-based tools (e.g., P12âs laminated paper protocols, P8âs field notebooks), underscoring a persistent misalignment between current automation and scientistsâ real-world needs. Two decades ago, Yeh et al. observed similar preferences among field biologists and created ButterflyNet, a hybrid physicalâdigital system to bridge this gap (Yeh et al., 2006). Today, with AI and technology far surpassing 2006 capabilities, this is an opportune moment to revisit hybrid system designs grounded in the realities of scientific practice. In addition, the growing flexibility and declining costs of modern AI systems make them feasible for ad-hoc appropriation. This is especially relevant in the scrappy, resource-limited environments our participants described. In such settings, bespoke digital tools are costly, niche, and time-consuming to build, risking obsolescence as research directions shift. Fortunately, todayâs AI models are becoming less expensive and easier to adapt, offering potential for AI tools that can be rapidly customized by end-users. While laboratory and field science provided the frame for our work, these findings might also offer insights for other high-stakes, embodied, and situated domains, including medical practitioners, artisanal craft workers, and field technicians that work in challenging physical settings. 8. Conclusion This paper is the first, to our knowledge, to report empirical accounts of how laboratory and field scientific practitioners in embodied and improvisational contexts envision the role of AI. Through on-site interviews and speculative design sessions at 12 scientific practitionersâ workplaces, we surfaced three barriers to adoption: experimental setups being too high-stakes to risk AI errors, constrained physical environments, and AIâs lack of embodied tacit knowledge. Thus, today AI is mostly used for desk-based science work, but our participants were able to sketch ideas for future tools that scaffold memory and attention during hands-on practice and extend across modalities from robots to lab-wide systems. These speculative designs highlight the importance of tacit knowledge (Polanyi, 1966), physical constraints, and spatially distributed workflows. In sum, we aim to shift the focus of AI for science from automating reasoning and computational work at the desk to infrastructure that sustains the embodied physical conditions upon which creative reasoning and discovery depend. Acknowledgements.This work was supported by the Alfred P. Sloan Foundation grant number G-2024-22587. We are grateful to the scientists and scientists-in-training who shared their time, trust, and space with us. References P. Agre (1997) Computation and human experience. Cambridge University Press. Cited by: §7.3. S. Amershi, D. Weld, M. Vorvoreanu, A. Fourney, B. Nushi, P. Collisson, J. Suh, S. Iqbal, P. N. Bennett, K. Inkpen, J. Teevan, R. Kikin-Gil, and E. Horvitz (2019) Guidelines for human-ai interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI â19, New York, NY, USA, p. 1â13. External Links: ISBN 9781450359702, Link, Document Cited by: §2.3. C. Arnold (2024) Can robotic lab assistants speed up your work?. Nature OUTLOOK. Cited by: §2.2. J. Auger (2013) Speculative design: crafting the speculation. Digital Creativity 24 (1), p. 11â35. Cited by: §3.2. J. Baek, S. K. Jauhar, S. Cucerzan, and S. J. Hwang (2024) Researchagent: iterative research idea generation over scientific literature with large language models. arXiv preprint arXiv:2404.07738. Cited by: §2.2. V. Braun and V. Clarke (2006) Using thematic analysis in psychology. Qualitative research in psychology 3 (2), p. 77â101. Cited by: §3.3. E. Brynjolfsson, B. Chandar, and R. Chen (2025) Canaries in the coal mine? six facts about the recent employment effects of artificial intelligence. Working Paper Technical Report Stanford Digital Economy Lab Working Paper, Stanford Digital Economy Lab. External Links: Link Cited by: §7.1. J. C. Carver, R. P. Kendall, S. E. Squires, and D. E. Post (2007) Software development environments for scientific and engineering software: a series of case studies. In 29th International Conference on Software Engineering (ICSEâ07), Vol. , p. 550â559. External Links: Document Cited by: §2.1. K. K. Cetina (1999) Epistemic cultures: how the sciences make knowledge. Harvard University Press. Cited by: §1, §2.1. M. Chang, S. Druga, A. Fiannaca, P. Vergani, C. Kulkarni, C. Cai, and M. Terry (2023) The prompt artists. External Links: Link Cited by: §2.3. H. M. Collins (1974) The tea set: tacit knowledge and scientific networks. Science studies 4 (2), p. 165â185. Cited by: §2.1. D. De Cremer, N. M. Bianzino, and B. Falk (2023) How generative ai could disrupt creative work. Harvard Business Review 13, p. 13. Cited by: §7.1. A. W. Ding and S. Li (2025) Generative ai lacks the human creativity to achieve scientific discovery from scratch. Scientific Reports 15 (1), p. 9587. Cited by: §3.4. [14] (2025)Emic and etic(Website) Note: Accessed: 2025-09-01 External Links: Link Cited by: §3.2, §3.2. R. Fok, N. Lipka, T. Sun, and A. F. Siu (2024) Marco: supporting business document workflows via collection-centric information foraging with large language models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI â24, New York, NY, USA. External Links: ISBN 9798400703300, Link, Document Cited by: §2.3. [16] (2025)GOFAI(Website) Note: Accessed: 2025-09-01 External Links: Link Cited by: §3.2. J. Gori, H. L. Han, and M. Beaudouin-Lafon (2020) FileWeaver: flexible file management with automatic dependency tracking. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology, p. 22â34. Cited by: §2.1. J. Gottweis, W. Weng, A. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger, K. Rong, R. Tanno, et al. (2025) Towards an ai co-scientist. arXiv preprint arXiv:2502.18864. Cited by: §1, §7.3. A. Guo, S. Sathyanarayanan, L. Wang, J. Heer, and A. X. Zhang (2025) From pen to prompt: how creative writers integrate ai into their writing practice. In Proceedings of the 2025 Conference on Creativity and Cognition, C&C â25, New York, NY, USA, p. 527â545. External Links: ISBN 9798400712890, Link, Document Cited by: §1, §2.3, §4.1. J. E. Hannay, C. MacLeod, J. Singer, H. P. Langtangen, D. Pfahl, and G. Wilson (2009) How do scientists develop and use scientific software?. In 2009 ICSE Workshop on Software Engineering for Computational Science and Engineering, Vol. , p. 1â8. External Links: Document Cited by: §2.1. J. Hoffman (2022) Speculative futures: design approaches to navigate change, foster resilience, and co-create the cities we need. North Atlantic Books. Cited by: §3.2. R. Huang, S. Ravi, M. He, B. Tian, S. Lerner, and M. Coblenz (2025) How scientists use jupyter notebooks: goals, quality attributes, and opportunities. arXiv preprint arXiv:2503.12309. Cited by: §2.1. D. Ippolito, A. Yuan, A. Coenen, and S. Burnam (2022) Creative writing with an ai-powered writing assistant: perspectives from professional writers. External Links: 2211.05030, Link Cited by: §2.3. [24] (2025) Is it ok for ai to write science papers? nature survey shows researchers are split. Nature. Note: Accessed: 2025-07-31 External Links: Document, Link Cited by: §3.4. K. M. Jablonka, Q. Ai, A. Al-Feghali, S. Badhwar, J. D. Bocarsly, A. M. Bran, S. Bringuier, L. C. Brinson, K. Choudhary, D. Circi, et al. (2023) 14 examples of how llms can transform materials science and chemistry: a reflection on a large language model hackathon. Digital discovery 2 (5), p. 1233â1250. Cited by: §2.2, §2.2. J. Jumper, R. Evans, âŚ, et al. (2021) Highly accurate protein structure prediction with alphafold. Nature 596 (7873), p. 583â589. Cited by: §1, §2.2, §7.3. S. Kernan Freire, M. Foosherian, C. Wang, and E. Niforatos (2023) Harnessing large language models for cognitive assistants in factories. In In ACM conference on Conversational User Interfaces (CUI â23),, p. 1â6. Cited by: §2.3. R. Khojah, M. Mohamad, P. Leitner, and F. G. de Oliveira Neto (2024) Beyond code generation: an observational study of chatgpt usage in software engineering practice. Proc. ACM Softw. Eng. 1 (FSE). External Links: Link, Document Cited by: §2.3, §4.1. I. K. Koukouvelas, G. Pappas, A. Kyriou, D. N. Apostolopoulos, and K. G. Nikolakopoulos (2023) UAV photogrammetry and in situ measurements for mass wasting monitoring in a steep cliff. In Earth Resources and Environmental Remote Sensing/GIS Applications XIV, Vol. 12734, p. 311â320. Cited by: §2.2. P. Laban, J. Vig, M. Hearst, C. Xiong, and C. Wu (2024) Beyond the chat: executable and verifiable text-editing with llms. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, p. 1â23. Cited by: §1. R. Lam, A. Sanchez-Gonzalez, M. Willson, P. Wirnsberger, M. Fortunato, F. Alet, S. Ravuri, T. Ewalds, Z. Eaton-Rosen, W. Hu, A. Merose, S. Hoyer, G. Holland, O. Vinyals, J. Stott, A. Pritzel, S. Mohamed, and P. Battaglia (2023) Learning skillful medium-range global weather forecasting. Science 382 (6677), p. 1416â1421. External Links: Document, Link Cited by: §1. B. Latour and S. Woolgar (1986) Laboratory life: the construction of scientific facts. Princeton University Press, Princeton, N.J.. External Links: ISBN 9780691028323 Cited by: §1, §2.1. Y. S. Lincoln (1985) Naturalistic inquiry. Vol. 75, Sage. Cited by: §3.2, §3.2. [34] (2025)List of research universities in the United States(Website) Note: Accessed: 2025-09-01 External Links: Link Cited by: §3.1, Table 1. M. X. Liu, A. Sarkar, C. Negreanu, B. Zorn, J. Williams, N. Toronto, and A. D. Gordon (2023) âWhat it wants me to sayâ: bridging the abstraction gap between end-user programmers and code-generating large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI â23, New York, NY, USA. External Links: ISBN 9781450394215, Link, Document Cited by: §1, §2.3, §4.1. Y. Liu, S. Chen, H. Cheng, M. Yu, X. Ran, A. Mo, Y. Tang, and Y. Huang (2024) How ai processing delays foster creativity: exploring research question co-creation with an llm-based agent. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI â24, New York, NY, USA. External Links: ISBN 9798400703300, Link, Document Cited by: §2.2. Y. Liu, P. Sharma, M. Oswal, H. Xia, and Y. Huang (2025) PersonaFlow: designing llm-simulated expert perspectives for enhanced research ideation. In Proceedings of the 2025 ACM Designing Interactive Systems Conference, DIS â25, New York, NY, USA, p. 506â534. External Links: ISBN 9798400714856, Link, Document Cited by: §2.2. C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha (2024) The ai scientist: towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292. Cited by: §2.2. M. E. Lynch (1979) Art and artifact in laboratory science: a study of shop work and shop talk in a research laboratory.. University of California, Irvine. Cited by: §2.1. Makery (2019) Andy quitmeyerâs wonderfully weird world of digital naturalism. Note: Accessed: 2025-08-05 External Links: Link Cited by: §7.2. N. Maslej, L. Fattorini, R. Perrault, Y. Gil, V. Parli, N. Kariuki, E. Capstick, A. Reuel, E. Brynjolfsson, J. Etchemendy, et al. (2025) Artificial intelligence index report 2025. arXiv preprint arXiv:2504.07139. Cited by: §7.1. I. Medicine (2024) Insilico medicine advances ai-designed drug to phase i clinical trials. Note: Accessed: 2025-09-04 External Links: Link Cited by: §1. M. R. Morris (2023) Scientistsâ perspectives on the potential for generative ai in their fields. arXiv preprint arXiv:2304.01420. Cited by: §2.2, §2.2, §3.4, §3.4, §4.1, §7.1. [44] (2024) Nobel prize in physics 2024: press release. Note: Accessed: 2025-09-01 External Links: Link Cited by: §1. G. OâBrien (2025) How scientists use large language models to program. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §2.2, §4.1. O. OâDonoghue, A. Shtedritski, J. Ginger, R. Abboud, A. E. Ghareeb, J. Booth, and S. G. Rodriques (2023) BioPlanner: automatic evaluation of llms on protocol planning in biology. arXiv preprint arXiv:2310.10632. Cited by: §2.2. [47] Outreach robotics. Note: Accessed: 2025-02-02 External Links: Link Cited by: §2.2. M. Polanyi (1966) The tacit dimension. University of Chicago Press. Cited by: §1, §5.3, §8, footnote 2. J. Prather, J. Leinonen, N. Kiesler, J. Gorson Benario, S. Lau, S. MacNeil, N. Norouzi, S. Opel, V. Pettit, L. Porter, et al. (2025) Beyond the hype: a comprehensive review of current trends in generative ai research, teaching practices, and tools. 2024 Working Group Reports on Innovation and Technology in Computer Science Education, p. 300â338. Cited by: §2.3. S. Pringle, M. Dallimer, M. A. Goddard, L. K. Le Goff, E. Hart, S. J. Langdale, J. C. Fisher, S. Abad, M. Ancrenaz, F. Angeoletto, et al. (2025) Opportunities and challenges for monitoring terrestrial biodiversity in the robotics age. Nature Ecology & Evolution, p. 1â12. Cited by: §2.2. K. Pu, D. Lazaro, I. Arawjo, H. Xia, Z. Xiao, T. Grossman, and Y. Chen (2025) Assistance or disruption? exploring and evaluating the design and trade-offs of proactive ai programming support. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, p. 1â21. Cited by: §2.3, §4.1. A. Quitmeyer (2024) Growing feral with contextual crafting: graphic essay. Note: Creative Commons Attribution-NoDerivs 4.0 International License External Links: Link Cited by: §7.2. M. C. Ramos, C. J. Collison, and A. D. White (2025) A review of large language models and autonomous agents in chemistry. Chemical science. Cited by: §2.2. J. T. Rapp, B. J. Bremer, and P. A. Romero (2024) Self-driving laboratories to autonomously navigate the protein fitness landscape. Nature chemical engineering 1 (1), p. 97â107. Cited by: §2.2. C. K. Reddy and P. Shojaee (2025) Towards scientific discovery with generative ai: progress, opportunities, and challenges. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 28601â28609. Cited by: §2.2. S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, Z. Liu, and E. Barsoum (2025) Agent laboratory: using llm agents as research assistants. arXiv preprint arXiv:2501.04227. Cited by: §2.2. E. Seifert, S. Seifert, H. Vogt, D. Drew, J. Van Aardt, A. Kunneke, and T. Seifert (2019) Influence of drone altitude, image overlap, and optical sensor resolution on multi-view reconstruction of forest images. Remote sensing 11 (10), p. 1252. Cited by: §2.2. N. Sherman (2025) Amazon boss says ai will replace jobs at tech giant. Note: BBC News External Links: Link Cited by: §7.1. Y. Shi, T. Gao, X. Jiao, and N. Cao (2023) Understanding design collaboration between designers and artificial intelligence: a systematic literature review. Proceedings of the ACM on Human-Computer Interaction 7 (CSCW2), p. 1â35. Cited by: §1, §2.3. C. Strong, S. Jones, A. Parker-Wood, A. Holloway, and D. D. Long (2011) Los alamos national laboratory interviews. University of California, Santa Cruz, Tech. Rep. UCSC-SSRC-11-06. Cited by: §2.1. H. Subramonyam, D. Thakkar, A. Ku, J. Dieber, and A. K. Sinha (2025) Prototyping with prompts: emerging approaches and challenges in generative ai design for collaborative software teams. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §2.3. L. A. Suchman (1987) Plans and situated actions: the problem of human-machine communication. Cambridge University Press. Cited by: §7.3. M. Suh, E. Youngblom, M. Terry, and C. J. Cai (2021) AI as social glue: uncovering the roles of deep generative ai during social music composition. In Proceedings of the 2021 CHI conference on human factors in computing systems, p. 1â11. Cited by: §2.3. S. Suh, M. Chen, B. Min, T. J. Li, and H. Xia (2024) Luminate: structured generation and exploration of design space with large language models for human-ai co-creation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â26. Cited by: §1, §2.3. Y. Tang, M. Ciancia, Z. Wang, and Z. Gao (2024) Whatâs next? exploring utilization, challenges, and future directions of ai-generated image tools in graphic design. External Links: 2406.13436, Link Cited by: §2.3. K. Tomlinson, S. Jaffe, W. Wang, S. Counts, and S. Suri (2025) Working with ai: measuring the occupational implications of generative ai. arXiv preprint arXiv:2507.07935. Cited by: §7.1, §7.1. S. Traweek (2009) Beamtimes and lifetimes. Harvard University Press. Cited by: §5.1. P. Vaithilingam, T. Zhang, and E. L. Glassman (2022) Expectation vs. experience: evaluating the usability of code generation tools powered by large language models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems, CHI EA â22, New York, NY, USA. External Links: ISBN 9781450391566, Link, Document Cited by: §1, §2.3. R. Vanukuru, P. Panda, X. Chen, A. E. Scott, L. Tankelevitch, and S. Rintel (2025) Designing interfaces that support temporal work across meetings with generative ai. In Proceedings of the 2025 ACM Designing Interactive Systems Conference, DIS â25, New York, NY, USA, p. 3600â3620. External Links: ISBN 9798400714856, Link, Document Cited by: §2.3. J. Vertesi (2014) Seamful spaces: heterogeneous infrastructures in interaction. Science, Technology, & Human Values 39 (2), p. 264â284. Cited by: §2.1. K. B. Wagman, M. T. Dearing, and M. Chetty (2025) Generative ai uses and risks for knowledge workers in a science organization. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, p. 1â17. Cited by: §2.2, §2.2, §3.4, §3.4, §3.4, §4.1, §7.1. I. Wallach, M. Dzamba, and A. Heifets (2015) AtomNet: a deep convolutional neural network for bioactivity prediction in structure-based drug discovery. External Links: 1510.02855, Link Cited by: §1. R. Yeh, C. Liao, S. Klemmer, F. Guimbretière, B. Lee, B. Kakaradov, J. Stamberger, and A. Paepcke (2006) ButterflyNet: a mobile capture and access system for field biology research. In Proceedings of the SIGCHI conference on Human Factors in computing systems, p. 571â580. Cited by: §2.1, §7.3. Z. Zheng, K. Ning, Q. Zhong, J. Chen, W. Chen, L. Guo, W. Wang, and Y. Wang (2025) Towards an understanding of large language models in software engineering tasks. Empirical Software Engineering 30 (2), p. 50. Cited by: §1, §2.3. J. Zhu, A. Liapis, S. Risi, R. Bidarra, and G. M. Youngblood (2018) Explainable ai for designers: a human-centered perspective on mixed-initiative co-creation. In 2018 IEEE Conference on Computational Intelligence and Games (CIG), Vol. , p. 1â8. External Links: Document Cited by: §2.3.