Paper deep dive
Quantifying the Expectation-Realisation Gap for Agentic AI Systems
Sebastian Lobentanzer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 2:57:12 PM
Summary
This paper quantifies the 'expectation-realisation gap' in Agentic AI systems, demonstrating that pre-deployment expectations for productivity and accuracy gains systematically exceed post-deployment outcomes. Through a review of controlled trials in software engineering, clinical documentation, and clinical decision support, the authors identify key drivers of this gap, including workflow integration friction, verification burdens, measurement construct mismatches, and treatment effect heterogeneity. The study argues for structured planning frameworks, such as the Agentic Automation Canvas, to explicitly quantify benefits and account for human oversight costs.
Entities (9)
Relation Signals (10)
Agentic AI Systems â exhibits â Expectation-Realisation Gap
confidence 95% · Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes.
GitHub Copilot â caused â 19% increase in completion time
confidence 92% · The measured outcome was a 19% increase in completion timeâa 43 percentage-point calibration error on the time-change scale...
METR â conductedtrialon â GitHub Copilot
confidence 90% · The sharpest illustration of the expectationârealisation gap comes from a randomised controlled trial conducted by METR... on 16 experienced open-source developers... AI assistance would reduce their completion time by 24%. The measured outcome was a 19% increase in completion time...
Verification Burden â contributesto â Expectation-Realisation Gap
confidence 90% · These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not.
Workflow Integration Friction â contributesto â Expectation-Realisation Gap
confidence 90% · These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not.
Epic Sepsis Model â reportedperformance â AUROC 0.76-0.83
confidence 90% · developer-reported model performance (AUROC 0.76â0.83 for Epicâs sepsis model) reflects evaluation choices that can systematically inflate apparent performance...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes. We review controlled trials and independent validations across software engineering, clinical documentation, and clinical decision support to quantify this expectation-realisation gap. In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error. In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note, and one widely deployed tool showed no statistically significant effect. In clinical decision support, externally validated performance falls substantially below developer-reported metrics. These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not. The evidence motivates structured planning frameworks that require explicit, quantified benefit expectations with human oversight costs factored in.
Tags
Links
- Source: https://arxiv.org/abs/2602.20292v2
- Canonical: https://arxiv.org/abs/2602.20292v2
Trouble viewing inline? Open PDF directly â
Full Text
27,952 characters extracted from source content.
Expand or collapse full text
Quantifying the ExpectationâRealisation Gap for Agentic AI Systems This manuscript (permalink) was automatically generated from slolab/agentic-expectation-realisation-gap@990c846 on February 25, 2026. Authors Sebastian Lobentanzer â https://orcid.org/0000-0003-3399-6695 · slobentanzer Institute of Computational Biology, Computational Health Center, Helmholtz Center, Munich, Germany; German Center for Diabetes Research, Munich, Germany â â Correspondence possible via GitHub Issues or email to Sebastian Lobentanzer <sebastian.lobentanzer@helmholtz-munich.de>. Abstract Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post- deployment outcomes. We review controlled trials and independent validations across software engineering, clinical documentation, and clinical decision support to quantify this expectationâ realisation gap. In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19%âa 43 percentage-point calibration error. In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note, and one widely deployed tool showed no statistically signicant eect. In clinical decision support, externally validated performance falls substantially below developer-reported metrics. These shortfalls are driven by workow integration friction, verication burden, measurement construct mismatches, and systematic variation in who benets and who does not. The evidence motivates structured planning frameworks that require explicit, quantied benet expectations with human oversight costs factored in. Introduction Agentic AI systemsâautonomous software agents that plan, reason, and execute multi-step tasks with limited human oversightâare being adopted across software engineering, clinical medicine, and customer operations with the expectation of transformative productivity gains. Vendor announcements routinely promise multi-minute time savings per encounter, double-digit percentage speedups, or near-expert-level decision accuracy. Procurement and investment decisions follow these expectations, committing substantial resources before deployment-grade evidence is available. Yet, as we explore below, a growing body of controlled trials and independent external validations reveals that realised outcomes frequently fall short of pre-deployment expectations. We term this discrepancy the expectationârealisation gap. It reects systematic patterns in how agentic systems interact with human workows, how performance is measured, and how benets are distributed across user populations. Understanding and quantifying this gap is a prerequisite for responsible deployment. Without structured, quantied expectations that account for real-world integration costs, organisations risk over-investing in systems that deliver marginal gains at best or impose net costs at worst. This review synthesises the strongest available empirical evidence for the expectationârealisation gap across three domainsâsoftware engineering, clinical documentation, and clinical decision supportâ identies the mechanistic drivers of shortfalls, and argues that structured planning frameworks with explicit benet quantication are necessary to close the gap between aspiration and reality. Evidence from controlled trials Software engineering copilots The sharpest illustration of the expectationârealisation gap comes from a randomised controlled trial conducted by METR (Model Evaluation & Threat Research) on 16 experienced open-source developers working in their own mature repositories across 246 real tasks [1]. Before each task, participants forecast that AI assistance would reduce their completion time by 24%. The measured outcome was a 19% increase in completion timeâa 43 percentage-point calibration error on the time-change scale, and a complete reversal in direction. Tasks took approximately 56% longer than developers expected. Intriguingly, AI-assisted developers also estimated 20% reduction in completion time after they had performed the task; the opposite of what had happened. Economics experts (N=34) and machine learning experts (N=54) overestimated the expected speedup even more dramatically, with 39% and 38%, respectively. This result contrasts instructively with a controlled trial of GitHub Copilot on developers recruited via Upwork, where treated participants completed a standardised, self-contained programming task 56% faster (95% CI 21â89%) [2]. In that experiment, participantsâ self-estimated productivity gains averaged approximately 35%, meaning they underestimated the realised speedup. Of note, the task here was the same for all participants; the development of an HTTP server in JavaScript. The divergence between these two trials can be explained by task complexity. On constrained, well-dened tasks, AI coding tools can exceed expectations, particularly for less experienced professionals; the study notes that the benet was greater for less experienced developers [2]. In contrast, in high-context, real- world repositories, the same class of tools can impose net costs that even senior developers fail to anticipate. These ndings are the rst instance of heterogeneous treatment eects of agentic AI (see the section on heterogeneity). Field experiments at Microsoft and Accenture provide complementary information: developers assigned to Copilot completed 12.9â21.8% more pull requests per week at Microsoft and 7.5â8.7% more at Accenture, though the authors emphasise imprecision and threats to inference including low compliance and organisational confounds [3]. However, these throughput metrics do not account for code quality: independent security analyses nd that 32.8% of Python and 24.5% of JavaScript snippets generated by Copilot are agged with security issues [4], and Copilot can replicate known- vulnerable code patterns at rates around 33%, although it improves on unassisted human developers [5]. Productivity gains that increase time for human oversight (such as review, remediation, and incident risk) are not net gains. Clinical documentation agents Ambient AI scribesâsystems that listen to clinical encounters and generate draft documentationâ represent one of the most actively deployed categories of agentic AI in healthcare. Vendor procurement narratives frequently frame benets in terms of âminutes saved per encounterâ; for instance, Microsoft publicised â5 minutes saved per clinician per encounter on averageâ for its DAX Copilot product [6]. Evidence from controlled trials partly contradicts these claims. A randomised controlled trial at UCLA across 238 physicians in 14 specialties compared two commercial ambient scribe tools (DAX and Nabla) against usual care, with approximately 24 000 encounters per arm [7]. Nabla reduced time-in- note by 9.5% relative to control (95% CI â17.2 to â1.8; P=0.02), while DAX showed no statistically signicant eect (â1.7%, 95% CI â9.4 to +5.9; P=0.66). Partial adoption was a key contextual factor: the tools were used in only approximately 30â34% of visits, and roughly 15% of treatment-group physicians never used their assigned scribe at all. Clinicians did not strongly endorse that generated notes were âat least as good as my ownâ; and âoccasionalâ clinically signicant inaccuracies potentially led to a signicant loss of time due to oversight activities [7]. A peer-matched cohort study of DAX in an integrated delivery system (99 providers, 12 specialties) found documentation EHR time fell from 5.3 to 4.5 minutes per patientâa saving of approximately 46 secondsâwhile after-hours EHR time worsened signicantly, suggesting time-shifting rather than uniform savings [8]. A pre/post study of the Abridge ambient listening tool across 332 physicians conrmed sub-minute savings: mean time in notes per note fell from 5.1 to 4.2 minutes (dierence 57 seconds, 95% CI 29â85) [9]. In this study, adoption increased from 15% to 50% of physicians in the study span of 8 weeks; however, the total number of notes created by the scribe only increased from 5% to 15% at the end of the study period. The perceptionâreality mismatch that was found in software engineering studies was replicated in a study of 252 physicians: 86.5% perceived that their documentation time had decreased, yet there was no overall association between perceived reductions and objectively measured time changes (OR 0.975, P=0.144) [10]. The objective eect was modest: each 10 percentage-point increase in AI scribe usage was associated with approximately 30 seconds lower documentation time per scheduled hour (P<0.001). Clinical decision support Where the preceding sections documented gaps in time-savings expectations, clinical decision support reveals equally stark expectationârealisation gaps in accuracy and qualityâthe benet dimensions most directly linked to patient outcomes. The Epic Sepsis Model, widely implemented across US hospitals, was externally validated in a large academic health system (38 455 hospitalisations) with an area under the receiver operating characteristic curve (AUROC) of 0.63 (95% CI 0.62â0.64), while Epic previously reported AUROC values of 0.76â0.83 [11]. At an operational alert threshold, the model achieved only 33% sensitivity (missing two thirds of septic patients), raising questions about clinical utility at scale. In oncology decision support, IBM publicised concordance rates as high as 96% for Watson for Oncology in lung cancer cases relative to a multidisciplinary tumour board [12]. A subsequent peer- reviewed retrospective study in Korea found strict concordance of 48.9% for colon cancer, with âacceptableâ concordance of 65.8% and strong heterogeneity by patient age (concordance dropping to approximately 20% among patients aged 70 and older) [13]. This discrepancy reected both denition dependenceâconcordance rose substantially when âfor considerationâ was treated as concordantâ and local constraint mismatches in guidelines, reimbursement, and patient demographics that prevent cross-site transferability. In both cases, the expectationârealisation gap follows the same pattern as for time savings: internally reported metrics set expectations that external, deployment- grade evaluation cannot reproduce. Why expectations overshoot The empirical evidence points to three recurrent mechanistic drivers that explain why expectations systematically exceed realised outcomes, consistent with planning fallacies widely reported in the psychology literature [14,15]. Workow integration friction and partial adoption. Agentic AI systems do not operate in isolation; they must integrate into existing workows, tools, and team practices. Clinical scribe evaluations repeatedly show partial adoptionâtools used in a minority of encounters, with non-trivial drop-o over time [7,8]. Even when per-use eects are real, intention-to-treat estimates are attenuated by low compliance, and the practical benet to an organisation depends on the adoption rate actually achieved, not the rate assumed during procurement. This is not always a simple, temporary onboarding issue; the UCLA RCTâs 30â34% utilisation rate was observed over the full study period [7]. Verication and review burden. Agentic systems generate outputs that require human verication, and this verication cost is rarely accounted for in pre-deployment projections. In the METR software engineering trial, the net slowdown occurred because the time spent reviewing, debugging, and integrating AI-generated code exceeded the time saved in initial generation [1]. In clinical documentation, neutral ratings on note quality and âoccasionalâ clinically signicant inaccuracies indicate non-trivial editing and review work that partially or fully osets time-in-note reductions [7]. The DAX cohortâs simultaneous reduction in documentation time and increase in after-hours EHR time concretely shows a shift of eort, as opposed to alleviation [8]. Measurement construct mismatch. Pre-deployment expectations are often framed in metrics that do not correspond to what deployment-grade evaluations actually measure. Vendor claims of âminutes saved per encounterâ refer to broader workow impacts, while trial outcomes measure âtime-in-noteââone slice of documentation burden [7]. Developer-reported model performance (AUROC 0.76â0.83 for Epicâs sepsis model) reects evaluation choices that can systematically inate apparent performance relative to development goals [11]. The gap between lab-task performance and eld performance in software copilots is a measurement construct problem at its core: bounded tasks estimate tool capability under low-context load, while eld trials estimate net productivity under realistic verication and integration costs [1,2]. These three drivers interact and compound over time, fueling inated expectations. Workow integration depends on existing competence: experienced professionals can restructure their work around agentic tools while less experienced users lack the mental models to do so eectively. Verication burden scales inversely with expertise: a senior developer can spot a awed code suggestion quickly, whereas a junior developer may accept it uncritically or spend disproportionate time reviewing it. And measurement construct mismatches extend to the time horizon of measurement itself: short-term productivity metrics cannot capture costs that materialise only over longer periods. For instance, the level of knowledge in the workforce is a slowly developing phenomenon relative to the speed at which agentic technologies are introduced. A randomised controlled trial in higher education found that students who used ChatGPT as a study aid scored signicantly lower on a surprise retention test 45 days later (57.5% vs 68.5%; Cohenâs d = 0.68) [16], suggesting that cognitive ooading can trade immediate task completion for degraded durable learning. A parallel study in software engineering found that AI use impaired the usersâ conceptual understanding, code reading, and debugging abilities, without delivering signicant eciency gains on average [17]. Ignoring these interactions when planning the deployment of an agentic system will systematically overestimate its net benets. Heterogeneity as the default Across every domain reviewed, treatment eects are not uniform. They are systematically moderated by baseline user eciency, task complexity, and local context. This implies that development of systems useful in practice requires careful planning that respects treatment heterogeneity. In clinical documentation, objective time savings from AI scribes concentrate among physicians with higher baseline documentation ineciency; ecient documenters derive minimal benet [10]. In customer support, a eld study of 5,172 agents found an average 15% productivity increase from a generative AI assistant, but gains were heavily concentrated among less experienced and lower-skilled workers, while the most experienced agents saw smaller gains and occasional quality declines [18]. This heterogeneity is not only observed inter-individually but also at the intra-individual level; given a single agent, gains from AI adoption are larger for relatively rare tasks, where human users have less baseline training and experience [18]. In software engineering, the METR trial specically selected experienced developers working in familiar repositoriesâprecisely the population most likely to have optimised their workows alreadyâand this is the population that was slowed [1]. In summary, there is currently no stable, globally positive treatment eect for agentic AI. Average headline gures (whether from vendors, lab trials, or even well-designed eld studies) will systematically misrepresent the benet realised by any specic user, team, or organisation. Planning that relies on average expected gains without modelling who benets and who does not will over- invest in low-yield deployments and under-invest in targeted high-yield ones. Implications for structured planning The evidence reviewed here converges on a clear conclusion: pre-deployment expectations for agentic AI systems are poorly calibrated, and the resulting expectationârealisation gap is large enough to undermine investment decisions, deployment strategies, and trust. This is not an argument against agentic AIâthe evidence also shows that real gains exist in specic contexts and for specic user populations. It is an argument for structured planning that takes the gap seriously. Several design principles follow directly from the empirical patterns. First, benet expectations must be explicit and quantied across all relevant dimensions, rather than framed as vague promises of eciency. The contrast between â5 minutes saved per encounterâ marketing and sub-minute measured reductions illustrates what happens when time-based expectations lack precision [7]; the gap between developer-reported AUROC (0.76â0.83) and externally validated AUROC (0.63) for the Epic Sepsis Model shows the same pattern for accuracy metrics [11]. Second, expectations should capture dual perspectivesâwhat users expect to gain and what developers assess as technically feasible. Miscalibration occurs on both sides: developers overshoot in internal validation; users overshoot in self-forecasts. The estimateâreality mismatch in the METR study, which persisted even after implementation, illustrates the cognitive biases at play [1]. Third, human oversight costs must be deducted from projected benets. Every controlled trial reviewed here shows that verication, review, and cleanup absorb a substantial fraction of the gross time savings; ignoring this yields unrealistic net benet estimates. Fourth, outcome metrics must link back to initial expectations in the same units and at the same level of granularity, enabling direct comparison rather than post hoc rationalisation. Ideally, plans are formalised early, versioned, and archived, in order to facilitate these later comparisons. Fifth, heterogeneity should be modelled explicitly by specifying which user populations and task types are expected to benet, rather than assuming uniform eects. One way to operationalise these principles is the Agentic Automation Canvas (AAC), a structured framework for designing, governing, and documenting agentic automation projects [19]. The canvas captures user expectations as quantied benet metrics across ve dimensionsâtime, quality, risk, enablement, and costâwith baseline values, condence levels from both user and developer perspectives, and explicit accounting for human oversight. This multi-dimensional structure reects the evidence reviewed here: the expectationârealisation gap manifests not only in time savings (as in software engineering and clinical documentation) but equally in accuracy and quality metrics (as in clinical decision support), and planning frameworks must accommodate all of these. The canvas formalises the bidirectional contract between stakeholders that the evidence reviewed here shows is necessary: without structured mechanisms for surfacing and testing expectations, the gap between aspiration and reality will persist. Conclusion The expectationârealisation gap in agentic AI systems is empirically documented, directionally consistent, and mechanistically explainable. Across software engineering, clinical documentation, and clinical decision support, pre-deployment expectations systematically overestimate realised benets in deployment settings, regardless of whether they hail from user forecasts, vendor claims, or developer-reported metrics. The drivers are complex, but not mysterious: workow integration friction, verication burden, measurement construct mismatches, and treatment eect heterogeneity are observable, predictable, and addressable in principle. Closing the expectationârealisation gap requires moving from ad hoc expectation-setting to structured, quantied planning that accounts for real-world integration costs, models heterogeneity across user populations, and links outcome measurement directly to initial benet projections. It also requires a mature interdisciplinary approach; mismatch from psychological bias cannot be countered by computer science methodology, and building better AI models will not solve all socio-technical problems in deployment and adoption. The alternative to closing the gap is to continue relying on articial benchmark results, marketing claims, and intuitive forecasts. In all likelihood, this will perpetuate a cycle of over-promise and under-delivery that erodes trust in systems that, when properly targeted and governed, can deliver genuine value. Glossary of experimental terms This review draws on evidence from controlled experiments across multiple domains. The following terms, standard in medicine, economics, and the social sciences, are used throughout. Randomised controlled trial (RCT) An experimental design in which participants are randomly assigned to either a treatment group or a control group. Random assignment ensures that observed dierences in outcomes can be attributed to the intervention rather than to pre-existing dierences between groups. In this review, RCTs include trials of AI coding assistants, ambient clinical scribes, and clinical decision-support models. Treatment and intervention The treatment (or intervention) is the condition being evaluated in an experimentâfor example, giving developers access to an AI coding assistant or equipping physicians with an ambient scribe. Participants who receive the treatment are referred to as treated participants or the treatment group. Control The control condition is the comparison group that does not receive the intervention. Control participants continue with their usual workow, providing a baseline against which the treatmentâs eect is measured. Treatment eect The treatment eect is the measured dierence in outcomes between the treatment group and the control group. For example, if treated developers complete tasks 20% faster than control developers, the treatment eect is a 20% reduction in completion time. A treatment eect can be positive (the intervention helps), negative (it hurts), or null (no detectable dierence). Intention-to-treat analysis Intention-to-treat (ITT) analysis includes all participants as originally assignedâwhether or not they actually used the tool. This preserves the validity of randomisation and reects real-world conditions, where not everyone who is oered a tool adopts it. ITT estimates are typically smaller than per-use estimates because non-adopters dilute the measured eect. Heterogeneous treatment eects Heterogeneous treatment eects means that the size (or direction) of the treatment eect varies across subgroups. For instance, less experienced developers may benet substantially from AI assistance while expert developers see no gain or a net slowdown. Recognising heterogeneity is critical for deployment planning: an average treatment eect can mask the fact that some users benet greatly while others are harmed. 1. 2. 3. 4. 5. 6. 7. 8. 9. 10. References Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity Joel Becker, Nate Rush, Elizabeth Barnes, David Rein arXiv (2025) https://doi.org/g9xnj6 DOI: 10.48550/arxiv.2507.09089 The Impact of AI on Developer Productivity: Evidence from GitHub Copilot Sida Peng, Eirini Kalliamvakou, Peter Cihon, Mert Demirer arXiv (2023) https://doi.org/hbqdb5 DOI: 10.48550/arxiv.2302.06590 The impact of generative AI on software developer productivity: evidence from two eld experiments Kevin Cui, Deepak Paramanand, Robert Sloyan MIT GenAI Impact (2025) https://mit-genai.pubpub.org/pub/v5iixksv Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical Study Yujia Fu, Peng Liang, Amjed Tahir, Zengyang Li, Mojtaba Shahin, Jiaxin Yu, Jinfu Chen arXiv (2023) https://doi.org/hbqdb7 DOI: 10.48550/arxiv.2310.02059 Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code? Owura Asare, Meiyappan Nagappan, N Asokan arXiv (2022) https://doi.org/hbqdb4 DOI: 10.48550/arxiv.2204.04741 DAX Copilot: new customization options and AI capabilities for even greater productivity Microsoft Microsoft Industry Blogs (2024-08-08) https://w.microsoft.com/en- us/industry/blog/healthcare/2024/08/08/dax-copilot-new-customization-options-and-ai- capabilities-for-even-greater-productivity/ Ambient AI Scribes in Clinical Practice: A Randomized Trial Paul J Lukac, William Turner, Sitaram Vangala, Aaron T Chin, Joshua Khalili, Ya-Chen Tina Shih, Catherine Sarkisian, Eric M Cheng, John N Ma NEJM AI (2025-12) https://w.ncbi.nlm.nih.gov/pmc/articles/PMC12768499/ DOI: 10.1056/aioa2501000 · PMID: 41497288 · PMCID: PMC12768499 The impact of nuance DAX ambient listening AI documentation: a cohort study Tyler Haberle, Courtney Cleveland, Greg L Snow, Chris Barber, Nikki Stookey, Cari Thornock, Laurie Younger, Buzzy Mullahkhel, Diego Ize-Ludlow Journal of the American Medical Informatics Association : JAMIA (2024-04-03) https://w.ncbi.nlm.nih.gov/pmc/articles/PMC10990544/ DOI: 10.1093/jamia/ocae022 · PMID: 38345343 · PMCID: PMC10990544 Ambient listening implementation in primary care and changes in electronic health record documentation metrics: Pre-post study of an ambient listening tool Frederick North, Marc R Matthews, Asif Iqbal, Jason A Post, Jon O Ebbert Digital health (2025-11-26) https://w.ncbi.nlm.nih.gov/pmc/articles/PMC12657781/ DOI: 10.1177/20552076251403211 · PMID: 41323090 · PMCID: PMC12657781 Subjective and objective impacts of ambulatory AI scribes. 11. 12. 13. 14. 15. 16. 17. 18. 19. Julia Adler-Milstein, Orianna DeMasi, Hossein Soleimani, Sarah Beck, Maria E Byron, Aris Oates, Robert Thombley, Jinoos Yazdany, Sara G Murray The American journal of managed care (2026-01) https://w.ncbi.nlm.nih.gov/pubmed/41592210 DOI: 10.37765/ajmc.2026.89869 · PMID: 41592210 The Epic Sepsis Model Falls ShortâThe Importance of External Validation Anand R Habib, Anthony L Lin, Richard W Grant JAMA Internal Medicine (2021-08-01) https://doi.org/g99mhd DOI: 10.1001/jamainternmed.2021.3333 · PMID: 34152360 At ASCO 2017, clinicians present new evidence about Watson cognitive technology and cancer care IBM IBM Newsroom (2017-06-01) https://uk.newsroom.ibm.com/2017-06-01-At-ASCO-2017- Clinicians-Present-New-Evidence-about-Watson-Cognitive-Technology-and-Cancer-Care Assessing Concordance With Watson for Oncology, a Cognitive Computing Decision Support System for Colon Cancer Treatment in Korea. Won-Suk Lee, Sung Min Ahn, Jun-Won Chung, Kyoung Oh Kim, Kwang An Kwon, Yoonjae Kim, Sunjin Sym, Dongbok Shin, Inkeun Park, Uhn Lee, Jeong-Heum Baek JCO clinical cancer informatics (2018-12) https://w.ncbi.nlm.nih.gov/pubmed/30652564 DOI: 10.1200/cci.17.00109 · PMID: 30652564 On the psychology of prediction. Daniel Kahneman, Amos Tversky Psychological Review (1973-07) https://doi.org/dc7c3k DOI: 10.1037/h0034747 The Planning Fallacy Roger Buehler, Dale Grin, Johanna Peetz Advances in Experimental Social Psychology (2010) https://doi.org/bkzd2w DOI: 10.1016/s0065-2601(10)43001-4 The hidden cost of AI: ChatGPT use impairs long-term knowledge retention in university students Muhammad Farrukh Shahzad International Journal of Educational Research Open (2025) https://w.sciencedirect.com/science/article/pii/S2590291125010186 How AI Impacts Skill Formation Judy Hanwen Shen, Alex Tamkin arXiv (2026) https://doi.org/hbqdbj DOI: 10.48550/arxiv.2601.20245 Generative AI at Work Erik Brynjolfsson, Danielle Li, Lindsey Raymond arXiv (2023) https://doi.org/hbqdb6 DOI: 10.48550/arxiv.2304.11771 The Agentic Automation Canvas: a structured framework for agentic AI project design Sebastian Lobentanzer arXiv (2026) https://doi.org/hbqdbk DOI: 10.48550/arxiv.2602.15090