Paper deep dive
Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 Course
Ruiwei Xiao, Runlong Ye, Xinying Hou, Jessica Wen, Harsh Kumar, Michael Liut, John Stamper
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/21/2026, 2:27:47 AM
Summary
This study presents a large-scale randomized controlled trial (N=979) investigating the impact of four instructional conditions based on the ICAP framework on students' prompting literacy in a CS1 course. The conditions varied in cognitive engagement intensity, ranging from passive text reminders to constructive practice with automated feedback. Results indicated that all interventions significantly improved prompting skills, with gains increasing progressively from Condition 1 to Condition 4. Higher learning gains in immediate post-tests predicted higher final exam scores, suggesting that pedagogical prompting serves as a transferable learning strategy. The study concludes that scalable, theory-informed prompting instruction can effectively transform GenAI classroom policies into actionable learning strategies.
Entities (8)
Relation Signals (6)
ICAP Framework ā informs ā Instructional Conditions
confidence 95% Ā· intervention design in our experimental conditions progressed from passive exposure... to constructive practice... Drawing on the ICAP framework
Condition 1 ā is ā Baseline
confidence 95% Ā· Condition 1: Baseline (Figure 2 top-left)
Condition 4 ā involves ā Constructive Practice
confidence 93% Ā· Condition 4: Select-then-Write... This condition combined selection with constructive practice
Pedagogical Prompting Framework ā basedon ā Self-Regulated Learning
confidence 92% Ā· synthesizes Self-Regulated Learning principles [52] with prompt engineering patterns
Condition 4 ā produces ā Highest Learning Gains
confidence 90% Ā· gains increasing progressively from Condition 1 to Condition 4
Prompting Literacy ā predicts ā Final Exam Score
confidence 88% Ā· higher learning gain in immediate post-test predict higher final exam score
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite universal GenAI adoption, students cannot distinguish task performance from actual learning and lack skills to leverage AI for learning, leading to worse exam performance when AI use remains unreflective. Yet few interventions teaching students to prompt AI as a tutor rather than solution provider have been validated at scale through randomized controlled trials (RCTs). To bridge this gap, we conducted a semester-long RCT (N=979) with four ICAP framework-based instructional conditions varying in engagement intensity with a pre-test, immediate and delayed post-test and surveys. Mixed methods analysis results showed: (1) All conditions significantly improved prompting skills, with gains increasing progressively from Condition 1 to Condition 4, validating ICAP's cognitive engagement hierarchy; (2) for students with similar pre-test scores, higher learning gain in immediate post-test predict higher final exam score, though no direct between-group differences emerged; (3) Our interventions are suitable and scalable solutions for diverse educational contexts, resources and learners. Together, this study makes empirical and theoretical contributions: (1) theoretically, we provided one of the first large-scale RCTs examining how cognitive engagement shapes learning in prompting literacy and clarifying the relationship between learning-oriented prompting skills and broader academic performance; (2) empirically, we offered timely design guidance for transforming GenAI classroom policies into scalable, actionable prompting literacy instruction to advance learning in the era of Generative AI.
Tags
Links
- Source: https://arxiv.org/abs/2602.16033v1
- Canonical: https://arxiv.org/abs/2602.16033v1
Trouble viewing inline? Open PDF directly ā
Full Text
60,375 characters extracted from source content.
Expand or collapse full text
Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 Course Ruiwei Xiao Carnegie Mellon University Pittsburgh, PA, United States ruiweix@cs.cmu.edu Runlong Ye University of Toronto Toronto, ON, Canada harryye@cs.toronto.edu Xinying Hou University of Michigan Ann Arbor, MI, United States xyhou@umich.edu Jessica Wen University of Toronto Mississauga Mississauga, ON, Canada jessica.wen@mail.utoronto.ca Harsh Kumar University of Toronto Toronto, ON, Canada harsh@cs.toronto.edu Michael Liut University of Toronto Mississauga Mississauga, ON, Canada michael.liut@utoronto.ca John Stamper Carnegie Mellon University Pittsburgh, PA, United States jstamper@cmu.edu Abstract Despite universal GenAI adoption, students cannot distinguish task performance from actual learning and lack skills to leverage AI for learning, leading to worse exam performance when AI use remains unreflective. Yet few interventions teaching students to prompt AI as a tutor rather than solution provider have been val- idated at scale through randomized controlled trials (RCTs). To bridge this gap, we conducted a semester-long RCT (N=979) with four ICAP framework-based instructional conditions varying in engagement intensity with a pre-test, immediate and delayed post- test and surveys. Mixed methods analysis results showed: (1) All conditions significantly improved prompting skills, with gains in- creasing progressively from Condition 1 to Condition 4, validating ICAPās cognitive engagement hierarchy; (2) for students with sim- ilar pre-test scores, higher learning gain in immediate post-test predict higher final exam score, though no direct between-group differences emerged; (3) Our interventions are suitable and scalable solutions for diverse educational contexts, resources and learners. Together, this study makes empirical and theoretical contributions: (1) theoretically, we provided one of the first large-scale RCTs ex- amining how cognitive engagement shapes learning in prompting literacy and clarifying the relationship between learning-oriented prompting skills and broader academic performance; (2) empiri- cally, we offered timely design guidance for transforming GenAI classroom policies into scalable, actionable prompting literacy in- struction to advance learning in the era of Generative AI. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conferenceā17, Washington, DC, USA Ā© 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X CCS Concepts ⢠Applied computingāComputer-assisted instruction; Inter- active learning environments;⢠Human-centered computing ā Empirical studies in HCI. Keywords AI in Education, Prompting Engineering, Prompting Literacy, Gen- erative AI ACM Reference Format: Ruiwei Xiao, Runlong Ye, Xinying Hou, Jessica Wen, Harsh Kumar, Michael Liut, and John Stamper. 2018. Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 Course. In . ACM, New York, NY, USA, 11 pages. https://doi.org/X.X 1 Introduction By 2025, generative AI (GenAI) had achieved near-universal adop- tion in higher education, with 88% of students reporting regular use in coursework [12]. However, experimental evidence reveals a concerning pattern: answer-seeking forms of AI use led to 17% worse assessment performance compared to no access [2], suggest- ing that typical usage patterns may actively undermine learning. In response, many instructors have adopted policies encouraging stu- dents to use GenAI āfor learning, not solutions,ā yet these policies rarely include concrete guidance on how students should differen- tiate learning from solution seeking, and corresponding prompts to achieve such goal [42]. As a result, a critical pedagogical gap has emerged: students possess powerful tools but lack the prompting strategies needed to deploy them productively. Addressing this gap requires scalable approaches for teaching pedagogically effective prompting. While recent work has begun to explore prompting literacy in educational contexts [48,51], ex- isting efforts often focus on professional prompt engineering or isolated technique demonstrations rather than systematic instruc- tion grounded in learning sciences principles. Moreover, these ap- proaches lack empirical validation in authentic classroom settings, leaving instructors without evidence-based guidance on how to arXiv:2602.16033v1 [cs.HC] 17 Feb 2026 Conferenceā17, July 2017, Washington, DC, USAXiao et al. teach prompting literacy effectively within realistic time and re- source constraints. To address this gap, we adapted the pedagogical prompting framework proposed by Xiao et al. [49], which synthe- sizes Self-Regulated Learning principles [52] with prompt engi- neering patterns (e.g., White[48]), and operationalized it into four scalable instructional conditions varying in cognitive engagement intensity. We deployed these conditions in a large-scale random- ized controlled trial (RCT) in an introductory computer science course (N = 979 students) at a public university in North America. Drawing on the ICAP framework [8], intervention design in our experimental conditions progressed from passive exposure (con- dition 1 or baseline: text reminder) through active engagement (condition 2,: scenario-based reading; condition 3: component selec- tion) to constructive practice (condition 4: select-then-write with automated feedback), allowing systematical examine on the rela- tionship between level of cognitive engagement in the intervention and learning outcomes. As the first RCT study to our knowledge examining pedagogical prompting instruction at scale, this work makes three primary con- tributions.First, we provide one of the first large-scale randomized controlled trials examining how levels of cognitive engagement shape learning in prompting literacy, an ill-defined domain. Second, we clarify the relationship between learning-oriented prompting literacy skills and broader academic performance, suggesting that prompting functions as a transferable learning strategy rather than a short-term content intervention. Lastly, we provide timely de- sign guidance to transform current GenAI classroom policies into scalable, actionable prompting literacy instruction. We address two main research questions (RQs): RQ1: Learning Gain: Does a pedagogical prompting interven- tion improve learning? RQ1a:Prompting Skill: Do students demonstrate improved ped- agogical prompting competencies, as measured by imme- diate and delayed post-tests? RQ1b:Transfer to Final Exam Performance: Do pedagogical prompt- ing skills improve broader computer science learning out- comes, as measured by final exam scores? RQ2: Instructional Methods: Do instructional approaches differ for teaching pedagogical prompting at scale? RQ2a:Learning Gain: Are there differences in learning gains across intervention conditions for teaching pedagogical prompting? RQ2b:Total Time Spent: Are there differences in total time spent across learning intervention conditions? RQ2c:Equity Across Learners: Do individual differences in mind- set, cognitive preferences, or perceived programming abil- ity differentiate learnersā learning gains? RQ2d:Learner Perceptions and Adoption: Are there differences in studentsā perceptions and reported frequency of peda- gogical prompting use across instructional approaches? 2 Related Work 2.1 GenAI Policy in Classroom In the era of generative AI, LLMs are rapidly entering higher educa- tion, reshaping how students study, write, and solve problems, while raising both opportunities and risks for human learning [3,12,15,27,32,33,46,50]. Institutional and instructor responses have often emphasized risk management, including integrity, au- thorship, and assessment validity, alongside broader discussions of opportunities and challenges for education [17,19,32]. These policy-centered responses are necessary, but provide limited guid- ance for how LLM use should be structured to reliably support learning, especially in large-enrollment courses where instructor bandwidth is constrained [17]. Empirical work in computing education and HCI shows that stu- dents already use LLMs for programming tasks such as generating code, debugging, and seeking explanations, often as a substitute for other help channels [10,16,23,36,37,41]. These practices can im- prove task completion, but can also be more likely to shift learners toward answer-seeking behaviors that reduce productive struggle and conceptual engagement, particularly for novices [3,38]. At the same time, custom-made AI assistance that embeds pedagogical value deployed at the course level demonstrates the potential to balance student agency with educator needs, highlighting the im- portance of embedding guardrails and learning-oriented interaction constraints rather than merely providing access [21, 31]. However, most existing responses scaffold on the system-side scaffolding: institutions publish policies and guidelines [17], and courses deploy classroom systems that embed guardrails in a partic- ular interface [21]. These solutions can be locally effective, but they are easy to route around because students retain low-friction access to more capable, general-purpose LLMs through consumer tools with rapidly declining barriers and costs [12,32,50]. This reveals a key gap: we lack scalable, theory-informed learner-facing scaffolds that shape how students use LLMs across use cases outside a con- trolled classroom tool [40]. Our work addresses this gap by treating pedagogical prompting as an instructional design layer embedded within a CS1 curriculum, shifting the focus from temporary system guardrails to a more durable, student-driven AI literacy [25]. 2.2 Prompting Literacy As LLMs become routine study tools, the ability to elicit, constrain and evaluate model outputs has been framed as āprompting literacyā and increasingly discussed as a curricular object [29]. However, em- pirical evidence suggests that prompting is nontrivial for novices. For example, non-expert users frequently fail to anticipate model be- havior, underspecify constraints, and struggle to iterate effectively, often leading to brittle or misleading outputs [51]. In educational settings, this difficulty is compounded by studentsā tendency to treat LLM responses as authoritative, increasing the risk of uncritical acceptance, shallow engagement, and over-reliance [3, 13, 38]. Existing approaches to teaching prompting often emphasize procedural heuristics or catalogs of reusable patterns (e.g., spec- ifying role, constraints, examples) [29,48]. While such guidance can improve output quality, it typically treats prompting as a tool- optimization skill rather than a learning strategy, and it rarely specifies how prompting practices should elicit the kinds of cogni- tive processes that produce durable understanding. In computing education, initial curricular integrations have begun to introduce explicit prompting tasks in introductory programming [22], and broader community syntheses highlight the need to rethink what Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 CourseConferenceā17, July 2017, Washington, DC, USA Figure 1: Experiment Process we teach when generative AI is present [10]. Still, most efforts focus on how to get better answers rather than how to use LLM interaction to produce better learning. This reveals a programmatic gap. Prompting literacy is currently framed as an individual competency, but the practical challenge is a design problem: how to structure LLM-mediated interaction so that large numbers of learners systematically engage in verification, ex- planation, and revision without requiring instructor-by-instructor prompt authoring or close monitoring [32]. We position pedagogi- cal prompting as a scalable scaffold that operationalizes learning objectives in the interaction itself. In particular, we employ Xiao et al. [49]ās learning-sciences-theory-based Pedagogical Prompting as the basis of our design, adapting it to support studentsā itera- tive refinement and evaluation behaviors in authentic classroom workflows. 3 Methods 3.1 Experiment Design After receiving approval from the institutional research ethics board, we conducted a large-scale randomized controlled trial in an intro- ductory computer science (CS1) course at a public North American university. The study followed a longitudinal design spanning the full semester (Figure 1). At the beginning of the semester, students completed a pre-survey collecting demographic information, self- reported AI usage frequency, and baseline psychological measures. The core learning intervention occurred during the āEthical Use of AI for Learning Lab,ā a mandatory weekly lab as part of the course module. During this intervention, students completed a pre-test, underwent the learning intervention according to their randomly assigned condition, and completed an immediate post-test. Finally, to assess long-term retention of prompting skills, a delayed post-test was administered as an extra credit question on the final exam. 3.1.1Pre- Post-Test Design. To assess studentsā proficiency in peda- gogical prompting, we administered three isomorphic assessments: a pre-test, an immediate post-test, and a delayed post-test. In each assessment, students were presented with a programming scenario where a fictional peer faced a specific coding struggle. Students were asked to write a prompt they would input into a Generative AI tool to help the peer learn, rather than simply providing the answer. 3.1.2 Survey Design. The pre-survey included several validated scales to assess learner characteristics that might moderate the interventionās effects, besides 3 prompt writing questions: ⢠Implicit Theories of Intelligence: We utilized the six- item scale from [11] to assess studentsā mindsets. Three items measured entity mindset (e.g., āI have a certain amount of intelligence, and I really cannot do much to change itā) and three measured incremental mindset. Responses were recorded on a 6-point Likert scale (1 = Strongly agree, 6 = Strongly disagree). Incremental items were reverse-scored to calculate a mean score, where higher scores indicate a stronger growth mindset. ⢠Need for Cognition (NFC): We employed the 18-item short form of the Need for Cognition Scale [7] to measure studentsā tendency to engage in and enjoy thinking. ⢠Confidence in CS Tasks: Students reported their self-efficacy regarding specific programming tasks, including code de- scription and code tracing, adapted from validated CS self- efficacy instruments [43]. 3.1.3Intervention Design. To operationalize effective help-seeking with GenAI, we adopted the framework of pedagogical prompt defined by Xiao et al. [49]based on principles of Self-Regulated Learning (SRL) [52] and prompt engineering patterns (e.g., [48]). In the intervention conditions (Condition 2ā4) in this study, students were taught to deconstruct a pedagogical prompt into five essential components: (1)Problem Identification: Effective help-seeking requires the learner to first diagnose their own knowledge gap [35]. We categorized these barriers into five distinct types derived from common CS1 struggles: ⢠Understand the Task Description ⢠Plan the General Idea ⢠Write the Code ⢠Fix a Bug with Error Message ⢠Fix Undesired Output (2) Problem Context: Novices often struggle to provide LLMs with sufficient grounding information, leading to hallucina- tions or generic advice [51]. We trained students to explicitly include code snippets or error logs to ground the AIās re- sponse. (3)Learning Method: To move beyond passive consumption, students must specify a pedagogical strategy (e.g., "Provide a worked example" or "Scaffold with hints"), aligning with the ICAP frameworkās emphasis on constructive engagement [8]. (4) Learner Level/Persona: Following the Persona prompt pat- tern [48], students were taught to contextualize the request for their proficiency (e.g., āBeginner Python Programmerā) to ensure the outputās complexity is appropriate. (5)Guardrails (Constraints): To mitigate the āillusion of com- petenceā where students mistake fluent AI answers for their own understanding [5], the framework requires explicit in- structions on what not to provide (e.g., āDo not provide the full solutionā). Conferenceā17, July 2017, Washington, DC, USAXiao et al. Figure 2: Four learning intervention designs across conditions. Condition 1: Baseline (top-left); Condition 2: Scenario-Based Reading (top-right); Condition 3: Select-Only (bottom-left); Condition 4: Select-then-Write (bottom-right). The instructional delivery varied by condition and was designed according to the ICAP framework [8], which posits that learning outcomes improve as cognitive engagement increases from passive (receiving information) to active (manipulating information) to con- structive (generating new outputs) to interactive (constructive level with turn-takings). This progression was further supported by the doer effect, which shows that completing interactive activities leads to substantially greater learning than passive reading or watching [26,45]. The design of each condition is demonstrated in Figure 2 and briefly described below. Condition 1: Baseline (Figure 2 top-left). Participants received a short text-based reminder emphasizing the ethical use of Generative AI as a tutor. This was supported by specific evidence suggesting that over-reliance on direct code generation hinders learning from Bastani et al. [3]. Condition 2: Scenario-Based Reading (Figure 2 top-right). Stu- dents engaged with four realistic, comic-style narratives depicting fictional students (e.g., āDavidā) encountering the problem types defined in our framework and successfully applying pedagogical prompts to address them. This condition primarily involved receiv- ing and processing worked examples, corresponding to the passive level of the ICAP framework. Condition 3: Select-Only (Figure 2 bottom-left). In this active learning condition, students analyzed the same scenarios but were required to select the most appropriate prompt components. For each step of the framework (e.g., Problem Type, Learning Method), students evaluated three options and selected the one best suited for the scenario. Condition 4: Select-then-Write (Figure 2 bottom-right). This con- dition combined selection with constructive practice by requiring learners to write prompts step by step, ensuring the generation of novel outputs. After selecting the correct components for a scenario, students manually wrote the full prompt. An automated system powered by an LLM (gpt-4o) validated the written prompts against predefined acceptance criteria and provided actionable feedback when critical components (e.g., guardrails) were missing. Iterating on the prompt with AI-generated adaptive feedback could further support multi-round, interactive practice. Notably, Condition 2 is categorized as passive reading not be- cause reading is inherently passive; learners can reach higher ICAP levels during reading (e.g., discussing the material with a partner would constitute interactive engagement). Rather, each condition is designed to specify the lower bound of cognitive engagement (e.g., Condition 4 ensures at least constructive engagement) without con- straining the possible upper bound of learner activity. Grounded in learning sciences literature, including the ICAP framework and the doer effect [26,45], we therefore hypothesize that both engagement and learning gains will progressively increase from Condition 1 to Condition 4. Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 CourseConferenceā17, July 2017, Washington, DC, USA 3.2 Participants and Data Cleaning Process A total ofķ=979 learners enrolled in the course were randomly as- signed to the four conditions. Participants were primarily first-year students from computing-related programs, with most reporting that they either majored in or intended to major in Computer Sci- ence (93%). Regarding gender, 68% identified as male, 30% as female, and 2% as non-binary. Approximately 34% were international stu- dents, with the remainder domestic. We performed rigorous data cleaning to retain only participants who provided valid responses across all study phases, including the pre-test, learning session, immediate post-test, and delayed post-test. This resulted in a final sample of 431 students (44% of the initial enrollment). The final sample distribution across conditions was: Condition 1 (ķ=102), Condition 2 (ķ=112), Condition 3 (ķ= 109), and Condition 4 (ķ= 108). 3.3 Analysis 3.3.1 Quantitative Analysis. We used multiple quantitative meth- ods to address the research questions. To measure learning ef- fectiveness and transfer, we computed descriptive statistics and conducted paired t-tests / Wilcoxon signed-rank tests to evaluate within-condition learning gains (pre-test vs. immediate and delayed post-tests). We then used one-way ANOVA with Tukey post-hoc tests to compare learning outcomes and final exam performance across conditions. To further examine transfer to course perfor- mance, we ran linear regression predicting final exam scores from prompting performance while controlling for pre-test scores. To measure time efficiency, equity, and adoption, we used one-way ANOVA to compare time-on-task, learner perceptions, and usage frequency across conditions. We conducted multiple regression analyses to test whether learner characteristics moderated learning gains (equity) and whether frustration predicted subsequent usage (adoption). Together, these analyses allowed us to evaluate learning gains, retention, transfer, time efficiency, equity, and learner adoption across the four instructional conditions. 3.3.2Qualitative Analysis: Human Grading and Inter-Rater Reliabil- ity. To evaluate the quality of student-generated prompts, we em- ployed a deductive content analysis approach [34] to grade prompts students wrote in pre-test, immediate and delayed post-test. We developed a quantitative coding scheme (rubric) derived directly from the five-component Pedagogical Prompting Framework. The grading process followed three phases: Codebook Development and Training. First, we iteratively devel- oped a comprehensive codebook defining the presence and quality of each component (e.g., Problem ID, Guardrails). Three Research Assistants (RAs) were trained on this rubric using a set of pilot data distinct from the final sample. Reliability Establishment. To establish Inter-Rater Reliability (IRR), the three RAs independently coded a random stratified sample of 600 student prompts (representing 10% of the total 6,000 prompts generated across the pre-test and immediate post-test). We calcu- lated reliability using Cohenās Kappa (ķ ) [9] to account for chance agreement. The team achieved aķ score of 0.86, indicating "almost perfect" agreement according to the benchmarks established by Landis and Koch [28]. Full Coding. Following the validation of the codebook, the re- maining 5,400 prompts were divided among the RAs for indepen- dent grading. To ensure consistency for the longitudinal compari- son, the delayed post-test responses (ķ=440) were graded by a single RA from the original team using the established codebook. 3.3.3Qualitative Analysis: Thematic Coding. To understand learn- ersā perceptions of the intervention in greater depth, we analyzed open-ended responses from the immediate post-survey using the General Inductive Approach [44]. The first author began by con- ducting a close reading of all responses to identify recurring pat- terns. Relevant text segments were then systematically labeled to generate initial codes capturing the nuances of studentsā perspec- tives. These codes were iteratively refined and aggregated into higher-level themes representing overarching patterns in the data. Throughout this process, the research team engaged in regular dis- cussions to validate and refine the emerging thematic structure (see subsubsection 4.2.4). 4 Results 4.1 RQ1 (Learning Outcomes): Pedagogical prompting instruction significantly improved prompting skills but did not transfer directly to final exam grade 4.1.1 RQ1a (Prompting Skill): All conditions (Condition 1-4) pro- duced significant immediate and delayed pedagogical prompting skill gains. As it shown in Table 1, the number of students in each con- dition was comparable, and there were no significant differences in pre-test scores, indicating a similar level of prior knowledge on pedagogical prompting for all 4 groups. All four conditions demonstrated statistically significant learn- ing gains both immediately post-intervention (see Table 1) and six weeks later (see Table 2). Immediate post-test gains increased progressively across conditions (ķ 1 = 0.08,ķ 2 = 0.39,ķ 3 = 0.45,ķ 4 = 0.51), with effect sizes ranging from medium (Condition 1, r = .53) to large (Conditions 2-4, r > .85). This progressive pattern replicated at the delayed post-test (ķ 1 = 0.06,ķ 2 = 0.15,ķ 3 = 0.17,ķ 4 = 0.24), demonstrating durable retention. Wilcoxon signed-rank tests were used for non-normal gain distributions (immediate: Conditions 1, 2, 4), while paired t-tests were employed for normally distributed gains (immediate: Condition 3, t(108) = -20.25; delayed: all condi- tions, Shapiro-Wilk p > .05); all comparisons reached significance. 4.1.2RQ1b (Transfer to Final Exam Performance): Modest Numerical But Not Significant Advantages for Conditions 2-4 on Final Exam, but Prompting Skill Predicts Final Exam Scores. As shown in Table 3, final exam scores six weeks post-intervention showed no significant between-group differences (F(3, 427) = 1.29, p = .278,ķ 2 = .009), despite numerical advantages for Conditions 2-4 over baseline (ķ 2 = 54.1%, ķ 3 = 53.3%, ķ 4 = 55.0% vs. ķ 1 = 50.4%). Although there is no direct between-condition effects, a linear regression model indicates that, after controlling studentsā pre-test score, their post-test score on pedagogical prompting is a significant predictor of their final exam scores (β= 0.090, t(428) = 3.01, p = Conferenceā17, July 2017, Washington, DC, USAXiao et al. ConditionN Pre-TestImmediate Post-TestImmediate Gain MeanSDMeanSDMeanSDEffect SizeķT-Statistics 11020.240.120.320.150.080.150.53< .001W = 976.50 21120.240.130.630.240.390.241.63< .001W = 53.00 31090.240.120.690.250.450.231.96< .001t = -20.25 41080.230.130.740.230.510.281.82< .001W = 125.00 Table 1: Descriptive statistics by condition on Immediate Learning Gain ConditionN Pre-TestDelayed Post-TestDelayed Gain MeanSDMeanSDMeanSDEffect SizeķT-Statistics 11020.240.120.310.140.060.160.38<.001t = -4.00 21120.240.130.390.160.150.200.75<.001t = -8.14 31090.240.120.400.190.170.200.85<.001t = -8.41 41080.230.130.470.200.240.221.09<.001t = -11.20 Table 2: Descriptive statistics by condition on Delayed Learning Gain .003). Specifically, for students with similar pre-test scores, every 1 percentage-point increase in immediate post-test scores, students gained approximately 0.09 percentage-point on the final exam. Cond.NMean SD Median95% CI 11020.670.170.73[0.64, 0.71] 21120.710.170.77[0.68, 0.74] 31090.690.180.72[0.66, 0.72] 41080.710.170.77[0.68, 0.74] Table 3: Final exam scores by condition. No significant between-group differences (ķ¹(3,427)=1.29,ķ= .278,ķ 2 = .009), ANOVA. 4.2 RQ2 (Instructional Methods): Instructors Can Balance Learning Effectiveness, Time Investment, and Learner Adoption When Selecting Interventions; All Approaches Benefit Different Learners Equitably 4.2.1RQ2a (Learning Gain): Learning Gains Increase Progressively from Condition 1 to 4, with Select-Then-Write (Condition 4) Achiev- ing Highest Gain. The select-then-write approach (Condition 4) achieved the highest learning outcomes across all measures (see Figure 3). More specifically, Condition 4 students demonstrated im- mediate learning gains of 0.51 (SD = 0.28), significantly exceeding Conditions 1 and 2 (both p < .001) and marginally outperforming Condition 3 (M = 0.45, p = .08). These advantages persisted at the 7-week delayed post-test, where Condition 4 maintained the high- est delayed gains of 0.24 (SD = 0.22), significantly surpassing all other conditions (p < .001) and representing 47% retention of initial improvements. Compared to the baseline condition, Condition 4 produced learning gains more than six times larger immediately (0.51 vs. 0.08) and four times larger at delayed follow-up (0.24 vs. 0.06), demonstrating both superior acquisition and retention. 4.2.2 RQ2b (Total Time Spent): Remarkably Low Barrier: Even <1 Minute Produces Significant Gains; Highest Learning Gain Achieved Within Half a Class Period. Following the pre-test, learners in Con- ditions 1ā4 spent a median of 0.89, 3.14, 11.11, and 36.92 minutes (Table 4, respectively, from initiating the learning intervention to logging out. As the required level of interactivity in the instruc- tional design increased across conditions, the median total time spent grew by roughlyĆ3 at each successive level. For Condition 3 and 4, the total time spent also included potential off-task inter- vals. For instance, some students were inactive for over 10 minutes before resuming. Cond. Duration Median (mins)95% CI 10.89[0.71, 1.07] 23.14[2.56, 3.62] 311.11[8.88, 12.31] 436.92[34.71, 42.78] Table 4: Duration of different learning interventions and learning efficiency (gain per minute) by condition. 4.2.3RQ2c (Equity Across Diverse Learners). : Interventionās effec- tiveness donāt vary from learnersā differences in mindsets, cognitive preferences, or prior abilities. Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 CourseConferenceā17, July 2017, Washington, DC, USA Figure 3: Distribution of Immediate and Delayed Learning Gain by Condition (Per Student) with Significance. ns:ķ ā„ .05; *: ķ< .05; **: ķ< .01; ***: ķ< .001 OLS regression analyses tested whether learner characteristics moderated instructional effects. Learner profiles examined included mindsets [11], Need for Cognition (NFC) [7], self-perceived pro- gramming abilities and prompting ability. Notably, these individual differences did not significantly predict learning outcomes, nor did they interact with experimental conditions to influence immedi- ate or delayed learning gains (all interaction terms p > 0.05). This suggests that the instructional interventions were equally effec- tive for learners regardless of their entering mindsets, cognitive preferences, or self-assessed abilities. 4.2.4 RQ2d (Learner Perceptions and Adoption): Constructive En- gagement Perceived as More Challenging, Yet Does Not Deter Sus- tained Adoption of Pedagogical Prompting. RQ2d-Frustration: Condition 4 reported significantly higher frus- tration caused by both desirable difficulties and extraneous load. Post- intervention surveys assessed frustration (1 = Not Frustrating at All, 5 = Very Frustrating). Condition 4 reported significantly higher frustration (ķ= 2.07) than Conditions 1ā3 (ķ= 1.46ā1.58;ķ¹(3, 427) = 7.05,ķ< .001; Tukey HSD: allķ< .005), with no differences among Conditions 1ā3 (allķ> .85). Notably, even Condition 4ās frustration remained below the scale midpoint. More specifically, by analyzing responses to āLet us know what makes you frustrated during the previous learning session!ā using the- matic coding, two categories (desirable difficulties and extraneous load) emerged (Table 5, Panel A): (1) Desirable difficulties reflected productive cognitive engage- ment inherent to the constructive task. Students reported struggling to craft prompts that elicit guidance rather than direct answers (e.g., āI had to think of different prompts that would not provide direct answersā ), difficulty diagnosing the fictional peerās specific prob- lem type (e.g., āit was hard to identify what kind of problem they encounteredā ), and the effortful restraint required to avoid answer- seeking behaviors (e.g., āstruggling in resisting the urge to ask for direct answersā ). (2) Extraneous load reflected system- or user-experience-level friction unrelated to the learning objectives, often caused by tech- nical or logistic issues. The most prominent source was the incon- sistency of the LLM-driven grading system, which occasionally rejected semantically correct prompts that lacked exact keyword matches (e.g., āa space being the difference between getting it wrong and getting it rightā ). Students also reported platform instability (e.g., ācrashesā, ālost progressā ), excessive repetition across the three structurally similar scenarios, and unclear instructions. However, these extraneous frustrations were relatively infrequent, account- ing for approximately 7% of all open-ended responses, suggesting that the majority of reported frustration stemmed from the cogni- tively productive aspects of the task rather than implementation issues. RQ2d-Usage: Regular adoption of pedagogical prompts exceeding once per week, with perceived benefits for self-regulated learning and help-seeking. Students also reported their pedagogical prompting usage in the delayed post-survey over the subsequent six weeks (1 = Never, 5 = Everyday) after the learning intervention. All conditions reported regular usage exceeding āonce a weekā (ķ= 3.49ā3.83). Condition 4ās usage (ķ= 3.77) did not differ from any other con- dition (allķ> .10), indicating that higher frustration did not deter sustained adoption. Conferenceā17, July 2017, Washington, DC, USAXiao et al. Primary ThemeSecondary ThemeRepresentative Examples Panel A: Student-Reported Sources of Frustration Desirable Difficulties Crafting learning-oriented promptsāI had to think of different prompts that would not provide direct answersā āIt was challenging at first to figure out how to phrase my prompts correctlyā Diagnosing problem typesāIt was hard to identify what kind of problem they encounteredā āI wasnāt quite sure how to determine what the problem was in what situationā Resisting answer-seeking impulsesāStruggling in resisting the urge to ask for direct answersā āIt is hard to change habit of using GPTā Extraneous Load Rigid LLM-driven validationāA space being the difference between getting it wrong and getting it rightā āI wrote the correct idea but it flagged incorrectā Platform instabilityāIt glitched out once and thus I had to restart the entire moduleā Excessive repetitionāEighty percent of the questions across all three sections are identicalā āSo many repetitive questions about prompting AIā Unclear instructionsāNot enough guidance on the website and I donāt know what to do at firstā Panel B: Student-Reported Benefits of Pedagogical Prompting Self-Regulated Learning Active thinking over passive answer-seekingāIt still left me with work to do to solve the problemā āI can think better and not just get the answers revealedā Transferable problem-solving skillsāIt helps to actively recall the steps which enforces your syntax memoryā āI was able to do similar (programming) problems in the future without GPTā Diagnosing knowledge gapsāOne prompt I used allows me to identify the gaps in my understandingā āI can know what kind of mistakes I make in my codingā Behavioral shift toward learning-oriented AI useāBefore lab 7.5 I would often look at the full solution ... After 7.5 I attempted to solve the issue on my ownā āI know the right way to ask AI nowā Accessible Help-Seeking Resource On-demand tutor-like supportāI could work through the issues with the bot like I would with a TA, and not have to wait until the next morningā āI can ask unlimited trivial and simple questions over and over until I understoodā Personalized pacing and level adaptationāMeets me at my level at my own time at my own discretionā āExplain to a 12 year old ... it breaks down the challenging topic into something legibleā Table 5: Thematic analysis of open-ended survey responses (ķ= 431). Panel A: student-reported benefits of pedagogical prompting (delayed post-survey). Panel B: student-reported sources of frustration (immediate post-survey). Quotes are lightly edited for brevity. Thematic analysis of delayed post-survey open-ended responses revealed two primary themes regarding the perceived benefits of pedagogical prompting (Table 5, Panel B): (1) Self-Regulated Learning, encompassed four secondary themes: students reported that pedagogical prompting promoted active thinking over passive answer-seeking, helped them develop trans- ferable problem-solving skills, enabled them to diagnose their own knowledge gaps, and triggered a behavioral shift from using AI as a solution provider to using it as a learning tool. For example, a majority of learners expressed a heightened awareness of the distinction between seeking answers and actively learning: āI can think better and not just get the answers revealedā, which motivated them to apply the pedagogical prompting strategies introduced during the intervention in their subsequent coursework. (2) Accessible Help-Seeking Resource, captured how students lever- aged pedagogical prompting to access on-demand, tutor-like sup- port unconstrained by office hours (e.g., āI could work through the issues with the bot like I would with a TA, and not have to wait until the next morningā ), and to receive explanations personalized to their proficiency level and pace (e.g., āExplain to a 12 year oldā ). Detailed examples can be found in Table 5. RQ2d-Frustration does not predict sustained adoption. Lin- ear regression confirmed that post-intervention frustration did not predict usage frequency over the subsequent six weeks (β= -0.010, t(416) = -0.245, p = .807, R2 < .001). This null relationship persisted when controlling for experimental condition in a multiple regres- sion model (frustration:β= 0.059, p = .475; overall model: F(7, 410) = 2.13, p = .039, R2 = .035). These findings demonstrate that the complexity of constructive instructional methods does not deter sus- tained adoption, encouraging instructors to confidently implement evidence-based approaches like select-then-write without excessive concern that student frustration will backfire on learning-oriented prompting adoption. 5 Discussion This paper examines the effectiveness of four types of prompting lit- eracy instruction (including a business-as-usual baseline) through a semester-long, large-scale randomized controlled trial in an in- troductory computer science course (N = 979). The results show Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 CourseConferenceā17, July 2017, Washington, DC, USA that all interventions significantly improved studentsā prompting skills; prompting ability predicted final exam performance, but improvements in prompting did not directly transfer to higher fi- nal exam scores (RQ1). We also identified factors, including time efficiency, learning effectiveness, equity, learner perceptions and adoption, that can inform future GenAI policies and instructional design in classrooms (RQ2). Detailed discussions of RQ1 and RQ2 are provided in subsection 5.1 and subsection 5.2 respectively. 5.1Robust Skill Acquisition, Potential Transfer to Final Exam Performance 5.1.1 Progressive Learning Gains Support the ICAP Framework in an Ill-Defined Domain. Learning gains in both the immediate and delayed post-tests increased progressively across conditions (ķ 1 = 0.08ā ķ 4 = 0.51). Because the four conditions were intentionally designed to reflect increasing levels of engagement based on the ICAP framework, the findings provide two-way validation between theory and design. First, the results extend ICAP to a new learning domain, prompting literacy. As engagement increased from passive to active, constructive, and interactive learning, time-on-task rose accordingly and learning gains and retention improved in parallel. This alignment between engagement level, time investment, and learning outcomes provides empirical support for ICAP in an ill- defined domain. Second, the results validate the intervention design itself. The systematic increase in learning gains and the strongest retention observed in Condition 4 (47%) suggest that the additional behavioral and cognitive effort required by higher-engagement activities translated directly into meaningful knowledge gains. In other words, the effort demanded by the interventions was not wasted; it was well aligned with the instructional goals. 5.1.2 Prompting Skill is Associated with Final Exam Performance. Although no significant between-group differences emerged on the final exam, prompting skill was positively associated with course performance. A linear regression model controlling for pre-test scores revealed that immediate post-test performance on pedagogi- cal prompting significantly predicted final exam scores (ķ½= 0.090, ķ”(428) = 3.01,ķ= .003): among students with comparable pre-test scores, each additional percentage point on the immediate post-test was associated with a 0.09 percentage-point increase on the final exam. This pattern suggests that the quality of prompting literacy ac- quired during the intervention, rather than condition assignment itself, is associated with end-of-semester academic performance. Students who more thoroughly acquired pedagogical prompting skills, regardless of which condition facilitated that acquisition, may have been better equipped to interact with GenAI tools in a learning-oriented manner rather than relying on them as short- cut solution providers [24,30]. This, in turn, could have generated more productive learning opportunities throughout the semester, ultimately contributing to stronger final exam performance. This interpretation is consistent with the view that structured prompting instruction mitigates the risks of unrestricted AI use identified by Bastani et al. [3], producing neutral-to-positive effects on course outcomes rather than the harmful effects observed under unguided AI access. However, this interpretation is one of several plausible explana- tions and should be treated with caution. The regression explains only a small proportion of variance in final exam scores, and the observational nature of this analysis precludes causal claims. Stu- dents who performed well on the post-test may have been more academically engaged or stronger learners overall. Future work in- corporating formal mediation analyses with additional co-variates (e.g., cumulative GPA, prior programming experience) would help clarify whether improved prompting skills causally contribute to course performance. 5.2 Instructional Design Recommendations To provide better design implications for future GenAI policy and instructions in classrooms, we examined 4 aspects: (1) time effi- ciency, (2) learning effectiveness, (3) equity, (4) learner perceptions and adoption. For time efficiency and learning effectiveness, we identified suitable instructional approaches for different combinations of available class time and desired learning outcomes. For instruc- tors with minimal time availability, Condition 1 (baseline reminder, <1 minute) demonstrates that brief policy-level reminders empha- sizing responsible AI use can meaningfully shift student prompting behaviors, producing significant and durable gains (r = .53, p < .001). For instructors who can allocate half a class period (about 37 minutes) and prioritize maximizing learning outcomes, Condi- tion 4 produces the highest gains we identified. This gradient of evidence-based options, spanning <1 minute to half a class period, all produce significant improvements and enable instructors to se- lect interventions aligned with their specific resource constraints and pedagogical goals, making pedagogical prompting instruction readily integrable into diverse curricular contexts without requiring extensive restructuring. For educational equity consideration, our findings reveal that all 4 types of pedagogical prompting instructions benefited all stu- dents equitably regardless of their entering characteristics. OLS regression analyses testing potential moderators, including implicit theories of intelligence [11], need for cognition [7], self-perceived programming abilities, and self-assessed prompting ability. This universal effectiveness pattern indicates that the instructional ap- proaches we examined do not differentially advantage or disad- vantage particular learner subgroups, addressing critical concerns about algorithmic fairness and equitable access to effective instruc- tion in technology-enhanced learning environments [1,14]. From a practical standpoint, this means instructors can confidently deploy any of the examined interventions across diverse student popula- tions without concern for exacerbating existing achievement gaps or requiring costly personalization systems to ensure fairness [39]. The equitable nature of these approaches stands in contrast to many educational technology interventions where benefits accrue pri- marily to already-advantaged learners [47], making pedagogical prompting instruction particularly valuable for large-enrollment courses serving heterogeneous student bodies. For learner perceptions and adoption, studentsā frustration during learning did not deter their subsequent use of pedagogical prompting, as students in Condition 4 reported higher frustration (p < .001), regression analysis confirmed frustration did not predict Conferenceā17, July 2017, Washington, DC, USAXiao et al. usage frequency (β= -0.010, p = .807). This suggests the challenge may represent productive struggle [6,18,20]ādesirable difficulty that enhances learning [4] while remaining tolerable when ped- agogical value is clear. Accordingly, GenAI instructional design should anticipate a moderate level of learner frustration as a natu- ral consequence of cognitive engagement rather than a barrier to adoption. Instructors can help students perceive difficulty from a negative signal into a positive indicator of learning, reinforcing the message that āfeeling challenged can mean learning is happening.ā 6 Limitations and future works With the initial success and findings, several limitations of the cur- rent study point the direction of future works. First, our study was conducted in a single introductory-level CS course at one institu- tion; replication across diverse disciplines and educational settings is needed to establish broader generalizability. Second, while we demonstrated that prompting skills predict course performance through mediation, we did not establish causal effects of improved prompting ability on domain-specific learning outcomes. Future experimental designs isolating prompting construction from ap- plication could clarify this mechanism. Finally, while we found no moderation by psychological characteristics, future research could investigate whether benefits remain equitable across more dimen- sions to ensure interventions do not exacerbate existing disparities. 7 Conclusion This study provides the first large-scale RCT evidence (N=979) that pedagogical prompting instruction, grounded in the ICAP frame- work, significantly improves studentsā ability to use GenAI as a learning tool rather than a shortcut. Our findings demonstrate that even minimal interventions produce durable gains, while construc- tive approaches like select-then-write maximize learning outcomes equitably across diverse learners. The complete mediation pathway linking prompting skills to course performance underscores that how students interact with AI matters more than whether they use it. We offer instructors a practical, evidence-based gradient of interventions adaptable to varying time and resource constraints, advancing the field from GenAI policy to GenAI pedagogy. References [1] Ryan S Baker and Aaron Hawn. 2022. Algorithmic bias in education. International Journal of Artificial Intelligence in Education 32, 4 (2022), 1052ā1092. [2]Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Ćzge Kabakcı, and Rei Mariman. 2024. Generative AI can harm learning. The Wharton School Research Paper (2024). [3]Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Ćzge Kabakcı, and Rei Mariman. 2025. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122, 26 (2025), e2422633122. [4]Elizabeth Ligon Bjork and Robert A Bjork. 2011. Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. Psychology and the real world: Essays illustrating fundamental contributions to society 2, 59-68 (2011). [5] Robert A Bjork. 1994. Memory and metamemory considerations in the training of human beings. Metacognition: Knowing about knowing (1994), 185ā205. [6]Zana BuƧinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-computer Interaction 5, CSCW1 (2021), 1ā21. [7]John T Cacioppo, Richard E Petty, and Chuan Feng Kao. 1984. The efficient assessment of need for cognition. Journal of personality assessment 48, 3 (1984), 306ā307. [8]Michelene TH Chi and Ruth Wylie. 2014. The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational psychologist 49, 4 (2014), 219ā243. [9] Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1 (1960), 37ā46. [10]Paul Denny, James Prather, Brett A Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing education in the era of generative AI. Commun. ACM 67, 2 (2024), 56ā67. [11]Carol S Dweck. 2013. Self-theories: Their role in motivation, personality, and development. Psychology press. [12]Josh Freeman. 2025. Student generative ai survey 2025. Higher Education Policy Institute: London, UK (2025). [13]Krzysztof Z Gajos and Lena Mamykina. 2022. Do people engage cognitively with AI? Impact of AI assistance on incidental learning. In Proceedings of the 27th International Conference on Intelligent User Interfaces. 794ā806. [14]Wayne Holmes, Jens Persson, Irene-Angelica Chounta, Barbara Wasson, and Vania Dimitrova. 2022. Ethics of AI in education: Towards a community-wide framework. International Journal of Artificial Intelligence in Education 32, 3 (2022), 504ā526. [15]Xinying Hou, Zihan Wu, Xu Wang, and Barbara J Ericson. 2024. Codetailor: Llm-powered personalized parsons puzzles for engaging support while learning programming. In Proceedings of the Eleventh ACM Conference on Learning@ Scale. 51ā62. [16]Xinying Hou, Ruiwei Xiao, Runlong Ye, Michael Liut, and John Stamper. 2026. Ex- ploring Student Choice and the Use of Multimodal Generative AI in Programming Learning. (2026). [17] Yueqiao Jin, Lixiang Yan, Vanessa Echeverria, Dragan GaÅ”eviÄ, and Roberto Martinez-Maldonado. 2025. Generative AI in higher education: A global perspec- tive of institutional adoption policies and guidelines. Computers and Education: Artificial Intelligence 8 (2025), 100348. [18] Manu Kapur. 2016. Examining productive failure, productive success, unproduc- tive failure, and unproductive success in learning. Educational Psychologist 51, 2 (2016), 289ā299. [19] Enkelejda Kasneci, Kathrin SeĆler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al.2023. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and individual differences 103 (2023), 102274. [20]Majeed Kazemitabaar, Oliver Huang, Sangho Suh, Austin Z Henley, and Tovi Grossman. 2025. Exploring the design space of cognitive engagement tech- niques with ai-generated code for enhanced learning. In Proceedings of the 30th international conference on intelligent user interfaces. 695ā714. [21] Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. Codeaid: Evaluating a classroom deployment of an llm-based programming assistant that balances student and educator needs. In Proceedings of the 2024 chi conference on human factors in computing systems. 1ā20. [22] Chris Kerslake, Paul Denny, David H Smith IV, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil. 2024. Integrating natural language prompting tasks in introductory programming courses. In Proceedings of the 2024 on ACM Virtual Global Computing Education Conference V. 1. 88ā94. [23]Hieke Keuning, Isaac Alpizar-Chacon, Ioanna Lykourentzou, Lauren Beehler, Christian Kƶppe, Imke de Jong, and Sergey Sosnovsky. 2024. Studentsā Percep- tions and use of generative AI tools for programming across different computing courses. In Proceedings of the 24th Koli calling international conference on comput- ing education research. 1ā12. [24]Sunnie SY Kim, Jennifer Wortman Vaughan, Q Vera Liao, Tania Lombrozo, and Olga Russakovsky. 2025. Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1ā19. [25]Nils Knoth, Antonia Tolzin, Andreas Janson, and Jan Marco Leimeister. 2024. AI literacy and its implications for prompt engineering strategies. Computers and Education: Artificial Intelligence 6 (2024), 100225. [26]Kenneth R Koedinger, Jihee Kim, Julianna Zhuxin Jia, Elizabeth A McLaughlin, and Norman L Bier. 2015. Learning is not a spectator sport: Doing is better than watching for learning from a MOOC. In Proceedings of the second (2015) ACM conference on learning@ scale. 111ā120. [27] Harsh Kumar, Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Xinyuan Wang, Huayin Luo, Joseph Jay Williams, Anna N Rafferty, John Stamper, et al. 2024. Supporting self-reflection at scale with large language models: Insights from randomized field experiments in classrooms. In Proceedings of the eleventh ACM conference on learning@ scale. 86ā97. [28]J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977), 159ā174. [29]Daniel Lee and Edward Palmer. 2025. Prompt engineering in higher education: a systematic review to help inform curricula. International Journal of Educational Technology in Higher Education 22, 1 (2025), 7. Transforming GenAI Policy to Prompting Instruction: An RCT of Scalable Prompting Interventions in a CS1 CourseConferenceā17, July 2017, Washington, DC, USA [30]John D Lee and Katrina A See. 2004. Trust in automation: Designing for appro- priate reliance. Human factors 46, 1 (2004), 50ā80. [31]Mark Liffiton, Brad E Sheese, Jaromir Savelka, and Paul Denny. 2023. Codehelp: Using large language models with guardrails for scalable support in programming classes. In Proceedings of the 23rd Koli calling international conference on computing education research. 1ā11. [32] Qian Liu, Anjin Hu, Tehmina Gladman, and Steve Gallagher. 2025. Eight months into reality: A scoping review of the application of ChatGPT in higher education teaching and learning. Innovative higher education 50, 5 (2025), 1677ā1700. [33]Wenhan Lyu, Yimeng Wang, Yifan Sun, and Yixuan Zhang. 2025. Will your next pair programming partner be human? an empirical evaluation of generative ai as a collaborative teammate in a semester-long classroom setting. In Proceedings of the Twelfth ACM Conference on Learning@ Scale. 83ā94. [34]Philipp Mayring. 2021. Qualitative content analysis: A step-by-step guide. (2021). [35]Sharon Nelson-Le Gall. 1981. Help-seeking: An understudied problem-solving skill in children. Developmental review 1, 3 (1981), 224ā246. [36]Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q Feldman. 2024. How beginning programmers and code llms (mis) read each other. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1ā26. [37]James Prather, Brent N Reeves, Paul Denny, Brett A Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. āItās weird that it knows what i wantā: Usability and interactions with copilot for novice programmers. ACM transactions on computer-human interaction 31, 1 (2023), 1ā31. [38] James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randri- anasolo, Brett A Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The widening gap: The benefits and harms of generative ai for novice programmers. In Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1. 469ā486. [39]Justin Reich. 2020. Failure to disrupt: Why technology alone canāt transform education. Harvard University Press (2020). [40]Ido Roll, Vincent Aleven, Bruce M McLaren, and Kenneth R Koedinger. 2011. Improving studentsā help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and instruction 21, 2 (2011), 267ā280. [41]James Skripchuk, John Bacher, and Thomas Price. 2024. An Investigation of the Drivers of Novice Programmersā Intentions to Use Web Search and GenAI. In Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1. 487ā501. [42]Meg Sullivan, Aaron Kelly, and Paul McLaughlan. 2023. ChatGPT in higher education: Considerations for academic integrity and student learning. Journal of Applied Learning and Teaching 6, 1 (2023). [43]Allison Elliott Tew and Mark Guzdial. 2010. Developing a validated assessment of fundamental CS1 concepts. In Proceedings of the 41st ACM technical symposium on Computer science education. 97ā101. [44]David R Thomas. 2006. A general inductive approach for analyzing qualitative evaluation data. American journal of evaluation 27, 2 (2006), 237ā246. [45] Rachel Van Campenhout, Kevin S Autry, Michelle W Clark, and Benny G Johnson. 2025. Scaling the doer effect: A replication analysis using AI-generated questions. In Proceedings of the Twelfth ACM Conference on Learning@ Scale. 24ā34. [46]Sierra Wang, Thomas Jefferson, Chris Piech, and John C Mitchell. 2025. The Effects of Chatbot Placement, Personification, and Functionality on Student Outcomes in a Global CS1 Course. In Proceedings of the Twelfth ACM Conference on Learning@ Scale. 73ā82. [47]Mark Warschauer. 2004. Technology and social inclusion: Rethinking the digital divide. MIT press. [48] J White. 2023. A prompt pattern catalog to enhance prompt engineering with ChatGPT. arXiv preprint arXiv:2302.11382 (2023). [49]Ruiwei Xiao, Xinying Hou, Runlong Ye, Majeed Kazemitabaar, Nicholas Diana, Michael Liut, and John Stamper. 2025. Improving Student-AI Interaction Through Pedagogical Prompting: An Example in Computer Science Education. arXiv preprint arXiv:2506.19107 (2025). [50] Lixiang Yan, Samuel Greiff, Ziwen Teuber, and Dragan GaÅ”eviÄ. 2024. Promises and challenges of generative artificial intelligence for human learning. Nature human behaviour 8, 10 (2024), 1839ā1850. [51]J Diego Zamfirescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang. 2023. Why Johnny canāt prompt: how non-AI experts try (and fail) to design LLM prompts. In Proceedings of the 2023 CHI conference on human factors in computing systems. 1ā21. [52] BJ Zimmerman. 2000. Attaining self-regulation: A social cognitive perspective. Self-regulation: Theory, research, and applications/Academic (2000).