Paper deep dive
To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks
Jessica Hutchison, Ian Tyler Applebaum, Kenneth Angelikas, Kush Rakesh Patel, Phuoc Nguyen, Antonio Lazaro, Nicholas Rucinski, Rahad Arman Nabid, Stephen MacNeil
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 7/5/2026, 4:23:21 AM
Summary
The paper introduces Clover, a specialized code completion tool designed to measure critical engagement in AI-assisted programming through behavioral signals and attention checks. By logging fine-grained interactions (like 'Tab-to-accept' vs. 'Slow-accept') and injecting semantically inconsistent suggestions (attention checks), the researchers developed a taxonomy of interaction behaviors. The study, conducted with CS1 students, found that higher 'Tab-accept' rates were strongly associated with lower performance on attention checks, suggesting a lapse in reflective engagement, whereas increased dwell time was associated with better attention check performance. The findings highlight the difficulty in finding a single perfect metric for responsible AI use in computing education.
Entities (7)
Relation Signals (4)
Clover → implements → Attention Checks
confidence 100% · Clover... additionally offers attention checks to probe reflective engagement
Clover → uses → Gemini 3 Large Language Model
confidence 100% · Clover is an integrated development environment (IDE) extension that uses the Gemini 3 Large Language Model
Tab Accept → negativelycorrelateswith → Attention Check Performance
confidence 90% · higher rates of tab accept were associated with lower attention check performance
Dwell Time → positivelycorrelateswith → Attention Check Performance
confidence 90% · increased dwell time was associated with higher attention check performance
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI code completion tools, such as Github Copilot, provide students with code suggestions to help them write programs. However, recent qualitative studies suggest that students fail to critically evaluate these suggestions. We present Clover, a code completion tool that logs students' interactions with code suggestions and additionally offers attention checks to probe reflective engagement during programming tasks. We also develop a taxonomy of behavioral interaction metrics for AI-assisted programming, informed by literature. We analyzed relationships between interaction patterns, engagement with attention checks, and task performance. We observed that higher rates of tab accept were associated with lower attention check performance, while increased dwell time was associated with higher attention check performance. We conclude by discussing how programming process data and attention checks might support reflective engagement in AI-assisted programming.
Tags
Links
- Source: https://arxiv.org/abs/2606.30549v1
- Canonical: https://arxiv.org/abs/2606.30549v1
Trouble viewing inline? Open PDF directly →
Full Text
44,528 characters extracted from source content.
Expand or collapse full text
To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks Jessica Hutchison Temple University Philadelphia, USA jessica.hutchison@temple.edu Ian Tyler Applebaum Temple University Philadelphia, USA ian.tyler@temple.edu Kenneth Angelikas Temple University Philadelphia, USA kenneth.angelikas@temple.edu Kush Rakesh Patel Temple University Philadelphia, USA tur46917@temple.edu Phuoc Nguyen Temple University Philadelphia, USA jaime.nguyen@temple.edu Antonio Lazaro Temple University Philadelphia, USA tut08369@temple.edu Nicholas Rucinski Temple University Philadelphia, USA nicholas.rucinski@temple.edu Rahad Arman Nabid Temple University Philadelphia, USA rahad.arman.nabid@temple.edu Stephen MacNeil Temple University Philadelphia, USA stephen.macneil@temple.edu Abstract AI code completion tools, such as Github Copilot, provide students with code suggestions to help them write programs. However, re- cent qualitative studies suggest that students fail to critically evalu- ate these suggestions. We present Clover, a code completion tool that logs students’ interactions with code suggestions and addition- ally offers attention checks to probe reflective engagement during programming tasks. We also develop a taxonomy of behavioral interaction metrics for AI-assisted programming, informed by lit- erature. We analyzed relationships between interaction patterns, engagement with attention checks, and task performance. We ob- served that higher rates of tab accept were associated with lower attention check performance, while increased dwell time was asso- ciated with higher attention check performance. We conclude by discussing how programming process data and attention checks might support reflective engagement in AI-assisted programming. Keywords Copilot, Attention Checks, Generative AI, Computing Education 1 Introduction AI-assisted programming is rapidly changing how software is writ- ten, with developers increasingly working alongside systems that generate code suggestions in real time. AI code completion tools, such as GitHub Copilot and Cursor, are increasingly integrated into professional workflows and, more recently, into computing courses as well [24,33]. Although student adoption of code assis- tants has lagged behind general purpose AI tools such as ChatGPT, recent surveys show their use by computing students is steadily increasing [10,15,16,25,32]. As these tools become standard, fun- damental questions arise about how students interact with AI code suggestions and how to meaningfully measure these interactions. © Jessica Hutchison, Ian Tyler Applebaum, Kenneth Angelikas, Kush Rakesh Patel, Phuoc Nguyen, Antonio Lazaro, Nicholas Rucinski, Rahad Arman Nabid, and Stephen MacNeil 2026. This is the author’s version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 31st ACM Conference on Innovation and Technology in Computer Science Education (ITiCSE 2026), http://dx.doi.org/10.1145/3803400. 3809394. Despite growing adoption, understanding students’ interactions with AI code completion tools remains a challenge. Existing stud- ies rely on qualitative methods including think-aloud protocols [1, 35], screen recordings [28,31], screenshots [21], and process jour- nals [30]. While these approaches provide rich insights they are difficult to scale and offer limited visibility into fine-grained real- time interactions with code suggestions. This measurement gap is important given emerging evidence that AI code assistants can impact learning behaviors and outcomes. Prior work has suggested that code completion tools can mislead students and disrupt their metacognitive processes [23,28]. They can also lead to unproductive debugging ‘rabbit holes’ [34]. These negative impacts tend to disproportionately affect students with less experience and lower self efficacy [2,9,20,26,36]. Therefore, it is critical to develop metrics and methods for evaluating how students interact with these tools. To address this gap, we introduce Clover, a code completion tool that instruments student interactions with AI-generated sug- gestions and introduces lightweight attention checks to probe reflective engagement during programming tasks. To inform the metrics Clover tracks, we developed a comprehensive Taxonomy of Interaction Behaviors with AI Code Suggestions guided by metrics used across prior literature [1,14,27,28,34]. We then conducted a lab study with Clover in a CS1 course and analyze relationships between interaction patterns, attention check performance, and task performance to investigate the following research questions: RQ1: How does CS1 students’ usage of a code completion tool relate to their task performance? RQ2:How does CS1 students’ engagement with attention checks relate to their task performance? In this paper, we make the following contributions: • Clover, an instrumented code completion tool that logs fine-grained interactions with AI-generated sugges- tions and integrates attention checks to probe reflective engagement. • A taxonomy of interactions with AI code suggestions for AI-assisted programming synthesized from prior work. arXiv:2606.30549v1 [cs.HC] 29 Jun 2026 Jessica Hutchison et al. •An empirical study demonstrating how interaction pat- terns relate to engagement and task performance. 2 Related Work Programming Process Data (PPD) captures how students inter- act with the integrated development environment while coding. PPD can be collected at varying levels of granularity, from coarse- grained artifacts, such as final submissions, to fine-grained logs of individual keystrokes [5,17]. PPD has been used to study how novice programmers code, identify sources of struggle, and support early intervention or integrity monitoring [3,5,7,13,17,29]. How- ever, this prior work has been focused on traditional programming environments without AI assistance. The introduction of AI code completion tools have surfaced di- verse interaction patterns and complex impacts on learning. Prior work suggests that programmers interact with AI tools differently depending on their goals and experience. For example, Barke et al. [1]conducted one of the earliest studies of how professional pro- grammers interact with Github Copilot identifying two modes: exploration where programmers use Copilot to explore options and acceleration where the programmer knows what to do and uses Copilot to get there faster. While this shows the flexibility of code assistants for experts, subsequent work suggests that novices may struggle to use these tools effectively. Based on an eye tracking study of students using Github Copi- lot, Prather et al. [28]found that less prepared students were often misled by AI suggestions and struggled to differentiate good sug- gestions from bad ones. This reflected a broader trend of unequal learning benefits associated with AI use [2,9,25,28,36]. Another study by Vaithilingam et al. [34]showed how students could get stuck in ‘debugging rabbit holes’ when using code generated by AI code assistants. Additional issues include students experienc- ing an ‘illusion of competence’ [28] where students have difficulty identifying gaps in understanding [20, 28]. However, understanding these interaction patterns remains chal- lenging. Most studies have used qualitative methods such as screen capture [31], eye-tracking [28], or manual observations [31]. While these approaches provide valuable insights, they are labor-intensive and difficult to scale. PPD offers a complementary, scalable approach for analyzing behaviors such as accepting or rejecting suggestions across larger populations. However, existing PPD approaches have not been fully adapted to capture these interaction patterns. 3 Methods 3.1 Clover Design and Implementation Clover is an integrated development environment (IDE) extension that uses the Gemini 3 Large Language Model to provide real-time, single-line code suggestions at the cursor position based on code already in the editor. While programming, students can accept these code suggestions or write their own code manually. To cap- ture how students engage with suggestions, Clover logs students’ interactions, which are described in detail in Section 3.3. 3.1.1Instrumenting the Copilot Experience. To maximize ecological validity, Clover was implemented as a Visual Studio Code extension, mirroring the interface and functionality of widely-adopted tools like GitHub Copilot. We kept familiar interaction models, such as Tab-to-accept, with code suggestions appearing automatically when a user pauses during typing. By maintaining this user experience, students likely behave as they would in a real-world programming environment, allowing us to collect authentic behavioral data. 3.1.2 Single-Line Code Suggestions. Large blocks of AI code sug- gestions can be difficult to understand and debug [34], so Clover provides single-line code suggestions to reduce cognitive load and provide granular detection of behavioral metrics. 3.1.3Attention Checks. Over-reliance on AI code completion tools can occur when users default to fast thinking, accepting suggestions without the critical evaluation required to determine their correct- ness [6,8,12]. In order to detect students’ critical engagement with AI code suggestions, we intentionally inject a controlled number of suggestions that are semantically inconsistent with the student’s immediate coding goal. We say ‘inconsistent’ instead of ‘incorrect’ because these attention checks could still produce functional code through later corrections [18]. For example, a student creating a counter might be showncount--, when what they need iscount++. If the student accepts this suggestion, they fail the attention check; and if the student rejects the suggestion, they pass. These attention checks are intended to be similar to misleading suggestions and hallucinations in tools like Github Copilot [19], but are deliberate in- stead of incidental. This makes them useful in determining whether the student is paying attention. This is similar to Harbarth’s method of giving incorrect routes as AI suggestions in a driving context [8]. 3.2 In-Person Classroom Study We conducted an initial study of Clover which was approved by the Institutional Review Board (IRB) at Temple University. 3.2.1 Course Context. This study was conducted in a first-year programming course that serves as the second course in the intro- ductory programming sequence. The in-person course is taught in Java and focuses on object-oriented programming and data struc- tures. The course contains mix of students in computer science and related majors. The course features a weekly programming lab that runs for 110 minutes. Students attend lab in four separate sections. The course instructor typically prohibits the use of generative AI, and instead encourages students to get help from peers or in- structional staff; however, they made an exception for this lab with the intention to give their students exposure to AI tools. 3.2.2Participants. We recruited students from the course, offering extra credit as compensation. Students who opted not to partici- pate could receive equivalent credit by completing the assignment without Clover. 56 students spread across four different sections of this course consented and participated in every part of the study. Details about the study were presented verbally alongside a digital consent form. Students were made aware that their participation was entirely voluntary. Data was anonymized using unique par- ticipant IDs and stored securely in a password-protected database. One participant was removed because their average dwell time was three standard deviations from the mean, resulting in 55 students included in the analysis. To Tab or Not to Tab 3.2.3Study Procedures. The study took place simultaneously across the four lab sections with each following the same procedure: 1) 15 minutes were allocated to explain the study and for participants to complete the consent form, 2) students were introduced to the tool and given a warm-up problem to become familiar with the system for another 15 minutes, 3) students accessed our system through URLs to GitHub Codespaces, a browser based VSCode, with Clover preinstalled, 4) participants were assigned the rainfall problem to complete within a 60-minute time frame. During the task, Clover intermittently produced deliberately incorrect suggestions as at- tention checks. Participants were not informed of this behavior in advance because this awareness would have undermined the purpose of the manipulation. This use of deception was reviewed and approved by the Institutional Review Board. 3.3 Data Collection We instrumented Clover to log fine-grained interaction events as participants worked on a programming task. These events capture changes in the code state that occur when a suggestion is generated, displayed, or acted upon, allowing us to observe both what decisions participants make and how those decisions are reached. Events are stored with a timestamp of when the event occurred. 3.4 Taxonomy of Interaction Metrics Prior work has characterized a range of interaction patterns with AI code completion tools, but these characterizations are fragmented across studies. For example, Prather et al. conducted a study on Github Copilot use by CS1 students and identified slow accept, backtrack, and adapt behaviors [27]. These behaviors could more broadly be attributed to two novice programmer types, drifters and shepherds. Similarly, Barke et al. observed two modes of inter- action, acceleration and exploration [1]. Other studies show that students respond to incorrect AI suggestions by either attempting to repair accepted code or deleting it entirely [34]. Shihab et al. displays how high efficiency in completing the ‘brownfield’ tasks can mask levels of comprehension. Accordingly, we include effi- ciency and effectiveness metrics to understand participants’ levels of comprehension [31]. These studies present overlapping but inconsistently defined behavioral constructs, making it difficult to compare results across systems and contexts. In some cases, similar behaviors are labeled differently, while in others, identical labels refer to different con- structs. To address this, we synthesize prior work into a unified taxonomy of behavioral metrics for AI code completion tools. 3.4.1Suggestion Generation Events. Clover logs the following three sequential events leading up to displaying a suggestion: GACGenerate_Attention_Check: Generating a deliberately mis- leading suggestion. GS Generate_Suggestion: Generating a code suggestion that is not an attention check. S Suggestion_Shown: Visually presenting a new code sugges- tion or attention check (AC) after it is generated. 3.4.2Acceptance Events. We consider the following two instances to indicate a suggestion was accepted: TATab_Accept: Pressing ‘tab’ to accept a suggestion, without immediate modification or deletion. SA Slow_Accept: Typing out a suggestion character by charac- ter without clicking tab. This behavior was first observed qualitatively in Prather’s study of GitHub Copilot [27]. 3.4.3 Rejection and Revision Events. Measuring rejection is non- trivial because suggestions may be modified or deleted after they are initially accepted. We therefore distinguish between different behaviors that result in partial or complete rejection of a suggestion. AMAccept_then_Modify: Editing a suggestion immediately af- ter tab_accepting it. This includes removing characters (but not all the characters). This behavior has been observed but not named [11], or previously referred to as accept and adapt [27], accept and edit [1], and modify [14]. There are also related metrics, such as accept_then_repair [34], which have the additional connotation of correctness. AD Accept_then_Delete: Deleting a suggestion immediately af- ter tab_accepting it. This behavior was observed qualita- tively as backtracking [27]. A similar behavior, delete and search, involves attempted debugging before deletion [34]. IIgnore: Dismissing a suggestion without directly interacting with it. This could include running code, typing something else, or clicking elsewhere. 3.4.4 Execution Events. We also tracked when students ran their code. Running code is consistently associated with higher task performance [4], and so this metric was important to include: RCRun_Code: This captures when students compile and run their code, and shows the tests that pass and fail. NRNumber_of_Runs Total number of times student ran their code throughout entire session. NR= count(푅퐶) 3.4.5 Derived Metrics. From the event-level logs, we derived per- suggestion metrics that characterize participants’ engagement and decision-making processes. These include when users Accept code without modification, Reject code by interacting with the system in any way other than accepting the suggestion, Fail an Attention Check by accepting a deliberately incorrect line of code, or the Dwell Time which is the time measured in milliseconds between when a suggestion is shown to the start of any user action. A Accept=푇퐴∨ 푆퐴 R Reject= 퐴푀∨ 퐴퐷∨ 퐼 FC Failed Attention Check=(푇퐴∨ 푆퐴)∧(퐺퐴퐶∧ 푆) DT Dwell_Time= 푡 action −푡 푆 , 푡 action ∈ 푇퐴,푆퐴,퐴푀,퐴퐷,퐼 3.4.6Aggregated Session Metrics. Based on the metrics described in the previous section, we also produce metrics that described more complex behaviors and aggregated behavior across the session. AR Acceptance_Rate=(푐표푢푛푡(푇퐴)+푐표푢푛푡(푆퐴))/푐표푢푛푡(푆) TAR Tab_Acceptance_Rate=푐표푢푛푡(푇퐴)/푐표푢푛푡(푆) SAR Slow_Acceptance_Rate=푐표푢푛푡(푆퐴)/푐표푢푛푡(푆) AMR Accept_then_Modify_Rate=푐표푢푛푡(퐴푀)/푐표푢푛푡(푆) ADR Accept_then_Delete_Rate=푐표푢푛푡(퐴퐷)/푐표푢푛푡(푆) FCR Failed_AC_Rate=푐표푢푛푡(퐹퐶)/푐표푢푛푡(퐺퐴퐶∧ 푆) R Reject_Rate=푐표푢푛푡(푅)/푐표푢푛푡(푆) ADT Average_Dwell_Time= Í 퐷푇/푐표푢푛푡(푆) Jessica Hutchison et al. Figure 1: Distribution of session-level interaction rates. 3.4.7Outcome Metrics. We also measured task-level outcomes to assess overall solution correctness. TPTask_Performance: Total number of test cases passed out of all possible test cases for programming problem. 3.5 Correlation Analysis After computing the metrics presented in the previous section, we summarized these values using box plots to show their distribution. To understand the relationships between the metrics, we computed a Spearman correlation matrix using session-level aggregates. To ensure a consistent level of abstraction, all event-derived measures were first aggregated at the session level prior to analysis to en- sure a consistent unit of analysis across variables. Our analyses focused on examining the relationships between session-level be- havioral measures, which include dwell time, acceptance rate, slow acceptance rate, accept-then-modify rate, and accept-then-delete rate, and outcome measures of task performance and incorrect suggestion acceptance (failed attention checks). 4 Results Figure 1 presents box plots that summarize the distribution of be- havioral metrics. We observed that participants accepted (A) an average of 18.4 suggestions (푆퐷=30.9). Participants tab_accepted (TA) an average of 17.1 suggestions (푆퐷=31.0), and 45 partici- pants TA at least once. Participants slow_accepted (SA) an aver- age of 1.3 times (푆퐷=1.5), and 34 participants SA at least once. On average, participants accept_then_modified (AM) 17.7 times (푆퐷= 29.8), and 49 participants AM at least once. Participants ac- cept_then_deleted (AD) an average of 0.6 times (푆퐷=1.0), and 19 participants AD at least once. Participants ignored (I) an average of 36.3 suggestions (푆퐷=13.3), and all participants I suggestions. Figure 2: Correlations between session-level metrics using Spearman’s Rho. (* p<.05 ** p<.01 *** p<.001) Average_dwell_time (ADT) was 12.2 seconds (푆퐷=8.1), the min- imum was 1.7 seconds, and the maximum was 39.1 seconds. We an- alyzed session-level metrics, including number_of_runs (NR) and task_performance (TP). The average NR was 15.5 (푆퐷=12.6), the minimum was 0, and the high was 52. The average TP was 10.4 tests (푆퐷=12.9), the minimum was 0, and the maximum was 26. Out of the 55 participants, 22 students successfully completed the task, passing all 26 test cases. 4.1 RQ1: Interactions and Task Performance As shown in Figure 2, interaction metrics were weakly or incon- sistently correlated with task performance. Code execution, dwell time, and failed attention checks were most strongly correlated though they remained moderate to weak. 4.1.1 Code Execution (Number of Runs). Among the examined metrics, NR showed the strongest positive correlation with TP (휌=0.26), although the relationship was modest. This finding aligns with previous studies, which suggest that running code is often associated with higher task performance [4]. However, NR was also negatively correlated (휌=−0.47) with ADT. The longer students look at suggestions before acting, the less they ran their code. One possible explanation is that some students quickly ac- cepted suggestions with minimal engagement only to use NR to verify the logic. 4.1.2Average Dwell Time. ADT negatively correlated (휌=−0.26) with FCR, which could mean the less time students spend on a suggestion, the more likely they are to have high FCR. This suggests that students who spent more time looking at a suggestion passed more attention checks. However, among all the metrics we logged, ADT had the strongest negative correlation (휌=−0.17) with TP, To Tab or Not to Tab suggesting that the less time students spend on a suggestion, the higher their TP. This indicates that DT does not seem to directly relate to critical engagement. 4.1.3 Takeaway: These findings suggest that there is no perfect measure of responsible AI use. Even behaviors often associated with more expert use, such as NR, were positively correlated with FCR, indicating some lapses in critical engagement. 4.2 RQ2: Attention Checks and Performance While no clear metric was strongly correlated with task perfor- mance, several metrics were correlated with unreflective engage- ment with suggestions (i.e.: failed attention checks). 4.2.1 Tab Accept Rate. TAR weakly negatively correlated (휌= − 0.10) with TP, suggesting that students with high TAR were slightly less likely to perform well. Furthermore, students who TAR were far more likely to fail attention checks (휌=0.73,푝<0.001). TAR was also strongly negatively correlated with ADT, (휌=−0.47) indicating that students with high TAR tended to have lower DT. 4.2.2Attention Checks and Task Performance. We investigated two measures of misuse: TP represents an implicit measure of potential misuse, whereas FCR offers a more explicit measure. FCR and TP had a weak negative correlation (휌= −0.17). This suggests that students who passed attention checks did not necessarily per- form better on the task. Therefore, FCR does not represent critical engagement, but it does suggest potential misuse. 4.2.3 Takeaway: While behaviors and attention checks were not strongly correlated with performance, we observed that some be- haviors, such as tab accept were strongly correlated with failed attention checks, suggesting a potential lapse in reflective engage- ment with suggestions. 4.3 Novel Behavioral Metrics Our log analysis also identified new behavioral metrics including: STA Slow Tab Accept: Starting a SA by typing out the suggestion manually, clicking ‘tab’ before completing the SA. SMSlow Accept then Modify: Completing the SA, but then imme- diately modifying the suggestion. SDSlow Accept then Delete: Completing a SA, but then deleting all or some of the suggestion. These metrics were somewhat common in our dataset. Of the 71 SA events, we observed 10 STA events and 53 SM events. STA and SM accounted for 75% of our SA, making it a majority. SD was not observed in our log data, but logically follows from the possibility of SM. These metrics highlight that a lot can happen in the process of accepting a suggestion. With the inclusion of these metrics, our taxonomy becomes even more robust in capturing how students interact with code suggestions. These behaviors indicate that suggestion acceptance is not a terminal action, but often part of a multi-step editing process in- volving revision and correction. Suggestions have a life that extends beyond when they are accepted as students may return to them at various points and make modifications or delete them. 5 Discussion Across recent studies of AI code completion tools, students’ behav- iors and strategies have been observed qualitatively [1,27,28]. Our work extends this literature by operationalizing these behaviors into a taxonomy of measurable interaction metrics and demon- strates how these metrics can be captured quantitatively at scale using Clover. Consistent with these previous qualitative findings, we observed substantial diversity in how students engage with AI code suggestions. Our findings also suggest that while metrics are weakly or inconsistently correlated with performance, some behaviors tend to be more consistently associated with unreflective engagement with suggestions (i.e.: failed attention checks). 5.1Interaction Metrics for AI Code Suggestions 5.1.1Interactions patterns varied across participants. We observed variation across all behavioral metrics including acceptance behav- iors, dwell time, code execution frequency, and post-acceptance revision behaviors. This heterogeneity aligns with insights from learning sciences that describe novice learning as non-linear and diverse in strategies, rather than a stable progression through stages of expertise. Students may shift fluidly between exploratory and exploitative strategies to solve the problem at hand. This variability has already been well documented in recent studies of AI code completion tools [28, 31]. These metrics also have the potential to operationalize multiple aspects. For example, average dwell time (ADT) might relate to both cognitive processing and expertise. From dual-process theory [12], shorter dwell times may indicate rapid, intuitive judgments, while longer dwell times may indicate more deliberate reflective reason- ing. Alternatively, considering expertise, short dwell times may reflect negative expertise [22], where experienced students rapidly recognize a low-quality suggestion and reject it. Novices may spend more time attempting to interpret the suggestion, or may engage superficially without meaningfully evaluating it. 5.1.2Interaction Metrics and Task Performance. Across the behav- ioral metrics studied, we did not observe a strong or consistent relationship with task performance. While several behaviors which are traditionally associated with productive debugging and learning, such as running code, had modest associations with performance, these relationships were generally weak or inconsistent. This sug- gests that behaviors typically associated with good performance may coexist with lapses in critical engagement. 5.1.3Interaction Metrics and Failed Attention Checks. In contrast, several behaviors showed more consistent patterns associated with potential lack of reflective engagement with suggestions. For exam- ple, high rates of tab acceptance were strongly associated with failed attention checks. This relationship was substantially stronger than any observed association between other interaction metrics and task performance. Slow accept was the only behavior nega- tively correlated with failed attention checks. Dwell time had a weaker but consistent pattern. Longer dwell times were associated with better attention check performance and slightly improved task performance. This suggests that looking more carefully at suggestions tends to improve performance. These findings suggest that interaction patterns, such as frequent tab acceptance and to a Jessica Hutchison et al. lesser extent shorter dwell times, may serve as indicators of reduced engagement with AI-generated code suggestions. 5.2 A Taxonomy of Interaction Metrics As shown in Section 3.4, there has been a need to systematize the behavioral metrics associated with AI code completion tools. These metrics often used different terms or were measured differently across studies. Our work contributes to this goal by formalizing a taxonomy of metrics capturing how students engage with AI code completion tools. This paper offers an example of how this taxonomy might be used to better understand student-AI interac- tions, but future work could more extensively interrogate sequential behaviors and patterns of interaction over time. 5.3 Limitations This study faces a number of limitations, which are appropriate given our goal to develop and prototype a taxonomy of behavioral metrics for how students use AI code completion tools. 5.3.1Single Problem and Task Performance. While the rainfall prob- lem used in this study has been commonly studied in computing education research, future work should investigate multiple diverse problem types. Task performance on the problem contained lim- ited variability, with some participants producing non-compiling solutions and others receiving fully correct outputs which resulted in a lack of intermediate performance levels. 5.3.2 Single Model and Latency. Our system relied on a single AI model, but there are significant differences in performance and latency based on the model selection. 5.3.3 Post-Accept Behaviors. We measure dwell time by calculat- ing the time between the suggestion being shown and the users’s first interaction. However, students may continue engaging with suggestions long after they have been accepted. These longer-term behaviors were not analyzed in this study. 5.3.4 Participant Sample. Participants varied in their experience with programming and with GitHub Copilot, but they only repre- sent the students from a single class within a single university. 5.3.5Potential for Cheating. This study was conducted in a class- room setting, and students received participation credit that did not depend on their performance. Instructions also emphasized the importance of genuine effort, but we cannot rule out the use of external tools such as Google or ChatGPT. 5.3.6 Suggestion Types and Attention Checks. Clover’s attention checks injected deterministic incorrect suggestions to probe critical engagement. While this approach allows controlled measurement of misuse, it does not fully replicate spontaneous AI hallucinations in real-world coding. Moreover, the study did not account for the type or complexity of suggestions students received, which may influence acceptance, modification, or rejection behaviors. 5.4 Future Work Our work opens a number of potential research threads. 5.4.1 Suggestion Lifecycles. Our work focuses on atomic interac- tion events, but students may continue to interact with suggestions long after they are accepted. Future work could investigate these longer-term interactions with suggestions. 5.4.2Partial Accepts. Our current implementation of slow accept only includes exact matches to code suggestions. Future work could explore partial matches by measuring the similarity between the suggested and accepted code using Levenshtein distance and cosine similarity. This will provide a more nuanced characterization of slow accepts and whether they reflect verification, partial reuse, or adoption of a new approach. 5.4.3 Attention Checks. In this study, attention checks were used as a behavioral probe to assess students’ engagement with code suggestions. Future work could explore shifting attention checks from a passive metric toward a feedback mechanism. Students could receive real-time or summative feedback about how they engage with the attention checks. 5.4.4 Incorporating Student Demographics. Prior work suggests that programming background, AI literacy, and self-efficacy may influence how students interact with AI assistants [20,28]. Future studies should investigate how learner characteristics and demo- graphics affect interaction patterns and behaviors. Incorporating demographic information with interaction logs could explain varia- tions and patterns in acceptance strategies, rejection behavior, and engagement for a diverse student population. 5.4.5 Longitudinal Studies. Our current study analyzes interac- tions within a single task, but over time, students’ behaviors may change. Studying interactions across multiple assignments or through- out the duration of a course could reveal fluctuations in students engagement with AI-generated suggestions and determine the ef- fect of problem type on students’ behaviors. 6 Conclusion In this work, we introduced Clover, an AI code completion tool that captures behavioral metrics and uses attention checks to probe critical engagement. Our study with CS1 students shows that en- gagement with code suggestions is highly variable. Metrics may reflect a complex mix of reflective reasoning, over-reliance, or nega- tive expertise. As a result, no single metric reliably captures critical engagement. By formalizing a taxonomy of behavioral metrics, we provide a framework for interpreting patterns of student-AI interaction which can create a path for future research. References [1] Shraddha Barke, Michael B. James, and Nadia Polikarpova. 2023. Grounded Copilot: How Programmers Interact with Code-Generating Models. Proc. ACM Program. Lang. 7, OOPSLA1, Article 78 (April 2023), 27 pages. https://doi.org/10. 1145/3586030 [2]Seth Bernstein, Ashfin Rahman, Nadia Sharifi, Ariunjargal Terbish, and Stephen MacNeil. 2025. Beyond the Benefits: A Systematic Review of the Harms and Consequences of Generative AI in Computing Education. In Proceedings of the 25th Koli Calling International Conference on Computing Education Research. https://doi.org/10.1145/3769994.3770036 [3]Neil C.C. Brown and Amjad Altadmri. 2014. Investigating novice programming mistakes: educator beliefs vs. student data. In Proceedings of the Tenth Annual Conference on International Computing Education Research. 43–50. https://doi. org/10.1145/2632320.2632343 [4]Adam S. Carter, Christopher D. Hundhausen, and Olusola Adesope. 2015. The Normalized Programming State Model: Predicting Student Performance in Com- puting Courses Based on Programming Behavior. In Proceedings of the Eleventh Annual International Conference on International Computing Education Research To Tab or Not to Tab (Omaha, Nebraska, USA) (ICER ’15). Association for Computing Machinery, New York, NY, USA, 141–150. https://doi.org/10.1145/2787622.2787710 [5]John Edwards, Arto Hellas, and Juho Leinonen. 2025. On the Opportunities of Large Language Models for Programming Process Data. In Proceedings of the 27th Australasian Computing Education Conference (ACE 2025). 105–113. https://doi.org/10.1145/3716640.3716652 [6] Jonathan St. B. T. Evans. 2008. Dual-processing accounts of reasoning, judgment, and social cognition. Annual Review of Psychology 59 (2008), 255–278. https: //doi.org/10.1146/annurev.psych.59.103006.093629 [7]Ge Gao, Samiha Marwan, and Thomas W Price. 2021. Early performance predic- tion using interpretable patterns in programming process data. In Proceedings of the 52nd ACM technical symposium on computer science education. 342–348. https://doi.org/10.1145/3408877.3432439 [8]Lydia Harbarth, Eva Gößwein, Daniel Bodemer, and Lenka Schnaubert. 2025. (Over) trusting AI recommendations: How system and person variables affect di- mensions of complacency. International Journal of Human–Computer Interaction 41, 1 (2025), 391–410. https://doi.org/10.1080/10447318.2023.2301250 [9] Irene Hou, Sophia Mettille, Owen Man, Zhuo Li, Cynthia Zastudil, and Stephen MacNeil. 2024. The Effects of Generative AI on Computing Students’ Help- Seeking Preferences. In Proceedings of the 26th Australasian Computing Education Conference (ACE ’24). ACM, 39–48. https://doi.org/10.1145/3636243.3636248 [10] Irene Hou, Hannah Vy Nguyen, Owen Man, and Stephen MacNeil. 2025. The Evolving Usage of GenAI by Computing Students. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2. 1481–1482. https://doi.org/10.1145/3641555.3705266 [11] Dhanya Jayagopal, Justin Lubin, and Sarah E. Chasins. 2022. Exploring the Learnability of Program Synthesizers by Novice Programmers. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. https://doi.org/10.1145/3526113.3545659 [12] Daniel Kahneman. 2011. Thinking, Fast and Slow. Farrar, Straus and Giroux. https://doi.org/10.1007/s00362-013-0533-y [13] Daniil Karol, Elizaveta Artser, Ilya Vlasov, Yaroslav Golubev, Hieke Keuning, and Anastasiia Birillo. 2025. KOALA: A Configurable Tool for Collecting IDE Data When Solving Programming Tasks. In Proceedings of the ACM Global on Comput- ing Education Conference 2025 Vol 1 (CompEd 2025). Association for Computing Machinery, 183–189. https://doi.org/10.1145/3736181.3747129 [14] Majeed Kazemitabaar, Xinying Hou, Austin Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman. 2024. How Novices Use LLM-based Code Gen- erators to Solve CS1 Coding Tasks in a Self-Paced Learning Environment. In Proceedings of the 23rd Koli Calling International Conference on Computing Edu- cation Research. 1–12. https://doi.org/10.1145/3631802.3631806 [15]Hieke Keuning, Isaac Alpizar-Chacon, Ioanna Lykourentzou, Lauren Beehler, Christian Köppe, Imke de Jong, and Sergey Sosnovsky. 2024. Students’ Percep- tions and Use of Generative AI Tools for Programming Across Different Comput- ing Courses. In Proceedings of the 24th Koli Calling International Conference on Computing Education Research. 1–12. https://doi.org/10.1145/3699538.3699546 [16]Shao-Heng Ko and Kristin Stephens-Martinez. 2025. Rethinking Computing Students’ Help Resource Utilization through Sequentiality. ACM Trans. Comput. Educ. 25 (2025), 1–34. https://doi.org/10.1145/3716860 [17]Juho Leinonen et al.2019. Keystroke data in programming courses. Department of Computer Science, Series of Publications A (2019). [18]Stephen MacNeil, James Prather, Rahad Arman Nabid, Sebastian Gutierrez, Silas Carvalho, Saimon Shrestha, Paul Denny, Brent N. Reeves, Juho Leinonen, and Rachel Louise Rossetti. 2025. Fostering Responsible AI Use Through Negative Expertise: A Contextualized Autocompletion Quiz. In Proceedings of the 30th ACM Conference on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2025). ACM. https://doi.org/10.1145/3724363.3729067 [19]Stephen MacNeil, Magdalena Rogalska, Juho Leinonen, Paul Denny, Arto Hel- las, and Xandria Crosland. 2024. Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students. In Proceedings of the 2024 on ACM Virtual Global Computing Education Confer- ence V. 1 (SIGCSE Virtual 2024). Association for Computing Machinery, 137–143. https://doi.org/10.1145/3649165.3690100 [20]Lauren E. Margulieux, James Prather, Brent N. Reeves, Brett A. Becker, Gozde Cetin Uzun, et al.2024. Self-Regulation, Self-Efficacy, and Fear of Failure In- teractions with How Novices Use LLMs to Solve Programming Problems. In Proceedings of the 2024 on Innovation and Technology in Computer Science Educa- tion V. 1 (ITiCSE 2024). ACM, 276–282. https://doi.org/10.1145/3649217.3653621 [21]Pratibha Menon. 2023. Exploring GitHub Copilot assistance for working with classes in a programming course. Issues in Information Systems 24, 4 (2023). [22] Marvin Minsky. 1997. Negative expertise. (1997). [23]Seong Min Park, Marco Ho, Michael Pin-Chuan Lin, and Jeeho Ryoo. 2025. Evaluating the Impact of Assistive AI Tools on Learning Outcomes and Ethical Considerations in Programming Education. In 2025 IEEE Global Engineering Education Conference (EDUCON). 1–10. https://doi.org/10.1109/EDUCON62633. 2025.11016517 [24] Leo Porter and Daniel Zingaro. 2024. Learn AI-Assisted Python Programming: With Github Copilot and ChatGPT. Simon and Schuster. [25]James Prather, Paul Denny, Juho Leinonen, Brett A. Becker, Ibrahim Albluwi, et al.2023. The Robots Are Here: Navigating the Generative AI Revolution in Computing Education. In Proceedings of the 2023 Working Group Reports on Innovation and Technology in Computer Science Education (ITiCSE-WGR ’23). Association for Computing Machinery. https://doi.org/10.1145/3623762.3633499 [26]James Prather, Juho Leinonen, Natalie Kiesler, Jamie Gorson Benario, et al.2025. Beyond the Hype: A Comprehensive Review of Current Trends in Generative AI Research, Teaching Practices, and Tools. In 2024 Working Group Reports on Innovation and Technology in Computer Science Education (Milan, Italy) (ITiCSE 2024). Association for Computing Machinery, New York, NY, USA, 300–338. https://doi.org/10.1145/3689187.3709614 [27]James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. “It’s Weird That it Knows What I Want”: Usability and Interactions with Copilot for Novice Programmers. ACM Trans. Comput.-Hum. Interact. 31, 1, Article 4 (Nov. 2023), 31 pages. https://doi.org/10.1145/3617367 [28] James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Ran- drianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Pro- grammers. In Proceedings of the 2024 ACM Conference on International Computing Education Research - Volume 1 (ICER ’24). Association for Computing Machinery, 469–486. https://doi.org/10.1145/3632620.3671116 [29] Thomas W. Price, David Hovemeyer, Kelly Rivers, Ge Gao, et al.2020. ProgSnap2: A Flexible Format for Programming Process Data. In Proceedings of the 2020 ACM Conference on Innovation and Technology in Computer Science Education (ITiCSE ’20). ACM, 356–362. https://doi.org/10.1145/3341525.3387373 [30] Anshul Shah, Anya Chernova, Elena Tomson, Leo Porter, William G. Griswold, and Adalbert Gerald Soosai Raj. 2025. Students’ Use of GitHub Copilot for Work- ing with Large Code Bases. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSETS 2025). Association for Computing Machinery, 1050–1056. https://doi.org/10.1145/3641554.3701800 [31] Md Istiak Hossain Shihab, Christopher Hundhausen, Ahsun Tariq, Summit Haque, Yunhan Qiao, and Brian Wise Mulanda. 2025. The Effects of GitHub Copilot on Computing Students’ Programming Effectiveness, Efficiency, and Processes in Brownfield Coding Tasks. In Proceedings of the 2025 ACM Conference on Interna- tional Computing Education Research V.1 (ICER ’25). Association for Computing Machinery, 407–420. https://doi.org/10.1145/3702652.3744219 [32] C. Estelle Smith, Kylee Shiekh, Hayden Cooreman, Sharfi Rahman, et al.2024. Early Adoption of Generative Artificial Intelligence in Computing Education: Emergent Student Use Cases and Perspectives in 2023. In Proceedings of the 2024 Innovation and Technology in Computer Science Education V.1 (ITiCSE 2024). ACM, 3–9. https://doi.org/10.1145/3649217.3653575 [33]Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024). ACM, 297–303. https://doi.org/10.1145/3649217.3653584 [34]Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22). Association for Computing Machinery, Article 332, 7 pages. https://doi.org/10.1145/3491101.3519665 [35]Michel Wermelinger. 2023. Using GitHub Copilot to Solve Simple Programming Problems. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2023). Association for Computing Machinery, New York, NY, USA, 172–178. https://doi.org/10.1145/3545945.3569830 [36]Cynthia Zastudil, Magdalena Rogalska, Christine Kapp, Jennifer Vaughn, and Stephen MacNeil. 2023. Generative AI in Computing Education: Perspectives of Students and Instructors. In 2023 IEEE Frontiers in Education Conference (FIE). 1–9. https://doi.org/10.1109/FIE58773.2023.10343467