Paper deep dive
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
Shalaleh Rismani, Su Lin Blodgett, Q. Vera Liao, Alexandra Olteanu, AJung Moon
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 3:03:44 AM
Summary
This study investigates how functional and structural mental models of AI writing assistants influence user behavior, control, and output quality. Through a controlled experiment with 48 participants using a modified CoAuthor platform, the researchers found that while structural mental models improve perceived ease of use, they can lead to a 'backfiring' effect where users exhibit overtrust, resulting in higher acceptance of erroneous AI suggestions and lower overall writing quality.
Entities (5)
Relation Signals (3)
CoAuthor → usedin → Cover Letter Writing Task
confidence 100% · asking them to complete a cover letter writing task using a writing assistant
Structural Mental Model → increases → Perceived Ease of Use
confidence 95% · participants in the structural mental model condition... judged the system as more usable
Structural Mental Model → influences → User Oversight
confidence 90% · participants in the structural mental model condition... produced letters with more grammatical errors
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI-based writing assistants are ubiquitous, yet little is known about how users' mental models shape their use. We examine two types of mental models -- functional or related to what the system does, and structural or related to how the system works -- and how they affect control behavior -- how users request, accept, or edit AI suggestions as they write -- and writing outcomes. We primed participants ($N = 48$) with different system descriptions to induce these mental models before asking them to complete a cover letter writing task using a writing assistant that occasionally offered preconfigured ungrammatical suggestions to test whether the mental models affected participants' critical oversight. We find that while participants in the structural mental model condition demonstrate a better understanding of the system, this can have a backfiring effect: while these participants judged the system as more usable, they also produced letters with more grammatical errors, highlighting a complex relationship between system understanding, trust, and control in contexts that require user oversight of error-prone AI outputs.
Tags
Links
- Source: https://arxiv.org/abs/2604.05166v1
- Canonical: https://arxiv.org/abs/2604.05166v1
Trouble viewing inline? Open PDF directly →
Full Text
138,355 characters extracted from source content.
Expand or collapse full text
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants Shalaleh Rismani ∗ McGill University Montréal, Canada Su Lin Blodgett Microsoft Research Montréal, Canada Q. Vera Liao University of Michigan Ann Arbor, Michigan, USA Alexandra Olteanu Microsoft Research Montréal, Canada AJung Moon McGill University Montréal, Canada Abstract AI-based writing assistants are ubiquitous, yet little is known about how users’ mental models shape their use. We examine two types of mental models—functional or related to what the system does, and structural or related to how the system works—and how they affect control behavior—how users request, accept, or edit AI sugges- tions as they write—and writing outcomes. We primed participants (푁=48) with different system descriptions to induce these men- tal models before asking them to complete a cover letter writing task using a writing assistant that occasionally offered preconfig- ured ungrammatical suggestions to test whether the mental models affected participants’ critical oversight. We find that while partici- pants in the structural mental model condition demonstrate a better understanding of the system, this can have a backfiring effect: while these participants judged the system as more usable, they also pro- duced letters with more grammatical errors, highlighting a complex relationship between system understanding, trust, and control in contexts that require user oversight of error-prone AI outputs. CCS Concepts • Human-centered computing→Empirical studies in HCI; User studies. Keywords Oversight, human-AI interaction, AI-based writing assistants, sys- tem safety, user control, mental models ACM Reference Format: Shalaleh Rismani, Su Lin Blodgett, Q. Vera Liao, Alexandra Olteanu, and AJung Moon. 2026. From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26), April 13–17, 2026, Barcelona, Spain. ACM, New York, NY, USA, 23 pages. https: //doi.org/10.1145/3772318.3791670 ∗ Corresponding author: shalaleh.rismani@mail.mcgill.ca This work is licensed under a Creative Commons Attribution 4.0 International License. CHI ’26, Barcelona, Spain © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2278-3/2026/04 https://doi.org/10.1145/3772318.3791670 1 Introduction As Artificial Intelligence (AI) systems are increasingly integrated into everyday tasks, what must users understand about these sys- tems in order to engage with them effectively and appropriately? Safe operation of complex, safety-critical technological systems (e.g., airplanes) typically demands their operators (e.g, pilots) to have a high degree of understanding of how a system works and how it can be used [45,66,68]. Prior work in human-computer interaction and system safety provides empirical evidence as to why users with more complete and accurate mental models—their understanding of how a system works and how it can be used—are better equipped to monitor system behavior, anticipate failures, and intervene in near-incident scenarios [23, 68, 77]. While these findings are well established in settings that have traditionally been deemed safety-critical (e.g., aviation), it remains unclear if, how, and when they translate to settings where gener- ative AI systems are being used for everyday tasks (e.g., writing). Today, a variety ofAIsystems (e.g., powered by Large Language Models (LLMs)) are integrated into writing platforms to provide writing support—for example, by generating next-sentence sug- gestions [43]—in order to improve efficiency and accessibility, es- pecially for more inexperienced users [24,30,49]. In contrast to more traditional safety-critical systems, however,AI-based writ- ing assistants differ in key ways: their failures can be subtle or less salient; the consequences of such failures often do not lead to severe harm immediately [8,35,49]; and developers often aim to make it frictionless and easier for users to write rather than encouraging agency and control [43,62]. This is not to imply that failures of these systems and their consequences in everyday writing tasks are inconsequential or that they should be accepted as the norm. Al- though subtle, failures and consequences ofAI-assisted writing can include biased [35], inaccurate or incoherent [70], and low quality writing outputs [2,48]. The use of such systems can also interfere with the writer’s voice [21,37,64]. These can all pose threats to a sense of authorship, ownership, and authenticity that writers con- tinue to demand and value even as they adoptAItools [33,37,64]. Existing work on recommendation systems [42] and predictive algorithms [4] has shown that more accurate and thorough mental models ofAIsystems can improve users’ satisfaction and trust, as well as support better decision making in task selection and prediction settings. By extension, one can hypothesize that users with more expansive mental models ofAI-based writing assistants should be able to more effectively use and control these systems towards desired outcomes. Yet the role of such mental models in arXiv:2604.05166v1 [cs.HC] 6 Apr 2026 CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. shaping interaction with LLM-based writing assistants remains underexplored, particularly in situations where outputs from such systems may be problematic, irrelevant, or erroneous. To address this gap, we design and conduct a in-person between-subjects exper- iment (푁=48) to explore how users with different mental models of the same system—shaped by alternative descriptions—engage withAI-based writing assistants. Our study and experiment design is guided by the following research questions: •RQ1: How do different mental models affect the way users exert control in the writing process? •RQ2: How do different mental models held by users influence the quality of the final written output? • RQ3: How do these mental models shape users’ overall ex- perience of writing with the AI-based writing assistant? To address these questions, we leveraged existing frameworks where mental models can be described as either functional or struc- tural [42]. While a functional mental model only involves an un- derstanding of how to use the system, a structural mental model builds on this understanding and includes knowledge of how the system works. For our experiment, we use a modified version of the open-source CoAuthor platform [44], which providesAI-based next-sentence suggestions and supports meta-prompting. Prior to a cover letter writing task, we primed participants’ mental models by showing them one of two videos: one on how to use the system (functional), and one also explaining how it works (structural). To better observe how participants exert control over the system, we intentionally introduced suggestions containing spelling or gram- matical errors. We then analyzed participants’ interactions with CoAuthor, final written work, and self-reported experiences from a post-task survey and semi-structured interview. Within this staged erroneous output context, our results reveal a key tension between system understanding, trust, and oversight. Participants with struc- turally richer mental models reported higher perceived ease of use, but they produced letters with higher grammatical errors and tended to accept a higher rate of flawed suggestions. This pattern may reflect a form of overtrust, in which increased confidence in understanding the system results in users placing greater reliance on its suggestions than is warranted. Contributions. First, we present an empirical case study of how different mental models about anAIwriting assistant affect users’ interactions with the system, particularly in the presence of low- quality suggestions. We further contribute a controlled experi- mental approach for manipulating and testing users’ functional and structural mental models through instructional priming under staged system failures. Second, we provide a baseline empirical account showing that while a deeper system understanding may increase perceived ease of use and foster greater trust, it may also lead users to uncritically accept staged flawed suggestions that contain spelling and grammatical errors, resulting in lower writing quality, motivating the need for future studies to further examine connection between mental models, system failure, oversight, and trust inAI-assisted writing. Third, drawing from qualitative ob- servations, we highlight the central role of interaction affordances in shaping perceptions of control and ownership as opposed to mental model understanding; users’ sense of agency may therefore be driven more by what systems allow them to do than by what they know about how those systems work. 2 Background This section reviews prior work on mental models, user control, and oversight from the perspectives of Human-Computer Interaction (HCI) and system safety, as well as research on in AI-based writing assistants, to contextualize our study. 2.1 Mental Models and Their Role in Human–AI Interaction Mental models have been a central concept in fields likeHCI[15,69] and system safety [9,17,72]. Mental models are internal represen- tations that people form of target systems (e.g., computer) when interacting with them. In cognitive psychology, mental models are understood to be partial, evolving, and sometimes inaccurate representations that people use to reason about complex systems under uncertainty [36,55]. Because these representations shape how users anticipate system states and interpret feedback, they directly influence trust, reliance, and decision making.HCIand sys- tem safety literature has long adopted this concept to explain how people learn, use, and reason about interactive systems [15,61], emphasizing that mismatches between users’ mental models and a system’s actual behavior can hinder effective interaction and lead to erroneous outcomes, as discussed further in the next section. Prior work shows that users’ mental models of AI systems are multifaceted and can reflect both functional expectations and deeper understanding about system behavior. People construct these mod- els by combining explicit information, such as model explanations, with inferences drawn from their use, often producing simplified representations that support everyday reasoning. When examining users’ mental models of recommendation systems, Kulesza et al. [42]distinguish between functional mental models, which represent how the system could be used, versus structural mental models, which describe how the system works. Gero et al. [27]examine the mental models users form ofAIagents in a collaborative game and conclude that users’ mental models of anAIagent have three components: its behavior at a large scale, its knowledge of various topics, and its behavior at the scale of individual output. Recent re- search has examined users’ mental models ofLLM- and generative AI-based systems [51,76]. For example, Mehmood et al. [51]docu- ment heterogeneous mental models ofLLMs at both individual and societal levels, ranging from optimistic views ofLLMs as helpful tools to more skeptical views of them as suspicious actors. Mental models influence how users interact with anAIsys- tem and shape their experience and the final output. Research in HCIand human–machine trust further shows that users’ expecta- tions about a system’s capabilities and limitations strongly affect how much they trust and rely on it, with inaccurate or incomplete models leading to both over-reliance and unwarranted skepticism [3,25,53,74]. Prior work on predictive algorithms has shown that persuasive explanations can lead users to form mental models that promote over-reliance or under-trust [4–6,56]. For example, Bansal et al. [4,5]argue that explanations should prioritize informative- ness over persuasion to help users develop calibrated trust and avoid inappropriate reliance. Recent studies ofLLMs show that From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain when users hold incomplete or inaccurate expectations about how these systems behave, they may rely on outputs in ways that are not well aligned with task needs, such as accepting suggestions too readily [10,26,39]. Work on recommendation systems has further explored how strengthening a user’s mental model can improve their experience; for example, Kulesza et al. [42]found that pro- viding structural scaffolding—offering insight into how a music recommendation system worked—helped users build more accurate mental models and led to higher user satisfaction with the system’s outputs. These works establish that mental models shape how users interact withAIsystems, the degree to which they trust and rely on these systems, and ultimately how these factors influence both their perceived experience and interaction outcomes. In the next section, we further elaborate on the connection between mental models and users’ ability to exercise control and oversight, focusing specifically on the case of AI-based writing assistants. 2.2 User Control, Oversight, and AI-based Writing Assistants AsAIsystems become integrated into a wide range of applica- tions, regulatory frameworks such as the EUAIAct increasingly emphasize the need for effective human oversight: a user’s abil- ity to notice, question, and intervene when system outputs may be inappropriate or harmful [22]. System safety research provides a longstanding foundation for understanding oversight, defining user control as the ability to intentionally influence, override, or intervene on system behavior to ensure that it operates safely and as intended [47]. Inconsistent, inadequate, and incomplete user mental models of a system are recognized as key factors in hinder- ing their ability to exert effective control [47]. Decades of work in safety-critical domains—such as aviation, medical devices, and au- tonomous vehicles—demonstrate that how well a user understands a technological system impacts how they interact with and control it [13,19,20]. For example, drivers who understand the operational logic of adaptive cruise control are better able to override unsafe behaviors than those who only know how to operate its interface [23]. Given these illustrations of how inadequate and inconsistent mental models can lead to harm, system safety emphasizes the importance of designing environments and feedback mechanisms that help users recognize hazardous conditions—system states or interactions that increase the likelihood of harm—, monitor their actions, and recover prior to occurrence of an incident [19, 47]. Given the importance of effective oversight, together with ex- tensive evidence that users’ mental models shape trust, reliance, outcomes, and user experience, it is critical to examine how men- tal models relate specifically to oversight of AI systems. In this work, we focus on this question in the context ofAI-based writ- ing assistants, which are now widely used across domains ranging from creative writing to professional and scientific communication [1,28,30,32,34,38,52,78]. In the domain ofAI-based writing as- sistants, these oversight challenges surface through cognitive and epistemic risks rather than physical ones. Specifically,AI-generated suggestions can influence what users write and how they represent themselves. For instance, Poddar et al. [60]show thatAI-based writing assistants can shift self-presentation in personal bios, and Jakesch et al. [35]demonstrate thatAIsuggestions can sway opin- ions in argumentative writing—even when misaligned with a user’s initial stance. Such influence may occur without users fully recog- nizing its extent, raising questions about users’ ability to oversee AI-generated content and maintain ownership. However, this line of work does not examine whether such influence varies as a function of users’ underlying mental models. Complementary work inHCIfurther shows that interface design can either support or undermine user control. Offering multiple parallel suggestions can aid idea generation but may also increase decision fatigue [14], while diegetic prompts can enable more intu- itive interaction [18]. Although these studies highlight the intrica- cies of designing writing assistant interfaces that support effective control, they primarily focus on interaction mechanics or user pref- erences rather than on how users’ mental models influence their engagement with the system or how different control mechanisms shape those mental models. Despite growing literature on influence, interaction patterns, and user experience inAI-based writing assistants, prior work has not examined whether users with different mental models demonstrate different levels of control and oversight. Building on Draxler et al.’s definition of objective control—“the degree of influence that users have over theAI-generated text, e.g., by employing interaction methods” [21]—we interpret these interaction behaviors as concrete indicators of user control and oversight. This aligns with system safety perspectives, where user control refers to a user’s ability to actively influence or override system behavior to ensure task success and mitigate errors before harm occurs. Building on this framing, we examine how mental models shape users’ ability to exercise control and oversight inAI-based writing assistants; the next two sections outline our study design. 3 Study Overview and Hypotheses Drawing on prior work in system safety andHCI, in this study we distinguish between functional (focused on how to use a system) and structural (focused on how the system works) mental models to design a controlled experiment [27,42]. Participants were asked to produce a high-quality cover letter under a time constraint using anAI-based writing assistant. We chose this task because cover letters are a high-stakes, consequential form of writing, where er- rors, tone, and content quality can affect real-world outcomes such as job opportunities. The time pressure was introduced to encour- age active engagement with the writing assistant. To better isolate the effects of mental model differences, we recruited participants with minimal prior experience usingAI-based writing assistants. To study how participants exercise oversight in a controlled and comparable way, we deliberately injected grammatical and spelling errors into the writing assistant’s suggestions. We chose this er- ror type, rather than issues such as inappropriate tone or factual issues, because grammatical and spelling mistakes are relatively clear, easy to fix, and less subject to disagreement among partici- pants. This allows us to see whether users actively scrutinize and refineAI-generated suggestions rather than passively accept them. Further details about the experimental procedure, the recruitment CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. process, the platform and the writing tasks are provided in Sec- tion 4. We present our research questions, expected outcomes, and corresponding hypotheses in Table 1. For control behavior (RQ1), we expect that participants with a structural mental model will demonstrate more deliberate control over the system. Following Draxler et al.’s definition of control as the user’s influence overAI-generated text [21], this behavior can manifest through (1) requesting, accepting, or editingAIsugges- tions in the text editor and (2) providing explicit instructions in the Instruction Box (refer to Section 4.4 for a description of the plat- form). Because structural participants understand how the system works and what it can and cannot do, we expect them — especially under time pressure to complete the task — to request more sugges- tions, steer those suggestions through targeted instructions, accept more of them, and then make edits to refine the output to their exact preferences, thereby ensuring higher quality in the final cover letter. We do not expect the overall proportion of text selected fromAI- generated suggestions versus written by participants differ across both conditions. For writing quality (RQ2), we hypothesize that par- ticipants with a structural mental model will produce higher-quality cover letters. Prior work suggests that users who better understand system functioning can more effectively guide outputs [42]. We therefore expect structural participants to exercise stronger over- sight, accept fewer erroneous suggestions, and ultimately produce cover letters with fewer errors and higher overall quality. For per- ceptions and experience (RQ3), we anticipate that structural partic- ipants will report more positive evaluations, including greater confi- dence in the final output, a stronger sense of control and ownership, and higher ratings of ease of use. All the corresponding constructs, data sources, and measures are outlined in Section 4.6 and Table 4. 4 Study Design and Methodology We conducted an in-person controlled experiment where we asked participants (푁=48) to write a cover letter using anAI-based writing assistant, a modified version of the CoAuthor platform [44]. Participants were randomly assigned to view one of two video descriptions of the assistant, which were designed to elicit dis- tinct mental models: (1) a functional description that focused solely on how to use the platform, and (2) a structural description that additionally explained how the platform works. For this, we delib- erately recruited individuals who self-reported limited familiarity withAI-based writing assistants, allowing us isolate the effects of our experimental manipulation. To evaluate participants’ ability to exercise critical oversight, we intentionally embedded erroneous suggestions (e.g., grammar and spelling mistakes) into theAI-based writing assistant’s suggestions at a fixed frequency. We compared participants’ behavioral data, cover letter quality, and self-reported experiences across the two mental model conditions. All partici- pants provided informed consent, and the study was approved by the Research Ethics Board of McGill University (REB #23-08-063; Mental Models and User Control When Interacting withAISystems). 4.1 Participant Recruitment Participants were recruited through posters, organizational mailing lists, and dedicated social media groups where individuals had pre- viously expressed interest in research studies. To inform our sample size, we conducted an in-person pilot study (n = 8) and calculated the ratio of accepted suggestions that contained grammatical errors to the total number of accepted suggestions. We observed a notable difference between the mental model conditions: for participants in the functional condition, 22% of accepted suggestions contained grammatical errors, versus 11% in the structural condition (effect size = 0.83). Based on this effect size, and assuming훼= .05 and power = .80, our power analysis indicated a minimum of 19 partici- pants per condition (38 total) for the main study. We recruited participants with prior experience with writing cover letters and regular use of English in academic or professional settings. We also selectively recruited individuals who reported limited understanding ofAI-based writing assistants, recognizing that experienced users often carry pre-existing mental models of such systems. By working with participants who had minimal prior understanding, we ensured that our experimental manipulation (two different system descriptions) could shape their mental mod- els, allowing us to study how these influenced engagement with the system. Eligibility was determined through four screening ques- tions: (1) experience writing cover letters (Yes/No); (2) daily use of English in professional or academic setting (Yes/No); (3) famil- iarity withAI-based writing assistants (Not at all familiar, Slightly familiar, Moderately familiar, Extremely familiar), and (4) level of understanding of howAI-based writing assistants work (Very poor, Poor, Fair, Very good, Excellent). Participants who answered “yes” to Questions 1–2, “not at all familiar” or “slightly familiar” for Ques- tion 3, and “very poor,” “poor,” or “fair” for Question 4 were eligible. Of the total of 295 individuals who completed the intake form, 72 met these criteria. From this pool, we scheduled 48 participants – 24 per mental model condition – for the in-person study, allowing for missing and outlier data points. Across the two conditions, the participants were demograph- ically comparable. The average age was 24.7 years (푀=24.7, 푆퐷=8.3) in the functional condition and 23.9 years (푀=23.9, 푆퐷=3.1) in the structural condition. Two-thirds of the participants identified as female (66.7%), 22.9% as male, 6.3% as non-binary, and 4.2% preferred not to answer, with nearly identical distributions be- tween conditions. Just over half of the participants (54.2%) reported English to be their native language. The functional condition had a higher proportion of native English speakers (66.7%) compared to the structural condition (41.7%). However, a chi-square test of inde- pendence indicated that this difference is not statistically significant, 휒 2 (1,푁=48)=3.021,푝= .082. Although this result falls short of conventional significance thresholds, the imbalance in native English speakers across conditions may still constitute a potential limitation. We note this here and considered it in subsequent anal- yses. Participants reported writing in English frequently (푀=4.67, 푆퐷=0.56, on a 5-point scale) and expressed high confidence in their English writing ability (푀=4.50,푆퐷=0.65). Self-assessed cover letter writing ability was rated between “fair” and “very good” on average (푀=3.60,푆퐷=0.68), and confidence in writing cover letters was rated as just above “fair” (푀=3.21,푆퐷=0.65). As expected, the use ofAI-based writing tools was generally infre- quent (푀=2.48,푆퐷=1.17), with no notable differences between conditions. From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain Table 1: Research questions, expected outcome, and hypotheses Research QuestionExpected outcomeHypotheses RQ1: How do different mental models affect the way users exert control in the writing process? Participants with a structural mental model will demonstrate more effective and deliberate control, steering the system rather than passively accepting suggestions. H1a: Structural participants will request more AI-generated suggestions. H1b: Structural participants will accept a greater proportion of suggestions. H1c: Structural participants will make more edits to the accepted suggestions. H1d: Structural participants will issue more instructions to steer suggestions. H1e: Structural and functional participants will contribute an equal amount of original text. RQ2: How do different mental models held by users influence the quality of the final written output? Participants with a structural mental model will produce higher-quality cover letters with fewer errors and better overall quality. H2a: Structural participants will achieve higher grammar correctness scores and accept fewer erroneous suggestions. H2b: Structural participants’ letters will receive higher overall quality ratings. RQ3: How do these mental models shape users’ overall experience of writing with the AI-based writing assistant? Participants with a structural mental model will report more positive perceptions of the writing process and system, including greater perceived control, ownership, and satisfaction. H3a: Structural participants will report higher perceived usefulness and ease of use. H3b: Structural participants will rate the quality of their final outputs more highly. H3c: Structural participants will report a stronger sense of control. H3d: Structural participants will report a greater sense of ownership. Introduction, consent form and demographics survey (5min) Familiarization period #1 (5 min) Functional or Structural video (5-10 min) Familiarization period #2 (5 min) Understanding survey (5 min) Writing task (20min) Post-task survey and interview (10-15 min) Figure 1: This figure illustrates the experimental design and procedure. White boxes indicate stages where participants completed various surveys. Gray boxes represent phases where participants either familiarized themselves with the platform or engaged in the writing task. The black box marks the manipulation point in the experiment. 4.2 Experimental Design Our experiment followed a between-subjects design with mental model (functional vs. structural) as the independent variable (see Figure 1). We operationalized this variable by priming participants with two different video descriptions about the modified CoAu- thor platform, one intended to prime a functional and the other a structural mental model. The study was conducted in person to ensure that participants engaged fully with the writing task without relying on external AI-based tools and to maintain experimental control by avoiding well-documented pitfalls of onlineHCIstudies such as reduced attention, multitasking, unmonitored tool use, and data-quality concerns in crowdsourced settings [41,57,58]. Participants first completed a survey capturing demographics, self-assessed English writing skills, and prior experience with cover letters andAI-based writing assistants. These data were collected to inform and contex- tualize our findings. For an initial familiarization, the experimenter briefly introduced the modified CoAuthor platform and guided the participants through the main functionalities such as writing in the editor and pressing tab to get suggestions. Next, participants were randomly assigned to watch one of two pre-recorded videos—one designed to elicit a functional mental model (how to use the system), and the other a structural mental model (how the system works), as described in Section 4.5. The functional video was 3:58 minutes, and the structural video was 8:54 minutes; although the videos differed in length, this reflected the additional explanation required to convey how the system works. Because our goal was to examine whether participants who received a more detailed explanation of the system’s underlying mechanisms would form different men- tal models and engage with the writing assistants differently, we accepted this difference in duration as appropriate for the manipula- tion. Afterward, participants explored the platform independently and could ask questions during a second familiarization period. Finally, before beginning the writing task, participants completed a brief survey to assess their understanding of the platform, which served as a manipulation check. At this point, participants in both conditions were presented with a job posting and had 20 minutes to write a cover letter using the modified CoAuthor platform, described in Sections 4.3 and 4.4. The allotted writing time was kept the same regardless of the length of the pre-recorded video across the two conditions, and participants in the structural condition were not under additional time pressure to complete the writing task. During the task, the platform injected grammatical and spelling errors at a predetermined frequency, al- lowing us to observe how participants with different mental models detected and addressed such mistakes. After the writing task, the participants completed a post-task survey to report on their writ- ing experience and perceptions of the platform and final output. This post-task survey was adapted from prior validated studies on writing assistants that examine perceived writing quality, user experience, sense of control, and ownership [14,18,21,37,44]. Lastly, the experimenter conducted a semi-structured interview with the participants to gather richer, qualitative insights into their experience of using the platform and writing the cover letter. The questions from the three surveys are included in Appendix A, B, and C. The experimenter script, including the post-task interview CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. questions, the mock job posting, the pre-recorded video slides, and the videos are all included in the supplemental material. 4.3 The Cover Letter Writing Task Participants were asked to write a high-quality cover letter for a Customer Service Representative position at a fictional company, Radiant Solutions—a technology solutions provider across various sectors. The cover letter writing task was chosen because it has a clear functional objective (to secure an interview) and audience (re- cruiter/potential employer). We asked participants to write a letter that illustrates their interest in the role and helps them secure an interview. Since all participants had prior experience writing cover letters, we were not prescriptive about what participants should consider as a high-quality letter. We instructed participants to write about themselves (rather than a fictional character), include at least one personal experience, and aim to fill the full span of the text editor, encouraging a letter of approximately 300 words. They were given 15 minutes to complete the task with the option to request an additional 5 minutes if needed. Lastly, the job posting was inten- tionally designed to be broad and accessible, enabling participants with diverse backgrounds, skill sets, and educational experiences to respond meaningfully. We initially generated the job posting using GPT-4o by prompting it to “create a generic job posting for a customer sales representative,” and then refined the output to align with the style and content of job postings returned when searching for customer sales representative positions on LinkedIn and Indeed. 4.4 Platform Description We adapted and extended the open-source CoAuthor platform [44] to create the writing assistant used in this study. In the modified plat- form, participants typed freely in the main text editor and pressed the Tab key whenever they wanted the system to propose continua- tions. The interface then displayed up to five candidate suggestions, and participants could choose to insert one into their draft or ig- nore them and continue writing. Participants could also optionally provide guidance through an Instruction Box that allowed them to customize the generated suggestions. The modified platform preserved CoAuthor’s core mechanism for generating suggestions. Each Tab press triggered an Applica- tion Programming Interface (API) call to a large language model, sending up to 4,000 tokens of recent editor text as a prompt. The model returned 15 candidate suggestions. We used the GPT-3.5- turbo-instruct model, which OpenAI recommended at the time of the study for sentence-level completions. After filtering duplicates and blocked content, five suggestions were randomly selected and displayed to the user (Figure 2). The modified platform was created by introducing two changes to the original CoAuthor system: (1) the controlled injection of suggestions containing grammatical and spelling errors at a fixed ratio, and (2) the addition of an Instruction Box, which provides a meta-prompting functionality allowing users to steer the generated suggestions. We describe these two modifications next. 4.4.1Modification 1: Grammatical and spelling errors in the sugges- tions. In this modification, CoAuthor was adapted to inject gram- matical and spelling errors into two of every five suggestions shown to participants. To do so, we introduced a secondAPIcall, prompt- ing the model to rewrite two suggestions with one grammatical or spelling error using a custom-designed prompt. The two sug- gestions were selected at random from the list of five suggestions. To ensure that the rewritten suggestions consistently contained a single deliberate error, we crafted prompts and piloted them to eval- uate the number and type of grammatical or spelling errors they produced. We iteratively refined the prompt wording to ensure the modified CoAuthor platform reliably introduced exactly one error in each of the two regenerated suggestions. Withgrammar_mistakes set to be 1, the final prompt we used was: Rewrite "\"" + suggestion + "\"" with exactly " + str( grammar_mistakes) + " grammatical or spelling mistakes, including a malapropism or a subject-verb agreement error within the first 8 words." 4.4.2Modification 2: The Instruction Box. For the Instruction Box, the user’s input was concatenated with the text editor’s content be- fore being sent to the model to guide suggestion generation. Specif- ically, we added logic such that if a user provided an instruction, CoAuthor constructed a new prompt by prepending the instruc- tion to the text in the editor (i.e., the prompt) using the following structure: Consider " + UserInput + " suggestions for the prompt: " + prompt The reformulated prompt was then passed to theAPIto generate completions as described earlier. We tested several variations and found that simple phrasing (e.g., “Consider”) was sufficient for the model to incorporate user instructions without altering their meaning. If no user instruction was provided, the platform defaulted to the editor text. With these modifications, participants could interact with CoAu- thor by asking for suggestions, accepting or ignoring suggestions, adding or deleting content directly in the text editor, and using the Instruction Box to adjust the suggestions. 4.5 Mental Model Manipulation and Validation To elicit the two mental models (functional and structural), we cre- ated a video for each condition, consisting of slides with a voiceover that provided descriptions and illustrated examples of the modified CoAuthor platform. The descriptions were developed iteratively, refined through pilot studies, and informed by key principles for supporting mental model formation inAIsystems, as described by Kulesza et al. [42] and Gero and Möller [27]. The functional description video provided an overview of the functionalities of the platform, such as how the user can ask for sug- gestions, how they can customize suggestions using the Instruction Box, and other basic usage information. The structural description video included all these slides but added more details about the inner workings of the platform. For example, every participant was told that CoAuthor generates suggestions when pressing Tab in the text editor. However, the structural description also covered that aLLMis used to generate a set of original suggestions, which are filtered before being shown to the user. The structural video also provided examples to illustrate how the platform works. For in- stance, it included examples of how different starter sentences in a From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain 1 2 Text editor Instruction box Figure 2: A screenshot of the AI-based writing assistant with the text editor where the main letter is written in the center, the Instruction Box is on the left side, and a basic summary of how to use the platform is on the right side. The interface is a modified version of the open-source CoAuthor platform created by Lee et al. [44]. Table 2: The content of the video description for the two experimental conditions ContentFunctionalStructural What is the CoAuthor platform?✓ What are the key features of the CoAuthor platform?✓ What are the main functionalities of the CoAuthor platform?✓ Usage of text editor (feature 1)✓ How does CoAuthor generate its suggestions?✗✓ Examples of how content in the text editor affects the suggestions generated✗✓ Usage of Instruction Box (feature 2)✓ How does the instruction from the Instruction Box customize the suggestions?✗✓ Examples of how content in the Instruction Box affects the suggestions generated✗✓ Table 3: The seven questions in the manipulation survey and their corresponding correct answers. Questions 5–7 assessed knowledge of internal system mechanisms described only in the structural video. #QuestionCorrect Answer 1CoAuthor is a writing assistant that does not rely on AI.False 2CoAuthor offers phrases and sentences when a user requests suggestions.True 3Where should the user write the main body of their cover letter on the CoAuthor platform?Text editor 4CoAuthor will automatically generate suggestions for the user when the user stops typing.False 5CoAuthor filters and trims the original list of suggestions generated by the large language model before showing it to the user. True 6CoAuthor gives suggestions based on what is a likely continuation of the last sentence in the text editor. True 7The content of the Instruction Box is sent to a different large language model than the text in the text editor. False cover letter would lead to varying suggestions based on the content of the starter sentences. To account for the fact that the participants might have a pre-existing mental model of commonly usedAI-based writing assistants, both functional and structural descriptions high- lighted how the modified CoAuthor platform used in the experiment differs from popular and commercially available tools such as Chat- GPT and Grammarly. Table 2 illustrates the content that was and was not covered in the functional and structural video descriptions. To validate the effect of our manipulation, participants were asked to fill out an understanding survey, which served as a manip- ulation check (see Figure 1), assessing participants’ understanding of the platform and the information presented in the instructional videos. As shown in Table 3, Questions 1–4 asked about general usage and functionalities of the platform. These aspects were cov- ered in both video descriptions, and therefore participants in both the functional and structural groups were expected to answer these items correctly. Questions 5–7 served as the manipulation check: they probed at how the system works as explained only in the structural video. As a result, participants in the structural condition were expected to answer these items correctly, whereas participants in the functional condition were expected to show lower accuracy, which could appear as either incorrect responses or selecting “I don’t know.” For all questions except Question 3, participants se- lected “True,” “False,” or “I don’t know.” For Question 3, they chose between “Text Editor,” “Instruction Box,” “Both,” or “I don’t know.” CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. 4.6 Data Analysis Guided by our research questions, we analyzed both quantitative and qualitative data collected from participants’ interactions with the writing assistant, their final cover letters, post-task surveys, and semi-structured interviews. Next, we describe the key measures, and the analysis approaches. 4.6.1 Quantitative Analysis. Our quantitative analysis was struc- tured around the three research questions, each paired with specific hypotheses and constructs. Table 4 provides an overview of how each hypothesis maps to key constructs, corresponding data sources, and specific measures. For RQ1, we expect participants with a structural mental model to exert more deliberate control by requesting, steering, accepting, and editing more suggestions, while also contributing original text. We study this research question through the following measures and hypotheses: • H1a (Requesting suggestions): Measured by the total num- ber of requested suggestions divided by the word count across the full document (suggestion requests); participants in the structural condition will request moreAI-generated suggestions than those in the functional condition. •H1b (Accepting suggestions): Measured by the total num- ber of accepted suggestions over the total number of re- quested suggestions across the full document (acceptance ratio); participants in the structural condition will accept a greater proportion of suggestions. •H1c (Editing accepted suggestions): Measured by sum- ming the number of deleted characters within the first 20 recorded edit actions after each suggestion is accepted (across the full document) and dividing by the total number of char- acters in all accepted suggestions (edit ratio); participants in the structural condition will make more direct edits to accepted suggestions. • H1d (Steering via instructions): Measured by the total number of instructions issued divided by the word count (instruction count) in the writing session; participants in the structural condition will issue more instructions to steer AI-generated suggestions. •H1e (Direct writing contribution): Measured by the total number of user-written characters divided by the total num- ber of characters in the final document (user character ratio); participants in the structural and functional condition will contribute an equal level of original text. For RQ2, we expect participants with a structural mental model to write higher-quality cover letters. We evaluated the letters using two complementary methods. •H2a (Grammar correctness): Measured by the Grammarly correctness score (normalized by total word count), which captures the number of suggested corrections related to grammar, spelling, and punctuation (Calibrated Corrections) [31]. Participants in the structural condition are expected to require fewer corrections in their final letter. As a comple- mentary exploratory analysis, we also examine the error ratio, measured as the total number of accepted erroneous sugges- tions divided by total number of accepted suggestions across the full document, and errors per word, measured as the to- tal number of accepted suggestions containing grammatical errors divided by the word count of the full document. •H2b (Overall quality): Measured by two researchers (both co-authors of this paper) independently assessing writing quality using a structured rubric with three dimensions: rel- evance of the skills and qualifications to the job posting, flow and clarity of ideas, and tone/style appropriateness. One re- searcher rated all of 48 letters, and a second rated 50% for reliability. Inter-rater agreement was high, with weighted Cohen’s Kappa scores of 0.79 for relevance, 0.70 for flow, and 0.85 for tone/style. The rubric for this qualitative quality assessment was grounded in participants’ own definitions of high-quality writing, including clearly aligning qualifi- cations with the job, using a persuasive and sincere tone, and maintaining a concise, well-structured format. Partic- ipants in the structural condition are expected to produce letters that score higher on these dimensions. The full rubric description is provided in Appendix D. For RQ3, we expect participants with a structural mental model to report more positive perceptions of the final outcome and the writing process. All measures for RQ3 were captured through 5- point Likert-scale items in the post-task survey. The survey was developed based on questions used in relevant prior work [14,18, 21, 37, 44]. •H3a (Perceived usefulness and ease of use): Measured by survey items on usefulness and ease of use; participants in the structural condition will report higher ratings on both. •H3b (Perceived quality of the final output): Measured by survey items on satisfaction with the final letter; participants in the structural condition will rate their letters more highly. •H3c (Perceived control): Measured by survey items on perceived influence over the writing process and system suggestions; participants in the structural condition will report a stronger sense of control. •H3d (Perceived ownership): Measured by survey items on authorship and attribution; participants in the structural condition will report higher ownership, indicating the letter reflects their own voice and intent. 4.6.2Qualitative Analysis. To enrich and contextualize our quanti- tative findings, we conducted two qualitative analyses: one focused on participants’ written instructions to customize suggestions on the platform, and the other on the responses from the post-task semi-structured interviews. First, to inform RQ1, we analyzed a total of 315 instructions written by participants during the cover letter revision task. To characterize how participants attempted to guide the AI-generated suggestions, we developed a coding rubric with three binary cat- egories: whether the instruction aimed to shift the tone or style of the letter, articulate relevant qualifications, or prompt a struc- tural change. Three researchers, who are all authors of this paper, collaboratively developed the rubric by jointly coding an initial set of 20 instructions. Two of the researchers then independently co-coded 40 instructions to establish consistency before the lead researcher completed coding the remainder. Inter-rater reliability From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain Table 4: Mapping of hypotheses to constructs and measures (quantitative analysis) HypothesisConstruct(s)Data SourceSpecific Measures (calculated per participant and letter) H1aRequesting suggestionsInteraction logsSuggestion requests: total number of requests for suggestions divided by word count. H1bAccepting suggestionsInteraction logsAcceptance ratio: accepted suggestions divided by total suggestion requests. H1cEditing accepted suggestionsInteraction logsEdit ratio: number of deleted characters within the first 20 recorded edit actions after each accepted suggestion, divided by the total number of characters in all accepted suggestions. H1dSteering via instructionsInteraction logsInstruction count: total number of instructions the user inputs divided by the word count. H1eDirect writing contributionInteraction logsUser character ratio: total number of user-written characters divided by total characters. H2aGrammatical correctness of the letter and accepted erroneous suggestions Grammarly scores of the final letter, Interaction logs Calibrated corrections: Grammarly correctness score, defined as the number of corrections suggested by Grammarly divided by word count; Error ratio: total number of accepted suggestions with errors divided by total accepted suggestions; Errors per word: total number of accepted suggestions containing grammatical errors divided by word count. H2bOverall cover letter qualityQualitative quality assessment Rubric-based quality ratings: a score between 1(poor) – 5(excellent) along three dimensions of relevance, flow, and tone/style. H3aPerceived usefulness and ease of usePost-task surveyUsefulness and ease-of-use: perceived helpfulness of the system and effort required to use it, assessed using multiple items rated on 5-point Likert scales. H3bPerceived quality of final outputPost-task survey Perceived quality: satisfaction with the quality of the final letter, assessed using multiple items rated on 5-point Likert scales. H3cPerceived controlPost-task surveyPerceived control: sense of influence over the writing process, assessed using multiple items rated on 5-point Likert scales. H3dPerceived ownershipPost-task survey Perceived ownership: extent to which the letter reflects the participant’s own voice and intent, assessed using multiple items rated on 5-point Likert scales. was high across all dimensions (휅=0.775 for tone/style shift, 0.773 for qualifications articulation, and 0.713 for structure prompting). This analysis provided insights into the type of instructions the participants used to steer theAI-generated suggestions. The full rubric description is provided in Appendix D. Second, to inform all three RQs, we conducted a reflexive the- matic analysis of the semi-structured interviews [11,12], which ex- plored participants’ perceptions of ownership, control, and writing quality. The coding process was explicitly oriented around themes relevant to our research questions. The lead researcher coded all interview transcripts, while a second researcher coded a quarter of them to support synthesis. Through iterative discussion, the two re- searchers compared interpretations, resolved discrepancies, and re- fined the coding structure until a shared understanding was reached. This process resulted in a set of themes describing how participants understood the role ofAI-based writing assistants, the degree of influence they felt over the evolving text, and the aspects of system interaction that shaped their sense of control and ownership. 5 Findings We begin by describing participants’ baseline writing practices and attitudes towardAI-based writing assistants, followed by the vali- dation results from our manipulation check to confirm the effective- ness of the mental model priming. We then present findings from both our quantitative and qualitative analyses, organized around the three research questions, and conclude with a reflection on study limitations and the researchers’ positionalities to contextualize our interpretations. Participants’ writing practices and perspectives onAIassistance. Participants in both mental-model conditions stated that they typi- cally begin by brainstorming and outlining, followed by iterative refinement, when completing writing tasks relevant to their pro- fessional and personal contexts. Many reported using comparable strategies for cover letters, though several emphasized relying on templates or examples, noting that they often “start with the same two sentences” (R2-S) 1 or “look for an example” (R2-S, R3-F). Most participants reported rarely usingAI-based writing assis- tants. Among participants who occasionally use these assistants, Grammarly and ChatGPT were most commonly mentioned. No- tably, some acknowledged thatAI-based assistants such as Gram- marly are “always on" (R1-F) and present “in the background" (R29-F), whether they actively sought them or not. Across both conditions, participants’ overall attitudes were cautious or negative towards the use ofAI-based writing assistants. Concerns included loss of authenticity (e.g., “when I read it back, it doesn’t feel authentic to me" R37-F), ethical implications (e.g.,“I feel like a fraud,” R3-F), pri- vacy risks (e.g.,“I have to be careful with where I store information," R27-F), and the reliability ofAI-generated content, especially in higher-stakes writing such as “academic [assignments]" (R1-F). Some participants mentionedAI’s usefulness for low-stakes or templated tasks. Among all participants, the most common reasons for using AI-based assistants were grammar and spelling checks, as well as phrasing improvements. A few participants also described usingAI for tasks like brainstorming ideas or adjusting tone, though these uses were less common. All participants valued maintaining control over their writing. Many did not seek external help and avoided usingAI-based assis- tants to preserve this control, while some welcomed support from peers, templates, orAI-based assistants for phrasing and high-level structure. Across both conditions, participants emphasized ensur- ing that their own voice remained clear and that they guided the ideas and overall direction. This preference extended to cover letter writing, though a few noted feeling less control “because they based [the letter] off of their job description” (R10-S). We observed a significant manipulation effect based on the manip- ulation check survey results. A Mann–Whitney U test indicates that participants in the structural condition (Mdn=6.0,Mean Rank= 30.56) outperformed those in the functional condition (Mdn=5.0, Mean Rank=18.44) on the three manipulation check questions 1 Participant IDs ending in –F correspond to the functional condition, while those ending in –S correspond to the structural condition. CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. (푈=142.50,푍= −3.15,푝< .01). This helped validate that our video-based manipulation influenced our participants’ mental mod- els. At the same time, the structural group performed more poorly than the functional group on Q1 (i.e., “CoAuthor is a writing assis- tant that does not rely on AI,” see Table 3), which assessed a basic factual claim about CoAuthor (푈=228.00,푍= −2.34,푝< .05). Although unexpected, this single-item difference is unlikely to indi- cate a true misunderstanding of the platform. Post-task interviews confirmed that all participants recognized CoAuthor as an AI-based system. We conjecture that Q1 was likely misread or overlooked (e.g., participants missed the negation in the question), given its early placement, the perceived complexity of the question, and the longer instructional video in the structural condition (8.54 minutes vs. 3.58 minutes). Looking across all manipulation-check items, the overall validation results still indicate that the structural framing shaped participants’ conceptual understanding as intended. Com- plete item-level descriptive statistics and test results are provided in Appendix E. 5.1 RQ1: How Do Different Mental Models Affect the Way Users Exert Control in Their Writing Process? We first investigated how participants with functional or structural mental model priming exercised control when interacting with the system. Below, we report the main findings; where needed, additional details on the statistical analyses are provided in the Appendices. We observed no significant differences in control behavior-related measures across conditions. To evaluate control behavior, we ana- lyzed interaction data that captured how participants engaged with the AI-based suggestions (described in Section 4.6 and outlined in Table 4), including the suggestion requests, the acceptance ratio, the edit ratio, the instruction count, and the user character ratio. Contrary to our hypotheses (H1a–H1d), structural participants did not request, accept, or edit suggestions more frequently, nor did they issue more instructions, and consistent with H1e, both groups contributed a similar amount of original text. As shown in Figure 3, there were no statistically significant differences between the functional and structural conditions when looking at suggestion requests (푝=0.409), acceptance ratio (푝=0.253), edit ratio (푝= 0.869), instruction count (푝=0.726), and user character ratio (푝= 0.305). User character ratio and acceptance ratio met normality assumptions and were analyzed using independent samples푡-tests; all other measures were evaluated using Mann–Whitney푈tests. Additional details about the descriptive and statistical analysis are in Appendix F. Qualitative interview responses provide cues to why we did not observe differences across conditions for different behavioral measures. Across both conditions, participants described usingAI- generated suggestions opportunistically—most commonly when they “did not know what [they] were going to say” (R5-F) or when they needed to “link different” ideas together (R26-S). Some used suggestions primarily for inspiration and “guidance” (R4-S) with- out selecting them, while others selectively adopted only specific elements such as a single “linking word” (R17-F); in some cases, suggestions were accepted as-is. Thus, suggestion acceptance de- cisions seem guided by stylistic fit (e.g., “sounded like something I would write,” R1-F) and local content relevance (e.g., “how relevant it was and how well it connected to the part of the sentence I had typed already,” R32-S), rather than by participants’ understanding of how the system generated the suggestions. Participants’ views on the quality of the suggestions were simi- larly mixed across the two conditions. Some found the suggestions helpful and appropriately toned (e.g., “helped me to form phrases that were appropriate to the context easily,” R41-F), while others described them as “generic” (R16-S), repetitive (e.g., “the suggestions were the same, with words placed differently,” R3-F), or “basic and as expected” (R10-S). All participants noticed that the suggestions contained grammatical and spelling errors. Experiences with the Instruction Box were similarly mixed and did not differ by condi- tion. Some participants appreciated that they “could write anything” in the Instruction Box (R11-F), while others preferred having “a little more structure” (R21-F). While some found the Instruction Box “useful” in shaping suggestions (R32-S), others reported lim- ited or inconsistent effects, sometimes needing to provide “a lot of instructions” (R23-F) before seeing meaningful changes. This suggests that perceptions of suggestion quality and the Instruction Box’s effectiveness were primarily driven by participants’ personal writing-style and meta-prompting preferences, rather than by dif- ferences in mental models. The types of instructions participants provided were similar across conditions, though participants in the structural condition wrote more descriptive instructions. To examine how participants used writ- ten instructions, we coded all instructions along three dimensions: tone/style shift, qualifications articulation, and structure prompting (see Section 4.6.1 for a description of our approach). Quantitative comparisons then showed no statistically significant differences in the types of instructions participants wrote across conditions. Specifically, chi-squared tests revealed no reliable association be- tween condition and tone/style shift (휒 2 (1,푁= 315) = 0.54,푝= 0.464), qualifications articulation (휒 2 = 1.30,푝= 0.254), or structure prompting (휒 2 = 0.06,푝= 0.802). Similarly, binomial tests, which we conducted to examine whether participants included some specific types of instructions more or less often than chance across the con- ditions, showed no deviation from a 50/50 distribution for tone/style shift (푝= 0.851), qualifications articulation (푝= 0.840), or structure prompting (푝= 0.590). Full instruction-level coding distributions and statistical test results are reported in Appendix G. While the coding of the instruction categories did not reveal significant differences between conditions, closer qualitative ex- amination of the instruction content shows some variation in how participants articulated tone and relevancy-related guidance. Over- all, participants in the structural condition tended to provide more descriptive and context-rich instructions such as “Emphasize and weave my roles in higher education, mental health, and clinician training” or “highlight using positive language that I am motivated and detail-oriented.” In contrast, participants in the functional con- dition used brief keyword-style cues such as “use positive language” or “highlight digital skills.” The difference between the conditions was modest and on the the order of several cases, but reflected more elaborate instruction phrasing in the structural condition. From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain Figure 3: Bar graphs of mean values (with error bars) for five control behavior measures—suggestion requests, acceptance ratio, edit ratio, instruction count, and user character ratio—across the functional and structural mental model conditions. Given that interview responses, across both conditions, revealed substantial variation in how participants experienced and used the Instruction Box (as explained earlier), it is not clear that these differences in descriptiveness of instructions are related to partici- pants’ understanding of the system. Nonetheless, the qualitative differences observed indicate the the relationship between mental models and instruction use needs to be further studied. Overall, both our quantitative and qualitative data suggest that participants’ moment-by-moment writing needs, perceived sugges- tion quality, and stylistic preferences primarily shaped how they exerted control during the writing process, with mental models playing a more limited role. 5.2RQ2: How Do Different Mental Models Held by Users Influence the Quality of the Final Written Output? To investigate whether different mental model framings influenced the quality of participants’ final cover letter, we assessed writing quality via two methods (i.e., grammatical correctness and qual- itative quality assessments, see Table 4 for an overview). Based on our hypotheses, we expected that participants with a struc- tural mental model would produce higher-quality cover letters, as judged via both automated grammar checks and qualitative quality assessments. Full results for grammatical correctness, erroneous suggestion acceptance and cover letter quality analyses are reported in Appendices H and I. Participants in the structural condition wrote letters with more grammatical, punctuation, and spelling errors. To evaluate grammat- ical correctness and surface-level writing quality (H2a), we used Grammarly’s count of correctness suggestions, which captures the number of grammatical, punctuation, and spelling errors in each letter. 2 Contrary to our expectations, a one-sided independent samples t-test revealed a statistically significant difference in the calibrated corrections identified using the Grammarly platform—i.e., the number of corrections suggested by Grammarly divided by word count—푡(46)=−2.316,푝<0.05, with a moderate effect size (푑=0.67). This indicates that participants in the structural condi- tion produced final letters with more grammatical errors compared to those in the functional condition. To better understand why the number of corrections required was higher for the structural con- dition, we conducted an exploratory analysis of how participants treated erroneous suggestions. As shown in Figure 4 and described next, participants in the structural condition accepted more sugges- tions containing grammatical errors and had a higher error ratio in accepted suggestions compared to those in the functional condition. Participants in the structural condition accepted more erroneous AI suggestions. To better understand the discrepancy in cover letter quality, we conducted an exploratory analysis of the interaction data. This revealed a significant difference: participants in the struc- tural condition accepted more suggestions containing grammatical errors (Figure 3). A Mann–Whitney U test showed that the total number of accepted succestions errors per word was significantly higher in the structural condition (푈=170.0,푝<0.05,푟= .352). The error ratio, defined as the proportion of accepted suggestions containing grammatical errors, was also higher in the structural condition, though this difference was only marginally significant 2 For brevity in this section, we use “grammatical errors” to refer to the combined count of grammatical, punctuation, and spelling errors captured in Grammarly’s correctness score. CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. Figure 4: Bar graphs of mean values, with error bars, for calibrated corrections identified by Grammarly in final letters, error ratio, and accepted erroneous suggestions, comparing the functional and structural mental model conditions. An asterisk (*) denotes statistically significant differences (푝< 0.05) for calibrated corrections and errors per word. Figure 5: Bar graph of mean values (and error bars) for rubric-based cover letter quality ratings across dimensions of relevance, flow, and tone/style for the functional and structural conditions. (푈=204.5,푝= .084,푟= .250). Because this analysis was ex- ploratory we did not apply corrections for multiple comparisons [7]. Given the marginally imbalanced distribution of participants whose self-reported native language was English across mental model conditions, we explored whether being a native language speaker moderated the effect of mental model prompting. We con- ducted two-way Analysis of Variance (ANOVA)s for both the error ratio and accepted suggestions with errors (calibrated). Although these variables did not meet the Shapiro–Wilk normality assump- tion, Levene’s tests indicated homogeneity of variance, justifying ANOVAuse. No significant interaction effects were observed be- tween mental model condition and native language for either the error ratio (퐹(1,44)=0.408,푝= .526) or accepted suggestions with errors per word count (퐹(1,44)=0.355,푝=0.554). No sig- nificant main effects of native language were observed for either outcome measures (푝=0.216 (error ratio) and푝=0.258 (accepted suggestions with errors)). From the semi-structured interviews and post-task surveys, we found that most participants across both conditions noticed gram- matical errors in theAI-generated suggestions, but adopted differ- ent strategies for handling them. Some participants reported editing errors after accepting suggestions (e.g., “I think once or twice, I did take one and then just remove the part that was wrong,” R36-S), others avoided selecting erroneous suggestions entirely (e.g., “I felt less inclined to choose the ones with the errors,” R19-F). Based on the observed error ratio and errors per word values, it appears that some participants may have accepted erroneous suggestions without subsequent changes, although no participant explicitly described this behavior. These strategies were observed in both mental model conditions, suggesting that participants relied on individual techniques to manage suggestion quality rather than following a condition-specific approach. As such, the qualitative data on handling grammatical errors does not conclusively explain why participants in the structural conditions have cover letters with higher-level grammatical errors. One possible interpretation is that the decision to accept a suggestion was often guided by writing From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain style and goal, and that participants in the structural condition paid more attention to the suggestion content than to surface-level errors. Qualitative quality assessment of letter quality indicates no signif- icant differences across groups. Contrary to H2b, qualitative qual- ity evaluation of the cover letters (described in Section 4.6) re- vealed no statistically significant differences in the overall writ- ing quality between the structural and functional conditions (Fig- ure 5). Mann–Whitney U tests yielded non-significant results for all three evaluation criteria: relevance (푝=0.398), flow and clarity (푝=0.869), and tone/style appropriateness (푝=0.954). These find- ings indicate that, aside from grammatical errors, participants’ final letters were judged to be of comparable quality across the mental model conditions. Reflecting on RQ2, we observed that participants in the struc- tural condition accepted more erroneous suggestions and produced letters with more grammatical errors, yet these differences were not reflected in the qualitative quality assessment of cover letter writing for factors such as relevance, fluency, and clarity, or tone/style ap- propriateness. This pattern indicates that a structural mental model may shape users’ attention or inattention to AI output—particularly in the presence of deliberate errors—without necessarily improving or degrading broader aspects of writing quality. 5.3 RQ3: How Do These Mental Models Shape Users’ Overall Experience of Writing With the AI-Based Writing Assistant? For H3a–H3d, we hypothesized that participants in the structural condition would report higher ratings of perceived quality, control, ownership, and usefulness of the system. With the exception of ease of use, where participants with a structural mental model gave significantly higher ratings, we found no statistically significant differences between conditions. Across both groups, participants reported a strong sense of ownership and control over the final letter, along with generally high ratings of quality and usefulness. Because we tested five related subjective experience measures (per- ceived quality, control, ownership, usefulness, and ease of use), we applied a Bonferroni correction, yielding a corrected significance threshold of훼= .01. Figure 6 summarizes selected quantitative results, and detailed analyses are provided in Appendix J. Structural understanding led to more positive perceptions of CoAu- thor’s usefulness and ease of use. We hypothesized that participants with a structural mental model would report a better experience (H3a), based on prior work suggesting that deeper system under- standing supports user satisfaction and usability [e.g.,42]. Our findings provide only tentative support for this hypothesis. Participants in the structural condition rated CoAuthor as easier to use than those in the functional condition. A one-sided indepen- dent samples푡-test revealed a difference in ease of use,푡(46)=−2.07, 푝< .05, with a moderate effect size (푑=0.60). This effect was statistically significant at the uncorrected훼= .05 level, but did not meet the Bonferroni-corrected threshold of훼= .01. Perceived usefulness was assessed using a four-item composite scale with acceptable internal consistency (Cronbach’s훼=0.76). Although the perceived usefulness score was on average slightly higher for the structural condition, the difference between conditions on this composite score was not statistically significant (푡(46)= −0.78, 푝= .220. Additionally, none of the individual item comparisons reached statistical significance. These comparisons are visualized in Figure 6. Qualitative interview data provided additional context for when participants’ perspectives appear to overlap and when they diverge. Across both conditions, participants described CoAuthor as use- ful for improving phrasing, overcoming writer’s block, generating ideas, and accelerating the writing process, with many completing a full draft within 20 minutes. Several also found the tool helpful for structuring their letters. At the same time, participants in both con- ditions raised similar concerns about generic or robotic-sounding suggestions and occasional grammatical errors. These shared per- ceptions of their interactions with the platform might explain why perceived usefulness ratings remained similarly high across both conditions. Where the conditions diverged was in how participants’ self- reports linked their understanding of the system with their confi- dence and perceived sense of control when using the platform. A few participants in the functional condition expressed that limited insight into how the system worked, constrained their trust and sense of control (e.g., “I need to know where it is getting the informa- tion from,” R21-F). In contrast, some participants in the structural condition explicitly linked their confidence to the explanatory on- boarding (e.g., “[The platform] works well and it’s very functional ...I am able to either use the suggestions effectively or dismiss them if they are bad,” R2-S). Together, these participants’ comments pro- vide further insights into why the structural understanding might have only selectively improved perceptions of ease of use, while the more quantitative evaluations of usefulness were minimally higher for the structural condition. We found no significant differences in perceived letter quality, con- trol, or ownership across conditions. Contrary to H3b–H3d, quan- titative analyses revealed no statistically significant differences between the functional and structural conditions in participants’ perceived letter quality, control, or ownership. Participants in the structural condition rated CoAuthor’s suggestions as slightly more error-free than those in the functional condition, although this dif- ference was not statistically significant (푡(46)= −1.48,푝= .072, 푑=0.43). They also reported marginally higher ratings for the final letter’s grammatical correctness, tone, and content, but these differences were likewise non-significant (푝= .801 and푝= .134, respectively). Perceived ownership was analyzed using a composite score across items (Cronbach’s훼= .714), whereas perceived con- trol was assessed using multiple items but is reported at the item level due to lower internal consistency (훼= .658). Both conditions reported similarly high levels of perceived control and ownership. Overall, we found no statistically significant differences across quality-, control-, or ownership-related measures. Mean ratings for all self-reported items are shown in Figure 6. Additional details are included in Appendix J. Interaction features, not mental models, appear to anchor partic- ipants’ perceptions of quality, control, and ownership. Qualitative findings help explain why perceived quality, control, and ownership remained consistently high across both mental model conditions. CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. Figure 6: Bar graphs of mean values (and error bars) for self-reported ratings across conditions for usefulness, ease of use, perceived quality, perceived ownership, and perceived control. A one-sided independent samples푡-test revealed a difference in ease of use, with a moderate effect size (푑=0.60). As marked on the left-most plot, this effect was statistically significant at the uncorrected 훼= .05 level, but did not meet the Bonferroni-corrected threshold of 훼= .01. Overall, participants noted that the modified CoAuthor platform felt different from other forms ofAI-based writing assistance. Many participants remarked that they felt a high level of control when using CoAuthor, describing it more as a supportive “friend” offering help (e.g., “I was basically just writing something and asking a friend how would you help me describe this,” R3-F) rather than a tool that did all the work for them (e.g., “I like [that] you still have to do the work [...]. For [tools like] ChatGPT [...], it just writes the whole thing and you don’t really do anything. This [...] you still having to do it yourself,” R8-S). Participants identified several concrete reasons for experiencing a high sense of control. The most common factors were that they could request suggestions whenever they wanted (e.g., “I could choose when to ask for suggestions really easily,” R24-S), directly edit the text in both the editor and the Instruction Box (e.g., “I always had the ability to go back and erase what I didn’t like,” R17-F), and choose whether or not to accept the suggestions (e.g., “I could reject all of them or I could choose an option,” R27-F). Moreover, participants noted that they felt in control because CoAuthor did not nudge or pressure them to use its features (e.g., “I like that I have to ask it for the suggestion,” R5-F)—this observation was coded four times more often in the functional condition compared to the structural condition. Participants further emphasized that they felt a sense of control because the suggestions were relevant to what they had written in the Instruction Box and text editor (e.g., “it gave relevant suggestions that helped me express my inner words and I just put my feelings in the instruction box,” R34-S). Additionally, the short length of the suggestions—typically less than a full paragraph—contributed to the sense of perceived control, as it required them to do more of the writing themselves. In the same vein, the fact that users had to enter one instruction at a time encouraged a more active and hands-on writing process. Notably, several participants across both conditions highlighted points where they felt some loss in control. Few participants in functional condition expressed that they lost some control when they “just kept asking for suggestions” (R1-F), often because they ran out of ideas or were under “time constraint” (R33-F). On the other hand, a participant in the structural condition expressed that having the ability to put in “more instructions [at single use of the Instruction Box]” (R18-S) would increase their sense of control. Participants primarily defined ownership in three ways. First, they described ownership as having their own ideas reflected in the writing (e.g., “writing reflects my thoughts, my arguments” R9-F). Second, some associated ownership with doing the writing them- selves (e.g., “I have really done the job of writing it myself,” R20-S). Third, others emphasized authenticity and personal voice (e.g., “it would have to be genuine and honest,” R14-S). Along these lines, participants suggested that a sense of ownership often came from feeling proud of what they had written. These conceptualizations of ownership were present across both mental model conditions. Several participants also noted that one could still maintain a sense of ownership when using help from others orAI-based writing assistance, depending on how much help was used and whether it aligned with these core criteria. Most participants expressed a strong sense of ownership because they felt in control—they could write their own text, were not nudged, could choose whether or not From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain to accept suggestions, and could request suggestions whenever they wanted. Many also felt ownership because they were able to bring in their own ideas and personal experiences. Several participants reported a strong but incomplete sense of ownership over the writ- ing, with seven characterizing their contribution at approximately 75%–85% (3 in the structural condition, and 4 in the functional one; e.g., “If I were to give it a percentage, I’d probably say around 80–85%, because it’s my experiences, but in terms of how the words are for- mulated, CoAuthor helped me formulate most of the words,” R35-F). However, some participants appeared conflicted. They reported a weaker sense of ownership when the writing did not feel authen- tic, did not reflect their personal style, or when they did not feel proud of the final output. When participants accepted too many suggestions, they found the draft letter felt less personalized. These observations were consistent across the conditions. Taken together, these qualitative accounts help contextualize and explain the quantitative patterns observed in participants’ overall experience of writing with the system. Regardless of participants’ mental model of how the system worked, perceptions of quality, usefulness, control, and ownership remained largely comparable across conditions, reflecting the central role of interaction features and affordances that supported user choice, idea contribution, and iterative editing. The qualitative data suggest that these affordances enabled participants to feel similarly empowered and satisfied with the system’s outputs, even when their conceptual understanding of the system differed. In contrast, perceived ease of use varied more directly with participants’ understanding of how the system oper- ated, as clearer mental models reduced uncertainty about available actions and system behavior, making interaction feel more intuitive and less effortful. 6 Limitations and Positionality This study has several limitations that shape how its findings should be interpreted. First, while we aimed to elicit distinct mental models of theAI-based writing assistant through carefully crafted videos providing functional and structural descriptions of the system, these framings represent static depictions of a user’s mental model at a single point in time. In practice, mental models are dynamic and may evolve with each interaction. Despite strict inclusion criteria targeting participants with limited prior experience, many partic- ipants likely entered the study with pre-existing beliefs aboutAI systems that could have influenced their perceptions and behav- iors. Furthermore, creating a truly structural mental model of a system as complex and opaque as anAI-based writing assistant is inherently challenging. We took deliberate steps to scaffold this understanding, but the extent to which participants internalized the structural framing likely varied. Another limitation concerns participant variability, particularly in writing proficiency. Although we attempted to recruit participants with comparable experience levels and provided a consistent task, underlying differences in writing skill may have shaped engagement with the system and perceptions of suggestion quality. In addition, while the prevalence of participants whose native language was not English did not dif- fer significantly across conditions, it was marginally imbalanced and may have introduced subtle variability in how participants interpreted and evaluated AI-generated writing. Our research team’s positionality also influenced the framing and interpretation of this work. The five researchers involved span both academic and industry settings in the Global North (Canada and the U.S.), with expertise in natural language processing, social computing, human–computer interaction, human–robot interac- tion, responsible AI, and system safety. These disciplinary perspec- tives shaped the study’s focus on user control, ownership, and safe system interaction. We acknowledge that our interpretations are in- formed by this interdisciplinary but Western-centric vantage point, and that broader perspectives—particularly those situated outside of North America or grounded in alternative epistemologies—are critical for expanding this line of inquiry. 7 Discussion In this study, we explored how users’ mental models of anAI-based writing assistant shaped interaction behavior, final writing qual- ity, and overall experience. Our manipulation was effective in that participants in the structural condition demonstrated a deeper con- ceptual understanding of how the system operated and performed better on factual knowledge questions. At the same time, many downstream outcome measures showed few or no statistically sig- nificant differences between conditions. Participants across both conditions displayed similar control behavior and produced let- ters that were similar in tone, clarity, and relevance. In addition, both groups reported a comparably strong sense of quality, owner- ship and control. These findings suggest that differences in users’ conceptual understanding of the system do not straightforwardly translate into broad differences in writing quality or perceived experience. Where differences did emerge, they should be inter- preted cautiously. Participants in the structural condition showed a consistent trend toward higher perceived ease of use and greater acceptance of erroneous suggestions, while also producing letters with a higher number of grammatical errors. Rather than indicating strong causal effects of structural understanding, these patterns suggest a subtle shift in how participants related to and used the sys- tem. Drawing on prior work on human–machine trust and reliance, one plausible interpretation is that structural framing functioned as a cue of system competence, fostering trust and shaping reliance tendencies, including greater acceptance of deliberately inserted erroneous suggestions. [54,65,73,79]. In this sense, overtrust is not presented here as a definitive explanation, but as a useful lens for interpreting why modest increases in ease of use may co-occur with a greater tendency to accept flawed suggestions and letters with more grammatical errors. Re-assessing the link between mental models and safe control in some human–AI interaction contexts. In traditional safety-critical systems, “hazardous situations” are often well-defined events with known consequences—such as a plane deviating from its flight path or a robot misinterpreting a command [46]. InAI-based writing as- sistants, hazards are subtler and more subjective. They may involve insertion of specific ideas via AI-generated text, misrepresentation of cultural content, or uniformization of writing styles [35,49,60]. Additionally, these risks are best understood as emerging from patterns of interaction over time rather than as isolated, single- point failures [40,63,75]. This implies that monitoring for such CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. hazards, requires oversight mechanisms over the course of these interactions. One common assumption in system safety is that a more ac- curate mental model should support more effective oversight and intervention [23,47]. Our findings complicate this narrative within the context of AI-assisted writing systems with preconfigured erro- neous AI outputs. Participants with a structural mental model did not consistently intervene to correct erroneous AI-generated sug- gestions and, if anything, tended to accept such suggestions more readily, resulting in letters with a higher number of grammatical errors. At the same time, they also reported the system as easier to use. This points to a key tradeoff that might be especially salient in lower-stakes, productivity-oriented domains like everyday writing where users may prioritize efficiency, fluency, and reduced cogni- tive effort over careful verification of each system output. From this perspective, what appears as reduced oversight may in fact just reflect a reasonable adaptation to task demands—e.g., optimizing efficiency over producing a polished writing outcome—rather than a failure of understanding. In settings where users might hold distinct conceptualizations of what constitutes a system error, failure, or harm, system designers may need to shift focus. Rather than providing users with a more in-depth description of how a system works, emphasis should be placed on scaffolding good interaction affordances: mechanisms that surface suggestion quality, prompt reflection, or highlight al- ternative options [14,37]. This resonates with system safety’s em- phasis on designing feedback-rich environments that allow users to recognize system states, monitor their own actions, and recover from errors [20]. Such design strategies may better support writers in exercising meaningful and context-sensitive control, particularly in situations where what counts as a “safe” action is not obvious. Rethinking mental model attunement in designingAI-based writ- ing systems. While our findings suggest that a more in-depth struc- tural mental model does not necessarily lead to better control or writing quality, this does not diminish the value of understand- ing how mental models shape AI use. In fact, our data show that mental models still meaningfully influence how users approach interaction, perceive system usability, and articulate their strategies for directing AI behavior. For example, participants in the struc- tural condition rated the system as significantly easier to use and were more descriptive in customizing tone or intent through the Instruction Box. These behaviors suggest that structural mental models can scaffold more purposeful interaction, even if they do not guarantee improved outcome quality. Importantly,AI-based writing systems are not static and non- deterministic. They evolve rapidly through updates, retraining, and interaction data, making it difficult to determine what a “cor- rect” or “complete” mental model even looks like. From a systems perspective, this introduces new challenges for maintaining safe user-system interaction: mental models may quickly become out- dated, incomplete, or misaligned with current system behavior [51,76]. From this perspective, mental models should not be treated as fixed educational outcomes, but as evolving resources that are continuously renegotiated through use. This has important design implications. Rather than relying on one-time instructional framing, systems should support ongoing recalibration through adaptive scaffolding [59]. This may include progressive disclosure of system behavior [16], just-in-time explanations tied to specific outputs [71], lightweight uncertainty indicators [67], and interaction techniques that encourage users to periodically reflect on whether AI sugges- tions align with their goals [50]. Supporting better mental models, in this sense, is less about transmitting backend technical detail and more about enabling sustained, critically informed engagement over time. Importance of affordances versus user understanding in mitigating potential sources of sociotechnical harms. In safety-critical systems, ensuring human control is paramount; users must be able to moni- tor, override, and intervene in ways that prevent harm. However, in creative or productivity-oriented domains like writing, the need for user control is more context-dependent. It depends on the nature of the writing task, the goals of the writer, and broader contex- tual factors such as whetherAIassistance is permitted or desirable [21,29]. Our findings suggest that participants experienced a strong sense of control and ownership when using CoAuthor, regardless of mental model framing. Notably, CoAuthor offered more flexible interaction affordances than many typical writing assistants: users could freely request suggestions, revise them, or reject them en- tirely. This flexibility appeared to support a high baseline sense of control and ownership [37]. Compared to the manyAItools that generate long outputs or push users toward passive acceptance, the modified CoAuthor’s interaction design may be a key reason participants reported feeling like authors, not just editors. This raises an important design implication: if the goal is to support user control in writing assistants, the priority may not necessarily be better system understanding, but better affordances. Scope, generalizability, and baseline framing. Our findings are grounded in scenarios where the system deliberately produced ob- servable errors in its suggestions. This design choice allowed us to directly examine whether users could detect and intervene in flawed outputs, but it also limits the generalizability of our conclu- sions. It remains an open question whether similar trust, reliance, and control dynamics would emerge under conditions in which no errors were deliberately introduced or in other forms of writ- ing support such as ideation, summarization, or stylistic rewriting. Structural mental models may play different roles in such settings, for example by shaping expectation management rather than error detection. At the same time, our investigation of user control was not lim- ited to error correction alone. Through analysis of how participants used the Instruction Box to adjust tone, structure, and intent, we also observed broader forms of intervention through which users attempted to steer and personalize the writing process. These be- haviors suggest that even when users show tendencies toward a lack of effective oversight at the level of individual suggestions, they may still engage in higher-level, directional forms of control over the overall output. This distinction is important for understand- ing ownership and control in AI-assisted writing beyond narrow notions of error detection. From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain 8 Conclusion In this study, we investigated how different mental model framings of anAI-based writing assistant influence users’ behaviors, writ- ing outcomes, and experiences. Participants exposed to structural explanations developed a deeper understanding of the system and reported higher perceived ease of use. However, they also showed a greater tendency to accept erroneous suggestions and produced cover letters with more grammatical errors. We posit that this pat- tern may reflect overtrust, whereby structural framing functioned as a cue of system competence, increasing reliance even when active oversight was warranted due to staged errors. Notably, participants across both conditions expressed a strong sense of ownership and control, largely fostered by the AI-based writing assistant’s inter- action features, which allowed users to actively write and edit. These results suggest that providing users with accessible means of control may be more impactful than deeper comprehension of the system alone when it comes to fostering a sense of ownership and control. Looking ahead, research should further examine how to cultivate appropriate mental models and interaction paradigms that empower users to exercise oversight when necessary. Acknowledgments We gratefully acknowledge all participants for their time and con- tributions to this research. We also thank Gauri Sharma, Kiara Wimbush, and Timothy Ko Lee for their invaluable assistance in running the pilot studies and setting up the research platform. References [1] Tazin Afrin, Omid Kashefi, Christopher Olshefski, Diane Litman, Rebecca Hwa, and Amanda Godley. 2021. Effective Interfaces for Student-Driven Revision Sessions for Argumentative Writing. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21, Article 58). Association for Computing Machinery, New York, NY, USA, 1–13. doi:10.1145/ 3411764.3445683 [2]Dhruv Agarwal, Mor Naaman, and Aditya Vashistha. 2025. AI suggestions homogenize writing toward western styles and diminish cultural nuances. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–21. doi:10.1145/3706598.3713564 [3] T. A. Bach, A. Khan, H. Hallock, G. Beltrão, and S. Sousa. 2024. A Systematic Literature Review of User Trust in AI-Enabled Systems: An HCI Perspective. International Journal of Human–Computer Interaction 40, 5 (2024), 1251–1266. doi:10.1080/10447318.2022.2138826 [4]Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019. Beyond accuracy: The role of mental models in human-AI team performance. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 7. Association for the Advancement of Artificial Intelligence (AAAI), Honolulu, Hawaii, USA, 2–11. doi:10.1609/hcomp.v7i1.5285 [5]Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21, Article 81). Association for Computing Machinery, New York, NY, USA, 1–16. doi:10.1145/3411764.3445717 [6]Kevin Bauer, Moritz von Zahn, and Oliver Hinz. 2023. Expl(AI)ned: The Impact of Explainable Artificial Intelligence on Users’ Information Processing. Information Systems Research 34, 4 (2023), 1582–1602. doi:10.1287/isre.2023.1199 [7] Ralf Bender and Stefan Lange. 2001. Adjusting for multiple testing–when and how? J. Clin. Epidemiol. 54, 4 (April 2001), 343–349. doi:10.1016/s0895- 4356(00) 00314-0 [8]Karim Benharrak, Tim Zindulka, and Daniel Buschek. 2024. Deceptive patterns of intelligent and interactive writing assistants. In Proceedings of the Third Workshop on Intelligent and Interactive Writing Assistants. ACM, New York, NY, USA, 62–64. doi:10.1145/3690712.3690728 [9] Peter Blokland and Genserik Reniers. 2020. Safety Science, a Systems Thinking Perspective: From Events to Mental Models and Sustainable Safety. Sustain. Sci. Pract. Policy 12, 12 (June 2020), 5164. doi:10.3390/su12125164 [10]Jessica Y Bo, Sophia Wan, and Ashton Anderson. 2025. To rely or not to rely? Evaluating interventions for appropriate reliance on large language models. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–23. doi:10.1145/3706598.3714097 [11] Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101. [12] Virginia Braun and Victoria Clarke. 2023. Toward good practice in thematic analysis: Avoiding common problems and be(com)ing a knowing researcher. Int. J. Transgend. Health 24, 1 (2023), 1–6. doi:10.1080/26895269.2022.2129597 [13]Michael J Burtscher and Tanja Manser. 2012. Team mental models and their potential to improve teamwork and safety: A review and implications for fu- ture research in healthcare. Saf. Sci. 50, 5 (June 2012), 1344–1354. doi:10.1016/ j.ssci.2011.12.033 [14]Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English Writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21, Article 732). Association for Computing Machinery, New York, NY, USA, 1–13. doi:10.1145/ 3411764.3445372 [15]John M Carroll and Judith Reitman Olson. 1988. Mental Models in Human- Computer Interaction. In Handbook of Human-Computer Interaction, Martin Helander (Ed.). North-Holland, Amsterdam, 45–65. doi:10.1016/B978-0-444- 70536-5.50007-5 [16]Karina Cortiñas-Lorenzo, Wanling Cai, and Gavin Doherty. 2025. Designing, implementing, and evaluating AI explanations: A scoping review of Explainable AI frameworks. ACM Trans. Comput. Hum. Interact. 32, 6 (Dec. 2025), 1–79. doi:10.1145/3769678 [17] Patrick Cox, Jörg Niewöhmer, Nick Pidgeon, Simon Gerrard, Baruch Fischhoff, and Donna Riley. 2003. The use of mental models in chemical risk protection: developing a generic workplace methodology. Risk Anal. 23, 2 (April 2003), 311–324. doi:10.1111/1539-6924.00311 [18] Hai Dang, Sven Goller, Florian Lehmann, and Daniel Buschek. 2023. Choice Over Control: How Users Write with Large Language Models using Diegetic and Non- Diegetic Prompting. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23, Article 408). Association for Computing Machinery, New York, NY, USA, 1–17. doi:10.1145/3544548.3580969 [19]Sidney Dekker. 2017. The Field Guide to Understanding “Human Error” (3rd ed.). CRC Press, Boca Raton, FL. doi:10.1201/9781317031833 [20] Roel Dobbe. 2022. System Safety and Artificial Intelligence. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22). Association for Computing Machinery, New York, NY, USA, 1584. doi:10.1145/3531146.3533215 [21]Fiona Draxler, Anna Werner, Florian Lehmann, Matthias Hoppe, Albrecht Schmidt, Daniel Buschek, and Robin Welsch. 2024. The AI Ghostwriter Ef- fect: When users do not perceive ownership of AI-generated text but self- declare as authors. ACM Trans. Comput. Hum. Interact. 31, 2 (April 2024), 1–40. doi:10.1145/3637875 [22]European Union. 2024. Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (AI Act). Official Journal of the European Union. https://eur-lex.europa.eu/eli/reg/ 2024/1689/oj Article 14 (Human Oversight). [23]John Gaspar, Cher Carney, Emily Shull, and William Horrey. 2020. The impact of driver’s mental models of advanced vehicle technologies on safety and performance. Technical Report. A Foundation for Traffic Safety and SAFER-SIM. [24]John Maurice Gayed, May Kristine Jonson Carlon, Angelu Mari Oriola, and Jeffrey S Cross. 2022. Exploring an AI-based writing Assistant’s impact on English language learners. Computers and Education: Artificial Intelligence 3 (Jan. 2022), 100055. doi:10.1016/j.caeai.2022.100055 [25]Biniam Gebru, Lydia Zeleke, Daniel Blankson, Mahmoud Nabil, Shamila Nateghi, Abdollah Homaifar, and Edward Tunstel. 2022. A Review on Human–Machine Trust Evaluation: Human-Centric and Machine-Centric Perspectives. IEEE Transactions on Human-Machine Systems 52, 5 (2022), 952–962. doi:10.1109/ THMS.2022.3144956 [26]Michael Gerlich. 2025. AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies (Basel) 15, 1 (Jan. 2025), 6. doi:10.3390/ soc15010006 [27]Katy Ilonka Gero, Zahra Ashktorab, Casey Dugan, Qian Pan, James Johnson, Werner Geyer, Maria Ruiz, Sarah Miller, David R Millen, Murray Campbell, Sadhana Kumaravel, and Wei Zhang. 2020. Mental Models of AI Agents in a Cooperative Game Setting. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–12. doi:10.1145/3313831.3376316 [28]Katy Ilonka Gero, Vivian Liu, and Lydia Chilton. 2022. Sparks: Inspiration for Science Writing using Language Models. In Proceedings of the 2022 ACM Designing Interactive Systems Conference (Virtual Event, Australia) (DIS ’22). Asso- ciation for Computing Machinery, New York, NY, USA, 1002–1019. doi:10.1145/ 3532106.3533533 CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. [29]Katy Ilonka Gero, Tao Long, and Lydia B Chilton. 2023. Social Dynamics of AI Support in Creative Writing. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23, Article 245). Association for Computing Machinery, New York, NY, USA, 1–15. [30]Steven M Goodman, Erin Buehler, Patrick Clary, Andy Coenen, Aaron Donsbach, Tiffanie N Horne, Michal Lahav, Robert MacDonald, Rain Breaw Michaels, Ajit Narayanan, Mahima Pushkarna, Joel Riley, Alex Santana, Lei Shi, Rachel Sweeney, Phil Weaver, Ann Yuan, and Meredith Ringel Morris. 2022. LaMPost: Design and Evaluation of an AI-assisted Email Writing Prototype for Adults with Dyslexia. In Proceedings of the 24th International ACM SIGACCESS Conference on Comput- ers and Accessibility (Athens, Greece) (ASSETS ’22, Article 24). Association for Computing Machinery, New York, NY, USA, 1–18. doi:10.1145/3517428.3544819 [31]Grammarly. 2019. How Correctness Keeps Your Writing Sharp | Grammarly Spotlight. https://w.grammarly.com/blog/product/correctness-dimension/. Accessed: 2025-6-17. [32]Alicia Guo, Shreya Sathyanarayanan, Leijie Wang, Jeffrey Heer, and Amy X Zhang. 2025. From pen to prompt: How creative writers integrate AI into their writing practice. In Proceedings of the 2025 Conference on Creativity and Cognition. ACM, New York, NY, USA, 527–545. doi:10.1145/3698061.3726910 [33]Angel Hsing-Chi Hwang, Q Vera Liao, Su Lin Blodgett, Alexandra Olteanu, and Adam Trischler. 2025. ’It was 80% me, 20% AI’: Seeking Authenticity in Co- Writing with Large Language Models. Proc. ACM Hum. Comput. Interact. 9, 2 (May 2025), 1–41. doi:10.1145/3711020 [34]Daphne Ippolito, Ann Yuan, Andy Coenen, and Sehmon Burnam. 2022. Creative Writing with an AI-Powered Writing Assistant: Perspectives from Professional Writers. arXiv:2211.05030 [cs.HC] arXiv preprint. [35]Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023. Co-Writing with Opinionated Language Models Affects Users’ Views. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23, Article 111). Association for Computing Machinery, New York, NY, USA, 1–15. doi:10.1145/3544548.3581196 [36] Philip N Johnson-Laird. 1983. Mental Models. Harvard University Press, London, England. [37]Kowe Kadoma, Marianne Aubin Le Quere, Xiyu Jenny Fu, Christin Munsch, Danaë Metaxa, and Mor Naaman. 2024. The role of inclusion, control, and ownership in workplace AI-mediated communication. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–10. doi:10.1145/3613904.3642650 [38] Jeongyeon Kim, Sangho Suh, Lydia B Chilton, and Haijun Xia. 2023. Metaphorian: Leveraging Large Language Models to Support Extended Metaphor Creation for Science Writing. In Proceedings of the 2023 ACM Designing Interactive Systems Conference (Pittsburgh, PA, USA) (DIS ’23). Association for Computing Machinery, New York, NY, USA, 115–135. doi:10.1145/3563657.3595996 [39] Sunnie S Y Kim, Jennifer Wortman Vaughan, Q Vera Liao, Tania Lombrozo, and Olga Russakovsky. 2025. Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–19. doi:10.1145/3706598.3714020 [40] Hannah Rose Kirk, Henry Davidson, Ed Saunders, Lennart Luettgau, Bertie Vidgen, Scott A Hale, and Christopher Summerfield. 2025. Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships. [41]Aniket Kittur, Ed H. Chi, and Bongwon Suh. 2008.Crowdsourcing User Studies with Mechanical Turk. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 453–456. doi:10.1145/1357054.1357127 [42] Todd Kulesza, Simone Stumpf, Margaret Burnett, Sherry Yang, Irwin Kwan, and Weng-Keen Wong. 2013. Too Much, Too Little, or Just Right? Ways Explanations Impact End Users’ Mental Models. In 2013 IEEE Symposium on Visual Languages and Human-Centric Computing. IEEE, San Jose, CA, USA, 3–10. doi:10.1109/ vlhcc.2013.6645235 [43]Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, Simon Knight, Seyed Parsa Neshaei, Antonette Shibani, Disha Shrivastava, Lila Shroff, Agnia Sergeyuk, Jessi Stark, Sarah Sterman, Sitong Wang, Antoine Bosselut, Daniel Buschek, Joseph Chee Chang, Sherol Chen, Max Kreminski, Joonsuk Park, Roy Pea, Eugenia Ha Rim Rho, Zejiang Shen, and Pao Siangliulue. 2024. A Design Space for Intelligent and Interactive Writing Assistants. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 1054, 35 pages. doi:10.1145/3613904.3642697 [44] Mina Lee, Percy Liang, and Qian Yang. 2022. CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities. In Pro- ceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22, Article 388). Association for Computing Machinery, New York, NY, USA, 1–19. doi:10.1145/3491102.3502030 [45]Nancy Leveson. 2004. A new accident model for engineering safer systems. Saf. Sci. 42, 4 (April 2004), 237–270. doi:10.1016/s0925-7535(03) 00047-x [46]Nancy Leveson and John Thomas. 2018.STPA Handbook.https:// psas.scripts.mit.edu/home/get_file.php?name=STPA_handbook.pdf . [47]Nancy G Leveson. 2012. Engineering a safer world: Systems thinking applied to safety. MIT Press, London, England. [48] Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin. 2024. How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?. In Proceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Lin- guistics, Miami, Florida, USA, 4849–4868. doi:10.18653/v1/2024.emnlp-main.279 [49]Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin. 2024. The Value, Benefits, and Concerns of Generative AI-Powered Assistance in Writing. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1048:1–1048:25. doi:10.1145/3613904.3642625 [50]Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2025. Towards human-AI deliberation: Design and evaluation of LLM-empowered deliberative AI for AI-assisted decision-making. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–23. doi:10.1145/3706598.3713423 [51]Khalid Mehmood, Katrien Verleye, Arne De Keyser, and Bart Larivière. 2025. From promises to practice: Unravelling users’ mental models about Large Language Models at individual and societal levels. doi:10.2139/ssrn.5598465 Preprint paper. [52]Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co- Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23, Article 355). Association for Computing Machinery, New York, NY, USA, 1–34. doi:10.1145/3544548.3581225 [53]Bonnie M. Muir. 1987. Trust Between Humans and Machines and the Design of Decision Aids. International Journal of Man-Machine Studies 27, 5–6 (1987), 527–539. doi:10.1016/S0020-7373(87) 80013-5 [54] Sheryl Wei Ting Ng and Renwen Zhang. 2025. Trust in AI chatbots: A sys- tematic review. Telemat. Inform. 97, 102240 (Feb. 2025), 102240. doi:10.1016/ j.tele.2025.102240 [55] D A Norman. 1987. Some observations on mental models. In Human-computer interaction: a multidisciplinary approach. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 241–244. [56]Mahsan Nourani, Chiradeep Roy, Jeremy E Block, Donald R Honeycutt, Tahrima Rahman, Eric Ragan, and Vibhav Gogate. 2021. Anchoring Bias Affects Mental Model Formation and User Reliance in Explainable AI Systems. In 26th Interna- tional Conference on Intelligent User Interfaces (IUI ’21). Association for Computing Machinery, New York, NY, USA, 340–350. doi:10.1145/3397481.3450639 [57]Aswati Panicker, Novia Nurain, Zaidat Ibrahim, Chun-Han (ariel) Wang, Se- ung Wan Ha, Yuxing Wu, Kay Connelly, Katie A Siek, and Chia-Fang Chung. 2024. Understanding fraudulence in online qualitative studies: From the researcher’s perspective. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–17. doi:10.1145/3613904.3642732 [58]Gabriele Paolacci and Jesse Chandler. 2014. Inside the Turk: Understanding mechanical Turk as a participant pool. Curr. Dir. Psychol. Sci. 23, 3 (June 2014), 184–188. doi:10.1177/0963721414531598 [59] Samir Passi, Shipi Dhanorkar, and Mihaela Vorvoreanu. 2025. Addressing over- reliance on AI. In Handbook of Human-Centered Artificial Intelligence. Springer Nature Singapore, Singapore, 1–34. doi:10.1007/978-981-97-8440-0_98-1 [60] Ritika Poddar, Rashmi Sinha, Mor Naaman, and Maurice Jakesch. 2023. AI Writing Assistants Influence Topic Choice in Self-Presentation. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI EA ’23, Article 29). Association for Computing Machinery, New York, NY, USA, 1–6. doi:10.1145/3544549.3585893 [61]Jens Rasmussen. 1987. Mental models and the control of action in complex envi- ronments. In Selected papers of the 6th Interdisciplinary Workshop on Informatics and Psychology: Mental Models and Human-Computer Interaction 1. North-Holland Publishing Co., NLD, 41–69. [62]Mohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry, Joseph Jay Williams, and Anastasia Kuzminykh. 2025. Co-writing with AI, on human terms: Aligning research with user demands across the writing process. Proc. ACM Hum. Comput. Interact. 9, 7 (Oct. 2025), 1–37. doi:10.1145/3757566 [63]Shalaleh Rismani, Renee Shelby, Leah Davis, Negar Rostamzadeh, and Ajung Moon. 2025. Measuring what matters: Connecting AI ethics evaluations to system attributes, hazards, and harms. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 8, 3 (Oct. 2025), 2199–2213. [64]Ronald E Robertson, Alexandra Olteanu, Fernando Diaz, Milad Shokouhi, and Peter Bailey. 2021. “I Can’t Reply with That”: Characterizing Problematic Email Reply Suggestions. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Ma- chinery, New York, NY, USA, Article 724, 18 pages. doi:10.1145/3411764.3445557 [65]G. Romeo and E. Conti. 2025. Exploring automation bias in human–AI collabo- ration: a review and implications for explainable AI. AI & Society 40, 3 (2025), 789–812. doi:10.1007/s00146-025-02422-7 From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain [66]Nadine B. Sarter, Christopher D. Wickens, Randall J. Mumaw, Steve Kimball, Roger Marsh, Mark I. Nikolic, and W. Q. Xu. 2003. Modern flight deck au- tomation: pilots’ mental model and monitoring patterns and performance. https://api.semanticscholar.org/CorpusID:107831159.Presented at the 12th In- ternational Symposium on Aviation Psychology. [67]Ulrike Schäfer, Lars Sipos, and Claudia Müller-Birn. 2025. ‘The AI is uncertain, so am I. What now?’: Navigating Shortcomings of Uncertainty Representations in Human-AI Collaboration with Capability-focused Guidance. Proc. ACM Hum. Comput. Interact. 9, 7 (Oct. 2025), 1–48. doi:10.1145/3757451 [68]Sathya S Silva and R John Hansman. 2015. Divergence Between Flight Crew Mental Model and Aircraft System State in Auto-Throttle Mode Confusion Acci- dent and Incident Cases. Journal of Cognitive Engineering and Decision Making 9, 4 (Dec. 2015), 312–328. doi:10.1177/1555343415597344 [69]Nancy Staggers and A F Norcio. 1993. Mental models: concepts for human- computer interaction research. Int. J. Man. Mach. Stud. 38, 4 (April 1993), 587–605. https://w.sciencedirect.com/science/article/pii/S002073738371028X [70]Yujie Sun, Dongfang Sheng, Zihan Zhou, and Yifei Wu. 2024. AI hallucination: towards a comprehensive classification of distorted information in artificial intelligence-generated content. Humanit. Soc. Sci. Commun. 11, 1 (Sept. 2024), 1–14. doi:10.1057/s41599-024-03811-x [71]Siddharth Swaroop, Zana Buçinca, Krzysztof Z Gajos, and Finale Doshi-Velez. 2025. Personalising AI assistance based on overreliance rate in AI-assisted deci- sion making. In Proceedings of the 30th International Conference on Intelligent User Interfaces. ACM, New York, NY, USA, 1107–1122. doi:10.1145/3708359.3712128 [72]Christopher L Tarola, Sameer Hirji, Steven J Yule, Jennifer M Gabany, Alessandro Zenati, Roger D Dias, and Marco A Zenati. 2018. Cognitive Support to Promote Shared Mental Models during Safety-Critical Situations in Cardiac Surgery (Late Breaking Report). In 2018 IEEE Conference on Cognitive and Computational As- pects of Situation Management (CogSIMA). Institute of Electrical and Electronics Engineers, Piscataway, NJ, USA, 165–167. doi:10.1109/COGSIMA.2018.8423991 [73]Mor Vered, Tali Livni, Piers Douglas Lionel Howe, Tim Miller, and Liz Sonenberg. 2023. The effects of explanations on automation bias. Artif. Intell. 322, 103952 (Sept. 2023), 103952. doi:10.1016/j.artint.2023.103952 [74]Kailas Vodrahalli, Roxana Daneshjou, Tobias Gerstenberg, and James Zou. 2022. Do Humans Trust Advice More if it Comes from AI? An Analysis of Human- AI Interactions. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (Oxford, United Kingdom) (AIES ’22). Association for Computing Machinery, New York, NY, USA, 763–777. doi:10.1145/3514094.3534150 [75] Hanna Wallach, Meera Desai, A Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas J Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaughan, Matthew Vogel, Hannah Washington, and Abigail Z Jacobs. 2025. Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge. In International Conference on Machine Learning. PMLR, Vancouver, Canada, 82232–82251. [76] Xingyi Wang, Xiaozheng Wang, Sunyup Park, and Yaxing Yao. 2025. Users’ Mental Models of Generative AI Chatbot Ecosystems. In Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’25). ACM, Cagliari, Italy. doi:10.1145/3708359.3712125 [77]Gesa Wiegand, Matthias Schmidmaier, Thomas Weber, Yuanting Liu, and Hein- rich Hussmann. 2019. I Drive - You Trust: Explaining Driving Behavior Of Autonomous Cars. In Extended Abstracts of the 2019 CHI Conference on Hu- man Factors in Computing Systems (Glasgow, Scotland Uk) (CHI EA ’19, Paper LBW0163). Association for Computing Machinery, New York, NY, USA, 1–6. doi:10.1145/3290607.3312817 [78] Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: Story Writing With Large Language Models. In 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 841–852. doi:10.1145/3490099.3511105 [79]Chunpeng Zhai, Santoso Wibowo, and Lily D Li. 2024. The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: a systematic review. Smart Learn. Environ. 11, 1 (June 2024), 1–37. doi:10.1186/s40561-024-00316-7 CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. A Demographics Survey This section lists the survey questions we administered to prospec- tive participants in order to select and stratify participants for our study. (1) What is your age? Single line text (2) What is your gender? Single choice • Female • Male • Non-binary • Prefer not to answer • Other (3) What is the highest level of education you have completed? Single choice • Primary school • High school • Post-secondary vocational institution (trade and technical school) or CEGEP • 3 or 4-year undergraduate university program • Master’s degree program • PhD program • None of the above • Prefer not to answer (4) How often do you write in English? 5-point Likert scale (5) How would you rate your confidence in writing in English? 5-point Likert scale (6) What is your native language? Single choice • English • French • Other (7) How would you rate your cover letter writing abilities? 5-point Likert scale (8) How would you rate your level of confidence in writing cover letters? 5-point Likert scale (9)Which one of the following AI-based writing assistants have you used? Multiple choice(can choose multiple) • ChatGPT • Grammarly • Wordtune • Notion • None of the above • Other (10) How often do you use any of the AI-based writing assistants stated in the previous question? 5-point Likert scale B System Understanding Survey This section lists the survey questions we used to probe partici- pants understanding of the AI-based writing assistant platform after they watched the instructions videos for one of the experimental conditions they were assigned to. (1)CoAuthor is a writing assistant that does not rely on Artificial Intelligence (AI). • True • False • I do not know. (2) CoAuthor offers phrases and sentences when a user requests suggestions. • True • False • I do not know. (3) Where should the user write the main body of their cover letter on the CoAuthor platform? • Text editor • Instruction box • Both • I do not know. (4) CoAuthor will automatically generate suggestions when the user stops typing. • True • False • I do not know. (5) CoAuthor filters and trims the original list of suggestions generated by the large language model before showing it to the user. • True • False • I do not know. (6) CoAuthor gives suggestions based on what is a likely con- tinuation of the last sentence in the text editor. • True • False • I do not know. (7)The content of the Instruction Box is sent to a different language model than the text editor. • True • False • I do not know. C Post-task Survey This section lists the questions asked in the post task survey in order to assess participants’ perceptions of their experience with the AI-based assistant. The responses were given on 5-point Likert scales. (1)Rate the following statements about the cover letter that was written using CoAuthor (5-point Likert Scale). •I am the main contributor to the content of the cover letter. • I am accountable for all aspects of the cover letter. • I feel like I am the author of the cover letter. • I feel like the cover letter represents me. (2)Rate the following statements about the cover letter that was written using CoAuthor (5-point Likert Scale). • I am the main contributor to the content of the cover letter. • I am accountable for all aspects of the cover letter. • I feel like I am the author of the cover letter. • I feel like the cover letter represents me. • I was able to customize the provided From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain (3)Rate the following statements about your experience of con- trolling CoAuthor when writing the cover letter (5-point Likert Scale). •The suggestions influenced the content of the cover letter. •I was able to customize the provided suggestions using the Instruction Box. • I felt in control of the quality and content of the generated suggestions. • I felt like I was writing the text and CoAuthor was assisting me. • I felt like CoAuthor was writing the text and I was assist- ing. • I felt in control of what was included in the final letter. • CoAuthor was easy to use. (4) Rate the following statements about the quality of the sug- gestions, the final letter, and CoAuthor (5-point Likert Scale). • The suggestions were relevant. • The suggestions were error-free. • The suggestions were helpful. • Customizing suggestions using the Instruction Box was helpful. • The final letter is error-free. •The final letter has the appropriate style and content for a cover letter. D Cover Letter Assessment and Instruction Coding Rubrics This section reports the assessment criteria we used to rate the overall quality of the cover letters (§D.1), as well as the coding rubric we used to assess the type of instructions participants provided in the instruction box (§D.2). D.1 Qualitative Quality Assessment Rubric Table 5 describes the three criteria the research team used to score the cover letters. Table 5: Cover letter assessment rubric used by evaluators. CriterionRating anchor (5 = Excellent), (1 = Poor) Demonstration of relevant skills and experience The letter clearly highlights the candidate’s skills, knowledge, and experiences that directly relate to the job posting. Examples are used to support claims. Flow and clarityThe flow of the letter is appropriate for a cover letter and the writing is clear. The content transitions smoothly from one idea to the next. Sentences and paragraphs are logically organized, making the letter easy to read and understand without confusion or abrupt shifts. Tone and styleThe writing has an appropriate and compelling tone and style for a cover letter. The letter has a professional tone and the candidate’s voice comes through as confident, enthusiastic, and sincere. D.2 Instruction Coding Rubric Table 6 includes a brief description of the criterion that the research team used to annotate each one of the instructions that participants wrote. E Manipulation Check Details This section includes further details about the manipulation check results. Table 7 provides full item-level descriptive statistics and Table 6: Instruction coding rubric used to annotate user- written instructions. CriterionResponse Tone or style shift Tone is focused on the emotional stance and attitude the writer brings through in their writing. Style is broader than tone, and captures how the writer expresses their message using different words, unique voice, and structure. For example, “use positive language” shifts tone; “make words professional” or “decrease monotony” shifts style. Yes / No Qualifications articulation Does the instruction articulate the candidate’s relevant qualities, quali- fications, or skills? For example, “add graduate degree at McGill Uni- versity” is specific, whereas “highlight great asset” is too broad. Yes / No Structure prompting Does the instruction cue the structure of the cover letter (e.g., ending, section transition)? For example, “conclusive sentence” or “write some- thing about company.” Yes / No nonparametric test results for all manipulation-check questions referenced in the main paper. It reports item-level descriptive statis- tics (mean, standard deviation, and median) for each manipulation- check question by condition, along with two-tailed Mann–Whitney U tests (U, Z, p). Items were coded as 1 = correct and 0 = incor- rect; means therefore represent proportions correct. The final row (“Total”) corresponds to the summed manipulation-check score reported in the main paper. Table 7: Descriptive statistics and Mann–Whitney U tests for manipulation check questions by condition. QFunctional (M = 1)Structural (M = 2)Mean RankUZp MeanSDMedianMeanSDMedianMM=1 M=2 11.000.001.000.790.411.0027.0022.00228.0 -2.34.019 2 1.000.001.000.960.201.0025.0024.00276.0 -1.00.317 30.960.201.000.960.201.0024.5024.50288.0 0.001.000 4 1.000.001.001.000.001.0024.5024.50288.0 0.001.000 50.210.410.000.790.411.0017.5031.50120.0 -4.00 <.001 6 0.750.441.000.920.281.0022.5026.50240.0 -1.53.125 70.250.440.000.670.481.0019.5029.50168.0 -2.87.004 Total5.170.925.006.081.066.0018.4430.56142.5 -3.15.002 F Control Behavior Measures Statistical Analysis Details This section elaborates on the results from the statistical analy- sis completed for various measures for control behavior. Table 8 reports the full descriptive and inferential statistics for the inter- action measures summarized in Figure 3. For each measure, we report group-level descriptive statistics for the functional and struc- tural conditions, along with the statistical test used, corresponding test statistics, exact푝-values, and effect sizes. Measures that met normality assumptions were analyzed using independent-samples 푡-tests, whereas measures that violated normality assumptions were analyzed using Mann–Whitney푈 tests. G Instruction Coding Statistical Analysis Details In this section, we provide more indepth information about two types of statistical analysis that we conducted to compare the in- structions that the participants provided across the mental model CHI ’26, April 13–17, 2026, Barcelona, SpainRismani et al. Table 8: Full statistical results for interaction measures. MeasureFunctional conditionStructural conditionTestStatistic푝Effect size Suggestion requests (calibrated) 푀= 0.071,푆퐷= 0.047 푀= 0.074,푆퐷= 0.036 Mann–Whitney푈 푈= 248,푍=−0.83 .409 푟= 0.12 Acceptance ratio푀= 0.651,푆퐷= 0.188 푀= 0.685,푆퐷= 0.161푡(46)푡=−0.67.253 푑=−0.19 Edit ratio푀= 0.173,푆퐷= 0.154 푀= 0.179,푆퐷= 0.153 Mann–Whitney푈 푈= 280,푍=−0.17 .869 푟= 0.02 Instruction count푀= 0.028,푆퐷= 0.021 푀= 0.024,푆퐷= 0.016 Mann–Whitney푈 푈= 271,푍=−0.35 .726 푟= 0.05 User character ratio푀= 0.609,푆퐷= 0.227 푀= 0.578,푆퐷= 0.187푡(46)푡= 0.51.305 푑= 0.15 conditions. Table 9 overviews the distribution of codes for partici- pants instructions across the functional and structural conditions. For each instruction dimension, the table shows the number of instructions in which the feature was absent (0) or present (1) for each condition, along with the results of chi-squared tests assess- ing whether instruction type was associated with experimental condition. Table 9: Distribution of coded instruction dimensions by condition (instruction-level). Instruction dimensionFunctional (0 / 1) Structural (0 / 1) Total 푁 휒 2 (1) 푝 Tone / style shift154 / 13133 / 153150.54.464 Qualifications articulation51 / 10534 / 1023151.30.254 Structure prompting137 / 30123 / 253150.06.802 Table 10 reports binomial tests evaluating whether the distri- bution of instruction types across conditions deviated from an ex- pected 50/50 split. For each instruction dimension, the table shows the number of instructions coded as present in each condition and the corresponding two-tailed binomial test results. Table 10: Binomial tests evaluating deviations from a 50/50 distribution across conditions for each instruction dimen- sion. Instruction dimensionFunctional (푛) Structural (푛) Total 푁 Binomial 푝 Tone / style shift (present)131528.851 Qualifications articulation (present)112108220.840 Structure prompting (present)302555.590 H Grammatical Error Statistical Analysis Details In this section, we provide more details on the exploratory statis- tical analysis we conducted to understand why the cover letters written by structural participants had more grammatical, spelling and punctuation errors. Table 11 reports exploratory two-way analyses of variance exam- ining whether differences in erroneous suggestion acceptance mea- sures persisted after controlling for participants’ native language. For each dependent variable, the table reports the main effects of condition and native language, as well as their interaction. Table 12 reports descriptive statistics and inferential test results for grammatical correctness and erroneous suggestion acceptance measures underlying Figure 4. The table summarizes group-level means and standard deviations for the functional and structural conditions, along with the statistical tests, test statistics,푝-values, and effect sizes used to evaluate differences between conditions. Table 11: Two-way ANOVA results for erroneous suggestion acceptance measures, controlling for native language. Dependent variableEffect퐹(1, 44) 푝Partial휂 2 Error ratioCondition (MMNumeric)2.36.132.051 Native language1.58.216.035 Condition× Language0.41.526.009 Errors per wordCondition (MMNumeric)3.37.073.071 Native language1.32.258.029 Condition× Language0.36.554.008 I Cover Letter Qualitative Quality Assessment Details This subsection provides additional detail on the qualitative eval- uation of cover letter outputs produced during the study. Table 13 reports descriptive statistics and nonparametric test results for qualitative evaluations of cover letter quality by condition. Letters were assessed and rated across three dimensions—relevance, flow and clarity, and tone/style appropriateness—and ratings distribu- tions were then compared between the functional and structural conditions using Mann–Whitney푈 tests. J Post-task Survey Statistical Analysis Details This subsection reports the full statistical details for the post-task self-report measures collected after participants completed the writ- ing task. Table 14 reports the full breakdown of the two composite self-report measures used in the study—Usefulness and Perceived ownership. For each composite, we list the constituent items, inter- nal reliability, and group-level descriptive statistics by condition, along with the corresponding independent-samples푡-tests. These results provide transparency into how each composite measure was constructed and confirm that there were no statistically significant differences between conditions at either the scale or item level. Table 15 presents descriptive and inferential statistics for all re- maining single-item self-report measures that were not included in composite scales. These items capture participants’ perceptions of system usability, suggestion quality, perceived control, and au- thorship. Consistent with the results reported in the main paper, none of these measures showed statistically significant differences between the functional and structural conditions. From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsCHI ’26, April 13–17, 2026, Barcelona, Spain Table 12: Grammatical correctness and erroneous suggestion acceptance measures by condition. MeasureFunctional conditionStructural conditionTestStatistic푝Effect size Calibrated corrections 푀= 0.016,푆퐷= 0.010 푀= 0.023,푆퐷= 0.010푡(46)푡=−2.32.013 푑= 0.67 Error ratio푀= 0.146,푆퐷= 0.119 푀= 0.226,푆퐷= 0.157 Mann–Whitney푈 푈= 204.5,푍=−1.73 .084 푟= 0.25 Errors per word푀= 0.006,푆퐷= 0.007 푀= 0.011,푆퐷= 0.008 Mann–Whitney푈 푈= 170,푍=−2.44 .015 푟= 0.35 Table 13: Cover letter quality ratings by condition. MeasureFunctional conditionStructural conditionTestStatistic푝Effect size Relevance (Q1) 푀= 3.67,푆퐷= 0.87 푀= 3.83,푆퐷= 0.96 Mann–Whitney푈 푈= 249,푍=−0.85 .398 푟=−0.12 Flow (Q2)푀= 4.04,푆퐷= 0.81 푀= 3.96,푆퐷= 0.91 Mann–Whitney푈 푈= 280.5,푍=−0.16 .869 푟=−0.02 Tone / style (Q3) 푀= 4.17,푆퐷= 0.64 푀= 4.04,푆퐷= 0.96 Mann–Whitney푈 푈= 285.5,푍=−0.06 .954 푟=−0.01 Table 14: Composite self-report scales and constituent items by condition. Scale / ItemTypeFunctional conditionStructural conditionTest statistic 푝 (one-sided) Usefulness (훼= .76) Usefulness (composite)Scale mean 푀= 3.89,푆퐷= 0.56 푀= 4.01,푆퐷= 0.55 푡(46)=−0.78.220 Suggestions were relevantItem푀= 3.88,푆퐷= 0.68 푀= 4.04,푆퐷= 0.46 푡(46)=−0.99.163 Suggestions were helpfulItem푀= 3.96,푆퐷= 0.69 푀= 4.13,푆퐷= 0.68 푡(46)=−0.84.202 Customizing suggestions using the Instruction Box was helpfulItem푀= 3.75,푆퐷= 0.79 푀= 3.79,푆퐷= 0.66 푡(46)=−0.20.422 Able to customize the provided suggestionsItem푀= 3.96,푆퐷= 0.96 푀= 4.08,푆퐷= 0.83 푡(46)=−0.48.315 Perceived ownership (훼= .71) Perceived ownership (composite)Scale mean 푀= 3.95,푆퐷= 0.64 푀= 3.93,푆퐷= 0.78 푡(46)= 0.10.460 I am the main contributor to the content of the cover letterItem푀= 4.29,푆퐷= 0.69 푀= 4.08,푆퐷= 0.93 푡(46)= 0.88.191 I am accountable for all aspects of the cover letterItem푀= 3.63,푆퐷= 1.13 푀= 3.67,푆퐷= 1.17 푡(46)=−0.13.450 I feel like I am the author of the cover letterItem푀= 3.88,푆퐷= 0.85 푀= 4.04,푆퐷= 0.96 푡(46)=−0.64.263 I feel like the cover letter represents meItem푀= 4.00,푆퐷= 0.89 푀= 3.92,푆퐷= 1.06 푡(46)= 0.30.384 Table 15: Single-item self-report measures by condition. MeasureFunctional conditionStructural condition 푡(46) 푝 (one-sided) CoAuthor was easy to use푀= 3.88,푆퐷= 1.04 푀= 4.38,푆퐷= 0.58 −2.07.022 The final letter is error-free푀= 3.38,푆퐷= 1.06 푀= 3.46,푆퐷= 1.22 −0.25.400 The final letter has appropriate style and content푀= 3.96,푆퐷= 1.00 푀= 4.25,푆퐷= 0.79 −1.12.134 The suggestions influenced the content of the cover letter푀= 3.17,푆퐷= 1.17 푀= 3.50,푆퐷= 0.93 −1.09.140 I felt in control of the quality and content of the suggestions 푀= 3.25,푆퐷= 1.07 푀= 2.92,푆퐷= 1.18 1.03.155 I felt in control of what was included in the final letter푀= 4.38,푆퐷= 0.92 푀= 4.46,푆퐷= 0.66 −0.36.360 I felt like I was writing the text and CoAuthor was assisting me 푀= 3.79,푆퐷= 0.66 푀= 4.00,푆퐷= 1.06 −0.82.209 I felt like CoAuthor was writing the text and I was assisting푀= 1.96,푆퐷= 1.00 푀= 2.08,푆퐷= 1.02 −0.43.335 The suggestions were error-free푀= 1.71,푆퐷= 0.75 푀= 2.13,푆퐷= 1.15 −1.48.072 The suggestions were helpful푀= 3.96,푆퐷= 0.69 푀= 4.13,푆퐷= 0.68 −0.84.202