Paper deep dive
MetaCues: Enabling Critical Engagement with Generative AI for Information Seeking and Sensemaking
Anjali Singh, Karan Taneja, Zhitong Guan, Soo Young Rieh
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/23/2026, 12:07:43 PM
Summary
MetaCues is a novel GenAI-based interactive tool designed to mitigate cognitive offloading and promote metacognitive engagement during information seeking. By providing dynamically generated metacognitive cues (Orienting, Monitoring, Broadening Perspectives, Consolidation, Source Engagement, Persistent Inquiry, and Independent Thinking) alongside AI responses, the tool encourages users to reflect on their search process. An online study (N=146) demonstrated that MetaCues increases user confidence in attitudinal judgments and fosters broader inquiry, particularly for less controversial and less familiar topics.
Entities (5)
Relation Signals (4)
Anjali Singh â developed â MetaCues
confidence 100% · We developed MetaCues, a novel GenAI-based interactive tool
MetaCues â provides â Metacognitive Cues
confidence 100% · MetaCues... delivers metacognitive cues alongside AI responses
MetaCues â utilizes â OpenAI GPT-4o
confidence 100% · The conversational AI interface uses OpenAI GPT-4o model
Metacognitive Cues â increases â Confidence in Attitudinal Judgments
confidence 90% · MetaCues leads to increased confidence in attitudinal judgments about the search topic
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generative AI (GenAI) search tools are increasingly used for information seeking, yet their design tends to encourage cognitive offloading, which may lead to passive engagement, selective attention, and informational homogenization. Effective use requires metacognitive engagement to craft good prompts, verify AI outputs, and critically engage with information. We developed MetaCues, a novel GenAI-based interactive tool for information seeking that delivers metacognitive cues alongside AI responses and a note-taking interface to guide users' search and associated learning. Through an online study (N = 146), we compared MetaCues to a baseline tool without cues, across two broad search topics that required participants to explore diverse perspectives in order to make informed judgments. Preliminary findings regarding participants' search behavior show that MetaCues leads to increased confidence in attitudinal judgments about the search topic as well as broader inquiry, with the latter effect emerging primarily for the topic that was less controversial and with which participants had relatively less familiarity. Accordingly, we outline directions for future qualitative exploration of search interactions and inquiry patterns.
Tags
Links
- Source: https://arxiv.org/abs/2603.19634v1
- Canonical: https://arxiv.org/abs/2603.19634v1
Trouble viewing inline? Open PDF directly â
Full Text
37,419 characters extracted from source content.
Expand or collapse full text
MetaCues: Enabling Critical Engagement with Generative AI for Information Seeking and Sensemaking Anjali Singh anjali.singh@ischool.utexas.edu The University of Texas at AustinUSA , Karan Taneja ktaneja6@gatech.edu Georgia Institute of TechnologyUSA , Zhitong Guan klarazt@utexas.edu The University of Texas at AustinUSA and Soo Young Rieh rieh@ischool.utexas.edu The University of Texas at AustinUSA Abstract. Generative AI (GenAI) search tools are increasingly used for information seeking, yet their design tends to encourage cognitive offloading, which may lead to passive engagement, selective attention, and informational homogenization. Effective use requires metacognitive engagement to craft good prompts, verify AI outputs, and critically engage with information. We developed MetaCues, a novel GenAI-based interactive tool for information seeking that delivers metacognitive cues alongside AI responses and a note-taking interface to guide usersâ search and associated learning. Through an online study (N=146N=146), we compared MetaCues to a baseline tool without cues, across two broad search topics that required participants to explore diverse perspectives in order to make informed judgments. Preliminary findings regarding participantsâ search behavior show that MetaCues leads to increased confidence in attitudinal judgments about the search topic as well as broader inquiry, with the latter effect emerging primarily for the topic that was less controversial and with which participants had relatively less familiarity. Accordingly, we outline directions for future qualitative exploration of search interactions and inquiry patterns. Information Seeking, Generative AI, Metacognitive Cues â isbn: 978-1-4503-X-X/2018/06â ccs: Information systems Users and interactive retrievalâ ccs: Human-centered computing Systems and tools for interaction designâ copyright: none 1. Introduction Generative AI (GenAI) search tools, such as Perplexity.ai and Google AI Overview, have been rapidly adopted and promoted by major tech companies. These tools are typically valued for their speed, convenience, and the ability to synthesize information from multiple relevant sources into coherent responses (Zhou and Li, 2024). However, the growing use GenAI for information seeking has raised several concerns. The design of GenAI search tools inherently encourages cognitive offloading (Singh et al., 2025b), which refers to the delegation of cognitive processes to an external system, such as AI (Risko and Gilbert, 2016). In particular, using GenAI for information seeking can bypass the cognitive processes of understanding, applying, analyzing, and evaluating information, potentially undermining learning and critical engagement with information (Narayanan Venkit et al., 2025; Singh et al., 2025b). Research suggests that information seeking with GenAI can lead to passive engagement and limit exposure to diverse perspectives, which can contribute to informational homogenization (Amer and Elboghdadly, 2024; Narayanan Venkit et al., 2025; Solaiman et al., 2023). Beyond influencing search behavior, the fluent, confident, and sycophantic nature of GenAI responses can also shape usersâ confidence in the judgments they form based on AI-generated content (Amer and Elboghdadly, 2024; Singh et al., 2025b). These concerns are particularly salient in the context of controversial topics, where people are more likely to align with information that reinforces their pre-existing beliefs (VedejovĂĄ and ÄavojovĂĄ, 2022). In such settings, the tendency of GenAI systems to exhibit sycophancy may further amplify confirmation bias. Given these challenges, effective use of GenAI tools for information seeking requires metacognitive engagement, which involves being aware of and regulating over oneâs thinking (Winne, 2017; Tankelevitch et al., 2024). Metacognitionâcommonly referred to as âthinking about thinkingââis essential not only for effectively prompting GenAI, which requires monitoring task goals and planning, but also for being vigilant of potential hallucinations in AI-generated responses (Tankelevitch et al., 2024). Recent work (Singh et al., 2025a) shows that metacognitive cuesâthat prompt users to pause, reflect, assess their comprehension and consider multiple perspectivesâdelivered while searching with GenAI tools can foster active engagement, broader exploration, and more thoughtful follow-up questioning. However, since that study employed a Wizard-of-Oz setup, where researchers manually delivered cues while monitoring participantsâ search behavior, it remains unclear whether such support can be provided autonomously. Figure 1. A snapshot of the MetaCues interface, and the process of generating metacognitive cues based on user-generated data. To address this gap, we designed and developed MetaCues (Figure 1), an interactive GenAI-based tool to support information seeking featuring three panels: a chat interface on the right where users query the AI and view responses with linked sources, and a notepad on the left, below which automatically generated metacognitive cues are displayed. MetaCues analyzes usersâ questions, AI responses, and notes to proactively deliver tailored cues that encourage active and critical user engagement during the information seeking process. To examine the effects of metacognitive cues on usersâ search behavior and resulting judgments across different search topics, we conducted an online study (N=146) comparing MetaCues to a baseline system, which was identical in every respect except for the absence of cues. Participants were randomly assigned one of two information-seeking tasks. This work describes the design and development of the MetaCues system and reports preliminary findings from this study addressing the following research question: What are the effects of metacognitive cues on: (i) usersâ search behavior, and (i) confidence in their attitudinal judgments regarding the assigned topic, during a GenAI-assisted information-seeking and sensemaking task, compared to not receiving any cues? 2. Related Work The integration of GenAI into search represents a shift from queryâ document matching to generating synthesized, context-aware responses (Trippas and Culpepper, 2025; Zhai, 2024). Generative Information Retrieval systems go beyond ranking results to summarizing, comparing, and explaining information, enabling conversational and adaptive user engagement. GenAI search tools promise efficiency and interactivity by generating synthesized, context-aware responses, and supporting conversational and adaptive user engagement (Trippas and Culpepper, 2025; Zhai, 2024). However, the fluent and confident nature of their responses can obscure inaccuracies and biases, which can cause overtrust and misinformation (Amer and Elboghdadly, 2024; Kaiser et al., 2025). Over-reliance on GenAI tools can cause cognitive offloading, reducing usersâ active engagement and critical evaluation of information (Risko and Gilbert, 2016; Fan et al., 2025; Lee et al., 2025). Consequently, GenAI interactions have been found to reinforce existing beliefs (Sharma et al., 2024) and marginalize alternative viewpoints (Amer and Elboghdadly, 2024). Given these concerns, the design of GenAI systems imposes metacognitive demands on users (Tankelevitch et al., 2024) and effective interaction with GenAI systems requires metacognitive engagement (Flavell, 1979; Winne, 2017). For GenAI-assisted information seeking, effective evaluation of AI responses requires users to accurately judge their topic knowledge and adapt prompting strategies as needed (Tankelevitch et al., 2024). However, this can be challenging for users who have misplaced confidence in their knowledge of a topic or prompting abilities (Singh et al., 2025a; Tankelevitch et al., 2024). Recent work shows metacognitive cues can significantly support information-seeking and sensemaking by promoting reflection, broadening exploration, and deepening inquiry (Singh et al., 2025a). Metacognitive cues draw on the concept of metacognitive prompting, an established instructional approach that has been found to support metacognition in educational contexts (Bannert and Mengelkamp, 2013; Lin and Lehman, 1999) and improve information seeking with traditional search engines (Hwang and Kuo, 2011; Stadtler and Bromme, 2008). Metacognitive cues are designed to direct peopleâs attention toward their own thought processes and toward understanding the activities in which they are engaged (Lin, 2001). Such cues are typically framed as thoughtful questions that are intended to support peopleâs monitoring and control of their information processing by inducing metacognitive and regulative activities, such as orientation, goal specification, planning, monitoring, and control as well as evaluation strategies (Bannert and Mengelkamp, 2013). Their use has been shown to significantly enhance the effectiveness of information seeking in traditional search processes (Zhou and Lam, 2019). More recently, emerging GenAI tools have begun to incorporate reflective scaffolds (Duelen et al., 2024; Gmeiner et al., 2025) and Socratic dialogue mechanisms (Favero et al., 2024) to increase usersâ awareness of their reasoning processes, promote critical engagement with information, and support the detection of misinformation. Building on this body of work, the present study investigates the automatic generation of metacognitive cues within MetaCues, a novel GenAI-based information seeking tool, and examines their effects on usersâ search behaviors during exploratory sensemaking tasks. 3. MetaCues System Design MetaCues is designed for GenAI-based information seeking with metacognitive guidance, and is served as a web application (Figure 1). The application acts as a learning companion that not only provides answers to user queries but also guides the sensemaking process through metacognitive cues that are generated based on userâs chat history, notes and click-stream data. Chat Interactions and Search The conversational AI interface uses OpenAI GPT-4o model with web search capabilities. Temperature is set to 0.8 to balance focus and diversity. Search context country is set to the U.S., and search context size is kept low to minimize response time. The chat uses an instructional LLM prompt111Supplementary material with LLM prompts and Cue Messages that defines the role of AI as a teaching assistant for providing comprehensive academic information. We prompt GPT to: (i) provide responses with clarity, conciseness, and minimal jargon, (i) mandatorily use information from web search with 5+ sources and include citations, (i) frame the response at a technical level of a Bachelorâs degree student, and (iv) structure the response in alignment with major themes in a Markdown format with headings, lists, and emphasis. We also prompt the model to respond with âSorry I canât help you with thatâ for off-topic queries to promote safety. Additionally, responses include visual link cards at the end to provide a compact summary of sources provided in the response. Cue Generation Process The cues implemented in MetaCues were informed by prior work by Singh et al. (2025a) on metacognitive support in GenAI-based search. Building on insights regarding the metacognitive demands of GenAI tools (Tankelevitch et al., 2024), GenAI-assisted search (Sharma et al., 2024; Narayanan Venkit et al., 2025), and search as a learning process (Rieh et al., 2016), Singh et al. initially proposed five types of metacognitive cues: Orienting, Monitoring, Comprehension, Broadening Perspectives, and Consolidation. Each cue was delivered according to predefined criteria based on observable user behavior. Following data collection, their study identified measurable indicators of critical thinking, called Persistent Inquiry, Independent Thinking, and Source Engagement, and recommended tailoring cues to support these behaviors. Guided by these insights, we adopted four cue types directly from Singh et al.âOrienting, Monitoring, Broadening Perspectives (BP), and Consolidationâand introduced three additional types of cues aligned with the identified indicators of critical thinking: Source Engagement (SE), Persistent Inquiry (PI), and Independent Thinking (IT). The Orienting and Monitoring cues, which help establish evaluative criteria for GenAI responses and prompt comparison with prior knowledge, respectively, were delivered at fixed intervals. The remaining cues were dynamically triggered based on user interactions. These cues, described in detail below, serve the following purposes: The PI cue encourages follow-up questions in pursuit of depth of understanding, the SE cue promotes active engagement with sources cited in the AI responses, the IT cue stimulates reflection and synthesis through note-taking, and the BP cue encourages consideration of unexplored perspectives. The SE, PI, and IT cues are instantiated in two variants: a regular variant that encourages under-exhibited desirable behaviors (e.g., âAre there parts of the AI response for which you need more details or evidence Consider going through the linked sourceâŠâ) and a reinforcement variant that acknowledges and strengthens desirable behaviors that are already demonstrated (e.g., âGreat job engaging with the sources! This is helpful for going beyond surface level understanding.â). This design choice was informed by feedback from pilot studies, in which participants expressed a preference for cues that recognized ongoing effective behaviors rather than redundantly prompting actions they believed they were already performing. While we prompt GPT-5footnote 1 to determine which of these variants to deliver, the actual cue messages are predefinedfootnote 1, for consistency and preserving the integrity of metacognitive supportâguiding usersâ thinking and search behavior without offering explicit search recommendations (Bannert and Mengelkamp, 2013). We explored dynamically generating cue messages via LLM prompting, but this approach produced noisy outputs that did not consistently meet established criteria for effective metacognitive scaffolding. Future work may investigate more advanced prompting strategies to support reliable dynamic cue generation. MetaCues currently delivers cues at predetermined intervals in a fixed sequence: an Orienting cue at session start, a Monitoring cue after the first query, followed by SE â IT â PI â BP. These four dynamic cues cycle in this order until the session ends. The initial SE cue is delivered 3-minutes after the session begins, and subsequent cues are triggered at 2.5-minute intervals. Future work may explore more adaptive cue scheduling strategies, including dynamically adjusting cue timing and gradually fading cues as users internalize desirable behaviors. The regular variant of the SE cue is sent if any AI response containing sources has zero source clicks, else its reinforcement variant is sent. If there are no sources in any of the responses so far, a special message is sent to encourage engagement with sources when they do appear. For the PI cue, MetaCues identifies if a user is asking relevant follow-up questions by prompting GPT with the chat history and search topic along with positive and negative examples of relevant follow-up queries. The regular variant is sent if it is determined that the user has not asked any relevant follow-up questions thus far, otherwise the reinforcement variant is sent. For the IT cue, MetaCues prompts GPT to compare the userâs notes with AI responses and content scraped from the sources in these responses, to determine if the notes contain novel viewpoints, such as questions the user may have about the information obtained from searching. If no novel viewpoints are found, the regular variant is sent, else the reinforcement variant is sent. If notes are empty, a special message is displayed to encourage the user to take notes and reflect on their prior knowledge and note any unanswered questions. Finally, the BP cue, which promotes exploration of overlooked perspectives, does not include a reinforcement variant, as achieving comprehensive exploration within the brief study duration is unlikely. Once a cue is triggered, it is queued to be displayed in an activity-aware process. MetaCues waits for natural pauses in user activity (3-second idle), and displays a new cue only when the interface is visible to the user, and there has been some recent user activity within the last 5 minutes. In case such an opportunity is not found, the cue is shown 60 seconds after generation. This approach minimizes distraction while ensuring that the cues are not missed. When a new cue is displayed, a pulsing glow around the cue icon draws attention. The pulsing effect stops when the user acknowledges the cue by clicking the thumbs-up button next to it. 4. Study Design We conducted an online between-subjects factorial experiment comparing (MetaCues) against a (Baseline) tool without metacognitive cues, across two information-seeking tasks. The Baseline tool consists of the chat box and notepad with identical functionality as MetaCues, but does not include the Cues box. The two search topics were selected to elicit participant interestâalbeit at varying levelsâbased on four primary criteria. Specifically, the topics needed to be: (i) timely, in order to foster authentic motivation to learn and ensure a baseline level of participant familiarity; (i) open-ended, to encourage the exploration of diverse perspectives; (i) readily understandable, to ensure participant engagement; and (iv) conducive to informed judgment, necessitating the consideration of multiple viewpoints. Additionally, the first topic was intentionally chosen to be more controversial than the second, enabling an evaluation of the impact of MetaCues across topics with differing levels of controversy. Accordingly, we selected the following two topics: âą Social Media: âConsidering how different social media platforms impact mental health in teenagers, should social media use be banned for individuals below the age of 16?â âą Music: âGiven the cognitive and physiological effects of listening to music while studying or test-taking, should students be allowed to listen to music during school exams?â Participants were randomly assigned to one of four topic-condition groups: Baseline-Social Media, MetaCues-Social Media, BaselineâMusic, or MetaCues-Music. Study Procedure After providing consent, participants answered demographic and LLM usage questions. Next, they were introduced to their assigned tool through a brief tutorial, and prompted to conduct research on the assigned topic while taking notes as if to prepare for writing an essay, using only the assigned tool. Participants were encouraged to gather evidence from reliable sources, understand the topic from multiple perspectives, and explore sources linked in the AI responses. Before starting the task, participants rated their familiarity with and interest in the topic on a 5-point Likert scale and were informed they had up to 25 minutes to complete the task. A timer was displayed in the chat panel, and participants could end the session early, otherwise it ended automatically after 25 minutes. Data on Attitudinal Judgments After completing the task, participants rated their attitudinal judgments regarding the topic and their level of confidence in their judgment, each on a 5-point Likert scale. For the social media topic, participants rated their agreement with banning its use in schools; for the music topic, they rated their agreement with allowing its use during school exams. Data on Search Behaviors We computed the following behavioral measures from system logs: (1) search duration, (2) time spent outside the interface, (3) total typing time, (4) number of queries, (5) average words per query, (6) number of sources clicked, (7) click-through rate, and (8) query divergence. Measure (2) estimates the time participants spent on the linked sources, which opened in a new tab. Measure (7) captures the ratio of unique sources clicked to the number of all unique sources linked in all AI responses during the session. Measure (8) quantifies how semantically divergent a userâs queries are within a given topic: lower divergence indicates more focused querying, whereas higher divergence reflects broader conceptual exploration. To measure query divergence, we first trained vectorized embedding representations separately for each topic. Each query was represented as a 384-dimensional embedding generated by the all-MiniLM-L6-v2 sentence-transformer model (Reimers and Gurevych, 2019). Further, queries were L2-normalized and cosine distance between two embeddings was used to capture semantic dissimilarity. For a user u, we computed a centroid of their queriesâ embeddings: u=(1/nu)ââic_u=(1/n_u) _iv_i. Then, we measured the cosine distance between each query iv_i and its centroid as di=1â(iâ u)/(âiâââuâ)d_i=1-(v_i·c_u)/(\|v_i\|\|c_u\|). Query divergence is the mean of these distances DÂŻu=(1/nu)ââidi D_u=(1/n_u) _id_i. In addition to this data, we evaluated participantsâ learning outcomes through a post-test, and captured artifacts reflecting their engagement during the information-seeking process, including notes, end-of-search summaries, and conversations with the assigned AI. The analysis of this data is reserved for future work. 5. Results Participants Overview Participants were recruited via Prolific222https://w.prolific.com and compensated at $15 per hour. Eligible participants were U.S.-based individuals currently enrolled in or holding a college degree. The study lasted approximately one hour. 175 participants completed the study, of which 146 passed all attention checks and constituted the study sample. Ages ranged from 18â59 (median: 18â24); 73 identified as male, 67 as female, 6 as non-binary. Kruskal-Wallis H tests revealed no significant differences between the four topic-condition groups in LLM use frequency (Hâ(3)=1.08H(3)=1.08, p=0.78p=0.78) and frequency of using LLMs for search purposes (Hâ(3)=0.152H(3)=0.152, p=0.99p=0.99). Topic Familiarity and Interest Mann-Whitney U tests revealed a significant difference in participantsâ perceived familiarity with the search topics (U=3459.50U=3459.50, p=0.001p=0.001), with participants reporting more familiarity with the social media topic (M=2.85M=2.85, SâD=0.88SD=0.88) than the music topic (M=2.36M=2.36, SâD=1.08SD=1.08). There was no significant difference in their level of interest between the social media (M=3.51M=3.51, SâD=1.04SD=1.04) and music (M=3.43M=3.43, SâD=1.18SD=1.18) topics (U=2737.00U=2737.00, p=0.770p=0.770). Effects on Search Behaviors & Confidence in Attitudinal Judgments Measure Social Media Music Baseline: M (SD) MetaCues: M (SD) Baseline: M (SD) MetaCues: M (SD) Search duration (s) 1299.27 (386.48) 1241.93 (365.43) 1186.39 (400.76) 1284.83 (405.38) Time spent outside interface (s) 150.79 (150.35) 370.01 (1103.84) 206.14 (217.77) 248.94 (303.02) Total typing time (s) 353.09 (213.67) 307.67 (222.27) 272.03 (143.03) 285.81 (173.26) Number of queries 7.39 (6.76) 7.35 (5.67) 5.95 (3.75) 9.25 (7.71) Average words per query 15.76 (6.60) 21.40 (19.45) 16.12 (12.83) 14.65 (6.62) Average words per AI response 280.06 (112.87) 264.24 (80.97) 210.12 (76.79) 218.60 (54.62) Number of sources clicked 2.61 (2.94) 2.76 (3.16) 2.49 (2.71) 2.78 (2.52) Click-through rate 0.21 (0.24) 0.24 (0.26) 0.23 (0.28) 0.27 (0.27) Query divergence 0.23 (0.17) 0.26 (0.15) 0.28 (0.14) 0.34 (0.15) Table 1. Descriptive statistics for each search behavior measure. Table 1 reports the mean and standard deviation for each search behavior measure across conditions and topics. To assess the effects of searching with versus without cues, we fit Generalized Linear Models (GLMs) with appropriate link functions as the data did not meet normality assumptions. Condition, topic, and their interaction were used as fixed effects. We used a negative binomial (log link) for count data, gamma (log link) for time-based measures and click-through rate, and Gaussian (identity link) for average words per query. For query divergence, we ran separate MannâWhitney U tests per topic as we trained distinct embedding models for each topic. Lastly, for attitudinal judgments, we conducted a two-way ANOVA with topic, condition, and their interaction as the independent variables. We now report the results of the statistical analyses. For the music topic, we found that usersâ query divergence in the MetaCues condition (M=0.34M=0.34, SâD=0.15SD=0.15) was significantly higher (U=511.5U=511.5, p=0.045p=0.045, d=0.40d=0.40) than Baseline (M=0.28M=0.28, SâD=0.14SD=0.14). For the social media topic, query divergence for MetaCues (M=0.26M=0.26, SâD=0.15SD=0.15) was slightly higher than Baseline (M=0.23M=0.23, SâD=0.17SD=0.17), but the difference was not statistically significant (U=610.5U=610.5, p=0.27p=0.27, d=0.16d=0.16). Additionally, to explore how queries differed semantically across conditions, we visualized their embeddings using UMAP (McInnes et al., 2018), projected into a two-dimensional latent space that preserved relative semantic distances. Figure 2 shows the Baseline and MetaCues groups side by side for both topics. For the music topic, the MetaCues group appears more widely dispersed, extending into regions that are sparse for the Baseline group, suggesting broader and more exploratory query formulation. For social media, the two groups show similar overall density, though the MetaCues group exhibits a slightly broader outer contour. These visualizations are consistent with the results of the statistical analyses reported above. For both time spent outside the interface (likely on linked sources) and click-through rate, participants in the MetaCues condition showed higher means across topics (see Table 1), though these differences were not significant. Turning to the remaining search behavior measures, mean time spent outside the interface and click-through rates were higher in the MetaCues condition, but neither difference was statistically significant. For the number of queries, the MetaCues condition showed a notably higher mean for the music topic, but this effect was also non-significant. No significant main effects of condition, topic, or their interaction were observed for overall search duration, total typing time, average words per query, and number of sources clicked. Detailed statistical analyses results can be found herefootnote 1. Regarding confidence in attitudinal judgments, a two-way ANOVA revealed a significant main effect of condition, (Fâ(1,142)=4.53F(1,142)=4.53, p=0.035p=0.035), indicating that participants in the MetaCues condition reported higher confidence than those in Baseline. There was no significant main effect of topic (Fâ(1,142)=0.03F(1,142)=0.03, p=0.857p=0.857), and no significant interaction between condition and topic (Fâ(1,142)=0.40F(1,142)=0.40, p=0.530p=0.530). Figure 2. UMAP visualizations of query embeddings for music (left) and social media (right) search topics. 6. Discussion This study demonstrates the feasibility of automatically generating metacognitive cues to guide GenAI-based information seeking, verifying the findings from Singh et al. (Singh et al., 2025a). However, the effects of MetaCues varied by search topic. The finding that MetaCues led to significantly greater query divergence for the music topicâon which participants reported lower prior familiarity compared to the social media topic, and which was less controversial in natureâsuggests that MetaCues may be more effective in supporting broad exploration for topics with which users have lower perceived familiarity and less strongly held opinions. In contrast, the social media topic was widely discussed at the time of the study, which may have contributed to participants holding more established or polarized views. As a result, MetaCues may not have exerted as strong an influence on search behavior for this topic. Future work involving qualitative and in-depth analyses of participantsâ search interactions and post-test outcomes is needed to provide further insight into how MetaCues shaped learning across both topics, and how these effects were mediated by participantsâ familiarity with and the controversial nature of the topics. Notably, MetaCues led to significantly higher confidence in the resulting attitudinal judgments compared to the baseline tool, which reflects potentially greater perceived epistemic grounding resulting from deeper engagement with sources, reflection, and inquiry. This suggests that metacognitive cues are effective for aiding the consolidation of understanding necessary for judgment formation (Reyna et al., 2003). However, level of confidence is not reflective of the actual impact on their learning outcomes. Therefore, future work should further examine whether such confidence is well-calibrated, how it evolves over longer-term or higher-stakes tasks, and also conduct qualitative exploration of participant interactions and inquiry patterns. 7. Conclusion We presented MetaCues, an interactive system that automatically generates metacognitive cues to support GenAI-based search. In an online between-subjects study, we found that it promotes more active and diverse inquiry than a baseline system without cues, particularly for topics that are less controversial and with which users are less familiar. Further, it leads to higher confidence in resulting attitudinal judgements. However, the studyâs modest sample size (N=146N=146) limits statistical power; future work with larger samples could further examine how metacognitive cues affect different types of search tasks. As cue generation in MetaCues was timed according to study constraints, subsequent research should explore more adaptive cue delivery and the gradual fading of metacognitive support as users gain experience. Future work should also include qualitative analyses of participant interactions and inquiry patterns. References E. Amer and T. Elboghdadly (2024) The end of the search engine era and the rise of generative ai: a paradigm shift in information retrieval. In 2024 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), p. 374â379. Cited by: §1, §2. M. Bannert and C. Mengelkamp (2013) Scaffolding hypermedia learning through metacognitive prompts. In International handbook of metacognition and learning technologies, p. 171â186. Cited by: §2, §3. A. Duelen, I. Jennes, and W. Van den Broeck (2024) Socratic AI Against Disinformation: Improving Critical Thinking to Recognize Disinformation Using Socratic AI. In Proceedings of the 2024 ACM International Conference on Interactive Media Experiences, IMX â24, p. 375â381. External Links: Document, Link, ISSN 979-8-4007-0503-8 Cited by: §2. Y. Fan, L. Tang, H. Le, K. Shen, S. Tan, Y. Zhao, Y. Shen, X. Li, and D. GaĆĄeviÄ (2025) Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance. 56 (2), p. 489â530. External Links: 2412.09315, ISSN 0007-1013, 1467-8535, Document, Link Cited by: §2. L. Favero, J. A. PĂ©rez-Ortiz, T. KĂ€ser, and N. Oliver (2024) External Links: 2409.05511, Document, Link Cited by: §2. J. H. Flavell (1979) Metacognition and cognitive monitoring: a new area of cognitiveâdevelopmental inquiry.. American psychologist 34 (10), p. 906. Cited by: §2. F. Gmeiner, K. Luo, Y. Wang, K. Holstein, and N. Martelaro (2025) Exploring the potential of metacognitive support agents for human-ai co-creation. In Proceedings of the 2025 ACM Designing Interactive Systems Conference, DIS â25, New York, NY, USA, p. 1244â1269. External Links: ISBN 9798400714856, Link, Document Cited by: §2. G. Hwang and F. Kuo (2011) An information-summarising instruction strategy for improving the web-based problem solving abilities of students. Australasian Journal of Educational Technology 27 (2). Cited by: §2. C. Kaiser, J. Kaiser, R. Schallner, and S. Schneider (2025) A new era of online search? a large-scale study of user behavior and personal preferences during practical search tasks with generative ai versus traditional search engines. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, p. 1â7. Cited by: §2. H. (. Lee, A. Sarkar, L. Tankelevitch, I. Drosos, S. Rintel, R. Banks, and N. Wilson (2025) The impact of generative ai on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §2. X. Lin and J. D. Lehman (1999) Supporting learning of variable control in a computer-based biology environment: effects of prompting college students to reflect on their own thinking. Journal of Research in Science Teaching: The Official Journal of the National Association for Research in Science Teaching 36 (7), p. 837â858. Cited by: §2. X. Lin (2001) Designing metacognitive activities. Educational technology research and development 49 (2), p. 23â40. Cited by: §2. L. McInnes, J. Healy, and J. Melville (2018) UMAP: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. External Links: Link Cited by: §5. P. Narayanan Venkit, P. Laban, Y. Zhou, Y. Mao, and C. Wu (2025) Search engines in the ai era: a qualitative understanding to the false promise of factual and verifiable source-cited responses in llm-based search. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, p. 1325â1340. Cited by: §1, §3. N. Reimers and I. Gurevych (2019) Sentence-bert: sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), Hong Kong, China. External Links: Link Cited by: §4. V. F. Reyna, F. J. Lloyd, and C. J. Brainerd (2003) Memory, development, and rationality: an integrative theory of judgment and decision making. Emerging perspectives on judgment and decision research, p. 201â245. Cited by: §6. S. Y. Rieh, K. Collins-Thompson, P. Hansen, and H. Lee (2016) Towards searching as a learning process: a review of current perspectives and future directions. Journal of Information Science 42 (1), p. 19â34. Cited by: §3. E. F. Risko and S. J. Gilbert (2016) Cognitive offloading. Trends in cognitive sciences 20 (9), p. 676â688. Cited by: §1, §2. N. Sharma, Q. V. Liao, and Z. Xiao (2024) Generative echo chamber? effect of llm-powered search systems on diverse information seeking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â17. Cited by: §2, §3. A. Singh, Z. Guan, and S. Y. Rieh (2025a) Enhancing critical thinking in generative ai search with metacognitive prompts. Proceedings of the Association for Information Science and Technology 62 (1), p. 672â684. Cited by: §1, §2, §2, §3, §6. A. Singh, K. Taneja, Z. Guan, and A. Ghosh (2025b) Protecting human cognition in the age of ai. arXiv preprint arXiv:2502.12447. Cited by: §1. I. Solaiman, Z. Talat, W. Agnew, L. Ahmad, D. Baker, S. L. Blodgett, C. Chen, H. DaumĂ© I, J. Dodge, I. Duan, et al. (2023) Evaluating the social impact of generative ai systems in systems and society. arXiv preprint arXiv:2306.05949. Cited by: §1. M. Stadtler and R. Bromme (2008) Effects of the metacognitive computer-tool met. a. ware on the web search of laypersons. Computers in Human Behavior 24 (3), p. 716â737. Cited by: §2. L. Tankelevitch, V. Kewenig, A. Simkute, A. E. Scott, A. Sarkar, A. Sellen, and S. Rintel (2024) The metacognitive demands and opportunities of generative ai. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â24. Cited by: §1, §2, §3. J. R. Trippas and J. S. Culpepper (2025) Report from the Fourth Strategic Workshop on Information Retrieval in Lorne (SWIRL 2025). Cited by: §2. D. VedejovĂĄ and V. ÄavojovĂĄ (2022) Confirmation bias in information search, interpretation, and memory recall: evidence from reasoning about four controversial topics. Thinking & Reasoning 28 (1), p. 1â28. Cited by: §1. P. H. Winne (2017) Cognition and metacognition within self-regulated learning. In Handbook of self-regulation of learning and performance, p. 36â48. Cited by: §1, §2. C. Zhai (2024) Large Language Models and Future of Information Retrieval: Opportunities and Challenges. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, p. 481â490. External Links: Document, Link, ISSN 979-8-4007-0431-4 Cited by: §2. M. Zhou and K. K. L. Lam (2019) Metacognitive scaffolding for online information search in k-12 and higher education settings: a systematic review. Educational technology research and development 67 (6), p. 1353â1384. Cited by: §2. T. Zhou and S. Li (2024) Understanding user switch of information seeking: from search engines to generative ai. Journal of librarianship and information science, p. 09610006241244800. Cited by: §1.