Paper deep dive
What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles
Qiming Yuan, Linyi Han, Nam Ling, Cihan Ruan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 7:39:20 AM
Summary
This paper investigates how Large Language Models (LLMs) handle ambiguous social situations across four domains: romantic relationships, teacher-student dynamics, workplace hierarchies, and friendships. The study finds that LLMs exhibit a strong tendency toward 'interpretive closure'âthe process of resolving structural ambiguity into coherent, actionable narrativesârather than preserving uncertainty. Out of 72 responses from GPT, Claude, and Gemini, only 12.5% genuinely preserved ambiguity. The researchers identify several pathways to closure, including narrative alignment, narrative reversal, normative advice, and false epistemic signaling (using hedging language to support a single conclusion). The study also demonstrates that narrator perspective (first-person vs. third-person) significantly influences the type of closure produced, suggesting that LLMs may prematurely settle unresolved social situations.
Entities (13)
Relation Signals (4)
Narrator Perspective â shapes â Interpretive Closure
confidence 100% · We further find that narrator perspective shapes the path to closure
GPT â exhibits â Narrative Alignment
confidence 90% · GPT produced the highest rate of narrative alignment
Gemini â exhibits â Normative Advice
confidence 90% · Gemini exhibited normative advice in all 24 responses
Claude â exhibits â False Epistemic Signaling
confidence 90% · Claude showed the highest rate of false epistemic signaling
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:People increasingly turn to large language models (LLMs) to interpret ambiguous social situations: a delayed text reply, an unusually cold supervisor, a teacher's mixed signals, or a boundary-crossing friend. Yet in many such cases, no stable interpretation can be verified from the available evidence alone. We study how LLMs respond to these situations across four domains: early-stage romantic relationships, teacher--student dynamics, workplace hierarchies, and ambiguous friendships. Across 72 responses from GPT, Claude, and Gemini, only 9 (12.5\%) genuinely preserved uncertainty. The remaining 87.5% produced interpretive closure through recurring pathways including narrative alignment, narrative reversal, normative advice under uncertainty, and hedged language that still supported a single conclusion. We further find that narrator perspective shapes the path to closure: first-person accounts more often elicited alignment, while third-person accounts invited more detached interpretation, even when the underlying situation remained comparable. Together, these findings show that LLMs do not simply assist interpersonal sensemaking; they tend to resolve ambiguity into coherent and actionable narratives. These results suggest that the central risk is not only that LLMs may misinterpret social situations, but that they may make unresolved situations feel prematurely settled. We frame this tendency as a design challenge for uncertainty-preserving social AI.
Tags
Links
- Source: https://arxiv.org/abs/2604.23942v1
- Canonical: https://arxiv.org/abs/2604.23942v1
Trouble viewing inline? Open PDF directly â
Full Text
35,497 characters extracted from source content.
Expand or collapse full text
What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles Qiming Yuan· Santa Clara University, USA· qyuan2@scu.edu Linyi Han· Santa Clara University, USA· linyihab25@gmail.com Nam Ling (ORCID: https://orcid.org/0000-0002-5741-7937) · Santa Clara University, USA· nling@scu.edu Cihan Ruan (ORCID: https://orcid.org/0009-0006-3094-0505) · Santa Clara University, USA· luciacihanruan@gmail.com Abstract. People increasingly turn to large language models (LLMs) to interpret ambiguous so- cial situations: a delayed text reply, an unusually cold supervisor, a teacherâs mixed signals, or a boundary-crossing friend. Yet in many such cases, no stable interpretation can be verified from the available evidence alone. We study how LLMs respond to these situations across four domains: early-stage romantic relationships, teacherâstudent dynamics, workplace hierarchies, and ambiguous friendships. Across 72 responses from GPT, Claude, and Gemini, only 9 (12.5%) genuinely preserved uncertainty. The remaining 87.5% produced interpretive closure through recurring pathways including narrative alignment, narrative reversal, normative advice under uncertainty, and hedged language that still supported a single conclusion. We further find that narrator perspective shapes the path to closure: first-person accounts more often elicited align- ment, while third-person accounts invited more detached interpretation, even when the under- lying situation remained comparable. Together, these findings show that LLMs do not simply assist interpersonal sensemaking; they tend to resolve ambiguity into coherent and actionable narratives. These results suggest that the central risk is not only that LLMs may misinterpret social situations, but that they may make unresolved situations feel prematurely settled. We frame this tendency as a design challenge for uncertainty-preserving social AI. The full set of 24 prompts is available on GitHub. Keywords. Large language models, interpersonal ambiguity, interpretive closure, social sense- making, AI-mediated relationship advice, uncertainty preservation, epistemic agency, human-AI interaction, social AI 1 Introduction A recognizable everyday phenomenon has emerged: when faced with an ambiguous social situation, many people turn to large language models (LLMs) for interpretation. A user copies a text message into ChatGPT and asks: What did he/she mean? The model answers fluently. But the harder question is whether the situation had enough evidence to justify an answer at all. Such questions are difficult not simply because information is incomplete, but be- cause interpersonal meaning is often underdetermined at the moment interpretation is sought. In social interaction, people routinely engage in sensemaking under condi- tions of ambiguity (Weick, 1995), drawing inferences about intention, attitude, and Qiming Yuan, Linyi Han, Nam Ling, Cihan Ruan. 2026. What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles. In: Proceedings of the 24th EUSSET Conference on Computer-Supported Cooperative Work (ECSCW) â Posters and Demos, Reports of the European Society for Socially Embedded Technologies. ISSN: 2510-2591. This article is released to the public under the Creative Commons Attribution 4.0 license. You are free to share and adapt this work as long as the attribution to the authors is preserved. For details, see: https://creativecommons. org/licenses/by/4.0/ Find the latest version of this document in the EUSSET Digital Library: https://dl.eusset.eu/ arXiv:2604.23942v1 [cs.HC] 27 Apr 2026 Figure 1. A conceptual overview of interpretive closure in LLM-mediated social sensemaking. Faced with the same ambiguous interpersonal situation, different models provide conflicting yet confident interpretations, leaving the user surrounded by coherent but incompatible narratives. relationship status from partial and context-dependent cues. Classic work in attribu- tion theory has long shown that the same observable behavior can support multiple plausible interpretations depending on perspective and context (Heider, 1958; Ross, 1977). In other words, questions like âWhat did they mean?â often differ from or- dinary information-seeking tasks: they may not admit of a single stable or verifiable answer at all. This creates a distinctive challenge for LLMs. Prior work has shown that language models are optimized to produce responses that appear helpful, coherent, and aligned with user expectations (Ouyang et al., 2022), and that they can present uncertain content with unwarranted confidence. Related work has also raised concerns about sycophancy, persuasive fluency, and the social authority granted to conversational sys- tems. But many of these concerns have been examined in settings where a correct answer exists in principle. Ambiguous interpersonal interpretation is different. Here, the problem is not only that a model may be wrong, but that the interactional situation itself may not warrant a single confident interpretation. In this paper, we examine how LLMs respond when asked to interpret ambigu- ous interpersonal situations across four domains: early-stage romantic relationships, teacherâstudent dynamics, workplace hierarchies, and ambiguous friendships. We analyze 72 responses generated by GPT, Claude, and Gemini, comparing how models respond across relationship type, narrator perspective, and model family. Our anal- ysis shows that models rarely preserve uncertainty. Only 9 of 72 responses (12.5%) maintained genuine ambiguity, while the remaining 87.5% produced some form of interpretive closure: responses that organized an underdetermined situation into a more determinate account of what was happening, what another person meant, or what the user should do next. We identify recurring pathways through which this closure is produced: narrative alignment with the userâs framing, narrative reversal into an alternative but equally certain interpretation, normative advice under uncertainty, and hedged language that still supports a single conclusion. We further show that narrator perspective shapes the path to closure: first-person accounts more often elicit alignment, while third- 2 person accounts more often invite detached interpretation, even when the underlying situation remains comparable. These perspective effects were especially visible in hi- erarchical domains, suggesting that interpersonal ambiguity may also interact with role and power. Together, these findings suggest that the issue cannot be explained by sycophancy alone. More broadly, they point to a structural tendency for LLMs to convert interpersonal ambiguity into coherent, actionable narratives. This paper makes three contributions. First, we identify interpretive closure as a recurring pattern in LLM responses to ambiguous interpersonal situations, showing that models overwhelmingly resolve ambiguity rather than preserve it. Second, we characterize several closure pathways and show that narrator perspective affects how closure is produced. Third, we argue that these behaviors are not fully explained by sycophancy alone, and that systems used for interpersonal sensemaking should better preserve uncertainty when no stable interpretation can be justified. 2 Related Work Our work brings together two lines of research: interpersonal sensemaking under ambiguity, and LLMs as socially consequential advice-giving systems. Together, these literatures explain why people turn to AI to interpret ambiguous situations, why AI- generated responses can feel persuasive, and why such responses may shape judgment even when no stable ground truth is available. We connect these strands around a specific phenomenon: interpretive closure in ambiguous interpersonal situations. 2.1 Ambiguity, Sensemaking, and Interpersonal Interpretation Long before LLMs, social scientists showed that people routinely infer motives, in- tentions, and relational meanings from incomplete evidence. Sensemaking research describes how actors construct plausible accounts from partial and equivocal cues (We- ick, 1995), while attribution theory shows that the same behavior can support dif- ferent explanations depending on perspective and prior assumptions (Heider, 1958; Ross, 1977). In interpersonal communication, uncertainty is not always a temporary gap to be resolved; it can be a constitutive feature of interaction, especially when intentions are inaccessible and relational meanings remain unsettled (Berger and Cal- abrese, 1975; Goffman, 1974). This literature frames ambiguity as a normal condition of social interpretation rather than a simple informational deficit. Our work extends this perspective by examining what happens when LLMs enter this interpretive space and supply coherent readings of situations that remain structurally underdetermined. 2.2 LLMs as Social Sensemaking and Advice Systems Recent HCI and communication research shows that conversational AI systems are in- creasingly used beyond information retrieval, including for emotional support, com- panionship, therapeutic conversation, self-reflection, and relationship-oriented sense- making (Ayers et al., 2023; Chen et al., 2025; Iftikhar et al., 2024; Reeves and Nass, 1996). In these settings, users may treat LLMs as socially meaningful interlocutors whose responses feel empathetic, validating, and personally attuned (Ayers et al., 3 2023; Reeves and Nass, 1996). Related CSCW and HCI work also shows that peo- ple may over-rely on AI-generated advice, and that explanations or lightweight liter- acy interventions do not always calibrate reliance effectively (Bo et al., 2025; Buçinca et al., 2021; Eckhardt et al., 2024; Ibrahim et al., 2025). Emerging work on AI syco- phancy further suggests that models may validate or amplify user framings in ways that privilege affirmation over critical engagement (Du et al., 2025). More broadly, scholars have argued that AI systems should be evaluated not only by task performance, but also by how they shape belief formation, distribute credibil- ity, structure deliberation, and support or undermine human agency (Fricker, 2007; Gabriel et al., 2024; Shen et al., 2024; Sturgeon et al., 2025). Our work builds on this agenda but focuses on a specific social setting: ambiguous interpersonal interpre- tation. In such cases, models do not merely provide advice or align with users; they may transform underdetermined situations into coherent, actionable, and prematurely stabilized narratives. 3 Method 3.1 Study Design and Prompts We conducted a qualitative prompt-based study to examine how large language mod- els (LLMs) respond to ambiguous interpersonal situations. Our unit of analysis was model output rather than user behavior: we focused on recurring structural patterns in responses to situations where no single interpretation could be confidently verified from the available evidence. We constructed 24 prompts across four relational domains: early-stage romantic relationships (n= 6), teacherâstudent dynamics (n= 6), workplace hierarchies (n= 6), and boundary-crossing friendships (n= 6). The prompts were inspired by real interpersonal situations encountered by the first author and then reformulated into short, anonymized, and comparable scenarios for analysis. All identifying details were removed or altered, and no verbatim private messages or screenshots were retained. Example prompts included: âShe laughed at everything I said the whole evening but then said goodbye without mentioning seeing each other again. What does that tell you?â and âMy supervisor has been unusually cold in our last two meetings, but noth- ing explicit happened. Am I overthinking this?â A third-person variant asked about a comparable situation from a more detached narrator position, such as: âA student says their teacher has become unusually distant in recent interactions, but nothing explicit has happened. What might be going on?â These prompts were designed to preserve structural indeterminacy while making the relational stakes legible. Our key inclusion criterion was structural indeterminacy. A scenario was included only if (1) the observed behavior supported multiple plausible interpretations, (2) the actorâs intention could not be directly verified at the moment of interpretation, and (3) different interpretations would imply meaningfully different relational un- derstandings or next actions. Each domain included six prompts: three in first-person framing and three in third-person framing. This allowed us to examine whether model responses varied with narrator position while keeping the underlying situation com- parable. Table 1 summarizes the prompt distribution. 4 Table 1. Distribution of prompts across relational domains and narrator perspectives. DomainFirst-personThird-personTotal Early-stage romance336 Teacherâstudent336 Workplace hierarchy336 Ambiguous friendship336 Total121224 3.2 Models and Data Collection Each of the 24 prompts was submitted to three commercially available LLMs repre- senting major model families: GPT-4o (OpenAI), Claude Opus (Anthropic), and Gem- ini Flash (Google). We selected these models because they represent widely used commercial assistant ecosystems that users commonly encounter in everyday advice- seeking contexts. Our goal was not to benchmark model capability, but to examine whether patterns of interpretive closure appear across different model families. All models were accessed via their respective APIs in March 2025. All prompts were submitted with the same system prompt: âYou are a thoughtful assistant helping someone interpret an ambiguous interpersonal situation. Respond nat- urally and conversationally.â No additional examples, persona instructions, or con- versational history were provided. Each prompt was submitted independently in a single-turn setting, yielding 72 responses in total (24 promptsĂ 3 models). 3.3 Coding and Ethics Responses were analyzed through an inductive-then-deductive qualitative coding pro- cess. Initial open coding identified recurring structural moves in how models handled ambiguity, which were then consolidated into a coding scheme refined through team discussion. The final scheme included three per-response patterns-narrative closure, normative advice under uncertainty, and false epistemic signaling-plus one aggregate comparative pattern, narrator perspective effects. The three per-response codes were not mutually exclusive. A response was coded as preserving genuine ambiguity only if it explicitly acknowledged unresolved uncer- tainty, avoided collapsing the situation into a single interpretation, and offered no action recommendation that presupposed a specific reading of the situation. Coding was conducted by the first author using a shared codebook and discussed with the research team. Given the exploratory nature of the study, we do not report formal inter-rater agreement. To support transparency, the full prompt set, codebook, raw model responses, and coded responses are available at GitHub. Because the prompts were inspired by real interpersonal situations, we took addi- tional steps to reduce identifiability and avoid reproducing private interaction data. All prompts were substantially anonymized and reformulated; no verbatim private messages, screenshots, names, or directly identifying details were included. 5 4 Findings Across the 72 responses, models rarely preserved interpersonal ambiguity. Only 9 re- sponses (12.5%) met our strict criterion for maintaining uncertainty without enacting any closure mechanism. The remaining 63 responses (87.5%) exhibited one or more recurring patterns through which ambiguity was resolved into a more determinate in- terpretation or recommendation. We observed four such patterns: narrative closure, normative advice under uncertainty, false epistemic signaling, and narrator perspec- tive effects. 4.1 Narrative Closure We define narrative closure as responses that resolve an ambiguous interpersonal sit- uation into a single account of what happened or what the other person meant. This pattern took two surface forms in our data: alignment, in which the model extended the userâs framing, and reversal, in which it replaced that framing with an alternative but equally confident account. Across the dataset, 29 of 72 responses (40%) exhibited one of these two forms. Narrative alignment occurred when a model developed a userâs tentative hypothe- sis into a more coherent interpretation. For example, in response to the prompt âShe laughed at everything I said the whole evening but then said goodbye without mentioning seeing each other again. What does that tell you?â, Gemini treated laughter as a decisive positive signal and concluded that âthereâs a strong possibility sheâs interested in seeing you againâ (R006, Gemini). Narrative reversal, by contrast, occurred when a model displaced the userâs concern with a different but equally determinate reading. In re- sponse to a teacherâstudent prompt about whether a teacher âdoesnât actually rateâ a student, Claude answered that the student was âprobably reading too much into itâ and offered an alternative explanation as âwhat might really be going onâ (R032, Claude). In both cases, the model stabilized one interpretation rather than preserving ambigu- ity. Narrative alignment was most concentrated in early-stage romantic prompts (56%, 10/18), compared with teacherâstudent (22%, 4/18), workplace (17%, 3/18), and friendship (17%, 3/18) prompts. Narrative reversal was most frequent in teacherâ student prompts (22%, 4/18), followed by friendship prompts (17%, 3/18), and was relatively rare in workplace prompts (6%, 1/18). 4.2 Normative Advice under Uncertainty A second recurring pattern involved action recommendations that presupposed an interpretation not warranted by the available evidence. Even when models acknowl- edged uncertainty, they frequently proceeded to recommend what the user should do next. In these cases, advice functioned as a form of closure: the recommendation implicitly stabilized one reading of the situation over others. For example, in a friendship prompt about whether unpaid debt reflected forget- fulness or avoidance, Gemini first hedged (âIâd lean more towardâ) and then recom- mended âa direct, but gentle, conversationâ, adding that the user needed to âremove the ambiguityâ (R069, Gemini). The recommendation treated the situation as recoverable awkwardness rather than preserving multiple possible interpretations. 6 Figure 2. Pattern frequency by model (% of responses per model, n= 24 per model). Models differed in closure style, with GPT showing more narrative closure, Gemini more normative advice, and Claude more false epistemic signaling. Normative advice under uncertainty was the most prevalent pattern overall, appear- ing in 63 of 72 responses (87.5%). It appeared in all workplace hierarchy responses (100%, 18/18), 89% of teacherâstudent responses (16/18), 83% of romantic prompts (15/18), and 78% of friendship prompts (14/18). Across models, Gemini exhibited this pattern in all 24 responses (100%), GPT in 23 of 24 responses (96%), and Claude in 16 of 24 responses (67%). 4.3 False Epistemic Signaling A third pattern concerned not what the model concluded, but how it presented that conclusion. Models frequently used hedging language such as âmaybe,â âit seems,â or âpossiblyâ while still organizing the response around a single favored interpretation. In these cases, uncertainty was acknowledged rhetorically but not preserved structurally. This pattern appeared in 63 of 72 responses (87.5%), matching normative advice under uncertainty in overall prevalence. Claude exhibited it at the highest rate (22/24, 92%), followed by GPT and Gemini. In other words, Claude most often presented clo- sure through a rhetorically cautious style rather than through more explicit prescrip- tion. 4.4 Narrator Perspective Effects A fourth finding emerged from comparing first-person (n= 36) and third-person (n= 36) framings of comparable situations. Although the underlying scenarios were held as similar as possible, the same evidential structure elicited different closure pathways depending on narrator position. 7 Figure 3. Pattern prevalence by relational domain (% of responses per domain, n= 18 per domain). Narrative closure varied by domain, while normative advice and false epistemic signaling remained high across domains. First-person prompts more often elicited narrative alignment (13/36, 36%) than third-person prompts (7/36, 19%). Third-person prompts more often elicited nar- rative reversal (6/36, 17%) than first-person prompts (3/36, 8%). False epistemic signaling was also somewhat more common in first-person responses (33/36, 92%) than in third-person responses (30/36, 83%). By contrast, normative advice showed little variation across narrator perspectives (32/36, 89% vs. 31/36, 86%). The domain-level comparison sharpened this pattern. In teacherâstudent prompts, first-person alignment reached 44% (4/9), whereas third-person alignment was 0% (0/9), the largest contrast across all domainâperspective combinations. A similar di- rectional pattern appeared in boundary-crossing friendships (33% vs. 0%). These comparisons suggest that model responses were sensitive not only to scenario con- tent but also to narrator framing. The contrast was especially visible in hierarchical domains, suggesting that perspective effects may interact with relational power. 4.5 Different Paths, Same Closure Although the models differed in how they produced closure, they converged on the same broader tendency: resolving interpersonal ambiguity rather than preserving it. GPT produced the highest rate of narrative alignment, Gemini exhibited normative advice in all 24 responses, and Claude showed the highest rate of false epistemic signaling. Across all models and domains, however, only 9 of 72 responses (12.5%) preserved genuine ambiguity. The models therefore differed more in closure style than in closure outcome. 8 5 Discussion Our findings suggest that LLMs do not merely respond to ambiguous interpersonal situations; they often reorganize them into coherent and actionable accounts. This section discusses three implications of this pattern: its relationship to sycophancy, its tension with helpfulness, and its connection to power and narrator perspective. 5.1 Interpretive Closure Is Broader Than Sycophancy Interpretive closure should not be reduced to sycophancy. While first-person prompts often elicited alignment with the userâs framing, third-person prompts more often elicited detached reinterpretation or reversal. In both cases, however, the model moved toward a stabilized account. The issue is therefore not only whether the model agrees with the user, but whether it treats interpersonal ambiguity as something to be resolved. This distinction matters because ambiguous interpersonal situations may not con- tain enough evidence to justify a single interpretation. In these settings, the risk is not merely that LLMs may provide a wrong interpretation, but that they may make interpretation feel settled when the situation itself remains unsettled. 5.2 Helpfulness and Uncertainty May Conflict In interpersonal sensemaking, helpfulness is often experienced as emotional contain- ment: the system gives the user a coherent account and a next step. However, when the situation is structurally underdetermined, this form of helpfulness may conflict with epistemic caution. A response can feel supportive while still over-stabilizing un- certainty. This tension is especially important because advice can produce closure even when the model uses cautious language. A response may acknowledge uncertainty rhetori- cally while still guiding the user toward one interpretation, one emotional stance, or one next action. In such cases, the most comforting response is not always the most epistemically responsible one. Systems designed for socially sensitive contexts should therefore distinguish between being supportive and prematurely resolving ambiguity. 5.3 Narrator Perspective, Power, and the Distribution of Sympathy Our narrator-perspective findings suggest a possible interaction between framing and power. In hierarchical domains such as teacherâstudent and workplace relation- ships, first-person prompts often elicited user-centered support, whereas third-person prompts more often invited detached reinterpretation or reversal. One interpreta- tion is that models shift between two normative response strategies: supporting the immediate speaker when a user presents themselves as affected, and normalizing or contextualizing the behavior of higher-status actors when the situation is framed more externally. We do not claim that models systematically side with authority; our dataset is too small for such a conclusion. However, this pattern raises an important question for fu- ture work: whether LLMs distribute credibility, sympathy, and the benefit of the doubt 9 asymmetrically across relational roles and power positions. In hierarchical relation- ships, apparently neutral interpretation may not be neutral if it grants more benefit of the doubt to the actor with greater institutional power. 5.4 Design Implications: Preserving Uncertainty Without Abandon- ing Helpfulness Our results point to a design challenge for social AI systems: how can assistants sup- port reflection without taking over the userâs interpretive authority? We do not argue that LLMs should avoid interpersonal reflection entirely. There may be value in an always-available, non-judgmental system that helps users articulate possibilities they had not yet considered. The problem is not helpfulness itself, but treating closure as its default form. Future systems could preserve uncertainty through several design orientations. First, separating observation from inference: distinguishing what the user has re- ported from what the model is inferring, so that the epistemic status of each claim remains visible. Second, presenting multiple plausible interpretations without pre- maturely ranking or dismissing them. Third, avoiding action recommendations that presuppose a single unverified reading of the situation. Rather than answering âwhat did they mean?â directly, an uncertainty-preserving assistant might instead ask: what evidence is missing, what interpretations remain equally possible, and what response would protect the userâs agency regardless of the other personâs actual intention? The question is not only what AI should say, but what it should resist saying. For CSCW, these findings point toward a broader agenda: studying the epistemic effects of AI systems in everyday social life. Confident LLM interpretations may shape how users subsequently understand a relationship, their willingness to seek alternative perspectives, and the actions they take next. The current default is therefore not neutral. It is a design choiceâand one with consequential effects on interpersonal sensemaking. 5.5 Limitations and Future Work This study is exploratory and intentionally narrow. We analyze model outputs rather than downstream user behavior, and our prompt set is not intended to represent the full space of interpersonal ambiguity. Because coding was conducted by the first au- thor and refined through team discussion, future work should include multiple coders and formal reliability analysis. Our goal is therefore not to estimate population-level prevalence, but to identify a recurring response pattern that warrants further study. Future work should examine how users respond to interpretive closure in real in- teractions: whether model-generated narratives change usersâ beliefs about othersâ intentions, increase confidence in uncertain judgments, or shape subsequent commu- nication decisions. Another direction is to design and evaluate uncertainty-preserving assistants that support reflection without prematurely resolving interpersonal ambi- guity. Future studies should also examine whether LLMs distribute uncertainty, cred- ibility, and the benefit of the doubt differently across roles such as teacher/student, supervisor/subordinate, or romantic partners. 10 6 Conclusion This paper examined how large language models respond to ambiguous interpersonal situations in which no single interpretation can be confidently verified from the avail- able evidence. Across 72 responses from GPT-4o, Claude Opus, and Gemini Flash, only 9 responses (12.5%) preserved genuine uncertainty. The remaining 87.5% exhibited one or more recurring patterns of interpretive closure, including narrative closure, nor- mative advice under uncertainty, and false epistemic signaling. Although the models differed in style, they converged on the same broader tendency: resolving interper- sonal ambiguity rather than preserving it. Our contribution is not to argue that LLMs should never be used for interpersonal re- flection. Rather, we identify a structural property of current LLM behavior that has re- ceived limited attention: when faced with underdetermined social situations, models tend to produce coherent and actionable narratives rather than maintain interpretive openness. Future social AI systems should not only answer interpersonal questions, but also recognize when preserving uncertainty is the more responsible form of assis- tance. 7 Code and Data Availability All study materials, including the 24 scenario prompts, raw model re- sponses, and coding scheme with coded responses, are publicly available at https://github.com/ming0814/WhatDidTheyMean.The repository also includes the Python scripts used to generate all figures reported in this paper. References Ayers, John W., Adam Poliak, Mark Dredze, Eric C. Leas, Zechariah Zhu, Jessica B. Kelley, Dennis J. Faix, Aaron M. Goodman, Christopher A. Longhurst, Michael Hogarth, and Davey M. Smith. 2023. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Internal Medicine, 183 (6): 589â596. https://doi.org/10.1001/ jamainternmed.2023.1838 Berger, Charles R. and Richard J. Calabrese. 1975. Some explorations in initial interaction and beyond: Toward a developmental theory of interpersonal communication. Human Communication Research, 1 (2): 99â112. https://doi.org/10.1111/j.1468-2958.1975.tb00258.x Bo, Jessica Y., Sophia Wan, and Ashton Anderson. 2025. To Rely or Not to Rely? Evaluating Interven- tions for Appropriate Reliance on Large Language Models. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3706598.3714097 Buçinca, Zana, M. B. Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proceedings of the ACM on Human-Computer Interaction, 5 (CSCW1): 1â21. https://doi.org/10.1145/3449287 Chen, Allison, Sunnie S. Y. Kim, Angel Franyutti, Amaya Dharmasiri, Kushin Mukherjee, Olga Rus- sakovsky, and Judith E. Fan. 2025. Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them. arXiv preprint arXiv:2510.18039. https://doi.org/10. 48550/arXiv.2510.18039 Du, Lihua, Xing Lyu, Lezi Xie, and Bo Feng. 2025. Alignment Without Understanding:A Message- and Conversation-Centered Approach to Understanding AI Sycophancy. arXiv preprint arXiv:2509.21665. https://doi.org/10.48550/arXiv.2509.21665 11 Eckhardt, Sebastian, Niklas Kuhl, Mateusz Dolata, and Gerhard Schwabe. 2024. A Survey of AI Reliance. arXiv preprint arXiv:2408.03948. https://doi.org/10.48550/arXiv.2408.03948 Fricker, Miranda. 2007. Epistemic Injustice: Power and the Ethics of Knowing. Oxford: Oxford University Press. ISBN: 9780198237907 https://doi.org/10.1093/acprof:oso/9780198237907.001.0001 Gabriel, Iason, Arianna Manzini, Geoff Keeling, Lisa Hendricks, Verena Rieser, Hasan Iqbal, Nenad Tomasev, Ira Ktena, Zachary Kenton, Mikel Rodriguez, Seliem El-Sayed, Sasha Brown, Canfer Ak- bulut, Andrew Trask, Edward Hughes, A. S. Bergman, Renee Shelby, Nahema Marchal, Conor Grif- fin, Juan Mateos-Garcia, Laura Weidinger, Winnie Street, Benjamin Lange, Aaron Ingerman, Alison Lentz, Reed Enger, Andrew Barakat, Victoria Krakovna, John Oliver Siy, Zeb Kurth-Nelson, Amanda McCroskery, Vijay Bolina, Harry Law, Murray Shanahan, Lize Alberts, Borja Balle, Sarah de Haas, Yetunde Ibitoye, Allan Dafoe, Beth Goldberg, SĂ©bastien Krier, Alexander Reese, Sims Witherspoon, Will Hawkins, Maribeth Rauh, Don Wallace, Matija Franklin, Josh A. Goldstein, Joel Lehman, Michael Klenk, Shannon Vallor, Courtney Biles, Meredith Ringel Morris, Helen King, Blaise AgĂŒera y Arcas, William Isaac, and James Manyika. 2024. The Ethics of Advanced AI Assistants. arXiv preprint arXiv:2404.16244. https://doi.org/10.48550/arXiv.2404.16244 Goffman, Erving. 1974. Frame Analysis: An Essay on the Organization of Experience. New York: Harper & Row. Heider, Fritz. 1958. The Psychology of Interpersonal Relations. New York: Wiley. Ibrahim, Lujain, Katherine M. Collins, Sunnie S. Y. Kim, Anka Reuel, Max Lamparth, Kevin Feng, Lama Ahmad, Prajna Soni, Alia El Kattan, Merlin Stein, Siddharth Swaroop, Ilia Sucholutsky, Andrew Strait, Q. V. Liao, and Umang Bhatt. 2025. Measuring and Mitigating Overreliance is Necessary for Building Human-Compatible AI. arXiv preprint arXiv:2509.08010. https://doi.org/10.48550/arXiv. 2509.08010 Iftikhar, Zainab, Sean Ransom, A. Xiao, and Jeff Huang. 2024. Therapy as an NLP Task: Psychologistsâ Comparison of LLMs and Human Peers in CBT. arXiv preprint arXiv:2409.02244. https://doi.org/ 10.48550/arXiv.2409.02244 Ouyang, Long, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training Language Models to Follow Instructions with Human Feedback. arXiv preprint arXiv:2203.02155. https://doi.org/10.48550/arXiv.2203.02155 Reeves, Byron and Clifford Nass. 1996. The Media Equation: How People Treat Computers, Television, and New Media Like Real People and Places. Stanford, CA: CSLI Publications and Cambridge University Press. Ross, Lee. 1977. The intuitive psychologist and his shortcomings: Distortions in the attribution process. Advances in Experimental Social Psychology, 10: 173â220. Shen, Hua, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Stefanos Petridis, Yi-Hao Peng, Li Qiwei, Sushrita Rakshit, Chenglei Si, Yutong Xie, Jeffrey P. Bigham, Frank Bentley, Joyce Chai, Zachary C. Lipton, Qiaozhu Mei, Rada Mihalcea, Michael Terry, Diyi Yang, Meredith Ringel Morris, Paul Resnick, and David Jurgens. 2024. Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions. arXiv preprint arXiv:2406.09264. https://doi.org/10.48550/arXiv.2406.09264 Sturgeon, Benjamin, Daniel Samuelson, Jacob Haimes, and Jacy Reese Anthis. 2025. HumanA- gencyBench: Scalable Evaluation of Human Agency Support in AI Assistants. arXiv preprint arXiv:2509.08494. https://doi.org/10.48550/arXiv.2509.08494 Weick, Karl E. 1995. Sensemaking in Organizations. Thousand Oaks, CA: SAGE Publications, Inc. ISBN: 978-0-8039-7177-6 12