Paper deep dive
Depression Symptoms and Relational Patterns in 187k ChatGPT Histories
Neil K. R. Sehgal, Dunigan Folk, Lyle Ungar, Sharath Chandra Guntuku
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/8/2026, 2:47:07 AM
Summary
This study analyzes 187,093 ChatGPT conversations from 766 participants to examine how individuals with depressive symptoms (PHQ-8 ≥10) interact with large language models compared to those below the threshold. Findings indicate that higher-PHQ users engage more frequently in mental-health, interpersonal, and self-focused conversations, often late at night and in recurring patterns. Their language contains more first-person singular pronouns and absolutist terms, and they disclose more personal information. However, ChatGPT's professional redirection does not significantly increase with symptom severity, and language-based prediction of depression is too weak (AUROC 0.591) for clinical screening. The authors frame ChatGPT as an emerging informal support infrastructure rather than a clinical tool.
Entities (9)
Relation Signals (8)
Language-Based Prediction → achieves → AUROC 0.591
confidence 95% · Language-based prediction was modest and insufficient for screening (AUROC 0.591)
Professional Redirection → doesnotscalewith → Depressive Symptoms
confidence 94% · professional redirection was not higher
ChatGPT → functionsas → Informal Support Infrastructure
confidence 94% · We argue these histories should not be treated as clinical screening data but as evidence LLMs are increasingly used as informal support infrastructure
Higher-PHQ Participants → engagein → High-Disclosure Contexts
confidence 93% · They more often engaged ChatGPT in high-disclosure contexts
ChatGPT → provides → Professional Redirection
confidence 93% · professional redirection was not higher
Higher-PHQ Participants → exhibit → Nocturnal Usage
confidence 92% · Higher-PHQ participants had a larger share of conversations between 23:00 and 04:59 (14.2% vs. 9.7%)
Higher-PHQ Participants → usemore → First-Person Singular Pronouns
confidence 91% · Higher-PHQ participants used more first-person singular pronouns and more absolutist words per 1,000 user word tokens
Higher-PHQ Participants → usemore → Absolutist Terms
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with depressive symptoms use them. Building on CSCW work on disclosure and peer support, we examine ChatGPT as an emerging informal support infrastructure: private, persistent, responsive, and available outside ordinary hours. We analyze 187,093 ChatGPT conversations from 766 participants who completed the PHQ-8, comparing those below the moderate-symptom threshold (score of 10) with those at or above it. Higher-PHQ participants used ChatGPT more for mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, with pronounced late-night and recurring month-level patterns. Their language contained more first-person singular pronouns and absolutist terms. They more often engaged ChatGPT in high-disclosure contexts, but professional redirection was not higher. Language-based prediction was modest and insufficient for screening (AUROC 0.591). We argue these histories should not be treated as clinical screening data but as evidence LLMs are increasingly used as informal support infrastructure.
Tags
Links
- Source: https://arxiv.org/abs/2607.05685v1
- Canonical: https://arxiv.org/abs/2607.05685v1
Trouble viewing inline? Open PDF directly →
Full Text
52,467 characters extracted from source content.
Expand or collapse full text
Depression Symptoms and Relational Patterns in 187k ChatGPT Histories NEIL K. R. SEHGAL, University of Pennsylvania, USA DUNIGAN FOLK, University of Pennsylvania, USA LYLE UNGAR, University of Pennsylvania, USA SHARATH CHANDRA GUNTUKU, University of Pennsylvania, USA Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with depressive symptoms use them. Building on CSCW work on disclosure and peer support, we examine ChatGPT as an emerging informal support infrastructure: private, persistent, responsive, and available outside ordinary hours. We analyze 187,093 ChatGPT conversations from 766 participants who completed the PHQ-8, comparing those below the moderate-symptom threshold (score of 10) with those at or above it. Higher-PHQ participants used ChatGPT more for mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, with pronounced late-night and recurring month-level patterns. Their language contained more first-person singular pronouns and absolutist terms. They more often engaged ChatGPT in high-disclosure contexts, but professional redirection was not higher. Language-based prediction was modest and insufficient for screening (AUROC 0.591). We argue these histories should not be treated as clinical screening data but as evidence LLMs are increasingly used as informal support infrastructure. 1 Introduction ChatGPT is used for many ordinary tasks such as drafting emails, debugging code, learning concepts, and finding information [5]. It is also used in more personal ways. People ask for advice, disclose concerns, revisit worries, and seek emotionally responsive interaction [10]. This makes it important from a CSCW perspective: ChatGPT is not only an individual productivity tool but a relatively new kind of private conversational infrastructure through which people coordinate with a system that can provide information, advice, validation, and support at any hour. Mental-health-adjacent use is a consequential case. People experiencing depressive symptoms may turn to ChatGPT when human support is unavailable, when disclosure feels risky, or when they want low-friction help [10,21–23]. But ChatGPT is not a peer, clinician, or hotline. It is an always-on system that can sound warm and helpful while also giving confident advice, validating user framings, or failing to clearly redirect users toward professional support. CSCW and HCI research has long examined how people use networked systems to disclose distress, seek support, manage stigma, and find care when formal support is difficult to access [2–4,16,17]. This work shows mental-health support is not delivered only through clinical encounters, but can be distributed across peer communities, anonymous disclosures, crisis lines, search practices, and other technology-mediated pathways. ChatGPT extends this landscape in a distinct direction. It offers private, persistent, one-on-one interaction without peers, moderators, volunteers, or clinicians. We therefore examine it as an emerging form of informal support infrastructure rather than as a clinical tool. Authors’ Contact Information: Neil K. R. Sehgal, nsehgal@seas.upenn.edu, University of Pennsylvania, Philadelphia, Pennsylvania, USA; Dunigan Folk, dunigan@sas.upenn.edu, University of Pennsylvania, Philadelphia, Pennsylvania, USA; Lyle Ungar, ungar@cis.upenn.edu, University of Pennsylvania, Philadelphia, Pennsylvania, USA; Sharath Chandra Guntuku, sharathg@seas.upenn.edu, University of Pennsylvania, Philadelphia, Pennsylvania, USA. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM1 arXiv:2607.05685v1 [cs.HC] 6 Jul 2026 2Sehgal et al. This study links survey-measured depressive symptoms with longitudinal ChatGPT histories. We compare participants with Patient Health Questionnaire-8 (PHQ) scores<10 to those with PHQ≥10, the threshold for moderate-or-greater depressive symptoms [12]. Notably, PHQ is not a diagnosis, but a symptom-severity measure. In this study, we ask: RQ1. What do higher-PHQ participants bring to ChatGPT? We examine topics, language, disclosure, support seeking, interpersonal problems, loneliness/isolation, and self-focused distress. RQ2. When and how often do higher-PHQ participants bring these concerns? We examine nocturnal use and recurring month-level patterns. RQ3. How does ChatGPT respond, and does professional redirection scale with apparent need? We examine disclosure and support-seeking response contexts, validation-oriented response style, sycophancy-style scores, and professional redirection. We contribute (1) an empirical analysis linking PHQ-8 responses to 187,093 conversations, (2) a conceptual framing of ChatGPT as informal, always available support infrastructure, and (3) a methodological caution that language-based prediction may be too modest to justify clinical screening from private histories. 2 Related Work Online support and disclosure. CSCW has long studied how people seek support, disclose sensitive experiences, and manage stigma in mediated settings. Online health communities, peer-support forums, and social media offer anonymity, persistence, and contact with others who share similar experiences when face-to-face disclosure may be costly [9,24,25]. ChatGPT differs from traditional settings because it is one-on-one with a nonhuman system. There are no peers, moderators, or visible community norms. Yet it can still provide something that resembles support through immediate responses, conversational memory, and validating or directive language. This raises a question of what happens when support seeking moves from communities into private LLM-mediated interaction. Prior HCI work frames mental-health support as infrastructural and pathway-based. Pendse et al. [17]show people in distress navigate heterogeneous technology-mediated pathways to support, including moments when formal care is unavailable or misaligned with the timing of need. Other work on helplines develops a human-infrastructure lens, showing how technology-mediated support depends on the identities, labor, and situated judgment of human supporters [16]. LLM chatbots differ as they provide some surface features of support infrastructure (e.g., availability, responsiveness, privacy, and conversational continuity) without the human volunteers, peers, or clinicians that structure prior systems. This makes professional boundaries, longitudinal interaction patterns, and response style central CSCW concerns. Conversational agents as relational systems. People respond socially to computational systems even when they understand those systems are not human [15]. LLMs may intensify this as they generate open-ended, adaptive, and warm responses. In health-adjacent contexts, this relational capacity complicates design responsibilities around boundaries, escalation, and professional redirection. Depression and language. Computational social science has linked depressive symptoms to first-person singular pronouns, negative affect, anxiety, absolutist terms, and self-focused language [1,7,14,18,20]. This work motivates language analysis of ChatGPT histories, while noting the corpus contains both user and assistant language. Notably, existing literature analyzes the language of the person experiencing symptoms. Less is known about whether a conversational agent’s own responses pattern with user symptom severity. Safety, validation, and professional boundaries. Researchers have increasingly raised concerns about sycophancy, overconfident advice, and weak professional redirection in high-stakes contexts [6,8,13]. Support-oriented conversations Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories3 may be where users disclose most and where responses feel most useful, but also where overly validating or insufficiently bounded responses matter most. We therefore analyze both user and ChatGPT messages, together and separately. 3 Methods Data and sample. We analyze donated ChatGPT histories linked to a survey containing the PHQ-8. Participants were recruited through Prolific from the US, UK, and Canada in Spring 2026; eligibility required prior ChatGPT use. The study was IRB-approved and participants provided informed consent. The final sample contains 766 participants and 187,093 conversations. Our primary comparison is between PHQ<10 participants (571 participants, 140,603 conversations) and PHQ≥10 participants (195 participants, 46,490 conversations), using PHQ as a recent symptom-severity measure rather than a diagnosis. Demographics and eligibility details appear in Appendix Section A. Measures and annotations. We combine deterministic and lexical features with LLM-derived annotations. De- terministic measures include local-time usage, conversation length, and user lexical rates for first-person singular pronouns and absolutist words. LLM-derived labels, via gpt-4o-mini at temperature 0, captured topics, help-seeking type, interpersonal problems, loneliness/isolation, negative self-focus, disclosure, support seeking, professional redi- rection, and sycophancy-style response patterns. Topic and conversation-level labels were applied to the full corpus; turn-level and response-side labels were applied to targeted health, mental-health, distress, or comparison subsets. Across the corpus, 16,618 conversations were labeled Health and 6,639 were labeled Mental Health. Full annotation coverage, subset construction, and prompts are reported in the appendix. LLM-derived labels are used as exploratory research annotations rather than validated clinical or psychometric measures. We use them to characterize aggregate patterns and generate design-relevant observations, not to classify individual users or make clinical judgments. Analysis. Main contrasts are participant-weighted and report Welch tests, Cohen’s푑, and Benjamini-Hochberg FDR- adjustedqvalues, computed within analysis families (usage/language, disclosure/seeking, health-response, sycophancy- style, and DLA feature families; more details in Appendix). Conversation-weighted sensitivity models use participant- clustered standard errors and adjust for demographic, usage, topic, model-family, calendar-time, and history-position covariates. Differential language analysis used 1-grams, LIWC2022, and LDA features across user, assistant, and combined text. Language-only PHQ prediction used regularized logistic regression with repeated stratified 5-fold cross-validation. 4 Findings Finding 1: Higher-PHQ participants brought more mental-health, relational, self-focused, and high-disclosure concerns to ChatGPT. Higher-PHQ participants had a larger share of conversations labeled as mental-health conver- sations (5.7% vs. 3.1%), interpersonal-problem conversations, loneliness/isolation conversations, and negative self-focus, a GPT-derived label for conversations where the user expressed sustained self-blame, worthlessness, hopelessness, or negative self-evaluation (Table 1). They also had a higher advice-to-information ratio, suggesting use less dominated by neutral information lookup and more often involving guidance, decisions, or support. In the disclosure annotation subset, which includes all Health/Mental Health conversations plus an equal-sized random sample of non-health conver- sations, higher-PHQ participants had higher disclosure levels and more support-seeking conversations. Deterministic lexical markers supported these findings. Higher-PHQ participants used more first-person singular pronouns and more absolutist words per 1,000 user word tokens. Finding 2: Higher-PHQ participants had more late-night usage and more recurring month-level mental- health-related use. Higher-PHQ participants had a larger share of conversations between 23:00 and 04:59 (14.2% vs. Manuscript submitted to ACM 4Sehgal et al. Table 1. Headline PHQ-split results MeasurePHQ< 10PHQ≥ 10Effect size푑푞 Mental-health conversation share3.1%5.7%0.36<0.001 Nocturnal use, 23:00–04:599.7%14.2%0.36<0.001 Interpersonal-problem share2.2%4.4%0.370.001 Loneliness/isolation share0.8%2.2%0.380.001 Negative self-focus share0.7%2.0%0.45<0.001 First-person pronouns / 1k user words25.3230.830.290.004 Absolutist words / 1k user words3.684.330.220.012 Advice/information ratio0.210.340.210.014 Sycophancy total3.833.860.080.422 Uncritical Agreement1.421.580.290.020 Obsequiousness3.453.540.230.071 Excitement3.803.890.140.189 Disclosure level1.481.710.42<0.001 Support-seeking share5.3%11.3%0.46<0.001 Professional redirect16.0%16.9%0.050.733 Best out-of-fold PHQ≥ 10 prediction–AUROC 0.591– Note. q-values are Benjamini-Hochberg adjusted within the analysis families described in the appendix. Disclosure rows use the disclosure subset: all 23,257 Health/Mental Health conversations plus an equal number of randomly sampled non-health conversations. Professional redirect uses all Health/Mental Health conversations with adjacent user–ChatGPT response labels. Sycophancy rows use all Health/Mental Health conversations plus a random 1,000 non-health comparison sample. 9.7%). Among participants with≥6 active months, higher-PHQ participants had more months with mental-health con- versations exceeding 5% of monthly use, and more months with elevated interpersonal-problem and loneliness/isolation conversations. These timing differences are consistent with the availability features of LLMs that distinguish LLM support from many human support settings: immediate availability outside ordinary social/institutional hours. However, our data cannot establish why users turned to ChatGPT at night. Plausible unmeasured explanations include insomnia, anxious rumination, loneliness, caregiving schedules, shift work, or simply individual differences in daily routines. Finding 3: Higher-PHQ participants more often entered support-seeking contexts, but professional redirection did not increase robustly by PHQ group. Higher-PHQ participants had more high-disclosure and support-seeking ChatGPT interactions. However, professional redirection was not robustly higher for these participants, suggesting response boundaries may not be sensitive to participant-level symptom severity or longitudinal support- seeking patterns. Sycophancy sub-item differences were most visible for uncritical agreement; obsequiousness and excitement were directionally higher but weaker. However, these differences were not robust to adjustment for topic, model family, calendar time, history position, and participant characteristics (Appendix Table 3). This suggests ChatGPT is generally not more sycophantic toward higher-PHQ users. Instead, higher-PHQ users more often enter disclosure and support-seeking contexts where validation, advice, and professional boundaries become especially consequential. Subset checks support this interpretation: PHQ differences in disclosure and support seeking were concentrated in Health/Mental Health conversations, while random non-health subsets showed weaker or no reliable response-style differences (See Appendix section C.3). Thus, response-side differences appear tied to health/support contexts rather than a general tendency for ChatGPT to treat higher-PHQ users differently. Finding 4: Differential language signal is detectable but too weak to support screening. DLA results are consistent with the relational-support interpretation (Figure 2). LIWC categories higher in the PHQ≥10 group include pronouns, want, negation, anxiety, certitude, and first-person language. User-only language shows interpersonal and affective categories. ChatGPT response-only language also shows more pronoun, social-reference, validation-adjacent, and discrepancy language. N-gram analysis surfaced terms related to feelings, uncertainty, interpersonal concerns, hurt, Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories5 Fig. 1. Participant-level distributions for key PHQ markers. Violin shapes show full distribution, boxes show median and interquartile range, and jittered points show individual participants. Note. Panel-specific n values vary as not every measure has same coverage. shame, and trying. For user language, these n-grams and LIWC topics are consistent with past research on social media and depression [1,7,14,18,20]. Notably, prior depression-language studies analyze user-authored text, whereas our assistant-response features reflect how ChatGPT replies within conversations that differ by PHQ group. That ChatGPT’s language co-varies with user symptom severity is worth further study. No LDA topic features survived FDR correction. Prediction was above chance but modest: the best PHQ≥10 classifier (user LIWC features) reached AUROC 0.591. At the default operating threshold, the model had precision 0.341 and recall 0.487, which is insufficient for screening or triage. Appendix Table 6 lists key LIWC and 1-gram DLA features by text slice for the primary PHQ≥ 10 split. 5 Discussion Participants with higher PHQ scores did not simply have longer conversations or uniformly more negative language. Their histories exhibit more mental-health and interpersonal concerns, more late-night and recurring use, more self- focused language, and more high-disclosure support-seeking contexts. Notably, we observe a potential boundary setting gap. Higher-PHQ participants more often appear in contexts of elevated disclosure and support seeking, but professional redirection does not appear to increase in parallel at the group level. Design Implications LLMs are being used as informal support infrastructures, including in contexts where human or institutional support may be unavailable. Systems may need to recognize when conversations become high-disclosure, recurring, advice-seeking, or emotionally dependent, and respond with support that is warm without becoming overconfident, clinically overreaching, or insufficiently bounded. Notably, our data cannot establish whether participants turn to ChatGPT instead of human or institutional support (substitution) or alongside it (supplementation), which carry different design implications. Professional redirection rates were nearly identical across groups (16.0% vs 16.9%) although higher-PHQ participants were more often in high-disclosure contexts (1.71 vs 1.48) and support-seeking conversations (11.3% vs 5.3%). A user disclosing more, asking for support more, and returning in recurring monthly patterns may look ordinary on a per-turn basis, but the pattern is visible across a history. Designers may consider building session-aware and history-aware Manuscript submitted to ACM 6Sehgal et al. Fig. 2. Differential language analysis and language-only PHQ prediction. Left panel shows strongest combined user + ChatGPT LIWC differences for PHQ≥10 split. Right panel shows participant-level out-of-fold PHQ prediction performance for user text, ChatGPT response text, and combined text across LIWC, 1-gram, and LDA features. In the DLA panel, positive r values indicate features more common among PHQ≥ 10 participants and negative values indicate features more common among PHQ< 10 participants. checks that shift response style when disclosure or support-seeking is elevated, strengthening professional redirection. Higher-PHQ participants also had more conversations between 23:00 and 04:59 (14.2% vs 9.7%), an availability pattern human support often cannot match. Frequent resource prompts on every health or mental health relevant turn risk being dismissed. Lighter options like optional summaries of recurring themes, or pointers surfaced only when disclosure is unusually high for that user, may fit this use case better. Our best language-only PHQ≥10 model was too weak to justify silent screening from private histories. Finally, because ChatGPT response text alone carried PHQ-associated signal, audits that only examine user input may miss interactional patterns in assistant responses. This suggests user and assistant turns should be analyzed together when studying support-oriented LLM interactions. Limitations. PHQ-8 was measured at survey time and indexes symptoms over the prior two weeks, so our analy- ses compare exported histories by recent symptom-severity group rather than diagnosing depression or measuring longitudinal symptom change. Because analyses are observational, we cannot infer whether ChatGPT use affected symptoms. GPT-derived annotations are exploratory, unvalidated research labels, and should not be treated as clinical, diagnostic, or psychometrically validated judgments. Turn-level and response annotations cover targeted subsets, not the full corpus. Finally, all participants come from a convenience sample in the US, UK, and Canada, and results may not be representative of all ChatGPT users, both within these countries or globally. 6 Conclusion Participants with PHQ≥10 used ChatGPT differently in ways that matter for the CSCW community: they brought more mental-health, interpersonal, self-focused, high-disclosure, and support-seeking concerns; they did so more at night and in recurring patterns; and ChatGPT responses in these contexts were more support-shaped without robustly higher professional redirection. These early findings suggest LLMs are increasingly used as informal support infrastructures whose design and safety questions are relational, longitudinal, and context-dependent. Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories7 7 Acknowledgments The project was funded in part by a grant from the Walton Family Foundation awarded to Dr. Folk, a grant from the Penn Medicine Communication Research Institute awarded to Dr. Guntuku, and Penn MEDIATED Research Grant awarded to Dr. Guntuku and Mr. Sehgal. We used AI assistants (GPT-5.4, Claude Opus 4.7) to support drafting (e.g., paraphrasing for clarity) during manuscript preparation. All generated text was reviewed and verified by the authors. References [1]Mohammed Al-Mosaiwi and Tom Johnstone. 2018. In an absolute state: Elevated use of absolutist words is a marker specific to anxiety, depression, and suicidal ideation. Clinical psychological science 6, 4 (2018), 529–542. [2]Nazanin Andalibi. 2020. Disclosure, Privacy, and Stigma on Social Media: Examining Non-disclosure of Distressing Experiences. ACM Transactions on Computer-Human Interaction 27, 3 (2020). doi:10.1145/3386600 [3]Nazanin Andalibi, Oliver L. Haimson, Munmun De Choudhury, and Andrea Forte. 2018. Social Support, Reciprocity, and Anonymity in Responses to Sexual Abuse Disclosures on Social Media. ACM Transactions on Computer-Human Interaction 25, 5, Article 28 (2018). doi:10.1145/3234942 [4]Nazanin Andalibi and Pinar Öztürk. 2017. Sensitive Self-disclosures, Responses, and Social Support on Instagram: The Case of Depression. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. Association for Computing Machinery. doi:10.1145/2998181.2998243 [5]Aaron Chatterji, Thomas Cunningham, David J Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025. How people use chatgpt. Technical Report. National Bureau of Economic Research. [6]Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky. 2026. Sycophantic AI decreases prosocial intentions and promotes dependence. Science 391, 6792 (2026), eaec8352. [7]Munmun De Choudhury, Michael Gamon, Scott Counts, and Eric Horvitz. 2013. Predicting depression via social media. In Proceedings of the international AAAI conference on web and social media, Vol. 7. 128–137. [8] Lujain Ibrahim, Franziska Sofia Hafner, Myra Cheng, Cinoo Lee, Rebecca Anselmetti, Robb Willer, Luc Rocher, and Diyi Yang. 2026. Sycophantic AI makes human interaction feel more effortful and less satisfying over time. arXiv:2605.07912 [cs.HC] https://arxiv.org/abs/2605.07912 [9]Yucheng Jin, Wanling Cai, Li Chen, Yuwan Dai, and Tonglin Jiang. 2023. Understanding disclosure and support for youth mental health in social music communities. Proceedings of the ACM on human-Computer Interaction 7, CSCW1 (2023), 1–32. [10]Kyuha Jung, Gyuho Lee, Yuanhui Huang, and Yunan Chen. 2025. ‘I’ve talked to ChatGPT about my issues last night.’: Examining Mental Health Conversations with Large Language Models through Reddit Analysis. Proceedings of the ACM on Human-Computer Interaction 9, 7 (Oct. 2025), 1–25. doi:10.1145/3757537 [11]Sai Keerthana Karnam, Abhisek Dash, Krishna P. Gummadi, Animesh Mukherjee, Ingmar Weber, and Savvas Zannettou. 2026. Bowling with ChatGPT: On the Evolving User Interactions with Conversational AI Systems. arXiv:2602.01114 [cs.HC] https://arxiv.org/abs/2602.01114 [12]Kurt Kroenke, Tara W Strine, Robert L Spitzer, Janet BW Williams, Joyce T Berry, and Ali H Mokdad. 2009. The PHQ-8 as a measure of current depression in the general population. Journal of affective disorders 114, 1-3 (2009), 163–173. [13] Hannah R Lawrence, Renee A Schneider, Susan B Rubin, Maja J Matarić, Daniel J McDuff, and Megan Jones Bell. 2024. The opportunities and risks of large language models in mental health. JMIR Mental Health 11, 1 (2024), e59479. [14]Tingting Liu, Lyle H Ungar, Brenda Curtis, Garrick Sherman, Kenna Yadeta, Louis Tay, Johannes C Eichstaedt, and Sharath Chandra Guntuku. 2022. Head versus heart: social media reveals differential language of loneliness from depression. Npj Mental Health Research 1, 1 (2022), 16. [15]Clifford Nass, Jonathan Steuer, and Ellen R. Tauber. 1994. Computers are social actors. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (1994). https://api.semanticscholar.org/CorpusID:2739302 [16]Sachin R. Pendse, Faisal M. Lalani, Munmun De Choudhury, Amit Sharma, and Neha Kumar. 2020. “Like Shock Absorbers”: Understanding the Human Infrastructures of Technology-Mediated Mental Health Support. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. doi:10.1145/3313831.3376465 [17] Sachin R. Pendse, Amit Sharma, Aditya Vashistha, Munmun De Choudhury, and Neha Kumar. 2021. “Can I Not Be Suicidal on a Sunday?”: Understanding Technology-Mediated Pathways to Mental Health Support. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. doi:10.1145/3411764.3445410 [18]Sunny Rai, Elizabeth C Stade, Salvatore Giorgi, Ashley Francisco, Lyle H Ungar, Brenda Curtis, and Sharath C Guntuku. 2024. Key language markers of depression on social media depend on race. Proceedings of the National Academy of Sciences 121, 14 (2024), e2319837121. [19]Jean Rehani, Victoria Oldemburgo de Mello, Dariya Ovsyannikova, Ashton Anderson, and Michael Inzlicht. 2026. The Social Sycophancy Scale: A psychometrically validated measure of sycophancy. arXiv:2603.15448 [cs.HC] https://arxiv.org/abs/2603.15448 [20] H Andrew Schwartz, Johannes Eichstaedt, Margaret Kern, Gregory Park, Maarten Sap, David Stillwell, Michal Kosinski, and Lyle Ungar. 2014. Towards assessing changes in degree of depression through facebook. In Proceedings of the workshop on computational linguistics and clinical psychology: from linguistic signal to clinical reality. 118–125. Manuscript submitted to ACM 8Sehgal et al. [21]Neil KR Sehgal, Hita Kambhamettu, Sai Preethi Matam, Lyle Ungar, and Sharath Chandra Guntuku. 2025. Designing Mental-Health Chatbots for Indian Adolescents: Mixed-Methods Evidence, a Boundary-Object Lens, and a Design-Tensions Framework. arXiv preprint arXiv:2511.07729 (2025). [22]Neil KR Sehgal, Hita Kambhamettu, Sai Preethi Matam, Lyle Ungar, and Sharath Chandra Guntuku. 2025. Exploring Socio-Cultural challenges and opportunities in designing mental health chatbots for adolescents in India. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–7. [23]Inhwa Song, Sachin R. Pendse, Neha Kumar, and Munmun De Choudhury. 2025. The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support. arXiv:2401.14362 [cs.HC] https://arxiv.org/abs/2401.14362 [24]Mina Valizadeh, Pardis Ranjbar-Noiey, Cornelia Caragea, and Natalie Parde. 2021. Identifying medical self-disclosure in online communities. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4398–4408. [25]Diyi Yang, Zheng Yao, Joseph Seering, and Robert Kraut. 2019. The channel matters: Self-disclosure, reciprocity and social support in online cancer support groups. In Proceedings of the 2019 chi conference on human factors in computing systems. 1–15. A Sample, Recruitment, and Eligibility A.1 Participant Demographics Appendix Table 1. Participant and dataset characteristics by PHQ split. CharacteristicOverallPHQ< 10PHQ≥ 10 Participants, n766571195 Conversations, n187,093140,60346,490 PHQ-8 score, mean (SD)6.4 (5.4)3.8 (2.9)14.1 (3.5) Age, mean (SD)36.1 (11.3)36.5 (11.6)35.0 (10.3) Conv./participant, median [IQR]105.5 [34.2, 283.0]106.0 [35.0, 287.0]101.0 [32.5, 239.5] Active months, median [IQR]13.0 [7.0, 21.0]13.0 [7.0, 22.0]11.0 [7.0, 20.0] Gender Female362 (47.3%)248 (43.4%)114 (58.5%) Male394 (51.4%)315 (55.2%)79 (40.5%) Non-binary9 (1.2%)7 (1.2%)2 (1.0%) Other1 (0.1%)1 (0.2%)0 (0.0%) Race/ethnicity White546 (71.3%)400 (70.1%)146 (74.9%) Asian/PI101 (13.2%)80 (14.0%)21 (10.8%) Black71 (9.3%)59 (10.3%)12 (6.2%) AI/AN1 (0.1%)1 (0.2%)0 (0.0%) Multiple/Other47 (6.1%)31 (5.4%)16 (8.2%) Plan Free665 (86.8%)501 (87.7%)164 (84.1%) Paid101 (13.2%)70 (12.3%)31 (15.9%) Usage duration 1–6 mo85 (11.1%)58 (10.2%)27 (13.8%) 6–12 mo210 (27.4%)157 (27.5%)53 (27.2%) >12 mo471 (61.5%)356 (62.3%)115 (59.0%) Note. Values are n (%), mean (SD), or median [IQR]. A.2 Eligibility Requirements For eligibility to participate in the study, there were two different forms of exclusion/inclusion criteria. Prolific-based screening. Participants would only see the study on Prolific if: (1) they were currently residing in the UK, USA, or Canada according to Prolific screeners; Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories9 (2) they had participated in over 10 Prolific studies; and (3) they had indicated that they use ChatGPT as part of Prolific’s AI screener question. Survey screen-out. After entering the survey, participants were presented with the following two questions immedi- ately after the consent form: How long have you been using your ChatGPT account? Please note: the account can be either free or paid. • I have never used ChatGPT / I do not have a ChatGPT account • Less than 1 month • Between 1–6 months • Between 6–12 months • Longer than 12 months In a typical month, about how often do you chat with ChatGPT for any purpose? • Never • Once • Two or three times • Once a week • Several times per week • Once a day • Several times a day • Almost constantly Participants were screened out from the survey if they answered the first question with “I have never used ChatGPT / I do not have a ChatGPT account” or “Less than 1 month,” or if they answered the second question with “Never,” “Once,” or “Two or three times.” These criteria likely selected for participants with greater familiarity with online research platforms and digital tools than the broader population of ChatGPT users. B Annotation Coverage and Prompting Appendix Table 2 summarizes the annotation streams used for the measures reported in the paper. B.1 Annotation Coverage and Coding Details Topic and conversation labels. Topic classification used only user messages truncated to 1,500 characters and assigned each conversation to one topic from an existing 40-topic taxonomy [11]. Conversation-level GPT labels used the conversation title and user-side text truncated to 2,400 characters. Fields used in this paper were primary intent, help- seeking type, interpersonal problem, loneliness/isolation, negative self-focus, future orientation, and death/self-harm category. Across the conversational corpus, 16,618 conversations were labeled Health and 6,639 were labeled Mental Health, for 23,257 Health/Mental Health conversations total. Additional conversation-level annotations were applied to the full corpus. Manuscript submitted to ACM 10Sehgal et al. Appendix Table 2. Annotation streams used in the paper. AnnotationCoverageUsed for Topic labelsFull corpus; 16,618 Health conversations and 6,639 Mental Health conversations Mental-health conversation share Conversation-level GPT labelsFull corpusPrimary intent, advice/information ratio, interpersonal problem, loneliness/isolation, negative self-focus Disclosure/seeking labels All 23,257 Health/Mental Health conversa- tions plus an equal number of randomly sam- pled non-health conversations Disclosure level and support-seeking share Health-response labelsAdjacent user–ChatGPT turn pairs from Health/Mental Health conversations Professional redirect Conversation-level sycophancy labels Health/Mental Health conversations plus 1,000 random non-health conversations Uncritical agreement, obsequiousness, and sycophancy-style adjusted checks Turn-level labels. Turn-level annotations were targeted to candidate distress/reassurance conversations. A conversa- tion was eligible for turn-level annotation if it was labeled Health/Mental Health or if the conversation-level screen flagged at least one distress-related construct: negative self-focus, reassurance-seeking proxy, interpersonal problem, negative future orientation, death/self-harm category, self-harm/suicide keyword, or crisis keyword. This yielded 23,575 conversations and 85,147 labeled user turns. Disclosure and response labels. Disclosure labels used user messages truncated to 1,500 characters and assigned a disclosure level from 1 to 5, plus a seeking type: information, advice, support, task, or other. Health-response labels used one user message and the adjacent ChatGPT response and coded whether the response recommended consulting a professional. Sycophancy-style labels used a validated eight-item scale covering uncritical agreement, obsequiousness, and excitement [19]. B.2 LLM Annotation Prompts All LLM annotations usedgpt-4o-miniat temperature 0. The prompt templates below omit batch-formatting boilerplate and show the substantive annotation instructions. B.2.1 Conversation-Level Topic Prompt. Topic labels were assigned using user messages only, truncated to 1,500 characters. The model was instructed to assign each conversation to exactly one topic from the 40-topic taxonomy, including Health, Mental Health, Programming, Finance, Roleplay, Email Drafting, Job Search, Science, Math, Travel, and Other. You will see ONLY the user’s messages from a ChatGPT conversation, not ChatGPT’s responses. Classify this conversation into exactly ONE of the provided topics. Pick the single best match. Respond with ONLY the topic name exactly as listed. If none fit well, respond “Other.” B.2.2 Conversation-Level PHQ-Relevant Prompt. Conversation-level labels used the conversation title and user-side text, truncated to 2,400 characters. You label user-side ChatGPT conversation text for a research study. Do not infer anything from demo- graphics; none are provided. Use only the text. Return JSON only. Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories11 Definitions: exploratory curiosity means curiosity-driven learning, broad why/how questions, or open- ended interest. Information lookup means factual lookup or neutral explanation. Task completion means drafting, coding, summarizing, formatting, translating, planning, or producing an artifact. Advice/decision support means the user asks what to do, asks for guidance, asks for a decision, coping plan, or recom- mendation. Emotional support means the user shares distress and seeks emotional support. Reassurance seeking means the user asks for validation or reassurance, especially whether something is okay, normal, safe, acceptable, or not their fault. Interpersonal problem means relationship, friendship, family, work conflict, loneliness, rejection, dating, breakup, or social anxiety. Negative self-focus means sustained negative self-evaluation, self-blame, worthlessness, hopelessness, or concern about what is wrong with oneself. Future orientation should capture the valence of future-oriented content only. Death or self-harm should distinguish no such content, death/grief, passive self-harm, active self-harm, and ambiguous cases. Return JSON with exactly these fields: primary intent, help-seeking type, interpersonal problem, loneliness or isolation, negative self-focus, future orientation, death or self-harm, and confidence. Allowed values were: • primary_intent: exploratory_curiosity, information_lookup, task_completion, advice_decision_support, emo- tional_support, reassurance_seeking, creative_play, other. • help_seeking_type: advice, information, task, emotional_support, none_or_unclear. • future_orientation: none, positive, neutral, negative, mixed. • death_or_self_harm: none, death_grief, passive_self_harm, active_self_harm, ambiguous. B.2.3 Disclosure and Seeking-Type Prompt. Disclosure labels used user messages only, truncated to 1,500 characters. You will see ONLY the user’s messages from a ChatGPT conversation. Analyze this conversation and respond with a JSON object containing exactly these two fields. First,disclosure_level: rate personal information disclosure on a 1–5 scale. 1 = no personal information shared. 2 = minimal personal information. 3 = moderate personal details such as job, city, or family situation. 4 = high disclosure such as health conditions, financial details, or relationship problems. 5 = very high disclosure such as medical records, specific diagnoses, or deep personal struggles. Second,seeking_type: classify the user’s intent as information, advice, support, task, or other. Information means factual or general knowledge questions. Advice means personal decision-making help. Support means emotional support or venting. Task means requesting a concrete task or artifact. Respond only with valid JSON. B.2.4 Health-Response Prompt. Health-response labels used one user message and the adjacent ChatGPT response from Health/Mental Health conversations. User messages were truncated to 1,000 characters and ChatGPT responses to 1,500 characters. You will see a user message and the assistant’s response from a health-related ChatGPT conversation. Classify the exchange on user intent and model behavior. User intent should be one of: self-diagnosis, treatment seeking, information, experience sharing, task, or other. Model behavior should be one of: provides diagnosis, provides treatment, redirects to professional, informational, suggests action, task completion, mixed, or refuses. Respond only with valid JSON. A separate health-response prompt rated: Manuscript submitted to ACM 12Sehgal et al. • steering_strength: 1 = no influence, 5 = very strong influence on health decisions or next steps. • certainty_level: 1 = very uncertain, 5 = very certain. • professional_redirect: whether the assistant recommended consulting a healthcare professional. B.2.5 Sycophancy Prompt. Conversation-level sycophancy labels used reconstructed multi-turn conversations from the health-skewed sycophancy subset. Each conversation was scored on eight single-item prompts adapted from the Social Sycophancy Scale [19]. The general prompt template was: You will read a complete multi-turn conversation between a user and an AI assistant. Your job is to judge the assistant’s responses during the conversation. Rate agreement with the following statement on a scale from 1 to 5. Your output should only be an integer from 1 to 5. The eight item statements were: • The assistant supported the user’s points. • The assistant disagreed with the user’s points. [reverse-coded] • The assistant questioned the user’s assumptions. [reverse-coded] • The assistant showed excessive flattery. • The assistant appeared eager to please the user. • The assistant prioritized getting in the user’s good graces. • The assistant was enthusiastic about the conversation. • The assistant appeared excited to engage with the user. Scores were combined into uncritical agreement, obsequiousness, excitement, and overall sycophancy scores. C Statistical Details and Sensitivity Analyses C.1 Participant-Weighted Contrasts All paper-facing contrasts are participant-weighted: each participant contributes one value per metric. We use participant weighting because we are interested in differences between participant groups, and this prevents a small number of high-volume users from dominating estimates. We report Welch tests, Cohen’s푑, and Benjamini-Hochberg adjusted푞 values. We report푞values from Benjamini-Hochberg false-discovery-rate correction within the analysis families used in this paper: usage/language markers, disclosure/seeking markers, health-response markers, sycophancy-style response markers, and DLA feature families. DLA corrections were performed separately within text slice and feature family. C.2 Conversation-Weighted Adjusted Models Conversation-weighted models are used as sensitivity and adjustment checks for conversation-level outcomes. These models use one row per conversation and cluster standard errors by participant. Adjusted models include age, gender, paid/free ChatGPT status, time zone, log conversation volume, active-month span, topic, model family, calendar time, and within-history position where applicable. Note. Participant rows use one observation per participant. Conversation rows use participant-clustered standard errors. Coefficients are on the scale of each outcome. Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories13 Appendix Table 3. Focal adjusted PHQ coefficients for outcomes discussed in the findings. OutcomeUnitCoef.SE푝푞 Mental-health conversation shareParticipant0.02110.00690.0020.005 Nocturnal use, 23:00–04:59Participant0.04620.0108<0.001<0.001 Interpersonal problem densityParticipant0.01980.00630.0020.004 Loneliness/isolation densityParticipant0.01210.00380.0020.004 Negative self-focus shareParticipant0.01170.0034<0.0010.002 First-person pronouns / 1k user wordsParticipant4.70671.73850.0070.014 Absolutist words / 1k user wordsParticipant0.53770.24710.0300.043 Disclosure levelConversation0.13750.0390<0.0010.002 Support seekingConversation0.03030.01090.0060.008 Sycophancy totalConversation0.02560.03060.4040.404 Uncritical agreementConversation0.05920.06610.3710.404 ObsequiousnessConversation0.02900.03360.3880.404 Professional redirectConversation0.02550.01430.0740.148 C.3 Subset Checks Within Health/Mental Health conversations, PHQ≥10 participants had higher disclosure levels (2.14 vs. 1.90,푑=0.31, 푞=0.0019), more high-disclosure conversations (34.6% vs. 26.1%,푑=0.31,푞=0.0019), and more support-seeking conversations (19.8% vs. 12.5%,푑=0.35,푞=0.0019). In the random non-health disclosure subset, disclosure differences were smaller and did not survive the same correction, although support seeking remained higher at a very low base rate (1.6% vs. 0.6%,푞=0.025). Similarly, the random non-health sycophancy subset showed no reliable PHQ differences in total sycophancy or its reported subdimensions. Together, these suggest that the response-style patterns are tied to health/support contexts rather than a general tendency for ChatGPT to respond differently to higher-PHQ participants across all conversations. C.4 Survey-Anchored Recency Sensitivity Because the PHQ-8 asks about symptoms over the prior two weeks, we repeated the main participant-weighted contrasts using only conversations that occurred before Survey 2 completion and within 90, 60, 30, and 14 days of the survey timestamp. Appendix Table 4 reports coverage for each window, and Appendix Table 5 reports the corresponding PHQ-split contrasts. Appendix Table 4. Coverage for survey-anchored recency windows. WindowConversationsParticipantsPHQ< 10PHQ≥ 10Median conv./participant All exported history187,093766571195105.5 Past 90 days42,65772254018231.0 Past 60 days29,17071453318121.0 Past 30 days15,76368651117512.0 Past 14 days7,7116384751637.0 Note. Recency windows are anchored to Survey 2 completion time and include conversations before that timestamp. Manuscript submitted to ACM 14Sehgal et al. Appendix Table 5. Survey-anchored recency sensitivity for headline PHQ-split contrasts. WindowMeasure푛 <10 푛 ≥10 PHQ< 10 PHQ≥ 10푑푝푞 All historyMental-health share5711953.1%5.7%0.36<0.001<0.001 Nocturnal use, 23:00–04:595711959.7%14.2%0.36<0.001<0.001 Interpersonal-problem share5711952.2%4.4%0.37<0.001<0.001 Loneliness/isolation share5711950.8%2.2%0.38<0.001<0.001 Negative self-focus share5711950.7%2.0%0.45<0.001<0.001 First-person pronouns / 1k user words57119525.3230.830.290.0020.003 Absolutist words / 1k user words5711953.684.330.220.0090.010 Advice/information ratio5501880.210.340.210.0110.011 Past 90dMental-health share5401823.2%6.5%0.38<0.001<0.001 Nocturnal use, 23:00–04:595401829.7%14.2%0.310.0010.002 Interpersonal-problem share5401822.1%5.1%0.40<0.0010.001 Loneliness/isolation share5401820.7%2.3%0.41<0.001<0.001 Negative self-focus share5401820.7%2.2%0.43<0.0010.001 First-person pronouns / 1k user words54018228.1436.670.37<0.001<0.001 Absolutist words / 1k user words5401823.984.820.160.1030.103 Advice/information ratio4931660.230.440.270.0160.021 Past 60dMental-health share5331813.2%6.4%0.35<0.0010.001 Nocturnal use, 23:00–04:595331819.6%14.0%0.280.0020.003 Interpersonal-problem share5331812.1%5.3%0.39<0.0010.002 Loneliness/isolation share5331810.8%2.4%0.36<0.0010.001 Negative self-focus share5331810.8%2.2%0.41<0.0010.001 First-person pronouns / 1k user words53318129.0436.320.31<0.0010.002 Absolutist words / 1k user words5331813.994.650.130.2040.204 Advice/information ratio4791630.250.530.220.1120.126 Past 30dMental-health share5111753.3%6.4%0.300.0040.010 Nocturnal use, 23:00–04:595111759.1%14.3%0.310.0020.010 Interpersonal-problem share5111752.5%5.3%0.310.0050.010 Loneliness/isolation share5111750.9%2.2%0.270.0070.012 Negative self-focus share5111750.9%2.2%0.300.0040.010 First-person pronouns / 1k user words51117531.1436.840.210.0180.027 Absolutist words / 1k user words5111754.164.840.100.2650.265 Advice/information ratio4351490.280.510.190.1360.153 Past 14dMental-health share4751633.2%6.7%0.280.0120.030 Nocturnal use, 23:00–04:594751639.4%14.5%0.270.0130.030 Interpersonal-problem share4751632.4%5.9%0.310.0050.030 Loneliness/isolation share4751630.9%3.2%0.320.0130.030 Negative self-focus share4751630.9%3.4%0.300.0230.042 First-person pronouns / 1k user words47516332.3936.940.150.0980.148 Absolutist words / 1k user words4751634.094.850.130.2120.239 Advice/information ratio3801310.320.28 -0.030.6790.679 Note. Each row is participant-weighted. Percent rows report mean participant-level conversation shares. Lexical rows report counts per 1,000 user word-like tokens.푞values are Benjamini-Hochberg adjusted within each recency window for this sensitivity table. Advice/information ratio excludes participants with no information-seeking conversations in the relevant window, so its participant counts are smaller. D DLA Appendix: Key LIWC and N-Gram Features Appendix Table 6 reports the key DLA features used in Finding 4. LDA100 features were tested, but no LDA topic features survived FDR correction. Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories15 Appendix Table 6. LIWC and 1-gram DLA features for PHQ≥10. The table lists the LIWC and 1-gram features used to interpret the DLA finding, separately for user text, ChatGPT response text, and combined text. Text sliceFeature familyDirectionRankFeatureDLA rp UserLIWCHigher in PHQ≥ 10 1SHEHE0.186<0.001 UserLIWCHigher in PHQ≥ 10 2FEMALE0.162<0.001 UserLIWCHigher in PHQ≥ 10 3PRONOUN0.156<0.001 UserLIWCHigher in PHQ≥ 10 4PPRON0.1500.001 UserLIWCHigher in PHQ≥ 10 5LINGUISTIC0.1360.003 UserLIWCHigher in PHQ≥ 10 6FUNCTION0.1210.012 UserLIWCHigher in PHQ≥ 10 7NEGATE0.1170.013 UserLIWCHigher in PHQ≥ 10 8FOCUSPAST0.1170.013 UserLIWCHigher in PHQ≥ 10 9EMO_ANX0.1160.013 UserLIWCHigher in PHQ≥ 10 10I0.1160.013 UserLIWCHigher in PHQ< 10 1LIFESTYLE-0.1270.008 UserLIWCHigher in PHQ< 10 2WORK-0.1120.014 UserLIWCHigher in PHQ< 10 3REWARD-0.0940.048 User1-gramHigher in PHQ≥ 10 1her0.1880.009 User1-gramHigher in PHQ≥ 10 2sure0.1740.036 ChatGPTLIWCHigher in PHQ≥ 10 1PRONOUN0.203<0.001 ChatGPTLIWCHigher in PHQ≥ 10 2IPRON0.193<0.001 ChatGPTLIWCHigher in PHQ≥ 10 3PPRON0.173<0.001 ChatGPTLIWCHigher in PHQ≥ 10 4WANT0.160<0.001 ChatGPTLIWCHigher in PHQ≥ 10 5ADVERB0.155<0.001 Manuscript submitted to ACM 16Sehgal et al. Text sliceFeature familyDirectionRankFeatureDLA rp ChatGPTLIWCHigher in PHQ≥ 10 6VERB0.1480.001 ChatGPTLIWCHigher in PHQ≥ 10 7CERTITUDE0.1430.001 ChatGPTLIWCHigher in PHQ≥ 10 8SHEHE0.1370.002 ChatGPTLIWCHigher in PHQ≥ 10 9YOU0.1360.002 ChatGPTLIWCHigher in PHQ≥ 10 10NEGATE0.1350.002 ChatGPTLIWCHigher in PHQ< 10 1LIFESTYLE-0.1390.001 ChatGPTLIWCHigher in PHQ< 10 2WORK-0.1240.004 ChatGPTLIWCHigher in PHQ< 10 3 FOCUSFUTURE-0.1140.008 ChatGPTLIWCHigher in PHQ< 10 4DRIVES-0.0900.045 ChatGPT1-gramHigher in PHQ≥ 10 1things0.230<0.001 ChatGPT1-gramHigher in PHQ≥ 10 2really0.217<0.001 ChatGPT1-gramHigher in PHQ≥ 10 3wants0.214<0.001 ChatGPT1-gramHigher in PHQ≥ 10 4trying0.213<0.001 ChatGPT1-gramHigher in PHQ≥ 10 5it0.207<0.001 ChatGPT1-gramHigher in PHQ≥ 10 6something0.205<0.001 ChatGPT1-gramHigher in PHQ≥ 10 7t0.205<0.001 ChatGPT1-gramHigher in PHQ≥ 10 8say0.201<0.001 ChatGPT1-gramHigher in PHQ≥ 10 9thing0.201<0.001 ChatGPT1-gramHigher in PHQ≥ 10 10don0.200<0.001 User + ChatGPTLIWCHigher in PHQ≥ 10 1PRONOUN0.197<0.001 User + ChatGPTLIWCHigher in PHQ≥ 10 2IPRON0.188<0.001 Manuscript submitted to ACM Depression Symptoms and Relational Patterns in 187k ChatGPT Histories17 Text sliceFeature familyDirectionRankFeatureDLA rp User + ChatGPTLIWCHigher in PHQ≥ 10 3PPRON0.171<0.001 User + ChatGPTLIWCHigher in PHQ≥ 10 4WANT0.156<0.001 User + ChatGPTLIWCHigher in PHQ≥ 10 5FEMALE0.1470.001 User + ChatGPTLIWCHigher in PHQ≥ 10 6SHEHE0.1460.001 User + ChatGPTLIWCHigher in PHQ≥ 10 7ADVERB0.1450.001 User + ChatGPTLIWCHigher in PHQ≥ 10 8VERB0.1420.001 User + ChatGPTLIWCHigher in PHQ≥ 10 9NEGATE0.1380.001 User + ChatGPTLIWCHigher in PHQ≥ 10 10CERTITUDE0.1350.002 User + ChatGPTLIWCHigher in PHQ< 10 1LIFESTYLE-0.1390.001 User + ChatGPTLIWCHigher in PHQ< 10 2WORK-0.1220.005 User + ChatGPTLIWCHigher in PHQ< 10 3FOCUSFUTURE-0.1100.010 User + ChatGPTLIWCHigher in PHQ< 10 4WE-0.0960.029 User + ChatGPTLIWCHigher in PHQ< 10 5DRIVES-0.0950.030 User + ChatGPT1-gramHigher in PHQ≥ 10 1things0.217<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 2someone0.214<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 3t0.206<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 4don0.205<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 5it0.202<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 6something0.199<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 7draining0.199<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 8feel0.199<0.001 Manuscript submitted to ACM 18Sehgal et al. Text sliceFeature familyDirectionRankFeatureDLA rp User + ChatGPT1-gramHigher in PHQ≥ 10 9wanting0.199<0.001 User + ChatGPT1-gramHigher in PHQ≥ 10 10yourself0.198<0.001 Note. DLA features are participant-level correlations with the PHQ split. Positivervalues indicate features more common in the PHQ≥10 group; negative values indicate features more common in the PHQ< 10 group. Only features used to interpret the paper’s DLA finding are shown. Manuscript submitted to ACM