Paper deep dive
Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling
Nadia Mehjabin, Henry Kautz, Subigya Nepal
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/14/2026, 4:32:12 AM
Summary
This study analyzes 369 journal entries from an 8-week passive sensing study using an LLM-based classification pipeline to determine which behaviors respond to AI journaling nudges. Key findings indicate that behavioral improvement depends primarily on social dependence rather than individual controllability; behaviors requiring other people (e.g., calls, social gatherings) improved in only 15-22% of cases, while individually controllable behaviors improved in 50-63%. Additionally, within intention entries, longer length, lower type-token ratio (focused vocabulary), and higher first-person singular usage were associated with follow-through, particularly for text messaging.
Entities (8)
Relation Signals (7)
MindScape → uses → LLM
confidence 95% · MindScape exemplifies this direction, combining passive sensing with large language models
MindScape → uses → Passive Sensing
confidence 95% · MindScape exemplifies this direction, combining passive sensing with large language models
Social Dependence → negativelypredicts → Behavioral Improvement
confidence 92% · Behaviors that depend on others improved in only 15 to 22% of cases
SMS → exhibitshighresponsiveness → AI Nudges
confidence 90% · For text messaging, the most immediately actionable behavior, intention language was linked to improvement
Phone Calls → exhibitslowresponsiveness → AI Nudges
confidence 90% · Behaviors that depend on others improved in only 15 to 22% of cases... phone call features (25 to 31%)
Elaboration Quality → positivelycorrelateswith → Follow-Through
confidence 88% · longer, more personal intention entries... tracked short-term follow-through
Individual Controllability → positivelypredicts → Behavioral Improvement
confidence 85% · behaviors a person can act on alone improved more often, up to 50 to 63%
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal entries from an eight-week passive sensing study. An LLM labeled each entry as expressing an intention to change a behavior or not, and we measured follow-through against 26 sensor features with a 3-day before/after comparison. Responsiveness depended most on whether a behavior involves other people. Behaviors that depend on others improved in only 15 to 22% of cases, while behaviors a person can act on alone improved more often, up to 50 to 63%, though unevenly. How users wrote mattered less. No single text feature separated improved from unimproved entries; writing carried signal only within specific behaviors, most clearly for text messaging and for longer, more personal intention entries. The sample is small, so we treat these as exploratory patterns that point to where AI journaling nudges are most likely to work.
Tags
Links
- Source: https://arxiv.org/abs/2608.12582v1
- Canonical: https://arxiv.org/abs/2608.12582v1
Trouble viewing inline? Open PDF directly →
Full Text
34,325 characters extracted from source content.
Expand or collapse full text
Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling Nadia Mehjabin Department of Computer Science University of Virginia Charlottesville, USA jqc7gj@virginia.edu Henry Kautz Department of Computer Science University of Virginia Charlottesville, USA rmw7my@virginia.edu Subigya Nepal Department of Computer Science University of Virginia Charlottesville, USA sknepal@virginia.edu Abstract—AI journaling tools can tailor prompts to a person’s own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal entries from an eight-week passive sensing study. An LLM labeled each entry as expressing an intention to change a behavior or not, and we measured follow- through against 26 sensor features with a 3-day before/after com- parison. Responsiveness depended most on whether a behavior involves other people. Behaviors that depend on others improved in only 15 to 22% of cases, while behaviors a person can act on alone improved more often, up to 50 to 63%, though unevenly. How users wrote mattered less. No single text feature separated improved from unimproved entries; writing carried signal only within specific behaviors, most clearly for text messaging and for longer, more personal intention entries. The sample is small, so we treat these as exploratory patterns that point to where AI journaling nudges are most likely to work. Index Terms—passive sensing, contextual journaling, mobile sensing, behavior change, large language models, mental health, mobile health, student well-being, human–AI interaction I. INTRODUCTION Behavioral sensing systems that integrate passive data col- lection with AI-generated prompts or nudges represent an emerging approach to support health behavior change. Passive sensing has been used to personalize health behavior interven- tions, and large language models (LLMs) have been applied to support journaling and self-reflection. Recent work has begun combining these two approaches, using continuous sensing data to generate LLM-based journaling prompts grounded in users’ own behavioral context. We conducted an 8-week study of the MindScape system with undergraduate students and demonstrated significant gains in well-being [1]. Yet, our system distributes prompts equally across all behavioral do- mains based on sensing data availability rather than behavioral responsiveness. In an initial analysis of the MindScape [1] dataset, we manually examined journal entries for statements of intent and compared them to subsequent behavioral sensing data. Several cases showed clear alignment: a participant who wrote about wanting to walk more showed increased walking episodes in the following days; another who mentioned reaching out to friends showed higher SMS counts in the same period. Accepted and presented at the HARMONY Workshop, IEEE/ACM Con- ference on Connected Health: Applications, Systems and Engineering Tech- nologies (CHASE 2026) These patterns were observable but not measurable; we had no systematic way to identify intention entries at scale, and no method to statistically connect journal content to sensing outcomes across the full dataset. This led us to build an LLM-based classification pipeline for this task, paired with a before/after sensing comparison to test follow-through at scale. Applying this pipeline to 369 journal interactions across 26 sensing features, we find that responsiveness depends mostly on whether a behavior involves other people: behaviors that depend on social context or other people improve in only 15 to 22% of cases, while behaviors a person can act on alone improve more often (up to 50 to 63%) but inconsistently. For text messaging, the most immediately actionable behavior, intention language was linked to improvement in a small set of entries. Within intention entries, richer elaboration was also linked to follow-through. We make two contributions. First, a responsiveness map across the 26 sensing features, in which dependence on other people – more than individ- ual controllability – separates responsive from unresponsive behaviors. Second, descriptive evidence that, in this sample, two things track short-term follow-through: intention language for individually controllable behaviors such as SMS, and richer elaboration within intention entries. Together, these help designers decide where to target nudges and what in a user’s response signals likely follow-through. I. RELATED WORK Journaling supports self-awareness [2], [3], emotional pro- cessing [4], and cognitive organization [5], with evidence showing benefits for mood, stress reduction, and mental well- being [5]–[7]. A key recent advance has been the shift from generic journaling prompts toward context-aware prompting grounded in users’ own behavioral data. MindScape [1] ex- emplifies this direction, combining passive sensing with large language models to deliver personalized journaling prompts, demonstrating meaningful well-being gains over 8 weeks. Similar systems including MindfulDiary [8] and Reflection Companion [9] have shown that AI-mediated journaling can support consistent reflective habits. Related work on daily planning prompts has likewise shown that prompting users to form concrete plans can shape routine behaviors[10]. However, MindScape distributes prompts across behavioral arXiv:2608.12582v1 [cs.HC] 12 Aug 2026 domains based on sensing availability, without accounting for whether those domains are equally responsive to nudges. A large-scale meta-analysis of over 200 studies found that while nudges produce reliable behavior change overall, effectiveness varies substantially by domain [11]. Personal informatics re- search has likewise found that self-monitoring and reflection do not affect all health behaviors equally. Frameworks for mo- bile behavior change stress that responsiveness depends on the behavior, the moment, and the person together [12]. Physical activity responds reliably to step-count feedback and goal- setting nudges, while sleep and social behaviors are harder to shift through brief app-based interventions [13]. Bhattacharjee et al. [14] further showed that behaviors embedded in social coordination are less amenable to individual-level prompting than behaviors under personal control. To the best of our knowledge, no prior study has empirically characterized which domains respond to journaling nudges and which do not. Even when a nudge reaches the right domain, whether it produces change depends on how the user responds. Goll- witzer’s work on implementation intentions showed that speci- ficity is the critical mediator: concrete plans predict follow- through more reliably than vague expressions of intent [15]. Prior journaling work has analyzed content for sentiment and linguistic patterns [1], [8], but has not connected the quality of users’ responses to subsequent behavioral outcomes. Englhardt et al. [16] showed that LLMs can reason over mobile behavioral data to generate clinical insights, but to our knowledge LLM-based intention classification has not been paired with before/after sensing validation. The pipeline developed in this paper addresses both gaps. I. METHODS This study is a secondary analysis of the MindScape longi- tudinal study [1]. Full details of the study design, participant recruitment, app architecture and prompt generation pipeline are described in our prior papers [1], [17]. We summarize the aspects most relevant to our analysis here. A. Dataset The MindScape study enrolled 20 undergraduate students at Dartmouth College over 8 weeks. During Weeks 1 through 6, participants received contextual AI-generated prompts derived from their passive sensing data. During Weeks 7 through 8, participants received generic prompts not grounded in behavioral data. Our analysis focuses on the first 6 weeks of contextual journaling, yielding 20 participants and 369 journal entries. B. Behavioral Signal and Improvement Operationalization For each of the four behavioral domains, physical fitness, sleep, digital habits, and social interaction, we identified sensing features representing the full 24-hour daily aggregate as the primary behavioral signal. The features included per domain are shown in Table I. For each journal entry, we identified the behavioral features most relevant to the journal content using the following TABLE I PASSIVE SENSING FEATURES PER BEHAVIORAL DOMAIN DomainFeaturesN Digital HabitsCommunication, Entertainment, and Social app use; Screen unlock duration and count 5 Social InteractionCall duration/count (In/Out), Conversation du- ration/count (Food venues, Home), Time at spe- cific locations (Greek spaces, Dorms, Study), SMS count (In/Out) 16 SleepSleep duration1 Physical FitnessGym/workout duration; Cycling, Walking, and Running episode durations 4 Total26 procedure. First, we searched the journal response text for explicit or implicit references to a sensing feature. For ex- ample, a mention of communication or texting was mapped to app Communication and smsinnum. If no feature was identifiable from the response, we examined the prompt text to determine the behavioral category and associated features. When multiple features were referenced, we selected the feature with the maximum absolute signal change as the primary improvement indicator for that entry. To measure whether a journal entry corresponded to behav- ioral change, we computed the mean value of each sensing feature over the 3 days before and 3 days after each entry. We evaluated 1, 3 and 7-day windows and selected the 3- day window as the primary analysis window. The overall improvement rates across the three windows are 43.8% (1- day), 45.8% (3-day), and 47.3% (7-day), a difference of only 3.5 percentage points across the full range. Domain-level rates are similarly stable, meaning the choice of window does not meaningfully change the pattern of results. We then assigned a binary improvement flag: 1 if the 3- day post-entry mean exceeded the 3-day pre-entry mean in the direction of improvement for that feature, 0 otherwise. The improvement direction was feature-specific and domain- informed: for example, we coded a decrease in phone usage duration as improvement for digital habits, while we coded an increase in walking duration as improvement for physical fitness. This yielded 26 features with valid improvement flags across 369 entries, which form the basis of all subsequent analyses. C. LLM Classification Pipeline Our initial exploration of the dataset used a time-series language model to detect intention language in journal en- tries at scale. That approach performed poorly, failing to reliably distinguish entries expressing behavioral intentions from general reflections. This motivated the development of a two-stage LLM classification pipeline (Fig.1) using llama-3.3-70b-versatile. Stage 1 - Prompt classification: Each system prompt was classified as REDIRECTIVE (explicitly targeting a behavioral domain and inviting the user to set a goal or plan) or RE- FLECTIVE (open-ended, inviting general reflection without a specific behavioral target). Stage 2 - Response classification: Journal entry + prompt Prompt classification REDIRECTIVEREFLECTIVE Stage 1 all entries Response classification INTENTIONNON-INTENTIONNOT APPLICABLE Stage 2 all entries Fig. 1. Two-stage LLM classification pipeline: Stage 1 classifies each sys- tem prompt as REDIRECTIVE (behavior-targeted) or REFLECTIVE (open- ended); Stage 2 classifies each journal response as INTENTION (a stated plan or commitment), NON-INTENTION (reflection without commitment), or NOT APPLICABLE. Each journal entry paired with its prompt was classified as INTENTION (the user expressed a plan, goal, or commitment to change a behavior), NON-INTENTION (the user reflected on past behavior without committing to change), or NOT APPLICABLE (the entry did not engage with the prompt’s behavioral focus). To validate the pipeline, we independently labeled a random sample of 30 entries, obtaining an aggregate agreement rate of approximately 70% across prompt and response classifications. This study is exploratory, and telling reflective from redirective prompts (and intention from non-intention responses) is partly subjective. We treat this agreement as adequate to proceed, note it as a limitation and encourage future work to build purpose-built annotation schemes. D. Linguistic Feature Extraction To examine whether the linguistic properties of journal en- tries predict behavioral follow-through, we extracted features at three levels. Surface features: We computed word count and type-token ratio (TTR) for each entry. TTR measures lex- ical diversity: a lower TTR in intention entries indicates more focused, repetitive vocabulary centered on a specific behavior, which we hypothesized would be associated with follow- through. Syntactic features: Using dependency parsing and morphological analysis, we extracted first-person singular pro- noun count (I, me, my, myself) as a proxy for personal agency and commitment, and future orientation as a measure of forward-looking language. LLM semantic scores: Each entry was scored by the LLM on three dimensions: behavioral concreteness (1 to 4, from no specific behavior mentioned to specific behavior with clear timing), planning depth (1 to 3, from awareness only to concrete plan), and emotional engage- ment (1 to 3, from detached to strong personal investment). We used these scores to check whether the INTENTION versus NON-INTENTION classification reflected qualitative differences in entry content, and whether semantic quality within INTENTION entries further predicted improvement. The full rubric and prompt are in the Appendix (Table I). IV. RESULTS A. Pipeline Validation As shown in Table I, REDIRECTIVE prompts elicited INTENTION responses in 49.5% of cases, compared to only 9.3% for REFLECTIVE prompts (Fisher’s exact p < 0.01). This pattern is expected, since redirective prompts explicitly invite users to set goals or make plans while reflective prompts invite open-ended reflection. Because the classifier picks up this difference, it appears to capture a real linguistic distinc- tion. This makes the INTENTION versus NON-INTENTION comparisons that follow meaningful. TABLE I RESPONSE TYPE DISTRIBUTION BY PROMPT TYPE. REDIRECTIVE PROMPTS ELICIT INTENTION RESPONSES AT A SIGNIFICANTLY HIGHER RATE THAN REFLECTIVE PROMPTS (FISHER’S EXACT P< 0.0001). Prompt typeINTENTIONNON-INTENTIONIntention rate REDIRECTIVE504849.5% REFLECTIVE252239.3% B. Finding 1: Social Dependence, More Than Individual Con- trollability, Predicts Responsiveness Improvement rates in 26 sensing features (see Table IV) reveal a pattern that is clearest at the extremes and tied most reliably to whether a behavior depends on other people. Behaviors primarily driven by individual decisions tended to improve more often: incoming SMS (62.9%), walking episodes (58.0%), and time spent at home (53.6%) improved in over half of cases. Behaviors involving some scheduling friction or environmental dependency improved in 40 to 50% of cases: phone usage duration (48.2%), running (43.5%), and communication app usage (42.5%). Behaviors most dependent on social coordination or physical co-presence with others rarely improved: phone call features (25 to 31%), food venue conversations (22%), and Greek space visits (15%). The low end is the most consistent part of this pattern: behaviors that require other people’s coordination, such as calls, food-venue conversations and Greek-space visits, were uniformly among the least responsive. The high end is nois- ier, and individual controllability did not by itself predict improvement. Among behaviors a person performs alone, walking improved in 58.0% of cases but gym/workout time in 20.0%, entertainment-app use in 18.8%, and cycling in 12.5% — near or below the socially dependent behaviors. Texting shows the same split: incoming SMS improved in 62.9% of cases but outgoing SMS in only 38.5%. We therefore read this result less as a clean controllability gradient than as a single robust asymmetry: behaviors that depend on other people are consistently unresponsive to journaling nudges, whereas behaviors under individual control vary widely and are not explained by controllability alone. Because we grouped behaviors by judgment rather than from a separate measure of controllability, and several features rest on small samples, we treat the ordering as descriptive. C. Finding 2: Intention Language Is Associated With Out- comes for Highly Responsive Behavior in a Small Exploratory Sample As noted earlier, an overall test across all entries found no link between text features and improvement. So we treat this subsection and the next as exploratory patterns within specific behaviors, meant to generate hypotheses. Among individual sensing features, SMS communication showed the clearest relationship between response type and behavioral outcome. For incoming SMS (or text messages), entries classified as INTENTION were followed by improvement in 100% of cases (8 of 8), compared to 52.2% for NON-INTENTION entries (Fisher’s exact p = 0.028). These 8 entries came from 7 different participants, so the result is not just one person repeated. Still, it rests on very few entries and should be read with caution. Outgoing SMS showed a directionally consistent pattern, with INTENTION entries exhibiting larger behavioral change than NON-INTENTION entries (Mann-Whitney U, p = 0.046). SMS is a plausible case for this association because it is immediately actionable: deciding to reach out to someone and sending a message can happen in the same moment as writing a journal entry. This sits at the most responsive end of the behaviors observed in the prior finding. We note, however, outgoing SMS is arguably more directly controllable than incoming SMS, as the sender can initiate contact unilaterally. The stronger result for incoming SMS may mean that journal- ing about social intentions makes people more open to others’ messages — responding and being available — rather than starting contact themselves. Both directions showed consistent patterns (incoming: Fisher’s exact p = 0.028; outgoing: Mann- Whitney U p = 0.046), though the small sample warrants caution in any mechanistic interpretation. D. Finding 3: Elaboration Quality Within Intentions Is Asso- ciated With Follow-Through Intention language alone is not sufficient to predict follow- through. Within INTENTION entries specifically, the quality of elaboration is further associated with whether behavioral improvement occurs. Restricting analysis to INTENTION en- tries only (improved: n = 27, not improved: n = 44), three linguistic features differentiated the two groups. Improved intention entries were longer (Fig.2, median word count 45 vs. 25.5, Mann-Whitney p = 0.012, r = 0.21). They also had lower type-token ratios indicating more focused vocabulary centered on a specific behavior (TTR median 0.79 vs. 0.85, p = 0.014, r =−0.26). But since improved entries were also longer, this lower diversity partly just reflects length, because TTR falls as texts get longer. They also used more first-person singular language (Fig. 2, median count 5 vs. 3, p<0.01). Short, vague entries were not linked to behavioral change; longer, more personal ones were. LLM semantic scoring showed that INTENTION entries differ from NON-INTENTION entries in behavioral con- creteness, planning depth, and emotional engagement (all p < 0.05), consistent with the classifier capturing qualitative differences in content. Within INTENTION entries, though, these scores did not separate improved from not-improved entries. This suggests the signal lies in how much someone elaborates and in their personal voice, not in the topic or category. V. DISCUSSION A. Relationship to the Original MindScape Findings In the main MindScape paper, we asked whether AI jour- naling improves well-being. In this secondary analysis, we ask which interactions within it drive change. Across all 369 entries, no single text feature predicted behavioral improve- ment (all p > 0.10), and intention and non-intention entries did not differ between the improved and unimproved groups (χ 2 p = 0.15). This interaction-level null fits the study-level well-being gains that the main paper reported [1]. Participants self- reported those gains, which built up over eight weeks of steady journaling, while we measure short-term behavioral change in passive sensing around each entry. A behavior does not have to move in the three days after an entry for journaling to help over two months. Our analysis adds resolution. The interactions that most often preceded near-term change involved behaviors a user can carry out alone and intentions the user elaborated, while behaviors that depend on other people rarely changed. The study-level result leaves this pattern implicit. B. Theoretical Grounding The most robust part of our result, that socially dependent behaviors resist change, lines up with two established bodies of theory. Self-determination theory [18] distinguishes between autonomously regulated behaviors: those driven by internal motivation and executable through individual decision, and those requiring external coordination. Food-venue conversa- tions and Greek-space events require other people’s presence and cooperation, and these externally coordinated behaviors were consistently the least responsive, as the theory would pre- dict. The autonomous end was less clean: walking responded strongly, but other self-directed behaviors such as cycling did not, so the theory explains the unresponsive floor better than it predicts the responsive ceiling. The elaboration finding echoes Gollwitzer’s work on imple- mentation intentions [15]: specific, concrete plans are more reliably followed through than vague expressions of intent. Our word count and first-person singular results mirror that pattern in naturalistic journaling data, extending a laboratory finding into a real-world mobile health context. C. Design and Research Implications These findings suggest several directions for both practi- tioners and researchers. ImprovedNot Improved 20 40 60 80 100 Word Count p=0.012 Word Count in INTENTION Entries ImprovedNot Improved 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Type-Token Ratio p=0.014 Lexical Richness in INTENTION Entries ImprovedNot Improved 0 2 4 6 8 10 12 Count p=0.005 First-Person Singular in INTENTION Entries Fig. 2. Word count and lexical richness (TTR) within INTENTION entries, split by behavioral outcome (improved n = 27, not improved n = 44). Improved entries are significantly longer (median 45 vs. 25.5, p = 0.012) and show lower TTR indicating more focused, repetitive vocabulary centered on a specific behavior (median 0.79 vs. 0.85, p = 0.014). Account for social dependence when targeting nudges. The most dependable design signal is negative: behaviors that depend on other people’s presence, such as calls, in-person conversations, group settings, were consistently unresponsive to journaling nudges, so systems should not expect prompts alone to move them. Behaviors a person can act on alone are better candidates, but not a guarantee: in our data, walking and incoming messaging responded well while other self-directed behaviors, such as cycling and entertainment-app use, did not. Which individually controllable behaviors respond, and why, is a question for prospective study. Detect and act on intention language. The pipeline proposed in this study can operate at inference time: when a user writes an intention response to a redirective prompt about a controllable behavior, that signal may be worth acting on, logging it as a commitment, triggering a follow-up, or building the next prompt on the stated plan. Follow up on short, vague intentions. Intention entries under∼25 words were not associated with behavioral change in this sample. A brief next-day follow-up asking whether the user acted could help close the gap between a passing thought and a committed plan, and is detectable from word count alone without additional LLM inference. Design for consistency, not per-prompt perfection. The aggregate pattern is the most consequential design signal in this paper. Aside from the exploratory, behavior-specific pat- terns above, no single interaction reliably predicted behavioral change at the aggregate level; sustained engagement over weeks did. For practitioners, features encouraging regular jour- naling habits such as streaks, gentle reminders, low-friction entry may matter more than optimizing any individual prompt. For researchers, it reframes the unit of analysis: the right question may not be what makes a good prompt, but what keeps users journaling long enough for cumulative effects to emerge. D. Limitations This work has several limitations. The sample is small and drawn from a single institution, limiting generalizability; all findings should be treated as exploratory rather than defini- tive. The SMS finding rests on 8 intention entries and the elaboration finding on 71, so both lack the power for strong conclusions. The controllability framing is interpretive and partly post-hoc; per-feature rates are reported descriptively and alternative orderings are possible. In particular, individual controllability does not cleanly order the per-feature rates: several self-directed behaviors (cycling, entertainment apps, outgoing SMS) rank among the least responsive, so the account holds mainly for the socially dependent low end. The 3-day binary window for measuring behavioral change is empirically justified but cannot capture timing or non- monotonic change, and response timescales likely vary by domain and individual. The correlational nature of passive sensing data means we cannot establish causality [19]; sensing reliability, behavioral inertia and academic calendar effects likely contribute. Finally, the LLM classification pipeline achieved about 70% agreement with manual labels on a 30- entry sample, introducing classification noise that can attenuate observed effects. VI. CONCLUSION Not all behaviors respond equally to AI journaling nudges, and not all intentions translate into behavioral change. This paper offers an early empirical look at where journaling-based nudges are most likely to produce proximal behavioral follow- through, and what in a user’s response signals that follow- through is likely. The findings are exploratory and limited by a small single-institution sample, but they suggest that how much a behavior depends on other people, along with elaboration quality, is a meaningful dimension for designing and evaluating AI health interventions. The next step is a prospective study that deliberately targets individually con- trollable behaviors, to test whether choosing the right domains improves outcomes beyond steady engagement alone. REFERENCES [1] S. Nepal, A. Pillai, W. Campbell, T. Massachi, M. V. Heinz, A. Kunwar, E. S. Choi, X. Xu, J. Kuc, J. F. Huckins et al., “Mindscape study: inte- grating llm and behavioral sensing for personalized ai-driven journaling experiences,” Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies, vol. 8, no. 4, p. 1–44, 2024. [2] D. Alt and N. Raichel, “Reflective journaling and metacognitive aware- ness: Insights from a longitudinal study in higher education,” Reflective Practice, vol. 21, no. 2, p. 145–158, 2020. [3] G. B. Williams, M. B. Gerardi, S. L. Gill, M. D. Soucy, and D. H. Taliaferro, “Reflective journaling: Innovative strategy for self-awareness for graduate nursing students,” International Journal of Human Caring, vol. 13, no. 3, p. 36–43, 2009. [4] J. M. Smyth, J. A. Johnson, B. J. Auer, E. Lehman, G. Talamo, and C. N. Sciamanna, “Online positive affect journaling in the improvement of mental distress and well-being in general medical patients with elevated anxiety symptoms: A preliminary randomized controlled trial,” JMIR Mental Health, vol. 5, no. 4, p. e11290, Dec. 2018. [5] M. Sohal, P. Singh, B. S. Dhillon, and H. S. Gill, “Efficacy of journaling in the management of mental illness: a systematic review and meta- analysis,” Family medicine and community health, vol. 10, no. 1, p. e001154, 2022. [6] K. N. Keech and P. G. Coberly-Holt, “Journaling for mental health,” in Strategies and Tactics for Multidisciplinary Writing. IGI Global, 2021, p. 39–44. [7] W. R. Miller, “Interactive journaling as a clinical tool,” Journal of mental health counseling, vol. 36, no. 1, p. 31–42, 2014. [8] T. Kim, S. Bae, H. A. Kim, S. woo Lee, H. Hong, C. Yang, and Y.-H. Kim, “Mindfuldiary: Harnessing large language model to support psy- chiatric patients’ journaling,” arXiv preprint arXiv:2310.05231, 2023. [9] R. Kocielnik, L. Xiao, D. Avrahami, and G. Hsieh, “Reflection compan- ion: a conversational system for engaging users in reflection on physical activity,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 2, no. 2, p. 1–26, 2018. [10] A. Cuadra, O. Bankole, and M. Sobolev, “Planning habit: daily plan- ning prompts with alexa,” in International Conference on Persuasive Technology. Springer, 2021, p. 73–87. [11] S. Mertens, M. Herberz, U. J. Hahnel, and T. Brosch, “The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains,” Proceedings of the National Academy of Sciences, vol. 119, no. 1, p. e2107346118, 2022. [12] F. Okeke, M. Sobolev, and D. Estrin, “Towards a framework for mobile behavior change research,” in Proceedings of the technology, mind, and society, 2018, p. 1–6. [13] D. Bakker and N. Rickard, “Engagement in mobile phone app for self- monitoring of emotional wellbeing predicts changes in mental health: Moodprism,” Journal of affective disorders, vol. 227, p. 432–442, 2018. [14] A. Bhattacharjee, D. Kulzhabayeva, M. Reza, H. Kumar, E. Seong, X. Wu, M. R. Rifat, R. Bowman, R. Kornfield, A. Mariakakis et al., “Integrating individual and social contexts into self-reflection technolo- gies,” in Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems, 2023, p. 1–6. [15] P. M. Gollwitzer, “Implementation intentions: Strong effects of simple plans.” American psychologist, vol. 54, no. 7, p. 493, 1999. [16] Z. Englhardt, C. Ma, M. E. Morris, C. Chang, X. O. Xu, L. Qin, D. McDuff, X. Liu, S. Patel, and V. Iyer, “From classification to clinical insights: Towards analyzing and reasoning about mobile and behavioral health data with large language models,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 8, no. 2, p. 1–25, May 2024. [17] S. Nepal, A. Pillai, W. Campbell, T. Massachi, E. S. Choi, X. Xu, J. Kuc, J. F. Huckins, J. Holden, C. Depp et al., “Contextual ai journaling: Integrating llm and time series behavioral sensing technology to promote self-reflection and well-being using the mindscape app,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2024, p. 1–8. [18] E. L. Deci and R. M. Ryan, “The” what” and” why” of goal pursuits: Human needs and the self-determination of behavior,” Psychological inquiry, vol. 11, no. 4, p. 227–268, 2000. [19] S. Zhu, H. Zhang, J. D. Chi, S. Nepal, and K. Saha, “Causal stories from sensor traces: Auditing epistemic overreach in llm-generated personal sensing explanations,” arXiv preprint arXiv:2605.08590, 2026. APPENDIX A LLM SCORING RUBRIC TABLE I LLM PROMPT USED FOR SEMANTIC FEATURE EXTRACTION. You are analyzing journal entries from a behavioral wellness study. For each entry, respond ONLY with a valid JSON object. No preamble, no markdown. Rate the entry on these dimensions: BEHAVIORAL CONCRETENESS (1–4): 1 = no specific behavior; 2 = general category; 3 = specific with some context; 4 = specific with clear timing PLANNINGDEPTH (1–3): 1 = awareness only; 2 = vague desire to change; 3 = concrete plan EMOTIONALENGAGEMENT (1–3): 1 = detached/factual; 2 = some reflection; 3 = strong personal investment Output:"behavioral_concreteness": <1-4>, "planning_depth": <1-3>, "emotional_engagement": <1-3> APPENDIX B FEATURE-LEVEL IMPROVEMENT RANKING TABLE IV FEATURE-LEVEL BEHAVIORAL IMPROVEMENT RANKING ACROSS ALL 26 SENSING FEATURES, RANKED BY IMPROVEMENT RATE. FeatureN Improvement Rate (%) FeatureN Improvement Rate (%) smsinnum3562.9appsocial3231.3 act walking5058.0locselfdormdur1428.6 loc homedur2853.6calloutduration2528.0 lochomeconvodur6048.3locfoodstill1526.7 unlock duration5648.2callinduration2326.1 lochomeconvonum6147.5locotherdormdur2725.9 sleepduration3844.7calloutnum3925.6 actrunning2343.5locfoodconvonum5822.4 locstudydur2842.9locfoodconvodur5422.2 appcommunication7342.5locworkoutdur2520.0 unlocknum2642.3appentertainment1618.8 smsoutnum3938.5locgreekdur2015.0 call innum3531.4actonbike812.5 APPENDIX C EXAMPLE JOURNAL ENTRIES BY CLASSIFICATION TYPE Prompt Type Examples REDIRECTIVE: (1) Curious, with gym time up but overall activity less, how might incorporating a quick outdoor walk boost your day and connections? (2) Your social interactions have lessened recently. Can you think of a person you haven’t caught up with in a while to reconnect with? How might that feel? REFLECTIVE: (1) With your sleep start time increasing, consider adjusting your evening routine for better rest. What change could you make tonight? (2) You’ve been clocking less screen time lately. What have you been doing instead that you’ve found rewarding or enjoyable? Response Type Examples INTENTION: (1) This would be good. Texting is exhausting and removes me from the present moment. (2) I think by putting my phone in my pocket and sitting by my friends letting them talk. I’m always afraid of missing out something online but I should focus more on what’s in front of me sometimes. NON-INTENTION: (1) I could probably find ways to stop the scroll, but I just feel so mentally exhausted that I fall into traps of doom scrolling. (2) I was carrying my phone in my hand a lot today, so the phone kept getting unlocked. I have been connected to friends in other places more recently through my phone. No effect on sleep. My sleep is erratic because of workload and finals week.