Paper deep dive
Depictions of Depression in Generative AI Video Models: A Preliminary Study of OpenAI's Sora 2
Matthew Flathers, Griffin Smith, Julian Herpertz, Zhitong Zhou, John Torous
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/23/2026, 12:06:02 PM
Summary
This study investigates how OpenAI's Sora 2 generative video model depicts 'depression' by comparing outputs from its consumer App and developer API. The research reveals that while both access points rely on a shared visual vocabulary (e.g., hoodies, windows, rain), the consumer App exhibits a significant 'recovery bias'âfeaturing brighter, more dynamic, and resolution-oriented narrativesâwhereas the API outputs are more static and somber. The findings suggest that platform-level product layers, rather than just the underlying model, significantly shape the mental health narratives presented to users.
Entities (5)
Relation Signals (3)
Developer API â providesaccessto â Sora 2
confidence 100% · examine whether depictions differ between the consumer App and developer API access points
Sora 2 â depicts â Depression
confidence 95% · This study characterizes how OpenAI's Sora 2 generative video model depicts depression
Consumer App â exhibitsbias â Recovery Bias
confidence 95% · App-generated videos exhibited a pronounced recovery bias
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generative video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions like depression. This study characterizes how OpenAI's Sora 2 generative video model depicts depression and examines whether depictions differ between the consumer App and developer API access points. We generated 100 videos using the single-word prompt "Depression" across two access points: the consumer App (n=50) and developer API (n=50). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Computational features across visual aesthetics, audio, semantic content, and temporal dynamics were extracted and compared between modalities. App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution, compared with 14% (7/50) of API outputs. App videos brightened over time (slope = 2.90 brightness units/second vs. -0.18 for API; d = 1.59, q < .001) and contained three times more motion (d = 2.07, q < .001). Across both modalities, videos converged on a narrow visual vocabulary and featured recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Figures were predominantly young adults (88% aged 20-30) and nearly always alone (98%). Gender varied by access point: App outputs skewed male (68%), API outputs skewed female (59%). Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, while platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge, and that patients may encounter such content during vulnerable periods.
Tags
Links
- Source: https://arxiv.org/abs/2603.19527v1
- Canonical: https://arxiv.org/abs/2603.19527v1
Trouble viewing inline? Open PDF directly â
Full Text
93,100 characters extracted from source content.
Expand or collapse full text
Abstract Background: Generative video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions like depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance. Objective: This study aimed to characterize how OpenAIâs Sora 2 generative video model depicts depression, and to examine whether depictions differ between the consumer App and developer API access points, which differ in their product-layer mediation. Methods: We generated 100 videos using the single-word prompt âDepressionâ across two access points: the consumer App (n=50) and developer API (n=50). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Inter-rater reliability was assessed using Cohenâs kappa, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, temporal dynamics) were extracted and compared between modalities using Welchâs t-tests with Benjamini-Hochberg false discovery rate correction. Results: App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution, compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (slope = 2.90 brightness units/second vs. -0.18 for API; d = 1.59, q < .001) and contained three times more motion (d = 2.07, q < .001). Across both modalities, videos converged on a narrow visual vocabulary: predominantly seated figures (93%), downward gaze (96%), and recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Transcript language in depressive phases emphasized weight and containment (âheavy,â âdrowning,â âroomâ), while recovery phases reversed these patterns: brightness increased 27% (d = 0.70, p < .001), gaze shifted upward in 68% of recovery videos, and terms like âlightâ and âbreathâ emerged. Figures were predominantly young adults (88% aged 20-30) and nearly always alone (98%). Gender varied by access point: App outputs skewed male (68%), API outputs skewed female (59%). Conclusions: Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, while platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge, and that patients may encounter such content during vulnerable periods. Depictions of Depression in Generative AI Video Models: A Preliminary Study of OpenAIâs Sora 2 Matthew Flathers1, Griffin Smith2, Julian Herpertz1,3, Zhitong Zhou1,4, John Torous5 1Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA, USA 2Rhode Island School of Design, Providence, RI, USA 3Department of Psychiatry and Neuroscience, CharitĂ© Berlin University Medicine, Campus Benjamin Franklin, Berlin, Germany 4Boston University, School of Public Health, Boston, MA, USA 5Department of Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA, USA Keywords: artificial intelligence; generative AI; depression; mental health representation; text-to-video; computer vision 1 Introduction Media representations of psychiatric conditions shape how the public understands mental illness, and how patients understand their own diagnoses [60, 2, 56, 22]. Increasingly, these representations are circulating through short-form video (SFV) platforms. Mental health has emerged as one of the most prevalent topics on TikTok, Instagram Reels, and YouTube Shorts, with hashtags like #mentalhealth accumulating over 44 billion views on TikTok alone [42]. SFV platforms have become embedded in daily life, with roughly six-in-ten U.S. teens visiting TikTok daily and one-third using at least one major platform almost constantly [14, 30]. Against this backdrop, generative artificial intelligence has introduced new capabilities for creating custom SFV content at scale. AI-generated videos can now be produced much faster and more cheaply than traditional video editing pipelines. OpenAIâs Sora 2, released in September 2025, is an early leader in this space [44]. Within 72 hours of launch, the Sora mobile application became the top-downloaded app on the iOS App Store, surpassing ChatGPT with over 164,000 downloads [46]. The app functions as a social network for AI-generated videos, featuring a vertical feed of 10-second clips with likes, comments, and remix tools [44]. In December 2025, OpenAI announced a partnership with Disney that will bring over 200 characters from Disney, Marvel, Pixar, and Star Wars to the platform [54]. OpenAIâs policies prohibit use by children under 13 [45], licensing characters from franchises with young audiences creates obvious pressure on these boundaries. As these capabilities expand beyond entertainment, AI videos will increasingly reach people who are seeking to understand their own mental health. Sora 2 is not alone. Multiple labs have released generative AI video models that can produce highly realistic outputs that viewers cannot always reliably distinguish from authentic footage [57]. Googleâs Veo 3 pairs high-fidelity video generation with native audio generation [21], Runwayâs Gen-4.5 provides sophisticated video editing capabilities [50], Meta has launched Vibes to integrate AI video generation into its social ecosystem [40], and ByteDanceâs Seedance demonstrates high-fidelity multi-shot video generation with consistent subject representation across scenes [20]. Most generative video tools build upon diffusion image model architectures, much like AI Image models. Unlike transformer-based language models that predict and generate words sequentially, diffusion models learn to transform noise into coherent imagery. They do this through a two-phase process: systematically adding noise to images until they become visual static (forward diffusion), then learning to reverse this process to reconstruct coherent images [52, 25]. Through training on billions of examples, these models learn the contours, patterns, and statistical regularities of their training data, storing this information across vast parameter spaces. Video diffusion models extend this into higher-dimensional space, maintaining consistency across hundreds of frames while generating plausible motion and transitions [26, 51]. The outputs of these systems occupy a unique categorical space [49]. The underlying models train on heterogeneous sources simultaneously: clinical documentation, Hollywood films, pharmaceutical advertising, video blogs and other user-generated content. The resulting synthesis reveals aggregate cultural understanding rather than any single authoritative source. When a user prompts a model with âDepression,â the resulting artifact is not summoned from nothing, nor is it copied from some original human source. The userâs prompt is mediated by a technical apparatus that constrains what can be produced to what its programming makes possible [17]. The output is âfoundâ by the AI model in the diffusion process, selected from the space of possible outputs that training made probable. Whether the patterns learned during training extend beyond what things look like to capture implicit associations, co-occurrences, and cultural conventions embedded in the training data (as has been observed in image [16], text [41], and embedding [7] models) remains an open question this study begins to address. This study examines the Sora 2 generative AI video system, but our methodological approach addresses a broader issue in clinical AI research. Much existing work treats âthe modelâ as the object of study without distinguishing between 1) the underlying models and 2) the products and consumer-facing apps/tools built upon them. When a user submits a prompt into the Sora app, they interact with a system that includes the model plus product layers. When a researcher accesses the same model through the API, fewer of these layers intervene (See Figure 1 for a generalized illustration of this distinction). Patterns observed in an app may therefore reflect product-level decisions (e.g. more safety features and filters) rather than model-level associations, or vice versa. Studying only the product tells us what users encounter but not what the model has learned; studying only the model tells us what the training has encoded but not what reaches users [15]. Comparing both access points reveals something about the safety and content moderation layers that companies have implemented for sensitive topics like depression. The distinction between model and product is also important for research on safety evaluation within the wider AI product ecosystem. Future developers will build applications on top of AI models accessed via APIs, not the consumer product layer. Sora 2 provides an opportunity to examine these dynamics because OpenAI has made both access points available: the consumer app and the developer API. Figure 1: Generalized system architecture for generative video AI platforms, illustrating the distinction between Developer API and Consumer App access pathways. Both pathways share core safety infrastructure (center), including input and output moderation. Consumer Apps additionally route prompts through an application layer (prompt processing, UX constraints, policy-aware transformations) before generation, and through a product surface layer (feed curation, watermarking, distribution controls) after. The Developer API bypasses these layers, accessing only the shared safety infrastructure and base model. When someone experiencing a depressive episode searches for content about depression and is served AI-generated videos on platforms like Sora 2, those outputs become part of their informational environment during a vulnerable period. The content might reinforce their experience, normalize or stigmatize it, or offer implicit models of recovery or self-harm [22, 34, 29]. Whether AI-generated depictions of depression draw on clinical understanding, popular media tropes, pharmaceutical advertising or some mixture remains an empirical question. The answer carries implications for how clinicians think about the information environment surrounding their patients. 2 Methods 2.1 Data Generation We generated 50 videos using each of two access points to OpenAIâs Sora 2 system: the consumer-facing mobile app and the developer API, yielding a total corpus of 100 videos. All videos were generated in portrait orientation (9:16 aspect ratio) with each platformâs standard output parameters: 10-second clips at 720p resolution from the app and 12-second clips at 720p resolution using the Sora 2 endpoint (sora-2) via the OpenAI Responses API. To control for potential model updates, all generations occurred within a single one-week window (November 21â28, 2025). All app videos were generated from a new Sora account with no prior video generation history. Each video was generated using the isolated prompt âDepressionâ with no additional qualifiers, contextual framing, or prompt engineering. Our minimal prompt design was a deliberate methodological choice. The prompt âDepressionâ is inherently ambiguous. It might refer to a clinical mood disorder, a colloquial feeling, a topological feature, or an economic phenomenon. Elaborate prompts resolve this ambiguity for the model; minimal prompts force the model to resolve it from its encoded associations. A prompt like âa depressed person sitting alone in a dark room cryingâ tests whether a model can render a specified scene but reveals little about what the model associates with depression as a concept, because the user has already supplied the interpretive content. Single-term prompts shift this interpretive labor to the model. When given only âDepressionâ with no contextual elaboration, the model must supply setting, figures, mood, color, motion, and narrative entirely from its encoded associations. Every visual choice in the output therefore reflects training data or product-level steering rather than user instruction, making minimal prompting a more sensitive probe of the modelâs default representations. 2.2 Qualitative Analysis Two authors (MF and Z) independently conducted a qualitative content analysis of all 100 videos using a directed approach, in which initial coding categories were derived from prior literature and the studyâs research questions [27]. Each coder tagged videos across six dimensions: (1) narrative arc (presence and direction of emotional shifts), (2) environments (physical settings with timestamps), (3) objects (visual symbols and props with timestamps), (4) figure demographics (gender, race/ethnicity, age), (5) figure states (posture, facial expression, alone in frame with timestamps), and (6) overall notes (AI artifacts and uncertainties). Raw codes were harmonized to standardized categories prior to analysis. Extended coding procedures and dimension rationale appear in Appendix A; the full codebook appears in Appendix B; and harmonization dictionaries appear in Appendix C. 2.2.1 Reliability Both authors independently coded the full dataset. One authorâs coding (MF) was designated as the primary dataset for analysis; the second authorâs (Z) served as inter-rater reliability verification. Cohenâs kappa (Îș) was calculated for presence/absence of categorical elements, and for dimensions involving temporal segments, we also calculated intersection-over-union (IoU) at one-second resolution to assess agreement on timing. Given the scale of the dataset, item-by-item consensus resolution was not feasible; reliability metrics verified that the coding scheme could be applied consistently across coders. 2.3 Quantitative Analysis To complement qualitative coding, we extracted computational features across four domains: visual aesthetics, audio properties, semantic content, and temporal dynamics. Visual features included brightness, saturation, colorfulness, color temperature, and GLCM texture measures [23], sampled at one-second intervals. Audio analysis extracted volume, pitch, and spectral centroid using the librosa library [38]. Semantic analysis of transcribed speech (OpenAI Whisper [47]) included VADER sentiment [28], custom light/dark lexicons based on conceptual metaphor theory [32], and linguistic markers including first-person pronoun ratio and negation frequency. Temporal dynamics were captured through dense optical flow (FarnebĂ€ck algorithm [13]), scene detection, and per-second feature trajectories. Features were selected based on documented connections to depression representation in prior research; technical specifications and literature rationale appear in Appendix D. 2.4 Ethical Considerations This study analyzed AI-generated video content produced by a commercially available system (OpenAI Sora 2) and did not involve human participants, human biological materials, or identifiable personal data. No human subjects were recruited, surveyed, or observed, and no personally identifiable information was collected or analyzed. Accordingly, institutional review board review was not required under federal regulations (45 CFR 46) or institutional policy. All analyzed content was generated by the research team using publicly available AI tools. 2.5 Statistical Analysis For each quantitative feature, we computed the mean and standard deviation for each access modality. Differences between app and API outputs were assessed using Welchâs t-tests [58], with Benjamini-Hochberg false discovery rate correction [3] to control for multiple comparisons across all 28 features tested. Effect sizes were quantified using Cohenâs d [10]. To characterize temporal dynamics, we computed the linear slope of each featureâs time series and compared slope distributions between modalities using the same inferential approach. 3 Results 3.1 Quantitative Findings We generated 50 videos from each access point (App and API), yielding a corpus of 100 videos for analysis. All comparisons used Welchâs t-tests with Benjamini-Hochberg false discovery rate (FDR) correction; we report FDR-corrected q-values and Cohenâs d effect sizes throughout. Table 1: Aggregate descriptive statistics. Mean ± SD for all visual, texture, motion, audio, speech, sentiment, and semantic features across App (n = 50) and API (n = 50) generated videos. Cohenâs d effect sizes and FDR-corrected q-values from Welchâs t-tests with Benjamini-Hochberg correction. * indicates q < .05. Category Feature App (M ± SD) API (M ± SD) d q Visual Brightness Mean 64.59 ± 13.35 53.48 ± 11.86 0.87 0.000* Saturation Mean 99.24 ± 26.61 92.48 ± 23.08 0.27 0.342 Colorfulness Mean 25.67 ± 7.48 22.41 ± 8.01 0.42 0.086 Color Temperature Mean â-13.38 ± 10.56 â-14.45 ± 9.66 0.10 0.716 Edge Density Mean 0.03 ± 0.01 0.03 ± 0.01 â-0.31 0.239 Texture GLCM Contrast 12.15 ± 5.77 7.78 ± 5.11 0.79 0.001* GLCM Dissimilarity 1.20 ± 0.33 1.12 ± 0.31 0.22 0.453 GLCM Homogeneity 0.72 ± 0.05 0.70 ± 0.05 0.51 0.032* GLCM Energy 0.17 ± 0.04 0.17 ± 0.04 â-0.03 0.908 GLCM Correlation 0.96 ± 0.02 0.97 ± 0.02 â-0.53 0.025* GLCM Entropy 6.98 ± 0.53 6.94 ± 0.51 0.09 0.730 Motion Optical Flow Mean 0.35 ± 0.15 0.11 ± 0.07 2.07 0.000* Scene Cut Count 3.10 ± 1.88 0.76 ± 1.42 1.39 0.000* Audio Audio Volume Mean 0.10 ± 0.01 0.09 ± 0.01 0.83 0.000* Audio Pitch Mean 115.51 ± 30.04 148.87 ± 40.78 â-0.92 0.000* Spectral Centroid Mean 1744.08 ± 395.02 1792.82 ± 284.42 â-0.14 0.633 Speech Word Count 31.40 ± 5.70 36.50 ± 8.01 â-0.73 0.001* Speech Rate (words/sec) 4.18 ± 0.64 4.07 ± 0.54 0.19 0.552 Sentiment Sentiment Compound 0.19 ± 0.33 0.24 ± 0.36 â-0.15 0.627 Sentiment Positive 0.09 ± 0.05 0.10 ± 0.07 â-0.17 0.554 Sentiment Negative 0.04 ± 0.05 0.04 ± 0.04 â-0.07 0.780 Sentiment Neutral 0.87 ± 0.06 0.86 ± 0.08 0.18 0.552 Semantic Light Word Count 0.02 ± 0.14 0.16 ± 0.37 â-0.50 0.035* Dark Word Count 0.58 ± 0.67 0.64 ± 0.52 â-0.10 0.716 Light/Dark Word Ratio 0.27 ± 0.27 0.27 ± 0.29 0.00 1.000 First Person Pronoun Ratio 0.04 ± 0.04 0.08 ± 0.05 â-0.78 0.001* Negation Ratio 0.02 ± 0.03 0.03 ± 0.03 â-0.12 0.701 Lexical Diversity 0.90 ± 0.05 0.89 ± 0.05 0.22 0.453 3.1.1 Visual Aesthetics App-generated videos were significantly brighter than API videos (M = 64.59 ± 13.35 vs. 53.48 ± 11.86; d = 0.87, q < .001). This difference emerged over time: App videos exhibited a pronounced brightening trajectory (slope = 2.90 ± 2.43 brightness units/second) while API videos remained essentially flat (slope = â-0.18 ± 1.24; d = 1.59, q < .001). Figure 2 illustrates this divergence, showing App videos beginning at comparable brightness levels to API videos but progressively lightening across their duration. Figure 2: Brightness trajectory over time. Mean brightness (0â255 scale) at each second for App (light blue, n = 50) and API (dark blue, n = 50) generated videos. Shaded regions indicate ± 1 SD. App videos exhibit a pronounced brightening trajectory (slope = 2.90 units/second), while API videos remain relatively flat (slope = â-0.18). The difference in brightness slope was statistically significant (d = 1.59, q < .001, Welchâs t-test with BH FDR correction). Chromatic properties showed complementary patterns. Although mean saturation did not differ significantly between modalities (q = .34), saturation trajectories diverged substantially: App videos desaturated over time (slope = â-2.99 ± 3.17), while API videos maintained stable saturation (slope = 0.24 ± 1.30; d = â-1.33, q < .001). Similarly, App videos warmed in color temperature over time (slope = 2.24 ± 2.26) compared to minimal change in API videos (slope = 0.18 ± 0.85; d = 1.21, q < .001). Together, these trajectories describe a characteristic App visual arc: beginning in cool, saturated tones and transitioning toward warm, desaturated, brighter imagery. Texture analysis revealed limited additional differences. App videos exhibited higher GLCM contrast (M = 12.15 ± 5.77 vs. 7.78 ± 5.11; d = 0.79, q < .001) and homogeneity (d = 0.51, q = .03), indicating sharper local intensity variation and more uniform textural regions. Other texture features remained largely consistent across modalities. 3.1.2 Temporal Dynamics The most pronounced differences emerged in motion characteristics. App videos contained approximately three times more optical flow than API videos (M = 0.35 ± 0.15 vs. 0.11 ± 0.07 pixels/frame; d = 2.07, q < .001), indicating substantially greater movement within scenes (Figure 3a). This difference persisted across the full video duration, with App videos maintaining elevated motion throughout rather than concentrating movement in particular segments. Scene structure differed dramatically between modalities. App videos contained a mean of 3.10 ± 1.88 scene cuts, while API videos contained only 0.76 ± 1.44 cuts (d = 1.39, q < .001). Figure 3b visualizes this difference as cumulative scene accumulation over time: by the end of App videos, viewers had encountered approximately 4 distinct scenes on average, compared to fewer than 2 for API videos. Figure 3: Temporal dynamics of motion and editing in App vs. API videos. (A) Motion intensity: Mean optical flow magnitude (pixels/frame; FarnebĂ€ck dense optical flow) at each second for App (light blue; n = 50) and API (dark blue; n = 50) videos. Shaded bands indicate ± 1 SD. App videos exhibit substantially higher motion across the clip duration (overall M = 0.35 vs. 0.11; d = 2.07, q < .001), indicating greater within-scene dynamism than the relatively static API outputs. (B) Scene accumulation: Mean cumulative scene number at each second, where Scene 1 denotes the opening scene and each cut increments the count by one. Shaded bands indicate ± 1 SD. App videos accrue scenes more rapidly, reaching 4 distinct scenes by the end versus <2 for API videos, consistent with higher scene-cut frequency (App M = 3.10 cuts vs. API M = 0.76; d = 1.39, q < .001). 3.1.3 Audio and Speech Audio properties diverged across modalities. App videos were slightly louder (M = 0.096 ± 0.009 vs. 0.087 ± 0.013; d = 0.83, q < .001) but featured lower-pitched audio (M = 115.51 ± 30.04 Hz vs. 148.87 ± 40.78 Hz; d = â-0.92, q < .001). Lower pitch in App videos may reflect the predominance of male narrators, different ambient sound profiles, or compositional choices in generated music. Speech content also differed. API videos used first-person pronouns at nearly twice the rate of App videos (M = 0.078 ± 0.048 vs. 0.044 ± 0.039; d = â-0.78, q < .001). API videos also contained more light-associated words (M = 0.16 ± 0.37 vs. 0.02 ± 0.14; d = â-0.50, q = .03). Sentiment valence did not differ significantly between modalities (q = .63), with both producing mildly positive compound sentiment scores (App: M = 0.19 ± 0.34; API: M = 0.24 ± 0.37). 3.1.4 Semantic Content Word frequency analysis revealed a coherent thematic vocabulary across the corpus (Figure 4). The most frequent terms clustered around emotional experience (âfeel,â n = 43; âheavy,â n = 25), temporal struggle (âday,â n = 28; âmorning,â n = 10), natural metaphors (âstorm,â n = 8; ârain,â n = 8; âcloud,â n = 8), and hope or release (âlight,â n = 13; âbreathe,â n = 8). Figure 4: Word frequency comparison. Twenty most frequent content words appearing in transcribed speech across App and API generated videos. Common terms cluster around themes of emotional experience (feel, heavy, weight), temporal struggle (day, morning), natural metaphors (storm, rain, cloud), and hope or release (light, breathe). This vocabulary reflects culturally prevalent metaphors for depression emphasizing weight, weather, darkness, and the possibility of relief. 3.2 Qualitative Findings 3.2.1 Inter-Rater Reliability Two coders independently coded all 100 videos. Reliability was strong for most dimensions: narrative arc (Îș = 0.69), environments (Îș = 0.91, IoU = 0.84), posture (Îș = 0.89), alone-in-frame (Îș = 0.94), and gender (Îș = 1.00). Object coding showed moderate overall agreement (Îș = 0.61) with variation by object type. Dimensions with insufficient reliability were excluded from analysis: facial expression (Îș = 0.49), race/ethnicity (Îș = 0.53), and apparent SES (Îș = 0.20). For objects, lower kappa values for ubiquitous items like beds (Îș = 0.44) and windows (Îș = 0.47) reflected asymmetric exhaustiveness between coders rather than substantive disagreement about object presence; the primary coder tagged these common items more exhaustively while the secondary coder focused on more distinctive elements. Full reliability results appear in Appendix E. 3.2.2 Narrative and Visual Patterns The most striking qualitative finding concerned narrative structure. After collapsing to binary classification (recovery vs. no recovery, combining deterioration and no shift), a dramatic divergence emerged between modalities. App videos frequently depicted recovery narratives with 39 of 50 videos (78%) containing a discernible shift toward hope or relief. API videos showed the opposite pattern: only 7 of 50 videos (14%) depicted recovery arcs. This aligns with quantitative trajectory analysis showing App videos brightening over time while API videos remained flat. Figure 5: Representative video stills from App and API outputs generated with the prompt âDepression.â (A) Consumer App video showing a typical recovery arc. The video progresses from a dark interior with a personal storm cloud (1 s), through floating debris labeled with terms like âfatigueâ and âhopelessnessâ (3â5 s), to the figure turning toward a bright window (7 s) and ending outdoors in sunlight, smiling (9 s). The coded recovery shift occurs at approximately 6 seconds. (B) Developer API video showing typical stasis. The figure remains seated on a bed in a dim bedroom throughout the full 12-second duration, wearing a grey hoodie with downward gaze, clasped hands, and minimal change in posture, lighting, or environment. Both videos were generated with identical single-word prompts and no additional parameters. Across both modalities, videos converged on a narrow visual vocabulary. Bedroom was the dominant setting, with App videos featuring somewhat more diverse environments including elevated outdoor spaces and urban settings. The most common objects were hoodies (n = 194), windows (148), beds (114), rain (83), and lamps (59). Generated figures were predominantly young adults aged 20â30 (88%) and nearly always alone in frame with downward gazeâ which in recovery videos often shifted upward. App videos showed greater physical mobility (66% containing standing, walking, or floating) compared to API videos (12%). Gender differed by modality: App outputs skewed male (68%), API outputs skewed female (59%). 3.2.3 Recovery Transition Patterns To examine whether recovery patterns extended beyond coded objects and environments, we analyzed transcript words and quantitative visual features before versus after coded recovery timestamps (n = 46 videos with recovery arcs). Analysis of environment appearance relative to recovery shifts (n = 46 videos) revealed systematic patterns. Environments appearing predominantly before recovery shifts included Bedroom (26 videos before vs. 3 after), Bathroom (7 vs. 0), and Urban Outdoor (12 vs. 4). Elevated outdoor settings (rooftops, hilltops, bridges) showed the opposite pattern (5 videos before vs. 8 after), appearing predominantly after recovery shifts. Objects followed parallel trajectories. Beds (32 vs. 1), rain (23 vs. 1), windows (32 vs. 8), and curtains (12 vs. 2) clustered in depressive phases, while sunlight and sunrise appeared predominantly after recovery shifts (3 vs. 17), as did birds (1 vs. 7). In terms of figure posture, gaze shifted from downward to upward in 68% of recovery videos compared to 16% of non-recovery videos. Transcript vocabulary reinforced these patterns. Words appearing predominantly before recovery timestamps included âheavyâ (n = 21), âfeelsâ (n = 25), âheavierâ (n = 10), âroomâ (n = 8), âworldâ (n = 14), ârainâ (n = 7), and âstormâ (n = 7). Words appearing predominantly after recovery timestamps included âlightâ (n = 12), âsmallâ (n = 7), âcloudsâ (n = 5), âbreathâ (n = 5), âcracksâ (n = 4), and âmomentâ (n = 4). Depression-phase vocabulary clustered around weight/pressure metaphors (heavy, heavier, pressing, chest) and water/confinement imagery (underwater, drowning, room), while recovery-phase vocabulary emphasized light, smallness, and atmospheric imagery. Quantitative visual features showed corresponding patterns. Comparing mean values before versus after recovery timestamps: brightness increased by 27.4% (d = 0.70, p < .001), saturation decreased by 13.5% (d = â-0.42, p < .001), and colorfulness increased by 10.8% (d = 0.28, p = .002). Motion showed no significant change (p = .98). Figure 6: Recovery-linked content shifts across objects, settings, language, posture, and low-level visual features. Panels AâD show recovery differentials comparing element presence (or rates) after vs. before the coded recovery timestamp in videos with recovery arcs (46 videos). Recovery differential was computed as (afterâ-before)/(after+before); positive values (light blue) indicate elements appearing predominantly after the recovery shift, and negative values (dark blue) indicate elements appearing predominantly before the shift. Bar labels indicate the number of videos containing each coded element (A, B, D) or total word instances (C). (A) Objects and (B) Environments meeting inclusion thresholds (n â„ 5) are shown; objects were additionally selected for thematic relevance. (C) Transcript words were chosen a priori by investigators for semantic relevance to depression and recovery themes and plotted using recovery differential based on before/after word rates. (D) Posture codes show recovery-linked changes in embodied state (e.g., lying down vs. standing/active positions). (E) Visual features summarize changes in saturation, colorfulness, brightness, and color temperature, expressed as effect size (Cohenâs d) for after-versus-before differences; positive values indicate increases and negative values indicate decreases (significance markers as shown). 4 Discussion Our findings indicate that depictions of depression generated by the Sora 2 App depend on more than the underlying model. How that model is embedded, constrained, and circulated in the product/app has a direct impact on the content, aesthetics, and narrative preferences present in outputs; in ways that diverge from base model behaviors. When prompted with the isolated diagnostic term âDepression,â the consumer-facing Sora app overwhelmingly produces recovery-oriented narratives, while the same model accessed through the API yields largely static, unresolved scenes. 78% of app-generated videos contained a discernible shift toward hope or relief, compared with only 14% of API outputs. The contrast between product/app and API outputs indicates that product-layer mediation shapes how depression is represented. The Sora app positions video generation within a social feed, governed by community guidelines that prohibit content perceived as promoting depression or depicting self-harm [44]. Whether this recovery shift bias stems from these explicit content safeguards, or emerges indirectly from platform dynamics that favor narrative progression and engagement cannot be determined from our analysis. Regardless of mechanism, our analysis suggests that the Sora 2 product layer introduces strong narrative constraints. Within this environment, outputs are disproportionately oriented toward recovery as a safe and legible endpoint. Clinically, this suggests that model outputs can be steered towards specific behaviors, and that clinical teams can play a role in improving AI video tools at the product level. Our results are consistent with broader concerns that generative systems may favor culturally overrepresented narrative endpoints (e.g., happy endings, resolved arcs, completed transformations) when responding to ambiguous prompts rather than sustaining unresolved middle states. Writers working with language models have described this tendency toward narrative overfitting [48], noting in particular a bias toward positive narrative endings [53]. In video generation, this bias manifests as a rapid exit from stasis into motion, light, elevation, and relief. This pattern may stem from training objectives, or it may encode the dominance of the ârestitution narrativeâ in human-generated content: the culturally preferred illness storyline in which sickness is a temporary interruption resolved through recovery [19]. This dynamic is not absent in the APIâs outputs, but it is substantially amplified when the model operates within the Sora 2 App, which is effectively a social distribution pipeline optimized for engagement, shareability, and moderation. Importantly, the visual and narrative grammars employed in these videos are not novel. The withdrawn figure in a dim bedroom, the rain-streaked window, and the gray hoodie all echo iconographies of melancholia traceable from Renaissance allegory and Enlightenment-era psychiatric imaging through to modern media depictions, consistent with patterns observed in text-to-image models [16]. The visual grammar also aligns with existing research on depression metaphors that identify darkness, descent, weight, and containment as near-universal conceptual frameworks [8, 39, 31]. Our findings seem to confirm these patternsâ persistence into AI video models. The dominance of solitary, seated figures, the clustering of beds, rain, and windows in depressive phases, and the systematic reversal toward light, birds, and elevation during recovery all make use of established symbols and motifs from media. Our findings suggest that Sora 2 does not invent new ways of depicting depression; instead, it compresses and recombines historical visual conventions and routes them through contemporary platform logics that privilege certain endpoints and user behaviors over others. Clinicians should be aware that AI-generated content about depression reflects training data and platform design rather than clinical expertise, and may need to help patients contextualize the simplified narratives and homogeneous imagery they encounter. For researchers and policymakers, these findings reiterate a methodological imperative: analyses of generative AI must attend to both base models, and the media forms and distribution systems through which AI model outputs are encountered by public userbases. This study has several limitations. The analysis is restricted to a single platform (Sora 2) and a single prompt term, limiting generalizability to other systems or prompt formulations. The disparity in default clip length between access points (10 seconds for App, 12 seconds for API) introduces a potential confound. Where possible, we mitigated this by analyzing per-second rates and slopes rather than raw totals, but the difference cannot be fully disentangled from product-layer effects. Coders were not blinded to access modality, as the generation interfaces differed visibly, which may have introduced expectation effects in coding. Finally, the sample of 50 videos per modality was chosen to establish preliminary baseline observations rather than to power specific inferential comparisons; effect sizes and confidence intervals should be interpreted accordingly. 5 Conclusion This study provides the first systematic analysis of how the Sora 2 generative video model depicts depression, revealing that its outputs are shaped as much by product-layer mediation as by underlying model associations. The consumer Appâs pronounced recovery bias (with 78% of videos progressing toward hope compared to only 14% of API outputs) demonstrates that platform constraints actively reshape which mental health narratives reach users. Across both access points, the visual and linguistic grammar proved remarkably consistent with centuries-old iconographic traditions: the solitary figure, the dim interior, the downward gaze, the metaphorics of weight and weather. These findings suggest that Sora 2 predominantly recirculates familiar cultural conventions for depicting depression rather than departing sharply from them. We expect that this pattern may well extend to other generative video models built on similar architectures and training data. As AI-generated content becomes increasingly prevalent in the informational environments of people experiencing psychiatric distress, clinicians should understand that these outputs reflect training data and platform constraints rather than clinical knowledge, and that they may shape patient expectations about recovery timelines and trajectories. Researchers and policymakers, in turn, should attend to the distinction between model behavior and product behavior when evaluating these systems. Acknowledgments Figure 1 was generated using Google Paper Banana, a research figure generative AI tool; all design decisions and content were specified by the authors. Claude Opus 4.6 (Anthropic) was used for proofreading and copy-editing during manuscript preparation. All scientific content, analysis, interpretation, and final text were authored and reviewed by the authors. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Conflicts of Interest None declared. Data Availability The analysis code and video dataset analyzed in this study are available from the corresponding author upon reasonable request. Authorsâ Contributions MF: Conceptualization, Methodology, Software Development, Formal Analysis, Investigation, Data Curation, Writing - Original Draft, Writing - Review & Editing, Visualization, Project Administration. GS: Conceptualization, Methodology, Writing - Review & Editing. JH: Investigation, Writing - Review & Editing. Z: Investigation, Data Curation. JT: Supervision, Writing - Review & Editing. Abbreviations API application programming interface FDR false discovery rate GLCM gray-level co-occurrence matrix IoU intersection over union SES socioeconomic status SFV short-form video VADER Valence Aware Dictionary and Sentiment Reasoner References [1] T. Amorese, M. Cuciniello, C. Greco, O. Sheveleva, G. Cordasco, C. Glackin, G. McConvey, Z. Callejas, and A. Esposito (2025) Detecting depression in speech using verbal behavior analysis: a cross-cultural study. Frontiers in Psychology 16, p. 1514918. Cited by: §D.3.3. [2] S. Armstrong, E. Osuch, M. Wammes, O. Chevalier, S. Kieffer, M. Meddaoui, and L. Rice (2025) Self-diagnosis in the age of social media: a pilot study of youth entering mental health treatment for mood and anxiety disorders. Acta Psychologica 256, p. 105015. Cited by: §1. [3] Y. Benjamini and Y. Hochberg (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 57 (1), p. 289â300. Cited by: §D.5.3, §2.5. [4] J. Brassine, J. Van den Eynde, T. R. Hubble, and J. Toelen (2020) Digital images of pediatric mental disorders do not accurately represent the conditions. Heliyon 6 (9). Cited by: Appendix A. [5] T. Brockmeyer, J. Zimmermann, D. Kulessa, M. Hautzinger, H. Bents, H. C. Friederich, W. Herzog, and M. Backenstrass (2015) Me, myself, and I: self-referent word use as an indicator of self-focused attention in relation to depression and anxiety. Frontiers in Psychology 6, p. 1564. Cited by: §D.3.3. [6] J. S. Buyukdura, S. M. McClintock, and P. E. Croarkin (2011) Psychomotor retardation in depression: biological underpinnings, measurement, and treatment. Progress in Neuro-Psychopharmacology and Biological Psychiatry 35 (2), p. 395â409. Cited by: §D.4. [7] A. Caliskan, J. J. Bryson, and A. Narayanan (2017) Semantics derived automatically from language corpora contain human-like biases. Science 356 (6334), p. 183â186. Cited by: §1. [8] J. Charteris-Black (2012) Shattering the bell jar: metaphor, gender, and depression. Metaphor and Symbol 27 (3), p. 199â216. Cited by: §4. [9] J. Cohen (1960) A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20 (1), p. 37â46. Cited by: §D.6.1. [10] J. Cohen (1988) Statistical power analysis for the behavioral sciences. 2nd edition, Lawrence Erlbaum Associates, Hillsdale, NJ. Cited by: §D.5.2, §2.5. [11] N. Dael, M. N. Perseguers, C. Marchand, J. P. Antonietti, and C. Mohr (2016) Put on that colour, it fits your emotion: colour appropriateness as a function of expressed emotion. Quarterly Journal of Experimental Psychology 69 (8), p. 1619â1630. Cited by: §D.1.1. [12] T. F. Dehcheshmeh, A. S. Majelan, and B. Maleki (2024) Correlation between depression and posture (a systematic review). Current Psychology 43 (33), p. 27251â27261. Cited by: Appendix A. [13] G. FarnebĂ€ck (2003) Two-frame motion estimation based on polynomial expansion. In Scandinavian Conference on Image Analysis, Berlin, Heidelberg, p. 363â370. Cited by: §D.4.1, §2.3. [14] M. Faverio and O. Sidoti (2024) Teens, social media and technology 2024. Note: Pew Research CenterAccessed: 2026 Jan 15 External Links: Link Cited by: §1. [15] M. Flathers, B. Dwyer, E. Rozenblit, and J. Torous (2025) Contextualizing clinical benchmarks: a tripartite approach to evaluating LLM-based tools in mental health settings. Journal of Psychiatric Practice 31 (6), p. 294â301. Cited by: §1. [16] M. Flathers, G. Smith, E. Wagner, C. E. Fisher, and J. Torous (2024) AI depictions of psychiatric diagnoses: a preliminary study of generative image outputs in Midjourney v. 6 and DALL-E 3. BMJ Mental Health 27 (1). Cited by: §1, §4. [17] V. Flusser (1983) Towards a philosophy of photography. Reaktion Books, London. Cited by: §1. [18] C. Forceville and S. Paling (2021) The metaphorical representation of depression in short, wordless animation films. Visual Communication 20 (1), p. 100â120. Cited by: Appendix A, Appendix A, Appendix A. [19] A. W. Frank (2013) The wounded storyteller: body, illness, and ethics. 2nd edition, University of Chicago Press, Chicago. Cited by: §4. [20] Y. Gao, H. Guo, T. Hoang, W. Huang, L. Jiang, F. Kong, H. Li, J. Li, L. Li, X. Li, and X. Li (2025) Seedance 1.0: exploring the boundaries of video generation models. arXiv preprint. External Links: 2506.09113 Cited by: §1. [21] Google DeepMind (2025) Veo 3 model card. Note: Accessed: 2026 Jan 19 External Links: Link Cited by: §1. [22] I. Hacking (1995) The looping effects of human kinds. In Causal Cognition: A Multidisciplinary Debate, D. Sperber, D. Premack, and A. J. Premack (Eds.), p. 351â383. Cited by: §1, §1. [23] R. M. Haralick, K. Shanmugam, and I. H. Dinstein (1973) Textural features for image classification. IEEE Transactions on Systems, Man, and Cybernetics SMC-3 (6), p. 610â621. Cited by: §D.1.2, §2.3. [24] D. Hasler and S. E. SĂŒsstrunk (2003) Measuring colorfulness in natural images. In Human Vision and Electronic Imaging VIII, Vol. 5007, p. 87â95. Cited by: §D.1.1. [25] J. Ho, A. Jain, and P. Abbeel (2020) Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33, p. 6840â6851. Cited by: §1. [26] J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet (2022) Video diffusion models. Advances in Neural Information Processing Systems 35, p. 8633â8646. Cited by: §1. [27] H. F. Hsieh and S. E. Shannon (2005) Three approaches to qualitative content analysis. Qualitative Health Research 15 (9), p. 1277â1288. Cited by: §2.2. [28] C. Hutto and E. Gilbert (2014) VADER: a parsimonious rule-based model for sentiment analysis of social media text. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 8, p. 216â225. Cited by: §D.3.1, §2.3. [29] A. Kleinman (1980) Patients and healers in the context of culture: an exploration of the borderland between anthropology, medicine, and psychiatry. University of California Press. Cited by: §1. [30] A. Klin and D. Lemish (2008) Mental disorders stigma in the media: review of studies on production, content, and influences. Journal of Health Communication 13 (5), p. 434â449. Cited by: §1. [31] Z. Kövecses (2010) Metaphor, language, and culture. DELTA: Documentação de Estudos em LinguĂstica TeĂłrica e Aplicada 26, p. 739â757. Cited by: §4. [32] G. Lakoff and M. Johnson (1980) Metaphors we live by. University of Chicago Press, Chicago. Cited by: §D.3.2, §2.3. [33] J. R. Landis and G. G. Koch (1977) The measurement of observer agreement for categorical data. Biometrics 33 (1), p. 159â174. Cited by: §D.6.1. [34] H. Leventhal, L. A. Phillips, and E. Burns (2016) The Common-Sense Model of self-regulation (CSM): a dynamic framework for understanding illness self-management. Journal of Behavioral Medicine 39 (6), p. 935â946. Cited by: §1. [35] A. Malkomsen, J. I. RĂžssberg, T. Dammen, T. Wilberg, A. LĂžvgren, R. Ulberg, and J. Evensen (2021) Digging down or scratching the surface: how patients use metaphors to describe their experiences of psychotherapy. BMC Psychiatry 21 (1), p. 533. Cited by: §D.3.2. [36] I. MartĂnez-NicolĂĄs, D. Criado, F. Gordillo, F. MartĂnez-SĂĄnchez, and J. J. MeilĂĄn (2025) Speech analysis for detecting depression in older adults: a systematic review. Frontiers in Psychology 16, p. 1715538. Cited by: §D.2. [37] M. Mauch and S. Dixon (2014) pYIN: a fundamental frequency estimator using probabilistic threshold distributions. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 659â663. Cited by: Table 20. [38] B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto (2015) Librosa â audio processing Python library. In Proceedings of the 14th Python in Science Conference, p. 18â25. Cited by: §D.2, §2.3. [39] L. M. McMullen and J. B. Conway (2002) Conventional metaphors for depression. In The Verbal Communication of Emotions: Interdisciplinary Perspectives, S. R. Fussell (Ed.), p. 167â181. Cited by: §4. [40] Meta (2025) Introducing Vibes: a new way to discover and create AI videos. Note: Accessed: 2026 Jan 19 External Links: Link Cited by: §1. [41] J. Moore, D. Grabb, W. Agnew, K. Klyman, S. Chancellor, D. C. Ong, and N. Haber (2025) Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, p. 599â627. Cited by: §1. [42] M. Motta, Y. Liu, and A. Yarnell (2024) âInfluencing the influencers:â a field experimental approach to promoting effective mental health communication on TikTok. Scientific Reports 14 (1), p. 5864. Cited by: §1. [43] J. C. Mundt, A. P. Vogel, D. E. Feltner, and W. R. Lenderking (2012) Vocal acoustic biomarkers of depression severity and treatment response. Biological Psychiatry 72 (7), p. 580â587. Cited by: §D.2. [44] OpenAI (2025) Sora 2 system card. Note: Accessed: 2025 Dec 18 External Links: Link Cited by: §1, §4. [45] OpenAI (2025) Terms of use. Note: Accessed: 2026 Jan 23 External Links: Link Cited by: §1. [46] S. Perez (2025) OpenAIâs Sora soars to no. 1 on Appleâs US App Store. Note: TechCrunchAccessed: 2026 Jan 3 External Links: Link Cited by: §1. [47] A. Radford, J. W. Kim, T. Xu, et al. (2022) Robust speech recognition via large-scale weak supervision. arXiv preprint. External Links: 2212.04356 Cited by: §D.3, §2.3. [48] J. W. Rettberg and H. Wigers (2025) AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by GPT-4o-mini. arXiv preprint. External Links: 2507.22445 Cited by: §4. [49] G. Rose (2016) Visual methodologies: an introduction to researching with visual materials. 4th edition, SAGE Publications Ltd, London. Cited by: §1. [50] Runway (2025) Introducing Runway Gen-4.5. Note: Accessed: 2026 Jan 19 External Links: Link Cited by: §1. [51] U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, and D. Parikh (2022) Make-a-video: text-to-video generation without text-video data. arXiv preprint. External Links: 2209.14792 Cited by: §1. [52] J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, p. 2256â2265. Cited by: §1. [53] P. Taveekitworachai, F. Abdullah, M. C. Gursesli, M. F. Dewantoro, S. Chen, A. Lanata, A. Guazzini, and R. Thawonmas (2023) What is waiting for us at the end? inherent biases of game story endings in large language models. In International Conference on Interactive Digital Storytelling, Cham, p. 274â284. Cited by: §4. [54] The Walt Disney Company (2025) The Walt Disney Company and OpenAI reach agreement to bring Disney characters to Sora. Note: Accessed: 2026 Jan 3 External Links: Link Cited by: §1. [55] P. Valdez and A. Mehrabian (1994) Effects of color on emotions. Journal of Experimental Psychology: General 123 (4), p. 394. Cited by: §D.1.1. [56] M. A. Wakefield, B. Loken, and R. C. Hornik (2010) Use of mass media campaigns to change health behaviour. The Lancet 376 (9748), p. 1261â1271. Cited by: §1. [57] J. Wang, W. Wu, Y. Zhan, R. Zhao, M. Hu, J. Cheng, W. Liu, P. Torr, and K. Q. Lin (2025) Video reality test: can AI-generated ASMR videos fool VLMs and humans?. arXiv preprint. External Links: 2512.13281 Cited by: §1. [58] B. L. Welch (1947) The generalization of âstudentâsâ problem when several different population variances are involved. Biometrika 34 (1â2), p. 28â35. Cited by: §D.5.1, §2.5. [59] L. Wilms and D. Oberfeld (2018) Color and emotion: effects of hue, saturation, and brightness. Psychological Research 82 (5), p. 896â914. Cited by: §D.1.1. [60] M. E. Young, G. R. Norman, and K. R. Humphreys (2008) Medicine in the popular press: the influence of the media on perceptions of disease. PLoS One 3 (10), p. e3552. Cited by: §1. Appendix A Extended Qualitative Methods Coding procedure. Each video was reviewed in three passes. The first pass was an orientation viewing without pausing to get an overall sense of narrative arc, mood, and content. The second pass involved detailed coding with pausing and frame-by-frame scrubbing as needed to tag all elements with accurate timestamps. The third pass was a verification review to confirm accuracy, fill gaps, and add notes about uncertainties. Coders were instructed to prioritize precision over interpretation (tagging what was directly observable rather than inferred) and to err toward over-tagging when uncertain. Coding dimensions. We coded six dimensions for each video, each selected based on established findings in the depression representation literature. Narrative arc captured whether the video presented a discernible shift in emotional trajectory and, if so, its direction: recovery (movement toward more positive, hopeful, or lighter presentation) or deterioration (movement toward more negative, darker, or heavier presentation). Shifts could be signaled through changes in lighting, color palette, music, figure posture, voiceover language, or environmental elements. Not all videos contained shifts; some presented a consistent mood throughout. This dimension addresses whether Sora 2 tends to depict depression as a static state, a worsening condition, or something that improves. Visual depictions of depression consistently employ trajectory metaphors, either escape from confinement or progressive entrapment, making directional shifts a key marker of how the condition is visually conceptualized [18]. Environments were coded with timestamps to capture all physical settings or spaces depicted in each video. We were interested in what kinds of environments Sora 2 associates with depression; whether predominantly indoor, outdoor, domestic, institutional, natural, urban, or abstract. Coders described environments as specifically as the video allowed (e.g., âsmall bedroom with drawn curtains,â âforest path with fog,â âgeometric void with floating shapesâ), using general categories only when more specific description was not possible. Confinement and containment are among the most frequently documented metaphors for depression, often rendered visually through enclosed or restrictive spaces [18]. Objects were inventoried with timestamps to capture what visual symbols and props Sora 2 associates with depression. Categories of interest included furniture and fixtures, weather and atmospheric elements, nature, personal items, domestic objects, and symbolically notable items such as empty chairs, wilting plants, or broken glass. Visual symbols of depression draw on embodied metaphor systems linking the condition to darkness, descent, weight, and fragility [18]. Figure demographics captured information about human figures depicted in the videos, including apparent gender presentation, race/ethnicity, age (coded by decade), and apparent socioeconomic status (inferred from clothing, grooming, and setting context). This dimension may help reveal whether the model reproduces stereotypical representations of who experiences depression, though as discussed in Appendix E, reliability for some demographic variables proved insufficient for substantive analysis. Figures were tagged even when only partially visible (hands, silhouette, back of head), with visibility limitations noted. Voiceover narration without a visible speaker was treated as a figure, with demographic fields marked âUnclearâ when not identifiable from voice alone. Stock imagery of depression disproportionately features young women and carries a more positive emotional tone when depicting men, patterns that may or may not resurface in AI generated videos [4]. Figure states captured how figures appeared and behaved over time, including posture (lying down, seated, standing, walking), facial expression, and whether the figure appeared alone in the frame. Because these attributes can change during a figureâs appearance (e.g., a figure who begins lying in bed looking sad and ends standing by a window looking neutral) this dimension allowed multiple timestamped entries per figure to capture transitions. Facial expressions were coded using standardized categories: sad, happy/smiling, angry, anxious/worried, neutral/flat, tearful, pensive, tired/exhausted, distressed, unclear, or not visible. Slumped posture and social isolation are characteristic of both clinical presentation and media clichĂ©; coding these features allows assessment of whether generated imagery reproduces narrow visual conventions [12]. Overall notes captured general observations, AI artifacts (visual glitches, morphing, distortions), and uncertainties requiring discussion. The full codebook with detailed definitions and edge case handling appears in Appendix B. Reliability. Two authors (MF and Z) independently coded the full dataset using the structured codebook. One authorâs coding (MF) was designated as the primary dataset for analysis; the second authorâs coding (Z) served as inter-rater reliability verification. Raw codes were harmonized prior to reliability analysis to reconcile minor vocabulary differences between coders. Environment descriptions were mapped to 11 standardized categories (Bedroom, Bathroom, Kitchen/Dining, Other Interior, At Window, Transit, Urban Outdoor, Natural Outdoor, Elevated Outdoor, Underwater, Abstract). Object descriptions were harmonized to standardized forms (e.g., âsweatshirt,â âhoodie,â âsweaterâ â âsweatshirt/hoodieâ); reliability was calculated at this harmonized object level. Facial expressions were collapsed into three affect categories: Positive (happy, peaceful, hopeful), Subdued/Low (tired, neutral, pensive), and Negative (sad, tearful, worried, distressed, angry). Posture was harmonized to six categories (Seated, Standing, Lying down, Walking, Floating/Swimming, Kneeling/Crouching). Minor variations in demographic coding (e.g., âMediumâ vs. âMiddleâ for SES) were reconciled during harmonization. The full harmonization dictionaries appear in Appendix C. Cohenâs kappa (Îș) was calculated for presence/absence of categorical elements within each video, including narrative direction, environment categories, harmonized objects, figure demographics, and figure states. For dimensions involving temporal segments (environments, objects, figure states), we also calculated intersection-over-union (IoU) at one-second resolution to assess agreement on when elements appeared. Coders frequently noted gaze direction in free-text annotations; agreement on gaze was assessed for videos where both coders documented directional looking behavior. Given the scale of the dataset, item-by-item consensus resolution was not feasible; reliability metrics verified that the coding scheme could be applied consistently across coders. Appendix B Qualitative Tagging Guide Study Goal This study examines how OpenAIâs Sora 2 video generation model visually interprets and represents the concept of depression. We generated 100 videos using the prompt âDepressionâ and are now systematically analyzing the visual, narrative, and demographic patterns that emerge. Your role as a qualitative coder is to capture timestamped observations about what appears in each video. This qualitative data will be paired with quantitative computer vision analysis to establish a baseline behavior profile for how Sora 2 represents depression. The ultimate goal is to understand whether the model reproduces, reinforces, or challenges existing stereotypes and media conventions around mental health. Overall Principles Precision over interpretation: Tag what you directly observe, not what you infer or interpret. If you see a person lying in bed, tag âperson lying in bedââdo not tag âperson who is exhaustedâ unless exhaustion is explicitly indicated. When uncertain whether something qualifies, tag it and note your uncertainty in the Notes column. Timestamps: All timestamps are in whole seconds from video start (0 = first frame). Round to the nearest second. Err toward over-tagging: It is easier to filter out unnecessary data during analysis than to re-watch all 100 videos. When in doubt about whether something is notable, tag it. Inter-rater process: Each video will be reviewed by at least two coders independently. Given the scale of the dataset, item-by-item consensus resolution is not feasible; and a primary coderâs data will be used for analysis, with inter-rater reliability reported separately. Insufficiently reliable fields will be excluded from substantive analysis. Watching Process First pass (orientation): Watch the full video without pausing to get an overall sense of the narrative arc, mood, and content. Take mental note of any shifts in tone or presentation. Second pass (detailed tagging): Pause and scrub as needed to tag all elements across all tabs (Narrative, Environments, Objects, Figures) with accurate timestamps. Third pass (verification): Review your tags against the video to confirm accuracy, fill in any gaps, and add notes about uncertainties or edge cases. Tab 1: Narrative This tab captures whether each video presents a narrative trajectory. Specifically, whether the emotional or thematic direction of the video changes over its duration. We are interested in whether Sora 2 tends to show depression as a static state, a worsening condition, or something that improves. Table 2: Column descriptions for Narrative tab. Column Description Video ID Filename of the video (e.g., âsora2_depression_001.mp4â) Shift Present? Yes / No: Does the video contain a discernible narrative shift? Shift Second Timestamp (in seconds) where the shift begins. Leave blank if Shift Present = No. Direction âRecoveryâ / âDeteriorationâ: The direction of the shift. Leave blank if Shift Present = No. Notes Any clarifying notes or description that are worth calling out Key Concept: Narrative Shift. A âshiftâ is a discernible change in the videoâs emotional or thematic direction. This could be signaled through changes in lighting, color palette, music, figure posture, voiceover language, or environmental elements. Not all videos will contain a shift. Some may present a consistent mood throughout. Direction Definitions. âą Recovery: A shift toward more positive, hopeful, or lighter presentation. Examples include: lighting warming or brightening; a figure standing up, going outside, or engaging with others; music becoming more uplifting; voiceover expressing hope or improvement. âą Deterioration: A shift toward more negative, darker, or heavier presentation. Examples include: lighting dimming or shifting to cooler tones; a figure withdrawing, lying down, or isolating; music becoming more somber; voiceover expressing despair or worsening. Indicators to watch for. When determining whether a shift has occurred and its direction, pay attention to the following types of signals. If any of these are particularly notable, describe them in the Notes field: âą Voiceover language: transitional phrases like âand yet,â âbut sometimes,â âuntil one dayâ âą Visual shifts: lighting changes, color warming/cooling, figure posture changes âą Audio shifts: music tone changes, volume shifts Tab 2: Environments This tab captures the physical settings or spaces depicted in each video. We want to understand what kinds of environments Sora 2 associates with depression. Are they predominantly indoor, outdoor, domestic, institutional, natural, urban, or abstract? Table 3: Column descriptions for Environments tab. Column Description Video ID Filename of the video (should be the same across all tabs for the same video) Start Second Timestamp when this environment first appears End Second Timestamp when this environment ends (or video ends) Environment Description of environment (see guidance below) Notes Any additional salient details about the environment/setting Environment Description Guidance. Describe the environment as specifically as the video allows. Be concrete and descriptive rather than interpretive. Good examples: âsubway car at night,â âhospital waiting room,â âforest path with fog,â âopen field at dusk,â âsmall bedroom with drawn curtains,â âgeometric void with floating shapes,â âbathroom with running shower.â When a more specific description isnât possible, use these general categories: âą Indoor: Generic interior space where specific room type is unclear âą Outdoor: Generic natural landscape (field, forest, sky, water) without more specific features âą Urban: Generic city or built environment (streets, buildings, infrastructure) âą Abstract: Non-representational space (geometric patterns, voids, surreal landscapes) If the setting is ambiguous or transitioning, describe what you observe and note uncertainty in the Notes column. Tab 3: Objects This tab captures non-human objects that appear in the videos. We are interested in what visual symbols and props Sora 2 associates with depression: medication bottles, rain, empty chairs, phones, etc. Table 4: Column descriptions for Objects tab. Column Description Video ID Filename of video Start Second Timestamp when the object first appears in frame End Second Timestamp when the object leaves the frame (or video ends) Object Name of the object Notes Adjectives or other relevant qualifiers (e.g., color, state, condition) Tagging Guidance. Tag all reasonably identifiable objects that appear in the video, including clothing, furniture, fixtures, and environmental elements. Use the Notes column to capture relevant qualifiers: state (lamp: âOnâ / âOffâ), color (sweatshirt: âGreyâ), condition (plant: âLivingâ / âWiltingâ), or other descriptors. If an object appears, disappears, and reappears, tag each appearance as a separate row. Object Reference List. This list is not exhaustive; describe freely as needed. It is meant to calibrate you to the types of objects we consider notable: âą Furniture & fixtures: bed, chair, couch, table, desk, window, mirror, door, lamp, stairs, bathtub, shower âą Weather & atmospheric: rain, raindrops on glass, clouds, storm, fog, mist, snow, sunshine, god-rays, lightning âą Nature: trees, forest, ocean, water, river, lake, flowers, plants, grass, field, mountain, sky, moon, sun, stars âą Personal items: phone, medication/pill bottle, book, journal, photographs, cup/mug, food, cigarette, headphones âą Domestic objects: blanket, pillow, curtains, clock, television, computer, dishes âą Symbolic/notable: empty chair, wilting plant, broken glass, candle, shadow, reflection Tab 4: Figures This tab captures information about human figures (and voices) depicted in the videos. We are interested in the demographics of who Sora 2 depicts when visualizing depression, as well as how these figures are portrayed (posture, expression, isolation). Table 5: Column descriptions for Figures tab. Column Description Video ID Filename of video Figure ID Figure 1, Figure 2, etc. (in order of first appearance) Start Second Timestamp when figure first appears End Second Timestamp when figure leaves frame Gender Observed gender presentation: Man / Woman / Nonbinary / Unclear Race/Ethnicity Observed race/ethnicity, or âUnclearâ if not identifiable Age Decade estimate (teens, 20-30, 30-40, 40-50, 50-60, 60+, unclear) Apparent SES Inferred socioeconomic status based on appearance: Low / Middle / High / Unclear (see guidance) Notes Visibility limitations, posture, or other relevant observations Tagging Guidance. âą Partial visibility: Tag figures even if only partially visible (hands, silhouette, back of head). Note the limitation in the Notes column (e.g., âonly hands visible,â âsilhouette onlyâ). âą Continuity across cuts: If the same figure is clearly continuous across a cut, use the same Figure ID. If itâs uncertain whether itâs the same person, assign a new ID and note the uncertainty. âą Voiceovers: Treat voiceover narration as a figure, using âVoice 1,â âVoice 2â as the Figure ID. Demographic fields may be marked âUnclearâ if not identifiable from voice. âą AI distortions: If AI artifacts distort a figure, still attempt to tag what you can observe and note the distortion. Apparent SES Guidance. This is inherently a rough inference. Base your judgment on observable cues such as: âą Clothing quality and style (worn/unkempt vs. well-maintained vs. expensive-looking) âą Grooming and personal presentation âą Setting context (sparse room vs. well-furnished home vs. affluent environment) âą Mark âUnclearâ if there are insufficient cues to make even a rough judgment. Tab 5: Figure States This tab captures how figures appear and behave over time: their posture, facial expression, and social context within the frame. Because these attributes can change during a figureâs appearance (e.g., a figure who begins lying in bed looking sad and ends standing by a window looking neutral), this tab allows multiple entries per figure to capture those transitions. Table 6: Column descriptions for Figure States tab. Column Description Video ID Filename of the video Figure ID Links to the figure in Tab 4 (e.g., âFigure 1â). Must match an existing Figure ID. Start Second Timestamp when this state begins End Second Timestamp when this state ends or figure leaves frame Posture Lying down / Seated / Standing / Walking / Unclear Facial Expression See options below Alone in Frame Yes / No: Is this figure the only person visible during this state? Notes Qualifiers or context (e.g., âhead in hands,â âlooking out window,â âback to camera,â âin bedâ) Tagging Guidance. Create a new row whenever a figureâs posture, expression, or social context (alone/not alone) meaningfully changes. Not every micro-movement requires a new entry. Use judgment. If a figure shifts from sad to neutral over a few seconds, thatâs worth capturing. If they blink or briefly glance away, itâs not. Facial Expression Guidance. Tag the predominant expression during the figureâs appearance. Options: âą Sad: Downturned mouth, furrowed brow, drooping features âą Happy/Smiling: Upturned mouth, raised cheeks, bright eyes âą Hopeful/Peaceful: Relaxed face, soft eyes âą Angry: Glaring eyes, clenched jaw, compressed lips, hard downward brow (intensity, aggression) âą Anxious/Worried: Raised inner eyebrows, wide or darting eyes, lip biting, tense but uncertain (unease, apprehension) âą Neutral/Flat: Blank or expressionless affect âą Tearful: Visible tears or crying âą Pensive: Thoughtful, contemplative, gazing distantly âą Tired/Exhausted: Heavy eyelids, slack features, weary appearance âą Distressed: Visible anguish, tension, or agitation âą Unclear: Face visible but expression hard to categorize âą N/A: Face not visible (back of head, silhouette, hands only) Tab 6: Notes This tab captures overall observations, uncertainties, edge cases, issues, or random notes that do not fit elsewhere in the tagging sheet. Every video should have exactly one row in this tab. Table 7: Column descriptions for Notes tab. Column Description Video ID Filename of video Date Coded Date you coded this video Notes Free-text field for overall observations Handling Edge Cases Abstract or non-representational content: Tag what you can. Environments can be âabstract.â If figures are distorted or surreal, attempt demographic coding with âUnclearâ where needed. AI artifacts and glitches: If visual glitches, morphing, or distortions make identification difficult, note this in Tab 6 but still attempt to tag what you observe. No figures present: If a video contains no human figures or voiceovers, leave the Figures tab empty for that video. This is valid data. Rapid scene changes: If environments or objects change very rapidly (e.g., montage sequences), tag the most salient or longest-duration elements. Note in Tab 5 that the video contained rapid cuts. Unclear or borderline cases: When genuinely uncertain, err on the side of tagging and note your uncertainty. Given the scale of the dataset, item-by-item consensus resolution is not feasible; reliability metrics will be used to verify coding consistency, and a primary coderâs judgments will be used for analysis. Appendix C Harmonization Dictionaries To reconcile minor vocabulary differences between coders and enable systematic analysis, raw codes were mapped to standardized categories. All mappings used case-insensitive substring matching unless otherwise noted. C.1 Environment Harmonization Notes: âunderwater bedroomâ mapped to Underwater. Close-up shots carried forward the previous environment. Table 8: Environment Harmonization Standardized Category Raw Code Patterns Bedroom âbedroomâ, âin bedâ Bathroom âbathroomâ, âbathtubâ, âin bathtubâ Kitchen/Dining âkitchenâ, âdiningâ, âcafeâ Other Interior âdeskâ, âliving roomâ, âhallwayâ, âbasementâ, âempty roomâ, âelevatorâ At Window âwindowâ, âat windowâ, ânext to windowâ, âin front of windowâ Transit âbusâ, âsubwayâ, âtrainâ, âcarâ, âtransitâ Urban Outdoor âstreetâ, âcityâ, âurbanâ, âunderpassâ, âparkingâ Natural Outdoor âparkâ, âfieldâ, âmeadowâ, âriverâ, âlakeâ, âlagoonâ, âpondâ, âwheatâ, âflowerâ, âgardenâ, âice flatâ, âoutsideâ Elevated Outdoor âroofâ, ârooftopâ, âhilltopâ, âmountainâ, âbalconyâ, âtop ofâ, âbridgeâ Underwater âunderwaterâ, âunder waterâ, âsubmergedâ, âin waterâ Abstract âspaceâ, âcloudâ, âabstractâ, âgalaxyâ, âskyâ, âdomeâ C.2 Object Harmonization Note: Objects not matching any pattern were retained as-is. Table 9: Object Harmonization Standardized Object Raw Code Patterns sweatshirt/hoodie âsweatshirtâ, âhoodieâ, âhoddieâ, âsweaterâ, âlong-sleeveâ, âlong sleeveâ, ât-shirtâ, âshirtâ jacket/coat âjacketâ, âcoatâ, ârainjacketâ, âraincoatâ, âblazerâ bed âbedâ (excluding âbedroomâ) couch/sofa âcouchâ, âsofaâ end table âend tableâ, ânightstandâ table/desk âtableâ, âdeskâ rain ârainâ, âheavy rainâ raindrops âraindropâ, ârain onâ cloud âcloudâ sunrise/sunset âsunriseâ, âsunsetâ sunlight/god rays âgod rayâ, âgod-rayâ, âsunlightâ standing water/puddles âstanding waterâ, âwater onâ, âpuddleâ bathtub âbathtubâ bubbles âbubbleâ plant/flowers âplantâ, âflowerâ, âsunflowerâ trees âTreeâ, âtreesâ birds âbirdâ phone âphoneâ mug/cup âmugâ, âcupâ lamp âlampâ window âwindowâ, âwindowsâ curtain âCurtainâ, âcurtainsâ tear âtearâ mirror âmirrorâ hands âhandsâ, âhandâ cityscape/buildings âcityscapeâ, âbuildingâ, âskylineâ cars âcarâ, âcarsâ picture/poster âpictureâ, âposterâ, âphotoâ, âpolaroidâ bench âbenchâ C.3 Figure State Harmonization Table 10: Posture Harmonization Standardized Category Raw Code Patterns Seated âseatedâ, âsittingâ, âsetedâ Standing âstandingâ, âstandngâ, âstadingâ Lying down âlyingâ Walking âwalkingâ Floating/Swimming âfloatingâ, âswimmingâ Kneeling/Crouching âkneelingâ, âcrouchingâ, âsquattingâ, âcrawlingâ Table 11: Facial Expression Harmonization Affect Category Raw Codes Positive âhappyâ, âhopefulâ, âpeacefulâ, âsurprisedâ Subdued/Low âtiredâ, âneutralâ, âpensiveâ, âflatâ, âfocusedâ, âsleepingâ Negative âsadâ, âtearfulâ, âcryingâ, âworriedâ, âdistressedâ, âanxiousâ, âangryâ, âdesperateâ, âdejectedâ Table 12: Alone-in-Frame Harmonization Standardized Raw Codes Yes âyesâ No ânoâ C.4 Figure Demographic Harmonization Table 13: Apparent SES Harmonization Standardized Raw Codes High âhighâ Medium âmediumâ, âmiddleâ Low âlowâ Table 14: Race/Ethnicity Harmonization Standardized Raw Codes Black âblackâ White âwhiteâ Latino/Hispanic âhispanicâ, âlatinoâ, âlatinaâ Asian âasianâ Arab/Middle Eastern âarabâ, âmiddle easternâ Indian âindianâ Table 15: Gender Harmonization Standardized Raw Codes Male âmaleâ Female âfemaleâ Table 16: Age Harmonization Standardized Raw Codes 10-20 âteensâ, â10-20â 20-30 â20-30â 30-40 â30-40â 40-50 â40-50â 50-60 â50-60â Table 17: Gaze Direction Harmonization Gaze Direction Note Patterns Down âlooking downâ, âlooking at handsâ, âlooking at the groundâ, âhead in handsâ, âface in handsâ, âeyes closedâ, âlooking at feetâ, âlooking at lapâ Up âlooking upâ, âlooking at the skyâ, âlooking at sunsetâ, âlooking at sunriseâ, âlooking at the cloudâ, âlooking at cloudsâ, âlooking at lightsâ, âlooking to the skyâ Out/Away âlooking outâ, âlooking outsideâ, âlooking at windowâ, âout of windowâ, âout windowâ, âlooking farâ, âlooking afarâ, âlooking into distanceâ, âlooking awayâ, âlooking aroundâ, âin front of windowâ, âchecking outsideâ Appendix D Extended Quantitative Methods All feature extraction was implemented in Python using a custom analysis pipeline. The analysis code and video dataset analyzed in this study are available from the corresponding author upon reasonable request. D.1 Visual Aesthetics We extracted frames at one-second intervals, yielding 11 frames for 10-second app videos and 13 frames for 12-second API videos. D.1.1 Color and Light Measures Each measure was selected based on established connections to mood and emotion in prior research. Brightness, measured as perceptual luminance, has established connections to mood representation [11]. Saturation captures color intensity; reduced saturation is associated with diminished emotional arousal [59]. Colorfulness was computed using the Hasler-SĂŒsstrunk metric [24], which correlates with human perception of chromatic vividness. Color temperature indexes the warm-cool balance; warm tones are associated with comfort and social connection, while cool tones connote isolation [55]. Table 18: Visual aesthetics measures Measure Computation Brightness Brightness=0.299âR+0.587âG+0.114âBBrightness=0.299R+0.587G+0.114B Saturation S channel of the HSV color space, computed as the difference between maximum and minimum RGB values divided by the maximum Colorfulness Colorfulness=Ïrâg2+Ïyâb2+0.3ĂÎŒrâg2+ÎŒyâb2Colorfulness= _rg^2+ _yb^2+0.3Ă _rg^2+ _yb^2 where râgrg and yâbyb represent opponent color channels (râg=RâG;yâb=0.5â(R+G)âB)(rg=R-G;\ yb=0.5(R+G)-B), and Ï and ÎŒ denote standard deviation and mean respectively. Color temperature Mean (RâB)+0.3Ă(RâG)(R-B)+0.3Ă(R-G) channel difference, with positive values indicating warmer colors and negative values indicating cooler colors. Edge Density Proportion of Canny edge pixels to total pixels: Edge Density=ââ[edgeâ(x,y)>0]WĂHEdge Density= 1[edge(x,y)>0]WĂ H where W and H are frame dimensions. Adaptive thresholds were set at 0.7Ă and 1.3Ă the median pixel intensity to account for varying image brightness. D.1.2 Texture Measures We also extracted texture features using Gray-Level Co-occurrence Matrices [23], which provide insight into environmental qualities. Smooth textures may suggest sterile spaces, while complex textures indicate detailed or chaotic environments. We computed the GLCM at distance=1 pixel across four angles (0°, 45°, 90°, 135°) and averaged the resulting features. where Pâ(i,j)P(i,j) is the normalized co-occurrence matrix. Table 19: GLCM texture measures Measure Formula Interpretation Contrast âi,j(iâj)2â Pâ(i,j) _i,j(i-j)^2· P(i,j) Local intensity variation Dissimilarity âi,j|iâj|â Pâ(i,j) _i,j|i-j|· P(i,j) Average difference between neighboring pixels Homogeneity âi,jPâ(i,j)1+(iâj)2 _i,j P(i,j)1+(i-j)^2 Closeness of distribution to diagonal Energy âi,jPâ(i,j)2 _i,jP(i,j)^2 Uniformity of gray level distribution Correlation âi,j(iâÎŒi)â(jâÎŒj)â Pâ(i,j)ÏiâÏj _i,j (i- _i)(j- _j)· P(i,j) _i _j Linear dependency of gray levels Entropy ââi,jPâ(i,j)â log2âĄ(Pâ(i,j))- _i,jP(i,j)· _2(P(i,j)) Textural randomness or complexity D.2 Audio Properties Audio was extracted from video files and analyzed using the librosa library [38]. These acoustic features have established links to depressive affect: depressed speech is characterized by reduced volume, lower pitch, reduced pitch variability, and lower spectral centroid compared to non-depressed speech [36, 43]. Table 20: Audio measures Measure Formula Interpretation Volume Root-mean-square (RMS) energy with frame length of 2048 samples and hop length of 512 samples Acoustic intensity Pitch Fundamental frequency (F0) extracted using the pYIN algorithm [37], which combines a probabilistic YIN estimator with a hidden Markov model. Mean F0 computed across all voiced frames Vocal register Spectral centroid Centroid=âkfkâ MkâkMkCentroid= _kf_k· M_k _kM_k where fkf_k is the frequency of bin k and MkM_k is the magnitude. âCenter of massâ of the frequency spectrum; higher values indicate brighter, more energetic audio, lower values suggest muffled or subdued soundscapes D.3 Semantic Content Spoken audio was transcribed using OpenAIâs Whisper model (base) [47], with word-level timestamps extracted for temporal analysis. We computed sentiment scores, custom lexicons, and linguistic markers selected based on prior research linking these features to depression and recovery discourse. D.3.1 Sentiment and Speech Measures We utilized a VADER [28] computed compound, positive, negative, and neutral sentiment ratings for each transcription. VADER is calibrated for social media text and handles informal language, emoticons, and emphasis markers. Table 21: Semantic content measures Measure Computation Word Count Total words per video Speech Rate Words per second of video duration Sentiment VADER compound, positive, negative, and neutral scores D.3.2 Light/Dark Lexicons We developed custom lexicons based on conceptual metaphor theory [32] and research showing that depression discourse frequently employs metaphors of weight, darkness, and drowning, while recovery narratives invoke lightness and release [35]. Using word-level timestamps from Whisper, we tracked when these terms appeared across the video timeline. Ambiguous terms that could appear in neutral contexts were excluded. Table 22: Light/dark lexicons Lexicon Terms Light (hope/recovery) hope, hopeful, hoping, relief, relieved, release, released, free, freedom, peace, peaceful, calm, calming, serene, serenity, heal, healing, healed, recover, recovery, recovering, joy, joyful, happy, happiness, smile, smiling, laugh, laughing, laughter, love, loved, loving, comfort, comforting, comfortable, safe, safety, warm, warmth, support, supported, supporting, embrace, embraced Dark (distress) heavy, heaviness, burden, burdened, weigh, weighing, crush, crushing, crushed, drown, drowning, drowned, suffocate, suffocating, choke, choking, trap, trapped, stuck, confined, cage, caged, alone, lonely, loneliness, isolated, isolation, abandoned, forgotten, empty, emptiness, void, hollow, numb, numbness, pain, painful, hurt, hurting, ache, aching, suffer, suffering, exhaust, exhausted, exhausting, drain, drained, draining, weary, fatigue, despair, despairing, hopeless, hopelessness, helpless, helplessness, worthless, worthlessness, fear, afraid, scared, terrified, anxiety, anxious, panic, dread D.3.3 Linguistic Pattern Measures We computed linguistic markers based on findings from computational studies of depression. First-person singular pronoun ratio indexes self-focused attention, which has been consistently linked to depressive rumination [5]. Negation frequency captures negative cognitive framing characteristic of depressive thought patterns [1]. Table 23: Linguistic pattern measures Measure Numerator Denominator First-person singular pronoun ratio Count of: I, me, my, mine, myself Total words Negation ratio Count of: no, not, never, nothing, none, nobody, nowhere, neither, nor, canât, couldnât, wonât, wouldnât, donât, doesnât, didnât, shouldnât, isnât, arenât, wasnât, werenât, havenât, hasnât, hadnât Total words Lexical diversity (TTR) Unique words Total words D.4 Temporal Dynamics We computed motion, scene structure, and feature trajectories to analyze whether videos depicted stasis, deterioration, recovery, or cyclical patterns over their duration. Motion magnitude is particularly relevant to depression representation because psychomotor retardation (slowed movement and activity) is a core diagnostic feature of major depression [6]. D.4.1 Motion Analysis Dense optical flow was computed using the FarnebĂ€ck algorithm [13] between consecutive frames at native framerate (30 fps). The algorithm estimates motion vectors for every pixel using polynomial expansion. Flow magnitude was computed as u2+v2 u^2+v^2 where u and v are horizontal and vertical flow components. Frame-level values were binned into per-second averages for trajectory analysis. Table 24: FarnebĂ€ck optical flow parameters Parameter Value Pyramid scale 0.5 Pyramid levels 3 Window size 15 Iterations 3 Polynomial neighborhood 5 Polynomial sigma 1.2 D.4.2 Scene Detection Scene cuts were detected by computing histogram correlation between frames sampled at one-second intervals. A cut was registered when correlation dropped below 0.3. Scene cut frequency provides insight into editing rhythm as rapid cuts may signal narrative progression or emotional intensity, while continuous shots suggest sustained states. D.4.3 Feature Trajectories For each video, per-second values were computed for all features. Linear slope was computed via ordinary least squares to characterize whether features increased, decreased, or remained stable over the video duration: slope=ât(tâtÂŻ)â(xtâxÂŻ)ât(tâtÂŻ)2slope= _t(t- t)(x_t- x) _t(t- t)^2 (1) D.5 Statistical Formulas We used standard inferential statistics to compare features between App and API outputs, with corrections for multiple comparisons across all features tested. D.5.1 Group Comparisons Welchâs t-test [58] was used for all between-group comparisons because it does not assume equal variances across groups; an appropriate choice given the substantial differences in output characteristics between modalities. t=XÂŻ1âXÂŻ2s12n1+s22n2t= X_1- X_2 s_1^2n_1+ s_2^2n_2 (2) D.5.2 Effect Sizes Cohenâs d [10] was used to quantify the magnitude of differences between App and API outputs, providing a standardized measure independent of sample size. d=XÂŻ1âXÂŻ2spooledd= X_1- X_2s_pooled (3) where spooled=(n1â1)âs12+(n2â1)âs22n1+n2â2s_pooled= (n_1-1)s_1^2+(n_2-1)s_2^2n_1+n_2-2 (4) Effect sizes were interpreted using conventional thresholds: |d|=0.2|d|=0.2 (small), 0.50.5 (medium), 0.80.8 (large). D.5.3 Multiple Comparison Correction To control the false discovery rate when conducting multiple comparisons, p-values were adjusted using the Benjamini-Hochberg procedure [3] which provides a principled balance between Type I error control and statistical power: The procedure operates as follows: Rank all m p-values from smallest to largest: p(1)â€p(2)â€âŠâ€p(m)p_(1)†p_(2)â€âŠâ€ p_(m) For each ranked p-value, compute the critical value: kmĂα kmĂα Find the largest k such that p(k)â€kmĂαp_(k)†kmĂα Reject all null hypotheses for tests with rank â€k†k. Adjusted p-values (q-values) were computed as: q(k)=minâĄ(mâ p(k)k,1)q_(k)= ( m· p_(k)k,1 ) (5) The raw transformed values were then made monotone by processing from rank m down to rank 1, replacing each q(k) with the minimum of its own value and the q-value at rank k+1, ensuring that adjusted p-values respect the ordering of the original p-values. We used α = 0.05 for all analyses. Where relevant throughout the results, we report FDR-corrected q-values rather than raw p-values. D.6 Qualitative Interrater Reliability Methods Two authors (MF and Z) independently coded the complete dataset of 100 videos using the structured codebook, producing parallel sets of timestamped codes across all dimensions. One authorâs coding (MF) was designated as the primary dataset for analysis. The second authorâs independent coding (Z) was used to calculate inter-rater reliability. D.6.1 Cohenâs Kappa Agreement on categorical dimensions was assessed using Cohenâs kappa (Îș) [9]: Îș=poâpe1âpeÎș= p_o-p_e1-p_e (6) where pop_o is the observed proportion of agreement and pep_e is the proportion of agreement expected by chance: pe=âkp1âkâ p2âkp_e= _kp_1k· p_2k (7) where p1âkp_1k and p2âkp_2k are the proportions of each category assigned by coders 1 and 2, respectively. Kappa was calculated for: âą Narrative shift presence (Yes/No) âą Narrative direction (Recovery/Deterioration) âą Environment categories (presence/absence per category per video) âą Object categories (presence/absence per harmonized object per video) âą Figure demographics (gender, race/ethnicity, age range, apparent SES) âą Figure states (posture, facial expression, alone in frame) âą Gaze direction (presence/absence of down, up, out/away per video) Kappa values were interpreted following Landis and Koch [33]: Table 25: Landis and Koch kappa interpretation Îș Value Interpretation < 0.00 Poor 0.00â0.20 Slight 0.21â0.40 Fair 0.41â0.60 Moderate 0.61â0.80 Substantial 0.81â1.00 Almost perfect D.6.2 Temporal Intersection over Union For dimensions involving timestamped segments (environments, objects, figure states), kappa captures only whether coders agreed that a category was present in a video, not whether they agreed on when it appeared. To assess temporal agreement, we computed Intersection over Union (IoU) at one-second resolution: IoU=|Aâ©B||AâȘB|IoU= |Aâ© B||AâȘ B| (8) where A and B are binary temporal masks indicating the seconds during which each coder marked a category as present. For each video, we created binary vectors with one element per second of video duration (11 elements for 10-second app videos, 13 elements for 12-second API videos) for each coderâs segment annotations, then computed the intersection and union across these vectors. IoU ranges from 0 (no temporal overlap) to 1 (perfect temporal agreement). We computed IoU only for categories where both coders agreed the element was present, providing a measure of boundary precision conditional on presence agreement. D.6.3 Timestamp Precision For narrative shift timestamps (a single moment rather than a segment), we computed the mean and median absolute difference between codersâ timestamps for videos where both identified a shift: Precision=|tMFâtZZ|Precision=|t_MF-t_Z| (9) Given the scale of the dataset (thousands of individual codes across six coding dimensions) item-by-item consensus resolution was not feasible. Reliability metrics were used to verify that the coding scheme could be applied consistently across coders, with the primary coderâs judgments (MF) used for all analyses. Appendix E Inter-Rater Reliability Results Two coders independently coded all 100 videos. The primary coderâs data (MF) was used for all analyses; the second coderâs independent coding (Z) was used to calculate reliability. E.1 Narrative Arc The three-way classification (recovery, deterioration, no shift) showed poor agreement on deterioration, so we collapsed to a binary classification of recovery vs. no recovery (no shift + deterioration), which achieved substantial agreement (Îș = 0.69, 85% overall agreement). When both coders identified narrative shifts, temporal agreement was strong (mean difference = 0.31 seconds, median = 0.00 seconds, n = 35 pairs). E.2 Environments Overall presence/absence agreement was excellent (Îș = 0.91) with strong temporal agreement (mean IoU = 0.84). Category-specific kappas ranged from substantial to perfect: Table 26: Environment inter-rater reliability Category Îș Agreement Bathroom 1.00 100% Kitchen/Dining 1.00 100% Elevated Outdoor 0.97 99% Urban Outdoor 0.97 99% Underwater 0.94 99% Transit 0.91 98% Natural Outdoor 0.89 97% Bedroom 0.82 96% Other Interior 0.78 94% Abstract 0.66 98% At Window 0.53 90% E.3 Objects Overall presence/absence kappa was moderate (Îș = 0.61) with strong temporal agreement when objects were mutually identified (mean IoU = 0.82). Agreement varied substantially by object type. High reliability objects (Îș > 0.70): mug/cup (Îș = 0.93), pen (Îș = 0.90), mirror (Îș = 0.90), bench (Îș = 0.86), birds (Îș = 0.81), lamp (Îș = 0.81), cityscape/buildings (Îș = 0.80), cloud (Îș = 0.77), plant/flowers (Îș = 0.71). Lower reliability objects: bed (Îș = 0.44), window (Îș = 0.47), sweatshirt/hoodie (Îș = 0.29), rain (Îș = 0.23). Examination of disagreements revealed an asymmetric pattern: the primary coder noted these ubiquitous items more exhaustively, while the secondary coder focused on more distinctive elements. This asymmetric pattern suggests the lower agreement values partly reflect a threshold difference in tagging exhaustiveness, though the resulting kappa values should be interpreted with this limitation in mind. E.4 Figure Demographics Gender achieved perfect agreement. Age showed high raw agreement (89%) but low kappa (Îș = 0.38) due to restricted range (88% of figures coded as 20-30). Race/ethnicity showed moderate agreement (Îș = 0.53), likely due to numerous bi-racial and ethnically ambiguous AI-generated figures; this dimension was excluded from analysis. Apparent SES showed poor agreement (Îș = 0.20) and was excluded from further analysis. Table 27: Figure demographics inter-rater reliability Variable Îș Agreement Gender 1.00 100% Age 0.38 89% Race/Ethnicity 0.53 68% Apparent SES 0.20 56% E.5 Figure States Posture and alone-in-frame status achieved strong reliability. Facial expression showed poor agreement (Îș = 0.49) due to ambiguous AI-generated expressions and was excluded from analysis. Table 28: Temporal agreement (IoU) Variable Îș IoU Posture 0.89 0.86 Alone in Frame 0.94 0.98 Facial Expression 0.49 0.63 E.6 Gaze Direction Coders frequently noted directional gaze in free-text annotations. When both coders documented gaze direction, agreement on dominant direction was high (96%, n = 84). The primary coder noted downward gaze more frequently (93 videos) than the secondary coder (79 videos), reflecting threshold differences in annotation. E.7 Dimensions Excluded from Analysis Based on reliability results, the following dimensions were excluded from substantive analysis: race/ethnicity (Îș = 0.53), apparent SES (Îș = 0.20), and facial expression (Îș = 0.49).