Paper deep dive
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
Brittany Harbison, Ashok K. Goel
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/3/2026, 1:52:43 AM
Summary
The paper introduces LEX-EC, a black-box audit framework for evaluating zero-shot Large Language Model (LLM) personality classification. It combines prevalence diagnostics, agreement metrics, and controlled lexical ablation to assess how trait associations persist across different text genres (essays, student introductions, Facebook statuses) and under content masking. The study finds that signal strength varies by genre and length, with function words and affective terms providing more stable evidence than topical content. It also examines how linguistic prompting influences model self-explanations.
Entities (19)
Relation Signals (13)
LEX-EC → evaluates → personality classification
confidence 95% · audit framework for interpretability of LLMs at the task of personality prediction
LEX-EC → uses → lexical ablation
confidence 95% · combining prevalence and agreement diagnostics with controlled lexical ablation
Pennebaker-King essays → analyzedby → LEX-EC
confidence 90% · The essay corpus is the Pennebaker–King stream-of-consciousness dataset... We report essay results
MyPersonality Facebook statuses → analyzedby → LEX-EC
confidence 90% · The Facebook corpus derives from the myPersonality project... We simultaneously balanced the five marginal trait distributions
LINEX → directs → justifications toward linguistic, stylistic, and affective evidence
confidence 90% · Linguistic+Explanation (LINEX), directing justifications toward linguistic, stylistic, and affective rather than topical or demographic evidence
BASICEX → directs → trait-level justifications
confidence 90% · Basic+Explanation (BASICEX), adding trait-level justifications
Claude Haiku 4.5 → usedin → LEX-EC
confidence 90% · evaluating five closed models... Claude Haiku 4.5
GPT-4o-mini → usedin → LEX-EC
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.
Tags
Links
- Source: https://arxiv.org/abs/2607.24435v2
- Canonical: https://arxiv.org/abs/2607.24435v2
Trouble viewing inline? Open PDF directly →
Full Text
69,029 characters extracted from source content.
Expand or collapse full text
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Brittany Harbison1 , Ashok K. Goel1 Abstract Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates item-level association and chance-corrected agreement, examines how both change under lexical restriction, and conducts a targeted audit of prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling. Code — https://github.com/Noeru-H/LEX-EC Introduction Large language models (LLMs) may assign personality traits (Goldberg 1990; McCrae and John 1992) to an author from ordinary text, and such inferences are increasingly used in education, hiring, and social computing (Vinciarelli and Mohammadi 2014; Matz et al. 2017). But a plausible label is not necessarily a valid one: a model may return the same confident judgment regardless of whether the text is rich or a single sentence, whether it reflects trait-relevant language or mere topical content, whether the inference tracks the author or demographic priors the model brings to it (Rao et al. 2023; Peters and Matz 2024), and whether aggregate scores mask near-chance item-level agreement. Prior work reports apparent success, as correlations with human raters or self-reports across traits and genres (Derner et al. 2024; Peters and Matz 2024; Schoenegger et al. 2025; Piastra and Catellani 2025; Wright et al. 2026), but the evidential basis of these black-box predictions remains difficult to interpret: do observable associations exist across text genres and quantities, do predictions persist under lexical content masking, and what effect does prompt engineering have on behavior? For closed or privately hosted models standard interpretability studies are impossible: without access to weights, gradients, or activations, interpretability methods that require internal inspection (Belinkov 2022; Olah et al. 2020) do not apply. Black-box interpretability operates under this constraint, using input perturbation, attribution, and behavioral testing to characterize models from inputs and outputs alone (Ribeiro et al. 2016; Lundberg and Lee 2017; Ribeiro et al. 2020). This study asks what zero-shot LLM Big Five Inventory (BFI) predictions rest on once they are examined across text conditions, prompt framings, distributional checks, and content-reduced inputs. This work is organized around three questions: • RQ1. Do zero-shot LLMs show BFI signal, and how does it vary with text length and genre? • RQ2. When topical and demographic content is masked, which predictions collapse and which persist as recoverable from the preserved lexical layer? • RQ3. Does linguistic prompting shift self-explanations toward stylistic and affective rationales? We approach these with LEX-EC (Lexical Evidence Channel auditing), a novel, reusable behavioral auditing framework for interpretability of LLMs at the task of personality prediction, applied across a text-length gradient, combining distributional and item-level agreement analysis, prompt comparison, and a content-masking ablation that tests signal survival after removal of topical and demographic vocabulary. Rather than treating LLM personality labels as measurements, they are treated solely as behavioral outputs. This framework can narrow causal hypotheses, motivate targeted follow-up interventions, and provide evidence informing decisions about under what conditions personality-labeling systems should be used. Related Work Personality Measures and the Big Five Inventory BFI models personality along Extraversion, Neuroticism, Agreeableness, Conscientiousness, and Openness through questionnaire responses (John et al. 1991; McCrae and Costa 1987; John and Srivastava 1999). Although these traits are ordinarily measured continuously, some reference datasets used in personality labeling historically provide only binary high/low labels. Prior psycholinguistic work has found modest, context-dependent associations between personality and function words, affective vocabulary, and other linguistic patterns (Pennebaker and King 1999; Yarkoni 2010; Schwartz et al. 2013), motivating the possibility of personality prediction from text. Lexicon-based Textual Analysis: LIWC, Emotion, NRC Lexicon-based methods such as LIWC map words to psychologically motivated categories including affect, cognition, and function-word use (Tausczik and Pennebaker 2010). Open resources such as Empath and NRC-EmoLex provide related affective and cognitive categories (Fast et al. 2016; Mohammad and Turney 2013). Notably, Empath’s categories, when compared against analogous LIWC categories, showed a reported average correlation of r=.906r=.906 (Fast et al. 2016). We use these resources to construct the preserved lexical layer in our masking ablation. Corpus-linguistic studies show that text length is systematically associated with register, linguistic-feature distributions, and communicative function rather than varying independently of the setting in which text is produced (Liimatta 2022, 2023; Öhman and Liimatta 2024), rendering these characteristics difficult to separate without confounding. LLM Personality Prediction Recent work suggests that LLMs can recover nonzero BFI signal from text, sometimes approaching individual human judgments (Derner et al. 2024; Schoenegger et al. 2025; Piastra and Catellani 2025; Wright et al. 2026). Prompting strategies can materially affect this performance (Yang et al. 2023). Reported performance varies substantially with prompt framing, text type, text quantity, and the evaluation target (Bhandarkar et al. 2025; Piastra and Catellani 2025). Models also exhibit demographic sensitivity, positivity bias, poor confidence calibration, and weaker agreement with self-reported traits than with individual human ratings (Rao et al. 2023; Peters and Matz 2024). In particular, prior studies commonly aggregate multiple messages or extended texts per author, leaving performance under limited-text conditions less well characterized;linguistic and affective evidence from topical or demographic cues are also rarely examined. We therefore evaluate predictions across a text-length gradient and test whether apparent signal persists after content masking. Complementing prompt-based studies, Maharjan et al. trained classifiers over BERT, RoBERTa, and OpenAI embeddings on the PANDORA dataset, reporting stronger performance than zero-shot prompting and associations with LIWC and NRC features, although psychometric reliability was only moderate (Maharjan et al. 2025). Figure 1: Document prevalence and corpus-level density of retained Empath and NRC categories. Prevalence is the percentage of documents containing at least one category term; density is the number of category occurrences per 100 word-tokens. Lexicon categories are not mutually exclusive. Labels identify selected high-prevalence, high-density, or corpus-distinctive categories. Black-box Interpretability Because closed LLMs do not expose weights or activations, behavior is often characterized through controlled input transformations. Common black-box interpretability methods include local feature-attribution approaches such as LIME and SHAP, which estimate the contributions of input features to individual predictions (Ribeiro et al. 2016; Lundberg and Lee 2017), as well as behavioral-testing approaches such as CheckList, which evaluate model responses under targeted input conditions and transformations (Ribeiro et al. 2020). Input-reduction studies similarly examine whether predictions persist as information is removed, while showing that behavior on transformed inputs need not provide a faithful account of the process that produced the original prediction (Feng et al. 2018). Some work additionally argues that interpretability methods should be evaluated in relation to the particular claims and contexts they are intended to support (Doshi-Velez and Kim 2018). Model-generated explanations present a related, but distinct, black-box signal. Fluent self-explanations may not faithfully reflect the computation producing a prediction (Jacovi and Goldberg 2020; Turpin et al. 2023), though they may still reflect some signal in aggregate (Madsen et al. 2024). Comparative evidence indicates that lexicon-based measures are useful in NLP because they provide transparent, computationally inexpensive, and directly inspectable representations of language (van der Veen and Bleich 2025). Related work in quantitative discourse analysis likewise emphasizes the value of explicit lexical evidence for improving methodological transparency and making analytical decisions easier to trace (Compton 2025). Relatedly, PsyTEx removes spans chosen from LLM-generated trait criteria and repeats prediction on the remainder, subsequently analyzing the selected text with LIWC (Bhandarkar et al. 2025). LEX-EC instead specifies the retained psycholinguistic channel independently to avoid self-report dependence and couples lexical restriction with item-level agreement audits. Datasets Three datasets were selected to span distinct genres and average text lengths: single MyPersonality Facebook statuses (Celli et al. 2013), graduate student forum introductions (Harbison et al. 2026), and Pennebaker-King essays (Pennebaker and King 1999; Celli et al. 2013). The texts average 17, 89, and 672 mean retained word-tokens 111Word-tokens: spaCy tokens on preprocessed text (URLs stripped, PII and program identifiers replaced), excluding punctuation, whitespace, and emoticon/emoji spans., respectively, representing micro-, short-, and middle-length conditions. This places the study at the lower end of the text quantities used in prior zero-shot personality-prediction work, which often aggregates multiple letters, statuses, or other texts per author. Binary high/low labels provide a common prediction and evaluation target across corpora. The corpora also differ in retained psycholinguistic content (Figure 1). Essays show broader document-level coverage of several affective categories; Facebook statuses show lower prevalence but sometimes substantial density when a category occurs; and student introductions contain comparatively frequent optimism-related language. Thus, the conditions vary in lexical composition as well as length and genre. Student Introduction Posts We collected BFI-44 responses (John et al. 1991) from students across two semesters of Georgia Tech’s OMSCS KBAI course. Of these, 226 also posted course-forum introductions, yielding complete paired survey–post records. Participation was voluntary and consented, procedures were IRB-approved, and names, emails, and links were removed before analysis. BFI-44 scores were converted to binary high/low labels using the midpoint procedure from prior work (Harbison et al. 2026), which is included in the Supplementary material. The cohort’s trait distributions were uneven, with marked right skew for all traits except Extraversion and Neuroticism; only Extraversion, Neuroticism, and Conscientiousness were sufficiently distributed for trait-level analysis. The introduction posts are a distinctive self-presentational genre: public educational forums encourage impression management (Hayes et al. 2024), and the icebreaker template solicited demographic and biographical details including location, courses, specialization, hobbies, motivation, and an interesting fact. 1. Matched prediction (RQ1+RQ2) Original inputs ⟶ 2. Explanation audit (RQ3) Content for explanation-enabled conditions ⟶ 3. Matched prediction (RQ1+RQ2) Ablated inputs ⟶ 4. Statistical evaluation (RQ1) Item-level association ρ and chance-corrected agreement κ ⟶ 5. Comparison (RQ2) Across evidence conditions, models, and prompts Figure 2: LEX-EC processing stages. Statistical analyses address RQ1, full-versus-masked comparisons address RQ2, and self-explanation analysis addresses RQ3. Free-form Essays The essay corpus is the Pennebaker–King stream-of-consciousness dataset (Pennebaker and King 1999): open-ended writing produced under a fixed-period, write-whatever-comes-to-mind protocol, paired with BFI labels. It serves as our long-form condition, averaging 672 word-tokens (≈3,295≈ 3,295 characters) per text after preprocessing. We report essay results at two scales. A 226-item subset, drawn by proportional stratification over joint five-trait label profiles with a fixed seed, preserves the full corpus’s trait-profile distribution and is used for pilot prompt and model comparisons; unlike the Facebook sample, it is a proportional downsample. The full corpus (n=2468) provides the more stable estimate and supersedes the pilot subset for all trait-level essay claims; the subset serves only to check that pilot behavior is consistent with sampling variability from the full set. Facebook Statuses The Facebook corpus derives from the myPersonality project (Celli et al. 2013), which linked Facebook status updates to user-level BFI questionnaire labels. After sampling one status per user with fixed seed 6767, the corpus provides a naturally occurring microtext condition averaging approximately 17 word-tokens (91 characters) per text after preprocessing. We simultaneously balanced the five marginal trait distributions using mixed-integer linear programming. Let P be the set of observed five-trait binary profiles, npn_p the number of users with profile p, and ajp∈0,1a_jp∈\0,1\ indicate whether profile p has a positive label for trait j. We selected the number zpz_p of users retained from each profile by solving max _z N=∑p∈zp N= _p z_p (1) s.t. zp∈0,…,np, z_p∈\0,…,n_p\, p∈, p , (0.5−δ)N≤∑p∈ajpzp, (5-δ)N≤ _p a_jpz_p, j∈, j , ∑p∈ajpzp≤(0.5+δ)N, _p a_jpz_p≤(5+δ)N, j∈, j , where T is the set of five traits and δ=0.05δ=0.05. With no fixed target size, the solution attained the largest feasible sample. We solved the program with scipy.optimize.milp using its HiGHS backend, sampled users without replacement within each profile, and used fixed seed 6767 for all sampling. The resulting dataset contains 164 unique users, with each trait’s positive-label proportion between 45% and 55%. Methodology Figure 2 summarizes the LEX-EC workflow. Across the three datasets, models assign binary BFI labels to original and content-masked texts under matched prompt conditions. We separately inspect item-level association and chance-corrected agreement, full-to-masked changes, and explanation content where collected. Each observation consists of one text from one author, and no model is fine-tuned on the target datasets. Model outputs in both conditions were collected through required tool calls.222Strict schema-constrained decoding in the API was avoided because it degraded association statistics in pilot tests. Unless otherwise stated, all results were each gathered via a single run. Model Selection Proprietary LLM evaluation presents a reproducibility problem: model families are revised frequently, older endpoints may be deprecated, and implementation details are generally unavailable. A result obtained from one model snapshot may therefore reflect transient provider- or version-specific behavior rather than a stable property of the task. We address this major reproducibility issue in the field as it applies to our work by evaluating five closed models spanning providers, model generations, and prediction paradigms: GPT-4o-mini (OpenAI 2024), GPT-4o (OpenAI 2024a), o3-mini (OpenAI 2025c), GPT-5.4-mini (OpenAI 2026b), and Claude Haiku 4.5 (Anthropic 2025).333Snapshots used: gpt-4o-mini-2024-07-18, gpt-4o-2024-11-20, o3-mini-2025-01-31, gpt-5.4-mini-2026-03-17, claude-haiku-4-5-20251001.444Models queried with temperature=0temperature=0 for non-reasoning models, however this was not possible for reasoning models. max_tokens=3000max\_tokens=3000, max_completion_tokens=3000max\_completion\_tokens=3000 used for reasoning models. Reasoning models were set to an effort of mediummedium. All used the default top=1top_p=1. No fixed seed was supplied. GPT-4o-mini and GPT-4o provide compact and larger-model baselines from the same pre-reasoning family; o3-mini and GPT-5.4-mini provide reasoning-oriented and newer-generation contrasts; and Claude Haiku provides a cross-provider comparison. The purpose is to test whether the main evidence-channel findings may persist across model snapshots that differ in scale, provider, and generation. Because the pilot comparisons showed similar dataset-level patterns, with limited descriptive variation in the direction and magnitude of Spearman’s r and Cohen’s κ across models, the full-corpus essay analysis was run with GPT-5.4-mini under the linguistic prompt rather than repeated for every model. Prompts Prompts directed the model to adopt an expert persona for this task; prompting was treated as a sensitivity factor rather than optimized for a single “best” formulation. We compared Basic (BASIC), requesting binary BFI classifications; Basic+Explanation (BASICEX), adding trait-level justifications; and Linguistic+Explanation (LINEX), directing justifications toward linguistic, stylistic, and affective rather than topical or demographic evidence. Although findings on the efficacy of expert personas are mixed, persona prompting can alter model output behavior, with effects that vary by use (Zheng et al. 2024; Kong et al. 2024; Luz de Araujo et al. 2025). BASICEX therefore matched LINEX in its expert-role and justification instructions, differing only in the instruction to prioritize linguistic evidence. A preliminary linguistic-only condition was dropped after showing no descriptive performance difference from its explanation-bearing counterpart, as was BASICEX, though it remained in use only as the explanation-bearing baseline for the targeted RQ3 self-explanation audit. Model Self-Explanations For explanation-enabled conditions, GPT-4o-mini provided a brief justification alongside each binary label. We analyzed explanations only for Extraversion in the student-introduction dataset, the sole post-level trait with an interpretable predictive signal; label imbalance or lack of trait discrimination made the remaining traits unsuitable for explanation analysis. This is therefore a targeted audit of one trait rather than a five-trait analysis. We manually coded, with a single annotator, GPT-4o-mini Extraversion explanations for the student introduction posts as linguistic or linguistic-adjacent (LIN), demographic or topical (DEM), combined (CMB), or other (OTH). We treat self-explanations as aggregate behavioral outputs only, as post hoc LLM explanations may be unfaithful (Jacovi and Goldberg 2020; Turpin et al. 2023; Madsen et al. 2024). We test whether the distribution of stated self-explanation types, such as linguistic or affective versus demographic or topical, shifts across prompt conditions. A shift toward linguistic justifications under the linguistic prompt indicates a change in explanation behavior, providing evidence of prompt sensitivity. Performance and Association Metrics Because the reference labels may be imbalanced in real-world data, we assess predictions using complementary measures of item-level association and chance-corrected agreement. Our primary measures are ρs=corr(rank(),rank(^)),κ=po−pe1−pe, _s=corr\! (rank(y),rank( y) ), κ= p_o-p_e1-p_e, (2) where y and y are the reference and predicted labels, respectively, pop_o is observed agreement, and pep_e is agreement expected from the observed marginals. For binary variables, Spearman’s ρ is equivalent to the phi coefficient, while Cohen’s κ measures agreement after accounting for marginal prevalence. We test dependence between predicted and reference labels using Fisher’s exact test for sparse 2×22× 2 tables, and Pearson’s χ2χ^2 test otherwise. We interpret p-values as evidence against independence rather than as measures of practical strength, which is assessed using the magnitudes of ρ and κ. LLM Personality Trait Prediction The first prediction stage used the unablated text: for each dataset, prompt condition, and model, the LLM produced binary BFI labels from the full text available in that condition, establishing pre-ablation performance, per Algorithm 1. Algorithm 1 Zero-shot BFI prediction 0: Texts X=xiX=\x_i\, model m, prompt condition c 0: Predictions y^i,t∈0,1 y_i,t∈\0,1\ and optional explanations ji,tj_i,t 1: T←O,C,E,A,NT←\O,C,E,A,N\ 2: Select prompts (sc,uc)(s_c,u_c) and format example FcF_c 3: Define required tool schema ScS_c for all t∈Tt∈ T, including justification when required by c 4: for each text xi∈Xx_i∈ X do 5: qi←uc++Fc++xiq_i← u_c +\!\!+F_c +\!\!+x_i 6: ai←ToolCall(m,sc,qi,Sc)a_i (m,s_c,q_i,S_c) 7: zi←Parsem(ai)z_i _m(a_i) 8: for each trait t∈Tt∈ T do 9: y^i,t←[low(zi[t].classification)=high] y_i,t \! [low(z_i[t]. classification)= high ] 10: if c requires explanations then 11: ji,t←zi[t].justificationj_i,t← z_i[t]. justification 12: end if 13: end for 14: Store predictions, optional explanations, model, prompt, dataset, and text identifier 15: end for 16: return (y^i,t,ji,t)\( y_i,t,j_i,t)\ Metric O C E A N Essays, full corpus: GPT-5.4-mini, LINEX ρ .168→.158[−.010].168→.158\;[-.010] .220→.121[−.099].220→.121\;[-.099] .187→.147[−.040].187→.147\;[-.040] .151→.150[−.001].151→.150\;[-.001] .115→.083[−.032].115→.083\;[-.032] κ .124→.145[+.021].124→.145\;[+.021] .220→.119[−.101].220→.119\;[-.101] .187→.144[−.043].187→.144\;[-.043] .135→.136[+.001].135→.136\;[+.001] .044→.026[−.018].044→.026\;[-.018] Facebook statuses: GPT-4o, BASIC (n=164n=164) ρ −.020→.066[+.086]-.020→.066\;[+.086] .163→.088[−.075].163→.088\;[-.075] .064→.028[−.036].064→.028\;[-.036] −.016→−.027[−.011]-.016→-.027\;[-.011] .184→−.058[−.242].184→-.058\;[-.242] κ −.020→.052[+.072]-.020→.052\;[+.072] .148→.044[−.104].148→.044\;[-.104] .063→.025[−.038].063→.025\;[-.038] −.016→−.022[−.006]-.016→-.022\;[-.006] .183→−.056[−.239].183→-.056\;[-.239] Posts: GPT-4o-mini, LINEX (n=226n=226) ρ −.045→−.035[+.010]-.045→-.035\;[+.010] −.026→−.111[−.085]-.026→-.111\;[-.085] .226→.062[−.164].226→.062\;[-.164] −.036→.045[+.081]-.036→.045\;[+.081] .024→.078[+.054].024→.078\;[+.054] κ −.031→−.032[−.001]-.031→-.032\;[-.001] −.023→−.068[−.045]-.023→-.068\;[-.045] .173→.058[−.115].173→.058\;[-.115] −.017→.044[+.061]-.017→.044\;[+.061] .009→.054[+.045].009→.054\;[+.045] Table 1: Primary association and agreement results before and after content ablation. Cells report Full → Ablated [Δ][ ], where Δ=Ablated−Full =Ablated-Full. O=Openness, C=Conscientiousness, E=Extraversion, A=Agreeableness, and N=Neuroticism. Shown are: full set essays with GPT-5.4-mini/LINEX after similar pilot results across models; posts with GPT-4o-mini/LINEX; Facebook with GPT-4o/BASIC, which showed the clearest pre-ablation signal and its subsequent attenuation. Deltas use coefficients rounded to three decimal places. Dataset (pairs) O C E A N Posts (10) 0→1;+.011/−.0010→1;\;+.011/-.001 0→0;−.071/−.0550→0;\;-.071/-.055 7→2;−.078/−.0937→2;\;-.078/-.093 0→0;+.081/−.0050→0;\;+.081/-.005 0→0;+.054/+.0220→0;\;+.054/+.022 Essays (10) 10→9;−.024/+.00410→9;\;-.024/+.004 4→1;−.044/−.0534→1;\;-.044/-.053 6→3;−.034/−.0396→3;\;-.034/-.039 8→7;+.014/+.0088→7;\;+.014/+.008 6→4;−.043/−.0496→4;\;-.043/-.049 Facebook (2) 0→0;+.069/+.0520→0;\;+.069/+.052 1→0;−.044/−.0741→0;\;-.044/-.074 0→0;+.006/.0000→0;\;+.006/.000 0→0;−.022/−.0200→0;\;-.022/-.020 2→0;−.173/−.1712→0;\;-.173/-.171 Table 2: Cross-model summary of changes after content ablation. Cells report significant Spearman associations, Full→ , followed by median Δρ/Δκ ρ/ κ across matched model–prompt pairs. Posts and reduced-set essays include five models under two prompts; Facebook includes GPT-4o under two prompts. Δ=Ablated−Full =Ablated-Full. Threshold for significance was considered to be ρ≈.1307ρ≈.1307 for Essays/Posts, ≈.153≈.153 for Facebook. Deltas use coefficients rounded to three decimal places; pairs with undefined ρ were omitted from median Δρ ρ. Content-Masking Ablation The preserved lexical layer was constructed from selected Empath Emotion Lexicon categories, selected to closely follow LIWC and the noted strong correlation between them. (Fast et al. 2016; Tausczik and Pennebaker 2010). The channel therefore draws on validated lexical categories and an established psycholinguistic framework, granting it a stronger a priori construct-validity basis than an ad hoc or generic bag-of-words method. NRC categories (Mohammad and Turney 2013) were additionally selected and added to the layer to address lexical coverage gaps of specific Emotion categories not granted by the Empath library. The retained psycholinguistic vocabulary is represented by, Vpsych=⋃c∈CEℓE(c)∪w∈WNRC:ℓN(w)∩K≠∅,V_psych= _c∈ C_E _E(c)\;∪\; \w∈ W_NRC: _N(w)∩ K≠ \, (3) where CEC_E contains the selected Empath affective and cognitive-style categories, K the retained NRC emotion and polarity categories, ℓE(c) _E(c) the Empath word set for category c, and ℓN(w) _N(w) the NRC tags assigned to w. Each text is transformed by, f=Λ∘σ∘Reconτ∘nlp∘ρprog∘ρpii∘ρtmpl∘ρurl,f= σ _τ _prog _pii _tmpl _url, (4) where ρurl _url removes URLs; ρtmpl _tmpl removes copied icebreaker questions after normalizing punctuation variants; ρpii _pii replaces PII placeholders; and ρprog _prog masks program and course identifiers, degree abbreviations, and course numbers with placeholder μ. Template removal applies only to student introductions, preventing copied prompts from being treated as authored evidence. After tokenization and POS tagging, tokens are mapped top-down by, τ(t)=t,E(t)∨Q(t),t,pos(t)∈ΠF∨low(t)∈WF,t,low(t),low(lem(t))∩Vpsych≠∅,μ,pos(t)=NUM,μ,otherwise.τ(t)= casest,&E(t) Q(t),\\ t,&pos(t)∈ _F (t)∈ W_F,\\ t,& \low(t),low(lem(t)) \∩ V_psych≠ ,\\ μ,&pos(t)=NUM,\\ μ,&otherwise. cases (5) Here, E(t)E(t) and Q(t)Q(t) identify emoji or emoticons and punctuation or whitespace; ΠF _F contains retained function-word POS classes; and WFW_F contains supplemental quantifiers, negations, degree and temporal adverbs, and wh-words. Numerals are masked because they may encode ages, years, or counts. ReconτRecon_τ restores original whitespace, σ collapses consecutive placeholders, and Λ removes empty or placeholder-only lines. The masking placeholder μ was instantiated as the literal underscore character ’_’ to avoid introducing lexical associations that may accompany replacement with unrelated words, or adjacency issues arising from simple word deletion. The resulting ablation is a retention test: all material outside the function-word, affective, cognitive-style, and structural channel is masked. Persistence indicates that an association remains recoverable from the retained channel; attenuation indicates that broader lexical content contributed to the unmasked association. Experiments Model Self-Explanations Under the basic explanation prompt, 113 of 226 explanations (50.0%) were coded LIN; under the linguistic prompt, this increased to 153 of 226 (67.7%). Explanations containing any demographic or topical reasoning (DEM or CMB) decreased from 113 of 226 (50.0%) to 72 of 226 (31.9%). Linguistic prompting shifted the content of the generated explanations without eliminating topical reasoning. Explanation type did not reliably distinguish correct from incorrect predictions: linguistic and demographic/topical justifications appeared in both groups. Linguistic prompting showed the same pattern. For example, sports and fitness references appeared equally often in correct-high and incorrect-high explanations under the basic prompt (14 each), and appeared across multiple outcome categories under the linguistic prompt. Prediction correlations also showed prompt sensitivity, but changes between the BASICEX and LINEX conditions varied across models, with no consistent advantage for either prompt. Prediction and Ablation Results Prediction patterns differed substantially across datasets (Tables 1 and 2). In the student introduction posts, Extraversion was the only trait with statistically detectable pre-ablation associations in a majority of model–prompt conditions (seven of ten). After masking, only two of ten Extraversion associations remained detectable. Both its association and agreement coefficients decreased in eight of ten conditions, with median paired changes of Δρ=−.078 ρ=-.078 and Δκ=−.093 κ=-.093. In the focal GPT-4o-mini/LINEX condition, Extraversion decreased from ρ=.226ρ=.226 to .062.062 and from κ=.173κ=.173 to .058.058. The other traits remained inconsistent across conditions, with generally negligible chance-corrected agreement and several constant or near-constant prediction distributions. Essays produced broader associations. In the reduced-set model comparison, positive pre-ablation coefficients appeared across multiple traits, although their magnitudes remained small. In the full-corpus GPT-5.4-mini analysis under the linguistic prompt, the Spearman associations for all five traits were statistically detectable both before and after masking. The largest decrease occurred for Conscientiousness (ρ: .220→.121.220→.121; κ: .220→.119.220→.119). Extraversion also decreased (ρ: .187→.147.187→.147; κ: .187→.144.187→.144), whereas Agreeableness changed little (ρ: .151→.150.151→.150; κ: .135→.136.135→.136). For Openness, the association was nearly unchanged (ρ: .168→.158.168→.158), while chance-corrected agreement increased from .124.124 to .145.145. Neuroticism had the weakest agreement both before and after masking (κ: .044→.026.044→.026). The Facebook results were weak across most model–prompt–trait combinations. Four of the 50 pre-ablation associations were statistically detectable: Conscientiousness under Claude Haiku/LINEX and GPT-4o/BASIC, and Neuroticism under GPT-4o/BASIC and GPT-4o/LINEX. Matched Facebook ablation analyses were available for the two GPT-4o prompts; after masking, none of their associations remained statistically detectable. In the focal GPT-4o/BASIC condition, Conscientiousness decreased from ρ=.163ρ=.163 and κ=.148κ=.148 to ρ=.088ρ=.088 and κ=.044κ=.044, while Neuroticism changed from ρ=.184ρ=.184 and κ=.183κ=.183 to negative values (ρ=−.058ρ=-.058, κ=−.056κ=-.056). Under the linguistic prompt, Neuroticism also decreased (ρ: .166→.062.166→.062; κ: .164→.062.164→.062). The remaining Facebook coefficients changed in both directions but were small before and after masking. Overall, masking most consistently attenuated Extraversion in the introduction posts and Conscientiousness in the full-corpus essay analysis. Openness and Agreeableness remained comparatively stable in the focal essay condition, while the limited pre-ablation Facebook associations were not maintained after masking. Across models and prompts, the reduced-set essay results varied in magnitude but showed the same broad trait-level pattern. In the Facebook condition, 59.8% of lexical tokens survived masking, comparable to the student posts. However, statuses were very short, the resulting inputs averaged only 10.2 retained tokens, including ≈ 1.9 psycholinguistic-vocabulary tokens. Post-ablation, mean retained word-token counts for Essays were 471, Posts 54, Facebook 10. Discussion Essays produced the broadest associations, but small chance-corrected agreement, indicating that statistical detectability did not translate into reliable individual-level classification. Student introductions yielded a narrower pattern centered on Extraversion, perhaps because this genre naturally elicits descriptions of social activities and preferences. Facebook performance remained weak despite balanced reference labels, showing a possible lower length/content boundary condition. The ablation results further suggest that different traits are consistent with dependence on different linguistic cues. In the full-corpus essay analysis, Conscientiousness was most attenuated, consistent with greater dependence on removed content-bearing vocabulary. Agreeableness was comparatively stable, while the near-stable Openness association accompanied by increased κ more likely reflects a change in prediction marginals than stronger trait information. In student introductions, the repeated decline in Extraversion across model–prompt conditions suggests sensitivity to comparatively explicit lexical cues. The disappearance of the isolated Facebook associations is likewise consistent with dependence on sparse lexical evidence, although the short retained inputs and limited matched ablation conditions preclude a stronger explanation. No model showed a stable advantage across datasets and conditions; model differences were less systematic than differences associated with genre, trait, and masking. This suggests that the pattern of findings may generalize to newer model families. Linguistic prompting increased linguistic explanations, while explanations containing demographic or topical reasoning decreased. Prompt framing therefore changed the distribution of stated justifications without directly establishing that those justifications reflected the evidence responsible for the classifications (Jacovi and Goldberg 2020; Turpin et al. 2023). The central result is not simply that performance was weak: associations differed in their sensitivity to lexical masking. Some remained observable from the retained function-word, psycholinguistic, and structural channel, whereas others attenuated when broader lexical evidence was removed. The broadly similar patterns across models suggest that weak performance in some conditions may reflect limitations of the available textual evidence rather than a deficiency specific to one model. These findings suggest that zero-shot personality prediction is best characterized as a property of a specific configuration of model, genre, trait, available word usage than as a general capability of a model. Persistence under lexical restriction is not inherently evidence of better personality prediction: it identifies signal recoverable from the retained channel, which may include differing stylistic contents. Personality-labeling systems should thus be evaluated separately for each intended text condition, rather than treated as possessing a general personality prediction capability. Limitations The explanation audit was manually coded by one annotator, so the reported explanation-category frequencies and prompt-related shifts lack an independent inter-rater reliability estimate. Exposure of closed models to the public personality corpora is unknown; results on those corpora cannot establish fully uncontaminated out-of-sample performance. The private student-introduction corpus was immune to this effect. Additionally, by removing all but psycholinguistic material during ablation, this may disrupt correspondence with data consumed during training. Conclusion Zero-shot BFI prediction from a single text is not a unitary capability: evidential profiles varied across corpora and traits, and lexical restriction produced heterogeneous changes. Linguistic prompting also shifted model self-explanations toward stylistic and affective evidence, indicating prompt sensitivity in explanation behavior. These results demonstrate the utility of LEX-EC as a task-specific black-box audit framework. By examining agreement, lexical-channel persistence, and explanation content together, our framework revealed information invisible to conventional statistical measures alone. AI Tool Use Disclosure The authors utilized GPT 5.5-5.6, Claude Opus 4.8, and Claude Fable 5 while editing text for clarity, grammar, and concision, and as non-decisional adversarial review aids to flag assumptions/methodological weaknesses in both the text and suggestions. Most non-copyediting suggestions were rejected as scientifically invalid/out-of-scope by the authors; suggestions that did inform revisions were human-verified. References Anthropic (2025) Claude Haiku 4.5 System Card. System Card Anthropic. Note: Available at https://w.anthropic.com/claude-haiku-4-5-system-card Cited by: Model Selection. Y. Belinkov (2022) Probing classifiers: promises, shortcomings, and advances. Computational Linguistics 48 (1), p. 207–219. External Links: Document Cited by: Introduction. A. Bhandarkar, R. Wilson, A. Swarup, G. D. Webster, and D. L. Woodard (2025) PsyTEx: a knowledge-guided approach to refining text for psychological analysis. In Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities, Albuquerque, USA, p. 151–178. External Links: Document, Link Cited by: LLM Personality Prediction, Black-box Interpretability. F. Celli, F. Pianesi, D. Stillwell, and M. Kosinski (2013) Workshop on computational personality recognition: shared task. Proceedings of the International AAAI Conference on Web and Social Media 7 (2), p. 2–5. External Links: Document Cited by: Facebook Statuses, Datasets. T. Compton (2025) Beyond the black box: integrating lexical and semantic methods in quantitative discourse analysis with BERTopic. External Links: 2508.19099, Document Cited by: Black-box Interpretability. E. Derner, D. Kučera, N. Oliver, and J. Zahálka (2024) Can ChatGPT read who you are?. Computers in Human Behavior: Artificial Humans 2 (2), p. 100088. External Links: Document Cited by: Introduction, LLM Personality Prediction. F. Doshi-Velez and B. Kim (2018) Considerations for evaluation and generalization in interpretable machine learning. In Explainable and Interpretable Models in Computer Vision and Machine Learning, H. J. Escalante, S. Escalera, I. Guyon, X. Baró, Y. Güçlütürk, U. Güçlü, and M. van Gerven (Eds.), The Springer Series on Challenges in Machine Learning, p. 3–17. External Links: Document, ISBN 978-3-319-98131-4 Cited by: Black-box Interpretability. E. Fast, B. Chen, and M. S. Bernstein (2016) Empath: understanding topic signals in large-scale text. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, p. 4647–4657. External Links: Document Cited by: Lexicon-based Textual Analysis: LIWC, Emotion, NRC, Content-Masking Ablation. S. Feng, E. Wallace, A. Grissom I, M. Iyyer, P. Rodriguez, and J. Boyd-Graber (2018) Pathologies of neural models make interpretations difficult. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, p. 3719–3728. External Links: Document Cited by: Black-box Interpretability. L. R. Goldberg (1990) An alternative “description of personality”: the Big-Five factor structure. Journal of Personality and Social Psychology 59 (6), p. 1216–1229. External Links: Document Cited by: Introduction. B. Harbison, S. Taubman, T. Taylor, and A. K. Goel (2026) Personality-enhanced social recommendations in SAMI: exploring the role of personality detection in matchmaking. In INTED2026 Proceedings, External Links: Document Cited by: Appendix A, Student Introduction Posts, Datasets. B. Hayes, A. Suleiman, and D. Watling (2024) Students’ impression management and self-presentation behaviours via online educational platforms: an archival review. First Monday 29 (3). External Links: Document Cited by: Student Introduction Posts. A. Jacovi and Y. Goldberg (2020) Towards faithfully interpretable NLP systems: how should we define and evaluate faithfulness?. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, p. 4198–4205. External Links: Document Cited by: Black-box Interpretability, Model Self-Explanations, Discussion. O. P. John, E. M. Donahue, and R. L. Kentle (1991) The Big Five Inventory—versions 4a and 54. Technical report University of California, Berkeley, Institute of Personality and Social Research, Berkeley, CA. Cited by: Personality Measures and the Big Five Inventory, Student Introduction Posts. O. P. John and S. Srivastava (1999) The big five trait taxonomy: history, measurement, and theoretical perspectives. In Handbook of Personality: Theory and Research, L. A. Pervin and O. P. John (Eds.), p. 102–138. Cited by: Personality Measures and the Big Five Inventory. A. Kong, S. Zhao, H. Chen, Q. Li, Y. Qin, R. Sun, X. Zhou, E. Wang, and X. Dong (2024) Better zero-shot reasoning with role-play prompting. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, p. 4099–4113. External Links: Link, Document Cited by: Prompts. A. Liimatta (2022) Do registers have different functions for text length? a case study of Reddit. Register Studies 4 (2), p. 263–287. External Links: Document Cited by: Lexicon-based Textual Analysis: LIWC, Emotion, NRC. A. Liimatta (2023) Register variation across text lengths: evidence from social media. International Journal of Corpus Linguistics 28 (2), p. 202–231. External Links: Document Cited by: Lexicon-based Textual Analysis: LIWC, Emotion, NRC. S. M. Lundberg and S. Lee (2017) A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30, p. 4765–4774. Cited by: Introduction, Black-box Interpretability. P. H. Luz de Araujo, P. Röttger, D. Hovy, and B. Roth (2025) Principled personas: defining and measuring the intended effects of persona prompting on task performance. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 26857–26886. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: Prompts. A. Madsen, S. Chandar, and S. Reddy (2024) Are self-explanations from large language models faithful?. In Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, p. 295–337. External Links: Document Cited by: Black-box Interpretability, Model Self-Explanations. J. Maharjan, R. Jin, J. Zhu, and D. Kenne (2025) Psychometric evaluation of large language model embeddings for personality trait prediction. Journal of Medical Internet Research 27, p. e75347. External Links: Document Cited by: LLM Personality Prediction. S. C. Matz, M. Kosinski, G. Nave, and D. J. Stillwell (2017) Psychological targeting as an effective approach to digital mass persuasion. Proceedings of the National Academy of Sciences 114 (48), p. 12714–12719. External Links: Document Cited by: Introduction. R. R. McCrae and P. T. Costa (1987) Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology 52 (1), p. 81–90. External Links: Document Cited by: Personality Measures and the Big Five Inventory. R. R. McCrae and O. P. John (1992) An introduction to the five-factor model and its applications. Journal of Personality 60 (2), p. 175–215. External Links: Document Cited by: Introduction. S. M. Mohammad and P. D. Turney (2013) Crowdsourcing a word–emotion association lexicon. Computational Intelligence 29 (3), p. 436–465. External Links: Document Cited by: Lexicon-based Textual Analysis: LIWC, Emotion, NRC, Content-Masking Ablation. E. S. Öhman and A. Liimatta (2024) Text length and the function of intentionality: a case study of contrastive subreddits. In Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities, Miami, USA, p. 1–8. External Links: Document, Link Cited by: Lexicon-based Textual Analysis: LIWC, Emotion, NRC. C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter (2020) Zoom in: an introduction to circuits. Distill 5 (3). External Links: Document Cited by: Introduction. OpenAI (2024) GPT-4o mini: advancing cost-efficient intelligence. Note: OpenAI model releaseAvailable at https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Cited by: Model Selection. OpenAI (2024a) GPT-4o System Card. System Card OpenAI. Note: Available at https://cdn.openai.com/gpt-4o-system-card.pdf Cited by: Model Selection. OpenAI (2026b) GPT-5.4 Thinking System Card. System Card OpenAI. Note: Appendix: GPT-5.4 mini, added March 17, 2026. Available at https://deploymentsafety.openai.com/gpt-5-4-thinking/gpt-5-4-thinking.pdf Cited by: Model Selection. OpenAI (2025c) OpenAI o3-mini System Card. System Card OpenAI. Note: Available at https://cdn.openai.com/o3-mini-system-card-feb10.pdf Cited by: Model Selection. J. W. Pennebaker and L. A. King (1999) Linguistic styles: language use as an individual difference. Journal of Personality and Social Psychology 77 (6), p. 1296–1312. External Links: Document Cited by: Personality Measures and the Big Five Inventory, Free-form Essays, Datasets. H. Peters and S. C. Matz (2024) Large language models can infer psychological dispositions of social media users. PNAS Nexus 3 (6), p. pgae231. External Links: Document Cited by: Introduction, LLM Personality Prediction. M. Piastra and P. Catellani (2025) On the emergent capabilities of ChatGPT 4 to estimate personality traits. Frontiers in Artificial Intelligence 8, p. 1484260. External Links: Document, Link Cited by: Introduction, LLM Personality Prediction. H. Rao, C. Leung, and C. Miao (2023) Can ChatGPT assess human personalities? a general evaluation framework. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, p. 1184–1194. External Links: Document Cited by: Introduction, LLM Personality Prediction. M. T. Ribeiro, S. Singh, and C. Guestrin (2016) “Why should i trust you?”: explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, p. 1135–1144. External Links: Document Cited by: Introduction, Black-box Interpretability. M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh (2020) Beyond accuracy: behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, p. 4902–4912. External Links: Document Cited by: Introduction, Black-box Interpretability. P. Schoenegger, S. Greenberg, A. Grishin, J. Lewis, and L. Caviola (2025) AI can outperform humans in predicting correlations between personality items. Communications Psychology 3, p. 23. External Links: Document Cited by: Introduction, LLM Personality Prediction. H. A. Schwartz, J. C. Eichstaedt, M. L. Kern, L. Dziurzynski, S. M. Ramones, M. Agrawal, A. Shah, M. Kosinski, D. J. Stillwell, M. E. P. Seligman, and L. H. Ungar (2013) Personality, gender, and age in the language of social media: the open-vocabulary approach. PLOS ONE 8 (9), p. e73791. External Links: Document Cited by: Personality Measures and the Big Five Inventory. Y. R. Tausczik and J. W. Pennebaker (2010) The psychological meaning of words: LIWC and computerized text analysis methods. Journal of Language and Social Psychology 29 (1), p. 24–54. Cited by: Lexicon-based Textual Analysis: LIWC, Emotion, NRC, Content-Masking Ablation. M. Turpin, J. Michael, E. Perez, and S. R. Bowman (2023) Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems 36, p. 74952–74965. Cited by: Black-box Interpretability, Model Self-Explanations, Discussion. A. M. van der Veen and E. Bleich (2025) The advantages of lexicon-based sentiment analysis in an age of machine learning. PLOS ONE 20 (1), p. e0313092. External Links: Document Cited by: Black-box Interpretability. A. Vinciarelli and G. Mohammadi (2014) A survey of personality computing. IEEE Transactions on Affective Computing 5 (3), p. 273–291. External Links: Document Cited by: Introduction. A. G. C. Wright, W. R. Ringwald, C. E. Vize, J. C. Eichstaedt, M. Angstadt, A. Taxali, and C. Sripada (2026) Assessing personality using zero-shot generative AI scoring of brief open-ended text. Nature Human Behaviour 10 (3), p. 541–555. External Links: Document, Link Cited by: Introduction, LLM Personality Prediction. T. Yang, T. Shi, F. Wan, X. Quan, Q. Wang, B. Wu, and J. Wu (2023) PsyCoT: psychological questionnaire as powerful chain-of-thought for personality detection. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, p. 3305–3320. External Links: Document, Link Cited by: LLM Personality Prediction. T. Yarkoni (2010) Personality in 100,000 words: a large-scale analysis of personality and word use among bloggers. Journal of Research in Personality 44 (3), p. 363–373. External Links: Document Cited by: Personality Measures and the Big Five Inventory. M. Zheng, J. Pei, L. Logeswaran, M. Lee, and D. Jurgens (2024) When “a helpful assistant” is not really helpful: personas in system prompts do not improve performances of large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 15126–15154. External Links: Link, Document Cited by: Prompts. Appendix A Appendix Midpoint for Student Introduction Posts Before use, the student post dataset was converted from the original continuous data to binary high/low. Each personality-trait score was dichotomized using the fixed scale midpoint of τ=3τ=3. For participant i and personality trait j, the binary classification cijc_ij was defined as, cij=(xij>3)=0,if xij≤3,1,if xij>3,c_ij=I(x_ij>3)= cases0,&if x_ij≤ 3,\\ 1,&if x_ij>3, cases (6) where xijx_ij denotes the original personality-trait score and I denotes the indicator function. Thus, scores less than or equal to the scale midpoint were assigned to class 0, whereas scores above the midpoint were assigned to class 1. Past work utilizing this dataset (Harbison et al. 2026), instead used a random midpoint procedure, which can introduce stochastic label noise, and was therefore avoided in this work. Additional Dataset Distributions This section includes additional tables. Table 3 demonstrates the distribution of the Student Introduction Post dataset, introduced in Section "Datasets" in the main body. Table 4 shows the distribution data for the marginally balanced single-post-per-user subset, while Table 5 and Table 6 show the distribution data for the full essay set and for the subset, respectively. Trait Value Proportion O 1 0.898230 (89.82%) 0 0.101770 (10.18%) C 1 0.840708 (84.07%) 0 0.159292 (15.93%) E 1 0.438053 (43.81%) 0 0.561947 (56.19%) A 1 0.871681 (87.17%) 0 0.128319 (12.83%) N 1 0.429204 (42.92%) 0 0.570796 (57.08%) Table 3: Distribution of binary Big Five personality labels in the student introduction-post dataset (N=226N=226). Trait abbreviations: O = Openness, C = Conscientiousness, E = Extroversion, A = Agreeableness, and N = Neuroticism. Trait Value Proportion cOPN 1 0.548780 (54.88%) 0 0.451220 (45.12%) cCON 1 0.542683 (54.27%) 0 0.457317 (45.73%) cEXT 1 0.518293 (51.83%) 0 0.481707 (48.17%) cAGR 1 0.548780 (54.88%) 0 0.451220 (45.12%) cNEU 1 0.451220 (45.12%) 0 0.548780 (54.88%) Table 4: Binary division of classification buckets for BFI traits from the myPersonality single-post-per-user subset (N=164N=164). Trait abbreviations follow those used by the underlying myPersonality dataset: cOPN = Openness, cCON = Conscientiousness, cEXT = Extroversion, cAGR = Agreeableness, and cNEU = Neuroticism. Trait Value Proportion cOPN 1 0.515397 (51.54%) 0 0.484603 (48.46%) cCON 1 0.508104 (50.81%) 0 0.491896 (49.19%) cEXT 1 0.517423 (51.74%) 0 0.482577 (48.26%) cAGR 1 0.530794 (53.08%) 0 0.469206 (46.92%) cNEU 1 0.499595 (49.96%) 0 0.500405 (50.04%) Table 5: Distribution of binary Big Five personality labels in the full, original Pennebaker-King Essays dataset (N=2468N=2468). Trait abbreviations follow those used by the underlying Pennebaker-King Essays dataset: cOPN = Openness, cCON = Conscientiousness, cEXT = Extroversion, cAGR = Agreeableness, and cNEU = Neuroticism. Trait Value Proportion cOPN 1 0.513274 (51.33%) 0 0.486726 (48.67%) cCON 1 0.508850 (50.88%) 0 0.491150 (49.12%) cEXT 1 0.517699 (51.77%) 0 0.482301 (48.23%) cAGR 1 0.535398 (53.54%) 0 0.464602 (46.46%) cNEU 1 0.500000 (50.00%) 0 0.500000 (50.00%) Table 6: Distribution of binary Big Five personality labels in the smaller Pennebaker-King Essays subset (N=226N=226, seed=67seed=67). Trait abbreviations: cOPN = Openness, cCON = Conscientiousness, cEXT = Extroversion, cAGR = Agreeableness, and cNEU = Neuroticism. Lexicon Categories for Ablation This section shows the full list of retained lexicon subcategories from Empath and NRC, shown in Table 7 Complete Results The full statistical output of the described experiments are included here. Lexicon source Category group Retained subcategories Empath Positive affect positive_emotion, joy, cheerfulness, contentment, affection, love, optimism, pride, zest Empath Negative affect negative_emotion, sadness, disappointment, suffering, torment, nervousness, fear, timidity, anger, rage, aggression, irritability, exasperation, hate, disgust, shame Empath Social-affective stance sympathy, emotional, warmth Empath Social/normative language swearing_terms Empath Cognitive style thinking, order, confusion, anticipation, deception NRC Core emotions and polarity joy, sadness, anger, fear, disgust, surprise, positive, negative Table 7: Psycholinguistic and affective lexicon categories retained during ablation. Examples of Pre-Ablated and Post-Ablated Text Below is a contiguous essay excerpt before and after content masking. The model received the complete essay in each condition; only the corresponding excerpt is shown. Each underscore represents one or more consecutive masked tokens. Original Essay Excerpt: I’m not saying that I’m perfect. but I’ve learned over the past few years about what I want out of life and what I don’t want. I’m living my life the way I want to. as stress-free as possible and as happy as possible. When I’m put into these stupid situations it just makes life that much harder and it sucks! Ablated: I’m not _ that I’m perfect. but I’ve learned over the _ few _ about what I _ out of _ and what I don’t _ . I’m _ my _ the _ I _ to. _ stress-_ as _ and _ happy as _ . When I’m _ into these stupid _ it just _ much harder and it sucks! Prompts Listing 1 shows the required format JSON example. Listing 2 shows the tool call definition that was then used in the model API call to return model outputs in the required structured format. The classification field was required in every condition. The justification field was required for the Basic+Ex and Linguistic+Ex conditions and omitted for the Basic condition. Listing 1: Unified JSON response format. The justification field is included only in explanation-required conditions. ⬇ 1expected_format = """ 2 3 "Openness": 4 "classification": "low/high" 5 , 6 "Conscientiousness": 7 "classification": "low/high" 8 , 9 "Extroversion": 10 "classification": "low/high" 11 , 12 "Agreeableness": 13 "classification": "low/high" 14 , 15 "Neuroticism": 16 "classification": "low/high" 17 18 19""" Listing 2: Abridged structured tool schema used in the Basic+Ex and Linguistic+Ex conditions. The remaining four traits used the same object structure as Openness. ⬇ 1tools = [ 2 3 "type": "function", 4 "function": 5 "name": "record_personality_profile", 6 "description": "Records the Big Five personality traits and a justification for each classification.", 7 "parameters": 8 "type": "object", 9 "properties": 10 "Openness": 11 "type": "object", 12 "description": "The analysis for the Openness trait.", 13 "properties": 14 "classification": 15 "type": "string", 16 "enum": ["low", "high"] 17 , 18 "justification": 19 "type": "string", 20 "description": "Reasoning based on a direct quote from the text." 21 22 , 23 "required": [ 24 "classification", 25 "justification" 26 ] 27 28 29 # Structure repeats for the other traits 30 , 31 "required": [ 32 "Openness", 33 "Conscientiousness", 34 "Extroversion", 35 "Agreeableness", 36 "Neuroticism" 37 ] 38 39 40 41] Table 8 shows the full prompts, system and user, used in the study. Condition Basic System message You are an expert in inferring Big-5 personality traits from text. User message Analyze this text for personality traits. Classify each Big Five trait as either ’low’ or ’high’. Return your results in the following JSON format without explanation: Condition Basic+Ex System message You are an expert in inferring Big-5 personality traits from text. Your justifications are always evidence-base and grounded in direct quotes from the text. User message Your task is to analyze the provided text, classify the author’s Big Five personality traits, and provide a detailed justification for your analysis. Classify each Big Five trait as either ’low’ or ’high’, no deviation allowed. Return your detailed justification and results in the following JSON format: Condition Linguistic+Ex System message You are an expert in inferring Big-5 personality traits from linguistic patterns in text. User message Analyze this text for personality traits based only on writing style, word choice, and language patterns. Classify each Big Five trait as either ’low’ or ’high’. If the linguistic evidence is insufficient, make your best inference based on available patterns, but always provide a classification, and note this in your explanation. Return your explanation and results in the following JSON format: Table 8: Prompt conditions used for zero-shot Big Five prediction. Prompts are reproduced as used. Each user message was followed by an expected JSON-format example and the target text sample. Outputs were collected through required structured tool calls. Model Self-Explanation Results Below is an example of an original student introduction text, the model’s prediction, and the explanation for its reasoning the model provided. Incidentally, this was a correct prediction, though the self-explanation listed topical content. This is one example of an explanation tagged ’DEM’. Introduction post: “Hello! My name is [NAME], and I currently live in [LOCATION]. This is my first [PROGRAM] course! I currently work as a software engineer at [ORGANIZATION]. Outside of work, I like to play sports and hang out with friends/family. From this class, I hope to develop a strong foundation in [FIELD]. One interesting fact about me is that I also have a [SECOND OCCUPATION].“ Prediction: High Extraversion Explanation: “The text expresses enjoyment of social activities, such as playing sports and spending time with friends and family, indicating sociability.“ Below is another example, this time tagged as ’LIN’. To preserve anonymous review, identifying spans in the displayed example were replaced with placeholders. The model received the complete original post. Interestingly, the model’s prediction was correct in this case, as well. Introduction post: “Hello, I am [NAME]. Currently, I am living in [LOCATION], and I am a [OCCUPATION] by day and a [DEGREE PROGRAM] student by night. I am specifically interested in this course, as I would like to learn more about [FIELD] and determine how to implement it in my current research. I took [COURSE] and [COURSE] last semester. That was a fun, occasionally stressful, and excellent learning experience. I am looking forward to this semester, and I hope to learn a lot. :)“ Prediction: Low Extraversion Explanation: “The text does not include social or outgoing language; it focuses more on personal academic interests rather than social interactions.“ Table 9 shows the counts for text content broad category markers for correct/incorrect and across total predictions. Additionally, it shows the confusion matrix of cue counts for text content. Prompt Explanation-tag counts Topical cues by prediction outcome Correct Incorrect Total Correct high Correct low Incorrect high Incorrect low Basic LIN 68; DEM 48; CMB 19 LIN 45; DEM 29; CMB 17 LIN 113; DEM 77; CMB 36 S/F 14; family (3), board games (2), friends (2), band (2), concerts (2), dog (2) S/F 3; kids (1), Netflix (1), coding (1), gaming (1) S/F 14; friends (4), gaming (4), concerts (3) S/F 10; family (3), traveling (2) Linguistic LIN 87; DEM 29; CMB 7; OTH 1 LIN 66; DEM 25; CMB 11; OTH 0 LIN 153; DEM 54; CMB 18; OTH 1 S/F 34; family (4), band (3), board games (3), concerts (3), friends (3), dog (2) S/F 19; reading (7), gaming (5), chess (2), dogs (2), Netflix (2) S/F 19; family (6), travel (5), friends (3), concerts (3) S/F 6; gaming (3), board/card games (3) Table 9: GPT-4o-mini Extroversion explanation audit for student introduction posts. Tag columns report explanation-category counts; topical-cue columns report sports/fitness and other cue frequencies. LIN = linguistic; DEM = demographic/topical; CMB = combined; OTH = indeterminate. Full Prediction and Ablation Results This section presents the remaining full Spearman r/ cohen’s k results across all models that were not included in the main body, for Posts, Essays, and Facebook. Only Ablation results for GPT-4o were included as other models either failed to output results reaching significance or failed to complete the task consistently at all. Pre-Ablated Predictions from Text Table 10 shows the full results on the non-ablated text, for all 3 datasets. Prompt Model O C E A N Posts (n=226n=226) BASIC GPT-4o -.039/-.024 .055/.027 .101/.088 .101/.074 -.058/-.009 GPT-4o-mini -.032/-.017 .055/.027 .209∗/.153∗ .105/.049 .077/.012 5.4-mini -.022/-.009 .033/.019 .162∗/.151∗ .105/.049 -.082/-.018 o3-mini -.032/-.017 -.058/-.033 .211∗/.178∗ -.026/-.009 -??? Claude Haiku -.022/-.009 -.041/-.017 .238∗/.165∗ -??? -??? LINEX GPT-4o -.056/-.044 .033/.019 .091/.088 .110/.094 -??? GPT-4o-mini -.045/-.031 -.026/-.023 .226∗/.173∗ -.036/-.017 .024/.009 5.4-mini -??? -.058/-.033 .138∗/.137 -.063/-.046 -.022/-.006 o3-mini .066/.045 -.058/-.033 .101/.088 -.058/-.039 -??? Claude Haiku -.022/-.009 .033/.019 .205∗/.136∗ -??? -??? Essays (n=226n=226) BASIC GPT-4o .221∗/.172∗ .139∗/.122 .102/.097 .200∗/.195∗ .133∗/.062 GPT-4o-mini .209∗/.142∗ .115/.097 .101/.100 .144∗/.137∗ .198∗/.142∗ 5.4-mini .263∗/.198∗ .149∗/.149∗ .151∗/.150∗ .193∗/.178∗ .102/.062 o3-mini .269∗/.178∗ .126/.106 .111/.106 .128/.125 .168∗/.097∗ Claude Haiku .132∗/.047 .114/.104 .183∗/.183∗ .165∗/.144∗ .124/.062 LINEX GPT-4o .215∗/.162∗ .150∗/.131∗ .180∗/.179∗ .198∗/.198∗ .119/.071 GPT-4o-mini .204∗/.133∗ .087/.071 .158∗/.157∗ .185∗/.164∗ .175∗/.097∗ 5.4-mini .263∗/.198∗ .115/.114 .151∗/.151∗ .165∗/.145∗ .133∗/.062 o3-mini .271∗/.188∗ .152∗/.131∗ .126/.122 .109/.106 .166∗/.106∗ Claude Haiku .180∗/.094∗ .126/.113 .147∗/.141∗ .185∗/.165∗ .089/.044 Facebook (per-user, n=164n=164) BASIC GPT-4o -.020/-.020 .163∗/.148 .064/.063 -.016/-.016 .184∗/.183∗ GPT-4o-mini -.015/-.015 .087/.051 .088/.088 .078/.078 .081/.078 5.4-mini -.025/-.024 .071/.061 .024/.024 .046/.046 .076/.076 o3-mini -.007/-.007 .078/.064 -.000/-.000 .049/.048 .117/.110 Claude Haiku -.060/-.056 .127/.119 .015/.014 -.042/-.040 .094/.094 LINEX GPT-4o .053/.051 .118/.099 .060/.060 .037/.037 .166∗/.164 GPT-4o-mini .056/.056 .071/.041 -.022/-.022 .116/.115 .067/.064 5.4-mini ??? .122/.106 -.038/-.038 -.010/-.009 .083/.082 o3-mini .051/.051 .062/.052 .041/.040 -.088/-.085 .098/.096 Claude Haiku .004/.004 .167∗/.149∗ .080/.064 -.058/-.053 .080/.080 Table 10: Cells: Spearman ρ/Cohen’s κ. Stars on Spearman p, on κ from χ2χ^2 p: p∗<.05^*p<.05, p∗<.01^**p<.01, p∗∗<.001^***p<.001. O=Openness, C=Conscientiousness, E=Extroversion, A=Agreeableness, N=Neuroticism. Facebook ablation run for GPT-4o only. Values of ??? indicate single class prediction only. Post-Ablated Predictions from Text Table 11 shows the final full results on the ablated text, for all 3 datasets. Prompt Model O C E A N Posts (n=226n=226) BASIC GPT-4o -.001/-.001 -.050/-.048 .017/.006 -??? .052/.023 GPT-4o-mini .015/.013 -.045/-.044 .103/.038 -??? .031/.014 5.4-mini .049/.036 -.048/-.048 .153∗/.139∗ .070/.056 .024/.009 o3-mini .145∗/.142 -.118/-.097 .139∗/.114 .082/.082 .052/.023 Claude Haiku -.026/-.025 -.057/-.057 .061/.033 -.036/-.017 .027/.012 LINEX GPT-4o -.045/-.045 -.004/-.004 .098/.097 .072/.072 .092/.062 GPT-4o-mini -.035/-.032 -.111/-.068 .062/.058 .045/.044 .078/.054 5.4-mini -.008/-.008 -.050/-.048 .075/.029 -??? .027/.012 o3-mini -.038/-.038 -.001/-.001 .110/.090 .059/.059 .051/.020 Claude Haiku -.026/-.025 -.057/-.057 .061/.033 -.036/-.017 .027/.012 Essays (n=226n=226) BASIC GPT-4o .224∗/.215∗ .086/.057 .076/.057 .073/.053 .030/.009 GPT-4o-mini .160∗/.139∗ .038/.023 .064/.062 .148∗/.148∗ .144∗/.053 5.4-mini .175∗/.146∗ .078/.077 .091/.088 .182∗/.176∗ .137∗/.080 o3-mini .171∗/.115∗ .045/.024 .033/.027 .112/.094 .110/.035 Claude Haiku .180∗/.094∗ .084/.071 .152∗/.152∗ .247∗/.241∗ .116/.027 LINEX GPT-4o .216∗/.183∗ .139∗/.092 .194∗/.172∗ .230∗/.223∗ .128/.044 GPT-4o-mini .245∗/.232∗ .106/.050 .069/.064 .096/.095 .144∗/.053 5.4-mini .161∗/.151∗ .087/.085 .144∗/.139∗ .231∗/.215∗ ??? o3-mini .128/.100 .071/.040 .119/.097 .132∗/.111 .096/.035 Claude Haiku .192∗/.104∗ .091/.072 .110/.102 .220∗/.214∗ .134∗/.035 Facebook (per-user, n=164n=164) BASIC GPT-4o .066/.052 .088/.044 .028/.025 -.027/-.022 -.058/-.056 LINEX GPT-4o .104/.083 .105/.055 .108/.098 .004/.004 .062/.062 Table 11: Cells: Spearman ρ/Cohen’s κ. Stars on Spearman p, on κ from χ2χ^2 p: p∗<.05^*p<.05, p∗<.01^**p<.01, p∗∗<.001^***p<.001. O=Openness, C=Conscientiousness, E=Extroversion, A=Agreeableness, N=Neuroticism. Facebook ablation run for GPT-4o only. Values of ??? indicate single class prediction only.