Paper deep dive
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
Yuan Yuan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 7/5/2026, 1:01:32 PM
Summary
The paper introduces a 'dual-nature' framework for LLM personas, demonstrating that persona expression consists of two dissociable components: frame-robust aggregated features (e.g., Big Five scores) and frame-dependent geometric features (e.g., SPD manifold structures). Using GPT-4o to simulate American and Chinese-American personas via IPIP-50, the researchers found that while aggregate scores are stable under question randomization, the geometric correlation structure collapses under frame misalignment (randomized order) but recovers significantly when a shared frame is reintroduced (bootstrap shared frame). This suggests that LLM persona geometry is a coordination pattern dependent on temporal alignment rather than an intrinsic, static trait.
Entities (7)
Relation Signals (5)
Big Five â istypeof â Aggregated Feature
confidence 100% ¡ aggregated features (Big Five scores) degrade under randomization (21% drop) but are frame-robust
SPD Manifold â istypeof â Geometric Feature
confidence 100% ¡ geometric features (SPD manifold) collapse under frame misalignment (42% drop)
IPIP-50 â measures â Big Five
confidence 100% ¡ Constructing within-instance correlation matrices from IPIP-50 responses... aggregated features (Big Five scores)
GPT-4o â simulates â American Persona
confidence 100% ¡ analyzing geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American personas.
GPT-4o â simulates â Chinese-American Persona
confidence 100% ¡ analyzing geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American personas.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. We test whether this geometric structure is intrinsic or frame-dependent. Constructing within-instance correlation matrices from IPIP-50 responses, we analyze geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American personas. We find that persona expression comprises two dissociable components: aggregated features (Big Five scores) degrade under randomization (21% drop) but are frame-robust; geometric features (SPD manifold) collapse under frame misalignment (42% drop) but recover substantially (to 84%) under shared frames, surpassing aggregated features (76%). This collapse-recovery pattern reveals that persona geometry is not intrinsic but a frame-dependent coordination pattern encoding information invisible to aggregation. Our findings establish a dual-nature framework for LLM personas, frame-dependent geometry versus frame-robust aggregates, necessitating frame-aware evaluation and challenging static trait conceptions.
Tags
Links
- Source: https://arxiv.org/abs/2607.02368v1
- Canonical: https://arxiv.org/abs/2607.02368v1
Trouble viewing inline? Open PDF directly â
Full Text
57,927 characters extracted from source content.
Expand or collapse full text
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry YUAN YUAN 1 yzy0014@auburn.edu Abstract Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. We test whether this geometric structure is in- trinsic or frame-dependent. Constructing within- instance correlation matrices from IPIP-50 re- sponses, we analyze geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American per- sonas. We find that persona expression comprises two dissociable components: aggregated features (Big Five scores) degrade under randomization (21% drop) but are frame-robust; geometric features (SPD manifold) collapse under frame misalign- ment (42%drop) but recover substantially (to 84%) under shared frames, surpassing aggregated features (76%). This collapse-recovery pattern reveals that persona geometry is not intrinsic but a frame-dependent coordination pattern encod- ing information invisible to aggregation. Our findings establish a dual-nature framework for LLM personasâframe-dependent geometry versus frame-robust aggregatesânecessitating frame-aware evaluation and challenging static trait conceptions. 1. Introduction The use of psychometric questionnaires (e.g., IPIP, Big Five) to investigate LLM personas is a well-established Yuan Yuan 1 Independent Researcher.Correspondence to: <yzy0014@auburn.edu>. This paper was submitted to ICML 2026 but has been withdrawn by the authors and is published on Arxiv as an independent preprint.>. Proceedings of the43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s). paradigm, consistently revealing systematic response pat- terns and biases (Safdari et al., 2023; Jiang & Zou, 2024; Argyle et al., 2023). Furthermore, the geometric factor struc- ture of personality is a psychological cornerstone, and this structure reliably emerges in aggregate LLM response data (Liu et al., 2023). However, a critical gap persists: pre- vailing research operates almost exclusively on aggregate feature averages (e.g., mean dimension scores), collapsing the within-instance correlation structure that defines indi- vidual differences. This renders the recovered âpersonaâ a sample-level artifact and obscures a fundamental question. This omission is especially pressing given the autoregressive nature of LLMs (Radford et al., 2019; Vaswani et al., 2017). The standard practice of using a fixed question order con- flates two potential sources of the observed geometry: is it a stable, intrinsic trait of the model, or is it merely an epiphe- nomenon of a specific, shared temporal frame during mea- surement? Consequently, the pivotal inquiry is not whether a geometric structure existsâit is mathematically givenâbut what its nature is, and why its frame-dependence has been systematically neglected. 1.1. Research Question and Hypotheses We hypothesize that the apparent stability of geometric bias structure is an artifact of fixed frames. To test this, we propose three competing hypotheses that capture distinct possibilities in the literature: 1.Intrinsic Structure (H1): Geometric features capture stable model properties, as assumed in trait-based per- sonality assessment (McCrae & John, 1992). 2.Measurement Artifact (H2): The apparent structure is spurious and destroyed by perturbation, analogous to order effects in survey methodology (Schuman & Presser, 1996). 3.Frame-Dependent Coordination (H3): Geometric features encode relational patterns that require tempo- ral alignmentâa novel hypothesis motivated by LLMsâ autoregressive nature (Radford et al., 2019). 1 arXiv:2607.02368v1 [stat.ML] 2 Jul 2026 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry Visual predictions. Figure 1 plots the predicted cluster- ing accuracy (y-axis) across our three analytical condi- tionsâFixed Order (FO), Random Order Native Frame (RO), and Random Order Bootstrap Shared Frame (RO- BTSP)âfor each hypothesis. H1 predicts a high, flat line; H2 predicts high accuracy only in FO with collapse in both randomized conditions; and H3 uniquely predicts a V-shaped pattern: collapse in RO (frame misalignment) followed by recovery in RO-BTSP (shared frame realign- ment). We systematically tests these hypotheses through controlled order manipulation and geometric analysis. Fixed Order (Frame-Aligned) Random Order (Frame Misaligned) RO-Bootstrap (Frame Realigned) Experimental Conditions 40 50 60 70 80 90 100 Predicted Clustering Accuracy (%) CollapseRecovery H1: Intrinsic Structure (Failed) H2: Measurement Artifact (Failed) H3: Frame-Dependence (Supported) Figure 1. Differential predictions for geometric bias features. Only frame-dependence (H3) predicts the observed collapse-recovery V-shape pattern across conditions (FOâ ROâ RO-BTSP). 1.2. Methodological Innovation and Dual-Nature Discovery To test these hypotheses and dissociate content from tem- poral structure effects, we introduce a novel methodolog- ical framework. We develop the Item-Dimension Matrix method, constructing within-instance correlation matrices from questionnaire responses, enabling geometric analy- sis on the manifold of symmetric positive definite (SPD) matrices. Critically, we systematically manipulate ques- tion ordering to test whether geometric features exhibit the invariance expected of intrinsic structures. Our investigation reveals a fundamental dissociation: all geometry-based features (SPD manifold, eigenvalues, eigenvectors) catastrophically collapse under frame mis- alignment, but substantially recover under shared frames. This reversible collapse indicates that what appears as âbias geometryâ is not a static property but a frame-dependent coordination pattern that can only be measurable in tem- poral alignment. In stark contrast, aggregation-based fea- tures (Big Five scores) show the opposite sensitivity: they are robust to frame misalignment but degrade under content randomization. This clean dissociation reveals that bias in LLM output space comprises two dissociable components: one tied to how dimensions coordinate during sequence processing (ge- ometry), and the other reflecting what values are typically generated (aggregation). This discovery challenges the pre- vailing monolithic view of bias, establishing a dual nature framework for understanding LLM bias: as simultaneously a frame-dependent coordination pattern and a frame-robust aggregate tendency. Our Contributions: â˘Empirical:We demonstrateâvia a novel item- dimension matrix and bootstrap protocolâthat the geometric structure of LLM bias is not intrinsic but frame-dependent: it collapses under misalignment but recovers under shared frames, revealing a dual nature distinct from aggregated tendencies. â˘Theoretical: We establish a dual-nature framework for LLM bias, distinguishing between frame-dependent coordination geometry and frame-robust aggregated tendencies. â˘Methodological: We propose a new standard for LLM evaluation that emphasizes frame-aligned analysis and explicit decomposition of order vs. frame effects. Be- yond substantive findings, we introduce: (1) a ma- trix approach based on item-dimensions within each LLM instance allowing geometric analysis of individ- ual LLM responses, (2) systematic manipulation of temporal frames to dissociate intrinsic structure from coordination artifacts, and (3) validation across sam- ple sizes (N â 100andN = 2000, see Appendix A) demonstrating robustness and clarifying optimal sam- ple selection for geometry analysis bias. 2. Related Work 2.1. LLM Personality Assessment Recent work has systematically applied personality inven- tories to LLMs (Safdari et al., 2023; Jiang & Zou, 2024). These studies typically adopt human psychometric assump- tions without questioning their applicability to autoregres- sive architectures. Our work directly tests these assumptions through sequence manipulation. 2.2. Order Effects in Measurement Human assessment shows modest order effects (typically < 10%) attributed to cognitive consistency mechanisms (Schuman & Presser, 1996). LLMs likely operate differently through context accumulation rather than self-consistency maintenance. 2 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry 2.3. Temporal Effects in LLMs While recent work has documented position-dependent bi- ases in LLMs due to attention mechanisms (Vaswani et al., 2017), these studies typically focus on local context effects (e.g., recency bias) rather than systematic geometric struc- tures emerging from sequential coordination. Our work extends this line by asking whether the global correlational geometry observed in persona assessments is itself frame- dependent, a question that has not been addressed in the prior literature on temporal effects. 2.4. Geometric Methods and Methodological Rationale We draw methodological inspiration from functional con- nectivity (FC) analysis in neuroscience, where correlation matrices capture how brain regions coordinate over time (Barachant et al., 2013). This framework better captures dynamic systems than static trait models. Similarly, we hypothesize LLM bias involves inter-dimensional coordi- nation during autoregressive generationâhow personality dimensions covary in sequenceâinvisible to dimension- wise aggregation. Correlation matrices naturally lie on the manifold of sym- metric positive definite (SPD) matrices, where Riemannian metrics (e.g., log-Euclidean (Arsigny et al., 2007)) provide principled distances respecting manifold geometry. SPD methods have proven effective in brain-computer interfaces (Barachant et al., 2013) and computer vision (Huang & Van Gool, 2017), but assume the geometric structure is in- trinsic. Our work tests this assumption for LLMs. Critically, recent studies document that LLMs exhibit position-dependent biases due to attention mechanisms and autoregressive processing (Vaswani et al., 2017). Unlike hu- man cognitive consistency effects, these architectural prop- erties may create frame-dependent structures. By sys- tematically manipulating temporal frames, we test whether bias geometry is intrinsic or an artifact of measurement alignmentâa question that has not been addressed in prior geometric analysis. 3. Method 3.1. Experimental Design and Data Generation Model and Instrument We used the OpenAI API (gpt-4o-2024-05-13, temperature=0.7) to collect re- sponses, targeting100LLM calls per cell 1 ; we retained only complete, well-formed answers (see the Appendix C.3 for 1 This sample size was selected to balance discriminable cul- tural signals with sufficient data for stable correlation estimation, while avoiding over-aggregation effects that dilute group differ- ences (see Appendix A for large-sample validation atN = 2000 demonstrating robustness of findings and rationale for this choice). details). After filtering, the final counts were FO (US=96, CA=97, total=193), RO (US=92, CA=95, total=187). This yields a balanced design with sufficient power for our hy- pothesis tests. Data Collection For each LLM call, we first simulated American or Chinese-American personas through cultural prompts (see the Appendix C.1). Following cultural induc- tion, we administer the 50-item International Personality Item Pool (IPIP-50) (Goldberg, 1992). Items were adapted to first-person statements for LLM comprehension (com- plete list in Appendix C.2). Cultural Persona Induction and Justification We se- lected cultural bias as an experimental platform because it provides a well-documented and robust signal for testing geometric representations. Previous work establishes that LLMs exhibit distinct response patterns when simulating American versus Chinese-American perspectives (Jiang & Zou, 2024; Santurkar et al., 2023). These established dif- ferences in bias content provide a strong signal for testing whether geometric representations vary independently of our core manipulation: temporal structure (question ordering). Although cultural identity is multifaceted, this binary classi- fication serves as a controlled testbed for frame-dependence mechanisms, not as exhaustive cultural representation. 3.2. Item-Dimension Matrix and Correlation Construction Methodological Rationale Our research question, whether geometric bias structures are intrinsic properties or frame-dependent artifacts, requires an analytic frame- work that captures within-instance correlation patterns while enabling systematic temporal frame manipulation. Traditional aggregate methods (Big Five means) collapse within-instance structure, while factor analysis on pooled data (Liu et al., 2023) recovers only sample level patterns. We develop the Item-Dimension Matrix approach to con- struct instance-specific correlation matricesC (Ď) that en- code how dimensions covary during sequential generation under orderingĎ. This conceptually parallels functional con- nectivity analysis in neuroscience (Barachant et al., 2013), where correlation matrices capture temporal coordination. Crucially, different orderingsĎproduce differentC (Ď) re- sponses from identical responses, directly operationalizing the frame variation while maintaining the content constant. Since correlation matrices reside on the SPD manifold, we map them to the tangent space at identity vialog(C)(Ar- signy et al., 2007), enabling standard Euclidean operations while respecting the manifold structure. 3 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry Constructing item-dimension matrix using Multivariate Time Series from Questionnaire ResponsesWe concep- tualize the sequential response process as a multivariate (5-channel) time series. Each channel corresponds to one of the Big Five personality dimensions. Time progresses with the presentation of each question in the order Ď. For a given instance with a response vectorr (i) âR 50 and a specific question orderĎ, we construct a10Ă 5Item- Dimension Matrix X (Ď) as follows: 1.Each of the50IPIP items is pre-mapped to one of the five dimensions via a fixed function dim(¡). 2. We iterate through the sequenceĎ. At time stept (corresponding to thet-th question inĎ, denotedĎ[t]), we obtain the numerical score r(Ď[t]). 3.We place this scorer(Ď[t])in the column ofX (Ď) that corresponds to the dimensionj = dim(Ď[t]). The score is appended to the next available row position within that column, preserving its temporal order of occurrence in Ď. 4.After processing all 50 questions inĎ, each of the 5 columns contains exactly 10 scoresâall responses for that dimensionâin the exact order in which they were encountered during the sequence Ď. Thus, the columnjofX (Ď) represents the temporal re- sponse sequence for dimensionjas it was sampled in- termittently during orderĎ. Different ordersĎproduce different temporal arrangements of the same 10 responses within each column. Computing Within-Instance Correlation Matrices From the matrixX (Ď) , we compute the5Ă5within-instance Pearson correlation matrix: C (Ď) = corr(X (Ď) ),(1) whereC (Ď) jk quantifies how the response sequence for dimen- sionjco-varies with the sequence for dimensionkunder the specific temporal frame defined by Ď. Operationalizing Temporal Frame Conditions This construction directly enables our three analytical conditions: â˘FO (Fixed Order): All instances useĎ std , aligning their temporal frames: C (i) FO = C (Ď std ) . â˘RO (Random Order, Native Frame): Each instance uses a uniqueĎ (i) , creating a frame misalignment situ- ation: C (i) RO = C (Ď (i) ) . â˘RO-BTSP (Random Order, Bootstrap Shared Frame): For each bootstrap iterationb, a randomĎ b is drawn and used to recomputeC (Ď b ) for all RO in- stances, imposing a shared frame: C (i) BTSP,b = C (Ď b ) . This methodology isolates the effect of temporal coordina- tion from the response content. 3.3. Instruments Selection and Methodology Validation InstrumentSelection Theconstructionofwell- conditioned correlation matrices for the LLM instanceC (Ď) requires a matrix of the full dimension of the itemX (Ď) . Many popular personality instruments fail this requirement. The Ten-Item Personality Inventory (TIPI) (Gosling et al., 2003) provides only 2 items per dimension, insufficient for a reliable correlation estimation within an instance. The IPIP-NEO-300 (Johnson, 2014) measures30facets with10 items each, producing a matrix10Ă 30with too few items per facet. The balanced structure of IPIP-50 (10 itemsĂ 5 dimensions) provides full-rankX (Ď) while maintaining comparability with previous research on the LLM persona (Goldberg, 1992). Validity of the Item-Dimension Matrix Approach The item-dimension matrix reorganizes the responses from the validated IPIP-50 (Goldberg, 1992) to allow the analysis of correlations within the instance. This is a data organization method, not a new psychometric instrumentâit preserves all information from original responses while enabling geo- metric analysis on SPD manifolds. Its validity rests on: (1) the established psychometric properties of IPIP-50 and (2) the full-rank structure (10Ă 5) that ensures well-conditioned correlation matrices. The approach is designed to test frame- dependence hypotheses by systematically manipulating tem- poral order while preserving response content. Advantages Over Traditional Approaches Traditional personality assessment relies on aggregate scores, discard- ing the within-instance correlation structure. Our item- dimension matrix enables geometric analysis at the instance level, analogously to functional connectivity analysis in neuroscience (Barachant et al., 2013). This approach is necessary because: (1) it preserves temporal coordination information lost in aggregation; (2) it allows testing of frame-dependence via order manipulation; (3) it provides a mathematically principled framework (SPD manifolds) for analyzing correlation structures. 3.4. Feature Extraction and Evaluation Mapping to SPD Manifold Tangent Space The correla- tion matricesC (Ď) lie in the SPD manifold. We map them to a local Euclidean tangent space using the logarithmic map at a chosen reference point. 4 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry Why the identity matrixI? We fix the reference point asIfor two reasons: (1) Theoretical canon: In the log- Euclidean framework,Iis the identity element of the SPD Lie group, serving as the natural origin (Arsigny et al., 2007). (2) Hypothesis alignment: A data-dependent reference (e.g. Riemannian mean) would itself vary with the temporal frame (FO vs. RO), conflating frame effects with reference shifts. UsingIprovides a fixed, frame-invariant origin that cleanly isolates the geometric impact of Ď. WithIas reference, the map simplifies to the matrix loga- rithm: Log I (C) = log(C),(2) yielding symmetric tangent-space matrices. We vectorize log(C)(exploiting symmetry) to obtain 10-D feature vec- tors for subsequent Euclidean analysis. This approach pre- serves the geometry of the manifold while ensuring that the observed differences directly reflect frame-dependent coordination. Feature Extraction From each correlation matrixC (Ď) , we extract four types of features that span the aggregation- geometry spectrum: ⢠Big Five Scores: Dimensional mean ofr (i) (5D). Frame-independent by construction. ⢠SPD Manifold Features: Tangent space representation vec(log(C (Ď) )) (10D after exploiting symmetry). ⢠Eigenvalues: Spectrum Îť(C (Ď) ) (5D). ⢠Top Eigenvector: v 1 (C (Ď) ) (5D). These features test our predictions: geometric features (SPD, eigenvalues, eigenvectors) should exhibit the collapse- recovery pattern under frame-dependence (H3), while ag- gregated features (Big Five scores) should not. Clustering Evaluation For the main study, we first re- duce the dimension of features using UMAP (McInnes et al., 2018) (nneighbors=15, mindist=0.1), then apply spectral clustering (Ng et al., 2002) with clustersk = 2. For large-sample validation (Appendix A.2), we apply k-means clustering directly on raw features for computational effi- ciency; PCA visualizations are provided for interpretability but not used in clustering. In both cases, clustering accu- racy is computed as the proportion of correctly assigned instances (maximizing over label permutations). We report clustering accuracy, silhouette scores (Rousseeuw, 1987), and AUC-ROC where applicable. 4. Results 4.1. Frame-Dependent Geometry: Collapse and Recovery We tested the three competing hypotheses outlined in Fig- ure 1 by analyzing clustering performance under analytical conditions of FO, RO, and RO-BTSP. Table 1 presents the clustering accuracy in conditions, revealing a striking disso- ciation between feature types. Table 1. Clustering Performance Across Conditions (B = 2000 bootstrap iterations) Feature Accuracy (%) SD RO-BTSP FORORO-BTSP Big Five96.8975.90â â SPD95.3452.9484.5013.7 Eigenvalues61.1450.2759.208.85 Eigenvector50.7850.2763.1010.6 â Big Five scores are frame invariant; RO-BTSP = RO (75.90%). Note: SPD features capture full correlation geometry; eigenval- ues lose phase information, and eigenvectors lack discrimina- tive power for this binary task. Full statistics in Appendix B. The Collapse-Recovery Pattern SPD manifold features collapse under native-frame randomization (RO, 52.94%) but recover substantially under shared frames (RO-BTSP, 84.50%), t(1999) = 102.69, p < .001, d = 2.30. However, in shared frames condition (RO-BTSP), SPD performance (M = 84.48%) exceeds Big Five scores (75.90%) computed from the same randomized responses, t(1999) = 27.94, p < .001, d = 0.63, with86.8%of bootstrap iterations showing the advantage 2 . These results demonstrate that temporal coordination pat- terns encode discriminative information invisible to aggrega- tion, a fundamental limitation of aggregate-based evaluation in autoregressive models. The superior performance of SPD features over eigenvalues/eigenvectors suggests that the full correlation geometry preserves discriminative information that spectral decompositions partially discard. Visualizing Geometric Collapse Figure 2 summarizes the differential sensitivity of geometric versus aggregated fea- tures under the three analytical conditions. Figures 3a and 3b visualize the effect through UMAP pro- jections. It is clear that under FO, SPD features show clear separation (Silhouette =0.69, AUC =0.98), while under RO, clusters overlap substantially (Silhouette =0.29, AUC =0.61), demonstrating the collapse of geometric discrim- inability when frames are misaligned. It is also worth noting 2 . A similar Collapse-Recovery-Surpass pattern of SPD fea- tures was also observed in large-sample replication, Table 5, Ap- pendix A. 5 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry that the LLM responses data collected under RO condition show a substantial increase in variances compared to the data from FO condition. AccuracyARIAUC Fixed OrderRandom Order (Pres Corr) Bootstrap Average (Shared Random) Fixed OrderRandom Order (Pres Corr) Bootstrap Average (Shared Random) Fixed OrderRandom Order (Pres Corr) Bootstrap Average (Shared Random) 0 25 50 75 100 0 25 50 75 0 30 60 90 Performance Value Feature Type Big FiveSPD IdentityEigenvaluesEigenvectors Fixed Order vs Random (Pres Corr) vs Bootstrap (Shared Random) with 95% CI MultiâMetric Performance Across Three Conditions Figure 2. Performance across three analytical conditions. Geo- metric features (SPD, Eigen.) collapse under frame misalignment (RO) but recover under shared frames (RO-BTSP), while aggre- gated features (Big Five) show opposite sensitivity. 4.2. Decomposing Order and Frame Effects To quantify distinct vulnerabilities, we analyze performance degradation from Fixed Order (FO) to Random Order Native Frame (RO). The total degradation is âTotal =|Acc RO â Acc FO |,(3) which could be further decomposed into two components: the order effect (OE): degradation due solely to content randomization when frames are aligned, âOE =|Acc RO-BTSP â Acc FO |,(4) and the frame effect (FE): additional degradation caused by frame misalignment beyond content randomization, âFE =|Acc RO â Acc RO-BTSP |.(5) The relative contributions ofâOE andâFE to the total degradation are the following. âOE% = âOE/âTotal,âFE% = âFE/âTotal. (6) Table 2. Decomposition of Performance Degradation Relative Contribution (%) FeatureâTotalâOEâFEDominant Big Fiveâ20.991000Order SPDâ42.402674Frame Eigenvaluesâ10.871981Frame Eigenvectorsâ0.51âBalanced Note: OE = order effect, FE = frame effect. Percentages rounded; The Inversion: Geometry is Frame-Driven, Aggregation is Order-DrivenTable 2 reveals a stark dissociation: SPD features are predominantly frame-driven (74% FE, 26% OE), â2â101234 â4 â2 0 2 4 6 Big Five (Silhouette = 0.642, Acc = 96.9%) â FIXED ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification â4â202468 â4 â2 0 2 4 6 8 SPD Identity Matrix (Silhouette = 0.685, Acc = 95.3%) â FIXED ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification â4â202468 â10 â5 0 5 Eigenvalues (Silhouette = 0.721, Acc = 61.1%) â FIXED ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification â10â5051015 â4 â2 0 2 4 6 8 Eigenvectors (Silhouette = 0.819, Acc = 50.8%) â FIXED ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification (a) Fixed Order (FO) â2â1012 â4 â2 0 2 4 Big Five (Silhouette = 0.468, Acc = 70.1%) â RANDOM ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification â2â1012 â2 â1 0 1 2 SPD Identity Matrix (Silhouette = 0.288, Acc = 52.9%) â RANDOM ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification â2â1012 â4 â2 0 2 4 Eigenvalues (Silhouette = 0.102, Acc = 50.3%) â RANDOM ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification â3â2â10123 â4 â2 0 2 4 Eigenvectors (Silhouette = 0.212, Acc = 50.3%) â RANDOM ORDER UMAP Dimension 1 UMAP Dimension 2 American (True) Chinese American (True) Correct Classification Misclassification (b) Random Order (RO) Figure 3. UMAP visualizations of SPD features under (a) Fixed Order (clear separation) and (b) Random Order Native Frame (collapsed overlap). Colors indicate true cultural group (American vs. Chinese-American). whereas Big Five scores are purely order-driven (100% OE, 0% FE). This confirms that geometric representations are vulnerable to measurement misalignment, while aggregated features are affected only by content randomization. 4.3. Data Quality: Structure Persists Under Randomization A potential alternative explanation is that randomization destroys the underlying correlation structure, producing ran- dom noise matrices. We refute this using Random Matrix Theory (RMT): eigenvalue spacing in both FO and RO follows the WignerâDyson ensemble (Wigner, 1958), con- firming preserved non-random structure despite increased entropy (t(378) =â8.69,p < 10 â13 ). (Figure 4): ⢠Entropy: Item-level response entropy increases sig- nificantly under randomization (t(378) =â8.69,p < 10 â13 ), confirming effective perturbation. 6 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry Fixed (American) Random (American) Fixed (Chinese) Random (Chinese) 0.0 0.5 1.0 1.5 ItemâLevel Entropy (Response Distribution) Shannon Entropy (bits) Fixed (American) Random (American) Fixed (Chinese) Random (Chinese) 0.0 0.2 0.4 0.6 0.8 1.0 1.2 ItemâLevel Response Variance Variance Fixed (American) Random (American) Fixed (Chinese) Random (Chinese) 1.6 1.7 1.8 1.9 2.0 2.1 2.2 Correlation Matrix Entropy (Eigenvalue Distribution) Entropy (bits) Maximum Entropy Fixed (American) Random (American) Fixed (Chinese) Random (Chinese) 0.35 0.40 0.45 0.50 0.55 0.60 First Eigenvalue Variance Explained (Structure Strength) Proportion of Variance Uniform (20%) Eigenvalue Distribution Fixed Order Eigenvalue Frequency 0.00.51.01.52.02.5 0 20 40 60 80 100 Eigenvalue Distribution Random Order Eigenvalue Frequency 0.00.51.01.52.02.53.0 0 10 20 30 40 50 60 70 Eigenvalue Spacing Distribution Fixed Order Normalized Spacing (s) Probability Density 01234 0.0 0.5 1.0 1.5 Fixed Order Data Poisson (Random) WignerâDyson (GOE) Eigenvalue Spacing Distribution Random Order Normalized Spacing (s) Probability Density 01234 0.0 0.2 0.4 0.6 0.8 Random Order Data Poisson (Random) WignerâDyson (GOE) 0.00.51.01.52.02.53.0 0 1 2 3 4 QâQ Plot: Eigenvalue Spacing Fixed vs Random Order Fixed Order Spacing Quantiles Random Order Spacing Quantiles Figure 4. Persistence of correlation structure under randomiza- tion. Top: Increased response entropy confirms effective order perturbation. Bottom: Eigenvalue spacing follows Wigner-Dyson ensemble, indicating preserved correlations despite randomization. â˘Eigenvalue Spacing: The distribution of eigenvalue spacings in both FO and RO conditions follows the Wigner-Dyson ensemble, indicating preserved system- wide correlation patterns characteristic of non-random matrices (Wigner, 1958). The preserved eigenvalue spacing indicates that random- ization perturbs, but does not erase, the underlying cor- relational geometry, ruling out the alternative explanation that RO collapse is due to structural destruction rather than frame misalignment. Thus, the correlational structure is not erased by randomization; the collapse in RO is due to misalignment, not structural dissolutionâconsistent with the frame-dependence hypothesis. 5. Discussion 5.1. The Dual Nature of LLM Persona Our results reveal that what appears as âpersonalityâ in LLMs comprises two dissociable components. Geometric features (SPD manifold, eigenvalues, eigenvectors) show strong frame-dependenceâcollapsing under misaligned question orders (RO) but recovering substantially when frames are realigned (RO-BTSP). In contrast, aggregate scores (Big Five means) remain largely order-stable. This clean dissociation demonstrates that bias in LLM outputs is not unitary: it emerges both as frame-dependent co- ordination geometry and as frame-robust aggregated tendencies. 5.2. Why Geometry Outperforms Aggregation SPD geometry surpasses Big Five scores under shared frames (84.50%vs.75.90%,p < .001), revealing that in- ter dimensional coordination encodes aggregation-invisible information (Pennec et al., 2006). This parallels functional connectivity in neuroscience: patterns emerge only with temporal alignment (Friston, 2011). We claim that SPD captures emergent computational connectivityâa transient coordination states that require frame alignment for consis- tency (Vaswani et al., 2017). 5.3. LLMs Lack Stable Trait Structure The catastrophic collapse of geometric features under frame misalignment (SPD: 95.34%â52.94%) indicates that what is measured as âpersonalityâ in LLMs is largely a measure- ment artifact of the fixed-order protocol, not an intrinsic, order-invariant structure analogous to human traits (Costa & McCrae, 1992; Roberts et al., 2007). Human personal- ity exhibits cross-situational consistency; LLM responses emerge through context-conditioned autoregressive gener- ation (Brown et al., 2020). Thus, the apparent âpersonaâ is better understood as temporally scaffolded response co- herenceâa pattern that emerges only when measurement frames are aligned, not as a stable trait property. 5.4. Implications for Evaluation Our findings necessitate a shift toward frame-aware evalu- ation. Current fixed-order protocols risk conflating frame artifacts with stable traits (Safdari et al., 2023). Rigorous assessment should vary temporal frames, decompose or- der/frame effects, and report alignment-condition perfor- mance. Consequently, âLLM personalityâ scores should be interpreted as measurement-contingent regularities, not as revealed intrinsic traits. 6. Limitations and Future Work Scope and GeneralizabilityOur study uses GPT-4o with a focused sample size (N â 400) optimized for signal preservationâa choice validated by large-sample replica- tion (N = 2000) showing consistent effects. However, several scope boundaries warrant consideration: â˘Model scope: The frame-dependence mechanism may vary across architectures (e.g., Llama, Gemini) and model scales. While we hypothesize it is inher- ent to autoregressive generation, this requires cross- architectural verification. â˘Cultural scope: We use American versus Chinese- American personas as a well-documented testbed (Jiang & Zou, 2024). Although this binary provides a clean signal, it does not capture the full spectrum of cultural variation. Future work should test collectivist vs. individualist cultures across diverse regions. â˘Bias domain: Our findings may generalize to other bias dimensions (political, gender, etc.), but this needs 7 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry empirical confirmation. Different bias types may ex- hibit distinct frame-sensitivity patterns. Mechanistic Underpinnings and Theoretical Pathway Our study establishes frame-dependence as a fundamental property of LLM persona measurement, opening a theo- retical pathway toward understanding how autoregressive architectures produce temporally scaffolded coherence. Fu- ture work should examine: â˘Semantic frame effects: Whether ordered (e.g., by valence) vs. random orderings elicit different coordina- tion patterns. ⢠Architectural causes: How positional encodings and attention dynamics produce frame sensitivity (Vaswani et al., 2017). â˘Neural correlates: Whether the behavior of SPD geometry mirrors functional connectivity in in- ternal activations.Recording layer-wise snap- shots and computing neuron/attention-head correla- tion matrices could test if neural SPD manifolds show the same collapse-recovery pattern, grounding frame-dependence in the transformerâs computational substrate (Friston, 2011; Barachant et al., 2013; Sporns, 2013). This would establish the dependence of the frame on the computational substrate of the trans- former, revealing how the functional computational connectivity emerges from the dynamics of atten- tion/feedbackâbridging behavioral measurement with mechanistic interpretability research. Evaluation Implications If bias partly reflects dynamic coordination patterns, mitigation may need to target sequence-generation processes beyond output distributions. Developing standardized frame-aware evaluation proto- colsâreporting performance under multiple orderings and decomposing order versus frame effectsâwould improve fairness auditing and model comparisons. 7. Conclusion We demonstrate that LLM âpersonalityâ is not unitary but comprises two dissociable components: geometric struc- ture (frame-dependent, 78% degradation from misalign- ment) and aggregated tendencies (frame-robust, 100% order- driven). Unlike human traits (Costa & McCrae, 1992), geometric representations are coordination artifacts of au- toregressive generation (Brown et al., 2020), not intrinsic structures. This dual nature necessitates frame-controlled evaluation: valid bias assessment requires distinguishing sta- ble tendencies from ephemeral coordination patterns. Our framework provides a rigorous foundation for robust LLM evaluation and AI safety. Accessibility Upon acceptance, the datasets, code, and documentation necessary to reproduce the core findings of this study will be made publicly available in accordance with the conference guidelines. This includes the response data, the analysis pipelines and the experimental protocols used in both the main experiments and the validation studies. Impact Statement This paper presents work whose goal is to advance the field of machine learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here. Acknowledgments We thank anonymous reviewers for their helpful feedback. This work was not supported by an external Funding Source. References Argyle, L. P., Busby, E. C., Gubler, J. R., Howe, T., Rytting, C., and Sorensen, T. Out of one, many: Using language models to simulate human samples. Political Analysis, 31 (3):337â351, 2023. doi: 10.1017/pan.2023.2. Arsigny, V., Fillard, P., Pennec, X., and Ayache, N. Logarith- mic maps and exponentials in the set of positive definite symmetric matrices: A survey. Journal of Mathematical Imaging and Vision, 31(2):93â105, 2007. Barachant, A., Bonnet, S., Congedo, M., and Jutten, C. Classification of covariance matrices using a riemannian- based kernel for bci applications. Neurocomputing, 112: 172â178, 2013. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33: 1877â1901, 2020. Costa, P. T. and McCrae, R. R. Normal personality assess- ment in clinical practice: The NEO personality inventory. Psychological Assessment, 4(1):5â13, 1992. Friston, K. J. Functional and effective connectivity: a review. Brain Connectivity, 1(1):13â36, 2011. Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., and Ahmed, N. K. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3):1097â1179, 2024. 8 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry Goldberg, L. R. The development of markers for the big-five factor structure. Psychological assessment, 4(1):26â42, 1992. Goldberg, L. R. A broad-bandwidth, public domain, person- ality inventory measuring the lower-level facets of several five-factor models. Personality Psychology in Europe, 7 (1):7â28, 1999. Gosling, S. D., Rentfrow, P. J., and Swann Jr, W. B. A very brief measure of the big-five personality domains. Journal of Research in personality, 37(6):504â528, 2003. Huang, Z. and Van Gool, L. A riemannian network for spd matrix learning. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAIâ17, p. 2036â2042, 2017. Jiang, G. and Zou, J. Cultural personality in llms: A cross- linguistic analysis. arXiv preprint arXiv:2401.x, 2024. Johnson, J. A. Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Devel- opment of the IPIP-NEO-120. Journal of Research in Per- sonality, 51:78â89, 2014. doi: 10.1016/j.jrp.2014.05.003. Liu, Y., Liska, A., Gallego, A., Dhamala, J., Jyothi, P., and Gurevych, I. LLM-Factor: A statistical framework for uncovering latent structure from large language models. arXiv preprint, 2023. URLhttps://arxiv.org/ abs/2310.14791. arXiv:2310.14791. McCrae, R. R. and John, O. P. An introduction to the five- factor model and its applications. Journal of Personality, 60(2):175â215, 1992. McInnes, L., Healy, J., and Melville, J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1â35, 2021. doi: 10.1145/3457607. Ng, A., Jordan, M., and Weiss, Y. On spectral clustering: Analysis and an algorithm. Advances in Neural Informa- tion Processing Systems, 14, 2002. Pennec, X., Fillard, P., and Ayache, N. A riemannian frame- work for tensor computing. International Journal of Computer Vision, 66(1):41â66, 2006. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019. Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., and Goldberg, L. R. The power of personality: The com- parative validity of personality traits, socioeconomic sta- tus, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4): 313â345, 2007. Rousseeuw, P. J. Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20:53â65, 1987. Safdari, M., Serapio-Garcia, G., Crepy, C., Fitz, S., Romero, P., Sun, L., Abdullahi, M., Faust, A., and Matari Ě c, M. Personality traits in large language models. arXiv preprint arXiv:2307.00184, 2023. Santurkar, S., Durmus, E., Ladhak, F., Lee, C., and Liang, P. Whose opinions do language models reflect? arXiv preprint, 2023. URLhttps://arxiv.org/abs/ 2303.17548. arXiv:2303.17548. Schuman, H. and Presser, S. Questions and Answers in Atti- tude Surveys: Experiments on Question Form, Wording, and Context. Sage Publications, 1996. Sporns, O. Structure and function of complex brain net- works. Dialogues in Clinical Neuroscience, 15(3):247â 262, 2013. Suresh, H. and Guttag, J. V. A framework for understanding sources of harm throughout the machine learning life cy- cle. Proceedings of the 2021 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO â21), 2021. doi: 10.1145/3465416.3483305. Also available as arXiv:1901.10002. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser,Ĺ., and Polosukhin, I. Atten- tion is all you need. Advances in Neural Information Processing Systems, 30, 2017. Wigner, E. P. On the distribution of the roots of certain symmetric matrices. Annals of Mathematics, p. 325â 327, 1958. 9 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry A. Pilot Study: Large Sample Experiments We conducted pilot experiments with larger samples, following the same 2 x 2 factorial design, targeting a totalN = 2000 LLM API calls, with 500 API calls for each cell. The large sample experiment generated a total of1931complete LLM responses, withN FO = 960for fixed order condition (N US = 473,N CA = 487), andN RO = 971for random order condition (N US = 485,N CA = 486). Our purpose was to assess the sample size effects on feature discriminability. A.1. Descriptive Statistics Fixed-Order ConditionTable 3 presents the mean scores, standard deviations, and results of independent samplest-tests for each Big Five dimension under the Fixed-Order condition, under American (US) cultural prompt and Chinese-American (CA) cultural prompts, respectively, with a large sample size (N = 960). Table 3. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Fixed-Order Condition, N US = 473, N CA = 487) M US SD US M CA SD CA tCohenâs d Extraversion3.900.093.580.1442.04 â 2.71 Agreeableness4.010.134.020.12-1.55-0.10 Conscientiousness3.920.143.940.10-1.83-0.12 Neuroticism3.030.093.020.151.650.11 Openness3.580.083.550.123.90 â 0.25 Random Order Condition Table 4 presents the mean scores, standard deviations, and results of independent samples t-tests for each Big Five dimension under the Random-Order condition, under American (US) cultural prompt and Chinese- American (CA) cultural prompts, respectively, with a large sample size (N = 971). Table 4. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Random-Order Condition, N US = 485, N CA = 486) M US SD US M CA SD CA tCohenâs d Extraversion3.870.213.610.2517.55 â 1.13 Agreeableness4.030.164.040.18-0.62-0.04 Conscientiousness3.930.173.990.16-5.81 â -0.37 Neuroticism3.020.172.940.187.42 â 0.48 Openness3.500.153.400.199.03 â 0.58 Key Findings and ImplicationsThe large-sample experiments reveal two critical patterns. First, in fixed order, increased aggregation attenuates Big Five cultural differences: the size of the extraversion effect decreases fromd = 2.93(main study) to d = 2.71 (large sample), while agreeableness, Conscientiousness, and Neuroticism become non-significant. This aligns with established concerns in fairness research: aggregation can obscure different group differences (Suresh & Guttag, 2021; Mehrabi et al., 2021) and produce models that are âoverly general or representative only of the majority groupâ (Gallegos et al., 2024). Second, geometric features demonstrate superior robustness to aggregation: SPD clustering accuracy remains high (FO: 87.60%, RO: 76.21%) despite the tenfold sample increase, and the frame-dependence pattern (collapse-recovery) persists (Appendix A.2). This dissociationâaggregated features degrade under averaging while geometric coordination remains detectableâinformed our selection ofN â 100per condition for the main study, balancing signal preservation with statistical adequacy. Implications for Sample Size Selection Large-sample results (N â 2000) reveal two key patterns: (1) Increased aggregation attenuates Big Five cultural differences (Extraversion:d = 2.93 â 2.71; three dimensions become non- significant), consistent with established concerns that averaging obscures distinct groups (Suresh & Guttag, 2021; Mehrabi et al., 2021; Gallegos et al., 2024). (2) Geometric features demonstrate superior robustness: SPD maintains 85-88% accuracy 10 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry across conditions, and the collapse-recovery pattern persists (Appendix A.2). This dissociation informed our selection of N â 100 per condition, balancing signal preservation with statistical power. A.2. Frame-Dependence Under Increased Aggregation To assess whether the frame-dependence pattern persists under increased aggregation, we conducted additional experiments with the large sample dataset (N = 500per cell; totalN = 2000)). This allows us to test: (1) whether the collapse-recovery pattern generalizes to larger samples, and (2) how sample size affects the relative robustness of geometric versus aggregated features. Table 5. Large Sample Clustering Performance Across Conditions (N = 2000, with B = 200 bootstrap iterations for RO-BTSP) Clustering Accuracy (%) SD RO-BTSP (%) FeatureFORORO-BTSP Big Five91.5673.33 â â SPD87.6076.2185.858.38 Eigenvalues75.2154.0759.758.42 Eigenvectors63.3363.2355.188.20 â Big Five scores are frame-invariant regardless of question ordering or random shuffling. ACC for Big Five in RO-BTSP is a constant (i.e. 73.33%, same as RO), and SD undefined. Full descriptive statistics in Appendix A.1. Collapse-Recovery Pattern and Differential Frame Sensitivity Table 5 presents the clustering performance in all three analytical conditions (FO, RO, RO-BTSP). The results reveal the characteristic dissociation between aggregated and geometric features: The aggregated features (Big Five) show order-dependence (91.56%â73.33%, 18.23% degradation) but frame indepen- dence. On the other hand, the geometric features (SPD) show frame dependency (87.60%â76.21%, 11.39% degradation) and substantial bootstrap variance (SD=8.38%). The geometric features also show the collapse-recovery pattern under shared-frame condition (76.21%â85.85%, for SPD features from RO to RO-BTSP), demonstrating that the geometric coordination remains detectable and recoverable even under frame perturbations. This differential bootstrap sensitivity provides strong methodological validation of the frame-dependence hypothesis (H3). Table 6. Large Sample: Decomposition of Performance Degradation (N = 2000) FeatureTotal â (%)OE%FE%Dominant Big Fiveâ18.231000Order SPDâ11.391585Frame Eigenvalues â21.147327Order Eigenvectors â0.10âNegligible Effect Decomposition at Large ScaleTable 6 decomposes the performance degradation into order effects (OE) and frame effects (FE) for the large sample. Consistent with the main study, Big Five scores show pure order effects (100% OE, 0% FE), while SPD features are predominantly frame-driven (85% FE). Notably, SPDâs total degradation is substantially smaller at large scale (â11.39% vs.â42.40% in main study), suggesting that geometric coordination structures become more resistant to frame misalignment as aggregation increasesâthough the fundamental frame-dependence mechanism persists. Implications for Sample Size SelectionThese large-sample results provide two key insights: (1) The frame-dependence mechanism generalizes across sample sizesâSPD features consistently exhibit collapse-recovery patterns. (2) The relative robustness of geometric versus aggregated features inverts at larger samples: geometric coordination becomes more preserved than simple aggregates under increased aggregation. Combined with the finding that large samples attenuate Big Five cultural differences, these results validate our selection ofN â 100per condition: this size balances discriminable cultural signals with sufficient data for geometric analysis, 11 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry while avoiding over-aggregation that would dilute both aggregate differences and obscure the dissociation between frame- dependent and frame-robust components. A.3. Visualization Figure 5 visualizes the large-sample patterns under Fixed Order (left) and Random Order (right) conditions. Consistent with the main study (Figures 3a and 3b), SPD features preserve clear group separation despite tenfold sample increase. In contrast, Big Five features show substantial overlap, reflecting the attenuation of cultural differences under increased aggregation. â20246 â2 â1 0 1 2 3 4 Big Five (5PCâ>Spectral) (Acc=50.1%, ARI=â0.000) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â4â2024 â4 â2 0 2 4 SPD Identity (10PCâ>Spectral) (Acc=85.2%, ARI=0.495) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â4â2024 â3 â2 â1 0 1 2 3 Eigenvalues (5PCâ>Spectral) (Acc=50.8%, ARI=â0.000) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â4â3â2â101 â2 â1 0 1 Eigenvectors (5PCâ>Spectral) (Acc=53.8%, ARI=0.005) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â4â202 â2 0 2 4 Big Five (5PCâ>Spectral) (Acc=50.2%, ARI=â0.000) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â4â2024 â4 â2 0 2 4 SPD Identity (10PCâ>Spectral) (Acc=90.9%, ARI=0.670) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â4â2024 â4 â2 0 2 4 Eigenvalues (5PCâ>Spectral) (Acc=78.0%, ARI=0.312) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) â20246 â2 0 2 4 Eigenvectors (5PCâ>Spectral) (Acc=62.9%, ARI=0.066) PC1 PC2 American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect) Figure 5. PCA visualizations of four features under Fixed Order (left) and Random Order (right) conditions with large sample (Nâ 2000). Colors indicate cultural group (American vs. Chinese-American). SPD features maintain clear separation, while Big Five features show substantial overlap due to attenuated cultural differences. 12 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry B. Main Study Descriptive Statistics B.1. Fixed-Order Condition Descriptive Statistics and Group Differences Table 7 presents the mean scores, standard deviations, and results of independent samplest-tests for each Big Five dimension under the Fixed-Order condition, under American (US) cultural prompt and Chinese-American (CA) cultural prompts, respectively. Table 7. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Fixed-Order Condition, N US = 96, N CA = 97) M US SD US M CA SD CA tCohenâs d Extraversion3.910.103.570.1420.37 â 2.93 Agreeableness4.020.123.970.112.79 â 0.40 Conscientiousness3.940.123.960.09-0.81-0.12 Neuroticism3.030.082.960.114.63 â 0.66 Openness3.560.063.520.142.43 â 0.35 B.2. Random-Order Condition Table 8 presents the mean scores, standard deviations, and results of independent samplest-tests for each Big Five dimension under the Random-Order condition, under American (US) cultural prompt and Chinese-American (CA) cultural prompts, respectively. Table 8. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Random-Order Condition, N US = 92, N CA = 95) M US SD US M CA SD CA tCohenâs d Extraversion3.890.213.620.257.77 â 1.13 Agreeableness4.030.184.030.190.030.00 Conscientiousness3.900.164.000.16-4.24 â -0.62 Neuroticism3.020.172.900.184.88 â 0.71 Openness3.500.143.410.213.28 â 0.48 Descriptive Statistics Across Order Conditions Tables 7 and 8 present the descriptive statistics for the Fixed-Order and Random-Order conditions, respectively. The increased standard deviations in the Random-Order condition visually corroborate the entropy increase reported in the main text (Figure 4). Notably, while mean differences exist under both conditions, they follow different patterns (e.g., the sign of the Conscientiousness difference flips), and the effect sizes (Cohenâs d) are substantially larger in the Fixed-Order condition due to its markedly reduced variability. 13 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry C. Experimental Materials All API calls used the following parameters unless otherwise specified: model: gpt-4o-2024-05-13 temperature: 0.7 max_tokens: 150 stop: None No system prompt was used; all instructions were provided in the user prompt (Appendix C.1) as shown below. C.1. Prompt Templates for Cultural Persona Induction This appendix provides the complete prompt templates used to induce American and Chinese-American cultural perspec- tives. The functioncreatepromptamerican()(and its counterpart for Chinese-American) generates the following structure, where [ITEMLIST] is replaced by the ordered list of 50 adapted items (see Appendix C.2). American Persona Prompt Template: You are an American person. Please answer the following personality questionnaire as an American would, reflecting typical American cultural values, attitudes, and perspectives. Please choose from the following options to identify how accurately each statement describes you as an American person. Respond ONLY with letters (A,B,C,D,E) for each question, one letter per line, in the exact order of the statements. Do not add any other text, numbers, or explanations. Do not stop early. If you reach the end, continue with the next line until all 50 statements are answered. Rating: A=Very Accurate, B=Moderately Accurate, C=Neutral, D=Moderately Inaccurate, E=Very Inaccurate Statements: [ITEM_LIST] Your responses as an American person (50 letters only, one per line): Chinese-American Persona Prompt Template: The template is identical in structure, with the opening instruction replaced by: You are a Chinese American person. Please answer the following personality questionnaire as a Chinese American would, reflecting the unique blend of Chinese and American cultural values, attitudes, and perspectives that characterizes the Chinese American experience. The remaining instructions, rating scale, and formatting constraints are the same. 14 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry C.2. Adapted IPIP-50 Item List This appendix lists all 50 items from the International Personality Item Pool (IPIP-50) inventory (Goldberg, 1999), sourced from the official IPIP websitehttps://ipip.ori.org/new_ipip-50-item-scale.htm. Each item was prefixed with the subject âIâ and adjusted for grammatical correctness. Items that are reverse-scored according to the standard IPIP-50 scoring key (available at the aforementioned URL) are marked with an asterisk (*) after the statement. 1. I am the life of the party. 2. I feel little concern for others. * 3. I am always prepared. 4. I get stressed out easily. * 5. I have a rich vocabulary. 6. I donât talk a lot. * 7. I am interested in people. 8. I leave my belongings around. * 9. I am relaxed most of the time. 10. I have difficulty understanding abstract ideas. * 11. I feel comfortable around people. 12. I insult people. * 13. I pay attention to details. 14. I worry about things. * 15. I have a vivid imagination. 16. I keep in the background. * 17. I sympathize with othersâ feelings. 18. I make a mess of things. * 19. I seldom feel blue. 20. I am not interested in abstract ideas. * 21. I start conversations. 22. I am not interested in other peopleâs problems. * 23. I get chores done right away. 24. I am easily disturbed. * 25. I have excellent ideas. 26. I have little to say. * 27. I have a soft heart. 28. I often forget to put things back in their proper place. * 29. I get upset easily. * 30. I do not have a good imagination. * 31. I talk to a lot of different people at parties. 32. I am not really interested in others. * 33. I like order. 34. I change my mood a lot. * 35. I am quick to understand things. 36. I donât like to draw attention to myself. * 37. I take time out for others. 38. I shirk my duties. * 39. I have frequent mood swings. * 40. I use difficult words. 41. I donât mind being the center of attention. 42. I feel othersâ emotions. 43. I follow a schedule. 44. I get irritated easily. * 45. I spend time reflecting on things. 46. I am quiet around strangers. * 47. I make people feel at ease. 48. I am exacting in my work. 49. I often feel blue. * 50. I am full of ideas. C.3. Data Collection Protocol We collected100valid API calls per cell. Responses were validated for: (1) exactly50rating characters (A-E), (2) no missing items, and (3) no explanatory text or formatting. Invalid responses were discarded and replaced. Final sample sizes are reported in Section 3.1. All conditions used identical user prompts (Appendix C.1) with no system prompt. 15