Paper deep dive
When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing
Taehyeon An, Jaehyeong Park, Donghyuk Shin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/30/2026, 2:35:29 AM
Summary
The paper introduces Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric designed to evaluate the quality of persona-conditioned large language model (LLM) simulations in survey contexts. PCI operates by modeling personas as a similarity graph and using Local Moran's I to detect local spatial coherence in response deviations. The study demonstrates that selecting a subset of personas based on PCI scores significantly improves the recovery of latent value structures (measured via Confirmatory Factor Analysis on the PVQ-RR) compared to random selection or response-stability baselines, suggesting PCI is an effective internal diagnostic for screening synthetic respondents.
Entities (8)
Relation Signals (5)
Persona-Conditioned Informativeness → uses → Local Moran's I
confidence 95% · PCI uses Local Moran’s I to quantify local spatial coherence and extract compact persona subsets
K-EXAONE-236B-A23B → generates → Simulated Survey Responses
confidence 94% · Each persona answers all 57 Portrait Values Questionnaire-Revised (PVQ-RR) items using K-EXAONE-236B-A23B
Portrait Values Questionnaire-Revised → usedinevaluationof → Persona-Conditioned Informativeness
confidence 93% · we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR)
Persona-Conditioned Informativeness → evaluates → Simulated Survey Responses
confidence 92% · PCI provides an unsupervised internal diagnostic that quantifies whether item-level response deviations exhibit local alignment
Persona-Conditioned Informativeness → improves → Confirmatory Factor Analysis
confidence 90% · a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines.
Tags
Links
- Source: https://arxiv.org/abs/2608.22438v1
- Canonical: https://arxiv.org/abs/2608.22438v1
Trouble viewing inline? Open PDF directly →
Full Text
22,024 characters extracted from source content.
Expand or collapse full text
When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion SensingConference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 07–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM ’26), November 07–11, 2026, Rome, ItalyDOI: 10.1145/3799682.3840062ISBN: 979-8-4007-2539-5/2026/11CCS: Information systems Decision support systemsCCS: Computing methodologies Simulation evaluation Taehyeon An Affiliation: Korea Advanced Institute of Science and Technology , Seoul , Republic of Korea email: taeohy@kaist.ac.kr , Jaehyeong Park Affiliation: Korea Advanced Institute of Science and Technology , Seoul , Republic of Korea email: hyeong@kaist.ac.kr and Donghyuk Shin Note: Corresponding author. Affiliation: Korea Advanced Institute of Science and Technology , Seoul , Republic of Korea email: dhs@kaist.ac.kr 2026; © c Abstract. Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran’s I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-R). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines. Keywords: Large Language Models; Silicon Sampling; Synthetic Respondents; Survey Simulation; Computational Social Science †c-license: by-nc-nd 1. Introduction Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across social, commercial, and policy domains (2; 3). The goal is to study how different synthetic groups respond to diverse survey questions. Yet plausible responses are not necessarily informative: LLMs can readily generate differentiated answers even when persona information has little systematic influence on the output. For synthetic survey analysis, the central question is therefore not whether responses vary, but whether that variation is meaningfully grounded in the personas being simulated. This challenge arises because persona profiles are typically rich and multidimensional, spanning demographics, socioeconomic status, personal values, and other attributes, whereas individual survey questions engage only particular aspects of those profiles (6). Environmental values, for example, may meaningfully shape a stance on carbon taxation, while unrelated background traits provide little consistent guidance. If an LLM effectively conditions on relevant profile cues, we should observe a coherent pattern: semantically proximate personas should shift in concordant directions rather than displaying isolated or arbitrary variation. Existing evaluation paradigms provide limited support for identifying such coherence. Most studies evaluate synthetic samples either by comparing aggregate response distributions against external human benchmarks (2) or by measuring response consistency across repeated runs (7). However, macro-level alignment can mask unconditioned individual responses (5), while persona effects may differ substantially across survey items (10). Moreover, external human data are often unavailable during exploratory simulations or survey pre-testing. An effective internal diagnostic should therefore determine whether response variation is systematically organized with respect to persona profiles, allowing weakly grounded simulations to be identified before downstream analysis. Four-panel diagram of the PCI workflow: a persona similarity graph, item response deviations on the graph, local coherence scores, and the selected informative persona subset. Figure 1. Overview of the PCI workflow: constructing a persona similarity graph from profile embeddings, computing centered item response deviations, scoring local spatial coherence with Local Moran’s I, and selecting an informative persona subset.Four-panel diagram of the PCI workflow: a persona similarity graph, item response deviations on the graph, local coherence scores, and the selected informative persona subset. Table 1. CFA comparison and graph sensitivity for 10% persona subsets. Category Selection / graph χ2↓χ^2 df CFI ↑ SRMR ↓ RMSEA (90% CI) ↓ Baseline Random mean (500 subsets) 2623.78 1367 .880 .106 .079 (.074, .084) Response stability 2711.10 1367 .877 .111 .082 (.077, .086) Proposed PCI, k=740k=740 (50%-N) 2485.15 1367 .925 .086 .075 (.070, .079) Sensitivity PCI, k=444k=444 (30%-N) 2481.24 1367 .924 .091 .074 (.070, .079) PCI, k=592k=592 (40%-N) 2489.01 1367 .925 .091 .075 (.070, .079) PCI, k=888k=888 (60%-N) 2633.03 1367 .915 .090 .079 (.075, .084) PCI, k=1036k=1036 (70%-N) 2671.06 1367 .910 .093 .081 (.076, .085) Note. All rows evaluate 10% subsets (n=148n=148). Random entries are means across 500 subsets. Boldface marks the best value for each metric (including ties). To address this need, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised internal diagnostic based on local spatial coherence in response deviations. Motivated by the manifold smoothness assumption (4), PCI models personas as a similarity graph and survey responses as graph signals. For each item, it measures whether semantically similar personas deviate from the sample-wide response baseline in concordant directions using Local Moran’s I (1), and aggregates this evidence to rank personas by informativeness. We evaluate PCI on 1,480 synthetic personas answering the 57-item Portrait Values Questionnaire-Revised (PVQ-R) (9; 8). Without using construct labels during selection, a PCI-selected 10% subset substantially improves latent construct recovery over response-stability and random baselines (ΔCFI=+.045 =+.045, p=.002p=.002), supporting local spatial coherence as a useful internal signal for screening synthetic respondents. 2. Persona-Conditioned Informativeness 2.1. From Item Relevance to Local Coherence Figure 1 illustrates the PCI workflow. PCI operationalizes persona-conditioned informativeness through two observable properties: (1) there must be between-persona variation around the item-level baseline, separating persona-specific variation from invariant consensus; and (2) semantically proximate personas in profile space must deviate in concordant directions, distinguishing shared conditioning from isolated sampling variance. PCI rewards a response only when both conditions are satisfied, downweighting invariant consensus and isolated deviations that lack neighborhood support. Formally, let P=1,…,nP=\1,…,n\ denote the set of personas and I=1,…,mI=\1,…,m\ the set of survey items. For persona p, item i, and generation trial r∈1,…,Rr∈\1,…,R\, let ypiry_pir be the generated score. We compute the Monte Carlo mean y¯pi=1R∑r=1Rypir y_pi= 1R _r=1^Ry_pir across R trials to reduce stochastic token variance. While repeated-response variance is not an input to PCI, we use it to construct our response-stability baseline. 2.2. Persona Graph and Response Signals Each persona possesses a narrative profile covering demographics, socioeconomic status, family structure, occupation, cultural identity, and preferences. Rather than discretizing these attributes into rigid categorical features, we embed the complete profile text with a dense representation model to capture complex, cross-attribute cues. We construct a weighted similarity graph G=(P,E,W)G=(P,E,W), where nodes represent personas and edges encode k-nearest-neighbor relations under cosine similarity. Raw edge weights are cosine similarities clipped at zero. We define the graph topology via mutual k-nearest-neighbor linkage and row-standardize the resulting weight matrix so that neighbor weights sum to one: (1) ∑q∈Pwpq=1,wpq=0 if (p,q)∉E. _q∈ Pw_pq=1, w_pq=0 if (p,q)∉ E. We use k=740k=740 (50%-N for n=1,480n=1,480) as the reference specification and evaluate sensitivity across k∈30%,…,70%k∈\30\%,…,70\%\ in Section 3. Graph density k governs the extent of neighborhood support, while the selection budget controls sample selectivity. For each item i, we mean-center the aggregated responses across the simulated population: (2) μi=1n∑p∈Py¯pi,zpi=y¯pi−μi. _i= 1n _p∈ P y_pi, z_pi= y_pi- _i. The item mean μi _i removes the shared sample-wide response level, making the centered deviation zpiz_pi capture between-persona variation around this baseline. 2.3. Local Association and Persona Ranking To quantify whether zpiz_pi is corroborated by semantic peers, we apply Anselin’s Local Moran’s I (1). Let si2=n−1∑p∈Pzpi2s_i^2=n^-1 _p∈ Pz_pi^2 denote the second central moment of item i, where the zero-mean property follows from centering. When si2>0s_i^2>0, the local spatial association is defined as: (3) Ipi=zpi∑q∈Pwpqzqisi2.I_pi= z_pi _q∈ Pw_pqz_qis_i^2. The weighted sum z~pi=∑q∈Pwpqzqi z_pi= _q∈ Pw_pqz_qi denotes the local spatial lag. A positive IpiI_pi signifies local clustering of concordant deviations: either high–high (zpi>0,z~pi>0z_pi>0, z_pi>0) or low–low (zpi<0,z~pi<0z_pi<0, z_pi<0), whereas negative IpiI_pi indicates spatial discordance (high–low or low–high). If si2=0s_i^2=0, all personas produce an identical response, yielding Ipi=0I_pi=0. Because our target construct is neighbor-aligned conditioning, PCI retains only the non-negative component PCIpi=max(Ipi,0)PCI_pi= (I_pi,0). Truncating negative values ensures that discordant or isolated deviations do not contribute to informativeness, conservatively requiring neighborhood support. Aggregating across items yields the persona score PCIp=m−1∑i∈IPCIpiPCI_p=m^-1 _i∈ IPCI_pi. Personas with high PCIpPCI_p exhibit robust, neighbor-aligned response patterns across items, and we select the top 10% (n=148n=148) as the primary PCI subset. 3. Construct-Recovery Evaluation 3.1. Data and Construct-Recovery Target We draw 1,480 profiles from NVIDIA’s Singapore persona dataset.11 1 https://huggingface.co/datasets/nvidia/Nemotron-Personas-Singapore Each persona answers all 57 Portrait Values Questionnaire-Revised (PVQ-R) items using K-EXAONE-236B-A23B.22 2 https://huggingface.co/LGAI-EXAONE/K-EXAONE-236B-A23B We generate R=10R=10 responses per persona--item pair on a 1--6 scale and embed profiles with llama-nemotron-embed-1b-v2.33 3 https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2 The evaluation tests whether selected responses better preserve the established PVQ-R construct structure without using construct information during selection. The PVQ-R measures 19 distinct human values (9; 8), organized into four motivational quadrants: self-enhancement (power, achievement), self-transcendence (benevolence, universalism), openness to change (self-direction, stimulation), and conservation (security, conformity, tradition). In Schwartz’s circular continuum, adjacent values represent compatible motivational goals, while opposing values represent conflicting psychological priorities. Successfully recovering this theoretical continuum through confirmatory factor analysis provides evidence that PCI enriches for structured, persona-aligned response variation rather than arbitrary subset selection. We fit a confirmatory factor analysis (CFA) model with 19 correlated value factors and an orthogonal common response factor to control for scale acquiescence. We report the Index of Quality IoQv=|Corr(Cv,ηv)|IoQ_v= |Corr(C_v, _v) |, where CvC_v denotes the observed composite score for value v and ηv _v its corresponding latent factor (IoQv2IoQ_v^2 represents the reliable composite variance). These metrics evaluate relative construct recovery across selection strategies under an identical structural model rather than absolute model fit. Table 2. Index of Quality across the 19 PVQ-R constructs. Value Random Response stability PCI Self-direction-thought .923 .924 .962 Self-direction-action .901 .900 .936 Stimulation .974 .977 .992 Hedonism .959 .965 .977 Achievement .958 .939 .983 Power-dominance .622 .497 .758 Power-resources .935 .913 .971 Face .793 .711 .831 Security-personal .699 .713 .811 Security-societal .875 .897 .897 Tradition .949 .938 .979 Conformity-rules .945 .952 .974 Conformity-interpersonal .948 .936 .979 Humility .902 .879 .962 Benevolence-care .709 .746 .842 Benevolence-dependability .536 .649 .480 Universalism-concern .840 .893 .791 Universalism-nature .916 .928 .911 Universalism-tolerance .904 .916 .939 Mean .857 .856 .893 Median .904 .913 .939 Note. All values use 10% subsets. Random entries are means across 500 subsets. Boldface marks the best value in each row (including ties). Figure 2. CFI across subset sizes with the 50%-N graph fixed. The 10% random subset distribution contains 500 subsets.Line chart of CFI by subset fraction for PCI, response stability, and random selection, where PCI achieves the highest CFI from 10\% through 80\% subsets. 3.2. Selection Rules and Statistical Tests We compare three 10% subsets (n=148n=148): PCI selects the highest PCIpPCI_p scores; response stability selects the lowest generation variance Vp=m−1∑iVarr(ypir)V_p=m^-1 _iVar_r(y_pir); and the random baseline uses 500 size-matched random subsets as a chance-level reference for the selection strategies. We evaluate graph density across k∈30%,…,70%k∈\30\%,…,70\%\-N around the reference 50%-N specification, alongside subset fractions swept from 10% to 100% to isolate budget and density effects. We compute empirical one-sided p-values against the 500 size-matched random subsets and apply Bonferroni correction across the five graph-density specifications. 3.3. Results As shown in Table 1, the primary 10% PCI subset achieves a higher CFI (.925) than the response-stability baseline (.877) and significantly exceeds the 500-subset random baseline mean of .880 (95% empirical random-subset interval: .865–.896; p=.002p=.002, Bonferroni-adjusted p=.010p=.010). PCI similarly improves residual fit, reducing SRMR from .106 to .086 (p=.002p=.002, adjusted p=.010p=.010) alongside consistent reductions in χ2χ^2 and RMSEA. As reported in Table 2, construct-level evaluations reveal consistent advantages alongside construct-specific nuances. PCI attains the best or tied-best IoQ for 16 of 19 constructs, achieving an average of .893.893 compared to .857.857 for random selection and .856.856 for response stability. Gains are particularly pronounced for power-dominance (from .622 to .758) and security-personal (from .699 to .811). Conversely, response stability performs better on several constructs, including benevolence-dependability and two universalism constructs, indicating that generation consistency and persona-conditioned local coherence capture distinct properties of synthetic responses. Across subset fractions, PCI consistently outperforms both baselines, maintaining superior CFI from 10% through 80% fractions and peaking at .927 at a 20% fraction (n=296n=296, Figure 2). In addition, sensitivity analysis demonstrates that these improvements remain robust across graph densities: CFI forms a stable plateau of .924.924–.925.925 between 30% and 50%-N, remains elevated at .910.910–.915.915 up to 70%-N, and preserves highly consistent persona rankings throughout (Spearman ρ∈[.961,.990]ρ∈[.961,.990]). 4. Discussion and Conclusion In computational social science, survey pre-testing, and pluralistic opinion sensing, plausible and repeatable outputs can mask unconditioned model consensus. PCI provides an unsupervised internal diagnostic that quantifies whether item-level response deviations exhibit local alignment among semantically proximate personas, without requiring construct labels. The observed construct-recovery gains support local spatial coherence as a useful internal signal of structured response variation in synthetic respondents. For research and practical applications, PCI suggests a three-stage screening workflow: (1) unsupervised auditing of raw synthetic generations to flag items exhibiting weak persona-aligned structure, (2) subset selection to extract compact cohorts that retain structurally informative variation for cost-effective exploratory surveys, and (3) targeted validation triage directing costly human validation toward items or subgroups exhibiting low informativeness or discordance. Rather than sampling personas blindly, this workflow concentrates auditing effort where needed. Because PCI is closed-form given fixed generations and graph construction, it remains computationally lightweight without requiring model training. Several limitations qualify our findings and highlight promising directions for future research. First, internal construct recovery provides evidence of coherent profile–response structure, but serves as an internal structural diagnostic rather than a guarantee of external population validity; combining PCI with external calibration remains essential for representative polling. Second, our primary graph construction embeds full narrative profiles into a single global metric space. Because survey items may engage different subsets of attributes, constructing subspace-specific or multi-relational persona graphs represents a promising direction to prevent irrelevant profile cues from diluting local neighborhood structures. Finally, while demonstrated on the 57-item Portrait Values Questionnaire-Revised with K-EXAONE-236B, evaluating cross-model generalizability across diverse model families and multilingual instruments will further delineate the boundaries of synthetic opinion sensing. Acknowledgements. This work was supported by the Sponsor Institute of Information & Communications Technology Planning & Evaluation (IITP) https://w.iitp.kr-Global Data-X Leader HRD program grant funded by the Sponsor Korea government (MSIT) https://w.msit.go.kr (Grant #IITP-RS-2024-00440626). GenAI Disclosure Generative AI supported translation, editing, and LaTeX formatting. References Anselin (1995) L. Anselin Local indicators of spatial association—LISA. Geographical Analysis 27 (2), p. 93–115. External Links: Document Cited by: §1, §2.3. Argyle et al. (2023) L. P. Argyle, E. C. Busby, N. Fulda, J. R. Gubler, C. Rytting, and D. Wingate Out of one, many: Using language models to simulate human samples. Political Analysis 31 (3), p. 337–351. External Links: Document Cited by: §1, §1. Arora et al. (2025) N. Arora, I. Chakraborty, and Y. Nishimura AI–human hybrids for marketing research: Leveraging large language models (llms) as collaborators. Journal of Marketing 89 (2), p. 43–70. External Links: Document Cited by: §1. Belkin et al. (2006) M. Belkin, P. Niyogi, and V. Sindhwani Manifold regularization: a geometric framework for learning from labeled and unlabeled examples. Journal of Machine Learning Research 7 (85), p. 2399–2434. Cited by: §1. Bisbee et al. (2024) J. Bisbee, J. D. Clinton, C. Dorff, B. Kenkel, and J. M. Larson Synthetic replacements for human survey data? The perils of large language models. Political Analysis 32 (4), p. 401–416. External Links: Document Cited by: §1. Hu and Collier (2024) T. Hu and N. Collier Quantifying the persona effect in LLM simulations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 10289–10307. External Links: Document Cited by: §1. Reusens et al. (2025) M. Reusens, B. Baesens, and D. Jurgens Are economists always more introverted? Analyzing consistency in persona-assigned LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025, p. 11268–11287. External Links: Document Cited by: §1. Schwartz et al. (2012) S. H. Schwartz, J. Cieciuch, M. Vecchione, E. Davidov, R. Fischer, C. Beierlein, A. Ramos, M. Verkasalo, J. Lönnqvist, K. Demirutku, O. Dirilen-Gumus, and M. Konty Refining the theory of basic individual values.. Journal of Personality and Social Psychology 103 (4), p. 663–688. External Links: ISSN 1939-1315(Electronic),0022-3514(Print), Document Cited by: §1, §3.1. Schwartz and Cieciuch (2022) S. H. Schwartz and J. Cieciuch Measuring the refined theory of individual values in 49 cultural groups: Psychometrics of the revised portrait value questionnaire. Assessment 29 (5), p. 1005–1019. External Links: Document Cited by: §1, §3.1. Taday Morocho et al. (2026) E. E. Taday Morocho, L. Cima, T. Fagni, M. Avvenuti, and S. Cresci Assessing the reliability of persona-conditioned llms as synthetic survey respondents. In Companion Proceedings of the ACM Web Conference 2026, W Companion ’26, New York, NY, USA, p. 320–329. External Links: Document, ISBN 979-8-4007-2308-7 Cited by: §1.