Paper deep dive
Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach
Xiaoyou Qin, Zhihong Li, Xiaoxiao Cheng
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/10/2026, 3:41:10 AM
Summary
This paper addresses the 'average persona' effect in LLM-based social simulation, where models collapse social diversity into homogeneous outputs. The authors propose 'audience segmentation' as a systematic methodological framework to restore heterogeneity. By testing six segmentation configurations across Llama 3.1-70B and Mixtral 8x22B using U.S. climate-opinion data, the study evaluates performance across distributional, structural, and predictive fidelity. Findings indicate that while moderate segmentation improves fidelity, excessive granularity can worsen performance, and the choice of selection logic (theory-driven, data-driven, or instrument-based) determines which fidelity dimension is prioritized.
Entities (5)
Relation Signals (3)
Audience Segmentation â restores â Heterogeneity
confidence 95% · This study introduces audience segmentation as a systematic approach for restoring heterogeneity in LLM-based social simulation.
Llama-3.1-70B â exhibits â Heterogeneity Masking
confidence 90% · In addition, all LLMs showed residual overregularization, indicating persistent difficulty in fully recovering real-world population heterogeneity.
Audience Segmentation â improves â Simulation Fidelity
confidence 90% · Moderate enrichment can improve performance, but further expansion does not reliably help and can worsen structural and predictive fidelity.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable "silicon samples" that can approximate human data. However, current simulation practice often collapses diversity into an "average persona," masking subgroup variation that is central to social reality. This study introduces audience segmentation as a systematic approach for restoring heterogeneity in LLM-based social simulation. Using U.S. climate-opinion survey data, we compare six segmentation configurations across two open-weight LLMs (Llama 3.1-70B and Mixtral 8x22B), varying segmentation identifier granularity, parsimony, and selection logic (theory-driven, data-driven, and instrument-based). We evaluate simulation performance with a three-dimensional evaluation framework covering distributional, structural, and predictive fidelity. Results show that increasing identifier granularity does not produce consistent improvement: moderate enrichment can improve performance, but further expansion does not reliably help and can worsen structural and predictive fidelity. Across parsimony comparisons, compact configurations often match or outperform more comprehensive alternatives, especially in structural and predictive fidelity, while distributional fidelity remains metric dependent. Identifier selection logic determines which fidelity dimension benefits most: instrument-based selection best preserves distributional shape, whereas data-driven selection best recovers between-group structure and identifier-outcome associations. Overall, no single configuration dominates all dimensions, and performance gains in one dimension can coincide with losses in another. These findings position audience segmentation as a core methodological approach for valid LLM-based social simulation and highlight the need for heterogeneity-aware evaluation and variance-preserving modeling strategies.
Tags
Links
- Source: https://arxiv.org/abs/2604.06663v1
- Canonical: https://arxiv.org/abs/2604.06663v1
Trouble viewing inline? Open PDF directly â
Full Text
84,943 characters extracted from source content.
Expand or collapse full text
Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach Xiaoyou Qin School of Journalism, Fudan University Zhihong Li College of Media and International Culture, Zhejiang University Xiaoxiao Cheng College of Media and International Culture, Zhejiang University Corresponding author: xxcheng21@zju.edu.cn Abstract Large language models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable âsilicon samplesâ that can approximate human data. However, current simulation practice often collapses diversity into an âaverage persona,â masking subgroup variation that is central to social reality. This study introduces audience segmentation as a systematic approach to restoring heterogeneity in LLM-based social simulation. Using US climate-opinion survey data, we compare six segmentation configurations across two open-weight LLMs (Llama 3.1-70B and Mixtral 8x22B), varying segmentation identifier granularity, parsimony, and selection logic (theory-driven, data-driven, and instrument-based). We evaluated simulation performance by using a three-dimensional evaluation framework covering distributional, structural, and predictive fidelity. The results show that increasing identifier granularity does not produce consistent improvement: Moderate enrichment can improve performance, but further expansion does not reliably help and can worsen structural and predictive fidelity. Across parsimony comparisons, compact configurations often match or outperform more comprehensive alternatives, especially in structural and predictive fidelity, while distributional fidelity remains metric dependent. Identifier selection logic determines which fidelity dimension benefits most: Instrument-based selection best preserves distributional shape, whereas data-driven selection best recovers between-group structure and identifierâoutcome associations. Overall, no single configuration dominates all dimensions, and performance gains in one dimension can coincide with losses in another. In addition, all LLMs showed residual overregularization, indicating persistent difficulty in fully recovering real-world population heterogeneity. These findings position audience segmentation as a core methodological approach to valid LLM-based social simulation and highlight the need for heterogeneity-aware evaluation and variance-preserving modeling strategies. Keywords: LLM-based social simulation; silicon sample; segmentation; fidelity; heterogeneity 1 Introduction Large language models (LLMs) have rapidly transformed social science research, creating new possibilities and capabilities for simulating human attitudes [undefaj, undefaz] and modeling complex social dynamics [undefaad, undefaq]. Their applications now extend far beyond conventional tasks such as literature review, text classification, and survey design [undefax]. Of particular note is the emergent field of LLM-based social simulationâa promising yet methodologically demanding domain. A pivotal development in this field is [undefaj]âs âsilicon samplingâ paradigm, which leverages LLMs to generate synthetic data capable of substituting for traditional human respondents. This brings numerous advantages, including dramatic reductions in research cost and time, as well as the ability to design and implement experiments that were previously infeasible due to practical or ethical constraints [undefax, undefaae]. In this line of research, LLM-based simulation largely falls into two main categories. The first integrates LLMs as agents within multi-agent systems to emulate complex decision-making processes, emergent collective behavior, and the corresponding social interaction patterns [undefaad, undefaak, undefaan]. The second employs LLMs as tools for measuring public opinion and directly replicating, estimating, and predicting public attitudes on issues such as climate change, political preferences, and policy support [undefaz, undefaaq]. Nevertheless, while these advances are exciting and promising, they also bring into sharp relief a fundamental challenge: preserving heterogeneity. The performance of LLM-based simulation hinges on several widely recognized factors, such as training data quality, model transparency, prompt engineering, and benchmark selection [undefaj, undefar, undefap]. However, a more insidious problem remains. LLM architectures, by design, tend to compress social diversity into their most probableâor modalârepresentations [undefaam]. This âaverage personaâ effect [undefaan, undefaae], wherein model responses converge toward central tendencies, largely obscures and even erases the very differences at the subgroup level that are essential for rich, meaningful, and credible social simulation. This phenomenon is not just a technical curiosity; it directly undermines attempts to accurately capture the distributional reality of social-grouping attitudes. Genuine social simulation, as [undefaac] warn, requires models that can both distinguish and generate responses from multiple subgroups, rather than models that reduce variation and heterogeneity to undifferentiated averages and homogeneity. The erasure of heterogeneity, which refer to as âheterogeneity masking,â is multifaceted. While the technical compression partially stems from LLMsâ inherent training mechanisms that favor averaged representations [undefaam], the absence of heterogeneity thinking in simulation practices exacerbates the issue. Methodologically, prompt engineering efforts tend to focus on refining instruction clarity and persona consistency yet rarely addresses how to systematically encode population heterogeneity in simulations [undefan]. This is largely due to a lack of robust frameworks for translating group variance into research design. Epistemologically, there is often an implicit bias toward generalizability, in which ârepresentativeâ means âaverageâ, rather than accurate reproductions of real-world diversity and distributional richness. Despite growing interest in demographic profiling and persona conditioning, systematic guidance on how segmentation choices affect simulation performance remains limited. Prior research has typically examined whether demographic characteristics should be included 2023OutOfOneMany,Santurkar2023WhoseOpinions, but has paid less attention to which variables (hereafter termed âsegmentation identifiersâ) matter, the level of granularity required, and the selection logic used. As a result, existing social simulation studies risk privileging modal responses and obscuring meaningful differencesâoften unintentionally. This is where our study intervenes. We advocate for audience segmentation, an approach that has been well established within communication science and marketing research [undefaag, undefaah] as a systematic way to restore heterogeneity in LLM-based simulation. Audience segmentation partitions populations into subgroups that are theoretically and empirically meaningful [undefau]. This study employs climate opinion survey data collected from the United States to rigorously examine how different segmentation strategies affect LLMsâ capacity to replicate human attitude distributions while preserving subgroup differences. To rigorously evaluate this proposition, we assessed six segmentation configurations that varied along the dimensions of identifier granularity, parsimony, and selection logic, and evaluated simulation performance using distributional, structural, and predictive fidelity. Our empirical analysis challenges the intuitive assumption that increasing the detail of persona prompts inherently yields more realistic social variation. Instead of improving fidelity, the indiscriminate accumulation of segmentation identifiers paradoxically amplifies the modelâs tendency toward over-regularization, which distorts the authentic relational structures among subgroups. We demonstrate that the successful restoration of heterogeneity depends fundamentally on informative parsimony. Furthermore, the underlying logic for identifier selectionâwhether theory-driven, data-driven, or based on empirical instrumentsâdirectly determines which specific dimension of simulation performance is most effectively preserved. Ultimately, these findings reposition audience segmentation from a surface-level prompt-engineering tactic into a core methodological decision for valid LLM-based research. 2 From Typological Thinking to Population Thinking The masking of heterogeneity epitomizes a deeper epistemological divide in the social sciences concerning how scholars conceptualize social variations. This divide is rooted in two philosophical traditions that have long shaped scientific inquiry: typological thinking and population thinking. Typological thinking, rooted in the Platonic tradition, treats variation as unwanted noise around an essential âtype.â Deviations from the type are regarded as confounding factors or exogenous disturbances that need to be eliminated in the pursuit of universal laws [undefaao]. A key assumption underlying this tradition, which has worked well in natural science, is homogeneity: Once we understand a type of phenomenon, we can generalize that knowledge to individual cases. Within the social sciences, typological thinking often manifests in attempts to identify modal personalities, average voters, or singular ârepresentativeâ agents. Individual differences are interpreted as error variance that obscures the underlying âtrueâ pattern. At its very core, this âsocial physicalâ way of thinking naively essentializes population averages, making the âaverage manâ the primary object of exploration. While appealing for its simplicity and promise of generalizability, this approach is limited in terms of modeling social reality, which is characterized by structured variation across social groups. In contrast, population thinking emerged from the Darwinian revolution and was later introduced to the social sciences by Francis Galton. In this mode of thinking, deviations from the mean are not scientifically trivial; they constitute the very basis of evolution and are, therefore, an intrinsic property of distributions rather than departures from an idealized type [undefaao, undefaap]. Instead of privileging âtypicalâ cases, population thinking treats heterogeneityâincluding subgroups, outliers, and the full spectrum of individual responsesâas meaningful signals that reveal the underlying structure of social phenomena. Accordingly, heterogeneity, not homogeneity, becomes central to explanations of social change. This distinction is far from philosophical; it carries direct methodological implications for social simulation. Population heterogeneity, that is, the authentic diversity of attitudes, identities, and behaviors, constitutes a precondition for valid social simulation. Research on complex systems has established that agent diversity drives behavioral evolution and enables emergent phenomena that homogeneous populations cannot generate [undefaan]. Recent work has further underlined the relevance of population thinking for LLM-generated data. [undefaj], for instance, demonstrated that demographically conditioned prompts can approximate survey responses for distinct subgroups. However, such attempts often assume, without sufficiently testing, whether conditioning truly preserves meaningful heterogeneity or merely produces tailored âtypesâ or âaverage menâ that still fail to capture the breadth of human diversity [undefaac]. This points to the central problem in LLM-based social simulation: Despite the rhetorical embrace of population-level accuracy, current practices often inadvertently revert to typological logic. Unconditioned or ill-specified prompts tend to yield statistically averaged, âtypicalâ responsesâtypes that erase the distributional richness central to social scientific inquiry. What is at stake, then, is not merely a technical limitation but an epistemological misalignment; to be specific, methods implicitly built on typological premises are at odds with the population-thinking ethos that underlies many sociological theories. Without deliberate intervention, simulations risk perpetuating and (re)producing the very âaverage personaâ effect that undermines meaningful group-level and societal inference. Addressing this misalignment requires more than technical tuning; it calls for population thinking that is operationalized and baking it into simulation designâspecifically, through systematic audience segmentation strategies. Treating heterogeneity as a signal to be detected, rather than as noise to be minimized, not only aligns simulations with complex real-world realities but also enables more nuanced inquiry into mechanisms, subgroup dynamics, and social structure. 3 Sources of Heterogeneity Masking As outlined above, heterogeneity masking is a crucial yet understudied limitation in current LLM-based social simulation. Without deliberate methodological attention, simulated populations risk being reduced to overly simplistic âaverage personas.â This flattening of diversity does not occur by chance; it is shaped by technical, methodological, and epistemological choices made throughout the simulation process. Recognizing these mechanisms is a necessary first step toward a population-thinking approach. 3.1 Technical compression Heterogeneity is often erased by mechanism embedded in LLM training and optimization. Mainstream LLMs are typically trained with methods such as maximum likelihood estimation (MLE), which incentivizes models to produce the most probableâand thus often the most centralâresponse to a given prompt; this means those less-typical, minority, or outlier responses are naturally suppressed, thereby making LLMs struggle with preserving Significant population heterogeneity [undefay]. Specifically, likelihood-based loss functions encourage high-probability generations, which compress diversity across subgroups [undefaam]. Reinforcement learning from human feedback can exacerbate this tendency: by aligning model outputs with what human annotators judge as âhigh qualityâ or the âright answer,â LLMs are nudged toward consensual, moderate, and âsafeâ responses, obscuring authentic social variation across social groups. [undefaaf] describe this process as âvalue alignment,â which prioritizes certain value systems and systematically reduces the expression of diverse viewpoints. The empirical consequences of these mechanisms have been documented in recent work. [undefal] tested LLMsâ capacity to replicate survey responses across 687 sociodemographically defined subpopulations; even with explicit sociodemographic prompts, the resulting answer distributions remained highly concentrated across subgroups. Importantly, their analysis demonstrated that this prediction error cannot be attributed to a single, stable pattern of social bias; instead, errors shifted across groups and survey questions, implying that heterogeneity compression is not merely inherited from societal inequalities in training corpora; rather, it arises from representational compression in model architectures. Taken together, these technical mechanisms create a foundational layer of heterogeneity masking that operates regardless of the researchersâ intentions. 3.2 Methodological neglect Methodological choices can compound technical compression when heterogeneity is not operationalized systematically. While proper conditioning is widely recognized as essential for simulation fidelity [undefaj], prompt engineering is often treated as technical tuning rather than as a process with profound implications for data validity [undefaae]. This produces âad hoc segmentationâ simulation practices that include identity cues without a clear theoretical rationale. Some studies exhibit unreflective methodological inertia toward demographic conditioning, routinely including age, gender, education, and race without considering the domain dependence of segmentation identifiers or specific research contexts. However, key variables affecting simulation performance vary by research topic Indiscriminately piling demographic information does not necessarily improve fidelity and may introduce irrelevant noise [undefai]. Recent empirical work offers a cautionary illustration: Despite conditioning prompts on five standard sociodemographic variables, LLMs failed to generate meaningfully differentiated responses across subgroups [undefal]. This suggests that surface-level demographic characteristics function as âidentity labelsâ rather than âattitudinal anchorsâ; they inform the model of who someone is, but not how that person is likely to think or respond within a specific domain. More fundamentally, social identities are not merely individual attributes but expressions of group relations embedded in social structures [undefaal]. Treating prompt variables as independent individual labels ignores in-group/out-group biases [undefaw] and strips group identity of its meaning within specific social contexts. 3.3 Epistemological misalignment Perhaps the most insidious layer of heterogeneity masking is the epistemological layer. This manifests in two ways. First, researchers often pursue model generalizability and representativeness in ways that implicitly privilege averages. Models are expected to produce âtypicalâ opinions rather than capture distributional variation. This carries over typological thinking-based sampling logic that is misaligned with simulation goals that require the preservation of social variations. Second, prevailing evaluation practices reinforce this bias. Many studies validate simulation fidelity using aggregated summary statistics while sidelining granular subgroup-level comparisons. Success is declared when model outputs âmatchâ real-world distributions on overall means or medians, regardless of how well subgroup variation is preserved. This evaluative blind spot has tangible costs. [undefaam] compared LLM outputs against 3,200 human participants across 16 demographic identities and found that LLMs consistently âflattenâ groups. Their responses occupy a far narrower semantic space than those of actual humans. Across four models and multiple diversity metrics, this pattern held for nearly every demographic group tested. More troubling still, LLM personas often resemble out-group imitations rather than authentic in-group voices, particularly for marginalized communities such as non-binary individuals and people with disabilities. The risk, then, is a drift toward superficial prediction rather than explanatory depth [undefav]. Social mechanisms, group dynamics, and causal processes are drowned in the ânoiseâ of averages [undefat]. 4 Audience Segmentation as a Pathway to Restore Heterogeneity The preceding section reveals that heterogeneity masking operates through interlocking technical, methodological, and epistemological mechanisms. Addressing the âaverage personaâ problem therefore requires more than technical refinement; it demands an integrated and systematic framework that embeds population thinking into simulation design from the outset. Here, we propose audience segmentation as such a framework. A clarification is necessary at the outset: The heterogeneity we seek to restore is not an exact replication of real-world population distributions, which would be neither feasible nor, for many research purposes, necessary. Rather, our goal is to recover theoretically meaningful segmentsâsubgroups that are conceptually grounded and empirically distinguishable according to established social scientific frameworks. Preserving such theoretical segmentation is essential for understanding group-level mechanisms and dynamics, even if perfect demographic fidelity remains unattainable. Segmentation, in this sense, is not merely a technical procedure but the explicit operationalization of population thinking. It treats heterogeneity and variation as signals rather than noise, thereby realigning simulation practices with the realities of differentiated societies. 4.1 Segmentation analysis Segmentation analysis originated in audience-oriented communication research and marketing science, where it serves to identify structural differences across psychological, behavioral, and value dimensions, thereby partitioning audiences into subgroups that are internally cohesive yet externally differentiated [undefaag, undefaah]. Rather than arbitrary labeling, this method is a systematic effort to remain faithful to the diversity that lready exists in social structures. It recognizes that individual behavior is jointly shaped by social environments, perceived norms, and cognitive processes [undefau]. The epistemological foundation of segmentation analysis aligns directly with population thinking. Rather than seeking a single âaverageâ profile, segmentation emphasizes carefully selecting segmentation identifiers to incorporate socialization pathways, including institutional positions, cultural resources, and social networks, thereby revealing the deep logic of social differentiation. This approach is particularly well-suited to LLM-based social simulation, that is, for LLMs to authentically simulate differential responses across groups in various social contexts, they must understand and reflect how social identity and positionality influences human viewpoints, reactions, and behaviors [undefaam]. Segmentation is well suited to uncover heterogeneity that is masked by averaging. A compelling illustration comes from [undefas]âs study of South African public attitudes toward science and technology. Through a hierarchical cluster analysis on 3,183 respondents, they identified six subgroups characterized by distinct configurations of socioeconomic status, cultural capital, information sources, and science attitudes. These patterns were not visible in the overall averages. Only through segmentation analysis could the unique cultural distances and differentiated responses to science communication be revealed. This example highlights segmentationâs ability to restore heterogeneity at an operational levelâprecisely what is needed to counteract the masking effects. 4.2 Segmentation identifiers in prior LLM-based simulation research Although segmentation has not been systematically formalized as a theoretical concept in LLM-based social simulation, multiple studies have exhibited practical awareness of segmentation-style thinking. A review of existing work suggests that identifier selection roughly follows three logics borrowed from traditional segmentation research, though often without explicit theoretical justification. First, demographic identifiers (e.g., age, gender, education, etc.) are the most commonly used segmentation identifiers, largely because they are easy to obtain and widely available in survey data [undefaah, undefao]. These variables dominate political science and public opinion simulation studies [undefak, undefaai, undefaaj]. For instance, [undefaj] combined GPT-3 with demographic background variables from the American National Election Studies through silicon sampling and found that model-generated âsilicon samplesâ closely matched survey data on vote choice, party identification, and ideological distribution. This demonstrates that demographic characteristics can enhance simulation performance. Second, psychographic identifiers (e.g., attitudes, values, and issue-specific cognitions) have been incorporated to address the limitation that demographics alone are often insufficient to distinguish attitudinal differences within seemingly homogeneous groups. For example, adding environmental policy stances and scientific-consensus cognitions to prompts improved predictive accuracy in climate change research [undefaz, undefau]. Similarly, [undefaab] introduced the Hofstede individualism index as a cultural segmentation variable and showed that LLMs can capture not only demographic and psychological dimensions but also deep self-concept differences in cross-national settings. Third, behavioral identifiers can be used to segment populations based on observed behavioral patterns. In human-computer interaction, [undefaaa] used user click, input, and submission paths to construct behavioral segments, enabling LLMs to reproduce interactive scenarios in ways that more closely reflect differences amongst users. 4.3 From identifiers to research design: Key methodological challenges and considerations While these categoriesâfrom broad demographics to specific psychographic and behavioral tracesâprovide raw materials for persona conditioning, simply identifying them does not resolve the deeper implementation problem. Moving from these markers to an operational research design raises significant yet often overlooked methodological challenges. The first challenge is identifier granularity. Existing LLM-based simulations tend to operationalize identifier inclusion as an all-or-nothing decision: Either a variable is incorporated into the persona prompt or it is omitted entirely [undefaj, undefaaj, undefaz]. What these practices leave unaddressed is how much descriptive detail is needed to characterize a synthetic persona in ways that reflect true social positionality. This oversight stems from a failure to recognize that the restoration of heterogeneity is not merely a matter of whether to condition a model, but of how deeply to describe the social self. If we accept that social identity is rooted in layered social structures, simulation practices should move beyond âthinâ descriptions based on isolated (demographic) labels toward âthickâ descriptions that incorporate multidimensional identities. By conceptualizing descriptive detail as a spectrum of granularity, we can examine the point at which an LLM shifts from generating homogenized responses to capturing authentic subgroup variance. Accordingly, we ask the following research question (RQ): RQ1: How does the granularity of segmentation identifiersâranging from singleâdimensional to layered, multidimensional conditioningâaffect the performance of LLM-based simulation? The second challenge, closely related to granularity, is the issue of quantitative parsimony. Shifting from âadding identity informationâ to âdesigning segmentation as an experimental factorâ foregrounds a classic methodological trade-off: determining whether additional identifiers contribute meaningful signal or instead introduce redundancy, noise, or over-constraint instead. Emerging evidence [undefai] indicates a non-linear relationship between informational complexity and simulation performance; while insufficient descriptive detail may produce âaverage personaâ bias, indiscriminately accumulating segmentation identifiers risks introducing noise that obscures subgroup variance [undefai]. Effective simulation therefore requires a balanceâan âinformative parsimonyââthat anchors the model without being over-regularized by irrelevant identifiers. Therefore, we ask the following research question: RQ2: What trade-offs emerge between parsimony and comprehensiveness as the number of identifiers used for persona conditioning increases, and how do these manifest in simulation performance across evaluation metrics? The third challenge concerns identifier selection logic. The field of LLM social simulation still lacks a unified framework for defining primary segmentation dimensions, making the rationale for choosing specific identifiers unclear and prone to inconsistent interpretation. For example, [undefaz] concatenated demographic characteristics with climate-related attitudinal variables to condition LLM predictions of public opinion. Although this improved predictive accuracy relative to demographics-only conditioning, such ad hoc data-driven configurations provide limited theoretical justification for why some identifiers are included while others are excluded. Without an explicit selection principle, performance gains may reflect opportunistic prompt tailoring that is practically useful but difficult to interpret as evidence of specific psychological or behavioral mechanisms. To move beyond case-by-case trial and error, identifier-selection logic should be made explicit and evaluated as an experimental factor. In this study, we operationalized selection logic into three distinct approaches: theory-driven selection, purely data-driven selection, and an important middle path, pre-validated instrument-based selection. Under this intermediate logic, identifiers were drawn from established survey instruments that were developed through theory-informed design and subsequently validated on large samples. A representative example is the Yale Program on Climate Change Communication (YPCCC) âSix Americasâ framework [undefam], which offers a validated segmentation instrument (question set) for assigning respondents to audience segments. By design, this approach is neither purely theoretical nor purely data-mined; it operationalizes theoretically meaningful constructs while retaining demonstrated predictive and classification performance. Therefore, we ask the following research question: RQ3: How do theory-driven, prevalidated instrumentâbased, and data-driven identifier selection logics compare in their effects on simulation performance? 5 Evaluating Heterogeneity Restoration To evaluate whether segmentation strategies restore population heterogeneity in theory-informed LLM-based social simulation, we use a comprehensive framework that diagnoses simulation performance across multiple dimensions. Existing studies have employed various metrics, but these methods were ad hoc and fragmented, with limited explicit connection to heterogeneity masking. Building on this gap, we propose a three-dimensional evaluation framework, informed by recent empirical studies [undefaj, undefal, undefak, undefaz, undefaab], to operationalize heterogeneity restoration across distributional, structural, and predictive fidelity (see Table 1). Each dimension targets a distinct facet of the extent to which social variation is preserved or masked. Table 1: Three-dimensional framework for evaluating heterogeneity restoration Dimension Theoretical focus Key metrics Distributional fidelity Overall distribution alignment Accuracy, Precision, Recall, F1 score; Mean Absolute Error (MAE); Kullback-Leibler divergence (KLD) Structural fidelity Subgroup variance preservation Within-group: Standard Deviation (SD), Coefficient of Variation (CV); Between-group: Normalized Earth moverâs distance (nEMD), Multidimensional scaling (MDS), Procrustes distance Predictive fidelity Relational pattern correspondence CramĂ©râs V 5.1 Distributional fidelity Distributional fidelity examines whether LLM-generated responses match the empirical distributions of the outcome variables in terms of both central tendency and distributional shape.Many prior studies have relied on mean-based comparisons to evaluate the aggregate similarity between simulated and human responses. For ordered categorical outcomes, especially Likert-style responses, the mean absolute error (MAE) is a useful complementary measure because it captures the absolute numerical deviation between simulated and empirical responses when response categories are treated as quasi-continuous. When responses are evaluated as discrete categories, classification metrics such as accuracy, precision, recall, and F1 score are used to assess agreement between simulated and empirical responses across response options. It is noteworthy, however, that agreement in means or category-level allocations does not necessarily indicate recovery of the full distributional structure. Prior work similarly cautions that matching aggregate statistics can still miss subgroup-level distributional patterns [undefal]. This suggests that LLMs may fit aggregate statistics by generating âaverageâ individuals, completely missing meaningful subgroup clusters [undefaan]. To address this limitation, distribution distance metrics such as Kullback-Leibler divergence (KLD), which quantifies the information loss when one distribution approximates another, offer a complementary assessment. 5.2 Structural fidelity Structural fidelity evaluates whether differences in variance across social subgroups, together with the relational patterns within and between them, are accurately preserved. This dimension directly addresses heterogeneity masking, specially, the tendency of LLMs to homogenize responses across social subgroups that should exhibit distinct patterns of variation. Structural fidelity can be decomposed into two complementary aspects (within-group variance and between-group structure). Within-group variance examines whether subgroups maintain internal diversity. This aspect is assessed using subgroup-level standard deviation (SD) and the coefficient of variation (CV = SD/Mean), which capture the degree of internal variation within social segments. These values are then compared with human benchmarks to evaluate whether LLM-generated subgroups are overly homogeneous. Between-group structure evaluates whether LLMs preserve systematic differences across potential social segments. We use the normalized Earth moverâs distance (nEMD) to quantify discrepancies between distributions across all subgroup pairs; higher values indicate stronger sensitivity of model outputs to subgroup labels [undefal]. Additionally, multidimensional scaling (MDS) is used to represent subgroup relations in a two-dimensional relational (identity) space, while Procrustes distance provides a complementary measure of alignment between simulated and empirical subgroup structures. 5.3 Predictive fidelity Social science is fundamentally concerned with understanding relationships between variables and phenomena rather than merely describing them in isolation. Predictive fidelity extends beyond simple distributional fit to examine relational pattern correspondence, that is, whether LLMs can accurately capture and reproduce the relational logic linking social positions (x) to attitudinal outcomes (y) as observed in real-world data. This dimension evaluates the extent to which LLMs have internalized authentic associational patterns that enable meaningful differentiation. Widely used measures of relational structure encompass correlation coefficients, such as Pearsonâs correlation for continuous variables and Spearmanâs correlation for ordinal data, as well as tetrachoric correlation for binary variables [undefaj]. CramĂ©râs V serves as a primary diagnostic for categorical associations that quantifies whether models faithfully reproduce structured preference patterns. By comparing CramĂ©râs V values for the same variable pairs (e.g., political affiliation and climate attitudes) across human and simulated responses, researchers can assess whether models preserve authentic relational structures or collapse responses into homogenized averages. If human data show moderate association but simulations produce near-deterministic patterns, this signals over-regularization, that is, the model has imposed exaggerated dependencies unsupported by empirical reality. Ultimately, evaluating predictive fidelity reveals whether models learn genuine relational patterns or merely spurious associations. However, high relational correspondence indicates successful âheterogeneity restorationâ only when both distributional and structural fidelity are adequately established. In summary, this three-dimensional framework integrates commonly used metrics from the LLM simulation literature into a coherent structure specifically designed to diagnose heterogeneity restoration. The dimensions are mutually reinforcing: Distributional fidelity captures whether simulated outputs reproduce aggregate response patterns, structural fidelity assesses whether subgroup heterogeneity and between-group structure are preserved, and predictive fidelity evaluates whether these subgroup differences support meaningful downstream associations. In what follows, we employ this evaluation framework to systematically assess the simulation performance of various segmentation configurations along three key dimensions: granularity (the richness of information included), parsimony (the number of segmentation identifiers), and variable selection logic (theory-driven, data-driven, or pre-validated instrument-based approaches). 6 Methodology 6.1 Data We employ a comparative design centered on climate opinion attitudes in the United States, leveraging both human and synthetic âsilicon sampleâ data. Human benchmark samples were collected in October 2025 through carefully stratified surveys distributed via Prolific, a leading online research platform trusted for diverse and representative participant pools. Sampling quotas were aligned with national census benchmarks to ensure demographic representativeness across key dimensions, including gender, age, and region. Strict protocols were implemented: Only unique (via IP filtering), attentive respondents were included, with invalid submissions filtered via response-time metrics and manual inspection. Following quality control, we retained 594 valid samples for further analysis. Respondentsâ baseline segment memberships were assigned using the Yale Program on Climate Change Communicationâs Six Americas Super Short Survey, which classifies individuals into one of six audience segments based on four climate-related attitudinal items. The six segments are alarmed, concerned, cautious, disengaged, doubtful, and dismissive (for details, see [undefam]). In our survey, each respondent completed the four-item questionnaire, which we used to assign an individual-level segment label. For synthetic silicon samples, we employed Llama 3.1-70B and Mixtral 8x22B, which were selected as high-performing open-weight models available at the time of data collection after benchmarking several alternatives (e.g., Qwen, Gemma3, and Bloom) against simulation-relevant criteria. Our criteria included (1) task completion success, (2) the absence of extraneous assistant commentary, (3) internal logical consistency under persona conditioning, and (4) diversity of attitudinal expression [undefaac]. These criteria, especially criterion 4, were deliberately designed to prioritize heterogeneity preservation rather than merely average performance. LLM completions were then systematically generated using persona conditioning aligned with each segmentation configuration. 6.2 Segmentation configurations To systematically examine how segmentation choices affect simulation performance, we manipulated three aspects of segmentation identifiers. Detailed overviews of the six configurations and their exact identifier wording are provided in Appendix A2âA3 (Tables A2âA8): Granularity. This aspect captures the richness of information used to construct persona prompts. We operationalized granularity by identifying a core set of identifiers from different dimensions (sociodemographic, psychological, and behavioral), and then progressively increasing the amount and dimensional breadth of conditioning information. These combinations were then used to craft prompts at varying levels of granularity, enabling us to parse the incremental simulation performance gains from including finer-grained, multidimensional segmentation identifiers. This design isolates how layered social identity information translates into more nuanced LLM subgroup distinctions and their contributions to improve or undermine simulation performance. Parsimony. This aspect addresses the trade-off between comprehensiveness and parsimony in selecting segmentation identifiers. We systematically compared the impact of multi-item identifiers with that of more streamlined, parsimonious identifiers. The comprehensive condition included a wide array of identifiers from each dimension (i.e., sociodemographic, psychographic, and behavioral), whereas the parsimonious condition was distilled into a reduced core subset of high-impact identifiers. We tested whether additional identifiers yielded diminishing or negative returns in terms of simulation performance. Identifier selection logic. This aspect contrasts with three distinct methodological approaches to selecting which identifiers to include in segmentation profiles. Theory-driven condition draws upon established theoretical models of climate attitudes and behavioral change, selecting identifiers with direct theoretical justification and prior empirical validation. In contrast, data-driven condition prioritizes statistical efficiency, applying machine learning techniques (e.g., Gradient Boosting Machines, GBM) to rank candidate identifiers and select those with the greatest discriminative power for segment differentiation. Given that our data focus on attitudes toward climate change, we further incorporated an empirical logic (i.e., a prevalidated instrumentâbased condition) that leverages the standardized SASSY instrument [undefam]. There are two versions of SASSY for audience segmentation, i.e., a detailed 15-item instrument (Item-15) and a shorter four-item version (Item-4). We implemented both versions for comparison. The comparison across and within identifier selection approaches aimed to elucidate the trade-offs between theoretical interpretability, statistical performance, and efficiency in constructing effective segmentation strategies. Based on these three aspects, we derived six segmentation configurations. We addressed the three RQs by juxtaposing and comparing these configurations (see Table 2 for details). The configurations are as follows: âą Demo: Baseline profiles built exclusively from standard demographic variables (i.e., gender, age, education, income, and ethnicity), paralleling common social simulation defaults. âą Demo+Theory-59: Profiles integrating demographic variables with 59 theoretically derived psychological and behavioral indicators relevant to climate attitudes, providing the richest and most comprehensive configuration. âą Demo+Theory-15: A streamlined, reduced version of the theory-informed profile including only the 15 most consequential predictors as identified by prior climate communication research. âą Data-driven: A profile using the top 15 high-impact identifiers as determined by gradient boosting machine-based feature selection, reflecting the data-driven approach to optimal subgroup discrimination. âą Item-15: Replicates the SASSY instrument, using 15 empirically validated climate-related attitudinal questions/items as identifiers (i.e., climate change beliefs, risk perceptions, and policy preferences) for segmentation. âą Item-4: Applies a condensed four-item version of the SASSY instrument for efficient, low-burden segmentation. Table 2: Segmentation configurations Segmentation configuration Granularity Parsimony Identifier selection logic Addressed RQs Demo Low (single-dimensional) High (5 identifiers) Empirical default Baseline for comparison Demo+Theory-59 Very high (multidimensional) Low (59 identifiers) Theory-driven RQ1, RQ2 Demo+Theory-15 High (multidimensional) Medium (15 identifiers) Theory-driven RQ1âRQ3 Data-driven High (multidimensional) Medium (15 identifiers) Data-driven (GBM) RQ3 Item-15 Medium (attitude-focused) Medium (15 identifiers) Pre-validated instrument-based RQ2, RQ3 Item-4 Medium (attitude-focused) High (4 identifiers) Pre-validated instrument-based RQ2, RQ3 6.3 Prompting strategies and responses All simulations were conducted using a zero-shot, Q&A prompt template. This format was chosen to construct an immersive context for the model to adopt a specified persona, while the zero-shot approach prevents the introduction of confounding variables that can arise from few-shot exemplars [undefaaf]. To encourage response diversity while maintaining coherence, a decoding temperature of 0.8 was employed in conjunction with a top-p value of 1.0 to prevent artificially constraining the output distribution [undefaj]. Each prompt elicited attitudinal responses (i.e., outcome items in this study) toward climate change under specified persona conditions. Specifically, the model was asked to rate its perception of climate change on three seven-point Likert scales: (1) the degree to which climate change is perceived as pleasant or unpleasant (pleasant), (2) its overall favorability or unfavourability (favorable), and (3) its general positivity or negativity (positivity). The corresponding questionnaire items are Q25, Q26, and Q27, respectively (see Table A1 in Appendix A for details). The modelâs task was to respond to each item by providing only the corresponding integer value (1â7), thereby ensuring standardized quantitative outputs suitable for statistical comparison with human survey data. 6.4 Evaluation metrics To evaluate how different segmentation strategies restore heterogeneity, we operationalized our evaluation framework along three complementary dimensionsâdistributional, structural, and predictive fidelityâand applied it consistently to all six segmentation configurations. Distributional fidelity was assessed using MAE, accuracy, precision, recall, F1 score, and KLD. MAE captured outcome itemâlevel numerical deviations between empirical and simulated mean responses, treating the 7-point Likert scale as quasi-continuous. Accuracy, precision, recall, and F1 score summarized agreement patterns at the response-category level (treating each Likert point as a discrete class), whereas KLD quantified information loss when the simulated distribution approximated the empirical response distribution. Because outcome variables were measured on a 7-point Likert scale and empirical response categories were highly imbalanced, we report weighted precision, weighted recall, and weighted F1 score to ensure that performance reflected observed category frequencies rather than being driven by rare response options. For each outcome item, we computed these measures by comparing the empirical frequency distribution with the corresponding LLM-generated frequency distribution and then summarized the results across outcomes to obtain configuration-level estimates of distributional fidelity. Structural fidelity was assessed at two levels: within-group variance and between-group structure. For within-group variance, we calculated subgroup-level SD and CV and compared these values with the corresponding empirical human benchmarks. Between-group structure was evaluated separately for each segmentation configuration and each LLM. Within each configuration, we represented each subgroup with its outcome distribution and computed nEMD for each subgroup pair, yielding a subgroup-by-subgroup distance matrix for the empirical data and a corresponding distance matrix for each simulated LLM output. Pairwise nEMD captured distributional differences between subgroups. To summarize between-group structure, we first computed the median pairwise nEMD within each outcome item and then averaged these outcome itemâlevel medians across the outcome variables to obtain a configuration-level estimate. To assess whether subgroup differences were organized similarly in the empirical and simulated data, we applied MDS to each distance matrix. MDS projects each subgroup distance matrix into a two-dimensional map in which subgroups with response distributions that are more similar appear closer together, whereas subgroups that are more distinct appear farther apart. In this way, the MDS map provides a visual representation of the overall pattern of subgroup relationships. Because two MDS maps could represent the same underlying structure while differing in orientation, position, or scale, they could not be compared directly using raw coordinates. We therefore used a Procrustes transformation to align each simulated map with the empirical map before comparison. The resulting Procrustes distance quantified the remaining mismatch between the two subgroup structures after these geometric differences were removed. Predictive fidelity examined whether segmentation identifiers retained their empirical association with outcome items in the simulated data. We operationalized this dimension via CramĂ©râs V, computed from contingency tables between each segmentation identifier and each outcome item. CramĂ©râs V provides a standardized effect-size estimate of association strength, enabling direct comparison of whether a configuration underestimates, matches, or inflates empirical subgroupâoutcome linkages. We summarize predictive fidelity by aggregating association estimates across identifiers and outcomes. By jointly evaluating segmentation strategies across distributional, structural, and predictive fidelity, we were able to distinguish configurations that yielded comprehensive heterogeneity restoration from those that merely optimized a single metric while degrading other properties of the empirical system. 7 Results This section evaluates simulation performance along three complementary dimensions. We present the main findings in relation to the three RQs. 7.1 RQ1: Increasing granularity does not produce consistent improvements in simulation performance To answer RQ1, we compared three segmentation configurations: Demo, Demo+Theory-15, and Demo+Theory-59. In terms of distributional fidelity (see Appendix A4, Table 9 and Figure 1), the results show that increasing granularity does not result in uniform improvement across metrics. Moving from Demo to Demo+Theory-15 generally improved performance, especially in KLD and several category-level classification metrics. However, moving further to Demo+Theory-59 did not provide additional gains and, in some respects, worsened performance. The Demo configuration performed worst in distributional shape for both LLMs (KLD: 2.72 for Llama and 6.58 for Mixtral) and yielded the lowest weighted F1 scores (.46 and .39, respectively), indicating that demographics-only conditioning departs substantially from the empirical distribution. Demo+Theory-15 improved this baseline in both LLMs. For Llama, KLD decreased from 2.72 to .68, F1 increased from .46 to .51, and accuracy/recall also improved, although MAE rose modestly from .38 to .46. For Mixtral, MAE decreased from .89 to .80, KLD decreased from 6.58 to 1.51, and F1 increased from .39 to .40. Cross-LLM averages reinforced this pattern. Relative to Demo, Demo+Theory-15 reduced average KLD from 4.65 to 1.10 and increased average F1 from .43 to .46, while average MAE remained nearly unchanged (.64 vs. .63). In contrast, Demo+Theory-59 showed weaker average performance (KLD = 2.76, F1 = .45, MAE = .69). Category-level metrics changed only modestly across configurations: Average accuracy increased from .52 (Demo) to .555 (Demo+Theory-59) and .560 (Demo+Theory-15); average precision from .39 to .40 and .42; average recall from .52 to .555 and .56; and average F1 from .425 to .445 and .455. Taken together, these comparisons indicate that additional granularity beyond the intermediate level did not improve overall distributional fidelity. Figure 1: Comparison of distributional fidelity metrics across segmentation configurations and LLMs. In terms of structural fidelity (see Appendix A4, Table 10 and Figure 2), the results suggest that intermediate granularity (Demo+Theory-15) generally outperformed both the Demo baseline and the more granular Demo+Theory-59 configuration. Specifically, Demo+Theory-15 preserved substantially more within-group variance than the Demo configuration in both LLMs: SD increased from .46 (Demo) to .69 (Demo+Theory-15) in Llama and from .02 to .60 in Mixtral, with CV rising correspondingly from .27 to .37 and from .02 to .39. On average across both LLMs, SD increased from .24 under the Demo configuration to .63 under Demo+Theory-15, and CV rose from .14 to .38. In contrast, the Demo+Theory-59 configuration did not sustain this improvement. In Llama, SD fell back to .44 and CV to .34, both below the Demo+Theory-15 configuration and closer to the Demo baseline; in Mixtral, SD and CV shifted to .63 and .32, respectively. Cross-LLM averages also revealed that Demo+Theory-15 slightly outperformed the Demo+Theory-59 configuration (SD: .54; CV: .33), suggesting that adding more identifiers did not consistently improve the recovery of within-group variance. Figure 2: Comparison of structural fidelity metrics across segmentation configurations and LLMs. Between-group structure exhibited a similar limit to increasing granularity. Although the Demo+Theory-15 configuration (median nEMD = .10 for Llama, .06 for Mixtral; average = .08) and the Demo+Theory-59 configuration (median nEMD = .04 for Llama, .05 for Mixtral; average = .04) yielded relatively higher median nEMD values than the Demo configuration (median nEMD = .07 for Llama, .00 for Mixtral; average = .03), both remained well below the human benchmark of .19. This pattern indicated that subgroup pairs in the simulations remained less differentiated from one another than they were empirically, and that increasing granularity beyond an intermediate level did not strengthen between-group structure. Figure 3 complemented Appendix A4, Table 10 by visualizing these subgroup patterns directly against the human benchmark. Procrustes distance provided an additional summary of how closely the simulated subgroup map resembled the empirical one. In Llama, Procrustes distance rose from .10 under the Demo baseline to .49 under Demo+Theory-15 and to .91 under Demo+Theory-59; in Mixtral, it decreased from .96 to .50 and then increased again to .60. While the LLM-specific baseline distances varied, the cross-LLM average Procrustes distance improved from .53 under the Demo baseline to .49 under Demo+Theory-15; however, expanding to Demo+Theory-59 worsened the average distance to .75. Taken together, these results suggested that increasing the granularity of segmentation identifiers beyond an informative threshold did not bring the simulated subgroup structure closer to the empirical one; instead, subgroup differences remained weaker than in the human data, and the overall subgroup pattern became less stable. Figure 3: MDS maps of empirical and simulated subgroup structures across segmentation configurations and LLMs. Note. Colors denote subgroup identity and are held constant across panels, allowing direct comparison between empirical and simulated subgroup locations. Distances between points reflect differences in subgroup response distributions derived from pairwise nEMD, such that closer points indicate more similar subgroup distributions. The reported nEMD in each panel summarizes the overall magnitude of between-group structure within that segmentation configuration. MDS Dimensions 1 and MDS Dimensions 2 do not have direct substantive interpretations; thin gray lines connect each simulated subgroup to its empirical counterpart only to facilitate visual comparison and should not be interpreted as between-group distances themselves. In terms of predictive fidelity, Table 3 and Figure 4 show that, among the demographic and theory-based configurations, Demo was consistently furthest from its human benchmark in both LLMs, with a cross-LLM average simulated CramĂ©râs V of .17 against a human benchmark that ranged from .19 to .26. The relative performance of the Demo+Theory-15 and Demo+Theory-59 configurations, however, was unstable across LLMs. In Llama, Demo+Theory-15 produced a simulated CramĂ©râs V of .23 against a human benchmark of .26 (a difference of .03), while Demo+Theory-59 reached only .19 against its benchmark of .25 (a difference of .05), making the intermediate configuration the better one. In Mixtral, this order is reversed: Demo+Theory-59 yielded a smaller difference from the human benchmark (.03) than Demo+Theory-15 (.08). Overall, these results indicated that the point at which additional identifiers contributed informative signals rather than noise was not fixed across LLMs and could not be naively determined by the granularity of identifiers. Table 3: Summary of CramĂ©râs V values Segmentation configuration Human benchmark Llama Mixtral Demo .19 .27 .08 Demo+Theory-59 .25 .19 .21 Demo+Theory-15 .26 .23 .18 Data-driven .28 .29 .33 Item-15 .34 .24 .21 Item-4 .39 .40 .44 Figure 4: Comparison of CramĂ©râs V across segmentation configurations and LLMs. Brackets and numbers indicate the differences from the human benchmark. 7.2 RQ2: Informative parsimony outperforms comprehensiveness To address RQ2, we compared two pairs of configurations: a theory-driven pair (Demo+Theory-15 vs. Demo+Theory-59) and a pre-validated instrument-based pair (Item-4 vs. Item-15). Because the theory-driven contrast has already been discussed in the previous section, here, we focus here on the instrument-based comparison Regarding distributional fidelity, the instrument-based pair exhibited a more conditional pattern than the theory-driven pair (see Appendix A4, Table 9 and Figure 1). Item-15 clearly improved KLD in Llama (.17 vs. 1.04 for Item-4) and matched Item-4 in Mixtral (both .76), yielding a lower cross-LLM average KLD overall (.47 for Item-15 vs. .90 for Item-4). However, MAE showed the opposite pattern across LLMs: Item-15 was better in Llama (.17 vs. .38), whereas Item-4 was better in Mixtral (.34 vs. .66), and Item-4 was slightly better in terms of the cross-LLM average (.36 vs. .42). Category-level classification metrics differed only modestly between the two configurations (all differences †.08), indicating that the main distributional contrast was concentrated in KL divergence and MAE. Overall, these results suggested that parsimony did not inherently impair distributional fidelity compared to more comprehensive configurations. Structural fidelity (see Appendix A4, Table 10 and Figure 2) did not mirror the pattern observed for distributional fidelity. In general, Item-15 better recovered within-group variance, whereas Item-4 tended to preserve somewhat stronger between-group differentiation; neither configuration reliably recovered between-group relational geometry. For within-group variance, Item-15 was consistently closer to the human benchmark (SD = 1.19, CV = .53) than Item-4 in both LLMs. The contrast was strongest in Llama, where Item-15 reached SD = 1.10 and CV = .52, compared with SD = .47 and CV = .30 for Item-4. Cross-LLM averages showed the same pattern: Item-15 reached SD = .99 and CV = .54, whereas Item-4 reached SD = .74 and CV = .41. For between-group structural differentiation, Item-4 was slightly stronger overall between-group structural differentiation, with a cross-LLM average median nEMD of .15, compared with .10 for Item-15. In Mixtral, Item-4 (.20) was also the only configuration of the two to reach the human benchmark (.19), whereas Item-15 remained far below (.07). In Llama, both remained below the benchmark (Item-15: .12; Item-4: .09). Procrustes results (see Figure 3) further showed weak relational recovery between simulated and empirical subgroup structures. Relational recovery between simulated and empirical subgroups remained weak: Distances remained high for both segmentation configurations (Item-15: .64â.75; Item-4: .87â.91), with cross-LLM averages of .70 and .89, respectively. Taken together with the theory-driven pair results, these findings indicated that greater identifier comprehensiveness did not yield consistent structural gains; more broadly, structural performance appeared to depend not simply on the number of added identifiers, but on whether those identifiers carried discriminative information relevant to subgroup structure. Predictive fidelity showed the clearest advantage of Item-4 over Item-15 (see Table 3 and Figure 4). Under the Item-4 configuration, the human benchmark was .39, while the simulated CramĂ©râs V values were .40 (Llama) and .44 (Mixtral), corresponding to gaps of .01 and .05, respectively. Under the Item-15 configuration, the human benchmark was .34, while the simulated CramĂ©râs V values were .24 (Llama) and .21 (Mixtral), with much larger gaps of .10 and .13. At the aggregated level, the cross-LLM average gap was therefore substantially smaller for Item-4 (.03) than for Item-15 (.11), indicating better preservation of empirical identifierâoutcome associations under the more parsimonious configuration. Combined with the theory-driven comparisonâwhere Demo+Theory-59 also departed further from the benchmark than Demo+Theory-15, at least in Llamaâthese findings suggested that adding identifiers beyond an informative threshold might impair predictive fidelity rather than improve it. In summary, all the evidence for RQ2 therefore lent stronger support to informative parsimony than to comprehensiveness; in other words, adding more identifiers did not reliably improve simulation performance, whereas a more compact and informative segmentation configuration tended to yield more robust overall performance. 7.3 RQ3: The logic of identifier selection matters To answer RQ3, we compared three segmentation configurations that differed in identifier selection logic while holding the number of identifiers constant: Demo+Theory-15, Data-driven, and Item-15. Across the theory-driven, data-driven, and prevalidated instrumentâbased selection logics, the resulting configurations exhibited distinct fidelity profiles. It was found that in general, no single configuration performed best across all three evaluative dimensions and that the dimension that benefited most depended on how the identifiers were selected. In terms of distributional fidelity, no single selection logic dominated all metrics (see Appendix A4, Table 9 and Figure 1). Item-15 performed best on KLD in both LLMs (.17 in Llama; .76 in Mixtral), yielding the lowest cross-LLM average KLD (.47), compared with Data-driven (.63) and Demo+Theory-15 (1.10). Moreover, the Item-15 and Data-driven configurations performed better on MAE overall, with a lower cross-LLM average MAE (.42 for Item-15; .37 for Data-driven) than Demo+Theory-15 (.63). Category-level classification metrics were relatively similar across the three configurations in both LLMs, with accuracy ranging from .48 to .57 and F1 score from .40 to .51. Taken together, these results suggested that the pre-validated instrument-based segmentation configuration provided the strongest distributional alignment with the empirical data. Structural fidelity (see Appendix A4, Table 10 and Figure 2) did not indicate a single uniform winner across all metrics. Instead, the three configurations showed different strengths. For within-group variance, Item-15 performed best overall. Its cross-LLM averages (SD = .99, CV = .54) were closest to the human benchmarks (SD = 1.19, CV = .53), compared with Data-driven (SD = .87, CV = .37) and Demo+Theory-15 (SD = .64, CV = .38). Data-driven was intermediate in terms of SD but remained farther from the benchmark in terms of CV, especially in Mixtral (CV = .31). For between-group structure, the pattern reversed. Data-driven showed the strongest differentiation overall, with a cross-LLM average median nEMD of .18, which was closest to the human benchmark (.19), while Demo+Theory-15 and Item-15 remained lower (.08 and .10, respectively). This advantage was driven mainly by Mixtral, wherein Data-driven reached .25; in Llama, all three configurations remained below the benchmark (Demo+Theory-15: .10, Data-driven: .11, Item-15: .12). Geometric alignment (i.e., structural relation between simulated and empirical subgroups) showed a similar tendency: Data-driven had the lowest cross-LLM average Procrustes distance (.48), compared with Demo+Theory-15 (.49) and Item-15 (.70). Taken together, these results suggested that the Data-driven configuration yields the strongest between-group structural recovery, whereas Item-15 better captures within-group variance. Predictive fidelity revealed a different pattern (see Table 3 and Figure 4). Because the three segmentation configurations had different human benchmark CramĂ©râs V values (Data-driven: .28; Demo+Theory-15: .26; Item-15: .34), the most informative comparison was each configurationâs absolute deviation from its own benchmark. Data-driven showed the smallest gaps in both LLMs (Llama: .02; Mixtral: .05), while Demo+Theory-15 had an intermediate gap (Llama: .03; Mixtral: .08), and Item-15 had the largest gap (Llama: .10; Mixtral: .13). The cross-LLM average gaps confirmed this pattern: .03 for Data-driven, .05 for Demo+Theory-15, and .12 for the Item-15 configuration. These results indicated that the Data-driven configuration best preserved empirical identifierâoutcome association strength, whereas the Item-15 configuration failed to recover its higher benchmark associations despite having the highest empirical CramĂ©râs V. Taken together, the evidence across the three fidelity dimensions indicated that identifier selection logic determined which dimension benefited most: Item-15 best preserved distributional shape, Data-driven best recovered between-group structure and identifierâoutcome associations, and Demo+Theory-15 showed intermediate performance without a clear advantage in any single evaluative dimension. 7.4 Summary of findings Across RQ1âRQ3, the results show that simulation performance depends less on the sheer number of identifiers than on whether those identifiers are informative in terms of the target task. First, the single-dimensional demographics-only configuration was an unreliable foundation across all three fidelity dimensions. It produced the greatest distributional divergence, weakened subgroup structure, and distorted association patterns. However, it is noteworthy that increasing granularity did not produce consistent improvements. Moving from Demo to Demo+Theory-15 improved several metrics, but further expansion to Demo+Theory-59 did not consistently improve performance and worsened structural and predictive fidelity. Second, the parsimonyâcomprehensiveness trade-off is fidelity specific. Parsimonious configurations often match or outperform more comprehensive alternatives, especially for structural and predictive fidelity, while distributional fidelity remains mixed depending on the metric (e.g., KLD vs. MAE). Third, when the number of identifiers is held constant, selection logic determines which fidelity dimension benefits most: The instrument-based configuration best preserves distributional shape, the data-driven configuration best recovers between-group structure and identifierâoutcome associations, and the theory-driven configuration is generally intermediate. In short, no single segmentation configuration dominates all evaluative dimensions; gains in one dimension can come with weaker performance in another. Table 4 summarizes these patterns. Table 4: Summary of simulation performance Distributional fidelity Structural fidelity Predictive fidelity Granularity More â better More â better More â better Parsimony Mixed (metric-dependent) Fewer generally better Fewer generally better Identifier selection logic Instrument-based best Data-driven and instrument-based best Data-driven best 8 Discussion This study demonstrates that restoring heterogeneity in LLM-based social simulation is both necessary and feasible but only when audience segmentation is designed deliberately. Moving beyond surface-level representational accuracy requires a population-thinking perspective and a robust method for operationalizing social variation. Through systematic experimentation and rigorous evaluation, our results suggest that audience segmentation can play this role by embedding individual- and subgroup-level heterogeneity into persona conditioning in a systematic and testable way. Three key conclusions arise. First, increasing identifier granularity does not produce consistent improvements: Moving from single-dimensional to moderately enriched conditioning can improve simulation performance, but further expansion does not reliably help and may impair structural and predictive fidelity. Second, there is a clear parsimonyâcomprehensiveness trade-off: Additional identifiers often bring diminishing returns, and parsimonious segmentation configurations frequently match or outperform more comprehensive ones, especially on structural and predictive fidelity. Third, identifier selection logic determines which fidelity dimension benefits the most. Specifically, instrument-based identifier selection best preserves distributional fidelity, whereas data-driven logic better recovers between-group structure and identifierâoutcome associations; theory-driven selection is generally intermediate. These findings shift the key focus from âhow many identifiersâ to âwhich identifiers, and why.â In short, simulation performance depends less on identifier quantity than on the theoretical or empirical relevance of the identifiers used. This echoes longstanding principles from audience segmentation and population-thinking traditions [undefaag, undefaao], which treat variation as meaningful social signals rather than residual noise. These findings also help make sense of the mixed results in prior LLM-based social simulation studies [undefaj, undefak, undefaam]. Some studies reported an âaverage personaâ tendency (variance suppression) [undefaam, undefaae, undefaan], while others reported unstable or uneven subgroup differentiation [undefaj, undefaak]. A plausible explanation is that these differences arose from discrepancies in LLM architecture, the intensity and objectives of reinforcement learning from human feedback, and persona prompting design [undefaac, undefaw, undefaam]. From this perspective, audience segmentation functions as a partial corrective in both directions: It can reintroduce social structure when variance is overly suppressed, and it can impose informative constraints when LLM responses become unrealistically extreme. Methodologically, this study advances LLM evaluation practice by proposing and using a heterogeneity-aware three-dimensional fidelity framework rather than relying on surface-level average-matching. By jointly assessing distributional, structural, and predictive fidelity, we provide a more rigorous basis for assessing simulation performance. This is important as overreliance on averages risks missing the texture and complexity of social phenomena and can conceal substantial errors in subgroup differentiation [undefaao, undefaam, undefaac]. Our analyses also highlight several limitations. All segmentation configurations, even the best-performing ones, show residual over-regularization; that is, simulated responses remain more orderly than observed human data, and subgroup differences are still compressed or restructured. Therefore, while audience segmentation is a substantial improvement, it is not a panacea. Further progress likely requires variance-preserving advances in model training, alignment, and calibration [undefaac, undefaam]. Looking forward, we recommend three priorities for further research. First, segmentation should be treated not as an afterthought or mere technicality, but as a central experimental factor that is methodologically deliberate and informed by theory rather than simply âadded onâ for realism [undefaah, undefao, undefau]. Careful documentation and justification of segmentation logic will facilitate cumulative progress in the field. Second, the explicit use of heterogeneity-aware metrics should become the new gold standard for simulation evaluation. Only by quantifying and reporting the preservation (or masking) of variation can we meaningfully improve LLMs and prompt design. Third, cross-LLM comparisons should be a research priority, and attention should be paid to how different LLMs encode (or erase) social diversity [undefaw, undefaam, undefaac]. In conclusion, high-fidelity LLM-based social simulation is not only an engineering challenge of data scale and model training but also a methodological challenge of representing population heterogeneity. Our findings position audience segmentation as a foundational component of this agenda and offer an empirical basis for heterogeneity-centered simulation design. References [undef] Jacy Reese Anthis et al. âPosition: LLM social simulations are a promising research methodâ In Proceedings of the 42nd International Conference on Machine Learning 267, Proceedings of Machine Learning Research PMLR, 2025, p. 81005â81034 URL: https://proceedings.mlr.press/v267/anthis25a.html [undefa] Lisa P. Argyle et al. âOut of one, many: Using language models to simulate human samplesâ In Political Analysis 31.3, 2023, p. 337â351 [undefb] James Bisbee et al. âSynthetic replacements for human survey data? The perils of large language modelsâ In Political Analysis 32.4, 2024, p. 401â416 [undefc] Julien Boelaert et al. âMachine bias. How do generative language models answer opinion polls?â In Sociological Methods & Research 54.3, 2025, p. 1156â1196 [undefd] Brooke Chryst et al. âGlobal warmingâs âsix Americas short surveyâ: Audience segmentation of climate change views using a four-question instrumentâ In Environmental Communication 12.8, 2018, p. 1109â1122 DOI: 10.1080/17524032.2018.1508047 [undefe] Thomas Davidson and Daniel Karell âIntegrating generative artificial intelligence into social science research: Measurement, prompting, and simulationâ In Sociological Methods & Research 54.3, 2025, p. 775â793 [undeff] Sally Dibb and Lyndon Simkin âImplementation rules to bridge the theory/practice divide in market segmentationâ In Journal of Marketing Management 25.3-4, 2009, p. 375â396 [undefg] D. Dillion, N. Tandon, Y. Gu and K. Gray âCan AI language models replace human participants?â In Trends in Cognitive Sciences 27.7, 2023, p. 597â600 DOI: 10.1016/j.tics.2023.04.008 [undefh] C. Gao et al. âLarge language models empowered agent-based modeling and simulation: A survey and perspectivesâ In Humanities and Social Sciences Communications 11.1, 2024, p. 1â24 [undefi] Igor Grossmann et al. âAI and the transformation of social science researchâ In Science 380.6650, 2023, p. 1108â1109 [undefj] Lars Guenther and Peter Weingart âPromises and reservations towards science and technology among South African publics: A culture-sensitive approachâ In Public Understanding of Science 27.1, 2018, p. 47â58 [undefk] Peter Hedström and Petri Ylikoski âCausal mechanisms in the social sciencesâ In Annual Review of Sociology 36.1, 2010, p. 49â67 [undefl] Donald W. Hine et al. âAudience segmentation and climate change communication: Conceptual and methodological considerationsâ In Wiley Interdisciplinary Reviews: Climate Change 5.4, 2014, p. 441â459 [undefm] Jake M. Hofman, Amit Sharma and Duncan J. Watts âPrediction and explanation in social systemsâ In Science 355.6324, 2017, p. 486â488 [undefn] Tian Hu et al. âGenerative language models exhibit social identity biasesâ In Nature Computational Science 5.1, 2025, p. 65â75 [undefo] Bernard J. Jansen, Sun Jung and Joni Salminen âEmploying large language models in survey researchâ In Natural Language Processing Journal 4, 2023, p. 100020 DOI: 10.1016/j.nlp.2023.100020 [undefp] H.. Kirk, B. Vidgen, P. Râottger and S.. Hale âThe benefits, risks and bounds of personalizing the alignment of large language models to individualsâ In Nature Machine Intelligence 6.4, 2024, p. 383â392 [undefq] Seojin Lee et al. âCan large language models estimate public opinion about global warming? An empirical assessment of algorithmic fidelity and biasâ In PLOS Climate 3.8, 2024, p. e0000429 [undefr] Y. Lu et al. âLLM Agents That Act Like Us: Accurate Human Behavior Simulation with Real-World Dataâ In arXiv e-prints, 2025 arXiv:arXiv-2503 [undefs] Y. Luo, J.. Du and L. Qiu âOn the cultural sensitivity of large language models: GPTâs ability to simulate human self-conceptâ In 2024 11th International Conference on Behavioral and Social Computing (BESC) IEEE, 2024, p. 1â8 [undeft] Andrew Lyman et al. âBalancing large language model alignment and algorithmic fidelity in social science researchâ In Sociological Methods & Research 54.3, 2025, p. 1110â1155 DOI: 10.1177/00491241251330582 [undefu] Joon Sung Park et al. âGenerative agents: Interactive simulacra of human behaviorâ In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023, p. 1â22 [undefv] Luca Rossi, Kate Harrison and Irina Shklovski âThe problems of LLM-generated data in social science researchâ In Sociologica: International Journal for Sociological Debate 18.2, 2024, p. 145â168 [undefw] Shibani Santurkar et al. âWhose opinions do language models reflect?â In International Conference on Machine Learning, Proceedings of Machine Learning Research PMLR, 2023, p. 29971â30004 [undefx] Michael D. Slater âCommunication: The American Healthstyles audience segmentation projectâ In Journal of Health Psychology 1.3, 1995, p. 261â277 [undefy] Michael D. Slater âTheory and method in health audience segmentationâ In Journal of Health Communication 1.3, 1996, p. 267â284 [undefz] J. Suh et al. âLanguage model fine-tuning on scaled survey data for predicting distributions of public opinionsâ, 2025 arXiv:2502.16761 [undefaa] S. Sun et al. âRandom silicon sampling: Simulating human sub-population opinion using a large language model based on group-level demographic informationâ, 2024 arXiv:2402.18144 [undefab] Petter Törnberg, Daria Valeeva, Justus Uitermark and Christopher Bail âSimulating social media using large language models to evaluate alternative news feed algorithmsâ, 2023 arXiv:2310.05984 [undefac] John C. Turner, Rupert J. Brown and Henri Tajfel âSocial comparison and group interest in ingroup favoritismâ In European Journal of Social Psychology 9.2, 1979, p. 187â204 [undefad] Angelina Wang, Jamie Morgenstern and John P. Dickerson âLarge language models that replace human participants can harmfully misportray and flatten identity groupsâ In Nature Machine Intelligence, 2025, p. 1â12 [undefae] Z. Wu, R. Peng, T. Ito and C. Xiao âLLM-based social simulations require a boundaryâ, 2025 arXiv:2506.19806 [undefaf] Yu Xie âPopulation heterogeneity and causal inferenceâ In Proceedings of the National Academy of Sciences 110.16, 2013, p. 6262â6268 [undefag] Yu Xie âLocalization of sociology and reconsideration of quantitative researchâ In Academic Monthly 56.3, 2024, p. 120â128 [undefah] K. Yang et al. âAre large language models (LLMs) good social predictors?â, 2024 arXiv:2402.12620 References [undefai] Jacy Reese Anthis et al. âPosition: LLM social simulations are a promising research methodâ In Proceedings of the 42nd International Conference on Machine Learning 267, Proceedings of Machine Learning Research PMLR, 2025, p. 81005â81034 URL: https://proceedings.mlr.press/v267/anthis25a.html [undefaj] Lisa P. Argyle et al. âOut of one, many: Using language models to simulate human samplesâ In Political Analysis 31.3, 2023, p. 337â351 [undefak] James Bisbee et al. âSynthetic replacements for human survey data? The perils of large language modelsâ In Political Analysis 32.4, 2024, p. 401â416 [undefal] Julien Boelaert et al. âMachine bias. How do generative language models answer opinion polls?â In Sociological Methods & Research 54.3, 2025, p. 1156â1196 [undefam] Brooke Chryst et al. âGlobal warmingâs âsix Americas short surveyâ: Audience segmentation of climate change views using a four-question instrumentâ In Environmental Communication 12.8, 2018, p. 1109â1122 DOI: 10.1080/17524032.2018.1508047 [undefan] Thomas Davidson and Daniel Karell âIntegrating generative artificial intelligence into social science research: Measurement, prompting, and simulationâ In Sociological Methods & Research 54.3, 2025, p. 775â793 [undefao] Sally Dibb and Lyndon Simkin âImplementation rules to bridge the theory/practice divide in market segmentationâ In Journal of Marketing Management 25.3-4, 2009, p. 375â396 [undefap] D. Dillion, N. Tandon, Y. Gu and K. Gray âCan AI language models replace human participants?â In Trends in Cognitive Sciences 27.7, 2023, p. 597â600 DOI: 10.1016/j.tics.2023.04.008 [undefaq] C. Gao et al. âLarge language models empowered agent-based modeling and simulation: A survey and perspectivesâ In Humanities and Social Sciences Communications 11.1, 2024, p. 1â24 [undefar] Igor Grossmann et al. âAI and the transformation of social science researchâ In Science 380.6650, 2023, p. 1108â1109 [undefas] Lars Guenther and Peter Weingart âPromises and reservations towards science and technology among South African publics: A culture-sensitive approachâ In Public Understanding of Science 27.1, 2018, p. 47â58 [undefat] Peter Hedström and Petri Ylikoski âCausal mechanisms in the social sciencesâ In Annual Review of Sociology 36.1, 2010, p. 49â67 [undefau] Donald W. Hine et al. âAudience segmentation and climate change communication: Conceptual and methodological considerationsâ In Wiley Interdisciplinary Reviews: Climate Change 5.4, 2014, p. 441â459 [undefav] Jake M. Hofman, Amit Sharma and Duncan J. Watts âPrediction and explanation in social systemsâ In Science 355.6324, 2017, p. 486â488 [undefaw] Tian Hu et al. âGenerative language models exhibit social identity biasesâ In Nature Computational Science 5.1, 2025, p. 65â75 [undefax] Bernard J. Jansen, Sun Jung and Joni Salminen âEmploying large language models in survey researchâ In Natural Language Processing Journal 4, 2023, p. 100020 DOI: 10.1016/j.nlp.2023.100020 [undefay] H.. Kirk, B. Vidgen, P. Râottger and S.. Hale âThe benefits, risks and bounds of personalizing the alignment of large language models to individualsâ In Nature Machine Intelligence 6.4, 2024, p. 383â392 [undefaz] Seojin Lee et al. âCan large language models estimate public opinion about global warming? An empirical assessment of algorithmic fidelity and biasâ In PLOS Climate 3.8, 2024, p. e0000429 [undefaaa] Y. Lu et al. âLLM Agents That Act Like Us: Accurate Human Behavior Simulation with Real-World Dataâ In arXiv e-prints, 2025 arXiv:arXiv-2503 [undefaab] Y. Luo, J.. Du and L. Qiu âOn the cultural sensitivity of large language models: GPTâs ability to simulate human self-conceptâ In 2024 11th International Conference on Behavioral and Social Computing (BESC) IEEE, 2024, p. 1â8 [undefaac] Andrew Lyman et al. âBalancing large language model alignment and algorithmic fidelity in social science researchâ In Sociological Methods & Research 54.3, 2025, p. 1110â1155 DOI: 10.1177/00491241251330582 [undefaad] Joon Sung Park et al. âGenerative agents: Interactive simulacra of human behaviorâ In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023, p. 1â22 [undefaae] Luca Rossi, Kate Harrison and Irina Shklovski âThe problems of LLM-generated data in social science researchâ In Sociologica: International Journal for Sociological Debate 18.2, 2024, p. 145â168 [undefaaf] Shibani Santurkar et al. âWhose opinions do language models reflect?â In International Conference on Machine Learning, Proceedings of Machine Learning Research PMLR, 2023, p. 29971â30004 [undefaag] Michael D. Slater âCommunication: The American Healthstyles audience segmentation projectâ In Journal of Health Psychology 1.3, 1995, p. 261â277 [undefaah] Michael D. Slater âTheory and method in health audience segmentationâ In Journal of Health Communication 1.3, 1996, p. 267â284 [undefaai] J. Suh et al. âLanguage model fine-tuning on scaled survey data for predicting distributions of public opinionsâ, 2025 arXiv:2502.16761 [undefaaj] S. Sun et al. âRandom silicon sampling: Simulating human sub-population opinion using a large language model based on group-level demographic informationâ, 2024 arXiv:2402.18144 [undefaak] Petter Törnberg, Daria Valeeva, Justus Uitermark and Christopher Bail âSimulating social media using large language models to evaluate alternative news feed algorithmsâ, 2023 arXiv:2310.05984 [undefaal] John C. Turner, Rupert J. Brown and Henri Tajfel âSocial comparison and group interest in ingroup favoritismâ In European Journal of Social Psychology 9.2, 1979, p. 187â204 [undefaam] Angelina Wang, Jamie Morgenstern and John P. Dickerson âLarge language models that replace human participants can harmfully misportray and flatten identity groupsâ In Nature Machine Intelligence, 2025, p. 1â12 [undefaan] Z. Wu, R. Peng, T. Ito and C. Xiao âLLM-based social simulations require a boundaryâ, 2025 arXiv:2506.19806 [undefaao] Yu Xie âPopulation heterogeneity and causal inferenceâ In Proceedings of the National Academy of Sciences 110.16, 2013, p. 6262â6268 [undefaap] Yu Xie âLocalization of sociology and reconsideration of quantitative researchâ In Academic Monthly 56.3, 2024, p. 120â128 [undefaaq] K. Yang et al. âAre large language models (LLMs) good social predictors?â, 2024 arXiv:2402.12620