Paper deep dive
Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception
Neemias da Silva, Myriam Delgado, Rodrigo Minetto, Daniel Silver, Thiago H Silva
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/8/2026, 9:55:34 PM
Summary
This paper investigates how persona prompting influences the language generated by multimodal large language models (MLLMs) in urban perception tasks. Analyzing 59,808 annotations from 1,200 persona-conditioned agents, the study distinguishes between descriptive grounding (captions) and interpretive framing (justifications). Results show that captions exhibit strong convergence across diverse personas, while justifications display systematic variation driven primarily by socioeconomic and political attributes. The findings suggest that persona effects concentrate in interpretive language rather than descriptive content, highlighting the need for nuanced evaluation of LLMs as social simulators.
Entities (14)
Relation Signals (10)
Daniel Silver → affiliatedwith → University of Toronto
confidence 95% · Daniel Silver 2 ... 2 University of Toronto, Toronto, Canada
Neemias da Silva → affiliatedwith → Universidade Tecnologica Federal do Parana
confidence 95% · Correspondence: neemiasbuceli@alunos.utfpr.edu.br ... Universidade Tecnologica Federal do Parana
Justifications → represent → Interpretive Framing
confidence 95% · interpretive framing, reflected in justifications that evaluate and contextualize that content.
Captions → represent → Descriptive Grounding
confidence 95% · descriptive grounding, reflected in captions that describe visual content
Persona Prompting → shapes → Generated Language
confidence 95% · We study how persona prompting shapes language generated by multimodal large language models in an urban perception setting.
Multimodal Large Language Models → appliedto → Urban Perception
confidence 90% · We study how persona prompting shapes language generated by multimodal large language models in an urban perception setting.
Socioeconomic Attributes → drivevariationin → Justifications
confidence 90% · justifications display systematic variation associated with socioeconomic and political attributes
Political Orientation → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We study how persona prompting shapes language generated by multimodal large language models in an urban perception setting. Using 59,808 annotations from 1,200 persona-conditioned agents and two no-persona settings, we analyze captions, justifications, and perception tags across personas. Results indicate strong convergence in captions for different personas, whereas justifications display systematic variation associated with socioeconomic and political attributes, while perception tags show no statistically significant persona-related differences, though effect trends are observed. Topic analysis further reveals that personas emphasize different evaluative themes when interpreting the same scenes.
Tags
Links
- Source: https://arxiv.org/abs/2605.29064v1
- Canonical: https://arxiv.org/abs/2605.29064v1
Trouble viewing inline? Open PDF directly →
Full Text
48,568 characters extracted from source content.
Expand or collapse full text
Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception Neemias da Silva 1 , Myriam Delgado 1 , Rodrigo Minetto 1 , Daniel Silver 2 , Thiago H Silva 1,2 1 Universidade Tecnologica Federal do Parana, Curitiba, Brazil 2 University of Toronto, Toronto, Canada Correspondence: neemiasbuceli@alunos.utfpr.edu.br Abstract We study how persona prompting shapes lan- guage generated by multimodal large language models in an urban perception setting. Us- ing 59,808 annotations from 1,200 persona- conditioned agents and two no-persona settings, we analyze captions, justifications, and per- ception tags across personas. Results indicate strong convergence in captions for different personas, whereas justifications display system- atic variation associated with socioeconomic and political attributes, while perception tags show no statistically significant persona-related differences, though effect trends are observed. Topic analysis further reveals that personas em- phasize different evaluative themes when inter- preting the same scenes. 1 Introduction Large language models (LLMs) have increasingly been used as social simulators, where persona prompts are employed to approximate diverse per- spectives (Filippas et al., 2024; Aher et al., 2023; Argyle et al., 2023). This paradigm has been ex- plored for several tasks, such as studying social bias, raising growing interest in whether LLM- generated agents can model aspects of collective perception and behavior (Park et al., 2024; Li et al., 2025). Most prior systems evaluate persona prompt- ing using structured outputs like labels, choices, or task performance (Hu and Collier, 2024; Beck et al., 2024; da Silva et al., 2026). Much less is known about how personas shape their gener- ated language—particularly whether persona ef- fects emerge in descriptions, justifications, and perception tags, or whether generated language converges despite diverse agent attributes. This question is especially relevant in multimodal set- tings, where descriptive content (captions) and in- terpretive framing (justifications) may diverge. In urban scenes, perception is subjective and context- dependent (Lopes et al., 2023; da Silva et al., 2025). This suggests that models may agree on visual content while differing in interpretation. This distinction is important for interpreting persona- conditioned MLLMs and their use in social simula- tions. In this paper, we analyze a total of 59,808 an- notations describing and evaluating urban scenes. Of these, 59,708 were generated by 1,200 persona- conditioned MLLM agents, while the remaining 100 were produced under no-persona ablation set- ting. We conceptualize MLLM outputs as a func- tional distinction: (i) descriptive grounding, re- flected in captions that describe visual content, and (i) interpretive framing, reflected in justifications that evaluate and contextualize that content. Per- ception tags provide a structured semantic represen- tation of scene attributes, linking visual description to evaluative interpretation. Understanding this distinction allows us to ex- amine whether persona effects emerge primarily in visual descriptions or in evaluative interpretations. We test the hypothesis that persona conditioning mainly affects the interpretive stage while leaving descriptive grounding largely unchanged. The results reveal three main findings. First, cap- tions exhibit consistently high semantic similarity across personas, indicating strong convergence in descriptive grounding. Second, justifications show systematic variation aligned with socioeconomic and political attributes, revealing that persona ef- fects concentrate in interpretive framing rather than descriptive content. Third, topic-level analysis shows that personas emphasize different evaluative themes when interpreting the same scenes. These findings suggest that persona-conditioned MLLMs primarily affect interpretive framing rather than descriptive grounding. Our contributions are fourfold: (i) we introduce a functional framework distinguishing descriptive 1 arXiv:2605.29064v1 [cs.CL] 27 May 2026 grounding from interpretive framing in multimodal LLM outputs; (i) we provide empirical evidence that persona effects are substantially stronger in in- terpretive language than in descriptive content; (i) we show that interpretive framing effects are struc- tured primarily along economic status and political dimensions, with limited impact from gender and personality; (iv) we provide a publicly available GitHub repository 1 with all code and processed outputs to support reproducibility and enable fur- ther research. 2 Related Work Prior work shows that LLMs can reproduce styl- ized experimental findings and simulate decision- making processes (Aher et al., 2023; Filippas et al., 2024), supporting applications in surveys, experi- ments, and content analysis (Bail, 2024). However, important limitations remain, including bias, lim- ited validity, and difficulties in representing human diversity (Wang et al., 2024, 2025). A related line of research investigates LLMs as synthetic respondents and simulated populations. Studies show that LLMs can approximate human- like responses and support large-scale agent sim- ulations (Argyle et al., 2023; Park et al., 2024), while also revealing limitations in modeling belief structures and reasoning processes (Li et al., 2025; Barrie and Cerina, 2026). These findings suggest that evaluation should extend beyond predictive ac- curacy to include consistency and interpretability. Persona prompting plays a central role in this context. Conditioning models on demographic or psychological attributes can influence outputs, al- though effects are often limited or task-dependent (Hu and Collier, 2024; Beck et al., 2024). Most of this work evaluates persona effects on structured outputs — labels, ratings, or choices — leaving open whether and how personas shape the language of open-ended responses (Lutz et al., 2025; Malik et al., 2025). Our work also connects to research on multi- modal LLMs and urban perception. Prior studies show that urban image perception is subjective and socially contextualized, with perceptions of safety, disorder, beauty, and wealth varying systematically across observers and environments (He et al., 2026; Balsa-Barreiro et al., 2026; Lopes et al., 2023; Xu et al., 2025; Oliveira et al., 2020), and that mul- timodal LLMs can partially help capture some of 1 at: [it will be available after publication]. those aspects (Wu et al., 2026; da Silva et al., 2025). At the same time, recent work reports strong within- persona consistency but limited cross-persona vari- ation in outputs such as sentiment labels (da Silva et al., 2026), raising the question of whether per- sona effects emerge more clearly in interpretive language than in structured outputs. We address this gap directly, analyzing captions, justifications, and perception tags generated by multimodal LLMs to examine how persona effects emerge in descriptive grounding and interpretive framing. The key question is whether persona conditioning produces structured variation in how scenes are interpreted, even when descriptive con- tent converges. 3 Data and Methodology 3.1 Dataset and Conditions We analyze an annotation corpus publicly avail- able (da Silva et al., 2026), comprising 59,808 records generated by 1,200 multimodal LLM agents and two no-persona settings annotating50 urban scene images. Each record contains four out- puts: (i) a predicted sentiment label, (i) a caption describing the image, (i) one or more perception tags selected from a controlled vocabulary, and (iv) a justification explaining the assigned sentiment. Captions and justifications contain16.9± 5.0and 18.0± 4.9words on average, respectively. Persona- conditioned justifications are written as full prose explanations rather than simple label rationales, enabling linguistic analysis beyond sentiment la- bels. Figure 1 illustrates one annotation record per persona for the same urban image. All annotations have been generated using the Qwen3-VL:8B multimodal LLM (temperature= 0.1, seed= 42), following the standard configu- ration documented in the GitHub code ([it will be available after publication]). We consider three annotation conditions: •Persona. The main corpus consists of anno- tations produced by|P| = 24persona pro- files obtained from the Cartesian product of |D| = 4demographic and personality at- tributes: gender (Male, Female), economic status (Low, High), political orientation (Con- servative, Progressive), and personality (Prag- matic, Empathetic, Analytical). Each profile is instantiated asA = 50independent agents (24∗50 = 1,200agents total), and every agent 2 Persona: Female, Low Income, Conservative, PragmaticSentiment: SlightlyPositive Caption: “A tall brick building with a stone base, surrounded by greenery and a clear sky.” Perception tags: Urban environment, Cleanliness, Everyday environment, Old building, Well-lit environment Justification: “A clean, well-maintained city scene with orderly architecture, fitting a conservative view of practi- cal urban upkeep.” Persona: Female, Low Income, Progressive, EmpatheticSentiment: Positive Caption: “A university campus with tall brick buildings, a tree, and a clear sky with people near the entrance.” Perception tags: Pleasant environment, Everyday environment Justification: “This campus feels welcoming and well-maintained, aligning with values of accessible education.” Persona: Male, Low Income, Conservative, AnalyticalSentiment: Neutral Caption: “A tall brick building with a stone base, surrounded by trees and a clear sky.” Perception tags: Old building, Everyday environment, Cleanliness Justification: “Looks like a stable, traditional urban setting with maintained buildings and greenery.” Figure 1: Annotation example: 3 persona-conditioned agents annotate the same image, producing different perception tags, captions, and justifications. annotates all|I| = 50images using persona- conditioned prompts with extended reasoning enabled (think=True). Agents are indepen- dent in the sense that each annotation is gen- erated in a separate inference call under the same decoding configuration. The full list of profiles is shown along the rows of Figure 2 (Section 4.1). •No-Persona. A no-persona setting, in which the same model annotates the same images using a neutral observer prompt, includ- ing two variants: with extended reasoning (NPT,think=True) and without (NPNoT, think=False). The two variants produce one annotation per image (50each), provid- ing complementary baselines for the persona- generated corpus. The persona-conditioned setting exhibits a strongly bimodal sentiment label distribution (42.9% Negative and 34.5% Positive), with inter- mediate classes substantially underrepresented rel- ative to human annotations. In contrast, the two variants of no-persona setting skew toward Neg- ative and use the intermediate SlightlyNegative / SlightlyPositive classes more frequently than the persona-conditioned, producing a less polarized distribution. These differences provide additional context for the textual analyses that follow. Per- ception tags exhibit coherent semantic regularities, with frequent co-occurrences such as Abandoned and Lack of Maintenance, and Calm Environment and Peaceful Environment, indicating consistent structure in the generated perception tags. 3.2 Evaluating Within-Persona Stability Following our distinction between descriptive grounding and interpretive framing, we evaluate stability separately for captions, justifications, and perception tags. We first assess the internal sta- bility of agents sharing the same persona, i.e., the reproducibility of generated outputs when all demo- graphic and personality attributes are held constant. For each imagei ∈ Iand each persona profile p ∗ ∈ P, we compute pairwise similarities among all agents instantiated under that profilep ∗ and re- port the mean across pairs(p x , p y ). We evaluate stability across three output modalities: •Captions (descriptive grounding) and jus- tifications (interpretive framing): Semantic similarity is measured via cosine similarity on sentence embeddings. •Perception tags (intermediate semantic layer): Each annotation contains a set of per- ception tags drawn from a controlled vocabu- lary. Those tags form an intermediate seman- tic layer between descriptive grounding and interpretive framing, capturing structured as- pects of scene interpretation. We measure set overlap using Jaccard similarity (intersection over union). For each modality (captions, justifications, or perception tags), we compute a distribution of within-persona similarities and summarize it for each profile using the mean. The higher the values, the higher the indication that agents sharing the same persona converge on similar outputs. Topic structure is analyzed only for justifications, as cap- tions show high cross-agent similarity and lim- ited thematic discrimination. We initially explored topic modeling for both captions and justifications, but retained only justification topics due to the low semantic variability of captions. 3 3.3 Evaluating Cross-Persona Divergence We next evaluate whether different persona groups produce systematically distinct outputs. Rather than full persona combinations, we analyze diver- gence across persona dimensionsD(gender, eco- nomic status, political orientation, and personality). For each imagei ∈ Iand dimensiond ∈ D, we partition pairs of annotations into: •Within-group pairs: agents sharing the same level of the attribute (e.g., two high-income agents), regardless of other persona dimen- sions. • Cross-group pairs: agents with different lev- els of the attribute (e.g., high-income vs. low- income). We analyze persona effects for complementary modalities: • Captions (descriptive grounding) and jus- tifications (interpretive framing): We com- pute pairwise cosine similarity between em- beddings. •Perception tags (intermediate semantic layer): We compute set overlap using Jac- card similarity (intersection over union), as described in Section 3.2. •Topic distributions (interpretive framing): To assess whether personas influence the- matic structure, we compare topic distribu- tions across persona groups. For each image i ∈ Iand persona dimensiond ∈ D, we ag- gregate topic proportions within each group and compute divergence between groups us- ing Jensen–Shannon divergence. To visual- ize persona-specific framing, we construct a heatmap where rows correspond to the most frequent topics and columns correspond to per- sona groups, with values representing topic prevalence. For captions, justifications, and perception tags, we compare within-group and cross-group simi- larity distributions using Mann–WhitneyUtests (computed per image). Higher within-group sim- ilarity relative to cross-group similarity indicates that the attribute contributes a detectable persona- specific signal. For topic distributions, higher diver- gence between groups indicates that personas sys- tematically emphasize different semantic themes when interpreting the same scene. 4 Results 4.1Persona Effects from a Profile Perspective Figure 2 shows image-conditioned 24×24 inter- profile similarity matrices for captions, justifica- tions, and perception tags. For each profile pair (p x , p y ), similarity is averaged across shared im- ages to isolate persona-conditioned framing differ- ences. We begin by examining descriptive grounding, as captured by captions. Caption similarity spans a narrow range of0.85–0.90across all cells, includ- ing diagonal, with no discernible block structure. This indicates uniformly strong convergence in de- scriptive grounding: agents sharing the same pro- file produce only marginally more similar captions than agents drawn from different profiles, confirm- ing that descriptive content is largely invariant to persona conditioning. In contrast, justification similarity spans a wider range (0.44–0.70) and diagonal cells are consis- tently among the highest values in their respective rows and columns, confirming that agents within the same profile generate more similar justifications than cross-profile pairs. These results show that persona effects concentrate in interpretive framing rather than descriptive content. Perception tag Jaccard similarity is reported on a different scale and is not directly comparable to co- sine values; however, profiles that diverge most in justification also tend to diverge in perception-tag selection (Pearsonr = 0.67,p < 0.001, com- puted on all profile pairs). This suggests that dif- ferences in interpretive framing are partially re- flected in structured semantic selection. However, although profile-level divergence in justification correlates with divergence in perception-tag selec- tion, attribute-level tests do not reveal statistically significant separation (see Section 4.2). 4.2 Persona Effects from a Dimensional Perspective While Section 4.1 examined the full persona-profile structure, we now isolate the contribution of in- dividual persona dimensions across three output modalities: captions (descriptive grounding), jus- tifications (interpretive framing), and perception tags (intermediate semantic layer). Building on the convergence observed in Section 4.1, this analysis asks whether semantic similarity is systematically structured by persona attributes. Figure 3 compares within-group and cross-group 4 F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag 0.890.880.880.870.870.870.870.870.870.870.860.860.880.870.880.870.860.870.870.870.870.860.860.86 0.880.890.890.870.870.880.870.880.870.870.870.870.880.880.880.870.870.870.870.880.870.870.870.86 0.880.890.890.870.870.880.870.880.880.870.870.870.880.880.880.870.870.870.870.880.870.870.870.87 0.870.870.870.890.880.880.860.860.860.870.870.860.880.870.880.880.870.870.860.860.860.870.870.87 0.870.870.870.880.890.890.860.870.860.870.870.870.870.880.870.880.880.880.860.870.860.870.870.86 0.870.880.880.880.890.890.860.870.860.870.870.860.870.880.870.880.880.890.860.870.860.870.870.86 0.870.870.870.860.860.860.890.890.890.880.880.880.870.860.870.850.850.860.890.880.880.880.870.87 0.870.880.880.860.870.870.890.900.890.880.880.880.870.870.870.860.860.860.880.890.880.880.880.88 0.870.870.880.860.860.860.890.890.900.880.880.880.870.870.870.860.860.860.890.890.890.880.880.88 0.870.870.870.870.870.870.880.880.880.890.880.890.870.860.870.870.860.870.880.880.870.890.880.88 0.860.870.870.870.870.870.880.880.880.880.890.880.860.870.870.860.870.860.880.880.880.880.880.88 0.860.870.870.860.870.860.880.880.880.890.880.890.860.860.870.860.860.860.880.880.880.890.880.89 0.880.880.880.880.870.870.870.870.870.870.860.860.890.870.880.880.870.870.870.870.870.870.860.86 0.870.880.880.870.880.880.860.870.870.860.870.860.870.900.880.880.880.880.870.880.870.870.870.86 0.880.880.880.880.870.870.870.870.870.870.870.870.880.880.890.870.870.870.870.880.870.870.870.87 0.870.870.870.880.880.880.850.860.860.870.860.860.880.880.870.890.880.880.860.860.860.870.870.86 0.860.870.870.870.880.880.850.860.860.860.870.860.870.880.870.880.890.890.860.860.860.860.870.86 0.870.870.870.870.880.890.860.860.860.870.860.860.870.880.870.880.890.890.860.860.860.860.870.86 0.870.870.870.860.860.860.890.880.890.880.880.880.870.870.870.860.860.860.890.890.890.880.870.88 0.870.880.880.860.870.870.880.890.890.880.880.880.870.880.880.860.860.860.890.900.890.880.880.88 0.870.870.870.860.860.860.880.880.890.870.880.880.870.870.870.860.860.860.890.890.890.880.870.88 0.860.870.870.870.870.870.880.880.880.890.880.890.870.870.870.870.860.860.880.880.880.900.880.89 0.860.870.870.870.870.870.870.880.880.880.880.880.860.870.870.870.870.870.870.880.870.880.890.88 0.860.860.870.870.860.860.870.880.880.880.880.890.860.860.870.860.860.860.880.880.880.890.880.89 0.4 0.5 0.6 0.7 0.8 Mean cosine similarity (a) Captions F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag 0.660.610.620.600.570.600.570.520.490.530.520.520.650.610.610.600.570.600.560.520.480.530.520.51 0.610.670.590.610.620.610.560.570.490.540.560.530.590.650.570.610.620.610.540.570.470.540.560.52 0.620.590.640.560.540.580.580.530.510.520.510.520.620.590.620.560.550.580.570.530.500.520.520.52 0.600.610.560.700.660.680.540.540.470.580.570.550.600.620.560.690.670.670.530.530.450.580.580.54 0.570.620.540.660.700.650.520.550.460.560.590.540.560.630.530.650.690.640.500.550.440.560.590.53 0.600.610.580.680.650.690.550.540.490.570.570.560.600.620.580.670.650.680.540.540.470.580.570.55 0.570.560.580.540.520.550.620.570.560.550.540.560.570.560.580.530.520.550.610.570.550.550.550.56 0.520.570.530.540.550.540.570.640.550.550.590.570.520.560.520.530.550.540.550.620.520.550.580.56 0.490.490.510.470.460.490.560.550.600.490.510.530.490.490.510.470.470.490.560.540.580.500.510.53 0.530.540.520.580.560.570.550.550.490.690.640.630.530.540.520.570.570.560.550.540.480.680.650.62 0.520.560.510.570.590.570.540.590.510.640.670.630.510.560.510.570.590.560.540.580.490.640.670.62 0.520.530.520.550.540.560.560.570.530.630.630.650.510.530.520.540.550.550.560.570.520.630.630.64 0.650.590.620.600.560.600.570.520.490.530.510.510.660.610.620.600.570.600.560.520.480.530.520.51 0.610.650.590.620.630.620.560.560.490.540.560.530.610.660.580.620.630.620.540.560.470.540.560.53 0.610.570.620.560.530.580.580.520.510.520.510.520.620.580.640.560.540.580.570.530.500.520.520.52 0.600.610.560.690.650.670.530.530.470.570.570.540.600.620.560.690.660.670.520.530.450.570.570.54 0.570.620.550.670.690.650.520.550.470.570.590.550.570.630.540.660.700.650.510.550.440.570.590.54 0.600.610.580.670.640.680.550.540.490.560.560.550.600.620.580.670.650.690.540.540.470.570.570.55 0.560.540.570.530.500.540.610.550.560.550.540.560.560.540.570.520.510.540.620.560.560.550.540.56 0.520.570.530.530.550.540.570.620.540.540.580.570.520.560.530.530.550.540.560.620.520.550.580.56 0.480.470.500.450.440.470.550.520.580.480.490.520.480.470.500.450.440.470.560.520.590.480.490.52 0.530.540.520.580.560.580.550.550.500.680.640.630.530.540.520.570.570.570.550.550.480.680.650.62 0.520.560.520.580.590.570.550.580.510.650.670.630.520.560.520.570.590.570.540.580.490.650.680.62 0.510.520.520.540.530.550.560.560.530.620.620.640.510.530.520.540.540.550.560.560.520.620.620.64 0.4 0.5 0.6 0.7 0.8 Mean cosine similarity (b) Justifications F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag 0.490.440.470.450.420.450.430.420.430.410.410.420.480.440.480.430.420.440.440.420.430.420.400.41 0.440.480.450.450.460.450.410.430.410.410.430.410.440.470.450.430.450.440.420.430.420.420.420.41 0.470.450.490.440.420.450.430.420.430.410.410.410.460.450.480.430.420.440.440.420.430.410.400.41 0.450.450.440.500.470.470.410.420.410.420.420.420.450.450.450.480.460.460.420.420.410.430.420.41 0.420.460.420.470.500.470.400.420.400.420.440.420.420.460.430.460.480.460.410.430.410.420.430.41 0.450.450.450.470.470.500.410.430.420.420.430.420.440.460.460.460.480.490.430.420.420.430.430.42 0.430.410.430.410.400.410.470.440.470.440.420.450.430.420.430.400.400.410.470.440.470.440.420.44 0.420.430.420.420.420.430.440.490.460.440.450.450.420.430.430.410.430.420.450.460.460.440.450.44 0.430.410.430.410.400.420.470.460.500.450.430.460.430.420.430.390.410.410.480.450.490.440.430.45 0.410.410.410.420.420.420.440.440.450.480.450.470.410.410.420.410.420.420.450.430.450.470.450.46 0.410.430.410.420.440.430.420.450.430.450.480.460.410.420.410.410.430.420.440.440.430.450.460.44 0.420.410.410.420.420.420.450.450.460.470.460.490.420.410.420.410.410.420.460.450.470.470.450.47 0.480.440.460.450.420.440.430.420.430.410.410.420.490.440.480.440.420.440.440.420.440.420.410.42 0.440.470.450.450.460.460.420.430.420.410.420.410.440.500.460.440.460.460.430.430.430.420.420.41 0.480.450.480.450.430.460.430.430.430.420.410.420.480.460.520.440.430.460.450.430.450.430.410.42 0.430.430.430.480.460.460.400.410.390.410.410.410.440.440.440.480.460.460.410.410.400.420.420.40 0.420.450.420.460.480.480.400.430.410.420.430.410.420.460.430.460.510.470.420.420.410.420.430.41 0.440.440.440.460.460.490.410.420.410.420.420.420.440.460.460.460.470.500.430.420.420.430.430.42 0.440.420.440.420.410.430.470.450.480.450.440.460.440.430.450.410.420.430.510.450.490.460.430.46 0.420.430.420.420.430.420.440.460.450.430.440.450.420.430.430.410.420.420.450.470.450.440.440.44 0.430.420.430.410.410.420.470.460.490.450.430.470.440.430.450.400.410.420.490.450.510.460.430.46 0.420.420.410.430.420.430.440.440.440.470.450.470.420.420.430.420.420.430.460.440.460.490.450.46 0.400.420.400.420.430.430.420.450.430.450.460.450.410.420.410.420.430.430.430.440.430.450.480.44 0.410.410.410.410.410.420.440.440.450.460.440.470.420.410.420.400.410.420.460.440.460.460.440.48 0.4 0.5 0.6 0.7 0.8 Mean Jaccard similarity (c) Perception Tags Figure 2: Similarity between persona profiles. Cell(p x , p y )is the mean similarity (computed on the same image and averaged over all shared images) between profilep x and profilep y . Values on Figure c are on a different scale. cosine similarity across four persona dimensions, for captions (top row), serving as a descriptive base- line, justifications (middle row), capturing interpre- tive framing, and perception tags (bottom row), measured via Jaccard similarity. For captions, no persona dimension produces a statistically significant within/cross-group differ- ence (p > 0.05for all Mann–WhitneyUtests). This reinforces the result of Section 4.1: descriptive grounding is largely invariant to persona attributes, with agents describing the same visual content in highly similar ways regardless of gender, economic status, political orientation, or personality. In contrast, for justifications, economic status is the strongest effect among the justification-based comparisons: agents sharing the same income level generate more similar interpretive framings (∆ = +0.062 ,p < 0.001). Political orientation yields a comparable but smaller effect (∆ = +0.044, p < 0.001). Personality shows a positive but non- significant trend (∆ = +0.023,p = 0.052), and gender produces no detectable difference (∆ = +0.002, p = 0.702). Perception tags, by contrast, show no statistically significant within/cross-group difference for any persona dimension (economic status:∆ = +0.036, p = 0.070; political orientation:∆ = +0.017, p = 0.229; personality:∆ = +0.012,p = 0.275; gender:∆ = +0.003,p = 0.403). The absence of significant effects in perception tags should not be interpreted as evidence of equivalence: it may reflect the coarser granularity of perception tags (categorical labels) compared to free-text genera- tion. These results reveal a clear gradient of per- sona sensitivity across linguistic abstraction lev- els. Justifications are the primary locus of persona- conditioned variation, with effects concentrated along socioeconomic and political dimensions. Captions and perception tags, by contrast, show no significant attribute-level structure, indicating that neither descriptive content nor categorical label selection exhibits strong persona-dependent struc- ture in this setting. This pattern reinforces the cen- tral distinction of this paper: persona conditioning leaves a detectable imprint specifically in interpre- tive framing. 4.3 Persona Effects: a Topic Structure Analysis We further examine interpretive framing through topic structure, investigating whether personas sys- tematically emphasize different evaluative themes even when descriptive content converges. We ini- tially applied BERTopic to both captions and jus- tifications. However, captions exhibited high se- mantic uniformity and unstable topic structure, as caption embeddings were too similar for BERTopic to identify clearly differentiated topics. Therefore, subsequent topic analyses focus only on justifica- tions. Inspection of the most relevant topics (top 10) reveals a clear distinction between descriptive and interpretive language. Caption topics are domi- nated by descriptive motifs (e.g., indoor/workplace scenes, structural decay, traffic and accidents, urban skylines), whereas justification topics exhibit ex- plicitly evaluative structure (e.g., natural beauty, la- bor and development, accident and police response, rural quiet scenes, urban decay). The presence of evaluative concepts (neglect, beauty, progress, quiet) in justification topics—and their absence from caption topics—reinforces that persona con- ditioning operates primarily at the level of interpre- 5 Within Cross 0.6 0.7 0.8 0.9 Caption cosine sim. ns Gender Within Cross ns Economic status Within Cross ns Political orientation Within Cross ns Personality Within Cross 0.4 0.5 0.6 0.7 Justification cosine sim. ns Within Cross *** Within Cross *** Within Cross ns Within Cross 0.2 0.4 0.6 0.8 Perception Jaccard sim. ns Within Cross ns Within Cross ns Within Cross ns Figure 3: Within-group vs. cross-group cosine similarity for caption (top row), justification (middle), and perception (bottom) across persona attributes. Significance markers:∗ p < 0.001,ns= not significant (Mann–WhitneyU, per-image). tive framing rather than descriptive content. The full ranked topic list is provided in the supplemen- tary material. Figure 4 shows that justification topics align with sentiment in interpretable ways: accident/conflict topics skew negative, landscape/beauty topics skew positive, and labor/development topics exhibit mixed polarity, reflecting ambiguity in how agents interpret urban development scenes. 0%20%40%60%80%100% Proportion Natural landscape & beauty Work & development Rural & quiet road Streets & orderly urban Home & peaceful scene Typical urban scene City & nature Historical & preserved Accident & police response Destruction & conflict 100% 83% 70% 98% 100% 33% 27% 17% 30% 98% 66% 71% 99% 100% Sentiment Positive Neutral Negative Figure 4: Sentiment distribution across the top 10 justifi- cation topics generated by persona-conditioned agents. To assess whether topic prevalence varies sys- tematically across persona groups, Figure 5 re- ports the proportion of each justification topic within each persona, computed by aggregating justifications from agents in that group and nor- malizing within-group.Two patterns emerge. First, evaluative topics emphasizing neglect/decay or progress/development are unevenly distributed across economic status and political orientation: progressive and low-income personas place greater weight on themes related to inequality and ne- glect, while conservative and high-income personas emphasize development, order, and labor-related themes. Second, gender and personality groups exhibit more uniform topic distributions, consistent with the attribute-level results in Section 4.2. Some topics exhibit relatively stable prevalence across persona profiles, particularly ‘Historical & preserved’ and ‘City & nature’, suggesting that certain evaluative dimensions remain largely persona-invariant.In contrast, topics such as ‘Typical urban scene’ and ‘Natural landscape & beauty’ show stronger variation across socioeco- nomic and political dimensions. For example, low- income conservative personas allocate substantially greater weight to ‘Typical urban scene’, whereas progressive personas more frequently emphasize landscape- and beauty-related framing. These dif- 6 F/High/Cons/Anal F/High/Cons/Empa F/High/Cons/Prag F/High/Prog/Anal F/High/Prog/Empa F/High/Prog/Prag F/Low /Cons/Anal F/Low /Cons/Empa F/Low /Cons/Prag F/Low /Prog/Anal F/Low /Prog/Empa F/Low /Prog/Prag M/High/Cons/Anal M/High/Cons/Empa M/High/Cons/Prag M/High/Prog/Anal M/High/Prog/Empa M/High/Prog/Prag M/Low /Cons/Anal M/Low /Cons/Empa M/Low /Cons/Prag M/Low /Prog/Anal M/Low /Prog/Empa M/Low /Prog/Prag Persona Natural landscape & beauty Work & development Accident & police response Rural & quiet road Historical & preserved Destruction & conflict Streets & orderly urban City & nature Home & peaceful scene Typical urban scene Topic 0.200.220.210.220.240.220.170.160.110.180.230.190.180.190.190.200.260.210.160.170.090.190.220.18 0.140.140.130.140.140.140.130.150.140.150.150.150.140.140.130.150.140.140.130.140.150.150.140.16 0.120.120.130.120.120.110.140.150.150.140.150.150.130.120.130.110.100.100.140.140.140.130.130.14 0.140.130.130.080.130.100.140.140.140.120.110.120.140.140.130.080.090.110.140.140.140.130.120.13 0.070.070.070.070.070.070.070.070.070.070.070.080.070.070.070.080.070.070.070.070.070.070.070.07 0.070.070.070.070.070.070.070.070.070.080.080.080.070.070.070.080.070.070.070.070.060.070.070.07 0.070.070.070.070.060.070.090.080.050.080.070.040.080.070.080.080.070.070.080.080.070.060.070.06 0.060.060.070.070.060.070.070.060.070.070.060.060.070.070.070.080.070.070.060.070.070.080.070.07 0.060.070.050.070.060.070.070.070.070.070.070.060.060.060.070.080.070.070.060.060.070.070.060.06 0.060.050.090.060.050.070.060.050.140.030.020.080.080.060.070.070.040.080.070.050.140.040.040.06 0.1 0.2 Topic proportion Figure 5: Topic proportions (column-normalized) of the top justification topics across personas. Row labels show topic nicknames. ferences indicate that persona conditioning affects not only sentiment polarity, but also the thematic salience assigned to urban environments. These topic-level patterns provide a semantic explanation for the justification-level divergence observed in Figures 2 and 3. Together, these findings indicate that persona conditioning has a limited effect on what agents describe, but systematically shapes how they frame and evaluate the same scenes, with effects emerg- ing through differences in thematic emphasis and evaluative language. 4.4 Persona vs. No-Persona Effects To contextualize the role of persona conditioning, we compare two no-persona variants (NPT and NPNoT) against the persona-conditioned pool us- ing cosine similarity of Sentence-BERT embed- dings (all-MiniLM-L6-v2), computed per image. Figure 6 plots, for each imagei∈ I, two quan- tities: the x-axis shows the average cosine sim- ilarity among persona-conditioned outputs (per- sona–persona), while the y-axis shows the cosine similarity between the no-persona output and the mean representation of the persona-conditioned outputs. Each point, therefore, indicates how closely a no-persona response aligns with the persona pool relative to the internal agreement among persona- conditioned agents. This comparison allows us to assess whether persona conditioning produces systematic shifts in generated language beyond the model’s persona-persona variability. Note that this comparison is structurally asym- metric: the persona-conditioned pool comprises 50 agents per profile across 24 profiles (≈ 1,194anno- tations per image), whereas each no-persona setting produces a single annotation per image. The no- persona values should therefore be interpreted as indicative rather than statistically equivalent base- lines. For captions, both no-persona setting yield sim- ilarity levels close to the persona within-pool av- erage (NPT mean:0.865; NPNoT mean:0.862; persona within-pool mean:0.873). This is con- sistent with the strong convergence in descriptive grounding observed in Section 4.1, indicating that descriptions remain stable even without persona conditioning. In contrast, justification similarities are lower overall (NPT mean:0.550; NPNoT mean:0.539; persona within-pool mean:0.564), reflecting the greater semantic variability of interpretive framing relative to captions. While no-persona outputs re- main broadly aligned with the persona-conditioned pool, the reduced similarity is consistent with the broader pattern observed in Section 4.2, where per- sona conditioning introduces structured variation primarily in justificatory language. Beyond these mean differences, Figure 6 shows that this pattern holds consistently across images: caption similarities cluster tightly near the diago- nal, whereas justification similarities are both lower and more variable, reinforcing that persona effects primarily operate in interpretive language. These results reinforce that descriptive ground- ing remains largely invariant to persona condition- ing, while interpretive framing is both more vari- able and more sensitive to the presence of persona prompts. The no-persona case thus confirms that the observed differences in justifications reflect sys- tematic persona-conditioned effects, rather than stochastic variation in model outputs. 7 0.60.70.80.91.0 Persona pool mean sim. 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 NPT sim. to pool No-persona mean = 0.865 Persona pool mean = 0.873 Caption NPT 0.70.80.91.0 Persona pool mean sim. 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 NPNoT sim. to pool No-persona mean = 0.862 Persona pool mean = 0.873 Caption NPNoT 0.40.50.60.7 Persona pool mean sim. 0.40 0.45 0.50 0.55 0.60 0.65 0.70 0.75 NPT sim. to pool No-persona mean = 0.550 Persona pool mean = 0.564 Justification NPT 0.40.6 Persona pool mean sim. 0.3 0.4 0.5 0.6 0.7 NPNoT sim. to pool No-persona mean = 0.539 Persona pool mean = 0.564 Justification NPNoT Figure 6: Per-image cosine similarity of no-persona captions and justifications to the persona-conditioned pool. Each point corresponds to one image. 5 Conclusion We presented a linguistic analysis of how persona prompting shapes MLLM-generated language in an urban perception task, based on 59,808 annota- tions from 24 persona profiles and two no-persona setting agents. Results from the profile perspec- tive, attribute-level view, and topic-structure reveal a consistent pattern. Captions exhibit uniformly high semantic simi- larity across personas, with no systematic cluster- ing by persona, indicating strong convergence in descriptive grounding. In contrast, justifications show structured divergence primarily along eco- nomic status and political orientation, with no sta- tistically detectable effect of personality and no detectable effect of gender. Perception tags show weaker structure and do not produce statistically significant attribute-level differences under the cur- rent experimental conditions. Topic-level analyses reinforce this distinction: evaluative themes related to neglect, decay, development, and order align mainly with socioeconomic and political dimen- sions. These findings show that persona prompting does not substantially alter what agents describe, but systematically shapes how scenes are framed, evaluated, and interpreted. These results carry direct methodological impli- cations for the use of persona-conditioned MLLMs. For descriptive annotation tasks, such as scene cap- tioning or factual summarization, persona condi- tioning provides limited additional value, as both persona and no-persona agents converge on highly similar outputs. In contrast, evaluative outputs — including justifications and perception tags — exhibit structured interpretive variation, suggest- ing that persona prompting mainly changes how agents interpret and explain urban scenes, rather than changing what they perceive in the scenes themselves. More broadly, our findings suggest that persona- conditioned MLLMs may be better understood as generators of interpretive perspectives rather than faithful simulations of persona perception. Future work should investigate richer and more intersec- tional persona representations, compare agent out- puts against demographically matched human an- notations, and extend these analyses to additional models, languages, and cultural contexts where de- scriptive grounding and interpretive framing may differ. Limitations and Ethics Considerations This study has several limitations. First, persona conditioning relies on simplified attribute combina- tions (e.g., gender, economic status, political ori- entation, personality), which cannot fully capture the complexity and intersectionality of real human identities. As a result, observed differences across personas may reflect interactions between prompt design and model capabilities rather than genuine human social variation. Second, our analysis is based on a single multimodal model and a fixed set of urban images. Although the dataset is large in terms of generated annotations, it remains limited in visual diversity and in the range of evaluated multimodal models, limiting the extent to which the findings generalize across models and visual domains. Future work should consider additional models, datasets, and prompt formulations. Third, the comparison between persona and no-persona settings involves structural asymmetries (e.g., many persona-generated annotations versus single out- puts in baseline settings), which may influence per- formance comparisons. Fourth, the exclusive focus on urban scenes may limit the applicability of the findings to other visual domains, where persona ef- fects could differ under distinct semantic or cultural contexts. 8 From an ethical perspective, synthetic personas should not be interpreted as proxies for real de- mographic groups. Such representations risk over- simplifying or reinforcing stereotypes, particularly when used to simulate social behavior. More- over, while LLM-generated annotations can sup- port large-scale analysis, they may introduce sys- tematic biases and should not substitute human judgment without careful validation. Generative AI tools were used in the preparation of this paper to assist with language refinement and clarity of presentation. All methodological design, analysis, and interpretation were conducted by the authors, who take full responsibility for the content. Acknowledgments National Council for Scientific and Technological Development - CNPq (processes 314603/2023-9, 441444/2023-7, and 444724/2024-9) and INCT TILD-IAR (proc. 408490/2024-1). References Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Using large language models to simulate mul- tiple humans and replicate human subject studies. In Proc. of ICML, Honolulu, Hawaii, USA. JMLR.org. Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of one, many: Using language mod- els to simulate human samples. Political Analysis, 31(3):337–351. Christopher A. Bail. 2024. Can generative ai improve social science? PNAS, 121(21):e2314021121. Javier Balsa-Barreiro, Samin Rabbani, Djellel Eddine Difallah, and 1 others. 2026. A large-scale llm- generated dataset for exploring social interactions and urban experiences across 21 global cities. Dis- cover Data, 4(5). Christopher Barrie and Roberto Cerina. 2026. Synthetic personas distort the structure of human belief systems. OSF: osf.io/preprints/socarxiv/n7fq8_v1. Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2024. Sensitivity, performance, robust- ness: Deconstructing the effect of sociodemographic prompting. In Proc. of EACL, pages 2589–2615, St. Julian’s, Malta. Neemias B. da Silva, John Harrison, Rodrigo Minetto, Myriam R. Delgado, Bogdan T. Nassu, and Thiago H. Silva. 2025. Multimodal llms see sentiment. ArXiv: https://arxiv.org/abs/2508.16873. Neemias B da Silva, Rodrigo Minetto, Daniel Silver, and Thiago H Silva. 2026. Stable behavior, limited variation: Persona validity in llm agents for urban sentiment perception. In Proc. of IEEE DCOSS-IoT- UrbCom, Reykjavik, Iceland. Apostolos Filippas, John J. Horton, and Benjamin S. Manning. 2024. Large language models as simulated economic agents: What can we learn from homo silicus? In Proc. of EC, page 614–615, New Haven, CT, USA. ACM. Jun He, Yi Lin, Zilong Huang, Jiacong Yin, Junyan Ye, Yuchuan Zhou, Weijia Li, and Xiang Zhang. 2026. Urbanfeel: A comprehensive benchmark for temporal and perceptual understanding of city scenes through human perspective. In Proc. of ICRL, Rio de Janeiro, Brazil. Tiancheng Hu and Nigel Collier. 2024. Quantifying the persona effect in LLM simulations. In Proc. of ACL, pages 10289–10307, Bangkok, Thailand. ACL. Chance Jiajie Li, Jiayi Wu, Zhenze Mo, Ao Qu, Yuhan Tang, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Jinhua Zhao, and 1 others. 2025.Simulat- ing society requires simulating thought.ArXiv arXiv:2506.06958. Cesar Rafael Lopes, Rodrigo Minetto, Myriam Regat- tieri Delgado, and Thiago H Silva. 2023. Perceptsent - exploring subjectivity in a novel dataset for visual sentiment analysis. IEEE Transactions on Affective Computing, 14(3):1817–1831. Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. 2025. The prompt makes the person(a): A systematic evaluation of sociodemo- graphic persona prompting for large language models. In Proc. of EMNLP, pages 23212–23237, Suzhou, China. ACL. Ananya Malik, Nazanin Sabri, Melissa M. Karnaze, and Mai ElSherief. 2025. Are LLMs empathetic to all? investigating the influence of multi-demographic personas on a model’s empathy. In Proc. of EMNLP, pages 24938–24959, Suzhou, China. ACL. Wyverson Bonasoli de Oliveira, Leyza Baldo Dorini, Rodrigo Minetto, and Thiago H. Silva. 2020. Out- doorsent: Sentiment analysis of urban outdoor im- ages by using semantic and deep features. ACM Trans. Inf. Syst., 38(3). Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Ben- jamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109. Angelina Wang, Jamie Morgenstern, and John P Dick- erson. 2025. Large language models that replace hu- man participants can harmfully misportray and flatten identity groups. Nat Mach Intell, 7(3):400–411. 9 Pengda Wang, Huiqi Zou, Zihan Yan, Feng Guo, Tian- jun Sun, Ziang Xiao, and Bo Zhang. 2024. Not yet: Large language models cannot replace hu- man respondents for psychometric research. OSF: osf.io/preprints/osf/rwy9b_v1. Songtai Wu, Wenbing Wang, Chengzhi Zhang, Qisheng Zeng, Haiying Wang, Jinyao Lin, and Shaoying Li. 2026. Exploring multimodal large language models’ potential in simulating human perception of urban cycling environments: A street-view perspective. In- formation Geography, 2(1):100037. Yunzhe Xu, Yiyuan Pan, Zhe Liu, and Hesheng Wang. 2025. Flame: Learning to navigate with multimodal llm in urban environments. In Proc. of AAAI, vol- ume 39, page 9005–9013, Philadelphia, USA. 10