Paper deep dive
Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/18/2026, 12:12:33 PM
Summary
The paper introduces Anamnesis, an open-source, interactive web platform for large-scale survey simulation using Large Language Models (LLMs). It operationalizes the Anthology and Alterity frameworks by using structured narrative backstories to create diverse, demographically controllable virtual personas. The system supports multimodal inputs, probabilistic demographic resampling, and sequential context accumulation to ensure response consistency. Case studies demonstrate that Anamnesis replicates real-world survey data from the Pew Research Center's American Trends Panel and New Yorker Caption Contest more accurately than standard persona-prompting baselines.
Entities (11)
Relation Signals (7)
Anamnesis â implements â Anthology
confidence 95% ¡ Anamnesis operationalizes the Anthology and Alterity methodologies within a unified, interactive interface.
Anthology â uses â Narrative Backstories
confidence 95% ¡ Anthology methodology... utilizes rich, open-ended narrative backstories to condition model responses
Anthology â createdby â Moon et al.
confidence 90% ¡ Anthology (Moon et al., 2024) advances this line by conditioning on rich, LLM-generated narratives
Anamnesis â emulates â New Yorker Caption Contest
confidence 90% ¡ (2) emulating human preference in the New Yorker Caption Contest.
Anamnesis â implements â Alterity
confidence 90% ¡ Anamnesis operationalizes the Anthology and Alterity methodologies within a unified, interactive interface.
Anamnesis â replicates â American Trends Panel
confidence 90% ¡ We evaluate the system through two case studies: (1) replicating segments of Pew Research Center's American Trends Panel (ATP)
Anamnesis â outperforms â Synthetic Users
confidence 80% ¡ Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines... offering a transparent, reproducible, and open-source alternative to proprietary simulation services.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source, and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey instruments on virtual populations rather than real human subjects. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface. It supports open-ended generation, probabilistic demographic resampling, and multimodal (image and audio) surveys. We evaluate the system through two case studies: (1) replicating segments of Pew Research Center's American Trends Panel (ATP) on political typology and biomedical issues and (2) emulating human preference in the New Yorker Caption Contest. In both cases, Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines, offering a transparent, reproducible, and open-source alternative to proprietary simulation services.
Tags
Links
- Source: https://arxiv.org/abs/2607.10628v1
- Canonical: https://arxiv.org/abs/2607.10628v1
Trouble viewing inline? Open PDF directly â
Full Text
34,885 characters extracted from source content.
Expand or collapse full text
Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan University of California, Berkeley vaclis,josephsuh,serinac,davidchan@berkeley.edu Abstract We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source, and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey instruments on virtual populations rather than real human subjects. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface. It supports open-ended generation, probabilistic demographic resampling, and multimodal (image and audio) surveys. We evaluate the system through two case studies: (1) replicating segments of Pew Research Centerâs American Trends Panel (ATP) on political typology and biomedical issues and (2) emulating human preference in the New Yorker Caption Contest. In both cases, Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines, offering a transparent, reproducible, and open-source alternative to proprietary simulation services. Demo Video:https://w.youtube.com/ watch?v=j5yrnJl287g Platform site: https://simulate.group 1 Introduction Opinion surveys and social polling are foundational tools for understanding human behavior, public pol- icy, and societal trends. However, traditional human- subject research faces mounting challenges, including rising costs, declining response rates, and the logisti- cal difficulty of reaching specific demographic sub- populations. The emergence of Large Language Mod- els (LLMs) as âvirtual personasâ offers an alternative, promising the ability to prototype survey instruments and stress-test social hypotheses at a fraction of the time and cost of traditional methods. For these simu- lated surveys to be scientifically valid, however, mod- els must move beyond âaverageâ aggregate responses LLM Persona 1 Persona 2 Persona 3 Persona 4 ... Human Study Virtual Opinions Figure 1: Anamnesis is an interactive system for demo- graphically controllable survey simulation using large lan- guage models. It provides a non-technical interface for An- thology, a method which approximates large-scale human studies by conditioning LLMs to representative, consistent, and diverse virtual personas. Together, these systems enable rapid prototyping and stress-testing of survey instruments on diverse virtual populations using multimodal stimuli. and instead demonstrate the ability to faithfully simu- late the nuanced, idiosyncratic perspectives of diverse individuals (Kang et al., 2025; Moon et al., 2024). Previous efforts to simulate human populations have primarily relied on âpersona prompting,â where a model is given a short list of demographic attributes. While functional for basic tasks, this approach often yields stereotypical responses and lacks the psycho- logical depth required for complex opinion elicitation (Cheng et al., 2023). This limitation has been ad- dressed by the Anthology methodology (Moon et al., 2024) which utilizes rich, open-ended narrative back- stories to condition model responses, and the Alterity framework (Kang et al., 2025), which explores âdeep bindingâ to ensure LLMs simulate authentic in-group perspectives rather than out-group misperceptions (Wang et al., 2025). Despite these academic advances, the methodologies remain largely confined to siloed Python scripts. Meanwhile, commercial platforms such as Synthetic Users, Expected Parrot, and Artifi- cial Societies (Synthetic Users, 2026; Expected Parrot, 2026; Artificial Societies, 2026) offer similar 1 arXiv:2607.10628v1 [cs.CL] 12 Jul 2026 simulation capabilities but operate as closed-source, proprietary platforms that lack the transparency and reproducibility required for rigorous social science. In this paper, we present Anamnesis, an open- source, web-based platform designed to democratize access to high-fidelity persona simulation for non-technical users. Anamnesis operationalizes the Anthology and Alterity methodologies within a unified, interactive interface. Unlike previous implementations of these methods, Anamnesis is a platform which provides a non-technical survey builder, supports multi-modal inputs (image and audio), and is backed by a range of LLM inference providers. Together, these contributions make state-of-the-art research in persona approximation openly available to a wider range of users. We evaluate the Anamnesis system through case studies in political opinion elicitation and multimodal preference estimation. Specifically, we replicate segments of the Pew Research Centerâs American Trends Panel (ATP) (PewResearch, 2025), demon- strating that the platform can elicit opinions that align with real-world human response distributions more accurately than standard prompting baselines (Santurkar et al., 2023; Kim and Yang, 2025). We also use the platformâs multi-modal capabilities to simulate human vision-language preference in the New Yorker Caption Contest. Our results show that Anamnesis closely mirrors human sentiment across both language-only and vision-language problems, and can be a valuable tool for researchers prototyping and stress-testing human-study survey instruments. 2 Anthology: Narrative-based Virtual Persona Anamnesis is built upon the Anthology framework (Moon et al., 2024), which introduces a methodology for simulating diverse human respondents using LLM-conditioned virtual personas. Rather than relying on short demographic prompts (e.g., "Re- spond as if you are a 35-year-old Hispanic woman") (Santurkar et al., 2023), Anthology conditions language models on a rich, open-ended narrative backstory-a multi-paragraph life history that captures not just demographic attributes but also formative experiences, values, and worldview. Backstories are generated via sampling multi-turn life narratives from pretrained base language models (Kang et al., 2025). Specifically, the language model is conditioned on interview questions of the American Voices Project (Stanford Center on Poverty and Inequality, 2021) to complete realistic and diverse open-ended life narratives. Sampled backstories are labeled by their demographic information which is obtained by querying a multiple-choice demographic question to a language model conditioned on the backstory. In Anamnesis, we construct a database of pre-sampled backstories indexed by their demograph- ics so that practitioners interested in a subpopulation behavior can easily run a targeted simulation. To ensure demographic representativeness and a targeted simulation, Anthology pairs backstory gener- ation with a population-matching step. A practitioner often has a target distribution over demographic dimensions (e.g., age, race, political affiliation) they aim to simulate: to this end, Anthology samples from the entire pool of backstories so that the demographic distribution of the sampled pool matches a target demographic distribution. Anamnesis operationalizes these methodologies into an end-to-end platform. With a built-in database of indexed backstories, it offers automated backstory generation, demographic balancing based on the target demographics, and response collection with an arbi- trary question set to ask a language model conditioned on backstories, all within a single interactive interface. 3 Anamnesis: Accessible, Open-Source, Anthology Implementation 3.1 System Overview Anamnesis translates the Anthology methodology from research prototypes into a deployable open- source survey simulation platform. Rather than requiring researchers to manually generate backsto- ries or write sampling scripts, the platform enables demographic-constrained simulation over a large pool of pre-generated personas through an interactive inter- face. A typical user workflow consists of four stages: 1.Survey Construction: Users define multi- question survey instruments through a graphical builder. The system supports multiple-choice, multi-select, open-ended, ranking, and multimodal (image and audio) questions. 2.Demographic Targeting: Users specify target audience demographics and sample size, selecting from a pool of 35K pre-sampled backstories with probabilistic demographic distributions (§ 3.3). If a desired demographic dimension is not available, users may create new dimensions through an integrated demographic inference procedure(§ 3.4). 2 Figure 2: System overview of Anamnesis. Pre-sampled backstories generated via Anthology are stored as a persona pool, each associated with probabilistic demographic distributions. Users construct surveys and specify demographic constraints through an abstraction layer. Survey runs are executed via a dispatcherâqueueâworker architecture with sequential context accumulation per persona, enabling scalable and reproducible simulation. Results are aggregated post hoc. 3.Simulation Execution: Users select a language model and answering algorithm(§ 3.5). Each backstory completes the survey sequentially, with responses accumulated to maintain consistency (§ 3.2). 4.Result Analysis: Responses are automatically aggregated and visualized. Users may further filter results by demographic attributes post hoc for comparative analysis. 3.2 Execution Architecture As a publicly accessible platform, Anamnesis must support concurrent survey runs over a large and grow- ing persona pool. A survey evaluates each selected persona across all questions, resulting inO(SĂ Q) LLM calls per run, whereSis the sample size (num- ber of virtual personas) andQthe number of questions. For demographic surveys using repeated sampling (N- sample mode; § 3.4), this yields O(SĂQĂN) calls. This execution regime requires (1) per-persona state preservation across questions, (2) bounded concurrency under API/vLLM rate limits, and (3) reproducible run-level configuration. Anamnesis addresses these constraints through a dispatcherâqueueâworker architecture (Figure 2). Each survey run is snapshotted at launch time, record- ing its demographic filters, answering algorithm, model configuration, and concurrency bounds. Tasks are decomposed into personaâquestion units and published to a message queue; workers consume tasks asynchronously while executing questions sequentially per persona with incremental context accumulation. 3.3 Backstory Selection Anamnesis enables researchers to simulate surveys over their specified target populations (e.g., âwomen aged 18â24â or âvoters aged 25â44 with a college de- gree, evenly split between Democrat and Republicanâ) without manual preprocessing. In prior Anthology experiments, each backstory was paired with an actual human respondent from a completed real-world survey (e.g., American Trends Panel). Demographic attributes were directly observed. Balancing therefore reduced to deterministic assignment: given known labels and target quotas, one could apply greedy se- lection or Hungarian matching to choose respondents whose attributes exactly satisfied requested cells. In Anamnesis, this assumption no longer holds. Demographics are not observed labels but inferred 3 probability distributions stored per dimension. For each backstory b and dimension d, the system stores: p b,d (c), câC d , wherep b,d (c)denotes the inferred probability that bbelongs to categoryc. Demographic selection must therefore operate under uncertainty. To accommodate different user scenarios, Anamnesis provides two selection algorithms: Top-K Probability Ranking. Designed for sce- narios where researchers prioritize selecting personas that most strongly match the target demographic con- straints, effectively treating the filtered demographic set as a single group without internal balancing. Given a sample sizeSand demographic filters, each backstory is scored by its joint compatibility with the filter (multiplying probabilities across dimensions and summing over selected categories when applicable). Backstories are ranked by this score and the topS are selected. Balanced Demographic Matching. Designed for studies where representation across demographic subgroups must be explicitly enforced (e.g., equal allocation across ageĂgender cells). First, selected categories are expanded into their cross-product demographic cellsG. The total sample sizeSis divided into slotsK g for each cellgâG(uniformly or via user-specified weights). Each slot represents a required demographic target. Because our pre-sampled backstories include probabilistic demographics, selection becomes an assignment problem: choose backstories such that (i) each slot is filled, (i) each backstory is selected at most once, and (i) overall demographic compatibility is maximized. Algorithm 1 summarizes the procedure. To maintain interactive latency, the balanced match- ing procedure restricts the candidate space before solving the assignment problem. Without pruning, Hungarian matching over the full persona pool would incurO(S 3 )time complexity with a score matrix of size SĂ|P|, which is impractical for real-time use. We therefore retain only the top-Mcandidates per demographic cell (defaultM =50), and take the union of these candidates to form a shared candidate pool. Hungarian assignment is then applied over the resulting SĂ|C| matrix, where|C|âŞ|P|. In practice, this reduces matching complexity to O(S 3 )withS⤠50, ensuring responsive client-side computation without impacting backend execution or worker throughput. Algorithm 1 Balanced demographic matching Require: Persona poolP; filtersF; sample size S 1: G âcross-product of selected demographic categories 2: Allocate slotsK g for eachg â Gsuch that P g K g =S 3: Expand slots into target listT (|T|=S) 4: for all gâG do 5: Retain top-Mbackstories by one-hot score for g 6: end for 7:Build score matrix between targetsTand candidate backstories 8:Apply Hungarian assignment to maximize total compatibility 9: return matched backstories 3.4 Extending the Demographic Space Anthology already introduced demographic surveys over backstories, and our persona pool includes pre-populated demographic dimensions derived from that pipeline. Anamnesis extends this capability to a user-driven platform feature. Researchers may require attributes not originally annotated (e.g., marital status, political leaning, occupation). Instead of offline scripts, users define a new categorical dimension through the interface, and the system conducts a demographic survey over the persona pool, estimating for each backstory a probability distribution over categories. Distribution Modes. While prior research code relies on token log-probabilities from a self-hosted vLLM backend, many researchers do not operate such infrastructure. We therefore support two interchangeable modes: â˘Logprobs mode. When available (e.g., vLLM), a single constrained forward pass yields the full categorical distribution. ⢠N-sample mode. When logprobs are unavail- able, the system repeats the questionNtimes and estimates the empirical distribution from sampled responses. Both modes produce the same probabilistic abstrac- tion. Although smallNin N-sample mode may intro- duce sampling variance, the estimate converges asN increases. While this approach incurs higher inference cost than logprobs mode, it closes the practical gap for researchers without access to self-hosted vLLM. 4 3.5 Survey Answering Algorithms Beyond backstory selection, Anamnesis allows users to choose the answering algorithm used during inference, enabling controlled comparisons between simulation strategies. Anthology (default). Each backstory is prepended to the first survey question. After the model responds, the questionâanswer pair is appended to the context before the next question is posed, implementing sequential context accumulation. This mechanism en- courages the virtual persona to condition on its prior re- sponses, promoting cross-question belief consistency. Zero-shot baselines. Users may alternatively select a baseline mode that conditions only on a short demographic description (e.g., CLAIR-style prompts (Chan et al., 2023)) rather than the full narrative backstory. Running both modes side-by-side enables direct quantification of the contribution of backstory conditioning, serving as an ablation control. Responses are parsed through a two-tier pipeline: structured output (guided decoding on vLLM; JSON schema on OpenRouter) is attempted first; if unsuccessful, a lightweight parser LLM extracts the final answer from the raw response. 3.6 Post-Hoc Demographic Filtering For exploratory studies, researchers may execute a survey over the full persona pool without specifying demographic constraints upfront, and subsequently segment results by demographic attributes after the fact. The results dashboard supports interactive filtering and re-aggregation by any demographic dimension stored in the backstory metadata, without requiring re-execution of the survey run. 4 Case Studies We anticipate that practitioners will find diverse applications for Anamnesis, tailoring simulations to their specific needs. In the following section, we highlight two illustrative use cases and encourage the community to discover further possibilities. 4.1 Simulating Public Opinion Polls To validate that the Anamnesis platform replicates the Anthology method, we replicate the core experiment of Moon et al. (2024): approximating survey response distributions from the Pew Research Centerâs American Trends Panel (ATP). We consider three ATP waves covering distinct topics: Wave 34 (biomedical and food issues), Wave 92 (political typology), and Wave 99 (AI and human enhancement) (see Appendix C for details). Survey questions are multiple-choice items asked to all respondents and preserve the original wording and answer options. Using the Anamnesis survey builder, we construct each ATP wave as a multi-question session. Surveys are executed with sequential context accumulation (§3.2) over backstory pools matched to the survey respondentsâ demographics. We evaluate using the same metrics as Moon et al. (2024): average Wasserstein distance (WD) measuring representative- ness of the response distribution and the Frobenius norm between response correlation matrices (Fro.) measuring response consistency. Table 1 summarizes the results. Consistent with the original findings, backstory-conditioned simulation on the Anamnesis platform outperforms demographic list-based baselines across three waves. Reproducing these three experiments, spanning 20 survey questions and thousands of virtual respondents, the pipeline was configured and executed through the platformâs graphical interface. This highlights the primary utility of Anamnesis: a social scientist can draft a survey instrument and stress-test it against a demographically balanced virtual population before recruiting a single human participant. The platformâs interactive result viewer further supports post hoc filtering by demo- graphic subgroup, enabling targeted analysis (e.g., ex- amining whether response distributions diverge across age or race groups) without re-running the simulation. 4.2 Multimodal Alignment Prior experiment only focused on text-based surveys. To verify that backstory-based simulation remains meaningful under multimodal inputs, we evaluated alignment against real human preference data. We therefore conduct a case study on the New Yorker Caption Contest benchmark (Hessel et al., 2022; Jain et al., 2020), a multimodal task in which cartoon images are paired with caption candidates and ground-truth labels are derived from large-scale crowd voting. The dataset provides a simple but controlled test of whether persona-conditioned virtual populations exhibit measurable correlation with collective human judgments. Method. We evaluated 49 contests with randomized caption order. For each, Gemini 2.5 Flash (temper- ature 1.0) makes 20 choices under two answering algorithms: (i) Anthology (backstory-conditioned simulation) (i) Zero-shot demographic baseline. We report majority-vote accuracy with Wilson intervals 5 Table 1: Simulating American Trends Panel public opinion polls, based on the Anamnesis platform and two demographic- list prompting method BIO and QA (Santurkar et al., 2023). Please refer to Moon et al. (2024) for the details of method choices, including persona matching, and the definition of metrics (Wasserstein distance (WD) and Frobenius Norm (Fro.)). Model PersonaPersonaATP Wave 34ATP Wave 92ATP Wave 99 ConditioningMatchingWD (â) Fro. (â)WD (â) Fro. (â)WD (â) Fro. (â) LLaMA-3.1-8B BIOn/a0.2581.5560.3462.0780.2771.229 QAn/a0.2351.4810.3921.7190.1801.475 max weight0.1600.8370.2511.6030.1481.026 Anamnesis greedy0.1470.9640.2181.4140.1391.352 Human0.0570.4180.0910.4110.0810.327 and an exact McNemar test. To retain within-item information, we also compare the vote share assigned to the human winner using a paired bootstrap interval and exact sign-flip test. Results. Anthology achieves a majority-vote accuracy of 59.2% (95% CI: 45.2â71.8%), compared with 51.0% (95% CI: 37.5â64.4%) for the Zero-shot baseline. Narrative conditioning also increases the mean vote share assigned to the human-preferred caption from 52.0% to 59.8%, a paired improve- ment of 7.8 percentage points (95% CI: 3.2â12.8; p=0.0024). These findings suggest that Anthology shifts model preferences toward the human-preferred caption overall. 5 Related Work LLM Persona Conditioning. A growing body of work explores conditioning LLMs to simulate human perspectives. Early approaches supply language models with short demographic attribute lists, e.g., question-answer pairs about demographic indicators, and measure alignment with human survey responses (Santurkar et al., 2023; Hwang et al., 2023; Li et al., 2025). While effective as baselines, these methods tend to produce stereotypical or flattened outputs that fail to capture within-group variation (Cheng et al., 2023; Wang et al., 2025). Anthology (Moon et al., 2024) advances this line by conditioning on rich, LLM-generated narratives rather than attribute lists, demonstrating improved consistency on survey benchmarks; Alterity (Kang et al., 2025) further demonstrates the efficacy via reproducing in-group, out-group and meta-perception study results. Comparison to Existing Survey Platforms. A growing ecosystem of commercial and open-source platforms offers LLM-based survey simulation. Synthetic Users (Synthetic Users, 2026) generates AI personas for interviews and surveys using a multi- agent framework with optional retrieval-augmented generation to incorporate proprietary data. However, personas are defined by short attribute profiles rather than rich narratives, and the methodology is entirely closed-source. Artificial Societies (Artificial Societies, 2026) focuses on simulating a network-level social dynamics, such as content virality and collective decision-making, rather than structured opinion surveys, and likewise does not publish its conditioning methodology to simulate virtual personas. On the open-source side, Expected Parrotâs EDSL (Expected Parrot, 2026) provides a Python domain- specific language for constructing AI agents with trait dictionaries and administering surveys across multiple LLMs, but its persona conditioning reduces to the short-attribute prompting baseline that Anthology was designed to supersede; moreover, as a code library, it remains inaccessible to researchers without program- ming experience, as it requires explicit code-based specification of scenarios, user/agent models, tools, policies, and evaluation hooks. Park et al. (2024) demonstrate that two-hour qualitative interviews with real individuals can produce generative agents that replicate survey responses with high fidelity, though this approach requires costly human data collection that limits scalability. Anamnesis is, to our knowledge, the first open-source, GUI-based platform that combines narrative backstory conditioning, probabilistic demographic matching, sequential context accumulation, and multimodal survey support in a single deployable system accessible to researchers without programming expertise. 6 References Artificial Societies Artificial Societies. 2026. Artificial soci- eties â company profile.https://w.ycombinator. com/companies/artificial-societies . Accessed: 2026-02-26. David M Chan, Suzanne Petryk, Joseph E Gonzalez, Trevor Darrell, and John Canny. 2023. Clair: Evaluating image captions with large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, Singapore. Association for Computational Linguistics. Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023. Compost: Characterizing and evaluating caricature in llm simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10853â10875. Expected Parrot Expected Parrot. 2026. Expected parrot. https://w.expectedparrot.com/.Accessed: 2026-02-26. Jack Hessel, Ana Marasovi Ě c, Jena D Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi. 2022. Do androids laugh at electric sheep? humor "understanding" benchmarks from the new yorker caption contest. arXiv preprint arXiv:2209.06293. EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023. Aligning language models to user opin- ions. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5906â5919. Lalit Jain, Kevin Jamieson, Robert Mankoff, Robert Nowak, and Scott Sievert. 2020. The New Yorker cartoon caption contest dataset. Minwoo Kang, Suhong Moon, Seung Hyeong Lee, Ayush Raj, Joseph Suh, and David Chan. 2025. Deep binding of language model virtual personas: a study on approximating political partisan misperceptions. In Second Conference on Language Modeling. Jaehyung Kim and Yiming Yang. 2025. Few-shot personalization of llms with mis-aligned responses. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Compu- tational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 11943â11974. Ang Li, Haozhe Chen, Hongseok Namkoong, and Tianyi Peng. 2025. Llm generated persona is a promise with a catch. arXiv preprint arXiv:2503.16527. Suhong Moon, Marwa Abdulhai, Minwoo Kang, Joseph Suh, Widyadewi Soedarmadji, Eran Kohen Behar, and David M Chan. 2024. Virtual personas for language models via an anthology of backstories. In Proceedings of the 2024 conference on empirical methods in natural language processing, pages 19864â19897. Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Ben- jamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109. PewResearch. 2025.America trends panel waves.Retrieved February 06, 2025, from https://w.pewsocialtrends.org/dataset. Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In International conference on machine learning, pages 29971â30004. PMLR. Stanford Center on Poverty and Inequality. 2021. American voices project methodology. Accessed: 2025-03-23. Synthetic Users Synthetic Users. 2026. Synthetic users â company profile.https://w.syntheticusers. com/. Accessed: 2026-02-26. Angelina Wang, Jamie Morgenstern, and John P Dickerson. 2025. Large language models that replace human par- ticipants can harmfully misportray and flatten identity groups. Nature Machine Intelligence, 7(3):400â411. 7 Appendix The appendix is organized as follows: ⢠Appendix A discusses the limitations of our method. â˘Appendix B discusses some additional details of the New Yorker Caption Contest. ⢠Appendix C discusses some additional details of the American Trends Panel. A Limitations While Anamnesis provides a robust platform for persona-based survey simulation, several limitations inherent to the methodology and the underlying technology must be acknowledged. First, the quality of any simulation is fundamentally bounded by the diversity of the backstory pool. As identified in the development of the Anthology framework (Moon et al., 2024), LLM-generated backstories can exhibit skewed demographic distributions that reflect the inherent biases of their training data rather than a true census-representative population. This leads to the risk of âshallow binding,â where the model reflects an out-groupâs stereotypical perception of a demographic rather than the groupâs actual internal logic. Although Anamnesis implements methodologies from the Alterity framework (Kang et al., 2025) to deepen this binding through multi-turn interview transcripts, re- searchers should remain critical of results on sensitive social topics where models may still default to car- icatured personas. Moreover, because all backstories and simulations are conducted in English, linguistic and cultural variation is necessarily compressed into English-language reasoning patterns, potentially limit- ing the cross-cultural validity of represented personas. Additionally, while the platform enables multi- modal conditioning, current multi-modal LLMs (MLLMs) may lack the perceptual nuance of human subjects, potentially ignoring subtle visual or auditory cues. Finally, virtual personas are temporally static; they do not evolve in response to real-world current events unless their backstories are explicitly updated, which limits the platformâs utility for longitudinal tracking of rapidly shifting public opinion. B New Yorker Caption Contest The New Yorker Caption Contest Benchmarks dataset (Hessel et al., 2022) is a large-scale multimodal benchmark designed to evaluate computational âhumor understandingâ using cartoons from The New Figure B.1: Image-based caption ranking example. Given the cartoon image above, the model must select the funnier caption among the candidates: (A) âIt comes with sub-par schools but a world-class trauma center.â (B) âIf we time it right, I can get you in this house today.â (GT: B) Yorker Caption Contest. We evaluate our method on the âQuality Rankingâ task, which requires methods to choose the funnier caption between alternatives. Each instance includes the original cartoon image, two captions, and gold labels for which caption won the contest. The dataset supports image-based and text- based settings; we use the image-based version for our experiments. An example is given in Figure B.1. C American Trends Panel The American Trends Panel (ATP), administered by the Pew Research Center (PewResearch, 2025), is a nationally representative survey panel comprising U.S. adults. The panel covers a broad range of subjects, from politics and religion to internet use and online dating, among others. Our analysis draws on selected questions from three survey waves, focusing on items that were posed to all human participants. Notably, some questions in the original ATP surveys use Likert-scale response options whose ordering (e.g., ranging from positive to negative, or vice versa) was randomized across respondents. To mirror this design, we similarly randomize the sequence of these options when constructing prompts for LLMs. ATP Wave 34 is conducted from April 23, 2018 to May 6, 2018 with a focus on biomedical and food issues. The number of total respondents is 2,537. An example is provided in Figure C.1. ATP Wave 92 is conducted from July 8, 2021 to July 21, 2021 with a focus on political typology and 10,916 respondents. American Trends Panel Wave 99 is conducted from November 1, 2021 to November 7, 2021 with a focus on artificial intelligence and human enhancement. The number of total respondents is 10,260. 8 American Trends Panel Wave 34 Selected Questions Please answer the following question keeping in mind your previous answers. Question: How likely is it that genetically modified foods will lead to more affordably-priced food (A) Not at all likely (B) Not too likely (C) Fairly likely (D) Very likely Answer with (A), (B), (C), or (D). Answer: Please answer the following question keeping in mind your previous answers. Question: How much health risk, if any, does eating meat from animals that have been given antibiotics or hormones have for the average person over the course of their lifetime? (A) No health risk at all (B) Not too much health risk (C) Some health risk (D) A great deal of health risk Answer with (A), (B), (C), or (D). Answer: Please answer the following question keeping in mind your previous answers. Question: How likely is it that genetically modified foods will create problems for the environment (A) Not at all likely (B) Not too likely (C) Fairly likely (D) Very likely Answer with (A), (B), (C), or (D). Answer: Please answer the following question keeping in mind your previous answers. Question: How likely is it that genetically modified foods will lead to health problems for the population as a whole (A) Not at all likely (B) Not too likely (C) Fairly likely (D) Very likely Answer with (A), (B), (C), or (D). Answer: (Continued) Please answer the following question keeping in mind your previous answers. Question: How much of the food you eat is organic? (A) None at all (B) Not too much (C) Some of it (D) Most of it Answer with (A), (B), (C), or (D). Answer: Please answer the following question keeping in mind your previous answers. Question: How much health risk, if any, does eating food and drinks with artificial preservatives have for the average person over the course of their lifetime? (A) No health risk at all (B) Not too much health risk (C) Some health risk (D) A great deal of health risk Answer with (A), (B), (C), or (D). Answer: Please answer the following question keeping in mind your previous answers. Question: How much health risk, if any, does eating food and drinks with artificial coloring have for the average person over the course of their lifetime? (A) No health risk at all (B) Not too much health risk (C) Some health risk (D) A great deal of health risk Answer with (A), (B), (C), or (D). Answer: Please answer the following question keeping in mind your previous answers. Question: How much do you, personally, care about the issue of genetically modified foods? (A) Not at all (B) Not too much (C) Some (D) A great deal Answer with (A), (B), (C), or (D). Answer: Figure C.1: 8 questions sampled from ATP Wave 34. The prompts âPlease answer the following question keeping in mind your previous answersâ are included before asking each survey question, which are found to enhance the consistency of response from Moon et al. (2024). 9