Paper deep dive
Hybrid Panels: Toward Human-AI Collaboration in Survey Research
Julia Romberg, Tobias Gummer, Gabriella Lapesa, Tanja Kunz, Claudia Wagner
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporates both human participants and LLMs as fundamental elements of its design. In this research note, we introduce the concept of a hybrid panel by providing a definition and outlining an overarching framework, spanning data collection to data validation. We detail results from a first pilot study to illustrate (open) challenges that we identify for hybrid panels.
Tags
Links
- Source: https://arxiv.org/abs/2608.22582v1
- Canonical: https://arxiv.org/abs/2608.22582v1
Trouble viewing inline? Open PDF directly â
Full Text
52,004 characters extracted from source content.
Expand or collapse full text
Hybrid Panels: Toward HumanâAI Collaboration in Survey Research Julia Romberg, Tobias Gummer, Gabriella Lapesa, Tanja Kunz, Affiliation: GESIS - Leibniz Institute for the Social Sciences Affiliation: GESIS - Leibniz Institute for the Social Sciences Affiliation: GESIS - Leibniz Institute for the Social Sciences Affiliation: GESIS - Leibniz Institute for the Social Sciences Affiliation: Heidelberg University Affiliation: Heinrich Heine University of DĂŒsseldorf and Claudia Wagner Affiliation: GESIS - Leibniz Institute for the Social Sciences Affiliation: RWTH Aachen University Contact: firstname.lastname@gesis.org Abstract Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporates both human participants and LLMs as fundamental elements of its design. In this research note, we introduce the concept of a hybrid panel by providing a definition and outlining an overarching framework, spanning data collection to data validation. We detail results from a first pilot study to illustrate (open) challenges that we identify for hybrid panels. 1 Introduction Population surveys are a longstanding and imperative means to measure opinions, attitudes and values, with application in various fields, from measuring attitudes, values, and behavior for social science research over public opinion polling and to market research for product and service development. Despite their importance, survey research faces several persistent challenges, including declining response rates (3; 33; 7; 19), increasing data collection costs (20; 34), long delays between data collection and data provision, and the risk of nonresponse bias, particularly among important population subgroups and hard-to-reach populations (24; 27). Recent advancements in artificial intelligence (AI) and, in particular, large language models (LLMs) have sparked optimism regarding their potential to address challenges in survey research by promoting efficiency, accessibility, and flexibility (21). The use cases in which AI can support survey research are manifold and span the entire survey lifecycle (22), ranging from questionnaire design and survey item generation (15) to AI-assisted interviewing (35; 36) and the automated coding of open-ended responses (30). Yet, realizing the potential of AI in survey research also requires considering the infrastructures within which these applications would be embedded. A fundamental requirement for conducting traditional population surveys are data collection infrastructures, which provide the organizational and methodological backbone for collecting high-quality data and making them available for scientific use (e.g., to enable analyses of the general population over time). Over the past decades, such infrastructures have been established for major longitudinal national and international survey programs, including repeated cross-national studies (e.g., the US General Social Survey, the European Social Survey, the World Values Study) and panel studies (e.g., the Panel Study of Income Dynamics, the German Socio-Economic Panel, and Understanding Society). Beyond enabling the systematic collection of data over time, these infrastructures provide researchers with access to already recruited, probability-based samples and established methodological standards, protocols, and quality control procedures. Commercial providers such as Qualtrics, Xpolls, and PersonaPanels have already begun to build up structures to offer synthetic panels generated from AI models. These solutionsâalthough promising faster and less costly data collectionâface severe challenges related to transparency, accountability, and ethical compliance. These concerns are especially salient in social science research because open science, reproducibility, and the ability to independently evaluate data quality are essential requirements for scientific integrity. What is more, recent research has also called into question the overall readiness of the technology for full synthesis (9; 29; 31). In this research note, we introduce the concept of a hybrid panel for social science research by providing a definition and outlining an overarching framework to facilitate the systematic development and testing of AI-enabled survey infrastructures. We define a hybrid panel as a longitudinal survey infrastructure in which responses are provided by both human participants and AI models, for example through planned missingness designs where missing data is (in part) imputed via LLMs, the augmentation of human responses with AI-generated input, or human validation of AI-generated responses. The longitudinal design reflects the need for ongoing maintenance and human participation in developing suitable AI methods for imputation, ideally with continuous validation of AI-generated responses against data collected from the same human participants over time. The hybrid panel bridges the gap between traditional panels based on human samples and fully synthetic panels, acknowledging both the current deficiencies of synthetic approaches such as the mismatch with human behavior (9; 29; 31) and the promise of AI for survey research (2). We hereby extend on previous research that integrates responses from human participants and AI within theoretical frameworks, including prediction-powered inference for valid statistical analysis when human observations are supplemented by AI-generated responses (1; 18), mixed subject designs combining observations from both sources (4), and adaptive resource allocation via informed sampling methods that are frequently used in the area of active machine learning to reduce model errors (26). To illustrate our framework, we introduce our ongoing research endeavor of building up a hybrid panel for social science research. We describe our design of the hybrid panel for social science research that we started piloting, discuss the methodological and practical challenges that we encountered and present initial findings of our pilot study which covers the human participant recruitment step. Before turning to the remainder of the paper, note that the authors of this research note come from different disciplinary backgrounds (Survey Research, Computer Science, Computational Social Science, Computational Linguistics) and have non-overlapping positions on the very important debate on AI in survey research. What we share, however, is the firm conviction that AI cannot fully replace humans as the object of study of survey research, and at the same time that the extent of the applicability of a technology to a research domain is a matter of evaluation and experimentation and it requires an ambitious and rigorous research agenda with a long-term time plan. This is precisely the contribution that we want to make with the hybrid panel outlined in this research note, which can only be developed further through the feedback we hope to receive from a community that is, at the time of writing, extremely split and for good reasons. 2 Implementing a Hybrid Panel for Social Science Research 2.1 Definition Survey responses obtained exclusively from human participants and those generated exclusively by AI systems (the so-called in-silico panels) represent the two ends of a continuum. Our proposed hybrid panel falls within this continuum by integrating both human participants and LLMs into a common longitudinal survey infrastructure in which they collaborate, e.g., by supplementing or validating each otherâs input. More specifically, we define a hybrid panel as a group of human participants and AI models that repeatedly contribute to surveys or annotation tasks over time. Rather than treating AI as a replacement for human participants, the hybrid panel enables various forms of interplay between humans and AI. For example, human responses or human annotations can be used to train, calibrate, and validate AI models that then may support users in filling out a survey or completing an annotation task (e.g., by suggesting entries for items that users skip). Another option would be to test planned-missingness designs (e.g., split-questionnaire designs) where the questionnaire is divided into subsets of questions, and different respondents receive different subsets. The planned missing items can be answered by AI models and the assignment can be rotated across subsequent waves. Over time, respondents may eventually answer all modules, which allows one to validate AI-generated answers while each individual wave remains shorter. Such hybrid designs may allow novel adaptive survey designs, re-distribution of data collection budget, e.g., to invest into the recruitment of hard-to-reach populations, and advancing the state-of-the-art in both LLM alignment and its intersection to survey research. The survey research community may also benefit from practical evaluations on a more fundamental level, as it will allow to sharpen and adapt crucial notions such as the distinction between imputation and simulation, and an update of the notion of data quality. It is to be noted that our definition deliberately extends beyond traditional survey tasks to include data annotation tasks, the process of labeling data snippets according to predefined concepts: providing LLMs with access not only to opinions, attitudes or values on a more conceptual level but also with concrete evaluations of data snippets that provide textual manifestations of these high-level concepts. For example, a survey item on attitudes toward populism can be complemented with several short texts that respondents evaluate as populist or non-populist. This design moves beyond abstract self-reports by capturing how individuals apply the concept in concrete contexts. The resulting data provide a richer and more behaviorally grounded representation of respondentsâ perspectives, which is particularly well suited for LLM-based modeling due to its reliance on semantic judgments and language-based reasoning. From the survey participant perspective, integrating annotation in traditional surveys may also end up to be more engaging than ticking scales or producing open-ended answers. Annotation within our hybrid survey infrastructure can be implemented as a form of gamification which may increase the appeal of surveys for socio-demographic groups who perceive traditional survey formats as unengaging. 2.2 Overall Framework Figure 1: Phases of the hybrid panel. Figure 1 illustrates the different phases in the implementation of our hybrid panel. The first phase concerns the recruitment of human participants where the challenges summarized in the representation arm of the Total Survey Error Framework apply (16). Theories on survey participation decisions (8; 17) of potential respondents suggest that cost-benefit considerations play a key role in these decisions, taking into account the general context and the design of a survey. In the case of hybrid panels, we argue that it is important to consider the conditions under which participants are willing to collaborate with and share their data for AI model training. Factors such as public interest, trust, data privacy, security, and transparency are expected to positively affect their willingness to participate (32). The impact of further design optionsâsuch as the degree of autonomy of AI, the sensitivity of data shared by participants for AI learning, the permissible sensitivity of data that the AI outputs, and the amount of financial compensationâremains uncertain in the novel context of hybrid panels and is investigated in our pilot study (see Section 2.3). Although ethical considerations related to the potential replacement of human participants by LLMs and the resulting reduction or loss of compensation are relevant, we do not advocate for it nor do we anticipate this transition occurring in the near future. The second phase concerns the LLM(s) selection and documentation. To ensure transparency, accountability, and ethical compliance, preference should be given to open and transparent models whose training data, model architecture, and weights are publicly available for inspection. We will use existing documentation procedures such as the recently released LLM-checklist (12) to document the LLM selection and usage processes. Furthermore, the selected LLMs should be calibrated for the specific survey task. This is an open challenge, since so far no specialized model exists for survey research and general-purpose LLMs reveal many forms of undesirable response biases and inconsistencies. For example, LLM responses are sensitive to changes in the question wording and show social desirability biases (29). To what extent these biases can be reduced by fine-tuning LLMs on a large amount of survey data is an open empirical question (see, e.g., 6). The third phase focuses on designing and implementing what we call LLM adaptive data collection design. The goal is to determine which information should be imputed and which should be collected from which respondents to maximize estimator accuracy keeping the data collection costs within reasonable limits.11 1 Although the LLM adaptive survey design aims to optimize resource allocation, its goal is not to remove human input, but to reposition it where it adds the most value. In traditional panels, adaptive survey design varies survey procedures between groups of respondents during and after data collection based on information gathered about respondents, response patterns, costs, or sample composition (25). For hybrid panels, methods developed for machine learning settings (e.g., uncertainty sampling, expected error reduction) (26) âwhere large amounts of unlabeled data coexist with a smaller amount of labeled ground-truth dataâmay provide a promising complement. During the field phase, human and LLM data are collected. Based on the LLM adaptive data collection design, we decide for which humans we collect empirical data and for whom data is simulated. Empirical and synthetic data are used to create survey estimators for the variables. Statistical methods such as design-based supervised learning (11), prediction-powered inference (1) or confidence-driven inference (13) will be used to rectify the biased LLM estimators. Human-in-the-loop designs can be used to assess the validity of individual-level predictions, as evaluation by those whose perspectives are modeled offers the most effective and ethically grounded way to assess AI simulations. In practice, an LLM responds to selected survey items or question groups, and these responses (or a strategically sampled subset of responses) are presented to human participants for assessment rather than requiring them to complete the items themselves. This allows researchers to identify systematic discrepancies between human and LLM responses, quantify prediction uncertainty, and iteratively improve the alignment of the LLM models. At the same time, it gives participants an active role in evaluating how their views are represented, increasing transparency, accountability, and trust in AI-assisted survey methodologies. After the field phase, the selected LLMs will be aligned to the data that was collected during this iteration. Fine-tuning (i.e., further training a pre-trained model on a specific dataset and task) or in-context learning (i.e., supplying the model with examples or instructions directly in the input) are possible options. As discussed before, the longitudinal nature is a key characteristic of a hybrid panel. Having survey waves over time enables us to validate the LLM component of the hybrid panel in between waves. In the last phase, we assess the validity of LLM predictions at the distribution level by comparing them against human-informed survey estimates based on data that were unavailable during model training and evaluation. By relying on unpublished survey data22 2 GESIS is involved in several panel initiatives in which a considerable amount of time typically elapses between data collection and publication., we minimize the risk of data leakage33 3 Data leakage occurs when information that should not be available to the model accidentally influences its development (e.g., the survey to be used as test dataset). This can make the model appear more accurate than it truly is, resulting in misleading performance estimates. and ensure that the resulting performance estimates reflect genuine predictive capability. We will use the validation step to provide detailed performance reports of the LLM component of the panel (e.g., which models and which parameter settings worked well, for which subgroups do we get the lowest performance, which questions cannot be well predicted) and infer design decisions about the next panel waves from these reports. 2.3 Pilot Study We conducted a pilot study with the crowdworking platform Prolific44 4 https://w.prolific.com/ as a widely adopted and cost-efficient environment. In the first phase of the pilot, we address the step of human participant recruitment and validating the sample itself, before, in later phases, we will begin implementing the subsequent steps of the hybrid panel. The invitation to participate to the survey was sent to 6,032 eligible participants (German residents who are fluent in German). Data collection took place from June 1 to June 7, 2026. To ensure a diverse sample, we collected 300 respectively 30155 5 We opened 300 places for each of the four time points. A submission error led to the collection of 301 responses in one of the batches. responses at different time points: twice during weekdays (morning and evening), on a public holiday (noon), and on the weekend (morning). Participants were incentivized for participation with a German minimum wage (13.90 EUR per hour) and gave informed consent. A total of 1,260 participants started the survey, of whom 1,201 completed it, 45 exited at an early stage and 14 were filtered out by the platform due to exceeding the maximum time allowed (35 minutes) without completing the task. The average completion time among respondents with valid responses was 11 minutes and 33 seconds. Our main research questions are: âą RQ1. Do respondents already use AI in their work on Prolific? âą RQ2. Would respondents be willing to participate in a hybrid panel? âą RQ3. Which respondents would take part in a hybrid panel? RQ1 investigates whether respondents use ChatGPT and inter alia to answer survey items. Understanding such behavior is crucial for a controlled hybrid panel, otherwise it remains shallow what true human-written and what AI-generated responses are. To assess this non-optimal answering behavior, we implemented a self-reported survey question on whether respondents use AI support in their work on Prolific66 6 While we use the term âAIâ to simplify understanding for a general audience, the listed answer options refer to applications of LLMs.. Overall, we found that only 6% of participants reported using AI for response generation or multiple-choice selection. Most respondents, 73%, reported no AI usage, while the remainder reported to use AI for minor language revisions. An additional spot check of the writing style in free-text response supports the low numbers of self-reported AI-use for our study. A recent Prolific study provides additional evidence of low AI prevalence on crowdworking platforms (14). These findings suggest that our sample can be used to implement the further steps of our proposed hybrid panel pilot. RQ2 assesses the respondentsâ general willingness to participate in a traditional or a hybrid panel (see Figure 2): Whereas 83% of the respondents reported to likely participate in a traditional panel (somewhat likely: 41%, very likely: 42%), 69% reported that they likely would participate in a hybrid panel (somewhat likely: 43%, very likely: 26%). Although these results are positive overall, they indicate that participation rates are lower for a hybrid panel than for a traditional panel and that design decisions are needed to improve them. (a) Participation in a human panel. (b) Participation in a hybrid panel. Figure 2: Distribution of responses regarding the likelihood of participating in a regular human panel and a hybrid panel, as asked sequentially in the survey. (a) Low data sensitivity, low AI autonomy, and high financial compensation. (b) High data sensitivity, high AI autonomy, and low financial compensation. Figure 3: Distribution of responses regarding the likelihood of participating in two opposing vignette design options for the hybrid panel. Option (a) describes the most conservative scenario, option (b) the most progressive case. To better understand how specific design choices affect participation in the hybrid panel, we included a vignette experiment that varied the conditions of data sensitivity (low vs. high), AI autonomy (low vs. high), and financial compensation (low vs. high). Each participant was presented with two randomly assigned vignettes. Figure 3 illustrates the results of the two extremes. The combination of low data sensitivity, low AI autonomy, and high compensation was associated with a substantial increase in the likelihood of stated participation, while the opposite scenario of high data sensitivity, high AI autonomy, and low compensation was perceived more negatively, though a notable share of respondents still reported their willingness to participate. These findings indicate that supporting respondent information need to be developed to convey a high degree of transparency about AI-support in the hybrid panel. RQ3 then explores potential nonresponse biases when recruiting a hybrid panel. In order to approximate nonresponse bias, we compared those who reported that they likely would participate in a hybrid panel to those who reported to be unlikely with respect to their attitudes toward AI (sum factor of ATTARI-12 (28)), political leaning (11-point left-right scale), and socio-demographic variables (gender (0=male, 1=female), education (0=low, 1=high)). We found that respondents who were likely to participate in a hybrid panel had a more positive attitude toward AI (M=3.63) than those who were unlikely to participate (M=3.13, t=10.15, p<<0.001). We further found differences in their political leaning, with likely respondents leaning more towards the right (M=4.86 vs. M=4.26, t=4.4, p<<0.001). With respect to gender, we found a higher share of males among the likely respondents (Chi2=26.55, p<<0.001). Concerning their education, our data show no differences between likely and unlikely respondents (Chi2=1.37, p>>0.05). These findings suggest that nonresponse bias may be a concern when recruiting human participants for a hybrid panel. 3 Discussion Exploring the potential of AI for survey research is only at its beginning, and it needs a long-term research agenda that takes into account the different components: the human subjects with their preferences and rights (which we focus on in the pilot study), the AI models and their continuous evaluation (whose features we only briefly touch upon in this research note), and the research community on whose reaction we call with this research note (because the standards for data quality and for robust generalization can only come from a community effort where diverse perspectives and disciplines are represented). Leveraging AI models for survey research raises expectations of cost reduction and increased speed of data collection. More efficient data collection enables researchers to respond more swiftly to societal events and to adapt surveys in real time. Nonetheless, accurately capturing genuine human opinions remains particularly difficult, and the current body of research demonstrates insufficient quality of replicating self-reported human behavior fully synthetically (e.g., 9; 29; 31; 23; 10; 5). Discrepancies from the ground truth are particularly consequential in applications where the survey outcomes inform the public or policy makers. These limitations suggest that the benefits of AI may be best realized by selectively combining human responses with LLM estimators. Hybrid panels offer a practical framework for such selective integration. They enable selective delegation: where an LLM can reliably simulate human responses for a particular subpopulation or type of question content, a greater share of responses can be delegated to the model, while more complex content or less predictable cases remain in the hands of the humans about whom researchers aim to make generalizations. Our vision for a hybrid panel allows for in-process validation of AI performance through direct human-in-the-loop involvement. The simulation of opinions, attitudes, or values intrinsic to human beingsâphenomena for which ground truth is typically unavailableâis best evaluated by those whose perspectives are being modeled. This design strengthens validation, improves simulation quality through iterative feedback, and promotes transparency and accountability, thereby increasing trust in AI-assisted survey methodologies. By incorporating humans directly into the process, we can also allow them to have more control over what data the AI is permitted to learn from and what it is allowed to impute. Our framework enables a systematic perspective on the integration of human and AI in surveys tasks and offers a foundational concept upon which future AI-supported survey infrastructures can be created, classified, and evaluated. The effort required from respondents will decrease as LLMs become more aligned with human respondents. Lower respondent burden might improve participation rates, measurement error, and retention of respondents in panel settings. Validating a proposed response usually requires less effort than creating original responses from scratch, especially for complex survey items, such as free-text responses. Initially, human validation should encompass all AI-generated responses to ensure accuracy and model reliability. Hard-to-reach populations and systematic nonresponse when recruiting human participants for hybrid panels will remain an open challenge: obtaining at least few data and participants from underrepresented groups is indispensable for model alignment and validation, and we hope that hybrid panels will allow survey effort and budgets to be at least partially invested for this. If this is successful, hybrid panels can augment available human responses of small or hard-to-reach subgroups. Ethics Statement It is not our belief that the measurement of opinions, attitudes, and values inherent to human individuals can or should be fully automated by computational models. Significant concerns arise when synthesized results only are used to answer research questions about society or guide external stakeholders, such as policymakers or the general public, due not only to the severe limitations of current models affecting accuracy (such as hallucinated answers and uncontrolled or hidden bias) but also to the unclear effects on society (such as diminished public acceptance and declining trust in research). Our approach presents a solution to overcoming the persistent challenges in survey research (namely declining response rates, increasing data collection costs, and the risk of nonresponse bias) by leveraging the recent advancements of AI while preserving the underlying subject of observation in the process: the human individual.77 7 Synthesized survey responses may still be useful in laboratory pre-test scenarios, in which limitations can be counted in. The ongoing development of the hybrid panel is planned to be monitored by an ethics board and, broadly speaking, is to be seen as a community-driven effort for which this research note and the development of the proposed hybrid panel for Germany is the trigger and also an exchange platform through calls for shared-task-like evaluation (on the AI development side) and empirical reflection on what the desiderata of the survey community are (at which point would survey researchers consider the output of an LLM robust enough to draw generalizations on?). An important aspect to conclude this ethics section is the consent of human participants. We are aware that at present, despite the clear definitions in the Prolific survey, it may not be 100% clear to the participants that what they are consenting to is the fact that AI may partially take over their answer (therefore reducing their compensation). The extent of this understanding is to be further tested for participants who are in principle willing to take part in the hybrid panel. There are however additional panel design considerations that will make this case very unlikely. Should a participant belong to a group for whom the LLMs produce accurate enough predictions for a specific attribute, this participant will still be recruited again at least in the following cases: a) for the smaller validation of these attributes; b) for other attributes and constructs; c) for purely annotation use cases that will still be supported by the infrastructure. It will be a recruitment consideration to make sure that allocation of requests remains fair, and that âmore predictableâ subjects are not penalized as data workers. References Angelopoulos et al. (2023) A. N. Angelopoulos, S. Bates, C. Fannjiang, M. I. Jordan, and T. Zrnic Prediction-powered inference. Science 382 (6671), p. 669â674. External Links: Document Cited by: §1, §2.2. Argyle et al. (2023) L. P. Argyle, E. C. Busby, N. Fulda, J. R. Gubler, C. Rytting, and D. Wingate Out of one, many: using language models to simulate human samples. Political Analysis 31 (3), p. 337â351. External Links: Document Cited by: §1. Brick and Williams (2013) J. M. Brick and D. Williams Explaining rising nonresponse rates in cross-sectional surveys. The ANNALS of the American Academy of Political and Social Science 645 (1), p. 36â59. External Links: Document Cited by: §1. Broska et al. (2025) D. Broska, M. Howes, and A. van Loon The mixed subjects design: treating large language models as potentially informative observations. Sociological Methods & Research 54 (3), p. 1074â1109. External Links: Document Cited by: §1. BultĂ© and Terryn (2026) B. BultĂ© and A. R. Terryn LLMs and cultural values: the impact of prompt language and explicit cultural framing. Computational Linguistics 52, p. 407â494. External Links: ISSN 0891-2017, Document Cited by: §3. Ceron et al. (2026) T. Ceron, D. Nikolaev, D. Stammbach, and D. Nozza What is the political content in llmsâ pre- and post-training data?. External Links: 2509.22367, Link Cited by: §2.2. de Leeuw et al. (2018) E. de Leeuw, J. Hox, and A. Luiten International nonresponse trends across countries and years: an analysis of 36 years of labour force survey data. Survey Methods: Insights from the Field, p. 1â10. External Links: Document Cited by: §1. Dillman et al. (2014) D. A. Dillman, J. D. Smyth, and L. M. Christian Internet, phone, mail, and mixed-mode surveys: the tailored design method. 4 edition, John Wiley & Sons. Cited by: §2.2. Dominguez-Olmedo et al. (2024) R. Dominguez-Olmedo, M. Hardt, and C. Mendler-DĂŒnner Questioning the survey responses of large language models. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37, p. 45850â45878. External Links: Document Cited by: §1, §1, §3. Durmus et al. (2024) E. Durmus, K. Nguyen, T. I. Liao, N. Schiefer, A. Askell, A. Bakhtin, C. Chen, Z. Hatfield-Dodds, D. Hernandez, N. Joseph, L. Lovitt, S. McCandlish, O. Sikder, A. Tamkin, J. Thamkul, J. Kaplan, J. Clark, and D. Ganguli Towards measuring the representation of subjective global opinions in language models. External Links: 2306.16388, Link Cited by: §3. Egami et al. (2023) N. Egami, M. Hinck, B. Stewart, and H. Wei Using imperfect surrogates for downstream inference: design-based supervised learning for social science applications of large language models. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 68589â68601. External Links: Document Cited by: §2.2. Feuerriegel et al. (2026) S. Feuerriegel, C. Barrie, M. Crockett, L. K. Globig, K. L. McLoughlin, D. Mirea, A. Spirling, D. Yang, T. Althoff, M. Antoniak, L. P. Argyle, A. Ashokkumar, M. Atari, H. Bailey, K. Bauer, U. Bhatt, Y. Chai, T. Chakraborty, Y. Chandra, H. Chen, H. DaumĂ©, G. D. F. Morales, M. Dehghani, D. Dillion, J. C. Eichstaedt, K. Forster, D. Geissler, K. Gray, T. L. Griffiths, H. Jochen, O. Hauser, J. K. He, R. Hemrajani, F. Holzmeister, A. H. Hwang, T. Hu, A. A. Ivanova, N. Köbis, Y. Kyrychenko, H. Lakkaraju, J. Liu, A. Maarouf, S. Maier, L. Meincke, R. Mihalcea, B. Mittelstadt, S. M. Mohammad, M. Naaman, O. Netzer, A. Oh, D. C. Ong, F. Pierri, B. Plank, I. Rahwan, T. Rahwan, P. S. Rao, C. E. Robertson, D. M. Rothschild, M. J. Salganik, E. Schulz, C. Shah, Y. R. Shrestha, E. Shutova, A. A. Siegel, A. Simchon, H. Sun, M. Toetzke, J. J. V. Bavel, M. Vaccaro, J. W. Vaughan, E. Vayena, P. O. Vaz-de-Melo, B. Vecchione, A. Wang, R. West, R. Willer, D. U. Wulff, R. Zhang, S. Zhang, S. Rathje, and M. H. Ribeiro A reporting checklist for large language models in behavioural science. Nature Human Behaviour 10, p. 1182â1186. External Links: Document Cited by: §2.2. GligoriÄ et al. (2025) K. GligoriÄ, T. Zrnic, C. Lee, E. CandĂšs, and D. Jurafsky Can unconfident LLM annotations be used for confident conclusions?. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 3514â3533. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2.2. Gordon et al. (2026) A. Gordon, D. Rothschild, F. M. Affonso, J. Sulik, D. Hauser, K. Pepin, and S. Jones AI agent prevalence and data quality across multiple online sample providers. PsyArXiv. External Links: Document Cited by: §2.3. Götz et al. (2024) F. M. Götz, R. Maertens, S. Loomba, and S. Van Der Linden Let the algorithm speak: how to use neural networks for automatic item generation in psychological scale development.. Psychological Methods 29 (3), p. 494â518. External Links: Document Cited by: §1. Groves et al. (2004) R. M. Groves, F. J. Fowler Jr, M. P. Couper, J. M. Lepkowski, E. Singer, and R. Tourangeau Survey methodology. John Wiley & Sons. Cited by: §2.2. Groves et al. (2000) R. M. Groves, E. Singer, and A. Corning Leverage-saliency theory of survey participation: description and an illustration. The Public Opinion Quarterly 64 (3), p. 299â308. External Links: Link Cited by: §2.2. Krsteski et al. (2025) S. Krsteski, G. Russo, S. Chang, R. West, and K. GligoriÄ Valid survey simulations with limited human data: the roles of prompting, fine-tuning, and rectification. External Links: 2510.11408, Link Cited by: §1. Luiten et al. (2020) A. Luiten, J. Hox, and E. de Leeuw Survey nonresponse trends and fieldwork effort in the 21st century: results of an international study across countries and surveys. Journal of Official Statistics 36 (3), p. 469â487. External Links: Document Cited by: §1. Olson et al. (2021) K. Olson, J. D. Smyth, R. Horwitz, S. Keeter, V. Lesser, S. Marken, N. A. Mathiowetz, J. S. McCarthy, E. OâBrien, J. D. Opsomer, D. Steiger, D. Sterrett, J. Su, Z. T. Suzer-Gurtekin, C. Turakhia, and J. Wagner Transitions from telephone surveys to self-administered and mixed-mode surveys: aapor task force report. Journal of Survey Statistics and Methodology 9 (3), p. 381â411. External Links: Document Cited by: §1. Rothschild et al. (2024) D. M. Rothschild, J. Brand, H. Schroeder, and J. Wang Opportunities and risks of LLMs in survey research. Available at SSRN 5001645. Cited by: §1. Rothschild et al. (2025) D. M. Rothschild, T. D. Buskirk, S. Eckman, D. S. Hillygus, F. Kreuter, and D. Lazer Successfully navigating the disruption ai will bring to survey research. The Survey Statistician 92, p. 30â44. Cited by: §1. Santurkar et al. (2023) S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto Whose opinions do language models reflect?. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, p. 29971â30004. External Links: Link Cited by: §3. Schanze (2023) J. Schanze Response behavior and quality of survey data: comparing elderly respondents in institutions and private households. Sociological Methods & Research 52 (3), p. 1519â1555. Cited by: §1. Schouten et al. (2018) B. Schouten, A. Peytchev, and J. Wagner Adaptive survey design. CRC Press, Boca Raton, FL. Cited by: §2.2. Settles (2009) B. Settles Active learning literature survey. Computer Sciences Technical Report Technical Report 1648, University of WisconsinâMadison. External Links: Link Cited by: §1, §2.2. Stein et al. (2026) A. Stein, T. Gummer, E. Naumann, B. Rohr, H. Silber, R. Auriga, M. Bergmann, A. Bethmann, M. Blohm, C. Cornesse, P. Christmann, M. Coban, J. P. DĂ©cieux, B. Gauly, C. Hahn, S. Helmschrott, O. Hochman, J. Lemcke, D. Naber, S. Pötzschke, J. RoĂmann, J. Schanze, T. Schmidt, S. L. Schneider, H. Spangenberg, T. Rettig, M. Trappmann, M. Weinhardt, and B. WeiĂ Education bias in probability-based surveys in germany: evidence and possible solutions. International Journal of Social Research Methodology 29 (2), p. 197â214. External Links: Document Cited by: §1. Stein et al. (2024) J. Stein, T. Messingschlager, T. Gnambs, F. Hutmacher, and M. Appel Attitudes towards ai: measurement and associations with personality. Sci Rep 14, p. 2909. External Links: Document Cited by: §2.3. Tjuatja et al. (2024) L. Tjuatja, V. Chen, T. Wu, A. Talwalkwar, and G. Neubig Do LLMs exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics 12, p. 1011â1026. External Links: Document Cited by: §1, §1, §2.2, §3. von der Heyde et al. (2025a) L. von der Heyde, A. Haensch, B. WeiĂ, and J. Daikeler Using large language models for coding german open-ended survey responses on survey motivation. Survey Research Methods 19 (4), p. 355â370. External Links: Document Cited by: §1. von der Heyde et al. (2025b) L. von der Heyde, A. Haensch, and A. Wenz Vox populi, vox AI? using large language models to estimate german vote choice. Social Science Computer Review 44, p. 549â571. External Links: Document Cited by: §1, §1, §3. Waind (2020) E. Waind Trust, security and public interest: striking the balance: a review of previous literature on public attitudes towards the sharing, linking and use of administrative data for research. International Journal of Population Data Science 5 (3). External Links: Document Cited by: §2.2. Williams and Brick (2018) D. Williams and J. M. Brick Trends in us face-to-face household survey nonresponse and level of effort. Journal of Survey Statistics and Methodology 6 (2), p. 186â211. External Links: Document Cited by: §1. Wolf et al. (2021) C. Wolf, P. Christmann, T. Gummer, C. Schnaudt, and S. Verhoeven Conducting general social surveys as self-administered mixed-mode surveys. Public Opinion Quarterly 85 (2), p. 623â648. External Links: Document Cited by: §1. Wuttke et al. (2025) A. Wuttke, M. AĂenmacher, C. Klamm, M. M. Lang, Q. WĂŒrschinger, and F. Kreuter AI conversational interviewing: transforming surveys with LLMs as adaptive interviewers. In Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2025), A. Kazantseva, S. Szpakowicz, S. Degaetano-Ortlieb, Y. Bizzoni, and J. Pagel (Eds.), Albuquerque, New Mexico, p. 179â204. External Links: Link, Document, ISBN 979-8-89176-241-1 Cited by: §1. Xiao et al. (2020) Z. Xiao, M. X. Zhou, Q. V. Liao, G. Mark, C. Chi, W. Chen, and H. Yang Tell me about yourself: using an AI-powered chatbot to conduct conversational surveys with open-ended questions. ACM Trans. Comput.-Hum. Interact. 27 (3). External Links: ISSN 1073-0516, Document Cited by: §1. 4 Appendix A 4.1 AAPOR Disclosure Standards âą First data source: The survey data was collected by the authors. âą Data Collection Strategy: The survey data was collected with an online survey with participant recruitment on the provider Prolific. âą Research Sponsor and Conductor: The survey was designed and conducted by the authors and sponsored by their institution. This information will be specified in the non-blinded version of the manuscript. âą Measurement Tools/Instruments: Our study uses the nine questions from the questionnaire that are described in detail with instructions and response options in Appendix 5.1. âą Population Under Study: Participants had to be German residents and fluent in German. They were 18+, based on the properties of Prolific. âą Methods Used to Generate and Recruit the Sample: The sample selected is selected with a non-probability method. All participants on Prolific from the population under study were allowed to take part until the quota of 1200 participants was filled. Note that our final sample includes 1201 completed responses due to a technical issue. To diversify the sample, the study was opened at different times and days. Participants were contacted through Prolific. Participation was incentivised. âą Method(s) and Mode(s) of Data Collection: Web survey, in German language. âą Dates of Data Collection: The data was collected from June 1 through June 6 of 2026. âą Whether and How the Data Were Weighted: The data was not weighted. âą How the Data Were Processed and Procedures to Ensure Data Quality: All respondent had a 95% acceptance rate for microtasks on Prolific. Response times were monitored. We analyzed free-text answers to detect bot usage and added a question on AI usage in the questionnaire. Participants could only complete the survey once. âą IF A PANEL WAS USED: Panel Description: not applicable âą IF INTERVIEWERS OR CODERS WERE USED: Interviewer Details: not applicable âą IF ELIGIBILITY SCREENING WAS DONE: Screening criteria and process: No screening was conducted. âą Study Stimuli: not applicable âą Dispositions or Response or Participation Rates: The invitation to participate to the survey was sent to 6,032 eligible participants, of which 1260 started the survey. 1,201 participants completed the survey, while 45 exited at an early stage and 14 were filtered out by the platform due to exceeding the maximum time allowed and non-completion of task. âą Sample Sizes: Our sample size is 1201. âą Measurement and Model Specification: The analysis was conducted in Python with the pandas package for conducting the t-test and scipy for conducting the chi2-test. âą A General Statement Acknowledging Limitations of the Design and Data Collection: Our data collection is conducted with a non-probability method. Respondents recruited through Prolific are likely to be biased towards young, left-leaning and highly educated. 5 Appendix B 5.1 Survey Questions 5.1.1 AI usage I Instructions: Jetzt geht es um Ihre Arbeit fĂŒr Prolific im Allgemeinen. (Framing 1) Nutzen Sie bei Ihrer Arbeit fĂŒr Prolific ĂŒblicherweise KI-UnterstĂŒtzung? (Framing 2) Die Nutzung von KI-Tools ist bei Online-Arbeit inzwischen weit verbreitet und kann sehr unterschiedlich aussehen. Haben Sie in Ihrer Arbeit fĂŒr Prolific schon einmal KI-UnterstĂŒtzung verwendet? (Framing 3) Menschen nutzen bei Online-Studien unterschiedliche Hilfsmittel, darunter teilweise auch KI-UnterstĂŒtzung. Haben Sie in Ihrer Arbeit fĂŒr Prolific schon einmal KI-UnterstĂŒtzung verwendet? (Framing 4) FĂŒr die QualitĂ€t unserer Forschung ist eine möglichst genaue Beschreibung Ihrer tatsĂ€chlichen Arbeitsweise wichtig. Es gibt dabei keine richtigen oder falschen Antworten. Haben Sie in Ihrer Arbeit fĂŒr Prolific schon einmal KI-UnterstĂŒtzung verwendet? Bitte wĂ€hlen Sie alles Zutreffende aus. Bitte antworten Sie ehrlich. Ihre Antwort hat keinerlei Auswirkungen auf Ihre Teilnahme oder mögliche VergĂŒtung in dieser Studie und wird nicht an Prolific weitergegeben. Response options: Nein, ich nutze ĂŒblicherweise keine KI. - Ja, ich nutze KI ĂŒblicherweise, um Texte zu verbessern (z. B. Rechtschreibung, Stil oder sprachliche Ăberarbeitung). - Ja, ich nutze KI ĂŒblicherweise, um komplette Antworten generieren zu lassen (z. B. automatische Erstellung von Antworttexten). - Ja, ich nutze KI ĂŒblicherweise, um passende Antworten auszuwĂ€hlen (z. B. VorschlĂ€ge einer KI genutzt, um eine der vorgegebenen Antworten auszusuchen). - Ja, anderes: 5.1.2 AI usage I Instruction: Aus welchen GrĂŒnden nutzen Sie ĂŒblicherweise KI-UnterstĂŒtzung bei Ihrer Arbeit fĂŒr Prolific? Ihre ehrliche Antwort hat keinerlei Auswirkungen auf Ihre Teilnahme oder mögliche VergĂŒtung in dieser Studie. Es erfolgt keine RĂŒckmeldung an Prolific. Response options: open-ended question 5.1.3 Panel consent Instructions: Im Rahmen wissenschaftlicher Studien werden oft die gleichen Personen mehrfach befragt â man spricht dann von einem âPanelâ. Wie wahrscheinlich wĂ€re es, dass Sie an einem solchen Panel teilnehmen? Response options: sehr unwahrscheinlich - eher unwahrscheinlich - eher wahrscheinlich - sehr wahrscheinlich 5.1.4 Hybrid panel consent Instructions: Panels können durch KĂŒnstliche Intelligenz (KI) unterstĂŒtzt werden, z.B. indem KI beim Antworten hilft, Informationen ergĂ€nzt oder an Ihrer Stelle antwortet. Solche Formen werden als âhybride Panelsâ bezeichnet. Wie wahrscheinlich wĂ€re es, dass Sie an einem solchen hybriden Panel teilnehmen? Response options: sehr unwahrscheinlich - eher unwahrscheinlich - eher wahrscheinlich - sehr wahrscheinlich 5.1.5 Vignettes on hybrid panel consent Instructions: Nun stellen wir Ihnen verschiedene Formen eines hybriden Panels vor. Bitte lesen Sie diese Beschreibung sorgfĂ€ltig und geben Sie anschlieĂend Ihre persönliche EinschĂ€tzung ab. In einem hybriden Panel beantworten Teilnehmende Fragen zu gesellschaftlich relevanten Themen gemeinsam mit KI-Systemen. Wenn Sie teilnehmen, unterstĂŒtzt Sie die KI bei der Beantwortung von Fragen, ergĂ€nzt Informationen oder antwortet an Ihrer Stelle. Die KI verarbeitet dazu Informationen aus frĂŒheren Antworten und lernt aus Ihren Angaben. [Dimension 1: Sensitivity of data, low vs. high] [Low] Die KI verarbeitet nur allgemeine Informationen, zum Beispiel zu Ihrem Alter, Geschlecht, Ihren Hobbys oder Interessen. [High] Die KI verarbeitet sensible persönliche Informationen, zum Beispiel zu Ihrer Gesundheit, SexualitĂ€t, Religion oder Ihren politischen Einstellungen. [Dimension 2: AI autonomy & Sensitivity of simulation, low vs. high] [Low] Die KI wird nur auf Ihre direkte Anweisung aktiv und handelt nicht selbststĂ€ndig. Die KI macht Ihnen VorschlĂ€ge fĂŒr mögliche Antworten, die Sie annehmen oder ablehnen können. Sie selbst entscheiden, welche Fragen von der KI beantwortet werden. [High] Die KI handelt selbststĂ€ndig ohne Ihre direkte Anweisung. Die KI ĂŒbernimmt die Beantwortung von Fragen fĂŒr Sie. Die KI entscheidet, welche Fragen von ihr beantwortet werden. [Dimension 3: Financial compensation, low vs. high] [Low] FĂŒr Ihre Teilnahme erhalten Sie 13,90 Euro pro Stunde. Das entspricht dem Mindestlohn. [High] FĂŒr Ihre Teilnahme erhalten Sie erhalten 20,85 Euro pro Stunde. Das ist 50% mehr als der Mindestlohn. Ihre Daten werden nach hohen Datenschutzstandards verarbeitet. Sie werden ausfĂŒhrlich darĂŒber informiert, wie Ihre Daten gespeichert und genutzt werden. Sie können Ihre Teilnahme am hybriden Panel jederzeit beenden und die Löschung Ihrer Daten verlangen. Wie wahrscheinlich wĂ€re es, dass Sie an einem solchen hybriden Panel teilnehmen? Response options: sehr unwahrscheinlich - eher unwahrscheinlich - eher wahrscheinlich - sehr wahrscheinlich 5.1.6 Attitudes towards artificial intelligence scale (ATTARI-12) Instruction: Im Folgenden interessieren wir uns fĂŒr Ihre Einstellungen gegenĂŒber KĂŒnstlicher Intelligenz (KI). KĂŒnstliche Intelligenz kann Aufgaben ausfĂŒhren, die ĂŒblicherweise menschliche Intelligenz erfordern. Im Ă€uĂersten Fall befĂ€higt KI Maschinen dazu, selbststĂ€ndig und Ă€hnlich dem Menschen, ihre Umwelt wahrzunehmen, zu handeln, zu lernen und sich anzupassen. KĂŒnstliche Intelligenz kann Teil eines Computers oder einer Onlineplattform sein â man kann ihr aber auch in verschiedenen anderen technischen GerĂ€ten, wie etwa Robotern, begegnen. Bitte geben Sie an, inwieweit Sie den folgenden Aussagen zustimmen. Es gibt dabei keine richtigen oder falschen Antworten. Formulierung Facette Valenz 1 KĂŒnstliche Intelligenz wird die Welt verbessern. Kognitiv Positiv 2 Ich habe starke negative Emotionen gegenĂŒber kĂŒnstlicher Intelligenz. Affektiv Negativ (reverse-coded) 3 Ich möchte Technologien nutzen, die auf kĂŒnstlicher Intelligenz basieren. Behavioral Positiv 4 KĂŒnstliche Intelligenz hat mehr Nachteile als Vorteile. Kognitiv Negativ (reverse-coded) 5 Ich freue mich auf zukĂŒnftige Entwicklungen im Bereich kĂŒnstliche Intelligenz. Affektiv Positiv 6 KĂŒnstliche Intelligenz bietet Lösungen fĂŒr viele globale Probleme. Kognitiv Positiv 7 Ich bevorzuge Technologien, die keine kĂŒnstliche Intelligenz beinhalten. Behavioral Negativ (reverse-coded) 8 Ich fĂŒrchte mich vor kĂŒnstlicher Intelligenz. Affektiv Negativ (reverse-coded) 9 Ich wĂŒrde mich eher fĂŒr eine Technologie mit kĂŒnstlicher Intelligenz entscheiden als fĂŒr eine ohne. Behavioral Positiv 10 KĂŒnstliche Intelligenz verursacht eher Probleme, anstatt sie zu lösen. Kognitiv Negativ (reverse-coded) 11 Wenn ich an kĂŒnstliche Intelligenz denke, habe ich hauptsĂ€chlich positive GefĂŒhle. Affektiv Positiv 12 Ich möchte mit Technologien, die auf kĂŒnstlicher Intelligenz beruhen, lieber nichts zu tun haben. Behavioral Negativ (reverse-coded) Response options: stimme ĂŒberhaupt nicht zu â stimme eher nicht zu â weder noch â stimme eher zu â stimme voll zu 5.1.7 Political leaning Instruction: In der Politik reden die Leute hĂ€ufig von âlinksâ und ârechtsâ. Wo wĂŒrden Sie sich selbst einordnen? Response options: links 1 2 3 4 5 6 7 8 9 10 11 rechts - weiĂ nicht 5.1.8 Gender Instruction: Welches Geschlecht haben Sie? Response options: MĂ€nnlich - Weiblich - Divers 5.1.9 Education Instruction: Welchen höchsten allgemeinbildenden Schulabschluss haben Sie? Anmerkung: Wenn Sie Ihren höchsten Schulabschluss im Ausland gemacht haben, versuchen Sie sich bitte den vorgegebenen Kategorien zuzuordnen. Response options: Noch keinen â SchĂŒler/in - Von der Schule abgegangen ohne Schulabschluss - Hauptschulabschluss (Volksschulabschluss) oder gleichwertiger Abschluss - Polytechnische Oberschule der DDR mit Abschluss der 8. oder 9. Klasse - Realschulabschluss (Mittlere Reife) oder gleichwertiger Abschluss - Polytechnische Oberschule der DDR mit Abschluss der 10. Klasse - Fachhochschulreife - Abitur/Allgemeine oder fachgebundene Hochschulreife (Gymnasium beziehungsweise EOS, auch EOS mit Lehre) - Einen anderen Schulabschluss, und zwar: