Paper deep dive
Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice
Suyash Fulay, Prerna Ravi, Emily Kubin, Shrestha Mohanty, Michiel Bakker, Deb Roy
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/13/2026, 12:32:37 AM
Summary
Agora is an AI-powered platform designed to cultivate civic competence and consensus-finding skills by allowing users to interact with LLM-generated personas grounded in authentic human voice recordings. A study with 44 students demonstrated that users of the full interface, which provides access to voice explanations and dynamic feedback on policy proposals, reported higher levels of problem-solving and internal deliberation, and produced higher-quality consensus statements compared to a control group.
Entities (5)
Relation Signals (3)
Agora â uses â LLM
confidence 100% · We present Agora, an early-stage AI-powered platform that uses LLMs to organize authentic human voices
MIT Media Lab â developed â Agora
confidence 95% · Suyash Pradeep Fulay MIT Media Lab... We present Agora
Agora â fosters â Civic Competence
confidence 90% · helping users build consensus-finding skills... these initial findings point toward a promising direction for scaling civic education.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Deliberative democratic theory suggests that civic competence: the capacity to navigate disagreement, weigh competing values, and arrive at collective decisions is not innate but developed through practice. Yet opportunities to cultivate these skills remain limited, as traditional deliberative processes like citizens' assemblies reach only a small fraction of the population. We present Agora, an early-stage AI-powered platform that uses LLMs to organize authentic human voices on policy issues, helping users build consensus-finding skills by proposing and revising policy recommendations, hearing supporting and opposing perspectives, and receiving feedback on how policy changes affect predicted support. In a preliminary study with 44 university students, participants using the full interface (with access to voice explanations) reported higher levels of problem-solving skills, internal deliberation, and produced higher quality consensus statements compared to a control condition showing only aggregate support distributions. These initial findings point toward a promising direction for scaling civic education.
Tags
Links
- Source: https://arxiv.org/abs/2603.07339v1
- Canonical: https://arxiv.org/abs/2603.07339v1
Trouble viewing inline? Open PDF directly â
Full Text
31,169 characters extracted from source content.
Expand or collapse full text
by Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice Suyash Pradeep Fulay MIT Media Lab Massachusetts Institute of TechnologyCambridgeMassachusettsUSA sfulay@mit.edu , Prerna Ravi MIT Computer Science & Artificial Intelligence Laboratory Massachusetts Institute of TechnologyCambridgeMassachusettsUSA prernar@mit.edu , Om Gokhale MIT Media Lab Massachusetts Institute of TechnologyCambridgeMassachusettsUSA ogo@media.mit.edu , Eugene Yi Oxford Internet Institute University of OxfordOxfordUnited Kingdom eyi@mit.edu , Michiel Bakker MIT Media Lab Massachusetts Institute of TechnologyCambridgeMassachusettsUSA bakker@mit.edu and Deb Roy MIT Media Lab Massachusetts Institute of TechnologyCambridgeMassachusettsUSA dkroy@mit.edu (2026) Abstract. Deliberative democratic theory suggests that civic competenceâthe capacity to navigate disagreement, weigh competing values, and arrive at collective decisionsâis not innate but developed through practice. Yet opportunities to cultivate these skills remain limited, as traditional deliberative processes like citizensâ assemblies reach only a small fraction of the population. We present Agora, an early-stage AI-powered platform that uses LLMs to organize authentic human voices on policy issues, helping users build consensus-finding skills by proposing and revising policy recommendations, hearing supporting and opposing perspectives, and receiving feedback on how policy changes affect predicted support. In a preliminary study with 44 university students, participants using the full interface (with access to voice explanations) reported higher levels of problem-solving skills, internal deliberation, and produced higher quality consensus statements compared to a control condition showing only aggregate support distributions. These initial findings point toward a promising direction for scaling civic education. decision-making, artificial intelligence, collective intelligence, deliberation, education â journalyear: 2026â copyright: câ conference: Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems; April 13â17, 2026; Barcelona, Spainâ booktitle: Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA â26), April 13â17, 2026, Barcelona, Spainâ doi: 10.1145/3772363.3798888â isbn: 979-8-4007-2281-3/2026/04â ccs: Applied computing Voting / election technologiesâ ccs: Human-centered computing Empirical studies in collaborative and social computing 1. Introduction One assumption of deliberative democratic theory is that civic competence is not innate, but a practice of learning and refinement (McDevitt and Kiousis, 2006). The capacities that enable citizens to navigate disagreement, weigh competing values, and arrive at collective decisionsâwhat Kirlin (2003) terms âcivic skillsââare competencies that allow individuals to become participants in democratic processes rather than observers. These skills are not fixed traits; they must be learned and practiced (Kirlin, 2003). Framing deliberation as a skill rather than a disposition has important implications: if democratic capacities are learnable, question becomes how can we foster their development. Civic education highlights several key competencies: organization skills for collective tasks, communication skills for exchanging views, collective decision-making for working with others on shared goals, and critical thinking for evaluating competing claims (Kirlin, 2003). However, opportunities to practice these skills remain scarce (Zilouchian Moghaddam et al., 2015). While well-designed deliberative processesâcitizensâ assemblies, deliberative polls, structured forumsâcan improve the quality and consistency of participantsâ opinions (Fishkin, 2009), they reach only a small fraction of the public. We therefore explore whether technology can help close this gap through âAgoraâ, a platform that turns passive exposure to the opinion landscape on an issue into active skill practice: users propose a recommendation, hear authentic perspectives, and receive feedback as they continuously update their position. This iterative cycle mirrors with the core cognitive operations of deliberative skill development (Shaffer et al., 2017; Maia et al., 2024). Our approach builds on two ideas. First, drawing on Deweyâs pragmatist education philosophy, we argue democratic competencies develop through experience and feedbackânot instruction alone (Dewey, 1916). Second, we propose deliberative skills can be developed through mediated exposure to authentic diverse perspectives (Kriplean et al., 2012a; Kim et al., 2021), suggesting tools that provide LLM-organized real voices with scaffolding could cultivate these skills at scales beyond traditional face-to-face deliberation. We evaluate this approach through a study in which participants use the dynamic Agora interface to try to find consensus on two policy issues in the United States: the minimum wage and domestic vs. foreign hiring priorities. We assess the toolâs effects on fostering internal deliberation, perspective-taking, and learning, and improving the quality of generated consensus policies. In a randomized experiment with 44 students, the full version of the tool increased self-reported problem-solving skills, topic interest, internal deliberation, and perspective-taking, and produced consensus statements that were clearer, more coherent, and more specific. While this prototype is in its early stages and requires more rigorous evaluation for learning impact and consensus quality, these initial results are promising. 2. Background: HCI and Deliberation The HCI community has long explored interfaces for public deliberation. Early systems like Opinion Space used visualizations to help users navigate diverse viewpoints, showing that making opinion spectra visible and explorable can increase agreement with and respect for opposing views (Faridani et al., 2010). Agora builds on this key principle. Work on the Search Engine Manipulation Effect (SEME) also shows that how systems present viewpoints can shape opinion formation (Epstein and Robertson, 2015). Systems like Procid scaffolded group deliberation by structuring engagement with divergent viewpoints to support consensus-building (Zilouchian Moghaddam et al., 2015), underscoring the importance of balanced exposure to diverse perspectives in tools aimed at building deliberative capacity. A parallel HCI thread emphasizes structured reflection to improve deliberation quality (Yeo et al., 2025, 2024). Prompts to articulate pros/cons and engage with othersâ reasoning can foster deeper engagement (Kriplean et al., 2012a), and reflection nudges, like clarifying views or adopting an opponentâs perspective, can increase attitude certainty and willingness to express opinions beyond what information access alone achieves (Zhang et al., 2021). Recent work also uses LLM-based critical thinking tools to improve usersâ argument quality and open-mindedness in online deliberation (Pan et al., 2025; Khadar et al., 2025). These findings informed our design: we encourage users to engage with othersâ lived experiences and reasoning to build broader consensus. Prior work also shows that who is speaking matters for learning: platforms that surface opinions without stakeholder context can obscure lived consequences and intra-group variation (Kim et al., 2019). Listening-centered designs can increase communication satisfaction and willingness to participate (Kriplean et al., 2012b), and audio can be especially humanizing in polarized contexts (Schroeder et al., 2017). Accordingly, Agora pairs real participant profiles with voice clips to preserve tone and narrative and help users understand the reasoning behind different positions. 3. Tool Description Agora is built on a foundation of voice-based interviews conducted with participants who shared their lived experiences and beliefs related to policy issues. We process these interviews with LLMs to (1) predict each intervieweeâs level of support for a given policy and (2) create audio medleys of experiences and reasoning that support those predictions. 3.1. Data Collection: AI-Conducted Interviews Agora is grounded in a corpus of semi-structured voice interviews conducted with 90 participants on Prolific. Participants were required to have at least one year of work experience and be currently living in the United States. We aimed for political balance by using Prolificâs sampling settings to recruit a mix of liberal, moderate, and conservative participants. We adapted the AI interviewer system from Park et al. (Park et al., 2024), prioritizing low latency and voice-to-voice interaction to give participants the feeling of actually conversing with an interviewer (Appendix 3). The system uses OpenAIâs Whisper model for speech-to-text transcription (Radford et al., 2023), GPT-4o for generating contextually appropriate responses, and OpenAIâs tts-1 model for text-to-speech synthesis. The system automatically detects pauses from the participant, transcribes their speech, and generates and vocalizes the next interviewer utterance. Like a semi-structured interview, the LLM begins with general background questions and follows up with curiosity and respect about the intervieweeâs life. It then shifts to three policy topicsâminimum wage, race/gender in hiring, and domestic vs. foreign hiringâeliciting both personal experiences and beliefs (e.g., âHave you or someone close to you ever been impacted by immigration policy, especially around hiring?â, âHow do you feel about companies prioritizing hiring local applicants over foreign applicants?â). Capturing both enabled us to represent diverse perspectives grounded in lived experience. The AI interviewerâs questions can be adjusted to align with the specific policy topic selected for display in the main user interface below. Figure 1. Full Agora interface for treatment condition. Top image A shows how participants iterate and test their policies, Bottom image B shows the profile view participants see when clicking on the different avatars. Screenshot of the full Agora interface in the treatment condition. The top panel shows a policy input area on the left with text proposing to raise the federal minimum wage to $30 per hour, alongside a âCalculateâ button. The center displays a horizontal distribution of circular avatar icons depicting participant demographics positioned along a 0â100% support scale, with a vertical line marking average support at 45%. Avatars are grouped under âAgainst,â âOn the fence,â and âFor.â The bottom panel shows the same visualization with a side profile view opened for an avatar named Layla (20% support), including tabs for summary and quotes, an audio playback bar showing the medley, and a text explanation of her stance. 3.2. Agora User Interface The interface ( Figure 1A), presents users with a policy-drafting task. On the left, users can write and revise their own policies. On the right, a visualization displays avatars representing the interviewees, positioned along a horizontal axis by their predicted support for the userâs current policy (from 0% to 100%). The avatars were created with GPT-5 using the demographic information of each interviewee (age, race, and gender). When users click on an individual avatar, they can listen to a 60-90 second audio medley of that person explaining their perspective in their own voice, along with a text summary (Figure 1B). This medley presents experiences and reasoning supporting the modelâs predicted stance, helping users understand not just where people stand but why they hold their views. A leaderboard encourages users to iteratively maximize overall policy support. The LLM prompts are in Appendix 4. 3.3. Backend Implementation 3.3.1. Predicting Policy Support To position avatars along the support spectrum, we use GPT-4.1 to estimate each intervieweeâs level of support for a given policy. The model receives the interview transcript along with the policy text and is prompted to: (1) Rate the predicted support of the interviewee from 0â100, (2) Provide reasoning for its prediction, and (3) Give a confidence score in its prediction. While we did not use the modelâs reasoning directly in the interface, prompting for reasoning prior to giving an answer has been shown to increase prediction accuracy (Wei et al., 2023). We collected our intervieweesâ votes for the two policiesâraising the federal minimum wage to $â30 30/hour, and whether companies should prioritize domestic over foreign applicantsâvia a pre-survey prior to interaction with the AI interviewer. We then validated the LLM predictions against these initial stances reported by interviewees, finding an average accuracy of 82%. Figure 2. Agora interface for control condition Screenshot of the full Agora interface in the control condition. The left panel shows a policy input area with text proposing to raise the federal minimum wage to $30 per hour, along with a âCalculateâ button and topic selection dropdown. The center displays a horizontal distribution of circular blank avatar icons positioned along a 0â100% support scale, with a vertical line marking average support at 45%. Avatars are clustered across the spectrum from low to high support. Unlike the treatment condition, no profile side panel, audio playback, summaries, or quote details are visible in this view. 3.3.2. Generating Audio Medleys For each interviewee and policy proposal, we use GPT-4.1 to identify segments from the interview transcript that contain experiences and reasoning supporting the predicted stance. The model selects audio clips that ground the prediction in the intervieweeâs own words, prioritizing concrete personal experiences in line with research showing that such experiences are particularly effective at bridging moral and political divides (Kubin et al., 2021; Kessler et al., 2023). These clips are assembled into a 60â90 second medley that tool users can listen to when clicking on an avatar. We also included a âmeta-medleyâ feature that curates summary of clips from interviewees with low, medium, or high support, providing a quick overview of perspectives across the spectrum without clicking through individual profiles. These were generated by prompting GPT-4.1 to extract relevant snippets from a set of individual medleys. 3.3.3. Dynamic Feedback Loop A key feature is the dynamic feedback loop: when users revise their policy text and click âCalculate,â the system re-processes all interviewee transcripts through the LLM pipeline: (1) New policy text is sent to GPT-4.1 along with each intervieweeâs transcript. (2) Support predictions are regenerated for all interviewees based on the updated policy. (3) New audio medleys are generated that include experiences and reasoning relevant to the revised proposal. (4) Avatar positions shift along the horizontal axis to reflect updated support. (5) Individual profile views update with the newly generated medleys and summaries. Users can immediately see how different policy framings affect the support distribution across the population. For instance, a user might discover that adding state-specific considerations to a minimum wage proposal shifts several previously opposed interviewees toward support, and can click on those avatars to hear the specific experiences and reasons that make such exemptions resonate. 4. Experimental Setup and Evaluation We evaluated the tool with 44 university students in the United States in a fully online study approved as exempt under MITâs Institutional Review Board (IRB). We recruited these students through mailing lists, messaging platforms, and word of mouth. Participants were asked to draft optimal policies on two topics: what the minimum wage should be and whether companies should prioritize hiring domestic over foreign applicants. Participants spent approximately 30-45 minutes completing the task and received $10 for participation, with an additional $50 bonus awarded to those whose policy proposals received the most support. During the task, participants aimed to maximize in silico support from avatars grounded in the beliefs and experiences of the interview participants. Those in the treatment condition saw the full interface, enabling them to see and hear the reasons behind each avatarâs predicted support, revise their policies, and observe how changes affected support levels. Those in the control condition used the same drafting tool, but avatars appeared as generic icons without interactive capabilities. Control participants could see how the distribution of predicted support changed with new policies, but not why it changed (see Figure 2). We chose this control design to isolate the impact of profile exploration and to establish a baseline for participantsâ prior knowledge and idea generation without external stimuli. Video walkthroughs for both conditions are in Appendix 1. We evaluated the tool for its impact on participantsâ self-reported learning and the quality of the consensus statements via a 10 minutes post survey (Appendix 2). To evaluate learning impact, we adopted several subscales from Shroff et al. (2019), who designed and validated several scales measuring studentsâ perceptions of technology-enabled active learning. We used scales that measure problem-solving skills, interest, and feedback. We also measured whether the tool helped participants âdeliberate withinâ. Deliberation within is a concept coined by Goodin (2000), for describing the process of considering and weighing alternatives by oneself prior to interpersonal deliberation. Weinmann (2018) propose a scale to measure this concept, which we use to evaluate our tool. To evaluate our toolâs ability to help generate practical outcomes, we also measure the quality of the consensus statements written by participants. We use an LLM-as-a-judge approach (Gu et al., 2025), that uses language models to code large amounts of open-ended text with a rubric, to measure the coherence, clarity, evidence integration, specificity, and balanced treatment of uncertainty of the consensus statements. (a) Mean survey responses across conditions measuring self-reports of learning and internal deliberation. (b) LLM-judged quality of consensus statements in treatment and control conditions. Figure 3. Agora learning outcomes and consensus quality. Left panel (a) shows mean survey responses (1â7 scale) for five learning and deliberation measures: Problem Solving (Control 4.73, Treatment 5.48), Interest (Control 4.81, Treatment 5.16), Understanding Others (Control 4.72, Treatment 5.60), Deliberation (Control 5.53, Treatment 5.93), and Feedback (Control 5.00, Treatment 5.61). Control (n=17) is shown in gray and Treatment (n=27) in red, with individual participant points overlaid. Right panel (b) shows LLM-judged consensus quality scores (0â100 scale) across six dimensions: Overall Score (Control 62.1, Treatment 65.9), Clarity (Control 77.1, Treatment 78.7), Coherence (Control 73.1, Treatment 76.2), Evidence Integration (Control 53.3, Treatment 56.9), Specificity & Actionability (Control 59.3, Treatment 64.2), and Balance & Uncertainty (Control 47.6, Treatment 53.5). Across all measures, the treatment condition scores higher than the control condition. 5. Results Participants in the treatment condition reported greater development of problem-solving skills, interest, and perceived the feedback from the tool as more timely and relevant (3(a)). They also reported higher scores on measures of âdeliberating within,â suggesting that students felt the full tool was more effective at fostering internal deliberation compared to the control. However, these are self-reported perceptions of learning, not direct measure of skill acquisition, and may reflect engagement or satisfaction rather than true learning gains such as capacity to navigate opposing perspectives. We are currently developing objective measures of deliberative skill development: pre/post assessments of consensus-finding skill and topical knowledge. In terms of consensus statement quality (3(b)), we used the LLM-as-a-judge approach to evaluate clarity, coherence, specificity, evidence integration, and balanced treatment of uncertainty. Statements from the treatment condition received higher scores across these dimensions, particularly for specificity and balanced acknowledgment of uncertainty. We observed that participants in the control condition tended to default to vague, generic statements that garnered broad support but lacked actionable detail (as seen in 3(b) with respect to Specificity and Actionability). However, LLM-based evaluation has known limitations (prompt sensitivity, model bias, and uncertain alignment with human judgment), so these scores should be viewed as preliminary. We are currently conducting a validation study with human annotators to assess the reliability of these LLM-generated scores and identify dimensions where model and human judgments diverge. 6. Discussion Our findings suggest that dynamic engagement with personas grounded in human experience can foster deliberative skill development. Participants who used the full Agora interface, with access to authentic voices explaining their positions, reported greater gains in problem-solving skills, interest, and internal deliberation compared to those who only saw aggregate support distributions. This also translated into higher-quality consensus statements that were more clear, coherent, and specific. The differences between conditions suggest that the why behind othersâ positions matters: understanding the reasons people hold their views, not just the distribution of those views, appears central to developing the perspective-taking capacities that deliberative theorists identify as foundational to democratic competence (Kirlin, 2003). These findings point toward a promising direction for civic education. If deliberative skills can be practiced through mediated engagement with authentic perspectives, tools like Agora could help address deliberative democracyâs scalability challenge by widening access beyond the small number of citizens who join well-designed deliberative processes (Fishkin, 2009). Rather than replacing face-to-face deliberation, such tools could serve as preparation: building the cognitive and affective capacities citizens will need when encountering collective decision-making opportunities. Limitations and Future Work: Our study has several limitations. Our small sample consisted entirely of university students, and the topics, while contentious, were relatively familiar, so results may not generalize to broader public in their baseline deliberative capacities or more technical/emotionally charged issues. In addition, our control condition isolates profile exploration but cannot pinpoint which full-interface features drove effects; future work should vary access to voice clips vs written explanations, and dynamic feedback on policy updates to identify the active ingredients of skill development. The current visualization reduces policy support to a single axis to enable comparability and reflect how public opinion is typically summarized. However, this simplification overlooks the inherently complex and heterogeneous nature of real-world policy preferences, as individuals may endorse the same position for very different reasons. Future iterations of the tool could better represent the nuanced, multidimensional trade-offs underlying policy decisions. We also identify ethical concerns. The current tool does not fully address privacy (e.g., voice identity anonymization), and future versions should give users more agency over what they disclose in democratic processes. LLM biases may have shaped how different peopleâs perspectives were organized and how numerical policy support was updated dynamically; while our prompts are public and accessible, behavior of these models remains hard to interpret. Future work could involve soliciting intervieweesâ feedback on how their views are represented through selectively curated audio medleys, as well as how shifts in support are portrayed following dynamic policy updates. This would help more rigorously evaluate the credibility and faithfulness of algorithmic interpretations of individualsâ perspectives. Finally, the tool cannot fully replace genuine face-to-face deliberation and may risk creating scalable but shallow connections, similar to social media. References J. Dewey (1916) Democracy and education: an introduction to the philosophy of education. Macmillan, New York. Cited by: §1. R. Epstein and R. E. Robertson (2015) The search engine manipulation effect (seme) and its possible impact on the outcomes of elections. Proceedings of the national academy of sciences 112 (33), p. E4512âE4521. Cited by: §2. S. Faridani, E. Bitton, K. Ryokai, and K. Goldberg (2010) Opinion space: a scalable tool for browsing online comments. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, p. 1175â1184. Cited by: §2. J. S. Fishkin (2009) When the people speak: deliberative democracy and public consultation. Oxford University Press, Oxford. Cited by: §1, §6. R. E. Goodin (2000) Democratic deliberation within. Philosophy and Public Affairs 29 (1), p. 81â109. External Links: Document Cited by: §4. J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y. Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y. Wang, W. Gao, L. Ni, and J. Guo (2025) A survey on llm-as-a-judge. External Links: 2411.15594, Link Cited by: §4. D. Kessler, D. Dimitrakopoulou, and D. Roy (2023) Hearing personal experiences improves social evaluations compared to personal opinions, especially for polarized parties. SSRN Electronic Journal. External Links: Document, Link Cited by: §3.3.2. M. Khadar, D. Runningen, J. Tang, S. Chancellor, and H. Kaur (2025) Wisdom of the crowd, without the crowd: a socratic llm for asynchronous deliberation on perspectivist data. Proceedings of the ACM on Human-Computer Interaction 9 (7), p. 1â35. Cited by: §2. H. Kim, H. Kim, K. J. Jo, and J. Kim (2021) StarryThoughts: facilitating diverse opinion exploration on social issues. Proceedings of the ACM on Human-Computer Interaction 5 (CSCW1), p. 1â29. Cited by: §1. H. Kim, E. Ko, D. Han, S. Lee, S. T. Perrault, J. Kim, and J. Kim (2019) Crowdsourcing perspectives on public policy from stakeholders. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems, CHI EA â19, New York, NY, USA, p. 1â6. External Links: ISBN 9781450359719, Link, Document Cited by: §2. M. Kirlin (2003) The role of civic skills in fostering civic engagement. CIRCLE Working Paper (6). Cited by: §1, §1, §6. T. Kriplean, J. Morgan, D. Freelon, A. Borning, and L. Bennett (2012a) Supporting reflective public thought with considerit. In Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work, CSCW â12, New York, NY, USA, p. 265â274. External Links: ISBN 9781450310864, Link, Document Cited by: §1, §2. T. Kriplean, M. Toomim, J. Morgan, A. Borning, and A. J. Ko (2012b) Is this what you meant? promoting listening on the web with reflect. In proceedings of the SIGCHI conference on human factors in computing systems, p. 1559â1568. Cited by: §2. E. Kubin, C. Puryear, C. Schein, and K. Gray (2021) Personal experiences bridge moral and political divides better than facts. Proceedings of the National Academy of Sciences 118 (6), p. e2008389118. External Links: Document, Link, https://w.pnas.org/doi/pdf/10.1073/pnas.2008389118 Cited by: §3.3.2. R. C. Maia, G. Hauber, D. Cal, and A. Veloso LeĂŁo (2024) Teaching and developing deliberative capacities: an integrated approach to peer-to-peer, playful, and authentic discussion-based learning. Democracy & Education 32 (1), p. Article 5. External Links: Document, Link Cited by: §1. M. McDevitt and S. Kiousis (2006) Deliberative learning: an evaluative approach to interactive civic education. Communication Education 55 (3), p. 247â264. External Links: Document, Link, https://doi.org/10.1080/03634520600748557 Cited by: §1. Q. Pan, J. Zeng, J. Wang, J. Liu, Y. Qiu, K. Yuan, and Z. Peng (2025) AMQuestioner: training critical thinking with question-driven interactive argument maps in online discussion. Proceedings of the ACM on Human-Computer Interaction 9 (7), p. 1â48. Cited by: §2. J. S. Park, C. Q. Zou, A. Shaw, B. M. Hill, C. Cai, M. R. Morris, R. Willer, P. Liang, and M. S. Bernstein (2024) Generative agent simulations of 1,000 people. External Links: 2411.10109, Link Cited by: §3.1. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever (2023) Robust speech recognition via large-scale weak supervision. In International conference on machine learning, p. 28492â28518. Cited by: §3.1. J. Schroeder, M. Kardas, and N. Epley (2017) The humanizing voice: speech reveals, and text conceals, a more thoughtful mind in the midst of disagreement. Psychological science 28 (12), p. 1745â1762. Cited by: §2. T. J. Shaffer, N. V. Longo, I. Manosevitch, and M. S. Thomas (2017) Deliberative pedagogy: teaching and learning for democratic engagement. Michigan State University Press. External Links: ISBN 9781611862492, Link Cited by: §1. R. H. Shroff, F. S. T. Ting, and W. H. Lam (2019) Development and validation of an instrument to measure studentsâ perceptions of technology-enabled active learning. Australasian Journal of Educational Technology 35 (4). External Links: Link, Document Cited by: §4. J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou (2023) Chain-of-thought prompting elicits reasoning in large language models. External Links: 2201.11903, Link Cited by: §3.3.1. C. Weinmann (2018) Measuring political thinking: development and validation of a scale for âdeliberation withinâ. Political Psychology 39 (2), p. 365â380. External Links: ISSN 0162895X, 14679221, Link Cited by: §4. S. Y. Yeo, G. Lim, J. Gao, W. Zhang, and S. T. Perrault (2024) Help me reflect: leveraging self-reflection interface nudges to enhance deliberativeness on online deliberation platforms. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Cited by: §2. S. Yeo, Z. Jiang, A. Tang, and S. T. Perrault (2025) Enhancing deliberativeness: evaluating the impact of multimodal reflection nudges. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, p. 1â26. Cited by: §2. W. Zhang, T. Yang, and S. Tangi Perrault (2021) Nudge for reflection: more than just a channel to political knowledge. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI â21, New York, NY, USA. External Links: ISBN 9781450380966, Link, Document Cited by: §2. R. Zilouchian Moghaddam, Z. Nicholson, and B. P. Bailey (2015) Procid: bridging consensus building theory with the practice of distributed design discussions. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing, p. 686â699. Cited by: §1, §2.