Paper deep dive
How People Evaluate AI-, Expert-, and Peer-Style Financial Advice
Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/12/2026, 3:12:52 AM
Summary
This study investigates how people evaluate AI-generated financial advice compared to expert and peer-style advice, focusing on the roles of source attribution and communication style. Through a preregistered vignette experiment with 285 participants, the authors found that expert advice is generally rated more favorably than AI advice across multiple dimensions, even without source labels. However, mislabeling AI advice as expert significantly increased its perceived quality and situational fit, while correct labels had limited impact. The findings suggest that disclosure is not a neutral transparency mechanism but an interpretive frame that interacts with message-level cues to shape trust and reliance.
Entities (9)
Relation Signals (5)
Expert Advice → isratedmorefavorablythan → AI Advice
confidence 95% · Expert advice was rated more favorably than AI advice on 9 of 10 outcomes
Source Attribution → shapes → Financial-Advice Evaluations
confidence 92% · financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues
Mislabeling → increasesratingof → AI Advice
confidence 90% · mislabeling increased ratings of AI advice for situational fit and overall quality
Disclosure → isnot → Neutral Transparency Mechanism
confidence 90% · We position disclosure not as a neutral transparency mechanism, but as an interpretive frame
AI Advice → ismostresponsiveto → Displayed Attribution
confidence 88% · AI advice was most responsive to displayed attribution
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
Tags
Links
- Source: https://arxiv.org/abs/2608.09019v1
- Canonical: https://arxiv.org/abs/2608.09019v1
Trouble viewing inline? Open PDF directly →
Full Text
56,825 characters extracted from source content.
Expand or collapse full text
How People Evaluate AI-, Expert-, and Peer-Style Financial Advice ARYAN RAMCHANDRA KAPADIA, University of Illinois Urbana-Champaign, USA ESHWAR CHANDRASEKHARAN, University of Illinois Urbana-Champaign, USA KOUSTUV SAHA, University of Illinois Urbana-Champaign, USA As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (푁=285) in which substantive financial content—including facts, numerical values, recommendation direction, and core reasoning—was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|푑|=0.20–0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to푑=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (푑=0.42 for each) and attenuated the Expert advantage in situational fit (푑= −0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance. CCS Concepts:• Human-centered computing→ Empirical studies in HCI. Additional Key Words and Phrases: AI financial advice, source attribution, trust calibration, financial decision-making 1 Introduction The role of generative AI tools is no longer limited to general information-seeking; individuals are increasingly using them to seek guidance in high-stakes decision contexts, including medicine [1,24] and finance [16]. In such settings, AI-generated advice is not simply received; it is evaluated by users, shaping their trust and reliance [17–19]. Personal finance is one such consequential domain, where everyday decisions about debt, budgeting, investing, and financial planning can have lasting effects on financial wellbeing [13]. People seeking financial guidance encounter AI-generated advice alongside recommendations from professional advisors, friends and family, and online communities. Yet it remains unclear how users evaluate this AI-generated financial advice relative to advice from these other sources. Understanding these influences is important because users may accept or reject financial guidance not only based on its informational quality, but also on perception cues that may be indirectly related to advice credibility, safety, and trustworthiness. Prior work presents a mixed picture of how users respond to AI advice. In some contexts, people exhibit algorithm appreciation [11,23] and overreliance, preferring algorithmic recommendations despite associated errors or costs [10,20]; in others, people are reluctant to use algorithmic recommendations, particularly in uncertain decision domains, consistent with algorithmic aversion [5,14]. Prior work also indicates that reader preferences can vary with the communication styles of LLM- and human-authored explanations, motivating closer separation of message-level cues from source identity [27]. Recent comparisons of LLM and online-community responses have identified differences in communication Authors’ Contact Information: Aryan Ramchandra Kapadia, kapadia8@illinois.edu, University of Illinois Urbana-Champaign, Urbana, IL, USA; Eshwar Chandrasekharan, eshwar@illinois.edu, University of Illinois Urbana-Champaign, Urbana, IL, USA; Koustuv Saha, ksaha2@illinois.edu, University of Illinois Urbana-Champaign, Urbana, IL, USA. 1 arXiv:2608.09019v1 [cs.HC] 10 Aug 2026 2Kapadia et al. patterns, including more structured and neutral language in AI responses and greater use of conversational engagement, personalization, and lived experience in community responses [21,22,26]. More broadly, advice-taking and persuasion research suggests that evaluations may depend on both source cues—who the advice appears to come from—and message cues—how it communicates reasoning through tone, structure, and framing [2,3,8,25]. It therefore remains unclear whether differences in how users evaluate AI, expert, and peer advice arise from source attribution, source-specific communication cues, or their interaction. In this study, we examine how source-specific communication style and displayed source attribution independently and jointly shape evaluations of financial advice. Accordingly, we ask: (1) when substantive financial content is held constant, how do AI-, Expert-, and Online Community-style advice differ in evaluations of presentation and comprehension, safety and risk, source authority, trust, and willingness to follow; and (2) how do displayed source labels shape these evaluations when advice is correctly labeled, unlabeled, or mislabeled? We answer these questions through a preregistered vignette experiment with 푁= 285 U.S. adults across eight personal-finance scenarios. We standardize financial facts, numerical values, recommendation direction, and core reasoning across advice styles while independently manipulating displayed source attribution through correctly labeled, unlabeled, and mislabeled conditions. We find that expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|푑|=0.20–0.47). These differences were already visible without source labels: in the unlabeled condition, Expert advice outperformed AI advice on 8 of 10 outcomes, with effects as large as푑=0.60 for situational fit. Correct labels added limited differentiation beyond message-level communication cues, whereas mislabeling selectively increased ratings of AI advice for situational fit and overall quality (푑=0.42 for each). In complementary descriptive sensitivity analyses, AI-style content was most responsive to displayed attribution: presenting it with an Expert label rather than no label improved both situational fit and overall quality by푑=0.47. Conversely, differences between the underlying advice styles were most apparent when the advice was displayed with an AI label. Together, our work contributes a controlled three-source comparison that disentangles source-specific communication cues from displayed attribution, showing that apparent source effects reflect how advice is communicated rather than labels alone. These findings imply that disclosure is not a neutral provenance cue: financial AI interfaces should pair accurate attribution with support for evaluating advice reasoning, assumptions, and risks. 2 Study Design and Methods We conducted a randomized vignette-based survey experiment on Prolific with푁=285 U.S. adults (see Table 1 for demographics) after excluding incomplete responses, failed attention checks, and submissions completed in under two minutes. The study used eight realistic personal-finance scenarios arranged in a 2×2×2 factorial structure crossing stakes, external uncertainty, and verifiability (Table 2). Participants were assigned to one of two scenario groups and evaluated four scenarios within their assigned group. Within each group, participants were assigned to one of three counterbalanced advice versions. Each participant saw all three source styles at least once, with one source style appearing twice. For each scenario, advice versions preserved the same financial facts, numerical values, recommendation direction, and core reasoning, while varying source-specific communication style, tone, and reasoning format: neutral and analytical AI advice, structured and principle-based expert advice, and informal experience-based online community advice. 1 Table 3 summarizes the style specifications, and Figure 5 provides matched examples illustrating how the same 1 For concision, we refer to the three source-specific communication-style conditions as AI, Expert, and OC advice; these terms denote the underlying advice version rather than its displayed source attribution. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice3 financial backbone was rendered across conditions. Figure 1 provides an overview of the study design, including advice construction, scenario structure, advice style× source attribution manipulation, and participant evaluation. Participants were randomly assigned to one of three source-labeling arms: labeled, unlabeled, or mislabeled. The mislabeled arm used one of two counterbalanced label-swap schemas (Figure 2). After each vignette, participants rated the advice on ten 7-point Likert items capturing readability, reasoning clarity, situational fit, risk acknowledgment, financial harm risk, misleadingness, perceived source knowledge, overall quality, trust intention, and reliance intention (Table 4). We analyzed responses using mixed-effects regressions with participant random intercepts, scenario fixed effects, and scenario familiarity as a covariate (Table 5). The study was approved by the Institutional Review Board at our university. 3 Results 3.1 Evaluation Differences by Source-Specific Communication Style Expert advice was rated more favorably than AI advice on 9 of 10 evaluation dimensions, while online community (OC) advice varied systematically across certain dimensions (Table 6). Presentation and Comprehension. Both Expert and OC advice were rated as more readable than AI advice, with OC showing the largest advantage (Expert vs. AI:푑=0.47; OC vs. AI:푑=0.53). Reasoning clarity showed a similar pattern: Expert (푑=0.30) and OC (푑=0.17) outperformed AI, though the effect was stronger for Expert advice. Expert advice was also perceived as a significantly better fit for the protagonist’s situation than both AI (푑=0.39) and OC (푑=0.24). Risk and Safety Perception. Expert advice was perceived as safer than AI and OC advice. It was rated as posing a lower risk of financial harm than both AI (푑=−0.27) and OC advice (푑=−0.23), while the non-significant AI–OC contrast suggests that participants perceived comparable risk across both non-expert sources. AI advice was rated as more likely to mislead a person with limited financial knowledge than both Expert (푑=−0.30) and OC (푑=−0.26). One notable reversal emerged: AI advice was rated as acknowledging risks and uncertainties more than OC advice (푑=−0.17). Source authority and Behavioral intentions. Expert advice presented the clearest advantage on evaluative and be- havioral outcomes. Compared with AI, it was rated higher on perceived source knowledge (푑=0.20), overall quality (푑=0.27), trust intention (푑=0.24), and reliance intention (푑=0.23). It also outperformed OC across all four dimensions, with small-to-moderate effect sizes (푑s = 0.19 to 0.39). Strikingly, OC was rated significantly less knowledgeable than AI (푑=−0.19) and did not outperform it in overall quality, trust, or reliance. 3.2 Source Attribution and Advice-Style Interactions To evaluate whether displayed source attribution changed advice evaluations beyond the underlying advice styles, we modeled the source-labeling arm using the unlabeled condition as the reference (Table 7). In the unlabeled arm, participants could already distinguish the three advice styles from message-level communication cues alone. Expert advice was evaluated more favorably than AI advice on 8 of 10 dimensions, including readability, reasoning clarity, situational fit, financial harm risk, misleadingness, overall quality, trust intention, and reliance intention (|푑|=0.27–0.60). OC advice was also distinguishable from AI advice on readability (푑=0.43), situational fit (푑=0.28), misleadingness (푑=−0.27), and perceived knowledgeability (푑=−0.35), with OC rated as less knowledgeable than AI. 4Kapadia et al. Correct source labels added limited explanatory value beyond these message-level communication cues. The labeled- arm main effect was significant only for situational fit (푑=0.40), and no Labeled×Expert or Labeled×OC interaction reached significance. Mislabeling, however, produced selective shifts in evaluation. Relative to the unlabeled baseline, AI advice shown with a non-AI label received higher ratings for situational fit (푑=0.42) and overall quality (푑=0.42). The Mislabeled×Expert interaction for situational fit was also significant (푑=−0.36), indicating that incorrect labels attenuated the advantage of Expert over AI advice on this dimension. A Mislabeled×OC interaction (푑= −0.41) indicated that mislabeling altered the OC–AI difference in risk acknowledgment. To further visualize the interaction between displayed attribution and message-level communication cues, we conducted descriptive sensitivity analyses on participant-aggregated ratings. In the label-sensitivity analysis, we held the underlying advice style constant and varied the displayed source label (Figure 3). Label effects were most pronounced for AI advice: ratings varied across displayed-label conditions for situational fit (KW퐻=9.6) and overall quality (KW 퐻=9.4), with the clearest improvements when the same AI advice was presented with an Expert label rather than no label for both situational fit (푑=0.47) and overall quality (푑=0.47). Expert advice showed comparatively little sensitivity to relabeling, whereas OC advice showed a more localized label effect for readability (KW퐻=11.7). In the complementary advice-style sensitivity analysis, we held the displayed source label constant and varied the underlying advice style (Figure 4). Advice-style differences were most visible under the AI label: Expert advice was rated more favorably than AI advice on readability (KW퐻=7.1,푑=0.49) and misleadingness (KW퐻=8.3,푑=−0.47). Under the Expert label, advice-style sensitivity was concentrated primarily in readability (KW퐻=18.9), with Expert advice (푑=0.57) and OC advice (푑=0.74) rated as more readable than AI advice. Under the OC label, advice-style differences were again more localized, with OC advice rated as more readable than AI advice (KW퐻=8.7,푑=0.51), but also as posing greater financial harm risk than Expert advice (KW 퐻= 6.1, 푑= 0.41). 4 Discussion Our findings show how displayed source attribution and message-level communication cues jointly shape evaluations of AI financial advice, with direct implications for disclosure policy and system design. The AI evaluation gap is not solely label-driven. Expert advice was rated more favorably than AI advice across 9 of 10 dimensions, yet this hierarchy remained identifiable even without any source label: participants distinguished Expert from AI advice based on message-level communication cues. This suggests the AI evaluation gap is not purely a labeling artifact; it reflects perceived differences in how AI and human advice communicate. This gap is also dimension- specific: AI advice was rated as comparable to expert advice on risk acknowledgment, while most of the penalty is concentrated in perceived authority, safety, and source credibility. Participants therefore did not simply reject AI advice as insufficiently analytical. This aligns with prior work showing that people can evaluate AI-generated support favorably on communicative qualities such as sincerity and actionability, even while recognizing aspects of human support that AI may not replicate [4]. Overall, our findings suggest that AI systems adopting the structured, principle-based communication patterns associated with Expert advice may partially reduce the perceived evaluation gap. Disclosure effects are limited and accuracy-dependent. Correct source labels added limited differentiation beyond what message-level communication cues already conveyed. Participants could distinguish the underlying advice styles even without explicit attribution, suggesting that disclosure alone was not the primary driver of evaluation. However, inaccurate attribution was more consequential: mislabeling selectively weakened distinctions between advice styles that were otherwise visible from message-level communication cues. This asymmetry, in which correct labels offered How People Evaluate AI-, Expert-, and Peer-Style Financial Advice5 modest benefits while incorrect labels distorted calibration, suggests that the value of disclosure depends critically on its accuracy. When AI advice is misattributed, inadvertently or by design, inaccurate disclosure may undermine users’ ability to calibrate trust to advice quality. The AI label may heighten scrutiny. Differences between advice styles were most visible under the AI label. When advice was labeled as AI, participants distinguished Expert advice from AI advice on readability and perceived risk, suggesting that AI attribution may heighten attention to message-level communication cues. By contrast, under the Expert label, sensitivity to advice style was concentrated primarily in readability, while broader differences in quality, trust, and reliance were less apparent. This pattern suggests a scrutiny asymmetry: AI labels may induce greater vigilance, whereas Expert labels may encourage greater deference. Source authority, rather than AI identity alone, may therefore shape how carefully people evaluate advice [3, 25]. Design and policy implications. Our findings suggest that calibrated evaluation of AI financial advice requires more than source disclosure. Three design directions follow. First, interfaces should support advice-level evaluation through explicit reasoning, transparent assumptions, verifiable claims, and calculations, rather than treating source labels as proxies for quality [9]. Second, disclosure accuracy matters, not merely disclosure presence. When AI advice is presented as originating from a human expert or online community, users may calibrate their trust against an inaccurate provenance cue rather than the advice’s reasoning and quality. Third, the scrutiny asymmetry observed in our descriptive analyses suggests that AI labels may encourage more critical evaluation. Rather than minimizing AI attribution to reduce stigma, designers might consider how to preserve its scrutiny-activating properties while also reducing unwarranted discounting of high-quality AI content. Limitations. Our study has several limitations. First, advice style was manipulated as a bundled set of voice, tone, framing, structure, and reasoning features, so we cannot isolate the contribution of any single feature. Second, the stimuli were experimentally constructed and reviewed for information parity; they should not be interpreted as representative of all AI assistants, financial professionals, or online communities. Third, participants evaluated short, text-based vignettes and reported intended trust and reliance rather than making actual financial decisions, limiting behavioral and ecological validity. Fourth, although the scenarios varied in stakes, external uncertainty, and verifiability, the study was not designed to estimate generalizable effects of these dimensions. Finally, the U.S. Prolific sample limits broader generalizability, and the descriptive sensitivity analyses should not be interpreted as confirmatory causal evidence. Taken together, our results shift the question from whether AI advice should be disclosed to how source attribution interacts with message-level communication cues. In financial contexts, labels are not neutral provenance markers; they frame interpretation and selectively shape judgments of advice quality, trust, and willingness to rely on it. References [1]Oluwatobiloba Ayo-Ajibola, Ryan J. Davis, Matthew E. Lin, Jeffrey Riddell, and Richard L. Kravitz. 2024. Characterizing the Adoption and Experiences of Users of Artificial Intelligence–Generated Health Information in the United States: Cross-Sectional Questionnaire Study. Journal of Medical Internet Research 26, 1 (Aug. 2024), e55138. doi:10.2196/55138 [2]Silvia Bonaccio and Reeshad S. Dalal. 2006. Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences. Organizational Behavior and Human Decision Processes 101, 2 (Nov. 2006), 127–151. doi:10.1016/j.obhdp.2006.07.001 [3] Shelly Chaiken. 1980. Heuristic versus systematic information processing and the use of source versus message cues in persuasion. Journal of Personality and Social Psychology 39, 5 (1980), 752–766. doi:10.1037/0022-3514.39.5.752 [4]Vedant Das Swain, Qiuyue "Joy" Zhong, Jash Rajesh Parekh, Yechan Jeon, Roy Zimmerman, Mary Czerwinski, Jina Suh, Varun Mishra, Koustuv Saha, and Javier Hernandez. 2025. AI on My Shoulder: Supporting Emotional Labor in Front-Office Roles with an LLM-based Empathetic Coworker. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 6Kapadia et al. [5]Berkeley J. Dietvorst and Soaham Bharti. 2020. People Reject Algorithms in Uncertain Decision Domains Because They Have Diminishing Sensitivity to Forecasting Error. Psychological Science 31, 10 (Oct. 2020), 1302–1314. doi:10.1177/0956797620948841 [6]John Grable and Ruth Lytton. 2001. Assessing The Concurrent Validity Of The SCF Risk Tolerance Question. Journal of Financial Counseling and Planning 12 (Jan. 2001). [7]Simone Grassini. 2024. A Psychometric Validation of the PAILQ-6: Perceived Artificial Intelligence Literacy Questionnaire. In Nordic Conference on Human-Computer Interaction. ACM, Uppsala Sweden, 1–10. doi:10.1145/3679318.3685359 [8]Carl I. Hovland and Walter Weiss. 1951. The Influence of Source Credibility on Communication Effectiveness. The Public Opinion Quarterly 15, 4 (1951), 635–650. https://w.jstor.org/stable/2745952 JSTOR: 2745952. [9] Michelle Huang, Agam Goyal, Koustuv Saha, and Eshwar Chandrasekharan. 2026. Answer Bubbles: Information Exposure in AI-Mediated Search. arXiv preprint arXiv:2603.16138 (2026). [10]Artur Klingbeil, Cassandra Grützner, and Philipp Schreck. 2024. Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior 160 (Nov. 2024), 108352. doi:10.1016/j.chb.2024.108352 [11]Jennifer M. Logg, Julia A. Minson, and Don A. Moore. 2019. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes 151 (March 2019), 90–103. doi:10.1016/j.obhdp.2018.12.005 [12] Annamaria Lusardi and Olivia S. Mitchell. 2011. Financial literacy around the world: an overview. Journal of Pension Economics and Finance 10, 4 (Oct. 2011), 497–508. doi:10.1017/S1474747211000448 [13]Annamaria Lusardi and Olivia S. Mitchell. 2014. The Economic Importance of Financial Literacy: Theory and Evidence. Journal of Economic Literature 52, 1 (March 2014), 5–44. doi:10.1257/jel.52.1.5 [14]Hasan Mahmud, A. K. M. Najmul Islam, Syed Ishtiaque Ahmed, and Kari Smolander. 2022. What influences algorithmic decision-making? A systematic literature review on algorithm aversion. Technological Forecasting and Social Change 175 (Feb. 2022), 121390. doi:10.1016/j.techfore.2021.121390 [15]Melanie J. McGrath, Oliver Lack, James Tisch, and Andreas Duenser. 2025. Measuring trust in artificial intelligence: validation of an established scale and its short form. Frontiers in Artificial Intelligence 8 (May 2025), 1582880. doi:10.3389/frai.2025.1582880 [16]Tae-Young Pak. 2026. How individuals use generative AI for personal financial management. Journal of Behavioral and Experimental Finance 49 (March 2026), 101145. doi:10.1016/j.jbef.2026.101145 [17]Olivia Pal, Veda Duddu, Agam Goyal, Drishti Goel, and Koustuv Saha. 2026. Do We Know What They Know We Know? Calibrating Student Trust in AI and Human Responses through Mutual Theory of Mind. In Proceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems. 1–8. [18]Muhammad Raees, Vassilis-Javed Khan, Ioanna Lykourentzou, and Konstantinos Papangelis. 2026. Do People Appropriately Rely on AI-Advice? An Analytical Review of HCI Research on Human-AI Decision-Making. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). Association for Computing Machinery, New York, NY, USA, 1–24. doi:10.1145/3772318.3791467 [19] Muhammad Raees and Konstantinos Papangelis. 2026. Trust to Reliance: Measurement Constructs for Human-AI Appropriate Reliance. In Proceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA ’26). Association for Computing Machinery, New York, NY, USA, 1–7. doi:10.1145/3772363.3798835 [20]Emely Rosbach, Jonas Ammeling, Jonathan Ganz, Christof Albert Bertram, Thomas Conrad, Andreas Riener, and Marc Aubreville. 2026. Stuck on Suggestions: Automation Bias, the Anchoring Effect, and the Factors That Shape Them in Computational Pathology. Machine Learning for Biomedical Imaging 2026, MELBA–BVM 2025 Special Issue (March 2026), 126–147. doi:10.59275/j.melba.2026-87b1 [21] Koustuv Saha, Yoshee Jain, Chunyu Liu, Sidharth Kaliappan, and Ravi Karkar. 2025. AI vs. Humans for Online Support: Comparing the Language of Responses from LLMs and Online Communities of Alzheimer’s Disease. ACM Transactions on Computing for Healthcare (2025). [22]Koustuv Saha, Yoshee Jain, Violeta J Rodriguez, and Munmun De Choudhury. 2026. Linguistic comparison of AI-and human-written responses to online mental health queries. NPJ artificial intelligence (2026). [23]Aaron Schecter, Eric Bogert, and Nina Lauharatanahirun. 2023. Algorithmic appreciation or aversion? The moderating effects of uncertainty on algorithmic decision making. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (CHI EA ’23). Association for Computing Machinery, New York, NY, USA, 1–8. doi:10.1145/3544549.3585908 [24]Yeganeh Shahsavar and Avishek Choudhury. 2023. User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study. JMIR Human Factors 10, 1 (May 2023), e47564. doi:10.2196/47564 [25]Sundar, S. Shyam. 2008. The MAIN Model: A Heuristic Approach to Understanding Technology Effects on Credibility. In Digital Media, Youth, and Credibility. The MIT Press, Cambridge, MA, 73–100. doi:10.1162/dmal.9780262562324.073 [26] See Heng Yim, Dong Whi Yoo, Apostolos Polymerou, Yuqi Liu, and Koustuv Saha. 2026. Generative ai for eating disorders: linguistic comparison with online support and qualitative analysis of harms. International Journal of Eating Disorders 59, 3 (2026), 519–533. [27] Jiawei Zhou, Kritika Venkatachalam, Minje Choi, Koustuv Saha, and Munmun De Choudhury. 2026. Communication styles and reader preferences of LLM-and human-authored COVID-19 information explanations: a case study. BMC Artificial Intelligence 2, 1 (2026), 10. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice7 Table 1. Participant demographics and background characteristics. Continuous variables are reported as mean (SD); categorical variables are reported as 푛 (%). CharacteristicFinal analytic sampleMeasure source Sample size285— Age39.76 (SD = 12.22)— Gender Male: 143 (50.2%); Female: 137 (48.1%); Prefer not to say: 3 (1.1%); Missing: 2 (0.7%). — Household incomeMedian category: $50,000–$74,999.— Financial literacy (0–3)2.66 (SD = 0.68)"Big Three" [12, 13] Self-rated financial literacy Very low: 2 (0.7%); Low: 20 (7.0%); Moderate: 154 (54.0%); High: 90 (31.6%); Very high: 19 (6.7%). — Financial risk tolerance (1–4)No financial risks [1]: 36 (12.6%); Average risks [2]: 161 (56.5%); Above-average risks [3]: 77 (27.0%); Sub- stantial risks [4]: 11 (3.9%). SCF item [6] AI use frequency Never: 9 (3.2%); Rarely: 20 (7.0%); Occasionally: 42 (14.7%); Regularly: 80 (28.1%); Frequently: 134 (47.0%). — AI use for personal financeNo, and have not considered it: 53 (18.6%); No, but have considered it: 43 (15.1%); Yes, once or twice: 102 (35.8%); Yes, multiple times: 87 (30.5%). — Used AI for finance at least once 189 (66.3%)— General trust in AI (1–7)4.63 (SD = 1.50)S-TIAS [15] Perceived AI literacy (1–7)5.16 (SD = 0.99)PAILQ-6 [7] Table 2. Scenario-level dimensions in the 2× 2× 2 vignette design. DimensionDefinitionLow levelHigh level StakesMagnitude and long-term financial consequences of the decision. Smaller or more reversible decisions, such as travel insurance or short-term spending choices. Substantial or difficult-to- reverse decisions, such as graduate school, debt re- payment, or concentrated investment risk. External un- certainty Whether outcomes depend on un- predictable future conditions. Outcomes are relatively stable or calculable, such as fixed interest rates or known rent differences. Outcomes depend on exter- nal events, such as market volatility, weather disrup- tions, or future job-market conditions. VerifiabilityWhether advice quality can be eval- uated against established financial principles. Preference-sensitive deci- sions with no single dom- inant rule, such as hous- ing lifestyle trade-offs or graduate-school ROI. Decisionswithbroad expert consensus, such as paying high-interest debt ordiversifyingconcen- trated assets. 8Kapadia et al. Advice Design 1 Experimental Design 2 Participant Evaluation 3 Financial Scenarios Common Decision Backbone Three Advice Styles AI Advice Expert AdviceOC Advice Facts, calculations, principles, recommendation direction Neutral Analytical Conditional Professional Structured Principle-driven Conversational Anecdotal Heuristic-driven Varied across style: voice, tone, framing, structure, and reasoning presentation Decision Context Stakes High External Uncertainty High Verifiability High All 2×2×2 combinations represented (8 scenarios) Low Advice Style × Source Attribution Underlying Advice Style Displayed Source Label Correct label Mislabeled Unlabeled Advice styles counterbalanced across scenarios using 3 versions TASK 1: Evaluate 4 scenario-advice pairs (randomized order) Risk & Safety Perception Source authority & behavioral intentions 10 items on 7-point Likert scale TASK 2: Compare all 3 advice styles for 1 new scenario Rank all 3 advice styles for given scenario Select ranking criteria (accurate, relevant, safe, source attribution) Presentation & Comprehension 1 2 3 DemographicsFinancial LiteracyRisk ToleranceAI Literacy & UseGeneral AI TrustScenario Familiarity Participant Characteristics and Covariates Trustworthy FinAInce: Study Design Overview Fig. 1. Overview of the study design. Eight financial scenarios represented all combinations of stakes, external uncertainty, and verifiability. A common decision backbone was rendered as AI, Expert, and OC (Online community) advice while holding the underlying facts, calculations, financial principles, and recommendation direction constant. Participants evaluated advice under correctly labeled, unlabeled, or mislabeled source-attribution conditions. They first rated four scenario–advice pairs and then ranked all three advice styles for one additional scenario. Advice styles were counterbalanced across scenarios, and participant-level characteristics and covariates were measured. Note that participants also completed Task 2, but the present analyses focus exclusively on Task 1. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice9 Mislabeling Schemas for Task 1 Schema A Advice StyleDisplayed Label AI-style Expert-style OC-style OC label AI label Expert label Schema B Advice StyleDisplayed Label AI-style Expert-style OC-style Expert label OC label AI label Fig. 2. Mislabeling schemas used in Task 1. Participants in the mislabeled arm were assigned to one of two counterbalanced label-swap schemas. In Schema A, AI, Expert, and OC advice were displayed with OC, AI, and Expert labels, respectively. In Schema B, they were displayed with Expert, OC, and AI labels, respectively. Thus, no advice style was paired with its corresponding source label in the mislabeled condition. Only Task 1 contributes to the analyses reported in this paper. Table 3. Operationalization of source-specific communication styles. FeatureAI adviceExpert adviceOC advice VoiceImpersonal; no first-person identity Mild first-person professional authority First-person peer perspective ToneNeutral and analyticalProfessional and measuredInformal and conversational ReasoningExplicit analysis and conditional if–then reasoning Structured, principle-based reasoning oriented toward long-term planning Experiential reasoning supported by practical heuristics FramingBalances available options and evaluates financial risk exposure Connects the decision to financial foundations and broader goals Uses lived experience, subjective judgment, and relatable consequences StructureSituation and trade-off; analysis; conditional recommendation Quantified framing; trade-off explanation; actionable principle Anecdote or position; experiential lesson; direct recommendation Excluded by design Personal anecdotes, emotional validation, and unsupported assumptions Personal anecdotes, slang, and emotional storytelling Formal advisory authority and technical optimization language Note. We define advice style as a controlled bundle of message-level features involving voice, tone, framing, discourse structure, and reasoning presen- tation. Style was varied separately from substantive financial content: within each scenario, the financial facts, numerical values, relevant principles, recommendation direction, and core reasoning were held constant across versions. The three versions were constructed using source-specific drafting and prompting instructions and were manually reviewed for information parity and comparable length. The excluded features represent experimental constraints rather than claims that real-world sources never exhibit these characteristics. The participant-facing labels were “AI Financial Assistant,” “Certified Financial Planner,” and “Online Community Forum.” 10Kapadia et al. Table 4. Measurement items used to evaluate financial advice. No. Measurement Item (7-point Likert scale)Dimension Presentation and Comprehension 1The advice is easy to read and follow.Readability 2The advice provides clear reasoning for its recommendation.Reasoning clarity 3The advice fits well with X’s situation.Situational fit Risk and Safety Perception 4The advice acknowledges potential risks and uncertainties.Risk acknowledgment 5Following this advice could put X at financial risk.Financial harm risk† 6Someone with limited financial knowledge could misunderstand this advice. Misleadingness† Source Authority and Behavioral Intentions 7The source of this advice appears knowledgeable about financial decisions. Perceived source knowledge 8Overall, this is high-quality financial advice for X.Perceived overall quality 9I would trust this advice if I were in X’s situation.Trust intention 10If I were in X’s situation, I would feel comfortable following this advice.Reliance intention Note. X denotes the protagonist named in each financial scenario. Items marked with dagger (†) are negatively valenced: higher ratings indicate greater perceived financial harm or misleadingness. Table 5. Model specifications used in the analysis. AnalysisModel specification Overall advice-style differences푌 ∼ 퐶(advice_style)+퐶(scenario)+ familiarity+(1 | participant) Advice style× label arm푌 ∼ 퐶(advice_style)×퐶(label_arm)+퐶(scenario)+familiarity+(1| participant) Note.푌denotes each advice-evaluation outcome. All models include participant-level random intercepts to account for repeated ratings from the same participant. Scenario fixed effects control for differences across financial-advice scenarios, and familiarity denotes participants’ self-reported familiarity with the scenario. The overall advice-style model estimates average differences among AI, Expert, and OC advice across labeling conditions; corresponding results are reported in Table 6. The advice style×label arm model estimates whether these differences vary across the unlabeled, labeled, and mislabeled source-attribution conditions; corresponding results are reported in Table 7. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice11 Table 6. Overall differences by advice style across advice-evaluation outcomes. Presentation & ComprehensionRisk & Safety PerceptionSource Authority & Behavioral Intentions Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Expert vs. AI +0.410*** (0.47) +0.255*** (0.30) +0.322*** (0.39) -0.107 (-0.10) -0.334*** (-0.27) -0.403*** (-0.30) +0.263*** (0.27) +0.182** (0.20) +0.238** (0.24) +0.233** (0.23) OC vs. AI +0.468*** (0.53) +0.141* (0.17) +0.121 (0.15) -0.183* (-0.17) -0.049 (-0.04) -0.343*** (-0.26) +0.016 (0.02) -0.173* (-0.19) +0.048 (0.05) +0.067 (0.07) Expert vs. OC -0.059 (-0.07) +0.114 (0.13) +0.201** (0.24) +0.076 (0.07) -0.285** (-0.23) -0.060 (-0.04) +0.248*** (0.26) +0.355*** (0.39) +0.190** (0.19) +0.166* (0.17) 푅 2 푚 0.2070.1540.1130.0970.1010.0870.1050.0680.1240.155 푅 2 푐 0.3990.3740.3810.3520.4600.4050.3890.3140.3690.376 Note. This table reports unstandardized mixed-effects regression coefficients훽with standardized effect sizes푑in parentheses. Models estimate overall differences by advice style, averaged across labeling conditions, and include scenario fixed effects, scenario familiarity as a covariate, and participant-level random intercepts. AI is the reference category for Expert vs. AI and OC vs. AI; Expert vs. OC contrasts were computed from the fitted model covariance matrix.푁=285 participants and 1,140 vignette-level observations.푅 2 푚 and푅 2 푐 denote marginal and conditional Nakagawa-style 푅 2 , respectively. † Lower scores indicate more favorable evaluations for these outcomes. Positive coefficients indicate that the first-named advice style was rated higher than the second-named style; for daggered outcomes, negative coefficients indicate more favorable ratings for the first-named source. ∗ 푝< .05, ∗ 푝< .01, ∗ 푝< .001. 12Kapadia et al. Table 7. Advice style and label-arm effects on advice evaluations. Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Advice-style effects in the unlabeled arm Expert vs. AI +0.383*** (0.43) +0.257* (0.30) +0.496*** (0.60) +0.010 (0.01) -0.340* (-0.27) -0.436* (-0.32) +0.338** (0.35) +0.154 (0.17) +0.283* (0.29) +0.328* (0.33) OC vs. AI +0.379** (0.43) +0.111 (0.13) +0.232* (0.28) +0.087 (0.08) -0.113 (-0.09) -0.359* (-0.27) +0.205 (0.21) -0.315** (-0.35) +0.097 (0.10) +0.071 (0.07) Label-arm effects for AI advice Labeled vs. Unlabeled +0.102 (0.12) +0.076 (0.09) +0.328* (0.40) +0.106 (0.10) -0.216 (-0.17) -0.107 (-0.08) +0.301 (0.31) +0.037 (0.04) +0.142 (0.14) +0.229 (0.23) Mislabeled vs. Unlabeled +0.058 (0.07) +0.073 (0.09) +0.352** (0.42) +0.127 (0.12) -0.338 (-0.27) -0.322 (-0.24) +0.402* (0.42) +0.065 (0.07) +0.280 (0.28) +0.295 (0.29) Advice style× label arm interactions Labeled× Expert +0.096 (0.11) +0.074 (0.09) -0.221 (-0.27) -0.146 (-0.14) -0.031 (-0.03) -0.061 (-0.05) +0.056 (0.06) +0.171 (0.19) +0.055 (0.06) -0.021 (-0.02) Labeled× OC +0.160 (0.18) +0.055 (0.06) -0.114 (-0.14) -0.362 (-0.35) +0.194 (0.16) -0.042 (-0.03) -0.216 (-0.22) +0.198 (0.22) +0.038 (0.04) +0.042 (0.04) Mislabeled× Expert -0.018 (-0.02) -0.083 (-0.10) -0.300* (-0.36) -0.205 (-0.20) +0.051 (0.04) +0.162 (0.12) -0.282 (-0.29) -0.091 (-0.10) -0.191 (-0.19) -0.266 (-0.27) Mislabeled× OC +0.092 (0.11) +0.027 (0.03) -0.224 (-0.27) -0.430* (-0.41) -0.016 (-0.01) +0.093 (0.07) -0.348 (-0.36) +0.218 (0.24) -0.194 (-0.20) -0.067 (-0.07) 푅 2 푚 0.2110.1560.1230.1000.1070.0900.1170.0750.1290.162 푅 2 푐 0.3990.3750.3840.3570.4630.4060.3950.3190.3710.379 Note. This table reports unstandardized mixed-effects regression coefficients훽with standardized effect sizes푑in parentheses. Models estimate advice style, label arm, and their interaction, with scenario fixed effects, scenario familiarity as a covariate, and participant-level random intercepts. The reference condition is AI advice in the unlabeled arm. Thus, Expert vs. AI and OC vs. AI estimate advice style differences in the unlabeled arm; Labeled and Mislabeled coefficients estimate label-arm effects for AI-style advice; interaction terms indicate whether advice style differences change under labeled or mislabeled conditions relative to the unlabeled arm.푁=285 participants and 1,140 vignette-level observations.푅 2 푚 and푅 2 푐 denote marginal and conditional Nakagawa-style푅 2 , respectively. † Lower scores indicate more favorable evaluations for these outcomes. Positive coefficients indicate higher ratings for the first-named condition; for daggered outcomes, negative coefficients indicate more favorable ratings for the first-named condition. ∗ 푝< .05, ∗ 푝< .01, ∗ 푝< .001. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice13 Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Omnibus (KW H) AI Label vs. No Label Expert Label vs. No Label OC Label vs. No Label Expert Label vs. AI Label OC Label vs. AI Label OC Label vs. Expert Label 2.7 +0.14 t=0.99 +0.07 t=0.40 +0.28 t=1.69 -0.07 t=-0.42 +0.13 t=0.78 +0.22 t=1.05 3.1 +0.11 t=0.77 +0.19 t=1.10 +0.26 t=1.50 +0.07 t=0.41 +0.13 t=0.78 +0.07 t=0.33 9.6* +0.34 t=2.37 +0.47 t=2.93* +0.38 t=2.19 +0.11 t=0.68 +0.04 t=0.24 -0.07 t=-0.33 2.2 +0.07 t=0.47 +0.17 t=0.98 -0.00 t=-0.02 +0.11 t=0.62 -0.07 t=-0.39 -0.18 t=-0.86 6.0 -0.23 t=-1.60 -0.37 t=-2.12 -0.29 t=-1.61 -0.12 t=-0.73 -0.05 t=-0.30 +0.07 t=0.36 4.2 -0.14 t=-1.01 -0.22 t=-1.21 -0.33 t=-1.72 -0.06 t=-0.36 -0.17 t=-0.94 -0.12 t=-0.56 9.4* +0.26 t=1.80 +0.47 t=2.89* +0.44 t=2.56 +0.19 t=1.20 +0.17 t=1.01 -0.02 t=-0.08 2.3 +0.06 t=0.41 +0.25 t=1.41 +0.07 t=0.34 +0.16 t=0.97 +0.01 t=0.03 -0.16 t=-0.75 5.0 +0.13 t=0.90 +0.39 t=2.40 +0.28 t=1.55 +0.25 t=1.60 +0.15 t=0.84 -0.10 t=-0.46 6.2 +0.23 t=1.58 +0.42 t=2.57 +0.35 t=2.06 +0.19 t=1.15 +0.13 t=0.74 -0.06 t=-0.30 AI Advice (displayed label varied) Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Omnibus (KW H) AI Label vs. No Label Expert Label vs. No Label OC Label vs. No Label Expert Label vs. AI Label OC Label vs. AI Label OC Label vs. Expert Label 4.1 +0.26 t=1.44 +0.17 t=1.16 +0.04 t=0.26 -0.09 t=-0.51 -0.24 t=-1.16 -0.14 t=-0.82 1.6 +0.07 t=0.40 +0.11 t=0.78 +0.09 t=0.51 +0.04 t=0.22 +0.01 t=0.03 -0.04 t=-0.22 2.4 +0.21 t=1.22 +0.11 t=0.76 +0.13 t=0.73 -0.08 t=-0.48 -0.08 t=-0.38 +0.01 t=0.05 1.5 -0.23 t=-1.18 -0.09 t=-0.62 +0.03 t=0.18 +0.13 t=0.72 +0.26 t=1.24 +0.12 t=0.73 2.8 -0.34 t=-2.05 -0.11 t=-0.74 -0.23 t=-1.38 +0.22 t=1.38 +0.11 t=0.54 -0.12 t=-0.73 2.8 -0.27 t=-1.44 -0.08 t=-0.56 -0.08 t=-0.49 +0.19 t=1.04 +0.20 t=0.96 +0.00 t=0.00 5.5 +0.26 t=1.52 +0.30 t=2.09 +0.20 t=1.20 +0.06 t=0.33 -0.06 t=-0.28 -0.11 t=-0.63 4.4 -0.02 t=-0.10 +0.24 t=1.65 +0.10 t=0.58 +0.27 t=1.54 +0.13 t=0.61 -0.14 t=-0.84 2.0 +0.22 t=1.22 +0.12 t=0.84 +0.17 t=1.03 -0.10 t=-0.55 -0.05 t=-0.24 +0.05 t=0.31 1.6 +0.11 t=0.58 +0.08 t=0.57 +0.19 t=1.09 -0.02 t=-0.13 +0.07 t=0.35 +0.10 t=0.58 Expert Advice (displayed label varied) Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Omnibus (KW H) AI Label vs. No Label Expert Label vs. No Label OC Label vs. No Label Expert Label vs. AI Label OC Label vs. AI Label OC Label vs. Expert Label 11.7** +0.01 t=0.07 +0.42 t=2.38 +0.29 t=2.02 +0.47 t=2.24 +0.32 t=1.82 -0.18 t=-0.93 2.5 +0.05 t=0.28 +0.28 t=1.65 +0.16 t=1.09 +0.26 t=1.25 +0.12 t=0.66 -0.15 t=-0.84 3.0 +0.06 t=0.35 +0.33 t=2.00 +0.23 t=1.60 +0.28 t=1.35 +0.17 t=0.91 -0.12 t=-0.71 5.0 -0.29 t=-1.57 -0.34 t=-1.72 -0.22 t=-1.55 -0.07 t=-0.33 +0.07 t=0.38 +0.14 t=0.70 5.3 -0.35 t=-2.17 -0.23 t=-1.35 -0.01 t=-0.06 +0.13 t=0.61 +0.38 t=2.27 +0.24 t=1.38 2.1 -0.23 t=-1.39 -0.17 t=-0.92 -0.14 t=-0.96 +0.05 t=0.24 +0.10 t=0.57 +0.04 t=0.22 0.8 +0.11 t=0.67 +0.17 t=0.95 +0.14 t=0.99 +0.06 t=0.31 +0.02 t=0.11 -0.05 t=-0.26 4.8 +0.19 t=1.17 +0.35 t=2.01 +0.23 t=1.61 +0.19 t=0.90 +0.03 t=0.18 -0.17 t=-0.87 1.3 +0.13 t=0.75 +0.18 t=0.98 +0.20 t=1.39 +0.06 t=0.30 +0.08 t=0.43 +0.01 t=0.03 4.6 +0.29 t=1.80 +0.30 t=1.69 +0.28 t=1.95 +0.04 t=0.21 -0.01 t=-0.08 -0.06 t=-0.29 OC Advice (displayed label varied) −0.60 −0.30 0 +0.30 +0.60 Cohen's d Label Sensitivity by Advice Style Fig. 3. Label sensitivity by advice style. Each panel holds the underlying advice style constant and varies only the displayed source label. Cell color indicates Cohen’s푑for each displayed-label comparison, with positive values indicating that the first-named label condition received higher ratings than the second. Cell text reports Cohen’s푑and the Welch independent-samples푡statistic. Pairwise significance stars are based on Holm-adjusted푝-values within each outcome ( ∗ 푝< .05, ∗ 푝< .01, ∗ 푝< .001). The omnibus column reports Kruskal–Wallis퐻tests across displayed-label conditions. Ratings were first aggregated to the participant×advice-style× displayed-label level. Daggered outcomes (†) indicate measures for which lower scores are more favorable. 14Kapadia et al. Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Omnibus (KW H) Expert vs.AI OC vs.AI OC vs. Expert 7.1* +0.49 t=3.14** +0.29 t=1.89 -0.29 t=-1.41 2.1 +0.29 t=1.76 +0.11 t=0.66 -0.22 t=-1.05 2.9 +0.34 t=2.24 -0.04 t=-0.23 -0.40 t=-1.97 2.1 -0.19 t=-0.99 -0.23 t=-1.28 -0.04 t=-0.21 4.0 -0.39 t=-2.46* -0.25 t=-1.57 +0.17 t=0.80 8.3* -0.47 t=-2.59* -0.34 t=-2.02 +0.15 t=0.71 1.8 +0.31 t=1.96 +0.04 t=0.21 -0.30 t=-1.45 0.6 +0.11 t=0.69 -0.10 t=-0.57 -0.24 t=-1.16 5.1 +0.37 t=2.21 +0.09 t=0.55 -0.30 t=-1.47 2.1 +0.24 t=1.37 +0.14 t=0.90 -0.12 t=-0.56 Displayed AI Label (advice style varied) Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Omnibus (KW H) Expert vs.AI OC vs.AI OC vs. Expert 18.9*** +0.57 t=2.90** +0.74 t=3.62** +0.26 t=1.45 3.7 +0.29 t=1.60 +0.29 t=1.41 -0.00 t=-0.02 2.2 +0.18 t=1.06 +0.15 t=0.71 -0.04 t=-0.23 2.9 -0.16 t=-0.89 -0.37 t=-1.77 -0.24 t=-1.23 0.3 -0.05 t=-0.27 -0.01 t=-0.06 +0.03 t=0.20 1.9 -0.24 t=-1.36 -0.21 t=-1.01 +0.01 t=0.03 2.8 +0.19 t=1.09 -0.09 t=-0.44 -0.26 t=-1.35 1.9 +0.20 t=1.11 -0.07 t=-0.36 -0.26 t=-1.34 0.5 +0.04 t=0.26 -0.09 t=-0.44 -0.13 t=-0.67 0.5 +0.04 t=0.26 -0.00 t=-0.00 -0.04 t=-0.21 Displayed Expert Label (advice style varied) Read- ability Reasoning Clarity Situational Fit Risk Ack. Fin. Harm Risk † Risk of Being Misled † Overall Quality Source Knowledge Trust Intention Reliance Intention Omnibus (KW H) Expert vs.AI OC vs.AI OC vs. Expert 8.7* +0.27 t=1.28 +0.51 t=2.64* +0.26 t=1.54 0.5 +0.22 t=1.03 +0.09 t=0.49 -0.13 t=-0.75 1.0 +0.25 t=1.21 +0.08 t=0.44 -0.18 t=-1.08 2.1 +0.15 t=0.70 -0.10 t=-0.53 -0.24 t=-1.46 6.1* -0.25 t=-1.21 +0.15 t=0.85 +0.41 t=2.44* 0.2 -0.11 t=-0.55 -0.06 t=-0.32 +0.05 t=0.31 3.0 +0.10 t=0.45 -0.14 t=-0.75 -0.25 t=-1.43 3.3 +0.21 t=1.02 -0.08 t=-0.41 -0.32 t=-1.90 1.2 +0.19 t=0.89 +0.01 t=0.03 -0.20 t=-1.14 2.0 +0.21 t=1.00 +0.00 t=0.01 -0.21 t=-1.25 Displayed OC Label (advice style varied) −0.60 −0.30 0 +0.30 +0.60 Cohen's d Advice-Style Sensitivity Under Each Displayed Source Label Fig. 4. Advice-style sensitivity under each displayed source label. Each panel holds the displayed source label constant and varies the underlying advice style. Cell color indicates Cohen’s푑for each advice-style comparison, with positive values indicating that the first-named advice style received higher ratings than the second. Cell text reports Cohen’s푑and the Welch independent-samples푡 statistic. Pairwise significance stars are based on Holm-adjusted푝-values within each outcome ( ∗ 푝< .05, ∗ 푝< .01, ∗ 푝< .001). The omnibus column reports Kruskal–Wallis퐻tests across advice-style conditions. Ratings were first aggregated to the participant× displayed-label× advice-style level. Daggered outcomes (†) indicate measures for which lower scores are more favorable. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice15 Sample Scenario: Credit Card Bill or New Laptop? Mark is a 26-year-old professional earning about $6,500 per month after taxes. He currently has an $800 balance on a credit card with an annual interest rate of 22%. Mark recently received a $1,000 bonus. Around the same time, his laptop started having issues, and he has been considering replacing it with a newer model that would cost about $1,000. He could continue using his current laptop for a while, but it has been slowing him down at work and causing frustration. For now, this $1,000 is the only flexible money he feels comfortable using. Mark is deciding whether to use it now to replace the laptop or to pay off the credit card balance first and delay the purchase. The credit card interest rate is fixed, and there are no penalties for paying it off early. Mark asks: “Should I replace the laptop now, or pay off the credit card balance first?” Low stakesLow external uncertaintyHigh verifiability AI Financial Assistant (161 words) You earn $6,500 per month after taxes and have an $800 credit card balance at 22% interest. You also have a $1,000 bonus and are considering a $1,000 laptop. The trade-off is between eliminating an $800 balance at a 22% fixed rate or using the full $1,000 for a replacement device while carrying that debt. A 22% interest rate is high-cost debt. Paying it off provides a guaranteed return equal to the interest avoided. With no penalty for early repayment, removing the $800 balance immediately stops further interest accumulation and improves monthly cash flow. Your income appears sufficient relative to the balance, but carrying revolving debt at that rate increases risk and reduces flexibility. High-interest credit is generally treated as a priority before discretionary purchases. If your laptop is still functional, then paying off the credit card first is financially efficient. If it becomes unusable for work, then replacement may take priority. Prioritizing debt elimination supports more sustainable cash flow management. 1 Impersonal voice The advice directly addresses the protagonist without first-person identity or personal narrative. 2 Neutral, analytical tone The language is measured and non- emotional, focusing on facts rather than feelings. 3 Explicit financial reasoning The advice states the governing financial principle clearly (e.g., high-cost debt and the interest avoided). 4 Risk / flexibility framing Consequences are framed in terms of financial stability and flexibility (risk and reduced flexibility). 5 Conditional recommendation The recommendation is expressed through explicit if-then logic based on the laptop's condition. 6Constrained style The response avoids personal anecdotes, storytelling, emotional validation, and unsupported new assumptions. Advice Skeleton Facts $800 debt at 22% APR $1,000 bonus $1,000 laptop current laptop still usable No early repayment penalty Trade-off Immediate replacement benefit vs. continued high- interest debt Prioritize high-interest debt before discretionary purchases because repayment provides a guaranteed financial benefit. Pay the card first; save for the laptop afterward Exception: replace it sooner if it becomes unusable for work Financial Principle Recommendation We transformed the shared backbone through manual drafting, few-shot prompting, and constraint-based rewriting, followed by iterative review for information parity and comparable length. Fig. 5. Controlled construction of source-specific advice styles. A shared decision skeleton preserves the financial facts, governing principle, trade-off, recommendation direction, and boundary condition across all three responses. The same substantive backbone is rendered using the source-specific communication characteristics of an AI Financial Assistant, a Certified Financial Planner, and an Online Community Forum. Highlighted passages and numbered annotations identify the voice, tone, framing, reasoning presentation, and discourse structure operationalized in each condition. Page one presents the financial scenario, common advice skeleton, and annotated AI Financial Assistant response. 16Kapadia et al. Certified Financial Planner (158 words) At a 22% interest rate, your credit card balance is expensive debt. Even though $800 may seem manageable, that rate means you're paying a high cost to carry the balance each month it remains unpaid. I generally treat debt at that level as a priority because paying it off delivers a guaranteed benefit. Using your $1,000 bonus to eliminate the $800 balance immediately removes that interest expense and improves your monthly cash flow. Once the card is paid off, you can begin setting aside money for the laptop without interest working against you. Going forward, I'd aim to pay the full statement balance each month and build spending limits into your budget, so the card functions as a payment tool rather than a financing source. That approach reduces reliance on irregular income, like bonuses, to reset balances and helps keep your long-term savings on track. Small structural habits like this tend to make future financial decisions much easier. 1Quantified professional framing The response opens by identifying the key financial variable and framing the problem through its cost implications. 2Mild first-person authority The advice uses a restrained first-person planner voice to signal expertise without becoming personal or anecdotal. 3 Principle-based financial reasoning The recommendation is justified through clear financial logic linking debt reduction to improved cash flow. 4 Actionable guidance The response offers a practical next step that translates the recommendation into a concrete plan. 5 Long-term planning orientation The advice emphasizes sustainable habits, budgeting discipline, and alignment with longer-term financial goals. 6 Constrained expert style The response avoids personal anecdotes, slang, and emotional storytelling while maintaining a measured professional tone. Online Community Forum (146 words) Pay off the credit card. Full stop. 22% interest is brutal. A new laptop is really tempting, I understand. And trust me, I've been there. As soon as I started my first full-time job, every minor inconvenience felt like an excuse to upgrade. I kept telling myself I'd pay the card off soon, and that "small" balance with heavy interest ended up costing me way more than the upgrades were worth. Unless your laptop is completely unusable, upgrading while carrying 22% debt doesn't make sense. You've got $1,000. Wipe out the $800 balance and be done with it. After that, start putting money aside for the laptop and buy it without dragging high-interest debt behind you. Clearing that balance first just feels cleaner. It will give you back control, and I didn't realize how much it was stressing me out until it was finally paid off. 1Direct, emphatic opening The advice opens with a blunt recommendation and strong informal emphasis rather than measured professional framing. 2 Informal, conversational tone The language is casual and relatable, directly acknowledging temptation and speaking like a peer. 3First-person personal experience The response draws on lived experience and peer perspective instead of neutral or expert distance. 4 Anecdotal reasoning The core justification is supported through a personal story and lesson learned rather than formal financial analysis. 5 Direct practical recommendation The advice gives a straightforward action path in simple everyday language. 6 Experiential and emotional closing The response closes with subjective judgment and emotional relief, using relatable peer language instead of a formal advisory tone. Fig. 5. Controlled construction of source-specific advice styles (continued). Annotated Certified Financial Planner and Online Community Forum responses generated from the same substantive advice skeleton.