Paper deep dive
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
Michael Correll, Lucy Havens, Mahsan Nourani
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/17/2026, 3:05:08 AM
Summary
This paper investigates how 'trust junk'—superfluous or irrelevant but accurate data in explainable AI (XAI) visualizations—can lead to unjustified trust and support for highly discriminatory predictive models. Through a crowdsourced experiment, the authors demonstrate that increasing the volume of explanatory information, even when it lacks meaningful connection to model fairness or performance, significantly increases user agreement, trust, and satisfaction while reducing perceptions of bias. The findings highlight the rhetorical power of data visualizations and warn XAI designers against inadvertently 'fairwashing' biased models.
Entities (8)
Relation Signals (5)
Michael Correll → affiliatedwith → Northeastern University
confidence 99% · Michael Correll * Northeastern University
Trust Junk → leadsto → Unjustified Trust
confidence 95% · providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model
Trust Junk → fairwashes → Discriminatory Models
confidence 92% · trust junk may therefore fairwash models— in other words, produce unearned and unwarranted assumptions of trust or fairness in models that are, in reality, biased and inaccurate.
Everything Condition → resultsin → Higher Model Agreement
confidence 90% · participants in the Everything condition followed the model significantly more often (M = 82.6%) than the other conditions (M = 69.5%)
XAI Techniques → exacerbates → Automation Bias
confidence 88% · XAI techniques can exacerbate automation bias [4, 18, 4, 29]—a cognitive bias characterized by an overreliance on automated systems.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers need to consider the implicit or explicit rhetorics of their work, and beware of the potential of visualizations to imbue models with unearned trust.
Tags
Links
- Source: https://arxiv.org/abs/2607.14152v1
- Canonical: https://arxiv.org/abs/2607.14152v1
Trouble viewing inline? Open PDF directly →
Full Text
32,563 characters extracted from source content.
Expand or collapse full text
“Trust Junk” Leads to Unjustified Support for Highly Discriminatory Predictive Models Michael Correll * Northeastern University Lucy Havens † Northeastern University Mahsan Nourani ‡ Northeastern University B C E F D A Figure 1: An annotated version of a stimulus from our experiment. A decision (C) made by an automated system on whether or not a potential candidate (B) is likely to pass the bar exam, along with explanatory information that serves as trust junk [38] (ADEF). The model is highly discriminatory (it predicts failure for candidates who identify as Black men, and predicts all other candidates will pass). Yet, providing increasing amounts of (largely fairness-irrelevant) explanatory information makes participants agree with the model more, and be more likely to rate the model as fair and unbiased. We empirically demonstrate that trust junk can mislead XAI viewers even when the explanatory information it provides is not, strictly speaking, incorrect. ABSTRACT The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers need to consider the implicit or explicit rhetorics of their work, and beware of the potential of visualizations to imbue models with unearned trust. Supplemental Material is available at https://osf.io/wufqz/. Index Terms: XAI, Data Rhetoric, Information Visualization * e-mail: m.correll@northeastern.edu † e-mail: l.havens@northeastern.edu ‡ e-mail: m.nourani@northeastern.edu 1 INTRODUCTION As articulated by Kennedy et al. [21], data and their presentations can produce the (false) impression of “objectivity” as well as “trans- parency, scientific-ness and facticity” through forms of implicit rhetorical work. Likewise, Peck et al. [32] find that mass audiences often have intrinsic trust in data visualizations even from sources they think of as biased in other contexts. Therefore, Drucker [8] argues we must beware of the “persuasive and seductive rhetorical force of visualization.” The unjustified authority of data, and data visualizations, is highly pertinent to the design of XAI. A particular danger in XAI is if machine learning (ML) system explanations that appear authoritative and complete “soothe” [39] viewers while fail- ing to provide information needed to meaningfully assess the sys- tem’s accuracy, fairness, or performance. Wall et al. [38] expand on this threat through their concept of trust junk, where an XAI vi- sualization “that has no meaningful connection with the underlying model or data is employed to enhance trust.” In other words, expla- nations with trust junk may include visualizations that persuade by being beautiful, complicated, or detailed without being useful. Con- cerningly, trust junk may therefore fairwash [3] models— in other words, produce unearned and unwarranted assumptions of trust or fairness in models that are, in reality, biased and inaccurate. arXiv:2607.14152v1 [cs.HC] 14 Jul 2026 In this paper, we report a crowdsourced study exploring the im- pact of trust junk in ML model explanations. We find that, even for an intentionally unfair model, increasing the amount of seemingly useful (but in actuality irrelevant) data in an explanation increases perceived user trust and agreement with the model, while also re- ducing perceptions of bias or unfairness. These results suggest that XAI designers must take accountability not only for the accuracy of the data used in explanations, but also for the persuasive force of the explanations. Even without an intent to deceive, XAI techniques can fall into “explainability pitfalls” when considered in their con- text of use and interpretation by their intended audiences [9]. 2 BACKGROUND The goals and techniques associated with XAI are vast (as demon- strated in numerous surveys [1, 2, 27, 15, 19, 16, 35]). However, determining how to create effectively human-centered explanations remains an open area of study. Researchers have investigated how to choose the appropriate level of detail [22, 26, 33], type of expla- nation [23], and whether to show explanations at all [28]. While there is growing work on intentionally adversarial XAI [24, 7, 34, 36] (e.g., gaming fairness scores or explanatory metrics to hide unfavorable model information) and “dark pat- terns” [9], we focus on a wider space of sociotechnical failures and the extent that explanations, even accurate ones, can work rhetori- cally to “fairwash” [3] biased models by imparting them with (per- haps unearned) authority. XAI techniques can exacerbate automa- tion bias [4, 18, 4, 29]—a cognitive bias characterized by an overre- liance on automated systems. Merely referring to an AI system as a “Statistical Model” instead of an “Artificial Intelligence” model can impact the perceived complexity and competency of the sys- tem [25]. Other effects are more subtle. Cabitza et al. [5] observe the “XAI halo effect,” where a high- or low-quality explanation may induce users to think of a model as being of similar quality as its explanation, even with insufficient evidence to support this judg- ment. Of particular concern to our work are “placebic” [12] and “empty” [39] explanations, which contain no useful information but can persuade by giving the appearance of being informative. Our study looks at an intentionally extreme case, measuring how in- creasing the amount of information in an explanation can persuade a viewer to trust a model more than they should. 3 MOTIVATING SCENARIO & TECHNIQUES Our user study design is based on controversial uses of AI systems to assign students’ grades on the International Baccalaureate (IB) and General Certificate of Education (GCE) Advanced (A) Level exams [11, 31]. We use the Law School Admissions Bar Passage dataset [40], which contains demographic, academic, and bar exam result information for 20,000 candidates who took the bar exam from 1991–1997, to generate an intentionally biased model. Our model predicts failure for Black male candidates and passes for all other candidates. Yet, due to class imbalances in exam outcomes across race and gender, the model has 93.7% accuracy. Wall et al. [38] claim that certain strategies such as communicat- ing provenance information, providing transparency, or increasing the amount of data views in an XAI visualization can foster trust. This trust can be engendered regardless of the actual capabilities of the model: performing too much trust-building is therefore akin to turning an “evil knob” too far, raising the risk of “fair-washing” [3] a biased or inaccurate model. Based on techniques described, but not tested, in Wall et al., as well as adversarial XAI work, we cre- ated “junk” explanations for our user study. The explanations are themed around four techniques that we find to be misleading, even as they present data that is correct. Burying in Details: We provided an excessive amount of infor- mation for our model’s binary classification task. Our inclusion of large amounts of complex-looking but ultimately irrelevant details was intended to overwhelm user study participants rather than pro- vide genuinely useful data. This technique, which metaphorically numbs the user into accepting their own inability to fully understand the data, is referred to by Correll [6] as a “novocaine chart.” Appeals to Authority: We provided superficially impressive in- formation, such as the name of the prestigious university where computer scientists developed the model (Figure 1, A), dataset size, and the model’s accuracy score. This strategy is especially salient when the audience lacks a baseline: in our case, the accu- racy of our model, at 93.7%, is lower than the accuracy of a model that would have predicted all candidates passed regardless of their background, which would be 94.8%. Our choice to describe the prediction model as an “AI system” is also meant to encourage the reader to ascribe undue complexity or accuracy [25] to our model that is, at heart, a glorified if statement. Cherry Picking: Many metrics reveal the unfair reliance of our model on gender and race, such as aχ 2 test and feature importance scores. We intentionally omitted these, instead providing metrics that represent our model’s performance favorably (Figure 1, D ). We also provided a cohort-based explanation by showing the three most similar candidates to the main candidate based on Euclidean distance and displayed their ground truth exam outcomes from the training data, rather than the model’s predictions (Figure 1, F). This allowed us to better hide that all Black men would be predicted to fail, even though only a fraction did so in reality. Encouraging Folk Algorithms: In the absence of knowledge of al- gorithmic internals, people often develop “folk algorithms”— sim- plified and often incorrect understandings about how an algorithm operates [42]. To encourage the creation of inaccurate folk algo- rithms, we provided information about all features in the training data (Figure 1, B,E, andF), implying that the model used all of them in its decision-making process, when it only used 2 (i.e., race and gender). Likewise, by showing that a candidate was partic- ularly high or low in certain attributes, our explanations encouraged participants to create inaccurate causal stories (which are particu- larly pernicious in visualizations of relationships) about why the model made a particular prediction [41]. 4 USER STUDY We conducted a between-subjects crowdsourced experiment on “trust junk” in XAI explanations, investigating whether this “junk” could persuade users that a biased model was fair and useful. Our main manipulation was to increase the number of explanatory com- ponents, where each component was technically correct but unhelp- ful for assessing model efficacy or fairness (see section 3). This ma- nipulation is what Wall et al. [38] refer to as a “knob” that designers can manipulate to impact trust in a model. While our stimuli are in- spired by real XAI visualizations (our confusion matrix is based on work by Gomez et al. [13] and our cohort-based explanations are based on the “C-N visual explanations” of Szymanski et al. [37]), the resulting XAI dashboards are our own design, embodying a va- riety of explanatory techniques from prior work. Each of our explanations had one of three levels of trust junk corresponding to one of three study conditions. Participants in the Baseline condition saw an explanation with basic model prove- nance information, a bar exam candidate’s profile, and the model’s decision for that candidate ( A,BandC). In the Model condi- tion, participants saw an explanation with model accuracy statistics (D ) in addition to all of the information in the Baseline condition. In Everything, participants saw an explanation that, in addition to all information in the prior two conditions, also includes histograms of the candidate’s feature scores (E) and short profiles of similar candidates in the training data ( F). The study was approved by our institutional review board and conducted on Prolific via Qualtrics. Participants’ were compen- sated $15/hour. Participants were randomly assigned to one of the three conditions. Their main task was to review eight pro- files of bar exam candidates, presented in random order, and judge whether each candidate would pass the bar. We used this task to measure participants’ agreement with the model as a reliance met- ric [30, 33]). Then, participants responded to post-study question- naires, adapted from previous work [14, 16], so we could assess their perceptions of and trust in the explanations and model predic- tions. After removing data from participants who failed our atten- tion check, our final sample included 28, 27, and 28 participants in the Baseline, Model, and Everything conditions, respectively. We hypothesized that increasing amounts of seemingly detailed (but ultimately distracting or at least incomplete) information will “fairwash” [3] our model by fostering unearned and unwarranted trust in it. More specifically, for participants in the Everything and Model conditions, relative to those in our Baseline condition, we hypothesized there will be (1) greater agreement with the model, (2) greater perception that the model is fair and trustworthy, and (3) greater perception that the model and explanation are useful. To understand participants’ rationale for their responses throughout our user study, and gain insight into their mental models and “folk algorithms” of our model, we analyzed their free text responses to questions around model performance and fairness using qualitative coding. We were interested in whether participants noticed that the model was discriminatory with respect to race or gender, or that the model explanations were incomplete. Additional study details including the full survey instrument, stimuli, analyses, participant demographics, and qualitative coding procedures are available in our supplement at https://osf.io/wufqz/. 5 RESULTS Here, we provide an overview of our main quantitative and qualita- tive results, broken down by measure. Model Agreement and Task Performance: Participants correctly predicted the ground truth label 5.3/8 times (66.6% of the time) and agreed with the model 5.9/8 times (73.9% of the time). Condition had a significant impact on agreement (F(2, 80) = 6.4, p = 0.0025): participants in the Everything condition followed the model sig- nificantly more often (M = 82.6%) than the other conditions (M = 69.5%) (see Figure 2a). Condition also had a significant impact on rate of over-reliance (F(2, 80) = 5.1, p = 0.008): participants in the Everything condition were more likely to erroneously follow the model (M = 89.3% of errors) compared to the other conditions (M = 75.3% of errors). Model Trust: Condition had a significant impact on the Likert scale trust rating (F(2, 80) = 4.7, p = 0.0.012). A post-hoc test found that participants in the Everything condition rated higher trust in the model (M = 25.3) than those in the Baseline condition (M = 16.8), neither of which were significantly different from the trust rating of those in the Model condition (M = 20.1). Figure 2b shows this result in more detail. Explanation Satisfaction: Condition had a significant impact on satisfaction with the explanation (F(2, 80) = 7.3, p = 0.001). A post-hoc test found that participants in the Everything condition rated their satisfaction significantly higher (M = 34.2) than those in the other two conditions (M = 26.4) (see Figure 2c). Fairness: We asked participants to rate their agreement (from 1: Strongly disagree to 7: Strongly agree) with four additional ques- tions around fairness (see Figure 3), which were modeled after Goyal et al. [14]. For questions around perceived gender biases, overall bias, and overall ethics, we found no significant difference among conditions (for gender: F(2, 80) = 0.74, p = 0.48; for bias in general: F(2, 80) = 1.9, p = 0.16; and for ethics in general: F(2, 80) = 2.1, p = 0.13). For perceived fairness with respect to race, we did find a significant effect of condition (F(2, 80) = 3.5, p = 0.035): those in the Baseline condition rated the algorithm as the least fair across race (M = 3.9), followed by those in the Every- thing (M = 4.8) condition and then the Model (M = 4.8) condition. The most troubling result regarding participants’ perceptions of model fairness is that the percentage of participants who rated our unfair model as fair—either in general or specifically with respect to race and gender—was highest in the two conditions with the most explanatory information. Equally concerning is that these two conditions had the lowest percentage of participants who correctly identified the unfairness of the model (see Figure 3). Qualitative Findings: Participants in the Baseline condition re- ported noticing more issues with race or gender in the model’s pre- dictions (12/28 = 42.9%) and gave more responses indicating that important information was missing (8/28 = 28.6%) compared to their counterparts in the Model and Everything conditions (see Figure 4). Still, responses from only three participants (one from each condition) indicated a partial understanding of the information most important in guiding the model’s predictions (e.g., P3 and P36 noted “race” and “family income.”). No response was fully correct. We noted 19 instances where participants explicitly described our model as fair (e.g., P8 said, “It’s fair in that it has no racial or gender bias.”). Still, participants occasionally expressed unease with the model. P76 said they were concerned with “what kind of darkness people might use it for.” P52 stated: “As a current law student I do not think this model is very fair at all. While it is true that things like the LSAT, GPAs, law school ranking, class ranking, etc. can be used to predict the likeliness of someone passing the bar, they are not perfect. I would say law school ranking and class ranking probably provide the best indicators since better schools have better bar passage rates as a fact, and better students tend to understand the subjects tested on the bar better, but someone can be at low tier school, ranked near the bottom of their class and still pass. A person is not just their data.” We noted 15 instances of unease about the model’s ability to holis- tically understand candidates. 6 DISCUSSION Our user study validates our central premise, that increasing amounts of nominally explanatory data can engender unwarranted trust in a predictive model. We presented participants with explana- tions of an unfair and superficial model using only race and gender to make a decision about academic success, even when much more informative features (such as GPA or class rank) were available. Regardless of their assigned condition, participants agreed with the model’s predictions the majority of the time and often rated the model’s quality highly. This result persists even for our Baseline condition where the participants had, essentially, no information about the model other than eight predictions and the fact that it was made by computer scientists. Consistent patterns of per-condition differences in participants’ perceptions of our model, despite limited information, suggest the potential impact of “trust junk” is large. Our Everything condi- tion’s explanation mostly reiterated information about the model’s training data. It and the Model condition’s explanation include largely contextless global accuracy information. Concerningly, the presence of this information, none of which provides insight on model internals or potential biases, resulted in increased rates of participant agreement with and trust in the model. Only a minority of participants described the model as unfair or biased. The few participants who did report concerns with the model were unable to accurately articulate why the model was flawed, and, in the absence of crucial model information, resorted to informal, often incorrect reasoning to explain why the model might be wrong or unfair. 0% 25% 50% 75% 100% BaselineModelEverything Human/Model Agreement 0 10 20 30 40 50 BaselineModelEverything Trust in Model 10 20 30 40 50 BaselineModelEverything Explanation Satisfaction Figure 2: Aggregate scores from our scales of agreement with (left) and perceived trust in (middle) the AI model, as well as satisfaction with the explanatory information (right), based on scales used in Hoffman et al. [16]. Participants in the Everything condition, who were provided with information that was ultimately insensitive to fairness assessments, had significantly higher ratings even though the underlying model was identical (and patently unfair) in all conditions. Error bars are 95% t-confidence intervals of the mean. 0% 25% 50% 75% 100% BaselineModelEverything % of Ratings The algorithm was fair across different races. 0% 25% 50% 75% 100% BaselineModelEverything The algorithm was fair across different genders. 0% 25% 50% 75% 100% BaselineModelEverything The AI's procedures are free of bias. 0% 25% 50% 75% 100% BaselineModelEverything Rating 1 - Strongly Disagree 2 3 4 5 6 7 - Strongly Agree The AI's procedures uphold ethical and moral standards. Figure 3: Likert responses from our participants when asked to assess potential biases in race and gender, or overall assessments of the ethics and biases of the model. While there was diversity in responses, in all cases on average, increasing the amount of information resulted in marginally higher ratings of perceived fairness (or lack of bias), despite the information being irrelevant to model fairness and bias. 0% 25% 50% 75% 100% BaselineModelEverything % of P ar ticipants Reported Potential Race/Gender Bias Issue 0% 25% 50% 75% 100% BaselineModelEverything Rating No Unsure Yes Reported Missing Explanatory Information Figure 4: Frequency of our derived qualitative codes based on par- ticipants’ free-text responses. Specifically, whether participants re- ported missing any explanatory information such as feature impor- tance scores, fairness metrics, or model architecture (on the left); and whether participants reported any possible unfairness with re- spect to the race or gender (on the right). As a reminder, our model only used race and gender to make decisions. The addition of un- informative explanatory components across our conditions resulted in lower rates of participants thinking more information was needed and noticing the severe biases in model outcomes. We additionally find that the amount of information influenced fairwashing effects. Participants who saw all of the explanatory components (in the Everything condition) demonstrated greater agreement with and reported higher trust in the model compared to those who saw the fewest explanatory components (in the Base- line condition). They also reported fewer fairness concerns in their open-ended responses compared to their Baseline counterparts. In short, we found that trust junk worked: participants were either “soothed” by the irrelevant information [39] or “numbed” by the sheer amount of data [6] in the explanations. All of this manipu- lation and persuasion occurred in the context of a model that was superficial and deeply unfair, precisely the case where we’d hope XAI would empower lay audiences to make accurate judgments. Our tested conditions cover some, but not all, of the strategies that Wall et al. [38] claim foster trust in viewers of XAI visualiza- tions. Future work is needed to look at specifically how factors like disclosure of uncertainty information, aesthetic appeal, and other specific trust junk “knobs” interplay. What makes an XAI visualiza- tion persuasive, misleading, or overwhelming is likely to be a com- plex combination of many sociotechnical factors [10] not amenable to the sort of self-contained study presented in this work. While some of the issues we uncover can be addressed by in- creased data literacy in audiences, education alone is not enough (even self-described AI experts habitually misinterpret or fail to un- derstand XAI techniques [20]). We echo the assertion of Hullman et al. [17]: the success of AI explanations can only be fairly as- sessed when considering the goals of such explanations. We urge the community to attend to the rhetorical goals of XAI explanations and their manipulative power as intrinsically persuasive artifacts. ACKNOWLEDGMENTS We thank Lace Padilla for comments on a draft of this work. REFERENCES [1] A. Adadi and M. Berrada. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI). IEEE Access, 6:52138– 52160, 2018. doi: 10.1109/ACCESS.2018.2870052 2 [2] N. Al-Ansari, D. Al-Thani, and R. S. Al-Mansoori. User-centered evaluation of explainable artificial intelligence (xai): A systematic literature review.Human Behavior and Emerging Technologies, 2024(1):4628855, 1 2024. doi: 10.1155/2024/4628855 2 [3] U. A ̈ ıvodji, H. Arai, O. Fortineau, S. Gambs, S. Hara, and A. Tapp. Fairwashing: the risk of rationalization.In ICML, p. 161–170. PMLR, 2019. 1, 2, 3 [4] A. Bussone, S. Stumpf, and D. O’Sullivan. The role of explanations on trust and reliance in clinical decision support systems. In International Conference on Healthcare Informatics, p. 160–169. IEEE, 2015. doi: 10.1109/ichi.2015.26 2 [5] F. Cabitza, C. Fregosi, A. Campagner, and C. Natali. Explanations considered harmful: the impact of misleading explanations on accu- racy in hybrid human-ai decision making. In World conference on explainable artificial intelligence, p. 255–269. Springer, 2024. doi: 10.1007/978-3-031-63803-9 14 2 [6] M. Correll. Towards a Theory of Bullshit Visualization, 9 2021. arXiv:2109.12975 [cs]. 2, 4 [7] B. Dimanov, U. Bhatt, M. Jamnik, and A. Weller. You shouldn’t trust me: Learning models which conceal unfairness from multiple expla- nation methods. In ECAI, p. 2473–2480. IOS Press, 2020. 2 [8] J. Drucker. Humanistic theory and digital scholarship. Debates in the digital humanities, 150:85–95, 2012. doi: 10.5749/minnesota/ 9780816677948.003.0011 1 [9] U. Ehsan and M. O. Riedl. Explainability pitfalls: Beyond dark pat- terns in explainable AI. Patterns, 5(6), 2024. doi: 10.1016/j.patter. 2024.100971 2 [10] U. Ehsan, K. Saha, M. De Choudhury, and M. O. Riedl. Charting the sociotechnical gap in explainable AI: A framework to address the gap in XAI. CSCW, 7:1–32, 2023. doi: 10.1145/3579467 4 [11] U. Ehsan, R. Singh, J. Metcalf, and M. Riedl. The algorithmic im- print. In ACM FAccT, p. 1305–1317. ACM, June 2022. doi: 10.1145/ 3531146.3533186 2 [12] M. Eiband, D. Buschek, A. Kremer, and H. Hussmann. The impact of placebic explanations on trust in intelligent systems. In CHI Extended Abstracts, p. 1–6. ACM, 2019. doi: 10.1145/3290607.3312787 2 [13] O. Gomez, S. Holter, J. Yuan, and E. Bertini. Advice: Aggregated visual counterfactual explanations for machine learning model vali- dation. In IEEE VIS, p. 31–35, 2021. doi: 10.1109/vis49827.2021. 9623271 2 [14] N. Goyal, C. Baumler, T. Nguyen, and H. Daum ́ e Iii. The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features. In IUI, p. 155–180. ACM, 3 2024. doi: 10.1145/ 3640543.3645210 3 [15] S. U. Hamida, M. J. M. Chowdhury, N. R. Chakraborty, K. Biswas, and S. K. Sami. Exploring the landscape of explainable artificial intelligence (xai): A systematic review of techniques and applica- tions. Big Data and Cognitive Computing, 8(149), 2024. doi: 10. 3390/bdcc8110149 2 [16] R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman.Met- rics for explainable AI: Challenges and prospects. arXiv preprint arXiv:1812.04608, 2018. 2, 3, 4 [17] J. Hullman, Z. Guo, and B. Ustun. Explanations are a means to an end. arXiv preprint arXiv:2506.22740, 2025. 4 [18] M. Jacobs, M. F. Pradier, T. H. McCoy Jr, R. H. Perlis, F. Doshi-Velez, and K. Z. Gajos. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry, 11(1):108, 2021. doi: 10.1038/s41398-021 -01224-x 2 [19] K. Kalasampath, K. N. Spoorthi, S. Sajeev, S. S. Kuppa, K. Ajay, and A. Maruthamuthu. A literature review on applications of explainable artificial intelligence (xai). IEEE Access, 13:41111–41140, 2025. doi: 10.1109/ACCESS.2025.3546681 2 [20] H. Kaur, H. Nori, S. Jenkins, R. Caruana, H. Wallach, and J. Wort- man Vaughan. Interpreting interpretability: understanding data scien- tists’ use of interpretability tools for machine learning. In CHI, p. 1–14. ACM, 2020. doi: 10.1145/3313831.3376219 4 [21] H. Kennedy, R. L. Hill, G. Aiello, and W. Allen. The work that vi- sualisation conventions do. Information, Communication & Society, 19(6):715–735, 2016. doi: 10.1080/1369118x.2016.1153126 1 [22] T. Kulesza, S. Stumpf, M. Burnett, S. Yang, I. Kwan, and W.-K. Wong. Too much, too little, or just right? ways explanations impact end users’ mental models. In VL/HCC, p. 3–10. IEEE, 9 2013. doi: 10. 1109/VLHCC.2013.6645235 2 [23] S. Laato, M. Tiainen, A. Najmul Islam, and M. M ̈ antym ̈ aki. How to explain AI systems to end users: a systematic literature review and research agenda. Internet Research, 32(7):1–31, 2022. doi: 10.1108/ intr-08-2021-0600 2 [24] H. Lakkaraju and O. Bastani. “how do I fool you?” manipulating user trust via misleading black box explanations. In AIES, p. 79–85. AAAI/ACM, 2020. doi: 10.1145/3375627.337583 2 [25] M. Langer, T. Hunsicker, T. Feldkamp, C. J. K ̈ onig, and N. Grgi ́ c- Hla ˇ ca. “Look! It’s a Computer Program! It’s an Algorithm! It’s AI!”: Does Terminology Affect Human Perceptions and Evaluations of Algorithmic Decision-Making Systems? In CHI, p. 1–28. ACM, 4 2022. doi: 10.1145/3491102.3517527 2 [26] Q. V. Liao, D. Gruen, and S. Miller. Questioning the AI: informing design practices for explainable AI user experiences. In CHI, p. 1– 15. ACM, 2020. doi: 10.1145/3313831.3376590 2 [27] S. Mohseni, N. Zarei, and E. D. Ragan. A multidisciplinary survey and framework for design and evaluation of explainable AI systems. TiiS, 11(3-4):1–45, 2021. doi: 10.1145/3387166 2 [28] M. Nourani, S. Kabir, S. Mohseni, and E. D. Ragan. The effects of meaningful and meaningless explanations on trust and perceived sys- tem accuracy in intelligent systems. In HCOMP, vol. 7, p. 97–105. AAAI, 2019. doi: 10.1609/hcomp.v7i1.5284 2 [29] M. Nourani, C. Roy, J. E. Block, D. R. Honeycutt, T. Rahman, E. Ra- gan, and V. Gogate. Anchoring bias affects mental model formation and user reliance in explainable AI systems. In IUI, p. 340–350. ACM, 2021. doi: 10.1145/3397481.3450639 2 [30] M. Nourani, C. Roy, T. Rahman, E. D. Ragan, N. Ruozzi, and V. Gogate. Don’t explain without verifying veracity: an evaluation of explainable AI with video activity recognition. arXiv preprint arXiv:2005.02335, 2020. 3 [31] C. O’Neil. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Penguin Books, London, 2016. 2 [32] E. M. Peck, S. E. Ayuso, and O. El-Etr. Data is personal: Attitudes and perceptions of data visualization in rural pennsylvania. In CHI, p. 1–12. ACM, 2019. doi: 10.1145/3411764.3445315 1 [33] F. Poursabzi-Sangdeh, D. G. Goldstein, J. M. Hofman, J. W. Wort- man Vaughan, and H. Wallach. Manipulating and measuring model interpretability. In CHI, p. 1–52. ACM, 2021. doi: 10.1145/3411764 .344531 2, 3 [34] D. Pruthi, M. Gupta, B. Dhingra, G. Neubig, and Z. C. Lipton. Learn- ing to deceive with attention-based explanations.In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, eds., ACL, p. 4782–4793, July 2020. doi: 10.18653/v1/2020.acl-main.432 2 [35] G. Schwalbe and B. Finzel.A comprehensive taxonomy for ex- plainable artificial intelligence: a systematic survey of surveys on methods and concepts.Data Mining and Knowledge Discovery, 38(5):3043–3101, Sept 2024. doi: 10.1007/s10618-022-00867-8 2 [36] D. Slack, S. Hilgard, E. Jia, S. Singh, and H. Lakkaraju. Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods. In AIES, p. 180–186. AAAI/ACM, 2 2020. doi: 10.1145/3375627. 3375830 2 [37] M. Szymanski, M. Millecamp, and K. Verbert. Visual, textual or hy- brid: the effect of user expertise on different explanations. In IUI, p. 109–119. ACM, 4 2021. doi: 10.1145/3397481.3450662 2 [38] E. Wall, L. Matzen, M. El-Assady, P. Masters, H. Hosseinpour, A. En- dert, R. Borgo, P. Chau, A. Perer, H. Schupp, H. Strobelt, and L. Padilla. Trust junk and evil knobs: Calibrating trust in AI vi- sualization. In PacificVis, p. 22–31. IEEE, 2024. doi: 10.1109/ pacificvis60374.2024.00012 1, 2, 4 [39] A. Weller.Challenges for transparency.arXiv preprint arXiv:1708.01870, 2017. 1, 2, 4 [40] L. F. Wightman. LSAC National Longitudinal Bar Passage Study. LSAC Research Report Series. Law School Admission Council, 1998. 2 [41] C. Xiong, J. Shapiro, J. Hullman, and S. Franconeri. Illusion of causal- ity in visualized data. IEEE TVCG, 26(1):853–862, 2019. doi: 10. 1109/tvcg.2019.2934399 2 [42] B. Ytre-Arne and H. Moe. Folk theories of algorithms: Understanding digital irritation. Media, Culture & Society, 43(5):807–824, 2021. doi: 10.1177/0163443720972314 2