Paper deep dive
SketchXplain: Intuitive Visual Explanations of Image Classifiers with Sketches
Wencan Zhang, Mario Michelessa, Xuejun Zhao, Brian Y. Lim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/20/2026, 10:25:37 AM
Summary
SketchXplain is a novel Explainable AI (XAI) framework that generates intuitive, sketch-based visual explanations for image classifiers. Unlike traditional saliency maps that often lack semantic clarity, SketchXplain uses a deep neural network to create simplified, coherent sketches. The architecture combines a base image encoder, a concept bottleneck model for semantic coherence, and a sketch optimization process (inspired by CLIPasso) that aligns visual cues with both the input image and the predicted class label. The method aims to satisfy three desiderata: simplicity, coherence, and quick interpretation. It was evaluated on facial expression recognition and skin lesion diagnosis, demonstrating superior performance in supporting human interpretation compared to standard saliency maps.
Entities (8)
Relation Signals (6)
SketchXplain → appliedto → Face Expression Recognition
confidence 100% · Evaluating on face expression recognition, modeling and user studies showed that SketchXplain supported quicker interpretation...
SketchXplain → appliedto → Skin Lesion Diagnosis
confidence 100% · Further evaluation on skin lesion diagnosis found that SketchXplain more coherently visualized disease symptoms...
SketchXplain → combines → Concept Bottleneck Model
confidence 100% · Combining techniques in saliency maps, concept-bottleneck models, and sketch optimization, SketchXplain integrates...
SketchXplain → extends → CLIPasso
confidence 100% · We extend CLIPasso to additionally i) extract more detailed lines... ii) infer explanatory concepts... and iii) align strokes to visual cues...
Grad-CAM → usedfor → Cue Localization
confidence 100% · We implemented Grad-CAM for transformers [94] to generate one CAM per multi-label concept... and use the aggregated saliency map to mask the detailed lines...
SketchXplain → improvesupon → Saliency Map
confidence 90% · SketchXplain supported quicker interpretation with more aligned visualizations than saliency maps or simple drawings.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Saliency map visualizations explain image-based AI predictions by pointing to regions, but these are often unintuitive and semantically unclear, leaving an interpretability gap. We argue that AI explanations should be intuitive -- coherent to user knowledge, yet simple and selective to accelerate interpretation. Inspired by artistic drawings, we propose SketchXplain to generate sketch-based visual explanations for intuitive image-based explainable AI (XAI). Combining techniques in saliency maps, concept-bottleneck models, and sketch optimization, SketchXplain integrates saliency to select coherent observation artifacts, concepts for knowledge coherence, cues to represent them, and abstraction for simplicity. Evaluating on face expression recognition, modeling and user studies showed that SketchXplain supported quicker interpretation with more aligned visualizations than saliency maps or simple drawings. Further evaluation on skin lesion diagnosis found that SketchXplain more coherently visualized disease symptoms, better supporting lay diagnosis. Thus, this work illustrates the value of sketches for intuitive, simple, coherent, and quick image-based XAI visualizations.
Tags
Links
- Source: https://arxiv.org/abs/2606.17646v1
- Canonical: https://arxiv.org/abs/2606.17646v1
Trouble viewing inline? Open PDF directly →
Full Text
142,191 characters extracted from source content.
Expand or collapse full text
SketchXplain: Intuitive Visual Explanations of Image Classifiers with Sketches Wencan Zhang * Mario Michelessa * Xuejun ZhaoBrian Y. Lim † National University of Singapore a) Noneb) Saliencyc) Outlined) Edgese) CLIPasso g) Concepts Explanation only Brow lowerer Upper eyelid raiser Eyelid tightener Lip tightener Overlaid on photo Explanation only Border irregular Color varying Overlaid on photo f) SketchXplain * * * * * * Figure 1: Visualsaliency (b), sketch (c–f) and verbal explanations of image classifications for two domains, face expression (angry) and skin lesion (melanoma): a) None, b) Saliency map [1] c) Outlinefromfaciallandmarks[2]orlesionsegmentationmasks[3], d) Edges detected [4] from photo, e) CLIPasso [5] sketch that neglects explanatory cues, f) our SketchXplain that leverages sketches for intuitive explanations, g) Concepts labels [6]. We compared a subset * of visualizations in quantitative user studies. ABSTRACT Saliency mapvisualizations explain image-based AI predictions by point- ing to regions, but these are often unintuitive and semantically unclear, leaving an interpretability gap. We argue that AI explanations should be intuitive–—coherent to user knowledge,yet simple and selective to acceler- ate interpretation. Inspired by artistic drawings, we propose SketchXplain to generate sketch-basedvisual explanations for intuitive image-based explain- able AI (XAI). Combining techniques in saliency maps, concept-bottleneck models, and sketch optimization, SketchXplain integrates saliency to select coherent observation artifacts, concepts for knowledge coherence, cues to represent them, and abstraction for simplicity. Evaluating on face expression recognition, modeling and user studies showed that SketchXplain supported quicker interpretation with more alignedvisualizations than saliency maps or simple drawings. Further evaluation on skin lesion diagnosis found that SketchXplain more coherentlyvisualized disease symptoms, better sup- porting lay diagnosis. Thus, this work illustrates the value of sketches for intuitive, simple, coherent, and quick image-based XAIvisualizations. Keywords: Explainable AI, Interpretability, Sketch explanation 1 INTRODUCTION Artificial Intelligence (AI) has achieved impressive performance on image prediction tasks [7, 8]. However, the complexity of AI models has raised concerns about their lack of transparency [9, 10]. This has driven the grow- ing need for Explainable AI (XAI) [11–15], which aims to make model predictions more understandable to users. Saliency maps [1, 16–20] are among the most popular XAI techniques * These authors contributed equally to this work. † Corresponding author. e-mail: brianlim@nus.edu.sg to explain visual tasks. They attribute a model’s prediction to specific pixels [18] or intermediate neurons in the model [1,16]. While saliency maps faithfully represent the model’s attention by highlighting relevant regions, they often fail to convey meaningful semantics, rendering explanations unintuitive [21]. For instance, a user might know where to focus but fail to understand why that region is relevant (Fig. 1b). This limited intuitiveness increases cognitive effort [22], and makes it difficult to interpret [23, 24] or anticipate [25] model behavior. One intuitive way to convey how images are perceived is throughsimpli- fied abstractsketches that emphasize salient strokes, omitting less important contour lines [26],whileremainingcoherenttokeyconcepts. Since sketches are perceived as approximately realistic images [27], people canquickly and accurately categorize scenes and recognize objects from them [28].Be- yondartisticabstraction,thissimplicity,coherency,andquickness,enables sketchestobeusedasasharedboundaryobjecttocommunicatebetween disciplines[29]. This opens a new opportunity for communicatingsemantic explanationsvisually beyond verbalizing concept-based explanations [6, 30]. Thus, we identify three desiderata of Intuitive Explanations—Simplicity [10, 31, 32], Coherence [10, 33, 34] and Quick Interpretation [31, 32, 35]—to closetheinterpretabilitygapofsaliencymaps(Fig.2). We operationalize these desiderata todefinedesignobjectivesandevaluationmeasuresfor intuitive sketch explanations that align with human knowledge, reflect the visual observation, and incorporate saliency-informedsemanticcues. Guided by these desiderata, we propose SketchXplain, a deep neural net- work to generate simple yet coherent sketch explanations (Fig. 1f). SketchX- plain consists of 1) base image encoder, 2a) concept bottleneck model to predict intermediate concepts and combine them to predict the class label, 2b) cue localization to locate each concept in the image, 3a) stroke initializa- tion to extract edges from the image and filter them based on localization, and 3b) stroke optimization to refine the sketch to faithfully represent the im- age and prediction label. Although the sketch explanation is model-agnostic, like LIME [12], it is faithful to the model’s input and prediction while pro- viding plausible, human-accessible explanations—like AI rationalization arXiv:2606.17646v1 [cs.HC] 16 Jun 2026 User Observation AI UserSaliency Map CoherentSimple Intuitive SketchXplain Interpretability Gap Quick Figure 2: Interpretability gap in saliency maps is addressed by intuitive sketch-based explanations that are simple and coherent, leading to quicker interpretation. from human labels [36] or large language models [37]. Wefirst evaluated SketchXplain on a facial expression task using various intuitiveness measures. We compared SketchXplain explanations against saliency maps and other line-drawing methods across multiple studies(Sec- tion5): a)preliminary modeling studies assessingproxiesof explanation visual complexity and coherence; b) qualitative user study identifying inter- pretation benefits; c) quantitative user study evaluating quick interpretation under time constraints. For generalization, we extended SketchXplain to explain skin lesion predictions(Section6)andgeneralimages(inAppendix). Alluserstudieswereapprovedbyouruniversityinstitutionalreviewboard (IRB). Overall, we demonstrate that SketchXplain produces sketches that are coherent and sufficiently simple, enabling more intuitive visual explanations. Hence,weframeSketchXplainvisualizationsasarticulatingAIpredic- tionstointuitivelyguidefinaluserdecisions. Our contributions are: 1)First to use sketches as anintuitive explanatory modality for visual explanation to closethe interpretability gap ofsaliencymaps.Westeer sketchgenerationtowardexplanatorycoherence. 2)SketchXplain,anexplanationmethodtogenerateintuitivesketchvisu- alizations,satisfyingthedesiderataofsimplicity,coherence,andquick interpretationforimageclassifiers. 3) Demonstration of sketch explanations onmultipledomains(faces and skin lesions),withproxyanduserevaluations. 2 BACKGROUND AND RELATED WORK Wediscussthelandscapeofimage-basedXAIanditscognitivechallenges, providingthebasisforusingvisualsketchesasintuitiveexplanations. 2.1 Explainable AI for Computer Vision ManyvisualizationshavebeendevelopedtounderstandhowCNNsinfercon- ceptsandpredictionlabelsfromimages[38–43].Whilethesehelpstudents ordevelopersbetterlearnaboutordebugmachinelearning,theyareless accessibletolayusers.Instead,attributionandconcept-basedexplanations aremorepopulartosupportend-userunderstandingofAIdecisions. Attribution explanation methodsreport feature importance by attributing model outputs to input features [12, 13, 44], allowingusers to focus onsalient features for reasoning. In image-based AI prediction tasks, saliency map techniques highlight influential pixels using gradient [1], perturbation [19], decomposition [20], ablation [17], or relevance propagation [18]. Although pointing to specific regions can direct a user’s attention,saliencymapslack semanticcontextualization,limitingintuitiveinterpretation. In contrast, concept-based XAI examines the influence of human- interpretable concepts.Techniquesincludeverballyexplainingviaconcept bottlenecks[6],interactivelyinspectingposthocconcepts[30, 45, 46],and visualizinglearnedneuronsorfilters[15, 47]. While these explanations are more semantically-aligned withdomainknowledge,theyarenotvisually linkedtotherepresentationorcuesintheimage,leavingacoherencegap. In this work, we are the first to investigate the usability andusefulness of sketches as an explanation modality, moving beyond simple attribution of influential features or verbal concept descriptions. It is important to note that our approach does not share the same goal as techniques to explain sketch generation [48, 49]. By treating concepts and cues as foundational elements, we use saliency methods to adjust their combination, achieving coherent yet simple sketches that are intuitive for human interpretation. 2.2 Visualization for End-User Human Interpretability VariousXAIvisualizations,includingrulematrices[50],beamsearch trees[51],simplifiednetworkgraphs[39],networktraces[42],orinter- activedistributionplots[52],havebeendevelopedtoassistdatascientistsin debuggingdataissuesandmodelbehavior.Althoughsomehelpend-usersto understandmodelsthroughsaliencymaps[53]andclasstreehierarchies[22], thevisualformatsremaincomplexforlayusers. Inpractice,XAItechniqueshavefacedusabilitychallenges, including misleading or incomplete explanations [54,55], misinterpretations [56], and over-reliance [57, 58]. To address these issues, researchers have advocated for human-centered XAI [10, 59, 60]. Approachesincludemoderatingcog- nitiveloadandimprovingmemorability[31, 61],aligningAIarchitectures tohumanreasoningprocesses[24, 62, 63],anticipatinghumanmisconcep- tions[64],forcingdeliberation[57],progressivedisclosureondemand[65– 67],andselectiveortailoredexplanations[59, 68–70]. Incontrast,weproposeacomplementaryapproachthatleveragessketches asaselectivevisualabstractionforintuitivevisualexplanations.Wehypoth- esizethatsuchexplanationswouldengageSystem1thinking[71]tohelp usersrapidlyassimilateanunderstandingoftheAIdecisionforvisualtasks. 2.3Illustrative Visual Abstraction for Semantic Communication Unlike photographs with full visual detail, illustrations [72] are flexible me- dia used to explain or enrich ideas, concepts, or narratives, intentionally abstracting or stylizing information to guide attention [73], clarify con- cepts [74], or convey meaning beyond literal imagery [72],makingthem highlyeffectiveforvisualizations[75].Tocomputationallyreplicatethis communicativeabstraction,non-photorealisticrendering(NPR)[76]meth- odssynthesizeillustrationsfromphotographsor3Dmodels. Unlike tra- ditional rendering which focuses on realistic lighting and shading, NPR emphasizes artistic styles or structural information, such as contour extrac- tion [77], stroke-based rendering (SBR) [78] and stylization [79]. However, theseneglectcognitiveprinciplesforcommunicatingillustrations[75]. AligningwithAgrawalaetal.’sparadigmofabstraction[75],weleverage illustrativevisualsketchingnottoassertanartisticstyle,buttoemphasize semanticcues.Byusingasparsesetof strokes to emphasize a subject’s core semantic essence [27, 80], sketches bridge the gap between spatially depictive geometry and structurally descriptive semantic concepts [26].How- ever,traditionalsketchingguidelinesandclassicalNPRrenderingpipelines operatepurelyasstaticpost-processingonrawvisualgeometry,lacking thecapacitytodynamicallyextract,represent,orprioritizetheunderlying activationstoexplainpredictionsofimage-classifiermodels. 2.4 AI Sketch Generation from Photographs Sketch generation differs from edge-map extraction [81], which is purely geometric, by producing abstracted drawings thatprioritize semantics.Re- centAI-basedmethodsemploy image-to-image transformation between photos and sketches [82] or unpaired domain transfer [83, 84], which can be trained with cycle consistency losses [85]. Although these methods gen- erate sketches in various styles, they rely heavily on curated sketch datasets. However,targetsketchescorrespondingtosourceimagesmaynotexist. Instead,withoutrequiringsupervisedlearningviatargetsketches,graph- ical vector-based(vs.bitmap-based) approacheshavebeenproposedto generate sketches [86–88] via differentiable rendering [89]. For example, recent work [5, 90] define sketches as sets of B ́ ezier curves and optimizes stroke parameters and leverage CLIP [91] with a joint image-text latent space to align visual and textual representations.Whilethesemethodsdoachieve simplifiedsketches,faithfultothephotosandtextualdescriptions,theyne- glectexplanatorycuesandsemanticsimportantforuserunderstandingof image-AIpredictions,whichweintegrateseamlesslyinSketchXplain. 3DESIDERATAFORINTUITIVEVISUALEXPLANATIONS To bridge the interpretability gap of XAI and reduce misinterpretations, we argue that AI explanations should be more intuitive.Cognitivepsycholo- gistsdescribeintuitionasthinking“automaticallyandquickly,withlittle ornoeffort”[71],“focusingontherelevantanddeliberatelyignoringthe rest”[92],suchthat“onceexperiencedintuitivedecisionmakersseethe pattern,anydecisiontheyhavetomakeisusuallyobvious”[93]. Therefore, information that is coherent to user expertise, yet simple and focused, helps facilitate quick intuitive reasoning. Consolidating XAI literature, we target the following desiderata for intuitive visual explanations: 1) Simple [10, 31, 32] to convey salient information clearly and concisely. 2) Coherent [10, 33, 34] to align with beliefs, evidence and hypotheses. 3) Quick [31, 32, 35] to allow easy understanding and avoid confusion. From these three desiderata, we derive explanation design objectives, proxy modeling evaluation metrics, and user-study measures, which we apply to sketch-based explanations of image-based AI predictions. 3.1 Explanation Design Objectives In Section 4, we describe our technical approach toward intuitive sketch explanations. To support simplicity, we optimize sketches for abstraction to ignore less relevant fine details, with smooth strokes rather than jagged ones. To support coherence, we optimize sketches to be representative of the AI prediction task, explanatory concept labels from human knowledge, cues that visually represent concepts, and the visual observation of the instance. 3.2 Proxy Evaluation Metrics In Section 5.2, we evaluated each design objective with correspondingproxy metrics. For simplicity, we evaluated its opposite visual complexity in terms of information in the spatial domain (e.g., local entropy) and frequency domain (e.g., discrete cosine transform in JPEG XL). For coherence, we measured the alignment of the sketch’sspatialrepresentationtotheorigi- nalphotoandextractedvisualcues,andembeddingrepresentationtothe embeddingsfortheAIpredictionandexplanatoryconceptlabels. 3.3 User Study Measures Sinceexplanationsareforhumaninterpretation,weconducteduserstudies toevaluateexplanationintuitiveness. Sections 5.3 and 6.3 report qualitative user studies to examine the inherent coherence and usability of sketch expla- nations. Section 5.4reports a quantitative study under time constraints to assess if users can quickly interpretthesketchexplanationtorecallrelevant conceptsandinferpredictionlabels. 4 SKETCHXPLAIN: TECHNICAL APPROACH We propose SketchXplain, an explainable deep neural network with three main modules (Fig. 3): 1) base prediction, 2) concept explanation, and 3) sketch rationalization.Itsatisfiestwodesideratatogeneratesimplestrokes byabstractingvisualinformationintosmoothedcurves,whicharecoherent totheimagepredictiontaskandexplanatoryconceptsbyaligningvisual cuestosemanticembeddingrepresentationsandtheobservedphoto. 4.1 Base Prediction We begin with the base task (Fig. 3, Step 1) of predicting class labelˆy(e.g., expression) from input image x x x (e.g., face)usinganimageencoderM. 4.2 Concept Explanations We provide a base explanation of conceptsviaaconceptbottlenecknetwork (Fig. 3, Step 2a) andlocalizations of associated cues in the imageviacon- cept-specificsaliencymaps (Fig. 3, Step 2b),whicharesubsequentlyused toconstrainthecoherenceofthesketchexplanation. 4.2.1 Concept Bottleneck Explanations should be relatable [24] to concepts [6, 30]. SketchXplain first predicts concepts ˆ γ γ γvia multi-label binary classificationasmultipleconcepts canco-occur 1 .Conceptsarethenusedasintermediatefeatures in a concept bottleneck model F γ [6] to predict the label ˆyvia multi-class classification. 4.2.2 Cue Localization For image-based predictions, saliency maps are widely used to explain which pixelsare important. We implemented Grad-CAM for transformers [94] to generate one CAM permulti-label concept ˆ ε ε ε γ , whichhighlightsimportant pixelsforpredictingeach corresponding concept ˆ γ γ γ.Weextractedimpor- tanceweights ˆ w w w γ (i.e.,gradientsofpredictionlabelw.r.t.concepts)through thefullyconnectedlayersoftheconceptbottleneckF γ . Unlike concept predictions ˆ γ γ γthat communicate whether a concept is present, ˆ w w w γ explains how important each concept is to determine the class label. To ensure more sensible saliency maps, we regularize the saliency pre- diction by making Grad-CAM a secondary prediction task trained via self- supervision [95]. Specifically, wepenalize salient pixels that are far away from where each concept is located. 4.3 Sketch Rationalization With the semantic information of concept-based explanations, and the lo- calization information from saliency maps, we can develop more intuitive explanations using sketches. Specifically, our approach extracts and ab- stracts informative Concept lines relevant to concepts to communicate key details as explanations. This extraction and sketching process involves extra pathwaystorationalize beyond the scope of the original prediction model M, so it is not a mechanistic interpretability approach [96]. Nevertheless, explanation rationalization is a plausible approach [36] for integrating astute observations from a sketch expert (as for artists) with attentive feedback from the target predictor (as for patrons or clients). We generate sketch explanations with two steps: a) initialize informative strokes extracted from the photo (Fig. 3, Step 3a), b) draw the sketch starting from those strokes and optimize it to best represent the original photo and its prediction label (Fig. 3, Step 3b). Prior work, CLIPasso [5], draws a small set of curved lines that capture the semantics of a photo (via CLIP [91]) by optimizing the similarity between the final sketch and the input photo.This servesthegoalofsimplicity,butnotcoherencetoexplanation.Weextend CLIPassotoadditionallyi)extractmoredetailedlinesthanjustobjectout- lines,i)inferexplanatoryconceptsrelevanttothepredictionlabel,andiii) alignstrokestovisualcuesrepresentingcorrespondingconcepts.Thus,this servesthegoalofexplanationcoherence. 4.3.1 Strokes Initialization To generate the sketch, CLIPasso starts with an initial set of strokes whose locations are sampled from salient regions. However, CLIPasso’s stroke initialization often fails to produce semantically meaningful sketches (see Table 2, CLIPasso).Toaddressthis,weusedetailedlinestocapturefiner details,suchasskinfoldsanddynamicwrinklesinfaces,orroughnessin texturedsurfaces. Thesedetailedlines are extracted into a bitmap image ̃ x x x using the pretrained image-to-image generatorEby Chan et al. [85].See AppendixTables6and9,lasttwocolumns,foracomparisonbetween initializationwithdetailedlines ˆ z z z x andwithoriginalphotox x x. AsinCLIPasso,weprioritizethemostsalientlinesof ̃ x x xbyleveraging saliency maps ˆ ε ε ε γ and importance weights ˆ w w w γ fromSection4.2.2. Using Grad-CAM, we merge all CAMs as a weighted-sum ˆ ε ε ε = ReLU ˆ w w w ⊤ γ ˆ ε ε ε γ , and use the aggregated saliency map to mask the detailed lines based on their importance, i.e., ̆ x x x = ˆ ε ε ε⊙ ̃ x x x, where⊙is the Hadamard element-wise multiplication.Thisaimstoimprovetheexplanationcoherencetotheex- planatoryconceptualcues.Theseweighteddetailedlinesareusedinthenext stage(strokeoptimization)tosamplethestrokelines. 1 Multi-labelbinaryclassificationhandlesmultiplenon-exclusivebinarylabels, whilemulti-classclassificationonlyselectsone. For example, Facial action unit AU1 Inner Brow Raiser, AU2 Outer Brow Raiser, AU5 Upper Eyelid Raiser, and AU26 Jaw Drop co-occur for Surprised expression; see Table 1. 풚 " 풙 퐸 풙 & 휸 " 퐹 ! 풘 " ! 푀 2a Concept Bottleneck 1Base Prediction 3 SketchXplain Rationalization × ∘ 푆 풙 + 풔 - " 풙 " # 푅 풛 - # 풛 - " 풛 - $ 퐶 %&' 풙 + 풙 " # 퐶 %&' 3cSemantic Alignment 3b Strokes Optimization 퐶 ()( ℒ * 3a Strokes Initialization 2b Cue Localization 흐 - ! Figure 3: SketchXplain architecture comprising: 1) base prediction, 2) base explanations in terms of concepts (2a) and cues (2b), 3) sketch explanation to generate initial strokes (3a), and optimize them (3b) to align them with the predicted class label (3c). Instance shown for a face expression use case (Section 5). 4.3.2 Strokes Optimization As in CLIPasso, we seek to generate parametric strokes ˆ s s s x that can be updated using gradient-based optimization. We initialize these strokes by samplingSendpoints for quadratic B ́ ezier curves from ̆ x x x, treating it as a probability density function where more salient pixels have a higher prob- ability of being sampled. The parametric strokes ˆ s s s x can be rasterized into a sketch ˆ x x x s using a differentiable rasterizerR[89]. Initially, the rasterized sketch ˆ x x x s will resemble a sparse fragmented rendition 2 of ̆ x x x, but can be improved through iteratively updating ˆ s s s x during inference.Unlikemodel trainingwhereweightsareupdated,wefreezetheweightsinR,anduse iterativeoptimizationviagradientdescentatinferencetimetoupdatethe input ˆ s s s x ,likewithactivationmaximization[97]. To guide the sketch optimization, weperformSemanticAlignment(Fig.3, Step3c)ofthesketchexplanation ˆ x x x s totheinputphoto ̆ x x xandpredictedlabel ˆy. We leverage CLIP [91], which encodes images and text into a shared feature spacetoenabledirectcomparisonbetweenvisualandtextualrepre- sentations.Specifically,weobtainembeddingrepresentationsforthesketch ˆ z z z s ,inputphoto ˆ z z z x ,andpredictionlabel ˆ z z z y .TheoriginalCLIPassoonlyaligns thesketchestowardthevisualsemantics ˆ z z z x ,butweaddalignmenttoward thepredictionlabel ˆ z z z y .Wethenpenalizelargecosinedistancesbetween embeddingvectors.Thisaimstoimprovetheexplanationcoherencetothe imageandpredictionlabel.Toimprovetheexplanationsimplicity,wealso includeasmoothnessconstraintκ( ˆ s s s x )forlesscurvystrokes. Thus, the full training loss is L =− cos( ˆ z z z s , ˆ z z z x )− λ t cos( ˆ z z z s , ˆ z z z y )+ λ κ κ( ˆ s s s x ),(1) whereκ( ˆ s s s x ) = | ˆ s s s ′ x,1 ˆ s s s ′ x,2 − ˆ s s s ′ x,1 ˆ s s s ′ x,2 | ( ˆ s ′2 x,1 + ˆ s ′2 x,2 ) 3/2 is the curvature of a stroke ˆ s s s x = ( ˆ s x,1 , ˆ s x,2 )in 2D, with primes indicating first- and second-order derivatives.λ t andλ κ are the corresponding hyperparameters.Tables2and4showexampleSketchX- plainvisualizationsoffaceexpressionsandskinlesions,respectively. 4.4 Implementation Details We implemented SketchXplain in PyTorch. For concept inference, we fine- tuned the final layer of the concept bottleneck model with Adam [98] for 100 epochs (batch size 32), and used cross-entropy loss for classifications. For sketch optimization, we also used Adam, with a learning rate of 1.0 (applyingthefullgradientmagnitude) for stroke positions and10 −2 for 2 See Appendix Tables 6 and 9, column Init. strokes. stroke opacity.Like[5, 99], to improve stability, gradients were averaged over four sketchesasdataaugmentionsvia random affine transformations oneachsketchinstance. Hyperparameters in Eq. 1 were set toλ t = 0.1 andλ κ = 0.01based on grid search. To encourage finer details,λ t = 0was set during the final 30% of training. For each image, we optimized sketch parameters over 1000 iterations (takingabout10sec per image), as in [5, 99]. Toimprovesimplicity,intherasterizer[89],weenabledstrokeopacityto de-emphasizelessimportantlines,anddiscretizedthemtotwolevels 3 . All experiments were conducted on a server with 8 NVIDIA RTX 3090 GPUs. 5 EVALUATION ON FACE EXPRESSION Weinvestigatedtheusabilityandusefulnessofsketchexplanationsinuser studiesoftwoapplicationdomains:facialexpressionrecognition(thissec- tion)andskinlesioncancerdiagnosis(describedlaterinSection6) 4 .We chosefaceexpressionrecognitionsinceitisaccessibletolayusersandim- portantforapplicationslikehuman-AIcommunication[101],mentalhealth therapy[102],educationaltools[103].Also,facesketchesareauniversally familiarmedium[26].Toexplainfacialexpressions,weusedfacialAction Units(AU)asexplanatoryconcepts.AUsarespecificmusclemovements, suchasraisedinnereyebrowsandliptightener,toexpressemotions 5 [104], whichpeopleobserveascuestoinfertheemotionofthesubject[105, 106]. Researchquestions:Aresketchexplanationsmore... RQ1)Intuitive (simpler, more coherent) than baselinevisualexplanations (line drawings, descriptive sketches, saliency maps)? RQ2) Interpretable and coherent to human intuition (qualitatively)? RQ3) Intuitively(quickly) interpretable to understand AI predictions? Weanswerthesequestionswithapreliminarymodelingstudywithproxy metrics(RQ1),qualitativestudyofuserinterpretationofvariousvisualiza- tions(RQ2),andquantitativeuserstudyforquickinterpretation(RQ3). Furthermore,duetotheabstractioninsketchexplanations,sketchescan alsoprovidethebenefitofprivacyprotectionbynotexposingidentifiable informationintheoriginalphoto.Weinvestigatedthisinanotherquantitative userstudyand,forbrevity 6 ,presentdetailsinAppendixA.3. 3 Pilotevaluationsfound≥ 3levelslesslegible,withmessyinformationoverload. 4 We also investigated sketch explanations of general images; see Appendix C. 5 We focused on recognizing expressed emotions, not internal emotional states. 6 Quickresultsonprivacy:wefoundthatsketchexplanationscouldprotectiden- tity,andgenderandethnicityinformationalmostaswellasSaliencymaps. Table 1: Typical Action Units (AUs) are the basis for concept-based facial expression explanations. 12 AUs are arranged in two rows, where the first row shows upper face AUs, while the second row shows lower face AUs. Pictograms adapted from Blasberg et al. [100]. AU1 Inner Brow Raiser AU2 Outer Brow Raiser AU5 Upper Eyelid Raiser AU26 Jaw Drop Surprised AU6 Cheek Raiser AU12 Lip Corner Puller Happy AU1 Inner Brow Raiser AU4 Brow Lowerer AU15 Lip Corner Depressor SadDisgusted AU9 Nose Wrinkler AU15 Lip Corner Depressor AU16 Lower Lip Depressor Angry AU4 Brow Lowerer AU5 Upper Eyelid Raiser AU7 Eyelid Tightener AU23 Lip Tightener 5.1 Data Preparation and Selection We used the Binghamton-Pittsburgh 3D Dynamic Spontaneous Facial Ex- pression Database (BP4D-Spontaneous) [107] dataset to train and test all models. This contains videos of face expressions from 41 actors (23 Female, 18 Male) with diverse ethnicities (20White, 11 Asian,6Black, 4 Hispanic). Each video consists of a continuous series of images, totaling 147.5k images. This datasetis primarily used for predicting action units, but it is also very suitable for expression prediction. However, it lacks expression labels, so we used HSEmotion [108] to automatically annotate expressions. Furthermore, due to the transient face changes, many images do not fully represent expres- sions, so we filtered out images that had low prediction confidence (<70%). To obtain a balanced dataset with equal numbers of instances per expression, we excluded two classes with low frequency (Contemptuous, Fearful), re- sulting in 6 expressions—Surprised, Happy, Neutral, Sad, Disgusted, and Angry. Table 1 shows the typical associations between facial expressions and action units (AUs). Consequently, we selected a subset of 4900 images of face photos with pseudo-labels for 6 expression classes, which we split into 80% training and 20% test. Due to the sequential nature of image frames in the BP4D dataset, many test images are redundantly similar. Therefore, we selected two examples of each actor expressing a different emotion, arriving at 110 images. After balancing for ethnicity(White,Asian,Black) 7 and gender(Female,Male), we obtained 72 images(6expressions×3ethnicities×2genders×2 examples), which we used in our modeling and user studies. 5.2 Preliminary Modeling Study UsingpretrainedOpenGraphAU[109]asthebackboneM,ontheBP4D dataset,SketchXplainachievedgoodaccuracy82.0%andF1Score71.3 onfaceexpressionclassification,per-AUF1scores44.3–89.2(M=59.9) 8 . Weconductedamodelingstudyasapreliminarycheckontheextentthat SketchXplaingeneratessketchexplanationstowardthedesiderataofsim- plicityandcoherencewithproxymetrics. 5.2.1 Proxy Computational Metrics Webrieflyintroduceproxymetricsforsimplicityandcoherence.SeeAp- pendixA.1formoredetailsofhowspecificmetricswereselected. Visual complexity:Toestimatesimplicityacrossheterogeneousimage types,weexaminedvariousinformation-theoreticmetricsofvisualcom- plexity:i)LocalShannonEntropy[110]toindicatethevariabilityofpixel intensitieswhileaccountingforspatialrelationshipsbetweenpixelblocks; andii)JPEGXLfilesize[111, 112]toindicateinformationdensityin thespatial(lossless)andfrequency(lossy)domainstoaccountforhuman perceptionofspatialfrequency[113],andservingasaplausibleestimator ofhumancomplexityperception[111]. Coherence:Toestimateexplanationcoherencewithknowledge(AI Alignmenttoˆy,ConceptAlignmentto ˆ γ)andwithobservation(Cue Alignmentto ̆ x x x,andPhotoAlignmenttox x x),weencodedeachofthemina 7 WeomittedHispanicduetodatasparsity(toofewidentifies). 8 Comparable to prior work (M = 65.5 in [109]). jointvision-languageembeddingrepresentation(usinganimageandsketch- specificevaluatormodel,CLIP[91]andTASK-Former[114],respectively) 9 , andmeasuredtheircosinesimilarity[117]. 5.2.2 Visual Explanation Comparisons We compared SketchXplain against other visualizations: i)Saliency map (Grad-CAM [1]) which highlights importantexplanatory pixels but as neb- ulous blobs; i)Outlinedrawing which traces key facial features based on landmarks[2]; i)CLIPasso[5] which abstracts lines as a descriptive sketch butnotanexplanatoryone,sinceitdoesnotexplicitlyconsiderexplanatory concepts.Allsketcheswererenderedwithequal24strokes 10 .Weincluded iv)Salient Photo,whichoverlaysasaliencymaskontheoriginalimageto improveusability,butitleakssourceinformationoftheimagepixels,giving itanunfairadvantagetootherexplanations.v)Photowasincludedonlyasa referencebaselineforthegoldstandardofmaximuminformation,butitis notanexplanation;itistheinputphotoindependentoftheAIlogic. 5.2.3 Proxy Evaluation Results Fig.4showstheresultsacrossmetricsforthecomparedvisualizations. Visualcomplexity(Fig.4a).BothlocalentropyandJPEGXLsizeexhib- itedsimilartrends.PhotoandSalient Photoweremostcomplexwithpixel- leveldetails.Conversely,Outlinewasthesimplest,sinceithadthefewest strokesandvisualfeatures.Sketchvisualizations(CLIPasso,SketchXplain) occupiedamiddleground.SaliencyhadthelowestJPEGXLfilesizes,yet higherlocalentropythanlinedrawings;thisreflectssomeinconsistencyin themetrics.Nevertheless,SketchXplainwithintuitivefacialcuesissignifi- cantlysimplerthanthepopular,usableSalientPhoto. Coherencewithknowledge(Fig.4b).Alllinedrawingswerehighly alignedtotheAIpredictedlabelsˆyandConcepts ˆ γ ,CLIPassoand SketchXplainwerehighest,closetothePhotogoldstandard;whileSaliency wasleastaligned,likelyduetoitsamorphousshapes.SketchXplainhadthe bestConceptAlignment,indicatingitsexplanatorypower. Coherencewithobservation(Fig.4c).SketchXplainhadthebestCue Alignmentto ̆ x x x,consistentwithitsstrokeinitializationandoptimization constraints.Outlineretainedmoderatealignmentduetopreservedfacial structure,whileSaliencywaspoorlyaligned.ForPhotoAlignmentto x x x,Salient Photowashighestasexpectedduetopreservedphotopixels. CLIPassohadhigheralignmentthanSketchXplain,indicatingatrade-off forphotofidelityagainstexplanatorycues. Insummary,theseresultssuggestthatsketchexplanations—particularly SketchXplain—preservetask-relevantconceptsandcueswhilereducing visualcomplexitycomparedtopixel-basedvisualizations.However,weac- knowledgethattheproxymetricsmaynotfullycaptureperceivedsimplicity andsemantics,especiallyacrossdiverseimagemodalities.Therefore,we furtherevaluatedhumaninterpretationnext. 9 We also examined other evaluators (DINOv2 [115] and Grounding- DINO [116]), but they were insensitive to all AUs, likely due to limited pretraining. 10 See Appendix A.1.3 for our ablation study on stroke count. Table 2: Examples comparing visualizations (cols) for explaining face expressions (rows). Disgusted Nose wrinkler Lip corner depressor Lower lip depressor Happy Cheek raiser Lip corner puller NeutralNil Sad Inner brow raiser Brow lowerer Lip corner depressor Surprised Inner brow raiser Outer brow raiser Upper eyelid raiser Jaw drop Angry Brow lowerer Upper eyelid raiser Eyelid tightener Lip tightener SketchXplainConcepts (AUs)LabelPhotoSaliencySalient photoOutlineCLIPasso Visualization type Saliency Landmark CLIPasso SketchXplain Observation Alignment 0 0.5 1.0 Salient Photo Photo Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Observation Alignment Visualization type Saliency Landmark CLIPasso SketchXplain Observation Alignment 0 0.5 1.0 Cues Observation Cue Alignment Photo Alignment c) Visualization type Saliency Landmark CLIPasso SketchXplain Knowledge alignment 0 0.2 0.4 Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Knowledge Alignment b) 01 Visualization group Visualization type Saliency Landmark CLIPasso SketchXplain Salient Photo Photo Knowledge alignment 0 0.2 0.4 Saliency Landmark CLIPasso SketchXplain Salient Photo Photo AI Alignment Concept Alignment 01 Visualization group Visualization type Saliency Landmark CLIPasso SketchXplain Salient photo Photo Local entropy 0 1 2 3 4 Saliency Landmark CLIPasso SketchXplain Salient photo Photo JPEG XL fi le size (KiB) 0 25 50 75 100 01 Visualization group Visualization type Saliency Landmark CLIPasso SketchXplain Salient photo Photo Local entropy 0 1 2 3 4 Saliency Landmark CLIPasso SketchXplain Salient photo Photo JPEG XL fi le size (KiB) 0 25 50 75 100 Local Entropy JPEG (Lossless) JPEG (Lossy) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain a) Local Entropy JPEG XL file size (KiB) Figure 4:Resultsofmodelingproxyevaluationofvisualizationa)simplicityandCLIP-basedcoherencetob)knowledgeandc)observationacrosslinedrawingsandsaliency explanations.SeeAppendixFig.7forTASK-Formercosinesimilaritycoherenceresults. Saliency, Outline, and CLIPasso are baseline explanations,Photoisagoldstandard referenceandnotanexplanation,andSalientPhotoisahybridexplanationthatpartiallyincludesPhotoinformation. Error bars indicate 95% confidence intervals. 5.3 Qualitative User Study Having investigated that SketchXplain can provide coherent yet simple visual explanations for face expressions in the modeling study, we next aim tovalidate these effects with real users. We conducted a qualitative user study with the think-aloud protocol to understand: 1) how people inherently identify and interpret expressions from face photos, and 2) how various visualizations could help people to interpret face expressions. 5.3.1 Method and Procedure We recruited 19 participants from a university mailing list. They were 9 males, 10 females, with ages 19–25 (Median = 22.0). We conducted the study via a recorded Zoom session,withconsent. The experiment took 50–65 minutes and each participant was compensated $11.50 USD in local currency. Participants viewed separate face visualizations across 9 trials without time constraint. For each trial, they were asked to identify the facial ex- pression from a randomly selected visualization—Photo, Outline, CLIPasso, SketchXplain, Saliency, or Salient photo. In separate trials, visualizations were shown alone or overlaid on the photo, with different types presented for the same original image to support reflective comparison. Whileidentifying the expression in the visualization,participantswere instructed to think aloudaboutcuehelpfulnessanddifficulties.Furthermore,they could freely return to different trials for comparison. 5.3.2 Findings We conducted a thematic analysis on recorded interviews, focusing on: i) how people recognize face expressions from photos and i) how they use or misuse various visualizations to interpret expressions. In general, across different visualization types, participantsinterpreted expressions toconstituent facial features, such as inferring happy from “a smiling mouth with teeth showing and smiley eyes” [Participant P8], or identifying a sad face by “a droopy mouth, watering eyes with slanted eyebrows” [P8]. They also described expressions in terms of dynamic wrinkles and muscle contraction. For example, P7 perceived surprise from “the wrinkles on forehead and the widely opened eyes”, while P11 identified happiness based on “upward tension around his cheeks”. Some participants quickly interpreted expressions based on the shapes of facial features, as P8 noted “look at the shape of his mouth, especially over here. (the target’s mouth)”. Others performed a more in-depth analysis by referring to action units cues and weighing their contribution toward plausible expressions. Contemplating a sad face, P3 noted “his eyebrows are raised and his eyes are downturn, so I wasn’t sure whether this part suggested sad or disgust. However, the pout here is quite prominent.” Participants often pointed at or annotated on key areas of the face when they were unable to verbally articulate the cue. While pointing at the cheeks, P7 remarked that “the skin on her face is generally expanding outwards, though I don’t really know how to justify it”. P13 combined annotation with verbal explanation: “the differentiating factor is the mouth like this”, while drawing a curvy line around mouth, “when you’re disgusted after eating something, your mouth looks like this”. Take-away: Interpretation went beyond verbal Concepts and Saliency pointing, and required complex shapes and visual annotations that were hard to articulate. This justifies the need for sketch explanations. Participants identified various advantages and limitations when interpret- ing expressions from different visualizations.Salient photo, which masked out less important regions on the face, allowed participants to verify align- ment between their mental model and salient regions. P13 felt “I should be correct since I see the curvy mouth highlighted by the AI”. However, participants generally foundSaliency mapsunhelpful since they were not overlaid on the photo, as they “don’t see any semantics about the face or the expression” [P4].Outline, which captured the shape of facial features, were considered “concise and easy to interpret the surprised expression ... clearly captures the opened mouth and upturned eyebrows” [P7]. However, participants were also concerned about its limited details from the original face. P13 noted “[couldn’t] see the wrinkles around cheeks and nose”. and was also misled to think “the photo seems more surprised, but the visual- ization looks more happy to me.” Compared to the Outline,CLIPassohad more expressive lines that were evocative of expressions. P2 described it as “intuitive and giving a feeling of negative expressions”. However, the draw- ings were sometimes unfaithful to the original face, with P4 commenting “because the lines are incomplete and don’t align with facial features” and P13 noting “seems to match the face well, but too many lines make it hard to tell the surprised expression”. SketchXplainwasintuitivetousers;forexample,P3foundthat“based ontheSketchXplainvisualization,it’smoreobviousshe’sfeelinghappy”. Unlike CLIPasso, SketchXplain had morecoherent strokes that “highlighted what I was referring to when I identified happy from the photo” [P1]. This helped P7 “realize [a face] is not happy but angry if I have paid atten- tion to the combo of arched eyes and frowning eyebrows.” Like Saliency, participants appreciated thesimplicity to emphasize onsalientlines by SketchXplain, noting that “from these darkest strokes, I can tell it’s clearly a happy face because the upward motion [of her lip corner], the scrunched-up eyes and arched eyebrows indicate she’s smiling” [P1]. By sourcing from fine lines,SketchXplaincould also effectively convey subtle cues, such as “the darker wrinkle lines around nose are straightforward to indicate disgust” [P8]. This even helped to augment their perception, for example, P1 “didn’t notice these lines until I saw the visualization.” Meanwhile, participants also noted some limitations withSketchXplain, such as missing information like “I can’t tell the eyeballs and the gaze direction” [P14]. To convey relative importance between AUs, SketchXplain varied stroke tone darkness [118], but this confused some users who were “distracted by other shallow lines” [P12]. Take-away: SketchXplain successfully conveyed concepts through recognizable facial features (like Outline), cues through expressive lines (like CLIPasso), and relevance using darker strokes (like Saliency);thus supportingseveraldesiderataforexplanatoryintuitiveness. 5.4 Quantitative User Study Havingfoundthatparticipantsinthequalitativeuserstudycouldintuitively interpretsketchexplanations,wenextquantitativelyevaluateifsketchexpla- nationscanbeintuitivelyinterpreted,whereausercanquicklyassimilate relevantcuesandconcepts,andcorrectlyinfertheAI’spredictionlabel. Inspired by Kendall et al. [119] which evaluated rapid human recognition of icon-based face expressions,todeterminehowquicklytherelevantin- formationiscorrectlyinterpreted,wedisplayedvisualizationsinverybrief durations(<500ms). We aimed to answer the research question: How well can participants interpret Face Expression predictions when viewing different Visualizations, for varying Display Durations? 1000 ms100 ms100 ms푡 Figure 5: Experiment apparatus of image sequence shown to participants per trial in the quantitative user study: 1) Centered crosshair shown for 1000.0ms to focus the participant’s attention, 2) Random lines shown for 100.0ms as distraction, 3) Visual- ization of randomly chosen Visualization Type and expression shown for randomly chosen Display Durationt(33.3–500.0ms) as main effect test, 4) Random lines shown for 100.0ms as distraction, before showing questions to label expression and AUs. 5.4.1 Experiment Design and Apparatus We conducted a within-subjects experiment with three independent variables (IVs): Visualization Type (Photo only 11 , Saliency, Salient photo 12 , Outline, CLIPasso, SketchXplain 13 ; seeTable2forexamples), Target expressiony (Surprised, Happy, Neutral, Sad, Disgusted, Angry), and Display Durationt (33.3ms 14 , 66.7ms, 100.0ms, 133.3ms, 166.7ms,200.0–500.0ms 15 ).Fig.5 showstheexperimentapparatuswithtimedvisualization. Ourapparatus user interface measured the screen refresh rate of each user’s monitor to ensure the correct duration for each online participant.Weevaluatedwith the same 72instancesas in the modeling study. For dependent variables (DVs),aftereachviewedvisualization,wemea- suredAIAlignmentoftheperceivedlabel 16 (multiple-choicequestion)to theAIlabelˆy,andConceptRecall(multiple-responsequestionofupto threefaceAUsexplanationcues),calculatedasTP/(TP+FN),foreachAU concept ˆ γ . 5.4.2 Experiment Procedure Each participant completed the following procedure: 1) Introduction to the study. 2) Consent to participate. 3) Tutorial on interpreting all face visualizations. 4) Screening questions to ensure ability to correctly: a) match expressions to various straightforward photos, b) recognize AUs from another photo, and c) interpret visualizations. 5)Main study with 72 trials, each involving viewing a rapid series of images with the test instance (Fig. 5), andanswering questions. A reference of AUs to expressions is provided to aid users. 6) Answer demographic questions. 7) Acknowledge bonus calculations and exit. See Appendix Figs. 11–16 for questionnaire details. 5.4.3 Participants We recruited 466 participants from Prolific with high qualification rates (≥1000 completed HITs,>97% approval rate). 176 participants passed our 11 Weincludedtheoriginalphotoasagoldstandardvisualizationofmaximum information,thoughitisnotanAIexplanation,sinceitisjustthemodelinput. 12 WeincludedthisasitisapopularXAImethod,thoughthisleaksinformation fromtheinputphoto,leadingtounfairadvantage. 13 We mostly excluded visualizations overlaid on the photo, to compare information gained purely from the visualizationwithoutleakage from the input photo. 14 Thedisplaydurationswerechosentobeasrapidaspossibleformoderncomput- ers.Sincecomputermonitorsmayhaveaslowrefreshrateof30Hz[120],welimited theshortestdurationtobe1frame,i.e.,33.3ms = 1s/30. 15 Weincludedpilotresultsfor200.0–500.0mstoshowresultsareconsistenteven withlesstimepressure.Wehadconductedapreliminarystudyof80participantsfor awiderrangeofdurations33.3–500.0ms,andfoundnosignificantdifferencesfrom ≥166.7msonward,solimitedthemainstudyto≤166.7ms. 16 Participantswereaskedtorecognizefaceexpressionasconveyedbythevisualiza- tiontomeasureitscommunicationvalue,ratherthaninferorlearnwhattheAIwould havepredicted.ThisissubtlydifferentfromtheforwardsimulatabilitytaskinXAI userstudiesthatfocusonAIunderstandingthanexplanationcommunication. 0 0.2 0.4 0.6 AI Simulatability Saliency Landmarks CLIPasso SketchXplain Salient photo Photo Visualization type a) AI Simulatability Saliency Landmarks CLIPasso SketchXplain Visualization type AI Simulatability Salient photo Photo Visualizat ion type 0 0.1 0.2 0.3 AU Recall Saliency Landmark CLIPasso SketchXplain Visualization type 0 0.1 0.2 0.3 AU Recall Saliency Landmark CLIPasso SketchXplain Visualization type Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain AI Alignment Concept Recall (Upper Face AUs) Display Duration (ms) 33.3 66.7 100.0 200.0–500.0 166.7 133.3 Concept Recall (Lower Face AUs) Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain b)c) Figure 6: Resultsofthequantitativeuserstudyonfaceexpressionquickinterpretation across Visualization TypesandDisplayDuration for measures: a) AIAlignment,b) ConceptRecall(UpperFaceAUs),c)ConceptRecall(LowerFaceAUs).GrayVisualizationTypes:Photoisagoldstandardofmaximuminformation,notacompetitor explanationmethodsinceitistheinputimage,andSalientPhotoismoreusablethanSaliencybutunfairlyleaksphotoinputdata. Gray dotted line in (a) represents correctness of a random guess from 6 MCQ choices (16.7%)forAIAlignment. screening test. They were89males,84femalesand3preferrednottosay, with ages 21–71 (Median = 37.0). They completed the survey in median 32.8 minutes and were compensated £4.00. We incentivized effort with up to £1.00 for more correct expression labeling. 5.4.4 Statistical Analysis and Quantitative Results We fit a linear mixed effects model for each dependent variable as the response, Visualization Type, Expression and Display Duration, with other confounding variables as fixed effects, some interaction effects among the factors, and Participant as a random effect. See Appendix Table 7 for details. Participants’perceivedlabelsshowedhigherAIAlignment as the Display Duration increased across all visualizations, except for Saliency (Fig. 6a). When the Display Duration was extremely short (t = 33.3ms;lightpurple line), both Outline (p<.0011) and SketchXplain (p = .0070) achieved significantly higheralignment compared to Photo (Fig. 6a), indicating their quicknessinconveyinginformationabouttheAIpredictions. Photo was harder for participants to interpret at this short duration, perhaps due to excessive details,thoughatlongerdurations,thisgoldstandardservedasan upperlimitindicationofhowmuchinformationparticipantscouldacquireat eachtimeduration. In general, CLIPasso sketches were significantly poorer in supportingAI-aligneduserperception compared to SketchXplain (p< .0032).Saliencywastheworstinhelpinguserstoaligntheirperception,per- hapsduetothelackoffeaturalcontextofitsamorphousblobs.Conversely, withtheaddedcontextfromtheretainedregionsoftheinputphoto,Salient photosweremoreinterpretable(higherAIalignment)thanSaliencyalone. ConceptRecallalso improved with increased Display Duration, and had similar trends across Visualization Types as for AI Alignment(Fig.6b,c). Interestingly, participants could recall Lower Face AUs (about mouth, nose) better than Upper Face AUs (eyes, eyebrows);contrastt-test:p<.0001. WithsufficientDisplayDuration(t> 33.3ms), SketchXplain had the high- est recall among non-Photo Visualization Types (p<.0001), indicating the usefulness of its semantically-aligned, cue-aware sketches. Outline was weakerthanSketchXplain at supporting Lower Face AUs(p<.0001), pos- sibly due to the similar mouth shape (open or closed) across expressions. CLIPasso had even weaker concept recall(p<.0001) due to not explicitly encoding AU information. Saliency explanations weremost unhelpful(p <.0001).SalientPhotodidimproverecallforlongerdurations(> 100.0ms, p<.0001),butstillworsethanSketchXplain(p<.0001). 6 EVALUATION ON SKIN LESION CLASSIFICATION To test the generalizability of sketch explanations, wealso applied SketchX- plain to skin lesion classification, a domain with well-established, lay- accessible concepts—the ABCDE criteria [126] for melanoma diagnosis. Weinvestigatehow explainable sketches could help non-specialists identify suspicious melanoma at an early stage [127], enhance clinical trust and facilitate communication between clinicians and patients [128]. We describe the dermoscopic image preparation, the SketchXplain ex- tension to extractABCD 17 concepts, and a qualitativeuserstudytolearn insights into the potential use and benefits of sketch explanations for com- municating melanoma risk. We omitted a modeling study with CLIP scores, because CLIP was not trained on medical images, leading to unreliable out-of-distribution measurements. We also omitted a quantitative user study withtightdisplaydurations as ABCD concepts are not inherently familiar like facial action units and thus require substantial user training. 6.1 Data Preparation We used 11,720 images of skin lesions from the HAM10000 datasetwithdi- agnosislabels [129] andsegmentationmasksoflesionlocations[3]. We sim- plified the lesion prediction to binary classification (i.e., benign or melanoma) by grouping all benign classes 18 and excluding malignant classes 19 other than melanoma. Future work could extend explanations to other malignant lesions with distinct clinical features. 6.2 Concept Labeling and Cue Extraction We reused the SketchXplain architecture, with additional data labeling and feature extraction. Since the skin-lesion dataset lacked ABCD labels, we estimated these concepts using simple image-processing heuristics adapted from prior work [122–125]:Asymmetry(unevenhalves),Borderirregularity (raggededges),Colorvariation(multipleshades),Diameter(largerthan6 m).Table3describestheABCDconceptsandvisualcuefeatureextraction methods. The resulting ABCD scores served as pseudo-labels for training the concept extractor in our concept bottleneck model (MandF γ ).For inference, SketchXplain first predicts the concepts ˆ γ (Fig. 3, Step 2a) and then the classˆy. For cue localization (Fig. 3, Step 2b), we regularized the Grad-CAM with ABCD heuristic heatmaps.Weusedthesamepipelineto generatesketchesasforfaceexpressions 20 . For semantic alignment, we disabled text-based guidance (settingλ t = 0in Eq. 1) because CLIP is not trained on medical images. Table 4 compares visualization types (columns) across different cases (rows) with varying diagnoses and ABCD features. Like[130],wefined-tunedResNet18[131]asthebackboneM,onthe HAM1000dataset.SketchXplainachievedgoodaccuracy85.8%andF1 Score71.2onmelanomaclassification,andABCDper-criteriaF1scores 17 FullcriteriaisABCDEincludingEforevolution,whichrequiresrepeatedmea- suresofmultiplephotostoobserve,soweomitteditforoursingle-imageapplication. 18 Benign: Benign Keratoses, Dermatofibroma, Melanocytic Nevi, Vascular lesions. 19 Malignant: Basal Cell Carcinoma, Melanoma, and Pigmented Actinic Keratoses. 20 SeeAppendixTable9,columnInit.strokesforexamplesofinitialization. Table 3: ABCD criteria for melanoma diagnosis. Given a skin lesion, letAdenote its area,Pits perimeter andL L Lits outline. Higher values indicate more suspicious features. CriteriaConcept Description and Label Quantification MethodVisual Cue Feature Extraction Method Asymmetry A lesion islabeled symmetrical if it resembles its reflection across its major and minor axes.Wedefinenon-overlappingpixelsastheXOR(⊕)betweenthemask(M)andits reflection(M r )alongtheseaxes[121–123]andquantifyasymmetryastheproportionof non-overlappingpixels,i.e.,∥M⊕ M r ∥ 1 /A. Wehighlightnon-overlappingpixels(M⊕ M r )betweenthelesion maskanditsreflectionsalongmajorandminoraxes. Border Irregularity Perimeter lengthPis larger for lesions with irregular borders. We use the perimeter–area ratio P A , and the non-circularity index P 2 4π A to measure deviation from a circular shape [121, 122, 124]. We highlight boundary segments corresponding tohigh curvature alongL L Lwith a sliding window and marking boundary segments whose curvature exceeds a threshold. Color Variation More colorful lesions are more suspicious. We capture overall RGB color variation using the standard deviation of each channel [122]. We highlight lesion pixels whose color deviates from the mean lesion color by more than one standard deviation in any RGB channel. Diameter Larger lesions are more likely melanoma. We estimate the equivalent diameter for irregular shapes, i.e., 2 p A/π , which is the diameter of a circle with the same area [124, 125]. We highlight the lesion outline L L L as a cue to overall lesion size. Table 4: Examples comparing visualizations (cols) for explaining suspicious melanoma (rows). BenignNil BenignAsymmetry BenignBorder irregularity MalignantColor variation Malignant Border irregularity Color variation Diameter oversize SketchXplainConcepts (ABCDs)LabelPhotoSaliencySalient PhotoOutlineCLIPasso 71.6–82.1(M=78.5) 21 . Whiletheaforementionedresultsdemonstratedgoodin-distributionpre- dictionperformance,thepretrainedimage-to-imagegeneratorE[85]maybe out-of-distribution.Nevertheless,thegeneratedDetailedlines(seeAppendix Table9)werequalitativelyfaithfultothephotos.Next,wefurtherexamined theirplausibilityinaformativequalitativeuserstudy. 6.3 Qualitative user study We conducted a qualitative user study using the think-aloud protocol to investigate how different visualizations(Saliency,Salientphoto,Outline drawing,CLIPassosketch,orSketchXplain) could help people identify concepts and melanoma cases. 6.3.1 Visual Explanation Comparisons WecomparedSketXplainagainstthesamebaselinevisualizationsdescribed inSection5.2.2,withonedifference:Outlinewasobtainedfromtracingthe edgeofthesegmentationmask[3]. 6.3.2 Method and Procedure We recruited 10 non-clinical participants through university mailing lists and personal contacts. The sample included 3 femalesand7males, ages 23–34(Median=29.0). From images of skin lesions, they were asked to identify ABCD concepts and determine the melanoma type based on various visualizations. To prepare participants, we provided a brief tutorial 21 BetterthanfaceexpressionclassificationinSection5.2andpriorwork[109]. describing how ABCD concepts relate to melanoma, and familiarized them with 7–10 examples across all visualizations. In the main study, participants viewed 15–20 image instances, each with a random visualization. Since skin lesions are less familiar to lay people than faces, we overlaid visualizations on the photos for added context, rather than using sketches in isolation on a white background. For generality, this also allowed us to evaluate sketch explanations as annotations. Participants described what they saw, decided on the diagnosis, and explained their reasoning. All sessions were conducted over Zoom, and audio and screen interactions recorded with participant consent. The study took 32 minutes on average, and participants were compensated $6.20 USD in local currency for their time. 6.3.3 Findings Weperformed a thematic analysis on participant utterances and interactions, focusing on how participants used or misused visualizations to identify concepts and diagnose melanoma. Unlike participants who perceived facial action units inherently in face expressions, here, participants were more deliberative and analytical. Participant P2 focused on “balancing the number of abnormal concepts and the severity of each abnormality”. P1 “prioritized certain concepts such as color variation and diameter”.However,when viewingPhotoalone,withoutexplicitABCDindicatorsorannotations,par- ticipantsfoundithardtodistinguishbetweennormalfromabnormalfeatures. P4 found it “hard to decide whether this is small or larger diameter without a threshold”. P6 “[could] not tell whether this [lesion] [was] [as]symmetrical if I can’t even see a clear boundary”. Next, we articulate differences across visualization types. Saliency map orSalient photowere perceived as less helpful thanall line drawings, because they imprecisely depicted the boundary and size. P4 felt the smooth boundary “too blurred, I can’t see the true border”. P4 also foundSalient photomisleading, since “the[grayscale] uneven color caused by the mask misleads me into thinking the lesion has color variation”. In contrast, line drawing annotations on the photo gave participants a strong first impression and encouraged them to relate observations with the clinical ABCD concepts to make more informed judgments. P1 appreciated Outlinebecause [it] provided “simple and straightforward annotations on the boundary”. However, P2 disagreed, notingitslimitation that“italways hasthesameintensity.Itmightbegoodforasymmetryandborder,butinfact wecanseetheborderalonesoitisnotveryuseful.”Conversely,simplifying linespresentedotherconcerns;CLIPasso’s imprecise spatial annotations often misled interpretation: P7 commented that “the visualization cuts it in half, so it might be asymmetry” and P8 noted, “from the photo [the] left part of the lesion is bad, but it isn’t highlighted”. Nonetheless,CLIPassostimulated more reflection. P3 noted that its “strokes inside the region reminded me to examine the potential color ab- normality”. Furthermore,SketchXplainwas perceived as more aligned to the original photo and more coherent with the underlying concepts. P2 remarked, “the emphasized scribbles around the corner help me judge asym- metry” and P6 appreciated that “the disconnected lines help me confirm border irregularity”, whereasifhehadviewedOutline, he “would judge the border differently when viewing a smooth annotation”. However,the increased interpretability with bothCLIPassoandSketchXplaincame at the cost ofincreased cognitive load. P1 found that “the lines are more complex than the [Outline] contour, which takes time to comprehend”. Take-away: SketchXplain produced interpretable strokes that were more coherent with theexlanatory ABCD concepts andlesion observations, better supporting users to makeconfident judgments on the task. 7 DISCUSSION Having introduced sketch explanations as a new paradigm for intuitive visual explanations, we discuss their generalization,contextualization, and further development. 7.1 Extending Sketch Explanations OurworkwasthefirsttoproposesketchvisualizationsasexplanationsofAI predictions,andwehadinvestigateditforalimitedscope.Here,wediscuss extensionstoothervisualproperties,explanatoryconcepts,andapplication domains. ThroughSketchXplainwehadinvestigatedthesimplicity–coherence trade-offwiththevisualpropertiesofstrokecount,smoothnessandtone. However,otherstrokeproperties,suchasstrokelength,thickness,andta- pering,couldbeleveragedtoconveysalientexplanatorycues.Futurework couldevencomparethedifferenceinperceivedacuity,intuitivenessand explanationcuerecallacrossthesevisualfeatures. Wehadusedconceptbottleneckmodels(CBMs)[6]withpre-deteremined conceptsforfacialactionunits(AUs)andskinlesioncriteria,butthiscan beextendedtosemi-supervisedLabel-FreeConceptBottleneckModel(LF- CBM)[132]andopen-domaincuelocalizationwithwithOwl-ViT[133]and SegmentAnythingModel(SAM)[134]thatweinvestigatedinAppendixC.2 forgeneralimages. Other methods to obtain concepts include TCAV [30], crowdsourced elicitation [135] or unsupervised discovery [136]. While sketches are commonly used by artists to conveyfacial expressions, we have demonstrated that they can also be applied to other domains, such asannotatingskinlesions.Sincebiologyandmedicineheavilyemploy sketches[137, 138],sketchexplanationscanprovidenewopportunities toexplainpredictionsonmedicalimages[4] with sketched annotations on pathology or radiology slides to articulate fine physical or anatomical features, rather than merely pointing at them with saliency maps [139, 140]. Furthermore,futureworkcouldinvestigatethetunabilityofsketchestoward simplicityorcoherenceforeverydaylaytasksorscientific,medical,and engineeringtasks.Inthelattertasks,diagrammaticconstraintscouldalsobe addedtoenforcedomainalignment[63]. 7.2 From Analytical to Intuitive AI Explanations WhilemanyexplainableAI(XAI)methodsfacilitateanalyticalreasoning throughchartsornetworkvisualizations[12, 13, 44, 141–144],userstypi- callyconservecognitiveeffort[57]andsatisficetheirunderstanding[145], leadingtoprematuremisinterpretationsanddecisionerrors[59].Hence, XAImustbeintuitive. Priorworkshavesoughtthisthroughsimplicityviasparsity[12,146]and smoothness[31,141],yetthesepropertiesaloneareinsufficientforintuitive- ness.Othersindirectlyaddressexplanatorycoherenceviarelatability[24] andfaithfulness[12].Infeature-basedXAIontabulardata,methodsreplace technicaljargonwithhuman-interpretabletermsorconcepts[147],orsigni- fiericons[65].Moreover,concept-basedexplanations[6,30]couldbemade moreintuitivewithanalogiesusingpriorknowledge[148]. Incontrast,methodsfornon-tabulardatagobeyondsemanticremapping toadoptdomain-specificcuesandconventions.Forinstance,audio-based XAIcanbemaderelatablethroughcounterfactualcasesandcontrastive cues[24].Inimage-basedXAI,saliencymaps(e.g.,[1])remainthedomi- nantmodalityduetotheirintuitivehighlighting,yettheycanbeincoherent tomodelinference[149]orhumanbeliefpriors[150].Whilesomemethods improvecoherencethroughregularizingattributionpriors[97]anddiagram- maticconventions[63],weintroducethesketchasanewvisualmodality thatleverageslayfamiliaritytoenhancebothexplanationsimplicity[10], coherence[33],andinterpretationspeed[32, 35].However,ascurrentmeth- odsrarelyprioritizerapidinterpretation,weemployspeed-constrainedtests fromcognitivepsychology[119]topreciselyevaluatetheperceptualeffi- ciencyofexplanations. Therefore,byprioritizingsimplicity,coherence,andspeed,sketch-based XAIaccommodatesthehumantendencytosatisfice,therebylimitingcogni- tiveloadandstreamliningexplanationtransmissiontotheviewer.Thisset ofdesiderataprovidesabasisfordevelopingandevaluatingimage-based XAIthatshiftsthefocusfrompurelyanalyticalauditingtowardintuitive interpretation. 7.3 Standalone Abstraction vs. Annotative Overlay Sketches Intwodomains,wehaveexaminedtheusabilityandusefulnessofsketches asexplanationsindifferentpresentationparadigms.Sketchexplanationsof facialexpressionswereintuitivelyinterpretableasstandalone,sincehu- manspossessaninnatefamiliaritywithfacialgeometry,allowingthem todecodelineshapesandspatialrelationshipswithoutexternalcontext. Methodologically,evaluatingthesesketchesinisolationallowedustoex- aminetheirindependentexplanatoryvaluewhilepreventinginformation leakagefromsourcephotos(furtherstudiedinAppendixA.3). Conversely,forskinlesionsketchexplanations,toaccommodatelimited familiarityordomainexpertise,weincludedthephotoascontextualground- ingfortheoverlaidsketches.Thisdependencyisanalogoustofeatureattri- butionexplanationsa(x)thatarelessinterpretablewithoutcorresponding featurevaluesx,andtosaliencymapsthatareobscurewithouttheunder- lyingsourcepixelsbeingvisible.Consequently,itisvitalforvisualexpla- nationstobedemonstrablycoherent[33]oralignedwithobservation[63]. Overlayingsketchesonphotosfacilitatesthisbyallowinguserstoperform sanitychecksagainsttheinputgrounding[53, 63]. Futureworkshouldinvestigatethisextendedusageofoverlaidsketch explanationswithdomainexperts.Asalientresearchdirectionisthetrade- offbetweensketchabstractionandspatialalignmenttothephotograph. Whilehighlyabstractedsketchesmayfacilitaterapidinterpretation,they maycompromisethedetailedexaminationrequiredforscientific,clinical,or engineeringrigor,suggestinganeedfordynamiclevelsofdetailinsketch- basedXAIinterfaces. 7.4 Evaluating Intuitive Understanding We had evaluated visualizations under tight time limitsinthequantitative userstudyonquickperceptionoffaceexpressionvisualizations. This is common in human perception psychology studies [151], and allows precise measurement of users’ intuitive impression (System 1 thinking [71]). We do not argue that sketch explanations will only be beneficial under rapid expo- sures, which future work can evaluate, but differences across visualizations maydiminish when users rationalize slowly (System 2) [36]. While poorlabel recognition orconcept recall under time pressure could indicate visual complexity, we did not explicitly measure perceived visual complexity. Nevertheless, prior work has shown the congruency between the objective measure of JPEG file size and perception [111],supporting toourapproach. Furthermore, we did not measure participants’ subjective self-reported satisfaction of each visual explanation, or their ability to debug AI prediction errors. The latter is especially important for mitigating AI over-reliance [152, 153], which future work can explorewithtechnicalusers. Emojis and icon-based faces are popular for illustrating face expressions and can be quickly perceived by users [119]. However, we did not evaluate them as baselines since they are static per expression and act more like labels ratherthanexplanationscontextualtotheinstancefeaturesorobservation. Finally, although SketchXplain explanations convey conceptual infor- mation viavisual cues,theycouldbeconveyedverballybylabelingwhich conceptarepresent[6]. Future work could investigate this further, though we hypothesize that visual explanations will remain more intuitive because they exploit higher human visual bandwidth instead of slower general language understanding, andavoidAUterminology(e.g.,“lipcornerdepressor”)that is known to be difficult for lay users to understand [154, 155].Neverthe- less,text-basedconceptscanbecomplementarytosketches,andshouldbe investigatedinfuturework. 8 CONCLUSION We have introduced SketchXplain as a new XAI paradigm tovisually explain image-basedAIpredictions intuitively and faithfully. It incorporates concept- based and saliency explanations to generate cue-aware, semantically-aligned sketches from relevant source lines. Through modeling and user studies on face expressionandmelanoma image predictiontasks, we demonstrated that sketch explanations support intuitive interpretation,throughsimplerand morecoherentvisualexplanationsthatarequickertointerpretcomparedto saliencymapsandotheroutlinedrawingsornon-explanatorysketches. This work contributes to the diverse options ofvisual XAI by offering intuitive and expressive sketch rationalizations to improve interpretability. REFERENCES [1]R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based local- ization,” in Proceedings of the IEEE international conference on computer vision, 2017, p. 618–626. [2]V. Kazemi and J. Sullivan, “One millisecond face alignment with an ensemble of regression trees,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, p. 1867–1874. [3]P. Tschandl, C. Rinner, Z. Apalla, G. Argenziano, N. Codella, A. Halpern, M. Janda, A. Lallas, C. Longo, J. Malvehy et al., “Human–computer col- laboration for skin cancer recognition,” Nature medicine, vol. 26, no. 8, p. 1229–1234, 2020. [4]H.-P. Chan, R. K. Samala, L. M. Hadjiiski, and C. Zhou, “Deep learning in medical image analysis,” Deep learning in medical image analysis: challenges and applications, p. 3–21, 2020. [5]Y. Vinker, E. Pajouheshgar, J. Y. Bo, R. C. Bachmann, A. H. Bermano, D. Cohen-Or, A. Zamir, and A. Shamir, “Clipasso: Semantically-aware object sketching,” ACM Transactions on Graphics (TOG), vol. 41, no. 4, p. 1–11, 2022. [6]P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang, “Concept bottleneck models,” in International Conference on Ma- chine Learning. PMLR, 2020, p. 5338–5348. [7]S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 6, p. 1137–1149, 2016. [8]A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun, “Dermatologist-level classification of skin cancer with deep neural networks,” nature, vol. 542, no. 7639, p. 115–118, 2017. [9] Z. C. Lipton, “The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.” Queue, vol. 16, no. 3, p. 31–57, 2018. [10] T. Miller, “Explanation in artificial intelligence: Insights from the social sci- ences,” Artificial intelligence, vol. 267, p. 1–38, 2019. [11] B. Y. Lim, A. K. Dey, and D. Avrahami, “Why and why not explanations improve the intelligibility of context-aware intelligent systems,” in Proceedings of the SIGCHI conference on human factors in computing systems, 2009, p. 2119–2128. [12]M. T. Ribeiro, S. Singh, and C. Guestrin, “” why should i trust you?” explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, 2016, p. 1135–1144. [13]S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in neural information processing systems, 2017, p. 4765–4774. [14]A. Abdul, J. Vermeulen, D. Wang, B. Y. Lim, and M. Kankanhalli, “Trends and trajectories for explainable, accountable and intelligible systems: An hci research agenda,” in Proceedings of the 2018 CHI conference on human factors in computing systems, 2018, p. 1–18. [15]C. Olah, A. Mordvintsev, and L. Schubert, “Feature visualization,” Distill, vol. 2, no. 11, p. e7, 2017. [16] B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, p. 2921–2929. [17]H. G. Ramaswamy et al., “Ablation-cam: Visual explanations for deep convolu- tional network via gradient-free localization,” in The IEEE Winter Conference on Applications of Computer Vision, 2020, p. 983–991. [18]S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. M ̈ uller, and W. Samek, “On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation,” PloS one, vol. 10, no. 7, p. e0130140, 2015. [19]R. C. Fong and A. Vedaldi, “Interpretable explanations of black boxes by meaningful perturbation,” in Proceedings of the IEEE International Conference on Computer Vision, 2017, p. 3429–3437. [20]G. Montavon, S. Lapuschkin, A. Binder, W. Samek, and K.-R. M ̈ uller, “Ex- plaining nonlinear classification decisions with deep taylor decomposition,” Pattern Recognition, vol. 65, p. 211–222, 2017. [21] A. D. Selbst and S. Barocas, “The intuitive appeal of explainable machines,” Fordham L. Rev., vol. 87, p. 1085, 2018. [22]A. Boggust, H. Bang, H. Strobelt, and A. Satyanarayan, “Abstraction alignment: Comparing model-learned and human-encoded conceptual relationships,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, p. 1–20. [23]H. Kaur, E. Adar, E. Gilbert, and C. Lampe, “Sensible ai: Re-imagining interpretability and explainability using sensemaking theory,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022, p. 702–714. [24]W. Zhang and B. Y. Lim, “Towards relatable explainable ai with the perceptual process,” in CHI Conference on Human Factors in Computing Systems, 2022, p. 1–24. [25]A. Alqaraawi, M. Schuessler, P. Weiß, E. Costanza, and N. Berthouze, “Eval- uating saliency map explanations for convolutional neural networks: a user study,” in Proceedings of the 25th international conference on intelligent user interfaces, 2020, p. 275–285. [26]J. Fish and S. Scrivener, “Amplifying the mind’s eye: sketching and visual cognition,” Leonardo, vol. 23, no. 1, p. 117–126, 1990. [27]A. Hertzmann, “Why do line drawings work? a realism hypothesis,” Perception, vol. 49, no. 4, p. 439–451, 2020. [28]D. B. Walther, B. Chai, E. Caddigan, D. M. Beck, and L. Fei-Fei, “Simple line drawings suffice for functional mri decoding of natural scene categories,” Proceedings of the National Academy of Sciences, vol. 108, no. 23, p. 9661– 9666, 2011. [29]D. Retelny and P. Hinds, “Embedding intentions in drawings: How archi- tects craft and curate drawings to achieve their goals,” in Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing, 2016, p. 1310–1322. [30]B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas et al., “In- terpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),” in International conference on machine learning. PMLR, 2018, p. 2668–2677. [31]A. Abdul, C. Von Der Weth, M. Kankanhalli, and B. Y. Lim, “Cogam: measur- ing and moderating cognitive load in machine learning model explanations,” in Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, p. 1–14. [32] X. Wang, M. Yu, H. Nguyen, M. Iuzzolino, T. Wang, P. Tang, N. Lynova, C. Tran, T. Zhang, N. Sendhilnathan et al., “Less or more: Towards glanceable explanations for llm recommendations using ultra-small devices,” in Proceed- ings of the 30th International Conference on Intelligent User Interfaces, 2025, p. 938–951. [33] P. Thagard, “Explanatory coherence,” Behavioral and brain sciences, vol. 12, no. 3, p. 435–467, 1989. [34] M. Nauta, J. Trienes, S. Pathak, E. Nguyen, M. Peters, Y. Schmitt, J. Schl ̈ otterer, M. Van Keulen, and C. Seifert, “From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai,” ACM Computing Surveys, vol. 55, no. 13s, p. 1–42, 2023. [35]S. Swaroop, Z. Buc ̧inca, K. Z. Gajos, and F. Doshi-Velez, “Accuracy-time tradeoffs in ai-assisted decision making under time pressure,” in Proceedings of the 29th International Conference on Intelligent User Interfaces, 2024, p. 138–154. [36]U. Ehsan, B. Harrison, L. Chan, and M. O. Riedl, “Rationalization: A neural machine translation approach to generating natural language explanations,” in Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, 2018, p. 81–87. [37]Y. O. Gat, N. Calderon, A. Feder, A. Chapanin, A. Sharma, and R. Reichart, “Faithful explanations of black-box nlp models using llm-generated counterfac- tuals,” in The Twelfth International Conference on Learning Representations, 2024. [38]D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba, “Network dissection: Quantifying interpretability of deep visual representations,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, p. 6541–6549. [39] F. Hohman, H. Park, C. Robinson, and D. H. Polo Chau, “Summit: Scaling deep learning interpretability by visualizing activation and attribution sum- marizations,” IEEE Transactions on Visualization and Computer Graphics, vol. 26, no. 1, p. 1096–1106, Jan. 2020. [40]M. Liu, J. Shi, Z. Li, C. Li, J. Zhu, and S. Liu, “Towards better analysis of deep convolutional neural networks,” IEEE Transactions on Visualization and Computer Graphics, vol. 23, no. 1, p. 91–100, 2017. [41]M. Kahng, P. Y. Andrews, A. Kalro, and D. H. Chau, “Activis: Visual explo- ration of industry-scale deep neural network models,” IEEE Transactions on Visualization and Computer Graphics, vol. 24, no. 1, p. 88–97, 2018. [42]Z. J. Wang, R. Turko, O. Shaikh, H. Park, N. Das, F. Hohman, M. Kahng, and D. H. Polo Chau, “Cnn explainer: Learning convolutional neural networks with interactive visualization,” IEEE Transactions on Visualization and Computer Graphics, vol. 27, no. 2, p. 1396–1406, 2021. [43]Q. Wang, J. Yuan, S. Chen, H. Su, H. Qu, and S. Liu, “Visual genealogy of deep neural networks,” IEEE Transactions on Visualization and Computer Graphics, vol. 26, no. 11, p. 3340–3352, 2020. [44]J. Krause, A. Dasgupta, J. Swartz, Y. Aphinyanaphongs, and E. Bertini, “A workflow for visual diagnostics of binary classifiers using instance-level expla- nations,” in 2017 IEEE Conference on Visual Analytics Science and Technology (VAST), 2017, p. 162–172. [45]Z. Zhao, P. Xu, C. Scheidegger, and L. Ren, “Human-in-the-loop extraction of interpretable concepts in deep learning models,” IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 1, p. 780–790, 2022. [46]C. J. Cai, E. Reif, N. Hegde, J. Hipp, B. Kim, D. Smilkov, M. Wattenberg, F. Viegas, G. S. Corrado, M. C. Stumpe et al., “Human-centered tools for coping with imperfect algorithms during medical decision-making,” in Pro- ceedings of the 2019 chi conference on human factors in computing systems, 2019, p. 1–14. [47]M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in European conference on computer vision. Springer, 2014, p. 818–833. [48] Z. Qu, Y. Gryaditskaya, K. Li, K. Pang, T. Xiang, and Y.-Z. Song, “Sketchxai: A first look at explainability for human sketches,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, p. 23 327–23 337. [49]H. Bandyopadhyay, P. N. Chowdhury, A. K. Bhunia, A. Sain, T. Xiang, and Y.-Z. Song, “What sketch explainability really means for downstream tasks?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, p. 10 997–11 008. [50]Y. Ming, H. Qu, and E. Bertini, “Rulematrix: Visualizing and understanding classifiers with rules,” IEEE Transactions on Visualization and Computer Graphics, vol. 25, no. 1, p. 342–352, 2019. [51]H. Strobelt, S. Gehrmann, M. Behrisch, A. Perer, H. Pfister, and A. M. Rush, “Seq2seq-vis: A visual debugging tool for sequence-to-sequence models,” IEEE Transactions on Visualization and Computer Graphics, vol. 25, no. 1, p. 353– 363, 2019. [52] M. Kahng, N. Thorat, D. H. Chau, F. B. Vi ́ egas, and M. Wattenberg, “Gan lab: Understanding complex deep generative models using interactive visual experimentation,” IEEE Transactions on Visualization and Computer Graphics, vol. 25, no. 1, p. 310–320, 2019. [53]A. Boggust, B. Hoover, A. Satyanarayan, and H. Strobelt, “Shared interest: Measuring human-ai alignment to identify recurring patterns in model behavior,” in Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, 2022, p. 1–17. [54] G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro, and D. Weld, “Does the whole exceed its parts? the effect of ai explanations on complementary team performance,” in Proceedings of the 2021 CHI conference on human factors in computing systems, 2021, p. 1–16. [55] U. Ehsan, S. Passi, Q. V. Liao, L. Chan, I.-H. Lee, M. Muller, and M. O. Riedl, “The who in xai: how ai background shapes perceptions of ai explanations,” in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 2024, p. 1–32. [56] H. Kaur, H. Nori, S. Jenkins, R. Caruana, H. Wallach, and J. Wortman Vaughan, “Interpreting interpretability: understanding data scientists’ use of interpretabil- ity tools for machine learning,” in Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, p. 1–14. [57]Z. Buc ̧inca, M. B. Malaya, and K. Z. Gajos, “To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making,” Proceedings of the ACM on Human-computer Interaction, vol. 5, no. CSCW1, p. 1–21, 2021. [58]X. Wang and M. Yin, “Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making,” in Proceedings of the 26th International Conference on Intelligent User Interfaces, 2021, p. 318–328. [59]D. Wang, Q. Yang, A. Abdul, and B. Y. Lim, “Designing theory-driven user- centric explainable ai,” in Proceedings of the 2019 CHI conference on human factors in computing systems, 2019, p. 1–15. [60]Q. V. Liao and K. R. Varshney, “Human-centered explainable ai (xai): From algorithms to user experiences,” arXiv preprint arXiv:2110.10790, 2021. [61] J. Y. Bo, P. Hao, and B. Y. Lim, “Incremental xai: Memorable understanding of ai with incremental explanations,” in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 2024, p. 1–17. [62]H. Matsuyama, N. Kawaguchi, and B. Y. Lim, “Iris: Interpretable rubric- informed segmentation for action quality assessment,” in Proceedings of the 28th International Conference on Intelligent User Interfaces, 2023, p. 368– 378. [63]B. Y. Lim, J. P. Cahaly, C. Y. Sng, and A. Chew, “Diagrammatization: Rational- izing with diagrammatic ai explanations for abductive-deductive reasoning on hypotheses,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, p. 1–25. [64]Z. Buc ̧inca, S. Swaroop, A. E. Paluch, F. Doshi-Velez, and K. Z. Gajos, “Con- trastive explanations that anticipate human misconceptions can improve human decision-making skills,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, p. 1–25. [65]B. Y. Lim and A. K. Dey, “Design of an intelligible mobile context-aware application,” in Proceedings of the 13th international conference on human computer interaction with mobile devices and services, 2011, p. 157–166. [66] —, “Evaluating intelligibility usage and usefulness in a context-aware application,” in International Conference on Human-Computer Interaction. Springer, 2013, p. 92–101. [67]A. Springer and S. Whittaker, “Progressive disclosure: empirically motivated approaches to designing effective transparency,” in Proceedings of the 24th international conference on intelligent user interfaces, 2019, p. 107–120. [68]V. Lai, Y. Zhang, C. Chen, Q. V. Liao, and C. Tan, “Selective explanations: Leveraging human input to align explainable ai,” Proceedings of the ACM on Human-Computer Interaction, vol. 7, no. CSCW2, p. 1–35, 2023. [69] B. Y. Lim and A. K. Dey, “Assessing demand for intelligibility in context- aware applications,” in Proceedings of the 11th international conference on Ubiquitous computing, 2009, p. 195–204. [70]Q. V. Liao, D. Gruen, and S. Miller, “Questioning the ai: informing design practices for explainable ai user experiences,” in Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, p. 1–15. [71] D. Kahneman, “Thinking, fast and slow,” Farrar, Straus and Giroux, 2011. [72] A. Male, Illustration: a theoretical and contextual perspective.Bloomsbury publishing, 2017. [73]F. M. Dwyer, “Exploratory studies in the effectiveness of visual illustrations,” AV Communication Review, p. 235–249, 1970. [74]M. Cook, “Students’ comprehension of science concepts depicted in textbook illustrations,” The Electronic Journal for Research in Science & Mathematics Education, 2008. [75]M. Agrawala, W. Li, and F. Berthouzoz, “Design principles for visual commu- nication,” Communications of the ACM, vol. 54, no. 4, p. 60–69, 2011. [76]B. Gooch and A. Gooch, Non-photorealistic rendering.AK Peters/CRC Press, 2001. [77] A. Hertzmann, “Introduction to 3d non-photorealistic rendering: Silhouettes and outlines,” Non-Photorealistic Rendering. SIGGRAPH, vol. 99, no. 1, 1999. [78]R. H. Kazi, T. Igarashi, S. Zhao, and R. Davis, “Vignette: interactive texture design and manipulation with freeform gestures for pen-and-ink illustration,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 2012, p. 1727–1736. [79]J. E. Kyprianidis, J. Collomosse, T. Wang, and T. Isenberg, “State of the” art”: A taxonomy of artistic stylization techniques for images and video,” IEEE transactions on visualization and computer graphics, vol. 19, no. 5, p. 866–885, 2012. [80] R. Chamberlain and J. Wagemans, “The genesis of errors in drawing,” Neuro- science & Biobehavioral Reviews, vol. 65, p. 195–207, 2016. [81]J. Canny, “A computational approach to edge detection,” IEEE Transactions on pattern analysis and machine intelligence, no. 6, p. 679–698, 1986. [82] R. Yi, Y.-J. Liu, Y.-K. Lai, and P. L. Rosin, “Apdrawinggan: Generating artistic portrait drawings from face photos with hierarchical gans,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, p. 10 743–10 752. [83]P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, p. 1125–1134. [84]T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, A. Tao, J. Kautz, and B. Catanzaro, “High- resolution image synthesis and semantic manipulation with conditional gans,” in Proceedings of the IEEE conference on computer vision and pattern recog- nition, 2018, p. 8798–8807. [85]C. Chan, F. Durand, and P. Isola, “Learning to generate line drawings that convey geometry and semantics,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, p. 7915–7925. [86]D. Ha and D. Eck, “A neural representation of sketch drawings,” in Interna- tional Conference on Learning Representations, 2023. [87]L. S. F. Ribeiro, T. Bui, J. Collomosse, and M. Ponti, “Sketchformer: Transformer-based representation for sketched structure,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, p. 14 153–14 162. [88]Y. Chen, S. Tu, Y. Yi, and L. Xu, “Sketch-pix2seq: a model to generate sketches of multiple categories,” arXiv preprint arXiv:1709.04121, 2017. [89]T.-M. Li, M. Luk ́ a ˇ c, M. Gharbi, and J. Ragan-Kelley, “Differentiable vector graphics rasterization for editing and learning,” ACM Transactions on Graphics (TOG), vol. 39, no. 6, p. 1–15, 2020. [90]Y. Vinker, Y. Alaluf, D. Cohen-Or, and A. Shamir, “Clipascene: Scene sketch- ing with different types and levels of abstraction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, p. 4146– 4156. [91]A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning. PMLR, 2021, p. 8748–8763. [92] P. M. Todd and G. Gigerenzer, “Pr ́ ecis of simple heuristics that make us smart,” Behavioral and brain sciences, vol. 23, no. 5, p. 727–741, 2000. [93]G. Klein, The power of intuition: How to use your gut feelings to make better decisions at work. Crown Currency, 2004. [94]J. Gildenblat and contributors, “Pytorch library for cam methods,” https:// github.com/jacobgil/pytorch-grad-cam, 2021. [95]W. Zhang, M. Dimiccoli, and B. Y. Lim, “Debiased-cam to mitigate image perturbations with faithful visual explanations of machine learning,” in CHI Conference on Human Factors in Computing Systems, 2022, p. 1–32. [96]A. Conmy, A. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga- Alonso, “Towards automated circuit discovery for mechanistic interpretability,” Advances in Neural Information Processing Systems, vol. 36, p. 16 318– 16 352, 2023. [97]D. Erhan, Y. Bengio, A. Courville, and P. Vincent, “Visualizing higher-layer features of a deep network,” University of Montreal, vol. 1341, no. 3, p. 1, 2009. [98]D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Interna- tional Conference on Learning Representations, 12 2014. [99]K. Frans, L. B. Soros, and O. Witkowski, “Clipdraw: exploring text-to-drawing synthesis through language-image encoders,” in Proceedings of the 36th Inter- national Conference on Neural Information Processing Systems, ser. NIPS ’22. Red Hook, NY, USA: Curran Associates Inc., 2022. [100]J. U. Blasberg, M. Gallistl, M. Degering, F. Baierlein, and V. Engert, “You look stressed: A pilot study on facial action unit activity in the context of psychosocial stress,” Comprehensive Psychoneuroendocrinology, vol. 15, p. 100187, 2023. [101]M. S. Bartlett, G. Littlewort, I. Fasel, and J. R. Movellan, “Real time face detection and facial expression recognition: Development and applications to human computer interaction.” in 2003 Conference on computer vision and pattern recognition workshop, vol. 5. IEEE, 2003, p. 53–53. [102]A. Thieme, M. Hanratty, M. Lyons, J. Palacios, R. F. Marques, C. Morrison, and G. Doherty, “Designing human-centered ai for mental health: Developing clinically relevant applications for online cbt treatment,” ACM Transactions on Computer-Human Interaction, vol. 30, no. 2, p. 1–50, 2023. [103]H. U. Rahiman and R. Kodikal, “Revolutionizing education: Artificial intelli- gence empowered learning in higher education,” Cogent Education, vol. 11, no. 1, p. 2293431, 2024. [104] P. Ekman and W. V. Friesen, “Facial action coding system,” Environmental Psychology & Nonverbal Behavior, 1978. [105]M. A. Sayette, J. F. Cohn, J. M. Wertz, M. A. Perrott, and D. J. Parrott, “A psychometric evaluation of the facial action coding system for assessing spontaneous expression,” Journal of nonverbal behavior, vol. 25, p. 167–185, 2001. [106]J. F. Cohn, Z. Ambadar, and P. Ekman, “Observer-based measurement of facial expression with the facial action coding system,” The handbook of emotion elicitation and assessment, vol. 1, no. 3, p. 203–221, 2007. [107] X. Zhang, L. Yin, J. F. Cohn, S. Canavan, M. Reale, A. Horowitz, and P. Liu, “A high-resolution spontaneous 3d dynamic facial expression database,” in 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG). IEEE, 2013, p. 1–6. [108]A. Savchenko, “Facial expression recognition with adaptive frame rate based on multiple testing correction,” in International Conference on Machine Learning. PMLR, 2023, p. 30 119–30 129. [109]C. Luo, S. Song, W. Xie, L. Shen, and H. Gunes, “Learning multi-dimensional edge feature-based au relation graph for facial action unit recognition,” in Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22, 2022, p. 1239–1246. [110]Y. Wu, Y. Zhou, G. Saveriades, S. Agaian, J. P. Noonan, and P. Natarajan, “Local shannon entropy measure with statistical tests for image randomness,” Information Sciences, vol. 222, p. 323–342, 2013. [111]P. Machado, J. Romero, M. Nadal, A. Santos, J. Correia, and A. Carballal, “Computerized measures of visual complexity,” Acta psychologica, vol. 160, p. 43–57, 2015. [112]H. Yu and S. Winkler, “Image complexity and spatial information,” in 2013 Fifth International Workshop on Quality of Multimedia Experience (QoMEX). IEEE, 2013, p. 12–17. [113]M. B. Sachs, J. Nachmias, and J. G. Robson, “Spatial-frequency channels in human vision,” Journal of the optical society of America, vol. 61, no. 9, p. 1176–1186, 1971. [114] P. Sangkloy, W. Jitkrittum, D. Yang, and J. Hays, “A sketch is worth a thousand words: Image retrieval with text and sketch,” in European conference on computer vision. Springer, 2022, p. 251–267. [115]M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa et al., “Dinov2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023. [116] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang et al., “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in European conference on computer vision. Springer, 2024, p. 38–55. [117]Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, and N. A. Smith, “Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,” in Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, 2023, p. 20 406–20 417. [118] C. Lu, L. Xu, and J. Jia, “Combining sketch and tone for pencil drawing production,” in Proceedings of the symposium on non-photorealistic animation and rendering, 2012, p. 65–73. [119] L. N. Kendall, Q. Raffaelli, A. Kingstone, and R. M. Todd, “Iconic faces are not real faces: enhanced emotion detection and altered neural processing as faces become more iconic,” Cognitive research: principles and implications, vol. 1, p. 1–14, 2016. [120](2024) 4k display market drivers, opportunities, trends, and forecasts by 2031. The Insight Partners. [Online]. Available: https://w.theinsightpartners.com/ reports/4k-display-market [121]R. Garnavi, M. Aldeen, and J. Bailey, “Computer-aided diagnosis of melanoma using border-and wavelet-based texture analysis,” IEEE transactions on infor- mation technology in biomedicine, vol. 16, no. 6, p. 1239–1252, 2012. [122]Z. She, Y. Liu, and A. Damatoa, “Combination of features from skin pattern and abcd analysis for lesion classification,” Skin Research and Technology, vol. 13, no. 1, p. 25–33, 2007. [123] R. Kasmi and K. Mokrani, “Classification of malignant melanoma and benign skin lesions: implementation of automatic abcd rule,” IET Image Processing, vol. 10, no. 6, p. 448–455, 2016. [124]S. Majumder and M. A. Ullah, “Feature extraction from dermoscopy images for melanoma diagnosis,” SN Applied Sciences, vol. 1, no. 7, p. 753, 2019. [125]A.-R. Ali, J. Li, and S. J. O’Shea, “Towards the automatic detection of skin lesion shape asymmetry, color variegation and diameter in dermoscopic images,” Plos one, vol. 15, no. 6, p. e0234352, 2020. [126]N. R. Abbasi, H. M. Shaw, D. S. Rigel, R. J. Friedman, W. H. McCarthy, I. Osman, A. W. Kopf, and D. Polsky, “Early diagnosis of cutaneous melanoma: revisiting the abcd criteria,” Jama, vol. 292, no. 22, p. 2771–2776, 2004. [127]Z. Yu, J. Nguyen, T. D. Nguyen, J. Kelly, C. Mclean, P. Bonnington, L. Zhang, V. Mar, and Z. Ge, “Early melanoma diagnosis with sequential dermoscopic images,” IEEE Transactions on Medical Imaging, vol. 41, no. 3, p. 633–646, 2021. [128]T. Chanda, K. Hauser, S. Hobelsberger, T.-C. Bucher, C. N. Garcia, C. Wies, H. Kittler, P. Tschandl, C. Navarrete-Dechent, S. Podlipnik et al., “Dermatologist-like explainable ai enhances trust and confidence in diagnosing melanoma,” Nature Communications, vol. 15, no. 1, p. 524, 2024. [129]P. Tschandl, C. Rosendahl, and H. Kittler, “The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions,” Scientific data, vol. 5, no. 1, p. 1–9, 2018. [130]J. Hou, J. Xu, and H. Chen, “Concept-attention whitening for interpretable skin lesion diagnosis,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2024, p. 113–123. [131]K. He, X. Zhang, S. Ren, and J. Sun, “ Deep Residual Learning for Image Recognition ,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2016, p. 770–778. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/CVPR.2016.90 [132]T. Oikarinen, S. Das, L. Nguyen, and L. Weng, “Label-free concept bottleneck models,” in International Conference on Learning Representations, 2023. [133]M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Doso- vitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen et al., “Simple open- vocabulary object detection,” in European conference on computer vision. Springer, 2022, p. 728–755. [134]A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al., “Segment anything,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, p. 4015– 4026. [135]S. Mishra and J. M. Rzeszotarski, “Crowdsourcing and evaluating concept- driven explanations of machine learning models,” Proceedings of the ACM on Human-Computer Interaction, vol. 5, no. CSCW1, p. 1–26, 2021. [136]A. Ghorbani, J. Wexler, J. Y. Zou, and B. Kim, “Towards automatic concept- based explanations,” Advances in neural information processing systems, vol. 32, 2019. [137] L. Baldwin and I. Crawford, “Art instruction in the botany lab: A collaborative approach.” Journal of College Science Teaching, vol. 40, no. 2, p. 26–31, 2010. [138]D. B. Hay, D. Williams, D. Stahl, and R. J. Wingate, “Using drawings of the brain cell to exhibit expertise in neuroscience: exploring the boundaries of experimental culture,” Science Education, vol. 97, no. 3, p. 468–491, 2013. [139]P. Rajpurkar, J. Irvin, R. L. Ball, K. Zhu, B. Yang, H. Mehta, T. Duan, D. Ding, A. Bagul, C. P. Langlotz et al., “Deep learning for chest radiograph diagnosis: A retrospective comparison of the chexnext algorithm to practicing radiologists,” PLoS medicine, vol. 15, no. 11, p. e1002686, 2018. [140] R. Sayres, A. Taly, E. Rahimy, K. Blumer, D. Coz, N. Hammel, J. Krause, A. Narayanaswamy, Z. Rastegar, D. Wu et al., “Using a deep learning algorithm and integrated gradients explanation to assist grading for diabetic retinopathy,” Ophthalmology, vol. 126, no. 4, p. 552–564, 2019. [141]R. Caruana, Y. Lou, J. Gehrke, P. Koch, M. Sturm, and N. Elhadad, “Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmis- sion,” in Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining, 2015, p. 1721–1730. [142] J. Krause, A. Perer, and K. Ng, “Interacting with predictions: Visual inspection of black-box machine learning models,” in Proceedings of the 2016 CHI conference on human factors in computing systems, 2016, p. 5686–5697. [143]F. Hohman, A. Head, R. Caruana, R. DeLine, and S. M. Drucker, “Gamut: A design probe to understand how data scientists understand machine learning models,” in Proceedings of the 2019 CHI conference on human factors in computing systems, 2019, p. 1–13. [144] Y. Zhang, T. Ren, F. Wang, and B. Y. Lim, “Comparables xai: Faithful example- based ai explanations with counterfactual trace adjustments,” in Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026, p. 1–33. [145]H. Kaur, M. R. Conrad, D. Rule, C. Lampe, and E. Gilbert, “Interpretability gone bad: The role of bounded rationality in how practitioners understand machine learning,” Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CSCW1, p. 1–34, 2024. [146]F. Doshi-Velez and B. Kim, “Towards a rigorous science of interpretable machine learning,” arXiv preprint arXiv:1702.08608, 2017. [147] A. Zytek, I. Arnaldo, D. Liu, L. Berti-Equille, and K. Veeramachaneni, “The need for interpretable features: Motivation and taxonomy,” ACM SIGKDD Explorations Newsletter, vol. 24, no. 1, p. 1–13, 2022. [148]G. He, A. Balayn, S. Buijsman, J. Yang, and U. Gadiraju, “It is like finding a polar bear in the savannah! concept-level ai explanations with analogical infer- ence from commonsense knowledge,” in Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, vol. 10, 2022, p. 89–101. [149]J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim, “Sanity checks for saliency maps,” Advances in neural information processing systems, vol. 31, 2018. [150]G. Erion, J. D. Janizek, P. Sturmfels, S. M. Lundberg, and S.-I. Lee, “Improving performance of deep learning models with axiomatic attribution priors and expected gradients,” Nature machine intelligence, vol. 3, no. 7, p. 620–631, 2021. [151]J. Willis and A. Todorov, “First impressions: Making up your mind after a 100-ms exposure to a face,” Psychological science, vol. 17, no. 7, p. 592–598, 2006. [152]V. Chen, Q. V. Liao, J. Wortman Vaughan, and G. Bansal, “Understanding the role of human intuition on reliance in human-ai decision-making with explanations,” Proceedings of the ACM on Human-computer Interaction, vol. 7, no. CSCW2, p. 1–32, 2023. [153]J. Schoeffer, M. De-Arteaga, and N. Kuehl, “Explanations, fairness, and ap- propriate reliance in human-ai decision-making,” in Proceedings of the CHI Conference on Human Factors in Computing Systems, 2024, p. 1–18. [154]D. E. Pearson and J. A. Robinson, “Visual communication at very low data rates,” Proceedings of the IEEE, vol. 73, no. 4, p. 795–812, 1985. [155] G. Donato, M. Bartlett, J. Hager, P. Ekman, and T. Sejnowski, “Classifying fa- cial actions,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 21, no. 10, p. 974–989, 1999. [156]R. C. Gonzalez, Digital image processing, Chapter 11.Pearson education india, 2009. [157]J. Alakuijala, R. Van Asseldonk, S. Boukortt, M. Bruse, I.-M. Coms , a, M. Firsching, T. Fischbacher, E. Kliuchnikov, S. Gomez, R. Obryk et al., “Jpeg xl next-generation image compression architecture and coding tools,” in Applications of digital image processing XLII, vol. 11137. SPIE, 2019, p. 112–124. [158] H. Kobayashi and L. R. Bahl, “Image data compression by predictive coding i: Prediction algorithms,” IBM Journal of Research and Development, vol. 18, no. 2, p. 164–171, 1974. [159] Y. Wang, S. Shen, and B. Y. Lim, “Reprompt: Automatic prompt editing to refine ai-generative art towards precise expressions,” in Proceedings of the 2023 CHI conference on human factors in computing systems, 2023, p. 1–29. [160]X. Zhao, W. Zhang, X. Xiao, and B. Lim, “Exploiting explanations for model inversion attacks,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, p. 682–692. [161]J. Howard et al., “Imagenette,” URL https://github.com/fastai/imagenette, vol. 2, 2020. [162]G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie, “The inaturalist species classification and detection dataset,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 8769–8778. [163] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al., “Gpt-4 techni- cal report,” arXiv preprint arXiv:2303.08774, 2023. Visualization type Saliency Landmark CLIPasso SketchXplain Observation Alignment 0 0.5 1.0 b) Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Observation Alignment Visualization type Saliency Landmark CLIPasso SketchXplain Observation Alignment 0 0.5 1.0 Cues Observation Cue Alignment Photo Alignment Visualization type Saliency Landmark CLIPasso SketchXplain Knowledge alignment 0 0.2 0.4 a) Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Knowledge Alignment 01 Visualization group Visualization type Saliency Landmark CLIPasso SketchXplain Salient Photo Photo Knowledge alignment 0 0.2 0.4 Saliency Landmark CLIPasso SketchXplain Salient Photo Photo AI Alignment Concept Alignment Appendix Fig. 7.CoherencemeasuredwithTASK-Former[114]cosinesimilarity towards a) Knowledge, b) Observation alignment. Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo CLIP classi fi cation correctness 0 0.5 1.0 EthnicityGenderIdentity 01 Vis Wrap Visualization type Saliency Landmark CLIPasso SketchXplain Salient Photo Photo CLIP classi fi cation correctness 0 0.5 1.0 Saliency Landmark CLIPasso SketchXplain Salient Photo Photo 01 Vis Wrap Visualization type Saliency Landmark CLIPasso SketchXplain Salient Photo Photo CLIP classi fi cation correctness 0 0.5 1.0 Saliency Landmark CLIPasso SketchXplain Salient Photo Photo Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Appendix Fig. 8. Privacy measured as CLIP classification correctness (i.e. re- identification risk) toward Gender, Ethnicity and Identity (as the original Photo). A EVALUATIONS ON FACE EXPRESSION IMAGE DOMAIN WeprovideadditionaldetailsontheSketchXplainevaluationforfacialex- pressions,coveringthemodelingstudyusedasapreliminarycheck,theuser studyonIntuitiveInterpretationdiscussedinthemainpaper(Section5.4). Here,wealsodiscussanadditionaluserstudytoinvestigatethePrivacy Protectionbenefitsofsketchexplanations. A.1 Modeling Study InSection5.2,weevaluatedwhetherSketchXplaingeneratessketch-based explanationsthatsatisfythedesiderataofsimplicityandcoherenceusing proxymetrics.Wefurtherjustifytheseproxymetricsandprovideadditional analyses,includingaprivacyevaluation,anablationstudyonthenumberof strokes,andintermediateoutputsforillustration. A.1.1 Justifications for Evaluation Metrics WejustifythechoiceofproxymetricsforVisualComplexityandCoherence, whichweusedtoevaluatewhetherthevisualizationssatisfythedesiderata inSection5.2.1. Visual Complexity: To be more accessible, AI explanations should be visually simple for people to comprehend. However, there is no reliable computational metric to compare visual complexity across heterogeneous image domains (e.g., photos, saliency maps, and line drawings). Therefore, we used two entropy-based measures—local Shannon entropy and JPEG XL file size—as complementary proxy indicators of visual complexity. Previous work [156] used Shannon entropy to indicate the variability of pixel intensities, but it ignores spatial relationships between pixels. Hence, we used local Shannon entropy [110], which divides the image into blocks, computes the entropy of each block and aggregates their average. For our images of256× 256pixels, we chose a block size of11× 11, which can adequately contain variations in curved lines and shapes. Although local entropy captures variability among image blocks, it does not encode local spatial information within each block. Image compression, in contrast, can capture this information and represent it concisely [112], making it a convenient and plausible estimator of human-perceived complexity [111]. Following this approach [111, 112], we used JPEG file size to complement local entropy, assuming larger file sizes correlate with higher information density, which requires greater effort to interpret. However, standard lossy compression of JPEG is not suitable for line drawings, since the discrete cosine transform (DCT) method would require high-frequency coefficients to represent the hard edges, causing file sizes to be unfairly large. Therefore, we also used “modular” lossless JPEG XL [157] which encodes images in the spatial domain (not the frequency domain) using predictive coding [158]. This allowed us to better predict the value of each pixel using neighboring ones and compute the residual between the predictions and actual intensities. These pixel-based residuals represent the surprisingness of the image, which is captured via entropy encoding. Thus, the resulting file size represents the visual complexity in terms of the amount of surprise from predictable patterns. Furthermore, to account for the human ability to perceive spatial frequency in images [113], we used lossy JPEG XL as an additional metric by setting the quality parameter to 99/100 to switch to the “variable DCT” encoding regime. In summary, we used local Shannon entropy and lossless and lossy JPEG XL file sizes to estimate visual complexity. Coherence: An explanation visualization should becoherentwiththeAI predictionlabel,i.e.,haveAIAlignment.Beyondthis,itshouldalsoalign withtheunderlyingexplanatoryconcepts(ConceptAlignment)toavoidspu- riousreasoning.Moreover,explanationsshouldcorrespondtorepresentative visualcuesoftheconcepts(CueAlignment)andthesourceimage(Observa- tionAlignment)sothatuserscanrelatewhattheyseewithrelevantconcepts andtheirmentalknowledge. To assess coherence, we leveraged the CLIP model [91]toobtainjointvision-languageembeddingrepresentationsallow- ingalignmentevaluationacrossmodalities.Wedevelopedcoherencemetrics based on cosine similarity scores from CLIP [117]. These similarity scores were computed by embedding both the visualization and its comparator (e.g., baseline drawing) and then taking their cosine similarity, where 1 indicates identical semantics and 0 indicates complete dissimilarity. As in [159], those scores indicatewhetherthevisualizationalignswiththepredictedlabelˆy andunderlyingconcepts ˆ γ.SinceCLIPcanalsomapcross-domainimages intoasharedsemanticspace,weusedCLIPscorestoevaluatealignment withvisualcues ̆ x x xortheoriginalimagex x x. Forgenerality,wealsoperformedthesameanalyseswithTASK-Former (TextAndSKetchtransformer)[114]asanadditionalevaluatormetricfor coherence.Itismoresuitedtosketch-basedrepresentations,sinceitwas co-trainedonimages,sketchesandtext.SeeresultsinAppendixFig.7which isconsistentwithresultsforCLIPinFig.4b–c). A.1.2 Privacy Evaluation A side-effect of abstract sketches is the omission of less relevant details in the face photos. This is useful for privacy, as it helps mitigate privacy attacks that re-identify people by exploiting the explanations without seeing the original photos [160]. a) b)c) Stroke count 46122436 CLIP classi fi cation correctness 0 0.5 1.0 EthnicityGenderIdentity Stroke count 46122436 CLIP classi fi cation correctness 0 0.5 1.0 EthnicityGenderIdentity Stroke count 46122436 Local entropy 0 0.5 1.0 1.5 JPEG XL fi le size (KiB) 0 5 10 15 20 Local entropy JPEG-XL (Lossy) JPEG-XL (Lossless) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Local entropy JPEG-XL (Lossless) JPEG-XL (Lossy) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Observation Alignment Stroke count 4122436 Knowledge Alignment 0.4 0.5 Observation Alignment 0.2 0.3 0.4 0.5 PhotoConceptAI Knowledge Alignment Appendix Fig. 9. Ablation results across various SketchXplain stroke settings: a) Simplicity, b) CLIP-based coherence, c) Privacy in explanations measured as CLIP classification correctness (i.e. re-identification risk). Appendix Table 5. Examples of SketchXplain sketches with increasing stroke count (cols) for 6 expressions (rows). SketchXplain 46122436 Disgusted Nose wrinkler Lip corner depressor Lower lip depressor Happy Cheek raiser Lip corner puller Neutral - Sad Inner brow raiser Brow lowerer Lip corner depressor Surprised Inner brow raiser Outer brow raiser Upper eyelid raiser Jaw drop Angry Brow lower Upper eyelid raiser Eyelid tightener Lip tightener LabelPhotoConcepts (AUs) We evaluated re-identification risk by how similar the visualization was to sensitive attributes—gender, ethnicity, identity. Gender and ethnicity labels were encoded as text and then converted to CLIP embeddings. Identity was encoded from the CLIP embedding of the original face photo. We modeled a privacy attack as a classification, where we calculated the cosine similarity of the visualization CLIP embedding to all sensitive attribute labels, and chose the label with the highest CLIP similarity score. Appendix Fig. 8 shows the results of privacy attack classification cor- rectness.Photocontained the most sensitive information, followed by Salient Photothat still retained key details of faces. All other visualiza- tions provided better privacy protection (lower attack correctness). Thus, SketchXplainprovided stronger privacy protection than Salient Photo. A.1.3 Ablation Analysis We conducted an ablation analysis on stroke count. The stroke initialization in SketchXplain requires explicitly setting the stroke count. We compared 4, 6, 12, 24, and 36 strokes to assess their impact on sketch coherence, simplic- ity, and privacy. Appendix Fig. 9 shows quantitative results and Appendix Table 5 shows example sketches with varying strokes. As expected, adding strokes raised the visual complexity of SketchXplain (Appendix Fig. 9a). Photo and AI Alignment increased with stroke count, but the trend plateaued after 12 strokes (Appendix Fig. 9b). Strangely, Concept Alignment decreased with stroke count despite starting very low. Perhaps, because AU concepts are not well-learned in the pre-trained CLIP model, the fewest strokes were poorly related to the concepts and increasing stroke count made the sketch more distinct and differentiated from the unclear (to CLIP) concepts. Nonetheless, Concept Alignment could be improved by adding a loss term to regularize the sketch embedding to align with AU embeddings (in Fig. 3, Step 3c). There was a slight increase in privacy risk as stroke count increased for gender and ethnicity, but there was no effect on the more challenging identity recognition (Appendix Fig. 9c). Appendix Table 6. Examples of intermediate outputs of SketchXplain for 6 expressions (rows). Disgusted Happy Neutral Sad Surprised Angry Label Photo 풙 Sampled endpoints Saliency 흐# ! Detailed lines풙$ Weighted lines풙% Init.strokes 풔# " Output (lines)풙' 풔 Output (photo)풙' 풔 Appendix Table 7. Statistical analysis of responses due to effects one per row as linear mixed effects models forquantitativeuserstudyonquickinterpretationofface expressionvisualizations. All models had various fixed main and interaction effects (shown as one effect per row) and Participant as a random effect.Fandpvalues indicate ANOVA tests and R 2 indicates model goodness-of-fit. Response LinearMixed Effects Model (Participant as random effect) Fp>FR 2 AI Simulatability Visualization Type +6.6<.0001 .196 Expression Truth +44.3<.0001 Display Duration +42.5<.0001 Visualization Type× Display Duration +5.7<.0001 Visualization Type× Expression Truth +9.0<.0001 Display Duration× Expression Truth +11.7<.0001 Visualization Type× Display Duration× Expression Truth2.2<.0001 Recall (AUs) Visualization Type +22.5<.0001 .352 Expression Truth +117.9<.0001 Display Duration +85.3<.0001 AU Type +15.8<.0001 Visualization Type× Display Duration +9.3<.0001 Visualization Type× Expression Truth +52.4<.0001 Display Duration× Expression Truth +20.8<.0001 Expression Truth× AU Type +1237.7<.0001 Display Duration× AU Type +41.0<.0001 Visualization Type× Display Duration× AU Type +22.6<.0001 Visualization Type× Expression Truth× AU Type11.6<.0001 A.1.4 SketchXplain Intermediate Outputs ToillustratehowSketchXplaingeneratessketchexplanationsforfacialex- pressions,weshowexamplesofintermediateoutputsinAppendixTable6 correspondingtomodulararchitecturestepsinFig.3. A.2 Quantitative User Study on Quick Interpretation Weprovidemoredetailsofthestatisticalanalysisforthequantitativeuser studyreportedinSection5.4toevaluatequick,intuitiveinterpretation. A.2.1 Statistical Analysis Wefitalinearmixedeffectsmodelforeachdependentvariableasthere- sponse,Visualizationtype,ExpressionandDisplayDuration,withother confoundingvariablesasfixedeffects,someinteractioneffectsamongthe factors,andParticipantasarandomeffect.AppendixTable7reportsthe modelfit(R 2 )andstatisticaleffectsofthelinearmixedeffectsmodel. SaliencyOutlineCLIPassoSketchXplainSalient photoPhoto Visualization Type 33.3 66.7 100.0 133.3 166.7 ≥ 200.0 Display Duration AI Label( Expression ) Angry Disgusted Sad Neutral Happy Surprised Perceived Label( Expression ) Surprised Happy Neutral Sad Disgusted Angry Surprised Happy Neutral Sad Disgusted Angry Surprised Happy Neutral Sad Disgusted Angry Surprised Happy Neutral Sad Disgusted Angry Surprised Happy Neutral Sad Disgusted Angry Surprised Happy Neutral Sad Disgusted Angry Angry Disgusted Sad Neutral Happy Surprised Angry Disgusted Sad Neutral Happy Surprised Angry Disgusted Sad Neutral Happy Surprised Angry Disgusted Sad Neutral Happy Surprised Angry Disgusted Sad Neutral Happy Surprised Freq. 0% 25% 50% 75% 100% Appendix Fig. 10. Confusion matrices ofAIalignmentdetailedresultsfromthequantitativeuserstudyonquickinterpretationoffaceexpressionvisualizations across Visualization Type and Display Duration. The y-axis represents participants’perceivedlabels, while the x-axis showsAI-predictedlabels. Darker blue indicates a higher number of responses in the corresponding cell. A.2.2 Supplementary Results on Multiclass Recognition Toillustratewhichfaceexpressionsarepreservedorconfusedbyvisualiza- tions,wepresentconfusionmatricesbetweenparticipantperceivedlabels andAIpredictedlabelsacrossVisualizationTypeandDisplayDuration (AppendixFig.10). ParticipantsviewingSaliencyexhibitedastrongbiastowardratingex- pressionsasNeutral,regardlessofthetrueExpressionorDisplayDuration. ThiswassimilarforSalient Photo,butdiminishedforlongerDisplayDura- tions(t> 133.3ms)duetosupplementaryinformationfromPhoto.Interest- ingly,thisbiastowardNeuralexpressionwasalsoevidentforparticipants viewingPhotounderverytighttimeconstraint(t = 33.3ms). Incontrast,line-drawingvisualizationswerelesssusceptibletothis bias,likelybecausetheirsalient,simplestrokesweremoreintuitivethan detailedpixelsorvaguesaliencyblobs.Though,OutlineandCLIPAssostill demonstratedmildNeuralexpressionbiasattheshortestDisplayDuration (t = 33.3ms).Overall,SketchXplainachievedstrongrecognitionperfor- manceacrossallDisplayDurationsbycombiningthesimplicityneeded forquickinterpretationwiththesalientcoherencerequiredforexpression recognition. A.2.3 Survey Screenshots We present screenshots of the survey questionnaire in thequantitativeuserstudyonquickinterpretationoffaceexpressionvisualizations (Section 5.4). Appendix Fig. 11. Tutorial to clarify users’ tasks and showcase visualizations. (Shared for both quantitative studies) Appendix Fig. 12. Tutorial on Action Units (AUs) and their correlation with face expressions. Appendix Fig. 13. Screening questions on face expression and visualizations. Appendix Fig. 14. Screening questions to check users’ understanding on Action Units. 1000ms100ms100ms 푡 Displayed in sequence Appendix Fig. 15. Example main study per-image trial with SketchXplain visualization. After “click to show”, the participant views the first frame with a cross to focus attention, views random lines before and after target visualization to clear memory, views the SketchXplain with t∈33.3, 66.7, 100.0, 133.3, 166.7 ms. Appendix Fig. 16. Example main study per-image trial with Action Unit andface expression recognition questions with reference table attached. 0 0.5 1.0 Correctness( Identity ) Saliency Landmark CLIPasso SketchXplain Visualization Type 0 0.5 1.0 Correctness( Identity ) Salient Photo Photo Visualizat ion Type c) 0 0.5 1.0 Correctness( Gender ) Saliency Landmark CLIPasso SketchXplain Visualization Type 0 0.5 1.0 Correctness( Gender ) Salient Photo Photo Visualizat ion Type a) Female Male Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain 0 0.5 1.0 Correctness( Ethnicity ) Saliency Landmark CLIPasso SketchXplain Visualization Type 0 0.5 1.0 Correctness( Ethnicity ) Salient Photo Photo Visualizat ion Type b) Asian Black White Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Salient Photo Photo Visualization Type Saliency Outline CLIPasso SketchXplain Appendix Fig. 17. Resultsfromthefaceprivacyprotectionuserstudy of recognition correctness of sensitive attributes: a) gender, b) ethnicity, and c) identity. Grey dotted line represents the random guess correctness (50.0% for gender, 33.3% for ethnicity, 25% for identity). Appendix Table 8. Statistical analysis of responses due to effects one per row as linear mixed effects models for faceprivacyprotectionuserstudy. All models had various fixed main and interaction effects (shown as one effect per row) and Participant as a random effect. Rows with grey text indicate non-significant effects.Fandpvalues indicate ANOVA tests and R 2 indicates model goodness-of-fit. Response LinearMixed Effects Model (Participant as random effect) Fp>FR 2 Recall (Gender) Target Expression +1.0n.s. .208 Target Ethnicity +28.6<.0001 Target Gender +46.8<.0001 Visualization Type +68.6<.0001 Visualization Type× Target Expression +1.9.0042 Visualization Type× Target Ethnicity +3.6<.0001 Visualization Type× Target Gender22.4<.0001 Recall (Ethnicity) Target Expression +2.9.0126 .283 Target Ethnicity +4.9.0076 Target Gender +32.5<.0001 Visualization Type +140.3<.0001 Visualization Type× Target Expression +1.8.0093 Visualization Type× Target Ethnicity +15.0<.0001 Visualization Type× Target Gender4.7.0003 Correctness (Identity) Target Expression +3.8.0019 .273 Target Ethnicity +20.6<.0001 Target Gender +4.4.0358 Visualization Type +163.4<.0001 Visualization Type× Target Expression +1.4n.s. Visualization Type× Target Ethnicity +3.7<.0001 Visualization Type× Target Gender1.2n.s. A.3 Quantitative User Study on Face Privacy Protection We investigated in a quantitative user study how sketch explanations provide privacy protection by omitting identifiable information in the originalface photo. A.3.1 Experiment Apparatus and Measures The user task was to identify facial attributes (gender and ethnicity) and match identity (from 4 face photos) from various visualizations. Unlike the previous quantitative user study(Section5.4), images were displayed without a time limit. We conducted a within-subjects experiment with Visualization Type and Expression as independent variables. A.3.2 Participants We recruited another 89 participants from Prolific with the same qualification criteria. 37 passed our screening test; 21 were maleand16werefemale, with ages 25–66 (Median = 41). They completed the survey in median 23.7 minutes and were compensated UK £4.00. A.3.3 Experiment Procedure Participants followed a procedure similar to the previous quantitative study. For screening, they had to a) demonstrate ability to interpretall visual- izations, b) identify facial attributes and identity from photos (Appendix Figs. 18–19). In the main study, for each of 72 trials, participants were asked to identify facial attributes fromthe visualization on the first page (Appendix Fig. 20), identity on the next page (Appendix Fig. 21). A.3.4 Statistical Analysis and Quantitative Results For all dependent variables, we fit a linear mixed-effects model with Ex- pression, Ethnicity, Gender, and Visualization Type as fixed main effects, various fixed interaction effects, and Participant as the random effect. See Appendix Table 8 for details. Appendix Fig. 17 shows thatSketchXplain,CLIPassoandOutlinesup- pressed Gender, Ethnicity and Identity information almost as well as Saliency.Salient Photoleaked the most information throughexposedimage pixels. Interestingly,CLIPassoleaked more information about males than females. A.3.5 Survey Screenshots Wepresentscreenshotsofthesurveyquestionnaireinthequantitativeuserstudyonfaceprivacyprotection(AppendixSectionA.3). Appendix Fig. 18. Screening questions to check users’ understanding on gender and ethnicity. Appendix Fig. 19. Screening questions to check users’ understanding of identity; the question continues on the next page. Appendix Fig. 20. Example main study per-image trial with gender and ethnicity recognition questions. Appendix Fig. 21. Example main study per-image trial with identification questions. Appendix Table 9. Examples of intermediate outputs of SketchXplain for skin lesion instances (rows). Benign Benign Benign Malignant Malignant Label Photo 풙 Saliency 흐# ! Detailed lines풙$ Weighted lines풙% Init.strokes 풔# " Output (lines)풙' 풔 Output (photo)풙' 풔 B EVALUATION ON SKIN LESION IMAGE DOMAIN WeprovidedetailsonapplyingSketchXplaintoexplainingskinlesionim- ages(Section6). B.1 SketchXplain Intermediate Outputs AppendixTable9showsintermediateoutputsfrommodularstepsin SketchXplainfortheseveralskinlesionexampleinstances. Appendix Table 10. Concepts used in the Label-Free Concept Bottleneck model for different domains. We obtained the concepts for each class using GPT-4o with the prompt: “Please provide the most important visual attributes to recognize [class].” DomainClassConcepts Dog Breeds Australian Terrier Hair-covered eyes, Upright ears, Shaggy muzzle, Bushy eyebrows, Topknot on head Border TerrierFolded ears, Wiry coat, Narrow body, Short muzzle, Thick base tail Samoyed Fluffy ears, Curled bushy tail, Thick White coat, Dark almond eyes, Feathered legs BeagleLong floppy ears, Short coat, Large Brown eyes, Square muzzle, White- tipped tail Shih-TzuHairy floppy ears, Short flat muzzle, Long flowing coat, Topknot on head, Curled tail over back, Short legs English FoxhoundLong drooping ears, Deep chest, Strong straight legs, Long tapered tail Rhodesian RidgebackFloppy ears, Strong legs, Long muzzle, Tapered tail Dingo Pointed upright ears, Bushy tail, Slender long legs, Narrow muzzle, Sandy coat Golden RetrieverFluffy ears, Feathered tail, Long wavy coat, Muscular build Old English Sheepdog Hair-covered eyes, Fluffy drooping ears, Shaggy grey-white coat, Docked tail Land Animals BearRounded ears, Massive paws, Rounded head. BisonShaggy front, Large hump, Short horns, Massive face BullBroad body, Short legs, Horns BuffaloCurved horns, Straight back, Low posture, Large muzzle CheetahRound ears, Long straight tail, Long limbs ElephantLarge fan ears, Tusks, Trunk GiraffeLong neck, Spotted coat, Ossicones HorseElongated face, Flowing tail, Flowing mane ImpalaLyre-shaped horns, Sleek body LionLarge mane, Round ears, Tufted tail end RabbitLong ears, Short round tail, Compact round body SquirrelLarge bushy tail, Pointed ears with tufts, Arched back ZebraStriped coat, Mane stands upright, Elongated face Land Vehicles Race CarSpoiler, Racing stripes, Air intake, Wheels, Windshield Police CarSiren lights, Badge, Wheels, Windshield Luxury CarElegant design, Spoiler, Wheels, Side-mirrors, Windshield MinivanSliding doors, Large windows, Wheels, Side-mirrors, Windshield Dump TruckDump bed, Hydraulic arms, Heavy tires Fire TruckWater hose, Ladder, Emergency lights, Wheels, Windshield TruckCargo bed, Wheels, Side-mirrors, Windshield ScooterSmall wheels, Handlebar, Seat C EVALUATION ON GENERAL IMAGE DOMAINS To investigate the generalizability of sketch explanations beyond face ex- pressionsandskinlesions, we applied SketchXplain to three other domains (dog breeds, land animals, and land vehicles). These span visual granular- ity and explanatory concepts. We describe data preparation, extensions to SketchXplain, results from both modeling studies and a qualitative study. We defer the quantitative evaluation of general sketches to future work, due to the heterogeneity of user goals for open domain images. Nevertheless, we draw qualitative insights from a qualitative study for the potential uses and benefits of sketch explanations in general. C.1 Data Preparation We used 12,954 images of 10 dog breeds (Australian Terrier, Border Ter- rier, Samoyed, Beagle, Shih-Tzu, English Foxhound, Rhodesian Ridgeback, Dingo, Golden Retriever, Old English Sheepdog) from the ImageWoof dataset [161], 21,874 images of 13 land animals (Bear, Bison, Bull, Buffalo, Cheetah, Elephant, Giraffe, Horse, Impala, Lion, Rabbit, Squirrel, Zebra) from the iNaturalist dataset [162], and 9,199 images of 8 land vehicles (Race Car, Police Car, Luxury Car, Minivan, Dump Truck, Fire Truck, Truck, Scooter) from ImageNet. For simplicity, these classes were selected such that the class objects were clearly shown in the photos, and their concepts can be simply identified and drawn; further work is needed to extract and depict more difficult concepts. C.2 Adapting SketchXplain for General Images Due to the increased heterogeneity of open-domain, general images, we adapted the concept bottleneck model (Step 2a) and cue localization (Step 2b) in SketchXplain to segment diverse concepts(AppendixFig.22). C.2.1 Semi-supervised Concept-bottleneck Model Unlike the BP4D face expression dataset, the datasets of general images do not have concept labels. We prompted the large language model GPT- 4o [163] using the same prompt template as [132] to generate a list of relevant concepts for each class(AppendixTable10). Next, we trained a semi-supervised Label-Free Concept Bottleneck Model (LF-CBM) [132] that co-trained two tasks for unsupervised learning of the concept labels, and supervised learning on the class labels. At inference time, this model Appendix Fig. 22. SketchXplain architecture adapted to general images. The concept bottleneck modelMis substituted byaCLIPvisualencoder[91], and cue localizationL by OWL-ViT [133] and Segment Anything [134]. Other components are unchanged from Fig. 3. Appendix Fig. 23. Examples comparing visualizations (cols) for explaining Dog Breeds, Land Animals and Land Vehicles (rows). LabelPhotoSaliencySalient photoCLIPassoSketchXplainConcepts Australian Terrier Hair-covered eyes Shaggy muzzle Topknot on head Beagle Long floppy ears Short coat Golden Retriever Heavy legs Fluffy ears Long coat Impala Lyre-shaped horns Sleek body Straight legs Giraffe Long neck Ossicones Spotted coat African Buffalo Curved horns Large muzzle Straight back Dump Truck Dump bed Heavy tires Side mirrors Luxury Spoiler Headlights Wheels Scooter Seat Handlebars Wheels Dog breeds Land animals Land vehicles performs Step 2a to predict the concepts ˆ γ, then uses them to predict the class ˆy. C.2.2 Open-domain Cue Localization While AU concepts tend to have fixed locations for the forward-looking faces in the BP4D dataset, concepts are heterogeneously placed in general images. This makes cue localization more challenging. Thus, we substituted the prior saliency map approach with two sub-steps: i) open-domain object detection with Owl-ViT [133] to extract the bounding boxes for each concept (if present) from the input image, and i) open-domain segmentation with Segment Anything Model (SAM) [134] to segment a pixel mask of the concept within the bounding box. We found that without object detection, SAM would spuriously select distant locations for the concept, thus the first sub-step was necessary. The segments of all concepts are weighted by importance determined by the gradients from the concept bottleneck and aggregated. The weighted segments were used as a mask to select edges to initialize the sketch strokes (Step 3a). C.3 Modeling Results We trained SketchXplain on 80% of the dataset and tested it on the rest. Class prediction accuracy was 91.8% for Dog breeds, 88.1% for Land animals, and 82.3% for Land vehicles. As in Section 5.2, we conducted a modeling study for each dataset to compare the coherence and simplicity of SketchXplain against other vi- sualizations. Appendix Table 23 and Appendix Fig. 24 show example explanations and the quantitative results, respectively. In general, results for general images were similar to those for face expres- sions.Regardingcoherence, CLIP and TASK-Formerscoreshadsimilar trends, except forAppendixFig.24c.1–c.3 where SketchXplain had higher Concept Alignmentasmeasuredby TASK-Former(browncolorlines). SketchXplainsketch explanations were equally aligned to the photo, AI pre- diction and concept as other visualizations, except forSaliency, which was the lowest due to omitting fine visual details.Furthermore,Salient Photo alignment was relatively lower for general images than for face images, perhaps due to thesalientregionsbeingtoosmall,fragmented,andscattered (seeAppendixTable23)forreliableCLIPrepresentation.Regarding visual complexity,SketchXplainsketches hadlowcomplexity(highsimplicity) and similar toCLIPassosketches.SaliencyandSalient Photohadlower complexity,possiblyduetothesmallsegmentfragments. C.4 Qualitative User Study Having shown that SketchXplain can faithfully generate sketches of other domains, we next conducted a qualitative user study to explore the potential use cases, usefulness, and challenges in interpreting sketch explanations of general images. C.4.1 Method and Procedure We recruited 15 participants through university mailing lists and personal contacts. They were 8 females and 7 males, with ages 21–31 years old(Me- dian=24.0). Notably, 5 participants had a background in art or design,who haveexperiencewithsketches. We showed participants sketch explanations of various Visualization Types for12 image instances(4perdomain). Par- ticipants described what they saw, why it mattered, and how they might use or share it. Due to the diversity of image domains, unlike for face sketches, we asked open-ended questions such as: • “Why do you think the model focused on this part?”, • “What information here feels important to you?”, • “How would you use this explanation?”, and • “How would you share this explanation with others?”. Sessions were conducted over Zoom. With consent, we recorded audio utterances and screen interactions.Theexperimenttook25–35minutesand eachparticipantwascompensated$7.80USDinlocalcurrency. C.4.2 Findings We conducted a thematic analysis of the recorded interviews, focusing on i) perceived value of sketch explanations (SketchXplain and CLIPasso) and i) possiblefuture applications. In general,unlikePhotowhich captured many details, sketch explanations provided an abstraction of visual content that “should be easy to under- stand, focusing on two or three features at a time” [P10] for recognizing an object from the image.Sketches were also “much easier to remem- ber” [P13],and could facilitateAIunderstanding by “showing the thought and idea process” [P9]. Specifically,SketchXplainfocused oncoherent visualcues, such as floppy ears or a muzzle to “break down the dog into the right silhouette” [P1],thegrilleandheadlightsas“whatmakesita Jeep”[P1], or ossicones and neck to confirm “what makes it a giraffe” [P3]. Thisselectivesimplicity “[helped] to know which part to focus on” [P7], andfacilitatesquickinterpretation“[made]iteasiertoidentifytheanimal quickly”[P11]. In contrast,CLIPassotended to “capture as much detail as possible” [P13] that could confuse interpretation. However, participants noted the limitation thatSketchXplainsometimes tended to oversimplify details, mentioned by P10 “too abstract, [because] some important lines [were] missing”. This suggests that SketchXplain could be improved with fine-tuning to domain-specific concepts. Toformativelyexplorepotentialusesofsketchexplanations,weasked participantsabouthypotheticalusesforotherapplications. Sketches were seen as “good for anatomy or biology pathways” [P11], helpful to “identify dogs or birds for beginners” [P3], and clear enough to “show why it’s a swallow: tail and beak” [P3]. Creative uses included extracting “Jeep grille and headlights that carry brand identity” [P1] and “traits to share on a mood board so everyone sees the same key cues” [P1]. Everyday communication examples ranged from “a simple map for which subway exit to take” [P13] to “website layout sketches to organize rows of text and images” [P14]. Participants emphasized the modularity of sketches, which allowed refining or complementing them. Editing practices included “draw over directly” [P14], “use thicker strokes for changes” [P12], and “color- code corrections” [P12]. Annotations were also suggested: “combine a sketch with a short explanation” [P9], or “use arrows and brief text labels to make intent clear” [P9].Clearly,participantsdrewfromtheireveryday ortrainingexperiencestoarticulatehowsketchexplanationsshouldfollow visualizationcommunicationdesignprinciples[75]. 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Local entropy 0 2 4 6 Saliency CLIPasso SketchXplain Salient Photo Photo JPEG-XL fi le size (KiB) 0 30 60 90 120 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Local entropy 0 2 4 6 Saliency CLIPasso SketchXplain Salient Photo Photo JPEG-XL fi le size (KiB) 0 30 60 90 120 Local entropy JPEG (Lossless) JPEG (Lossy) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) Visualization type Landmark CLIPasso SketchXplain Saliency Salient Photo Photo Local entropy 0 1 2 3 4 JPEG XL fi le size (KiB) 0 25 50 75 100 Local entropyJPEG-XL (Lossy)JPEG-XL (Lossless) 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Local entropy 0 2 4 6 Saliency CLIPasso SketchXplain Salient Photo Photo JPEG-XL fi le size (KiB) 0 30 60 90 120 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Local entropy 0 2 4 6 Saliency CLIPasso SketchXplain Salient Photo Photo JPEG-XL fi le size (KiB) 0 30 60 90 120 a.1)a.2)a.3) 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 Dog breedsLand animalsLand vehicles Salient Photo Photo Visualization Type Saliency CLIPasso SketchXplain Salient Photo Photo Visualization Type Saliency CLIPasso SketchXplain 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (CLIP) 0.3 0.4 Saliency CLIPasso SketchXplain Salient Photo Photo AI Faithfulness (TASK-Former) 0 0.1 0.2 0.3 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo CLIP similarity score (concepts) 0 0.5 1.0 Saliency CLIPasso SketchXplain Salient Photo Photo TASK-Former similarity score 0 0.5 1.0 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo CLIP similarity score (concepts) 0 0.5 1.0 Saliency CLIPasso SketchXplain Salient Photo Photo TASK-Former similarity score 0 0.5 1.0 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo CLIP similarity score (concepts) 0 0.5 1.0 Saliency CLIPasso SketchXplain Salient Photo Photo TASK-Former similarity score 0 0.5 1.0 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo CLIP similarity score (concepts) 0 0.5 1.0 Saliency CLIPasso SketchXplain Salient Photo Photo TASK-Former similarity score 0 0.5 1.0 Salient Photo Photo Visualization Type Saliency CLIPasso SketchXplain 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo CLIP similarity score (concepts) 0 0.5 1.0 Saliency CLIPasso SketchXplain Salient Photo Photo TASK-Former similarity score 0 0.5 1.0 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo CLIP similarity score (concepts) 0 0.5 1.0 Saliency CLIPasso SketchXplain Salient Photo Photo TASK-Former similarity score 0 0.5 1.0 Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Photo Alignment (CLIP) 0 0.5 1.0 Photo Alignment (TASK-Former) 0 0.5 1.0 Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Photo Alignment (CLIP) 0 0.5 1.0 Photo Alignment (TASK-Former) 0 0.5 1.0 Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Concept Alignment (CLIP) 0.25 0.3 Concept Alignment (T-Former) 0.1 0.2 0.4 0.2 0 0.2 0.1 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Local entropy 0 2 4 6 Saliency CLIPasso SketchXplain Salient Photo Photo JPEG-XL fi le size (KiB) 0 30 60 90 120 01 Vis type wrap Visualization type Saliency CLIPasso SketchXplain Salient Photo Photo Local entropy 0 2 4 6 Saliency CLIPasso SketchXplain Salient Photo Photo JPEG-XL fi le size (KiB) 0 30 60 90 120 b.1)b.2)b.3) c.1)c.2)c.3) d.1)d.2)d.3) CLIP TA S K-Former Local entropy JPEG - XL file size (KiB) AI Alignment (CLIP) AI Alignment (TASK - Former) Concep t Alignment (TASK - Former) Appendix Fig. 24. Results of modeling evaluation across Visualization Types on visual complexity (row a) and coherence (rows: b, c, d) and for general images of Dog Breeds (first column: a.1–d.1), Land Animals (a.2–d.2), and Land Vehicles (a.3–d.3). For each row, all graphs have the same left y-axes and the same right y-axes.