Paper deep dive
AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations
Ying Huang, Wencan Zhang, Brian Y. Lim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/17/2026, 5:09:50 AM
Summary
The paper introduces AlignFace, an interpretable, human-aligned face similarity metric that incorporates cognitive psychology principles such as featural/configural attribute dependence, nonlinear psychophysical scaling, and own-group biases. It utilizes a Visual-Language Model (VLM), Gated Cross-Attention, Concept Bottleneck Modeling (CBM), and a Neural Generalized Additive Model (GAM). The authors also present the FACETS dataset, containing human perceptual annotations for face similarity.
Entities (10)
Relation Signals (9)
AlignFace → uses → Gated Cross-Attention
confidence 95% · gated cross-attention (CA) to extract attribute-specific facial difference representations
AlignFace → uses → Concept Bottleneck Modeling
confidence 95% · concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes
AlignFace → uses → Neural Generalized Additive Model
confidence 95% · neural generalized additive model (GAM) to model their nonlinear influence
AlignFace → uses → Visual-Language Modeling
confidence 95% · It employs visual-language modeling (VLM) to encode paired face images and text-based attributes
AlignFace → accountsfor → Own-group bias
confidence 90% · own-group biases... AlignFace significantly improves alignment with human subpopulation perceptions
AlignFace → isbasedon → Cognitive Psychology
confidence 90% · we leverage scientific findings from cognitive psychology of human face similarity perception
FACETS → isusedby → AlignFace
confidence 90% · We introduce the FACETS dataset and propose AlignFace
AlignFace → outperforms → SSIM
confidence 90% · AlignFace outperforms metrics like LPIPS and SSIM in triplet comparisons
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debugging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face similarity perception: dependence on facial featural and configural attributes, nonlinear psychophysical response scaling, and own-group biases. We introduce the FACETS dataset and propose AlignFace, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face images and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and human cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images.
Tags
Links
- Source: https://arxiv.org/abs/2608.14130v1
- Canonical: https://arxiv.org/abs/2608.14130v1
Trouble viewing inline? Open PDF directly →
Full Text
71,567 characters extracted from source content.
Expand or collapse full text
AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations Ying Huang ∗ Department of Computer Science, National University of Singapore Singapore, Singapore ying.h@u.nus.edu Wencan Zhang ∗ Department of Computer Science, National University of Singapore Singapore, Singapore wencanz@nus.edu.sg Brian Y. Lim Department of Computer Science, National University of Singapore Singapore, Singapore brianlim@nus.edu.sg a) AlignFace LPIPS SSIM Human c) ⋯ Δlip ColorΔPupillary Distance d) ReferenceABReferenceABReferenceAB Overall Own-group Other-group b) ⋯ ΔPupillary Distance ΔLip Color Figure 1: a) AlignFace provides a face-pair similarity score that better matches human perception. b) It decomposes face similarity into attribute-level comparisons, capturing featural (e.g., lip color) and configural (e.g., pupillary distance) cues for fine-grained explanations. c) Decisions are explained via learned nonlinear functions that map attribute similarity to overall face similarity, while capturing group-dependent effects (own- vs. other- group). d) AlignFace outperforms metrics like LPIPS and SSIM in triplet comparisons—whether A or B is more similar to a reference—improving alignment with human perception. Abstract Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human percep- tion. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debug- ging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face sim- ilarity perception: dependence on facial featural and configural ∗ These authors contributed equally to this work. This work is licensed under a Creative Commons Attribution 4.0 International License. M ’26, Rio de Janeiro, Brazil © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2213-4/2026/11 https://doi.org/10.1145/3767308.3836354 attributes, nonlinear psychophysical response scaling, and own- group biases. We introduce the FACETS dataset and propose Align- Face, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face im- ages and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottle- neck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace signifi- cantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and hu- man cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images. CCS Concepts • Computing methodologies→Computer vision; Image rep- resentations;• Human-centered computing→ User studies. Keywords Face Perception, Face Evaluation, Human-AI Alignment, Similarity, Interpretability, Explainable AI, Vision-Language Model arXiv:2608.14130v1 [cs.M] 14 Aug 2026 M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. ACM Reference Format: Ying Huang, Wencan Zhang, and Brian Y. Lim. 2026. AlignFace: Human- Aligned Face Similarity Metric with Interpretable Concept Relations. In Proceedings of the 34th ACM International Conference on Multimedia (M ’26), November 10–14, 2026, Rio de Janeiro, Brazil. ACM, New York, NY, USA, 16 pages. https://doi.org/10.1145/3767308.3836354 1 Introduction The study of faces in computer vision has given rise to many applications for machine understanding, including face recogni- tion [18,42,61,72], social media retrieval [14,84], expression recog- nition for affective computing [16,71], and descriptive caption- ing [32,37]. In these paradigms, faces are treated as data points for classification. However, innovations in generative face synthesis and automated image editing have redefined the prediction tasks. Recent works target human observers with visual content for hu- man consumption, instead of machine interpretation. These human- centric applications range from face editing [8,20,31,36,40,63], in- cluding digital makeup [17,38,78] and restoration [73,83] to obfus- cation for privacy protection [7,79,82]. While standard methods for evaluating model performance involve metrics and models [28,74], these data-centric approaches are poor proxies for human percep- tion [24], critical for estimating the true effects on people. Thus, a high-performing face verification model could deceptively un- derestimate privacy risk or identity shifts by identifying two faces as different when a human would still perceive them as similar, or vice versa. Recent efforts to embed human perceptual ratings [85], semantic constraints [50], and conceptual abstraction [49] lack transparency and are not cognitively grounded, risking spurious correlations [58] and being unfaithful to human perception. Instead, we argue that perception models should adhere to sci- entific cognitive effects to guide their alignment with humans. For face perception tasks, we focus on the psychology of human face recognition [10]. First, humans process faces based on featural attributes (face parts, e.g., eye color, nose shape) and configural at- tributes (spatial layout, e.g., pupillary distance, jaw width) [3,15,47]. This determines which features or concepts face perception models should encode explicitly or implicitly. Second, humans do not per- ceive stimuli linearly [66], e.g., perceived brightness varies by the square root of light intensity, and perceived differences depend on absolute values of each stimulus [25]. Hence, we hypothesize that human perception of attribute differences nonlinearly affects their perception of the similarity between two faces. Third, humans have varied backgrounds and abilities, leading to individual or group variances across a population. Notably, own-group bias has been observed, showing that people better recognize faces from their own demographic groups (ethnicity, gender, etc.) [45,48]. Thus, human alignment should steer model behaviors toward human subgroups rather than assume a universal model for all humans. We leverage the aforementioned cognitive principles with a data collection experiment and a new human-aligned model for face perception. We collected large-scale human perceptual annotations for both overall face and attribute-specific similarity ratings from two-alternative forced choice (2AFC) comparison tasks [9]. This produced 9.36k triplet ratings of face similarity from 78 participants, and 73.32k triplet ratings of 20 face attributes from 611 participants. Trained on these data and grounded in cognitive prin- ciples, we propose AlignFace, a human-aligned similarity metric model for face perception. It is ante-hoc interpretable with mod- ules for face attribute predictions and nonlinear attribute-overall perception mapping. Each face attribute is predicted from a contex- tual fusion of a pair of face images for comparison, and a textual prompt of the attribute (e.g., “eyebrow shape” ) [6]. The attributes are combined as a multi-label concept set in a Concept Bottleneck [34] to ensure downstream decisions in terms of these interpretable attributes. These attributes then serve as inputs into a neural Gen- eralized Additive Model (GAM) [12] to capture their nonlinear partial influence on the final overall perception score. In experiments, we evaluated the human-alignment of AlignFace against heuristic similarity, learned perceptual, and cross-modal baselines. We compared the correlation with human ground-truth ratings for overall and per-attribute similarity, for all humans and for specific ethnic groups, and performed an ablation study on the model modules. Results showed that AlignFace was more correlated to human perception than baselines overall, per-attribute, and per- group. Visual examination of partial dependence from the GAMs also revealed highly nonlinear relations, confirming the hypothesis of nonlinear attribute influence. Our contributions are: 1)Cognitive characteristics—featural and configural attributes, nonlinear perceptual scaling, and demographic (White/Asian) own-group bias—identified for human face similarity perception. 2)AlignFace, a cognitively-grounded, interpretable ante-hoc model for face perception that integrates the cognitive characteristics via Visual-Language encoding, Gated Cross-Attention, Concept Bottleneck and Neural GAM to enforce human alignment. 3)AlignFace2AFC, an explainable human-aligned face similarity metric that explains perception via the nonlinear contributions of cognitively-grounded attributes. 4)FACETS, a human-annotated dataset of face triplets with simi- larity judgments at both overall and attribute levels across 20 cognitively-grounded face attributes. 2 Related Work We examine why current face models fail to capture true human perception due to neglecting cognitive principles, and how AI in- terpretability methods can close this gap. Face Representation Models. Face models have been devel- oped for myriad tasks, such as identity recognition [18,61,72], expression analysis [89], landmark localization [77], and attribute prediction [43]. Many depend on face representation model back- bones, particularly margin-based architectures like ArcFace [18] and CosFace [72], or triplet-loss based models like FaceNet [61]. However, despite their reliance on human-annotated data, these models remain primarily data-centric, optimizing for categorical ground truths rather than accounting for the subjectivity and non- linear response scaling inherent in perceptual attributes. This limi- tation stems from face benchmarks relying on binary annotations that neglect the contrastive nuance and context dependence of hu- man perception [30,43,68]. In contrast, our model training collects human perception labels from pairwise 2AFC judgments of overall AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil face perception and multiple face attributes, and measures viewer demographics to account for perception biases. Perceptual Similarity Metrics. The advancement of image generation necessitates evaluation metrics faithful to human per- ception. Conventional metrics like PSNR [28] and SSIM [74] or transformer-based representations like CLIP [56] and DINO [52] fail to capture semantics or fine-grained perceptual nuances, offering poor proxies for human judgment. Learned metrics like LPIPS [85] or DreamSim [23] improve alignment, their black-box nature risks depending on spurious features and correlations [58]. Indeed, DNNs often exploit cues divergent from human perception [24]. While re- cent work attempts to enforce representational alignment through local-constrained global transform [50] or concept abstraction [49], these methods are data-driven and not cognitively grounded. Focusing on facial perception, humans prioritize a combina- tion of featural (face parts) and configural (spatial relations) at- tributes [15,47], and integrate them additively [5,70]. Furthermore, their effects are governed by nonlinear psychophysical laws [25,66] and modulated by observer-dependent variations, such as the own- group bias [45,48], characterized by superior recognition of one’s own demographic group. Unlike prior metrics that are data-driven and not cognitively grounded, we leverage these cognitive princi- ples to design a human-aligned face perceptual metric that explicitly structures facial similarity through facial attributes, with nonlin- ear psychophysical weighting and observer-dependent modulation, introducing cognitively grounded inductive biases. Explainable Vision Models. Ante-hoc interpretability has been shown to improve model performance while supporting human interpretability [59]. Unlike visual explainable AI (XAI) techniques, such as saliency maps [62,64] that can be spurious [86,88], such approaches explicitly encode cognitive principles [2,87] or domain knowledge [41,46] to improve human alignment. While concept- based explanations [33,34] can encode human-understandable con- cepts, they assume these concepts influence behavior linearly. In- stead, generalized additive models (GAM) [11,12] can explain the nonlinear influence of features on predictions. In this work, we developed an ante-hoc interpretable model, leveraging concept bot- tleneck [34] and GAM [12] together to align perceptual attributes nonlinearly to human perception. Furthermore, much of explainable AI (XAI) has focused on dis- crimination tasks, with few methods specifically for similarity tasks. They focus on showing comparative saliency maps with similar regions [91], saliency and similarity [42,90], and landmarks of posi- tive/negative contributions [21]. Within the facial domain, research similarly focuses on saliency maps [76,81], and more recently on grouped concepts of face attributes [67] or inferred concepts [54]. While prior methods focus on attribution of latent activations or features, we ground and constrain our approach on cognitive prin- ciples, leading to a more theoretically justified model, which we also show to be more accurately aligned with human perception. 3 Face Attributes for Comparative Evaluation with Triplet Similarity (FACETS) Dataset To develop a human-aligned face similarity metric, we collected human perception ratings of similarity between faces overall and per-attribute for 20 face attributes. Table 1: Cognitively-grounded featural and configural at- tributes of human face perception adapted 1 from [3]. CategoryAttributeScale (0–1) Featural Eyebrow shapeRounded–Straight Eyebrow thicknessThin–Thick Eye shapeNarrow–Round Eye sizeSmall–Large Eye colorLight–Dark Nose shapePointed–Flat Nose sizeSmall–Large Mouth sizeSmall–Large Lip thicknessThin–Thick Lip colorLight–Dark Hair colorLight–Dark Skin colorLight–Dark Skin textureSmooth–Rough Configural Face aspect ratioShort-wide–Tall-narrow Hair lengthBald–Long Forehead heightShort–Long Pupillary distanceSmall–Large Cheek shapeSunken–Puffy Jaw widthNarrow–Wide Chin shapePointed–Square 1 We omitted ear attributes due to occlusion by hairstyles, and added lip color due to variation in makeup. 3.1 Image Selection and Stimuli Preparation We constructed our face-pair stimuli using images from two face datasets spanning both lab-controlled and in-the-wild faces: CMU Multi-PIE [26] and CelebA [43]. To better balance demographic groups (gender, ethnicity) from the two datasets, due to the severely limited number of Asian identities (only 40), we subsampled White identities (to 80) to obtain 120 identities (see Appendix Table 3). For each identity, we selected four source images with varying poses and illumination. This balances intra-identity variability with inter-identity variability across demographic groups. To reduce po- tential confounding factors, we excluded images containing heavy facial hair or eyeglasses. This initial image set provides seed images that we use to generate edited faces for comparison. To emulate real-world scenarios, where changes are often subtle to preserve identity (face editing or retouching) or contextual infor- mation (privacy protection), we synthesized edited face images us- ing face image generation techniques to produce variations of seed face photos. We employed diffusion-based inpainting [55] for edit- ing featural attributes, landmark-based warping [44] for configural attributes, and GAN-based transfer [51] for hair-related attributes while preserving hairstyle consistency with demographics. Based on detected face landmark points, we estimated attribute values related to length, size, and shape using geometry-based heuristic face anthropometric methods [22]. Color-relevant attributes were estimated from image color histograms from cropped facial compo- nents. We ensured that the generated face had the same distribution of attribute values as the original faces (see Appendix Fig. 11). Fi- nally, we further masked the background from the face images to avoid distraction bias during human annotation. M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. ReferenceAB AB [Face]Whichfaceismoresimilartoreference? [Attribute]Onlylookingatlipcolor,whichismoresimilartoreference? Figure 2: Experiment apparatus of the 2AFC protocol for two user studies on [Face] overall similarity and per-[Attribute] similarity, with different survey questions based on the task. We constrained the edits to 20 face attributes identified as highly influential in face space [3], encompassing both featural and config- ural attributes (Table 1). Featural attributes describe discrete compo- nents or surface qualities that can be identified in isolation, such as the specific color of an eye or the smoothness of skin, regardless of their spatial position on the face. Configural attributes describe the spatial geometry and relational distances between features, ranging from internal gaps like pupillary distance to the external silhouette boundaries defined by attributes like cheek and chin shapes, and hair length. Each source image underwent two distinct edits, each editing 3–15 face attributes with randomly sampled magnitudes. 3.2 Psychophysical Similarity Measurements Having prepared face image data, similar to data collection studies of human similarity perception [23,85], we elicited human percep- tion ratings via the comparative two-alternative forced choice test (2AFC) [9]. Instead of rating the similarity between two items (A and Target), the participant chooses which of two choices (A or B) is more similar to a Reference [53,69]. This improves sensitivity to perceptual differences [65] and mitigates response biases [75]. We constructed 480 face triplets (4 random tuples from each of 120 identities)⟨푅,퐴,퐵⟩, where푅denotes the original unedited face, and퐴and퐵are two edited versions derived from푅. For each triplet, to assess overall similarity, we asked participants to choose “which face is more similar to the reference” ; to further assess per- attribute similarity, we asked participants to choose based on one randomly selected attribute out of 20. Fig. 2 shows the experiment apparatus, with different questions for two user studies on overall face similarity and per-attribute similarity. For both user studies, the experiment procedure is: introduction, consent, screening questions to verify correctness on trivially easy cases, main study with 120 trials each of randomly selected triplets over four sessions interspersed with 30-second breaks, conclusion with demographic questions. Each main study session included one attention check question similar to the screening question. We recruited online participants from Prolific. 78 participants 1 com- pleted the face similarity study in a median time of 23.2 minutes and were compensated £4.00, and 611 participants completed the attribute similarity study in a median time of 30.5 minutes and 1 Although modest, this participant sample yielded 8,880 ratings across 480 triplets, sufficient to fine-tune the pretrained VLM and achieve high agreement (87.0%) with human judgments. (see Fig. 6). Attribute Nose size Relative Attribute Selection -1+1 Relative Face Selection -0.2 0 0.2 0.4 Mouth size -1+1 Lip color -1+1 Skin color -1+1 Face aspect ratio -0.2 0 0.2 0.4 Forehead height Pupillary distance Cheek shape Other GroupOwn Group Face Aspect ratio Attribute Featural Configural Own-group Other-group Figure 3: Partial dependence plots of overall relative face selection by the 8 most salient attributes, split by own- and other-group. See Appendix Fig. 12 for full attributes. were compensated £4.00. Notable for our own-group analysis, we report the ethnicity distribution of our participants 2 : 377 White, 222 Asian, 55 Black, 11 Hispanic, and 24 Mixed-race. See Appendix Table 4 for participant details. The user studies were approved by our institutional review board. After collecting the human similarity ratings, to ensure data quality, we excluded 6,510 trials with few ratings (fewer than 3), had ambiguous ratings (M = 30–70% selecting either A/B), or were in the same session as a failed attention check question (all 30 responses excluded). Hence, from both user studies, we collected 8,880 triplet ratings of face similarity, and 72,450 triplet ratings of 20 face attributes. We randomly selected 80% of this dataset for model training, and 20% for testing. 3.3Cognitive Characteristics of Face Perception To understand which attributes influence face similarity percep- tion more than others, we fit Generalized Additive Models (GAM) with a logit link function. 2AFC selection (A or B) as response and difference in attribute change (R→A− R→B) as factors. Given the demographic imbalance in our participant sample, we restrict the ethnicity-based own/other-group analysis to White and Asian participants. We fit two GAMs for own-group and other- group demographics to account for subpopulation variance. See Fig. 3 for the results. Attribute Relevance. Human similarity judgments were un- evenly distributed across face attributes; certain features (e.g., face aspect ratio, forehead height) are significantly more influential than others (e.g., mouth size, lip color). Additive Nonlinear Scaling of Attribute Differences. Per- ceptual sensitivity to attribute differences is strongly nonlinear. Interestingly, as Relative Attribute Distance Difference increases, i.e., Face B is more different from the Reference than A, Relative Face Selection (toward B) increasingly increases. However, this trend is reversed toward Face A, where the Relative Face Selection effect diminishes with Relative Attribute Distance Difference. This suggests a side-choice bias [80]. 2 Achieving perfect balance remains challenging due to the inherent demographic skew of Western crowdsourcing platforms [19]. AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil 풛" ! 풁 $ "##$ 풛" % 풁 $ & Δ 퐸 '() 퐸 #*# 퐸 '() tanh휶 Q K V 푑 , 풅 $ "##$ 푓 + Δ푐 + 푓 , Δ푐 , ∑ ⋮ ⋮ Δ풄 " "##$ 풙 ! 풙 % 풙 "##$ “푃푢푝푖푙푎푟푦 푑푖푠푡푎푛푐푒” “퐿푖푝 푐표푙표푟” ⋮ Figure 4: AlignFace architecture. 1) vision-language model (VLM) encoding of face-pair images and text-based attributes into a shared representation, 2) attribute-gated cross-attention (CA) to extract attribute-specific visual features, 3) face-attribute concept bottleneck model (CBM) to constrain reasoning through disentangled face attributes, and 4) Attribute-Influence neural Generalized Additive Model (GAM) to represent each attribute’s nonlinear, additive contribution to the pairwise distance ˆ 푑. Demographic Variance (Own-Group Effect). For the White and Asian ethnic groups, different trends were found for own-group and other-group ratings. For example, humans were more sensitive to configural attributes Face aspect ratio and Forehead height, when perceiving faces sharing their own demographics (own-group). Con- versely, they were more sensitive to featural attributes Skin color and Cheek shape when perceiving other-group faces. This percep- tual divergence confirms that a single, universal similarity metric is insufficient for human-aligned evaluation, motivating the need for demographic-aware steering. These findings of cognitive effects directly motivate our technical approach, which we describe next. 4 AlignFace for Perceptual Similarity Metric We propose an interpretable, human-aligned face similarity metric model—AlignFace—for both overall and attribute perceptual com- parisons between faces (Fig. 4). Grounded in cognitive principles identified in Section 2 and observed in Section 3.3, AlignFace pre- dicts similarity in a face pair, accounting for multiple face attributes, their nonlinear scaling effects, and demographic-specific variance. AlignFace leverages 1) vision-language models (VLMs) [39,56] to map face images and extensible, open-ended attribute-based text prompts into a shared semantic embedding space, 2) gated cross- attention [4] to extract attribute-specific visual features, 3) concept bottleneck model (CBM) [34] to constrain reasoning through disen- tangled face attributes, and 4) neural Generalized Additive Model (GAM) [12] to resolve overall similarity into an additive compo- sition of nonlinear functions over individual attribute differences. See Appendix Table 8 for justifications of each module in AlignFace compared to standard alternatives. 4.1 Vision-Language Model (VLM) Encoding While a standard visual encoder extracts embeddings that capture overall facial information, these representations are often semantic- agnostic and thus fail to support attribute-targeted comparisons. We instead leverage a shared semantic representation space from a 풙 ! 푀 "# Δ풄 & !$ 푑 ( !$ 풙 % 푀 "# Δ풄 & %$ 푑 ( %$ 풙 &'( − Δ푑 ( Δ풄 & −푦 ) −풚 * − ℒ +,-.' Δ푓 / Δ푐 / Δ푓 0 Δ푐 0 ⋮ Figure 5: AlignFace2AFC metric using AlignFace (푀 AF ) with a triplet instance⟨푥 퐴 ,푥 Ref ,푥 퐵 ⟩to predict two sets of overall and attribute distances and compute their differences (Δ ˆ 푑,Δ ˆ 풄). Differences are learned via supervised training with binary 2AFC human labels of the overall face푦 푑 or of all attributes 풚 푐 . See푦 푑 –풚 푐 GAM-based relations in Fig. 3. pre-trained VLM to enable face comparison along specific attribute directions (Fig. 4.1). Specifically, we employ a shared vision encoder to extract representations for a face pair as⟨ ˆ 풛 퐴 , ˆ 풛 퐵 ⟩. We then apply relational reasoning [60] to model the differences between two face images, yielding ˆ 풁 Δ 3 . Next, we use the corresponding text encoder to extract descriptions of face attributes (e.g., “Pupillary distance” ) into ˆ 풁 푎푡푟 , which serves as an anchor guiding the comparison be- tween the two faces along a specific attribute dimension. 4.2 Attribute-Gated Cross-Attention (CA) Having obtained the visual and textual embeddings, we perform multi-modal fusion to extract attribute-specific information for 3 Concatenation퐴⊕ 푅, element-wise difference퐴− 푅, and Hadamard product퐴⊙ 푅. M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. face comparison. In particular, we employ Attribute-Gated Cross- Attention (AGCA) [4] on image-difference features as queries (Q) and textual features as keys (K) and values (V), like in latent diffu- sion models [57] 4 , (Fig. 4.2): ˆ 풁 ′ Δ = ˆ 풁 Δ + tanh(훼)· CrossAttn(푄= ˆ 풁 Δ ,퐾= ˆ 풁 attr ,푉= ˆ 풁 attr ), (1) where훼is a learnable parameter for gating factortanh(훼), control- ling the attribute-conditioned signal injected into ˆ 풁 Δ . The cross-attention operates analogously to a dictionary lookup, retrieving attribute-relevant visual features from ˆ 풁 Δ conditioned on the attribute context ˆ 풁 attr . The residual connection stabilizes optimization by preserving the original representation. 4.3 Face-Attribute Concept Bottleneck Model For a given face pair⟨푥 퐴 ,푥 퐵 ⟩, the attribute-gated CA produces one attribute distance measure per text-based attribute prompt. We denote these as interpretable multi-label binary conceptsΔ ˆ 풄 attr ∈ [− 1,+1], which also bottlenecks subsequent reasoning (Fig. 4.3), making the architecture a concept bottleneck model (CBM) [34]. However, unlike typical CBMs that input the concepts into an MLP for downstream reasoning, we leverage an interpretable mod- ule, described next. 4.4 Attribute-Influence Neural GAM To interpret the attribute influence while maintaining cognitive- grounding with nonlinear response scaling and additive attribute integration, we represent the relationship between overall distance and attribute distances with a Generalized Additive Model (GAM): 퐹(Δ ˆ 풄)= 휎(훽+ 푁 ∑︁ 푖=1 푓 푖 (Δ ˆ 푐 푖 )),(2) where each푓 푖 is a spline function implemented via a small neu- ral network, capturing both value-dependent attribute contribu- tions and their relative importance across factors.휎(·)denotes a sigmoid function 5 , and훽is the learned bias term. To ensure per- ceptual consistency, we further impose a monotonicity constraint, i.e., 휕푓 푖 (Δ ˆ 푐 푖 ) 휕Δ ˆ 푐 푖 ≥0, to ensure that higher attribute difference never decreases the overall face distance. We implement GAM as neural layers using Node-GAM [12] (Fig. 4.4). 4.5 AlignFace2AFC Perception Metric Learning We define AlignFace2AFC, a multi-task triplet metric for similarity perception of a face triplet풙= ⟨푥 퐴 ,푥 Ref ,푥 퐵 ⟩(Fig. 5). For each triplet, we apply AlignFace twice to compare face pairs⟨푥 퐴 ,푥 Ref ⟩ and⟨푥 퐵 ,푥 Ref ⟩separately to predict their pairwise distances and calculate the Relative Distance Difference for the overall face (Δ ˆ 푑= ˆ 푑 퐵푅 − ˆ 푑 퐴푅 ) and attributes (Δ ˆ 푐=Δ ˆ 푐 퐵푅 -Δ ˆ 푐 퐴푅 ). We collected human ground-truth labels of similar face selection for overall face푦 푑 ∈ +1,−1and per-attribute푦 푖 푐 ∈ +1,−1. Each 푦is +1 if face B푥 퐵 is selected as more similar to the reference face 푥 Ref , and -1 if face A푥 퐴 is selected instead. Note that the similarity label has the opposite sense of the distance that AlignFace predicts, 4 This inverted from cross-attention in text-guided VLMs [4], since we require attention in terms of face attributes (V) for downstream concept-bottleneck modeling. The image- diff query (Q) asks which attributes (K) explain the differences between the faces. 5 An inverse of the logistic link to model the bimodal distribution of human responses. PSNR SSIM SimCLR MoCo DINO LPIPS DreamSim CLIP FLIP FaceNet CosFace ArcFace AF-CLIP AF-FLIP HeuristicSemanticPerceptualVLMFace recognitionAlignFace 0% 50% 100% Face Agreement ✓ Figure 6: Model-human agreement on overall face percep- tion. The grey line represents random guessing. Error bars show 95% confidence intervals. Baselines without✓are sig- nificantly lower than AF-FLIP (Dunnett’s test, 훼= 0.05). so we need to flip it to Dissimilarity Label−푦 푑 and−풚 푐 . AlignFace is trained by minimizing the squared hinge loss [23] between the dissimilarity labels and relative distance differences, i.e., L 푖 푐 (Δ ˆ 푐 푖 ,푦 푖 푐 )= max 0, 푚−Δ ˆ 푐 푖 ·(−푦 푖 푐 ) 2 ,(3) L 푑 (Δ ˆ 푑,푦 푑 )= max 0, 푚−Δ ˆ 푑 ·(−푦 푑 ) 2 ,(4) where푚is a margin (set to푚=0.05), to enforce relative ranking via attract-and-repelling. To manage training instability, we first train the attribute-based difference prediction with per-attribute lossL 푖 푐 , freeze it, and then train the Neural GAM with overall distance lossL 푑 . 5 Experiments We evaluated AlignFace to investigate its alignment and attribution faithfulness to human perception, compared to baselines. 5.1 Experimental Settings 5.1.1 Implementation Training Details. We implemented AlignFace in PyTorch, freezing the text encoder and adapting the visual en- coder with LoRA [29] (푟=4,훼=8, dropout=0.5). The fusion module comprises cross-attention followed by a three-layer MLP (dropout=0.2), while NodeGAM [12] uses 50 knots. We trained both attribute- and face-level predictions with squared hinge loss (margin=0.05) using Adam (learning rate=10 −3 , weight decay =10 −4 , batch size=16) on NVIDIA RTX 3090 GPUs. AlignFace has 29.42M trainable parameters and required 157.16ms per face pair on a single RTX 3090 GPU. 5.1.2 Evaluation Criteria. To assess alignment with human 2AFC judgments, we evaluate each metric on held-out triplets by com- paring its predicted preference with human annotations. Given the relative distance differenceΔ ˆ 푑and human labels푦 푑 ∈ +1,−1 defined earlier, we compute agreement as: 푎= 1[Δ ˆ 푑 ·(−푦 푑 )> 0]. We evaluate alignment from two perspectives: 1)Behavioral Agreement: We calculate the agreement score be- tween averaged human ratings and binarized model preferences. 2)Reasoning Analysis: To gain deeper insights into the model behavior, we visualize and analyze partial dependence plots (PDPs) to compare the relationship between individual attribute similarities and overall face perception. AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Attribute Agreement a)b)c) Backbone AF - CLIP AF - FLIP Figure 7: Attribute perception agreement of a) averaged across all attributes, b) grouped into featural and configural categories, and c) at individual-level attributes. Featural Configural 0.94 0.71 0.96 0.94 -0.28 0.98 0.92 0.83 0.96 0.91 -0.51 0.97 0.91 0.52 0.99 0.94 0.92 0.94 0.95 0.97 0.99 0.93 0.91 0.94 휌: Face Aspect ratio nGAM Figure 8: Partial dependence plots (PDPs) of selected at- tributes, comparing human judgments with Linear, MLP and nGAM. The numbers in graphs are Pearson correlations휌be- tween each model and humans. High correlations for nGAM (0.94–0.98) indicate strong human alignment of AlignFace. 5.1.3 Baseline Comparators. For overall face similarity, we com- pare AlignFace against widely used metrics for assessing generative content quality. These include: 1) Heuristic metrics: PSNR [28] and SSIM [74]. 2)Self-supervised semantic visual representation models: SimCLR [13], MoCo [27] and DINOv2 [52]. 3) Visual-language models: CLIP [56], FLIP [39]. 4) Perceptual similarity models: LPIPS [85] and DreamSim [23]. 5)Face-specific models: FaceNet [61], CosFace [72] and ArcFace [18]. We also performed ablation studies to examine the benefits to overall and per-attribute alignment of the modules of AlignFace. For these comparisons, we initialize our fusion module with randomized parameters to provide a rigorous point of reference. 5.2 Alignment Evaluation Overall Face Perception Alignment. Heuristic approaches (PSNR, SSIM) and self-supervised representation models (SimCLR, MoCo, DINO) exhibit the lowest consistency with human ratings (Fig. 6a). In contrast, perceptual methods outperform both groups, with a) Face Agreement Attribute Agreement b) Own-group Other-group Figure 9: Cross-group perception agreement evaluations on a) face and b) averaged attributes. DreamSim achieving slightly higher agreement than LPIPS. Surpris- ingly, approaches based on pretrained contrastive visual-language models (VLMs) achieve relatively strong performance among the baselines. FaceNet, CosFace, and ArcFace baselines performed sig- nificantly worse, likely due to their smaller training datasets and lack of VLM pretraining. Overall, AlignFace outperforms all compet- ing methods. We also observed additional gains when using domain- specific encoders (i.e., AlignFace (FLIP) vs. AlignFace (CLIP)), sug- gesting that FLIP better captures facial semantics. Face Attributes Alignment. We investigate how well Align- Face aligns with human perception across face attributes 6 . Fig. 7a shows that AlignFace had reasonably high attribute agreement (푀=0.79 for FLIP backbone). This was lower than the overall face agreement, perhaps due to attribute diversity and data sparsity. Specifically, it had the highest agreement for Hair length, Eyebrow shape and Forehead height, and the lowest agreement for Nose shape, Jaw width, and Chin shape (Fig. 7c). Attribute Influence Alignment. Next, we examined whether AlignFace was “right for the right reason” [58] by comparing its attribute influence on overall perception against that of human perception. Specifically, we compared the individual GAM function shapes of AlignFace due to its Attribute-Influence Neural GAM module (Section 4.4) represented in the triple-based dual-pair com- parison (Fig. 5) against the human GAM (Section 3.2, Fig. 12). Fig. 8 shows how well AlignFace’s partial dependence plots align with human perception relations, achieving high Pearson correlations휌. See Appendix Fig. 13 for all attributes. 6 As a mediation check, we also validated that general-purpose VLMs could accurately encode objectively-verifiable, facial attributes by comparing AlignFace’s predicted at- tribute sizes and distances against ground-truth heuristic inter-landmark measures [44]. Appendix Table 9 shows high correlations (Median 휌= 0.947). M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. Table 2: Agreement under within- and cross-dataset eval- uation for AF-CLIP and AF-FLIP. Consistent performance indicates learned generalized perception across datasets. Setting AF-CLIPAF-FLIP FaceAttrFaceAttr Within-dataset79.9%77.9%81.3%78.3% Cross-dataset79.4%72.9%81.3%73.2% Subpopulation Alignment. Given the own-group bias observed for White and Asian partici- pants in Section 3.2, we next investigated its impact on AlignFace. Fig. 9 shows that AlignFace is in good agreement with human overall perception across the faces of the Own and Other-groups. Notably, agreement was stronger for Own-group perception than Other-group, perhaps because the model was trained on fewer Other-group faces due to imbalanced ethnicities in the face datasets. See Appendix Fig. 15 for results across all attributes. This finding suggests that the proposed approach benefits from tailoring the model to specific subpopulations. 5.3 Cross-Dataset Generalizability We investigated the impact of dataset domain shifts in face images on AlignFace’s performance. To this end, we retrained two Align- Face models, with CLIP and FLIP backbones, using targets from the CMU Multi-PIE [26] and CelebA [43] datasets separately, and evaluated them under both within-dataset and cross-dataset set- tings. Table 2 shows that AlignFace achieves consistent agreement scores across both settings. See Appendix Fig. 14 for per-attribute results. This result indicates strong robustness and generalizability across datasets, perhaps due to the VLM’s open-domain knowledge, the use of cognitively-grounded attributes, and the alignment of relationships between attributes and overall perception. 5.4 Ablation Studies We conducted ablation studies to evaluate the contribution of two key components in AlignFace: 1) Attribute-Gated Cross-Attention for attribute-conditioned representation learning, and 2) Attribute- Influence Neural GAM (nGAM) for aggregating attribute differences into the overall face distance perception. Results are in Fig. 10. Attribute-Gated Cross-Attention. We replaced the Attribute- Gated Cross-Attention module with a direct fusion variant, where the image-pair and attribute embeddings are concatenated and fed into an MLP. Fig. 10a,b show a clear drop in agreement with human perception at both the overall and attribute levels, with a greater degradation for attribute-level agreement. Even with su- pervision from human labels, the direct fusion variant fails to cap- ture attribute-specific perceptual cues as effectively. These results highlight the importance of attribute-conditioned visual-semantic interaction for modeling human-aligned similarity. Attribute-Influence Neural GAM. We further replaced the Neural GAM module with two alternative aggregation heads: a Linear model and a Multi-Layer Perceptron (MLP). The Linear model assumes a simple additive weighting across attributes, while the MLP captures nonlinear interactions in a black-box manner. Fig. 10 shows that the Linear model had the lowest agreement, indicating that simple additive weighting is insufficient to model Face Agreement Attribute Agreement Face Agreement a) b) c) w/o attn with attn 0% 50% 100% Condition w/o attn with attn w/o attn with attn 0% 50% 100% Condition with attn w/o attn Condition with attn w/o attn Figure 10: Results of ablation studies on Attribute-Gated Cross-Attention: a) overall face model-human agreement, b) average attribute agreement; and on Attribute-Influence Neural GAM: c) overall face agreement. perceptual similarity. The MLP achieved performance comparable to Neural GAM, but lacked interpretability. We also examined the attribute-influence trends using partial dependence plots for the ablated models. Fig. 8 shows the poor fit of the Linear model and MLP. Notably, the complexity of the MLP did not help its agreement compared to Linear. For example, MLP learned spurious trends of decreasing overall perceived difference with increasing relative attribute distance difference for Cheek shape. Hence, Neural GAM had the highest human alignment while being interpretable. 6 Discussion We have shown the benefits of grounding face similarity modeling with the cognitive principles of featural and configural attributes, nonlinear scaling, and own-group bias. While this work establishes a foundation for human-aligned face similarity, it focuses on con- trolled perceptual settings with a relatively compact dataset and a specific participant demographic, and excludes ambiguous tri- als with 30–70% rater agreement 7 . Future work could extend it to broader populations and more diverse generative scenarios, such as privacy-preserving obfuscation, makeup transfer, and restoration. Beyond faces, AlignFace suggests a general approach for semantic- aware visual similarity. Although we focus on attributes for face perception, the prompt-based VLM can model open-ended domain concepts, such as ABCD criteria for assessing skin-lesion disease progression [1]. Ultimately, our method advances toward human- aligned modeling that mirrors both behavioral responses and un- derlying cognitive reasoning, which has the potential to inspire other work on human-AI alignment across diverse perceptual tasks. 7 Conclusion We have presented AlignFace, a human-aligned face similarity metric to model both face and fine-grained attribute perception. Grounded in scientific cognitive principles, it achieves alignment across both predictive behavioral responses and underlying model reasoning. Experimental evaluations demonstrate that AlignFace significantly outperforms existing domain-free metrics in mirror- ing human judgment. We further provide an in-depth analysis of subpopulation influences, dataset-driven domain shifts, and the ar- chitectural impact. Ultimately, this work establishes a more faithful proxy for human face perception, providing a reliable foundation for evaluating generative facial content. 7 Although common in crowdsourced model training [23,85], this may omit Point of Subjective Equality (PSE) “hard examples” for fine-grained perceptual differences. AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Acknowledgments This research is supported by the Ministry of Education, Singapore (Award No: T2EP20121-0040), the National Research Foundation, Singapore and Infocomm Media Development Authority under its Trust Tech Funding Initiative (Award No: DTC-RGC-09), a Google Research Scholar Award, and the NUS Institute for Health Innova- tion and Technology (iHealthtech). The views expressed are those of the authors and do not necessarily reflect those of the funders. References [1]Naheed R Abbasi, Helen M Shaw, Darrell S Rigel, Robert J Friedman, William H McCarthy, Iman Osman, Alfred W Kopf, and David Polsky. 2004. Early diagnosis of cutaneous melanoma: revisiting the ABCD criteria. Jama 292, 22 (2004), 2771–2776. [2] Harshavardhan Sunil Abichandani, Wencan Zhang, and Brian Y Lim. 2025. Robust Relatable Explanations of Machine Learning with Disentangled Cue-specific Saliency. In Proceedings of the 30th International Conference on Intelligent User Interfaces. 1203–1231. [3] Noga Abudarham and Galit Yovel. 2016. Reverse engineering the face space: Discovering the critical features for face identification. Journal of Vision 16, 3 (2016), 40. [4]Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736. [5]Norman H. Anderson. 1968. A Simple Model for Information Integration. In Theories of Cognitive Consistency: A Sourcebook, Robert P. Abelson, Elliot Aronson, William J. McGuire, Theodore M. Newcomb, Milton J. Rosenberg, and Percy H. Tannenbaum (Eds.). Rand McNally, Chicago, 731–743. [6]Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision. 2425–2433. [7]Simone Barattin, Christos Tzelepis, Ioannis Patras, and Nicu Sebe. 2023. Attribute- preserving face dataset anonymization via latent code optimization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8001–8010. [8]Denis Bobkov, Vadim Titov, Aibek Alanov, and Dmitry Vetrov. 2024. The devil is in the details: Stylefeatureeditor for detail-rich stylegan inversion and high quality image editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9337–9346. [9]Rafal Bogacz, Eric Brown, Jeff Moehlis, Philip Holmes, and Jonathan D Cohen. 2006. The physics of optimal decision making: a formal analysis of models of performance in two-alternative forced-choice tasks. Psychological review 113, 4 (2006), 700. [10]Vicki Bruce and Andy Young. 1986. Understanding face recognition. British journal of psychology 77, 3 (1986), 305–327. [11]Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. 2015. Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. In Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining. 1721–1730. [12] Chun-Hao Chang, Rich Caruana, and Anna Goldenberg. 2022. NODE-GAM: Neu- ral Generalized Additive Model for Interpretable Deep Learning. In International Conference on Learning Representations. [13] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Interna- tional conference on machine learning. PmLR, 1597–1607. [14]Min Jin Chong, Wen-Sheng Chu, Abhishek Kumar, and David Forsyth. 2021. Retrieve in style: Unsupervised facial feature transfer and retrieval. In Proceedings of the IEEE/CVF international conference on computer vision. 3887–3896. [15]Stephan M Collishaw and Graham J Hole. 2000. Featural and configurational processes in the recognition of faces of different familiarity. Perception 29, 8 (2000), 893–909. [16]Feng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang, Dan Guo, Zhi Liu, and Meng Wang. 2025. Learning from heterogeneity: Generalizing dynamic facial expression recognition via distributionally robust optimization. In Proceedings of the 33rd ACM international conference on multimedia. 5587–5596. [17]Han Deng, Chu Han, Hongmin Cai, Guoqiang Han, and Shengfeng He. 2021. Spatially-invariant style-codes controlled makeup transfer. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition. 6549–6557. [18] Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4690–4699. [19]Benjamin D Douglas, Patrick J Ewell, and Markus Brauer. 2023. Data qual- ity in online human-subjects research: Comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA. Plos one 18, 3 (2023), e0279720. [20]Xiaoxiong Du, Jun Peng, Yiyi Zhou, Jinlu Zhang, Siting Chen, Guannan Jiang, Xiaoshuai Sun, and Rongrong Ji. 2023. Pixelface+: Towards controllable face generation and manipulation with text descriptions and segmentation masks. In Proceedings of the 31st acm international conference on multimedia. 4666–4677. [21]Oliver Eberle, Jochen Büttner, Florian Kräutli, Klaus-Robert Müller, Matteo Valle- riani, and Grégoire Montavon. 2020. Building and interpreting deep similarity models. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 3 (2020), 1149–1161. [22] Leslie G. Farkas. 1994. Anthropometry of the head and face. [23] Stephanie Fu, Netanel Y Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. 2023. DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data. In Advances in Neural Information Processing Systems. 50742–50768. [24]Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. 2018. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International conference on learning representations. [25] George A Gescheider. 2013. Psychophysics: the fundamentals. Psychology Press. [26]Ralph Gross, Iain Matthews, Jeffrey Cohn, Takeo Kanade, and Simon Baker. 2010. Multi-pie. Image and vision computing 28, 5 (2010), 807–813. [27]Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9729–9738. [28]Alain Hore and Djemel Ziou. 2010. Image quality metrics: PSNR vs. SSIM. In 2010 20th international conference on pattern recognition. IEEE, 2366–2369. [29] Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations. [30]Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. 2008. La- beled faces in the wild: A database for studying face recognition in unconstrained environments. In Workshop on faces in’Real-Life’Images: detection, alignment, and recognition. [31]Ziqi Huang, Kelvin CK Chan, Yuming Jiang, and Ziwei Liu. 2023. Collabora- tive diffusion for multi-modal face generation and editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 6080–6090. [32]Youngjoon Jang, Kyeongha Rho, Jongbin Woo, Hyeongkeun Lee, Jihwan Park, Youshin Lim, Byeong-Yeol Kim, and Joon Son Chung. 2023. That’s what i said: Fully-controllable talking face generation. In Proceedings of the 31st ACM Inter- national Conference on Multimedia. 3827–3836. [33]Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al.2018. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning. PMLR, 2668–2677. [34] Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pier- son, Been Kim, and Percy Liang. 2020. Concept bottleneck models. In International conference on machine learning. PMLR, 5338–5348. [35]J. Richard Landis and Gary G. Koch. 1977. The measurement of observer agree- ment for categorical data. Biometrics 33, 1 (1977), 159–174. [36] Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. 2020. Maskgan: Towards diverse and interactive facial image manipulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5549–5558. [37]Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang, Wenqi Ren, Jianhou Gan, and Xiaochun Cao. 2025. FaceInsight: A multimodal large language model for face perception. In Proceedings of the 33rd ACM International Conference on Multimedia. 11052–11061. [38] Tingting Li, Ruihe Qian, Chao Dong, Si Liu, Qiong Yan, Wenwu Zhu, and Liang Lin. 2018. Beautygan: Instance-level facial makeup transfer with deep generative adversarial network. In Proceedings of the 26th ACM international conference on Multimedia. 645–653. [39]Yudong Li, Xianxu Hou, Zheng Dezhi, Linlin Shen, and Zhe Zhao. 2024. Flip- 80m: 80 million visual-linguistic pairs for facial language-image pre-training. In Proceedings of the 32nd ACM International Conference on Multimedia. 58–67. [40]Hanbang Liang, Xianxu Hou, and Linlin Shen. 2021. SSflow: Style-guided neu- ral spline flows for face image manipulation. In Proceedings of the 29th ACM International Conference on Multimedia. 79–87. [41]Brian Y Lim, Joseph P Cahaly, Chester YF Sng, and Adam Chew. 2025. Diagramma- tization and Abduction to Improve AI Interpretability With Domain-Aligned Explanations for Medical Diagnosis. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–25. [42] Yu-Sheng Lin, Zhe-Yu Liu, Yu-An Chen, Yu-Siang Wang, Ya-Liang Chang, and Winston H Hsu. 2021. xcos: An explainable cosine metric for face verification task. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 17, 3s (2021), 1–16. [43]Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015. Deep Learning Face Attributes in the Wild. In Proceedings of International Conference on Computer Vision (ICCV). [44]Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, et al.2019. Mediapipe: A framework for building M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. perception pipelines. arXiv preprint arXiv:1906.08172 (2019). [45]Roy S Malpass and Jerome Kravitz. 1969. Recognition for faces of own and other race. Journal of personality and social psychology 13, 4 (1969), 330. [46] Hitoshi Matsuyama, Nobuo Kawaguchi, and Brian Y Lim. 2023. Iris: Interpretable rubric-informed segmentation for action quality assessment. In Proceedings of the 28th International Conference on Intelligent User Interfaces. 368–378. [47] Daphne Maurer, Richard Le Grand, and Catherine J Mondloch. 2002. The many faces of configural processing. Trends in cognitive sciences 6, 6 (2002), 255–260. [48] Christian A Meissner and John C Brigham. 2001. Thirty years of investigating the own-race bias in memory for faces: A meta-analytic review. Psychology, Public Policy, and Law 7, 1 (2001), 3. [49]Lukas Muttenthaler, Klaus Greff, Frieda Born, Bernhard Spitzer, Simon Kornblith, Michael C Mozer, Klaus-Robert Müller, Thomas Unterthiner, and Andrew K Lampinen. 2025. Aligning machine and human visual representations across abstraction levels. Nature 647, 8089 (2025), 349–355. [50]Lukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A Vandermeulen, Katherine Hermann, Andrew Lampinen, and Simon Kornblith. 2023. Improving neural network representations using human similarity judgments. Advances in neural information processing systems 36 (2023), 50978–51007. [51]Maxim Nikolaev, Mikhail Kuznetsov, Dmitry Vetrov, and Aibek Alanov. 2024. Hairfastgan: Realistic and robust hair transfer with a fast encoder-based approach. Advances in Neural Information Processing Systems 37 (2024), 45600–45635. [52]Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al.2024. DINOv2: Learning Robust Visual Features without Supervision. Transactions on Machine Learning Research (2024). [53]Devi Parikh and Kristen Grauman. 2011. Relative attributes. In 2011 International conference on computer vision. IEEE, 503–510. [54]Richard Plesh, Janez Križaj, Keivan Bahmani, Mahesh Banavar, Vitomir Štruc, and Stephanie Schuckers. 2024. Discovering interpretable feature directions in the embedding space of face recognition models. In 2024 IEEE International Joint Conference on Biometrics (IJCB). IEEE, 1–10. [55] Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. SDXL: Improving Latent Dif- fusion Models for High-Resolution Image Synthesis. In The Twelfth International Conference on Learning Representations. [56]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PMLR, 8748–8763. [57]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695. [58]Andrew Slavin Ross, Michael C. Hughes, and Finale Doshi-Velez. 2017. Right for the right reasons: training differentiable models by constraining their ex- planations. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (Melbourne, Australia) (IJCAI’17). AAAI Press, 2662–2670. [59] Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature machine intelligence 1, 5 (2019), 206–215. [60] Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. 2017. A simple neural network module for relational reasoning. Advances in neural information processing systems 30 (2017). [61]Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition. 815–823. [62]Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE interna- tional conference on computer vision. 618–626. [63]Yichun Shi, Xiao Yang, Yangyue Wan, and Xiaohui Shen. 2022. Semanticstylegan: Learning compositional generative priors for controllable image synthesis and editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11254–11264. [64]Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013). [65]Chaehan So. 2023. Measuring aesthetic preferences of neural style transfer: More precision with the two-alternative-forced-choice task. International Journal of Human–Computer Interaction 39, 4 (2023), 755–775. [66] Stanley S Stevens. 1957. On the psychophysical law. Psychological review 64, 3 (1957), 153. [67]Divyang Teotia, Agata Lapedriza, and Sarah Ostadabbas. 2022. Interpreting face inference models using hierarchical network dissection. International Journal of Computer Vision 130, 5 (2022), 1277–1292. [68]Philipp Terhörst, Daniel Fährmann, Jan Niklas Kolf, Naser Damer, Florian Kirch- buchner, and Arjan Kuijper. 2021. Maad-face: A massively annotated attribute dataset for face images. IEEE Transactions on Information Forensics and Security 16 (2021), 3942–3957. [69]Louis L Thurstone. 2017. A law of comparative judgment. In Scaling. Routledge, 81–92. [70]Anne M Treisman and Garry Gelade. 1980. A feature-integration theory of attention. Cognitive psychology 12, 1 (1980), 97–136. [71]Hanyang Wang, Bo Li, Shuang Wu, Siyuan Shen, Feng Liu, Shouhong Ding, and Aimin Zhou. 2023. Rethinking the learning paradigm for dynamic facial expression recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 17958–17968. [72]Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. 2018. Cosface: Large margin cosine loss for deep face recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 5265–5274. [73] Xintao Wang, Yu Li, Honglun Zhang, and Ying Shan. 2021. Towards real-world blind face restoration with generative facial prior. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9168–9178. [74] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 4 (2004), 600–612. [75]Eunike Wetzel, Jan R Böhnke, and Anna Brown. 2016. Response biases. (2016). [76]Jonathan R Williford, Brandon B May, and Jeffrey Byrne. 2020. Explainable face recognition. In European conference on computer vision. Springer, 248–263. [77] Jiahao Xia, Weiwei Qu, Wenjian Huang, Jianguo Zhang, Xi Wang, and Min Xu. 2022. Sparse local patch transformer for robust face alignment and landmarks inherent relation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4052–4061. [78] Chenyu Yang, Wanrong He, Yingqing Xu, and Yang Gao. 2022. Elegant: Exquisite and locally editable gan for makeup transfer. In European conference on computer vision. Springer, 737–754. [79] Zixuan Yang, Yushu Zhang, Tao Wang, Zhongyun Hua, Zhihua Xia, and Jian Weng. 2024. Once-for-all: Efficient visual face privacy protection via person- specific veils. In Proceedings of the 32nd ACM International Conference on Multi- media. 7705–7713. [80] Yaffa Yeshurun, Marisa Carrasco, and Laurence T Maloney. 2008. Bias and sensitivity in two-interval forced choice procedures: Tests of the difference model. Vision research 48, 17 (2008), 1837–1851. [81]Bangjie Yin, Luan Tran, Haoxiang Li, Xiaohui Shen, and Xiaoming Liu. 2019. To- wards interpretable face recognition. In Proceedings of the IEEE/CVF international conference on computer vision. 9348–9357. [82]Lin Yuan, Linguo Liu, Xiao Pu, Zhao Li, Hongbo Li, and Xinbo Gao. 2022. PRO- face: A generic framework for privacy-preserving recognizable obfuscation of face images. In Proceedings of the 30th ACM international conference on multimedia. 1661–1669. [83]Zongsheng Yue and Chen Change Loy. 2024. Difface: Blind face restoration with diffused error contraction. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 9991–10004. [84]Alireza Zaeemzadeh, Shabnam Ghadar, Baldo Faieta, Zhe Lin, Nazanin Rahnavard, Mubarak Shah, and Ratheesh Kalarot. 2021. Face image retrieval with attribute manipulation. In Proceedings of the IEEE/CVF international conference on computer vision. 12116–12125. [85]Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition. 586–595. [86]Wencan Zhang, Mariella Dimiccoli, and Brian Y Lim. 2022. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–32. [87]Wencan Zhang and Brian Y Lim. 2022. Towards relatable explainable AI with the perceptual process. In Proceedings of the 2022 CHI conference on human factors in computing systems. 1–24. [88]Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, and Ting Wang. 2020. Interpretable deep learning under fire. In 29thUSENIXsecurity symposium (USENIX security 20). [89]Yuhang Zhang, Chengrui Wang, Xu Ling, and Weihong Deng. 2022. Learn from all: Erasing attention consistency for noisy label facial expression recognition. In European conference on computer vision. Springer, 418–434. [90]Wenliang Zhao, Yongming Rao, Ziyi Wang, Jiwen Lu, and Jie Zhou. 2021. Towards interpretable deep metric learning with structural matching. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 9887–9896. [91]Sijie Zhu, Taojiannan Yang, and Chen Chen. 2021. Visual explanation for deep metric learning. IEEE Transactions on Image Processing 30 (2021), 7593–7607. AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil A Appendix A.1 Dataset and Participant Details Appendix Table 3. Demographic distribution of face targets selected during dataset curation. EthnicityWhiteAsian GenderFemaleMaleFemaleMale Multi-PIE20201010 CelebA20201010 Attribute Eyebrow shape Eyebrow thickness Eye shape Eye size Eye colour Nose shape Nose size Mouth size Lip thickness Lip colour Face proportion Forehead height Pupillary distance Cheek shape Jaw width Chin shape Hair colour Hair length Skin colour Skin texture Feature Value Norm 0 0.5 1.0 EditedOriginal Appendix Fig. 11. Statistical distribution of feature values for source and edited faces across 20 face attributes. These feature values are computed using geometry-based heuristic methods and serve as proxy evaluation metrics. Appendix Table 4. Demographic statistics of participants in our psychophysical similarity measurement studies. EthnicityGenderAge WhiteAsianBlackHispanicOtherFemaleMaleOtherRangeMedian Face perception35352243838221-7844.0 Attribute perception34218753920324286121-8444.0 M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. A.2 Statistical Analysis Appendix Table 5. Inter-rater agreement (Fleiss’ 휅) [35] for overall face and 20 facial attributes. Attribute Fleiss’ 휅 Agreement Overall0.604Moderate Eyebrow shape0.718Substantial Eyebrow thickness0.703Substantial Eye shape0.560Moderate Eye size0.605Moderate Eye color0.705Substantial Nose shape0.598Moderate Nose size0.604Moderate Mouth size0.577Moderate Lip thickness0.595Moderate Lip color0.538Moderate Hair color0.544Moderate Skin color0.537Moderate Skin texture0.578Moderate Face aspect ratio0.598Moderate Hair length0.736Substantial Forehead height0.613Moderate Pupillary distance0.457Moderate Cheek shape0.576Moderate Jaw width0.594Moderate Chin shape0.546Moderate Appendix Table 6. Dunnett’s test comparing AF-FLIP against baselines (훼= 0.05). MethodComparisonDiff. 푝-value PSNRAF-FLIP-0.184642< .0001 SSIMAF-FLIP-0.171484< .0001 SimCLRAF-FLIP-0.210958< .0001 MoCoAF-FLIP-0.145168< .0001 DINOAF-FLIP-0.145168< .0001 LPIPSAF-FLIP-0.1188520.0010 DreamSimAF-FLIP-0.0662210.2008 CLIPAF-FLIP-0.0793790.0736 FLIPAF-FLIP-0.0793790.0736 FaceNetAF-FLIP-0.171484< .0001 CosFaceAF-FLIP-0.145168< .0001 ArcFaceAF-FLIP-0.1188520.0010 AF-CLIPAF-FLIP-0.0327870.9382 AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Appendix Table 7. Attribute importance and 푅 2 of GAM models fitted to Own-group and Other-group human judgments. Attribute Own-group Other-group Eyebrow shape0.0350.033 Eyebrow thickness0.0270.029 Eye shape0.0430.079 Eye size0.0970.028 Eye color0.0760.056 Nose shape0.0200.063 Nose size0.0140.033 Mouth size0.0160.025 Lip thickness0.0160.026 Lip color0.0330.032 Hair color0.0410.030 Skin color0.0150.070 Skin texture0.0160.081 Face aspect ratio0.1590.098 Hair length0.2030.071 Forehead height0.1090.083 Pupillary distance0.0200.054 Cheek shape0.0290.042 Jaw width0.0160.033 Chin shape0.0150.034 Overall GAM 푅 2 0.5200.480 M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. A.3 Module Justification and Validation We justify the design rationale of each module in Section 4 against standard alternatives in Appendix Table 8. Appendix Table 8. Justifications for AlignFace modules compared to standard alternatives. ModuleAlternativeJustification w.r.t. alternative (Benefit) Vision-Language ModelFeature EngineeringMore scalable: support open-domain (label-free) attributes instead of hard-coded features. Attr-Gated Cross-AttentionFeature ConcatenationInterpretability and dimensionality reduction: constrain reason- ing via attention to focus on information most relevant to each attribute, rather than finding relationships across the full face pair without differencing or focusing. Concept Bottleneck ModelMLPInterpretability: regularizes reasoning with domain-specific at- tributes, avoiding spurious features. Neural GAMMLPInterpretability: regularizes reasoning independently and nonlin- early for each attribute, avoiding spurious relations. We validated that general-purpose VLMs encode featural and configural attribute information by comparing AlignFace’s predicted attribute distances against ground-truth heuristic inter-landmark distances [44]. Appendix Table 9 shows high Pearson correlations휌, indicating that AlignFace reliably tracks the intended physical attribute changes. Appendix Table 9. Pearson correlation휌between heuristic physical measurements and model-predicted attribute distances on single-attribute manipulations. CategoryAttribute Pearson correlation 휌 Featural Eyebrow thickness0.916 Eye size0.970 Nose size0.969 Mouth size0.840 Lip thickness0.989 Configural Face aspect ratio0.934 Forehead height0.963 Pupillary distance0.892 Jaw width0.947 A.4 Baseline Implementation Details We describe additional details of the baseline competitors used to evaluate AlignFace in Section 5.1.3. All models are implemented in PyTorch, except for PSNR/SSIM, which use scikit-learn. Images are scaled to 224×224 by default, except for face recognition models (112×112). We loaded the official pretrained checkpoint for each backbone during evaluation. Faces are preprocessed to fit each model’s input size and normalized according to its backbone. Feature embeddings are extracted and used to compute cosine similarity, which then determines the 2AFC triplet response. Appendix Table 10 shows the backbone and pretrained dataset for each baseline. Appendix Table 10. Baseline implementation details. BaselineBackbonePretrained PSNR/SSIM— SimCLRResNet-50ImageNet-1K MoCoResNet-50ImageNet-1K DINOv2ViT-B/32LVD-142M CLIPViT-B/32WIT-400M FLIPViT-B/32FLIP-80M LPIPSAlexNetImageNet-1K DreamSimEnsembleImageNet + synthetic FaceNetInception-ResNet-v1MS1MV3 CosFaceInception-ResNet-v1MS1MV3 ArcFaceInception-ResNet-v1MS1MV3 AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept RelationsMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil A.5 Supplementary Results Attribute Own-groupOther-group Appendix Fig. 12. Partial dependence plots of overall relative face selection by relative attribute selection across 20 attributes. Attribute 0.92 0.83 0.96 0.91 0.52 0.99 0.94 0.65 0.94 0.93 0.88 0.94 0.95 0.97 0.99 0.94 0.71 0.96 0.91 -0.51 0.97 0.94 -0.28 0.98 0.98 -0.79 0.99 0.94 0.93 0.95 -0.97 0.75 0.97 0.96 0.97 1.00 0.97 0.96 0.98 0.95 0.94 1.00 0.97 0.88 1.00 -0.97 0.84 0.99 0.91 0.92 0.98 0.97 0.99 0.99 0.97 0.91 0.98 0.91 0.99 0.92 nGAM Appendix Fig. 13. Partial dependence plots (PDPs) of overall relative face distance difference by relative attribute distance difference across 20 attributes. Numbers denote the Pearson correlation between human judgments and comparators. M ’26, November 10–14, 2026, Rio de Janeiro, BrazilHuang et al. Attribute Agreement b) a) Attribute Agreement Backbone AF - CLIP AF - FLIP Appendix Fig. 14. Cross-dataset perception agreement for a) attribute categories and b) individual attributes. b) a) Attribute AgreementAttribute Agreement CLIP FLIP Backbone Type Category Featural Con fi gural 0% 50% 100% 0% 50% CLIP FLIP Backbone Type Attribute Group / Attribute Eyebrow shape Eyebrow thickness Eye shape Eye size Eye color Nose shape Nose size Mouth size Lip thickness Lip color Hair color Skin color Skin texture Face aspect ratio Hair length Forehead height Pupillary distance Cheek shape Jaw width Chin shape FeaturalConfigural 0% 50% 100% 0% 50% Backbone AF - CLIP AF - FLIP Appendix Fig. 15. Cross-group attribute perception agreement for a) attribute categories and b) individual attributes.